The persistent COSMOS deployment at Amsterdam Ave and 120th St—equipped with street-level/bird's-eye cameras, wireless connectivity, and programmable edge/cloud resources—provides a unique testbed for city-scale intelligent systems. Vehicles, cyclists, and pedestrians interact under occlusion, changing lighting and weather, heterogeneous sensing conditions, and safety-critical constraints. The COSMOS smart-intersection research program combines programmable multi-view cameras, wireless connectivity, edge/cloud computing, and long-term real-world deployments to study the complete pipeline from urban sensing to intelligent decision making.
Our research has evolved from real-time object detection and tracking toward data-driven traffic modeling, digital twins and controllable simulation, semantic traffic understanding, and adaptive urban AI. Long-term observations from the COSMOS intersection support the development of robust perception models, realistic traffic simulations, and synthetic safety-critical scenarios. More recent work investigates vision-language and language-based reasoning for understanding complex traffic interactions, as well as adaptive methods for allocating sensing, communication, computation, and annotation resources across heterogeneous data sources.
Together, these capabilities make the COSMOS intersection a testbed not only for sensing and communications, but also for studying how AI systems can observe, model, generate, understand, and adapt to complex urban environments.
System architecture and capabilities overview of the COSMOS smart city intersection testbed, highlighting sensing, digital twins, edge intelligence, and semantic understanding.
Recent work investigates vision-language and language-based reasoning for understanding complex traffic interactions, leveraging zero-shot image tagging (Ho et al., 2026) to transform semantic descriptions into structured traffic evidence.
Early research at the COSMOS smart intersection established the sensing, computing, and perception infrastructure for intelligent urban applications (Raychaudhuri et al., 2020; Yang et al., 2020). Using street-level and bird's-eye cameras connected to COSMOS edge and cloud resources, we developed real-time pipelines for detecting and tracking vehicles, pedestrians, and cyclists at Amsterdam Avenue and West 120th Street (Turkcan et al., 2024; Ghasemi et al., 2023; Ghasemi et al., 2022).
Constellation Dataset. The dataset provides 13,314 annotations for high-altitude object detection, demonstrating robustness across variable lighting, night scenes, and adverse weather conditions.
Recent advances expand this perception foundation with calibration-free, view-agnostic 3D object detection across ground-level, infrastructure, and aerial viewpoints (Turkcan et al., 2026), alongside real-time video analytics optimized for deployment across edge nodes and end-user devices (Ghasemi et al., 2025). We distinguish between methods evaluated on recorded intersection data and demonstrated live deployments running on testbed hardware.
UrbanOmniDetect Pipeline. By formulating 3D detection as ordered keypoint regression, this framework enables view-agnostic, calibration-free detection across viewpoints.
Together, this work established the core sensing → communication → edge/cloud computing → perception pipeline of the COSMOS smart intersection. Persistent perception produces trajectories, which support spatial/temporal distributions and learned traffic models, enabling recent research to model and generate traffic behavior, understand complex interactions, and adapt data collection and computing to support robust urban AI.
Long-term observations from the COSMOS intersection enable us to build data-driven models of urban traffic and reproduce complex interactions in controlled environments. Rather than relying only on manually specified traffic rules, our work learns spatial, temporal, and trajectory distributions directly from real-world observations. Persistent perception produces trajectories, which support spatial/temporal distributions and learned traffic models, providing a foundation for realistic traffic simulation and digital-twin experimentation.
Our digital twin research is structured around three core pillars:
1. Data-Driven Trajectory Simulation: Recorded trajectories are modeled as spatial and temporal distributions, sampled to generate new traffic scenes, and refined using learned trajectory-prediction models (Zang et al., 2024).
Data-driven traffic simulation from real-world observations. Recorded trajectories are modeled as spatial and temporal distributions, sampled to generate new traffic scenes, and refined using learned trajectory-prediction models.
2. Data-Driven Trajectory Simulation: Recorded trajectories are modeled as spatial and temporal distributions, sampled to generate new traffic scenes, and refined using learned trajectory-prediction models (Zang et al., 2024).
Photorealistic synthetic streetscapes generation. Synthetic images of the intersections are generated under diverse environmental conditions, including fog, snow, rain, and nighttime, enabling scalable training and evaluation across conditions that are difficult to collect systematically in the real world.
3. Language-Guided Interaction Generation: High-level natural language descriptions of rare and safety-critical events are translated into behavioral segments and physically plausible trajectories (Zang et al., 2026).
Language-guided rare-event synthesis. A high-level instruction such as “a vehicle yields to a pedestrian” is translated into behavioral segments and refined into physically plausible agent trajectories within a realistic traffic scene.
Together, these efforts connect the physical COSMOS deployment with increasingly flexible digital representations, enabling researchers to replay observed traffic, synthesize new conditions, and systematically test intelligent transportation systems under controlled and safety-critical scenarios.
A large-scale urban sensing system produces far more data than can be continuously transmitted, processed, and annotated. Different cameras observe different viewpoints, traffic patterns, occlusions, and environmental conditions, making uniform resource allocation inefficient. To address this, our research is organized around a central decision question: Where should we spend limited sensing, bandwidth, compute, and annotation resources?
We group our edge intelligence initiatives by key system decision variables:
1. What to collect & label: Under a limited annotation budget, adaptive data collection algorithms continuously evaluate model performance and direct labeling resources toward camera views and conditions where new data is most valuable (Zang et al., 2025).
Adaptive data collection across heterogeneous camera views. Under a limited annotation budget, the system dynamically selects which camera to collect and label data from, improving robust model performance across viewpoints and operating conditions.
2. What to transmit & where to process: Adaptive video streaming and edge filtering reduce bandwidth and processing costs while preserving information for downstream tasks (Ghasemi et al., 2025).
Edge-cloud intelligence for the COSMOS smart intersection. Real-time camera streams are processed across distributed edge and cloud resources to support scalable urban sensing and AI applications.
3. How to manage resources: Distributed vision-language inference partitions model processing between edge devices and cloud servers to optimize execution costs and latency (Li et al., 2025).
Distributed vision-language inference across edge and cloud resources. Visual features are extracted locally at the edge and transmitted to a central server for language-model processing, reducing centralized computation while supporting scalable VLM deployment.
More broadly, this research treats data acquisition, communication, computation, and model training as interconnected components of the same system. The goal is to move from a passive sensing infrastructure toward an adaptive urban AI platform that determines what information should be collected, where it should be processed, and how limited resources should be used to improve system-wide performance.