EDBT 2026 Demo / reviewers in the wild / expert
Jiaqi Ma 0003
dblp:155/2199-3
· DBLP profile ↗
36ranked-venue papers
0as first author
34since 2021 · last 2026
0000-0002-8184-5157ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 24 since 2021Systems, architecture and hardware · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented ReasoningabstractUnderstanding and adhering to traffic regulations is essential for autonomous vehicles to ensure safety and trustworthiness. However, traffic regulations are complex, context-dependent, and differ between regions, posing a major challenge to conventional rule-based decision-making approaches. We present an interpretable, regulation-aware decision-making framework, DriveReg, which enables autonomous vehicles to understand and adhere to region-specific traffic laws and safety guidelines. The framework integrates a Retrieval Augmented Generation (RAG)-based Traffic Regulation Retrieval Agent, which retrieves relevant rules from regulatory documents based on the current situation, and a Large Language Model (LLM)-powered Reasoning Agent that evaluates actions for legal compliance and safety. Our design emphasizes interpretability to enhance transparency and trustworthiness. To support systematic evaluation, we introduce DriveReg Scenarios Dataset, a comprehensive dataset of driving scenarios across Boston, Singapore, and Los Angeles, with both hypothesized text-based cases and real-world driving data, specifically constructed and annotated to evaluate models’ capacity for regulation understanding and reasoning. We validate our framework on the DriveReg Scenarios Dataset and real-world deployment, demonstrating strong performance and robustness across diverse environments. Tianhui Cai, Zewei Zhou, Haoxuan Ma, Seth Z. Zhao, Zhiwen Wu, Xu Han 0014, Zhiyu Huang, Jiaqi Ma 0003 |
AAAI | 9 |
| 2026 | Exploring Factors Influencing Human-Autonomy Teams in Automated Driving System Operations
Camila Correa-Jullian, Marilia Abílio Ramos, Ali Mosleh 0001, Jiaqi Ma 0003 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2026 | Deep Activity Model: A Generative Deep Learning Approach for Human Mobility Pattern SynthesisabstractHuman mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations. Existing deep learning models tend to overlook the semantic interdependencies among activities and households and rely on restricted GPS data, while activity-based models depend on rigid assumptions and extensive data, making them costly and difficult to adapt to new regions, especially those with limited conventional travel data. We propose a generative Transformer model that uses socio-demographic and household attributes to synthesize daily activity chains, with a location module assigning spatial zones to produce complete daily trajectories. To address these limitations, we propose a generative Transformer model that uses socio-demographic and household attributes to synthesize daily activity chains, with a location module assigning spatial zones to produce complete daily trajectories. Trained on open-source and widely available household travel survey data and then fine-tuned with local data, the model captures national activity patterns and transfers effectively to California, Washington, and Mexico City. This approach offers significant potential for advancing synthetic human mobility modeling and provides urban planners and policymakers with improved tools for simulating transportation systems and supporting decisions in urban development and public health. Its practical utility is demonstrated through large-scale traffic simulations in Los Angeles County. Compared to the SCAG Activity-Based Model (ABM), the proposed method produces consistent spatial demand patterns, achieving an activity location cosine similarity of 0.997 and a network-level vehicle-miles-traveled Mean Absolute Percentage Error (MAPE) of 4.97%. Compared to real-world observations from Caltrans PeMS, the simulated traffic achieves MAPE of 5.85% for traffic volume and 4.36% for speed on California’s I-405 corridor. This demonstrates the practicality of learning-based synthetic mobility for regional simulation and planning. Xishun Liao, Qinhua Jiang, Brian Yueshuai He, Chenchen Kuai, Jiaqi Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionabstractVehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores temporal cues and temporal tasks (e.g., temporal perception and prediction). In this paper, we focus on the spatio-temporal fusion in V2X scenarios and design one-step and multi-step communication strategies (when to transmit) as well as examine their integration with three fusion strategies - early, late, and intermediate (what to transmit), providing comprehensive benchmarks with 11 fusion models (how to fuse). Furthermore, we propose V2XPnP, a novel intermediate fusion framework within one-step communication for end-to-end perception and prediction. Our framework employs a unified Transformer-based architecture to effectively model complex spatio-temporal relationships across multiple agents, frames, and high-definition maps. Moreover, we introduce the V2XPnP Sequential Dataset that supports all V2X collaboration modes and addresses the limitations of existing real-world datasets, which are restricted to single-frame or single-mode cooperation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods in both perception and prediction tasks. Zewei Zhou, Hao Xiang 0001, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Tianhui Cai, Johnson Liu, Maheswari Bajji, Xin Xia 0007, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
ICCV | 14 |
| 2025 | TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionabstractEnd-to-end training of multi-agent systems offers significant advantages in improving multi-task performance. However, training such models remains challenging and requires extensive manual design and monitoring. In this work, we introduce TurboTrain, a novel and efficient training framework for multi-agent perception and prediction. TurboTrain comprises two key components: a multi-agent spatiotemporal pretraining scheme based on masked reconstruction learning and a balanced multi-task learning strategy based on gradient conflict suppression. By streamlining the training process, our framework eliminates the need for manually designing and tuning complex multi-stage training pipelines, substantially reducing training time and improving performance. We evaluate TurboTrain on a real-world cooperative driving dataset, V2XPnP-Seq, and demonstrate that it further improves the performance of state-of-the-art multi-agent perception and prediction models. Our results highlight that pretraining effectively captures spatiotemporal multi-agent features and significantly benefits downstream tasks. Moreover, the proposed balanced multi-task learning strategy enhances detection and prediction. Zewei Zhou, Seth Z. Zhao, Tianhui Cai, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
ICCV | 6 |
| 2025 | Traffic Regulation-aware Path Planning with Regulation Databases and Vision-Language ModelsabstractThis paper introduces and tests a framework integrating traffic regulation compliance into automated driving systems (ADS). The framework enables ADS to follow traffic laws and make informed decisions based on the driving environment. Using RGB camera inputs and a vision-language model (VLM), the system generates descriptive text to support a regulation-aware decision-making process, ensuring legal and safe driving practices. This information is combined with a machine-readable ADS regulation database to guide future driving plans within legal constraints. Key features include: 1) a regulation database supporting ADS decision-making, 2) an automated process using sensor input for regulation-aware path planning, and 3) validation in both simulated and real-world environments. Particularly, the real-world vehicle tests not only assess the framework's performance but also evaluate the potential and challenges of VLMs to solve complex driving problems by integrating detection, reasoning, and planning. This work enhances the legality, safety, and public trust in ADS, representing a significant step forward in the field. Xu Han 0014, Zhiwen Wu, Xin Xia 0007, Jiaqi Ma 0003 |
ICRA | 4 |
| 2025 | CooperRisk: A Driving Risk Quantification Pipeline with Multi-Agent Cooperative Perception and PredictionabstractRisk quantification is a critical component of safe autonomous driving, however, constrained by the limited perception range and occlusion of single-vehicle systems in complex and dense scenarios. Vehicle-to-everything (V2X) paradigm has been a promising solution to sharing complementary perception information, nevertheless, how to ensure the risk interpretability while understanding multi-agent interaction with V2X remains an open question. In this paper, we introduce the first V2X-enabled risk quantification pipeline, CooperRisk, to fuse perception information from multiple agents and quantify the scenario driving risk in future multiple timestamps. The risk is represented as a scenario risk map to ensure interpretability based on risk severity and exposure, and the multi-agent interaction is captured by the learning-based cooperative prediction model. We carefully design a risk-oriented transformer-based prediction model with multi-modality and multi-agent considerations. It aims to ensure scene-consistent future behaviors of multiple agents and avoid conflicting predictions that could lead to overly conservative risk quantification and cause the ego vehicle to become overly hesitant to drive. Then, the temporal risk maps could serve to guide a model predictive control planner. We evaluate the CooperRisk pipeline in a real-world V2X dataset V2XPnP, and the experiments demonstrate its superior performance in risk quantification, showing a 44.35% decrease in conflict rate between the ego vehicle and background traffic participants. Mingyue Lei, Zewei Zhou, Jia Hu 0003, Jiaqi Ma 0003 |
IROS | 5 |
| 2025 | CooPre: Cooperative Pretraining for V2X Cooperative PerceptionabstractExisting Vehicle-to-Everything (V2X) cooperative perception methods rely on accurate multi-agent 3D annotations. Nevertheless, it is time-consuming and expensive to collect and annotate real-world data, especially for V2X systems. In this paper, we present a self-supervised learning framwork for V2X cooperative perception, which utilizes the vast amount of unlabeled 3D V2X data to enhance the perception performance. Specifically, multi-agent sensing information is aggregated to form a holistic view and a novel proxy task is formulated to reconstruct the LiDAR point clouds across multiple connected agents to better reason multi-agent spatial correlations. Besides, we develop a V2X bird-eye-view (BEV) guided masking strategy which effectively allows the model to pay attention to 3D features across heterogeneous V2X agents (i.e., vehicles and infrastructure) in the BEV space. Noticeably, such a masking strategy effectively pretrains the 3D encoder with a multi-agent LiDAR point cloud reconstruction objective and is compatible with mainstream cooperative perception backbones. Our approach, validated through extensive experiments on representative datasets (i.e., V2X-Real, V2V4Real, and OPV2V) and multiple state-of-the-art cooperative perception methods (i.e., AttFuse, F-Cooper, and V2X-ViT), leads to a performance boost across all V2X settings. Notably, CooPre achieves a 4% mAP improvement on V2X-Real dataset and surpasses baseline performance using only 50% of the training data, highlighting its data efficiency. Additionally, we demonstrate the framework’s powerful performance in cross-domain transferability and robustness under challenging scenarios. The code will be made publicly available at https://github.com/ucla-mobility/CooPre. Seth Z. Zhao, Hao Xiang 0001, Chenfeng Xu, Xin Xia 0007, Bolei Zhou, Jiaqi Ma 0003 |
IROS | 6 |
| 2025 | V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative PerceptionabstractModern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In recent years, a series of cooperative perception datasets have emerged; however, these datasets primarily focus on cameras and LiDAR, neglecting 4D Radar—a sensor used in single-vehicle autonomous driving to provide robust perception in adverse weather conditions. In this paper, to bridge the gap created by the absence of 4D Radar datasets in cooperative perception, we present V2X-Radar, the first large-scale, real-world multi-modal dataset featuring 4D Radar. V2X-Radar dataset is collected using a connected vehicle platform and an intelligent roadside unit equipped with 4D Radar, LiDAR, and multi-view cameras. The collected data encompasses sunny and rainy weather conditions, spanning daytime, dusk, and nighttime, as well as various typical challenging scenarios. The dataset consists of 20K LiDAR frames, 40K camera images, and 20K 4D Radar data, including 350K annotated boxes across five categories. To support various research domains, we have established V2X-Radar-C for cooperative perception, V2X-Radar-I for roadside perception, and V2X-Radar-V for single-vehicle perception. Furthermore, we provide comprehensive benchmarks across these three sub-datasets. Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Jiaqi Ma 0003, Zhiying Song, Ziying Song, Li Wang 0092, Yang Shen 0005, Chen Lv 0001 |
NeurIPS | 5 |
| 2025 | AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningabstractRecent advancements in Vision-Language-Action (VLA) models have shown promise for end-to-end autonomous driving by leveraging world knowledge and reasoning capabilities. However, current VLA models often struggle with physically infeasible action outputs, complex model structures, or unnecessarily long reasoning. In this paper, we propose AutoVLA, a novel VLA model that unifies reasoning and action generation within a single autoregressive generation model for end-to-end autonomous driving. AutoVLA performs semantic reasoning and trajectory planning directly from raw visual inputs and language instructions. We tokenize continuous trajectories into discrete, feasible actions, enabling direct integration into the language model. For training, we employ supervised fine-tuning to equip the model with dual thinking modes: fast thinking (trajectory-only) and slow thinking (enhanced with chain-of-thought reasoning). To further enhance planning performance and efficiency, we introduce a reinforcement fine-tuning method based on Group Relative Policy Optimization (GRPO), reducing unnecessary reasoning in straightforward scenarios. Extensive experiments across real-world and simulated datasets and benchmarks, including nuPlan, nuScenes, Waymo, and CARLA, demonstrate the competitive performance of AutoVLA in both open-loop and closed-loop settings. Qualitative results showcase the adaptive reasoning and accurate planning capabilities of AutoVLA in diverse scenarios. Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
NeurIPS | 7 |
| 2024 | V2X-Real: A Largs-Scale Dataset for Vehicle-to-Everything Cooperative Perception
Hao Xiang 0001, Zhaoliang Zheng, Xin Xia 0007, Runsheng Xu, Letian Gao, Zewei Zhou, Xu Han 0014, Xinkai Ji, Zonglin Meng, Mingyue Lei, Haoxuan Ma, Yunshuang Yuan, Yingqian Zhao, Jiaqi Ma 0003 |
ECCV (52) | 18 |
| 2024 | Breaking Data Silos: Cross-Domain Learning for Multi-Agent Perception from Independent Private SourcesabstractThe diverse agents in multi-agent perception systems may be from different companies. Each company might use the identical classic neural network architecture based encoder for feature extraction. However, the data source to train the various agents is independent and private in each company, leading to the Distribution Gap of different private data for training distinct agents in multi-agent perception system. The data silos by the above Distribution Gap could result in a significant performance decline in multi-agent perception. In this paper, we thoroughly examine the impact of the distribution gap on existing multi-agent perception systems. To break the data silos, we introduce the Feature Distribution-aware Aggregation (FDA) framework for cross-domain learning to mitigate the above Distribution Gap in multi-agent perception. FDA comprises two key components: Learnable Feature Compensation Module and Distribution-aware Statistical Consistency Module, both aimed at enhancing intermediate features to minimize the distribution gap among multi-agent features. Intensive experiments on the public OPV2V and V2XSet datasets underscore FDA’s effectiveness in point cloud-based 3D object detection, presenting it as an invaluable augmentation to existing multi-agent perception systems. The code is available at https://github.com/jinlong17/BDS-V2V. Xinyu Liu 0009, Runsheng Xu, Jiaqi Ma 0003, Hongkai Yu |
ICRA | 5 |
| 2024 | S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to RealityabstractDue to the lack of enough real multi-agent data and time-consuming of labeling, existing multi-agent cooperative perception algorithms usually select the simulated sensor data for training and validating. However, the perception performance is degraded when these simulation-trained models are deployed to the real world, due to the significant domain gap between the simulated and real data. In this paper, we propose the first Simulation-to-Reality transfer learning framework for multi-agent cooperative perception using a novel Vision Transformer, named as S2R-ViT, which considers both the Deployment Gap and Feature Gap between simulated and real data. We investigate the effects of these two types of domain gaps and propose a novel uncertainty-aware vision transformer to effectively relief the Deployment Gap and an agent-based feature adaptation module with inter-agent and ego-agent discriminators to reduce the Feature Gap. Our intensive experiments on the public multi-agent cooperative perception datasets OPV2V and V2V4Real demonstrate that the proposed S2R-ViT can effectively bridge the gap from simulation to reality and outperform other methods significantly for point cloud-based 3D object detection. Runsheng Xu, Xinyu Liu 0009, Qin Zou 0001, Jiaqi Ma 0003, Hongkai Yu |
ICRA | 6 |
| 2024 | End-to-end Cooperative Localization via Neural Feature SharingabstractCooperative driving automation attracts great attention for its potential to improve traffic safety. Knowing each vehicle’s accurate position serves as the cornerstone for the information fusion necessary in cooperative driving tasks. However, inherent errors within a vehicle’s self-localization system often necessitate correction to facilitate cooperative perception and downstream tasks. Leveraging intermediate features shared among other Connected Automated Vehicles (CAVs), we propose an end-to-end learning localization framework aimed at estimating the relative pose error between the ego vehicle and the CAV. We investigate factors that may influence learning performance and validate the algorithm’s performance using a simulation dataset. The proposed method is compared with the traditional point cloud matching-based relative localization method. Remarkably, our framework effectively corrects relative pose errors even when the vehicle exhibits significant initial localization inaccuracies, and it can be integrated into the cooperative perception system. Letian Gao, Hao Xiang 0001, Xin Xia 0007, Jiaqi Ma 0003 |
IV | 4 |
| 2024 | V2X-DSI: A Density-Sensitive Infrastructure LiDAR Benchmark for Economic Vehicle-to-Everything Cooperative PerceptionabstractRecent research has demonstrated that the Vehicle-to-Everything (V2X) communication techniques can fundamentally improve the perception system for autonomous driving by collaborating between vehicle and infrastructure sensors. LiDAR is the commonly-used sensor for V2X autonomous driving due to its robustness in challenging scenarios. However, the LiDAR sensor is expensive, so the cost of equipping LiDAR sensors to a large number of infrastructures on the large-scale roadway network is extremely high, which has limited the wide deployment of the V2X cooperative perception system. How to discover an economic V2X cooperative perception system is never been well studied before. Inspired by the cost difference of the various point cloud densities of LiDAR, we propose the first Density-Sensitive Infrastructure LiDAR benchmark for economic V2X cooperative perception, named V2X-DSI, in this paper. Using the proposed V2X-DSI benchmark, we analyze the effect of cooperative perception performance under different beam infrastructure LiDAR. We specifically assess three state-of-the-art methods, i.e., OPV2V, V2X-ViT, and CoBEVT, using our V2X-DSI dataset. The results indicate that varying beam infrastructure LiDAR sensors play a crucial role in influencing cooperative perception performance. Xinyu Liu 0009, Runsheng Xu, Jiaqi Ma 0003, Xiaopeng Li 0020, Hongkai Yu |
IV | 4 |
| 2024 | MSTF: Multiscale Transformer for Incomplete Trajectory PredictionabstractMotion forecasting plays a pivotal role in autonomous driving systems, enabling vehicles to execute collision warnings and rational local-path planning based on predictions of the surrounding vehicles. However, prevalent methods often assume complete observed trajectories, neglecting the potential impact of missing values induced by object occlusion, scope limitation, and sensor failures. Such oversights inevitably compromise the accuracy of trajectory predictions. To tackle this challenge, we propose an end-to-end framework, termed Multi-scale Transformer (MSTF), meticulously crafted for incomplete trajectory prediction. MSTF integrates a Multiscale Attention Head (MAH) and an Information Increment-based Pattern Adaptive (IIPA) module. Specifically, the MAH component concurrently captures multiscale motion representation of trajectory sequence from various temporal granularities, utilizing a multi-head attention mechanism. This approach facilitates the modeling of global dependencies in motion across different scales, thereby mitigating the adverse effects of missing values. Additionally, the IIPA module adaptively extracts continuity representation of motion across time steps by analyzing missing patterns in the data. The continuity representation delineates motion trend at a higher level, guiding MSTF to generate predictions consistent with motion continuity. We evaluate our proposed MSTF model using two large-scale real-world datasets. Experimental results demonstrate that MSTF surpasses state-of-the-art (SOTA) models in the task of incomplete trajectory prediction, showcasing its efficacy in addressing the challenges posed by missing values in motion forecasting for autonomous driving systems. Zhanwen Liu, Yang Wang 0015, Jiaqi Ma 0003, Xiangmo Zhao |
IV | 5 |
| 2024 | CooperFuse: A Real-Time Cooperative Perception Fusion FrameworkabstractCooperative perception algorithms based on fusing sensing data across multiple connected automated vehicles (CAVs) have shown promising performance to enhance the existing individual perception algorithm in terms of object detection tracking. However, existing cooperative perception algorithms are only developed offline given the constraints from data sharing and computational resources and none of them have been verified in real-time conditions. In this work, we propose a real-time cooperative perception framework called CooperFuse, which achieves cooperative perception in a late fusion scheme. Based on object detection and tracking results from individual vehicle, the late fusion cooperative perception algorithm considers object detection confidence score, kinematics, and dynamics consistency as well as scale consistency of detected objects. The algorithm computes the kinematic and dynamic consistency of the objects by solving for the energy consumption of inter-frame trajectories, and determines scale consistency by calculating inter-frame scale changes, enabling feature-based bounding box fusion. The experimental results demonstrate the real-time performance of the proposed algorithm and reveal its effective improvements in feature fusion and object detection accuracy when dealing with heterogeneous detection models across different cooperative intelligent agents. Zhaoliang Zheng, Xin Xia 0007, Letian Gao, Hao Xiang 0001, Jiaqi Ma 0003 |
IV | 5 |
| 2024 | Impact of Sensing Errors on Headway Design: From $\alpha$α-Fair Group Safety to Traffic ThroughputabstractHeadway, namely the distance between vehicles, is a key design factor for ensuring the safe operation of autonomous driving systems. There have been studies on headway optimization based on the speeds of leading and trailing vehicles, assuming perfect sensing capabilities. In practical scenarios, however, sensing errors are inevitable, calling for a more robust headway design to mitigate the risk of collision. Undoubtedly, augmenting the safety distance would reduce traffic throughput, highlighting the need for headway design to incorporate both sensing errors and risk tolerance models. In addition, prioritizing group safety over individual safety is often deemed unacceptable because no driver should sacrifice their safety for the safety of others. In this study, we propose a multi-objective optimization framework that examines the impact of sensing errors on both traffic throughput and the fairness of safety among vehicles. The proposed framework provides a solution to determine the Pareto frontier for traffic throughput and vehicle safety. ComDrive, a communication-based autonomous driving simulation platform, is developed to validate the proposed approach. Extensive experiments demonstrate that the proposed approach outperforms existing baselines. Wei Shao 0006, Zejun Fan, Chia-Ju Chen, Jiaqi Ma 0003, Junshan Zhang |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Toward Foundation Models for Inclusive Object Detection: Geometry- and Category-Aware Feature Extraction Across Road User CategoriesabstractThe safety of different categories of road users comprising motorized vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists is one of the priorities of automated driving and smart infrastructure services. Three-dimensional (3-D) LiDAR-based object detection has been a promising approach to perceiving road users. Despite accurate 3-D geometry information, the point cloud from LiDAR is usually nonuniform, and learning the effective point cloud abstract representations for diverse road users remains challenging for 3-D object detection, particularly for small objects such as VRUs. For inclusive object detection (IDetect), we propose a general foundation convolution component, called geometry-aware convolution (GA Conv) toward a foundation feature extraction model, to serve as basic convolution operations of the neutral network for inclusive 3-D object detection. Further, the GA Conv operations are then utilized as the elementary feature extraction layers to build a novel elegant and pyramid network for IDetect. It learns the effective geometric-related features from the unstructured point cloud data by implicitly learning the distribution property and geometry-related features from different categories of road users in particular for VRUs. The proposed IDetect is comprehensively evaluated on the large-scale benchmark Waymo open datasets with all categories of road users. The qualitative and quantitative experiment results demonstrate that IDetect can effectively consider the nonuniform distributed point clouds and learn the geometric features to assist the different categories of road user detection. In addition, the GA Conv has been integrated with other state-of-the-art neural networks and a performance boost for VRU detection has been demonstrated, showing the foundation functionality of the GA Conv and making it a general component in the future inclusive 3-D object detection foundation model. Zonglin Meng, Xin Xia 0007, Jiaqi Ma 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative PerceptionabstractModern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at research.seas.ucla.edu/mobility-lab/v2v4real/. Runsheng Xu, Xin Xia 0007, Hanzhao Li, Zhengzhong Tu, Zonglin Meng, Hao Xiang 0001, Rui Song 0007, Hongkai Yu, Bolei Zhou, Jiaqi Ma 0003 |
CVPR | 13 |
| 2023 | Optimizing the Placement of Roadside LiDARs for Autonomous DrivingabstractMulti-agent cooperative perception is an increasingly popular topic in the field of autonomous driving, where roadside LiDARs play an essential role. However, how to optimize the placement of roadside LiDARs is a crucial but often overlooked problem. This paper proposes an approach to optimize the placement of roadside LiDARs by selecting optimized positions within the scene for better perception performance. To efficiently obtain the best combination of locations, a greedy algorithm based on perceptual gain is proposed, which selects the location that can maximize the perceptual gain sequentially. We define perceptual gain as the increased perceptual capability when a new LiDAR is placed. To obtain the perception capability, we propose a perception predictor that learns to evaluate LiDAR placement using only a single point cloud frame. A dataset named Roadside-Opt is created using the CARLA simulator to facilitate research on the roadside LiDAR placement problem. Extensive experiments are conducted to demonstrate the effectiveness of our proposed method. Hao Xiang 0001, Xinyu Cai, Runsheng Xu, Jiaqi Ma 0003, Gim Hee Lee, Si Liu 0001 |
ICCV | 5 |
| 2023 | HM-ViT: Hetero-modal Vehicle-to-Vehicle Cooperative Perception with Vision TransformerabstractVehicle-to-Vehicle technologies have enabled autonomous vehicles to share information to see through occlusions, greatly enhancing perception performance. Nevertheless, existing works all focused on homogeneous traffic where vehicles are equipped with the same type of sensors, which significantly hampers the scale of collaboration and benefit of cross-modality interactions. In this paper, we investigate the multi-agent hetero-modal cooperative perception problem where agents may have distinct sensor modalities. We present HM-ViT, the first unified multi-agent hetero-modal cooperative perception framework that can collaboratively predict 3D objects for highly dynamic Vehicle-to-Vehicle (V2V) collaborations with varying numbers and types of agents. To effectively fuse features from multi-view images and LiDAR point clouds, we design a novel heterogeneous 3D graph transformer to jointly reason inter-agent and intra-agent interactions. The extensive experiments on the V2V perception dataset OPV2V demonstrate that the HM-ViT outperforms SOTA cooperative perception methods for V2V hetero-modal cooperative perception. Our code will be released at https://github.com/XHwind/HM-ViT. Hao Xiang 0001, Runsheng Xu, Jiaqi Ma 0003 |
ICCV | 3 |
| 2023 | Analyzing Infrastructure LiDAR Placement with Realistic LiDAR Simulation LibraryabstractRecently, Vehicle-to-Everything (V2X) cooperative perception has attracted increasing attention. Infrastructure sensors play a critical role in this research field; however, how to find the optimal placement of infrastructure sensors is rarely studied. In this paper, we investigate the problem of infrastructure sensor placement and propose a pipeline that can efficiently and effectively find optimal installation positions for infrastructure sensors in a realistic simulated environment. To better simulate and evaluate LiDAR place-ment, we establish a Realistic LiDAR Simulation library that can simulate the unique characteristics of different popular LiDARs and produce high-fidelity LiDAR point clouds in the CARLA simulator. Through simulating point cloud data in different LiDAR placements, we can evaluate the perception accuracy of these placements using multiple detection models. Then, we analyze the correlation between the point cloud distribution and perception accuracy by calculating the density and uniformity of regions of interest. Experiments show that when using the same number and type of LiDAR, the placement scheme optimized by our proposed method improves the average precision by 15%, compared with the conventional placement scheme in the standard lane scene. We also analyze the correlation between perception performance in the region of interest and LiDAR point cloud distribution and validate that density and uniformity can be indicators of performance. Both the RLS Library and related code will be released at https://github.com/PJLab-ADG/LiDARSimLib-and-Placement-Evaluation. Xinyu Cai, Runsheng Xu, Wenquan Zhao, Jiaqi Ma 0003, Si Liu 0001 |
ICRA | 5 |
| 2023 | V2XP-ASG: Generating Adversarial Scenes for Vehicle-to-Everything PerceptionabstractRecent advancements in Vehicle-to-Everything communication technology have enabled autonomous vehicles to share sensory information to obtain better perception performance. With the rapid growth of autonomous vehicles and intelligent infrastructure, the V2X perception systems will soon be deployed at scale, which raises a safety-critical question: how can we evaluate and improve its performance under challenging traffic scenarios before the real-world deployment? Collecting diverse large-scale real-world test scenes seems to be the most straightforward solution, but it is expensive and time-consuming, and the collections can only cover limited scenarios. To this end, we propose the first open adversarial scene generator V2XP-ASG that can produce realistic, challenging scenes for modern LiDAR-based multi-agent perception systems. V2XP-ASG learns to construct an adversarial collaboration graph and simultaneously perturb multiple agents' poses in an adversarial and plausible manner. The experiments demonstrate that V2XP-ASG can effectively identify challenging scenes for a large range of V2X perception systems. Meanwhile, by training on the limited number of generated challenging scenes, the accuracy of V2X perception systems can be further improved by 12.3% on challenging and 4% on normal scenes. Our code will be released at https://github.com/XHwind/V2XP-ASG. Hao Xiang 0001, Runsheng Xu, Xin Xia 0007, Zhaoliang Zheng, Bolei Zhou, Jiaqi Ma 0003 |
ICRA | 6 |
| 2023 | Model-Agnostic Multi-Agent Perception FrameworkabstractExisting multi-agent perception systems assume that every agent utilizes the same model with identical parameters and architecture. The performance can be degraded with different perception models due to the mismatch in their confidence scores. In this work, we propose a model-agnostic multi-agent perception framework to reduce the negative effect caused by the model discrepancies without sharing the model information. Specifically, we propose a confidence calibrator that can eliminate the prediction confidence score bias. Each agent performs such calibration independently on a standard public database to protect intellectual property. We also propose a corresponding bounding box aggregation algorithm that considers the confidence scores and the spatial agreement of neighboring boxes. Our experiments shed light on the necessity of model calibration across different agents, and the results show that the proposed framework improves the baseline 3D object detection performance of heterogeneous agents. The code can be found at this url. Runsheng Xu, Weizhe Chen 0004, Hao Xiang 0001, Xin Xia 0007, Lantao Liu, Jiaqi Ma 0003 |
ICRA | 6 |
| 2023 | Bridging the Domain Gap for Multi-Agent PerceptionabstractExisting multi-agent perception algorithms usually select to share deep neural features extracted from raw sensing data between agents, achieving a trade-off between accuracy and communication bandwidth limit. However, these methods assume all agents have identical neural networks, which might not be practical in the real world. The transmitted features can have a large domain gap when the models differ, leading to a dramatic performance drop in multi-agent perception. In this paper, we propose the first lightweight framework to bridge such domain gaps for multi-agent perception, which can be a plug-in module for most of the existing systems while maintaining confidentiality. Our framework consists of a learnable feature resizer to align features in multiple dimensions and a sparse cross-domain transformer for domain adaption. Extensive experiments on the public multi-agent perception dataset V2XSet have demonstrated that our method can effectively bridge the gap for features from different domains and outperform other baseline methods significantly by at least 8% for point-cloud-based 3D object detection. Runsheng Xu, Hongkai Yu, Jiaqi Ma 0003 |
ICRA | 5 |
| 2023 | Domain Adaptive Object Detection for Autonomous Driving under Foggy WeatherabstractMost object detection methods for autonomous driving usually assume a consistent feature distribution between training and testing data, which is not always the case when weathers differ significantly. The object detection model trained under clear weather might be not effective enough on the foggy weather because of the domain gap. This paper proposes a novel domain adaptive object detection framework for autonomous driving under foggy weather. Our method leverages both image-level and object-level adaptation to diminish the domain discrepancy in image style and object appearance. To further enhance the model’s capabilities under challenging samples, we also come up with a new adversarial gradient reversal layer to perform adversarial mining for the hard examples together with domain adaptation. Moreover, we propose to generate an auxiliary domain by data augmentation to enforce a new domain-level metric regularization. Experimental results on public benchmarks show the effectiveness and accuracy of the proposed method. The code is available at https://github.com/jinlong17/DA-Detect. Runsheng Xu, Jin Ma 0005, Qin Zou 0001, Jiaqi Ma 0003, Hongkai Yu |
WACV | 5 |
| 2023 | Pik-Fix: Restoring and Colorizing Old PhotosabstractRestoring and inpainting the visual memories that are present, but often impaired, in old photos remains an intriguing but unsolved research topic. Decades-old photos often suffer from severe and commingled degradation such as cracks, defocus, and color-fading, which are difficult to treat individually and harder to repair when they interact. Deep learning presents a plausible avenue, but the lack of large-scale datasets of old photos makes addressing this restoration task very challenging. Here we present a novel reference-based end-to-end learning framework that is able to both repair and colorize old, degraded pictures. Our proposed framework consists of three modules: a restoration sub-network that conducts restoration from degradations, a similarity network that performs color histogram matching and color transfer, and a colorization subnet that learns to predict the chroma elements of images conditioned on chromatic reference signals. The overall system makes uses of color histogram priors from reference images, which greatly reduces the need for large-scale training data. We have also created a first-of-a-kind public dataset of real old photos that are paired with ground truth "pristine" photos that have been manually restored by PhotoShop experts. We conducted extensive experiments on this dataset and synthetic datasets, and found that our method significantly outperforms previous state-of-the-art models using both qualitative comparisons and quantitative measurements. The code is available at https://github.com/DerrickXuNu/Pik-Fix. Runsheng Xu, Zhengzhong Tu, Yuanqi Du, Zibo Meng, Jiaqi Ma 0003, Alan C. Bovik, Hongkai Yu |
WACV | 7 |
| 2023 | Transportation 5.0: The DAO to Safe, Secure, and Sustainable Intelligent Transportation SystemsabstractIn 2014, IEEE Intelligent Transportation Systems Society established a Technical Committee on Transportation 5.0 with the mission of promoting and transforming the deployment of advanced and innovative technologies, especially Artificial Intelligence in transportation. This paper briefly summarizes our main research and findings over the last decade. Transportation Foundation Models, Transportation Scenarios Engineering, and Transportation Operating Systems have been identified as the main directions for the research and development of next-generation intelligent transportation systems. Fei-Yue Wang 0001, Yilun Lin 0002, Petros A. Ioannou, Ljubo Vlacic, Azim Eskandarian, Xiaoxiang Na, David Cebon, Jiaqi Ma 0003, Lingxi Li 0001, Cristina Olaverri-Monreal |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2022 | V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer
Runsheng Xu, Hao Xiang 0001, Zhengzhong Tu, Xin Xia 0007, Ming-Hsuan Yang 0001, Jiaqi Ma 0003 |
ECCV (39) | 6 |
| 2022 | OPV2V: An Open Benchmark Dataset and Fusion Pipeline for Perception with Vehicle-to-Vehicle CommunicationabstractEmploying Vehicle-to-Vehicle communication to enhance perception performance in self-driving technology has attracted considerable attention recently; however, the absence of a suitable open dataset for benchmarking algorithms has made it difficult to develop and assess cooperative perception technologies. To this end, we present the first large-scale open simulated dataset for Vehicle-to-Vehicle perception. It contains over 70 interesting scenes, 11,464 frames, and 232,913 annotated 3D vehicle bounding boxes, collected from 8 towns in CARLA and a digital town of Culver City, Los Angeles. We then construct a comprehensive benchmark with a total of 16 implemented models to evaluate several information fusion strategies (i.e. early, late, and intermediate fusion) with state-of-the-art LiDAR detection algorithms. Moreover, we propose a new Attentive Intermediate Fusion pipeline to aggregate information from multiple connected vehicles. Our experiments show that the proposed pipeline can be easily integrated with existing 3D LiDAR detectors and achieve outstanding performance even with large compression rates. To encourage more researchers to investigate Vehicle-to-Vehicle perception, we will release the dataset, benchmark methods, and all related codes in https://mobility-lab.seas.ucla.edu/opv2v/. Runsheng Xu, Hao Xiang 0001, Xin Xia 0007, Xu Han 0014, Jiaqi Ma 0003 |
ICRA | 6 |
| 2022 | Too Afraid to Drive: Systematic Discovery of Semantic DoS Vulnerability in Autonomous Driving Planning under Physical-World Attacks
Ziwen Wan, Junjie Shen 0001, Jalen Chuang, Xin Xia 0007, Joshua Garcia, Jiaqi Ma 0003, Qi Alfred Chen |
NDSS | 6 |
| 2022 | Evaluating the Effectiveness of Integrated Connected Automated Vehicle Applications Applied to Freeway Managed LanesabstractThe purpose of this study is to define an operational concept involving connected automated vehicle (CAV) operation on freeway managed lanes. Despite the low projected market penetration of CAVs during the next decade, the use of managed lane facilities has the potential to support the realization of increased mobility benefits by their very nature. The proposed CAV operation involves platoons of equipped vehicles governed by integrated CAV applications, including cooperative adaptive cruise control (CACC), cooperative merge, and speed harmonization. This study proposes an algorithm for integrating CAV applications. Through microscopic simulation, the study particularly examines the effectiveness of CACC, CACC plus cooperative merge, and the addition of speed harmonization under different penetration rates. Simulation results show the effectiveness of the bundled application to enhance system throughput and reduce delay, even with low CAV penetration rates. The speed harmonization shows the greatest effects on delay reduction at medium-to-high penetration rates and some benefits even at low penetration rates. The conclusions provide operational insights and guidance for traffic management centers to implement CAV-based traffic control in the future. Jiaqi Ma 0003, Edward M. Leslie, Zhitong Huang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Cooperative Perception for Estimating and Predicting Microscopic Traffic States to Manage Connected and Automated TrafficabstractReal-time traffic state estimation and prediction are of importance to the traffic management systems. New opportunities are enabled by the emerging sensing and automation technologies to manage connected and automated traffic, particularly in terms of controlling trajectories of automated vehicles. Traffic information from connected and automated vehicles (CAV) and roadside detectors (RSD) can be fused and has great potential for providing detailed microscopic traffic states (i.e., vehicle speeds, positions) of all vehicles. In this paper, we propose a cooperative perception framework for this purpose. The proposed framework based on particle filtering is developed to provide an accurate estimation and prediction of the microscopic states of partially observed traffic systems, while accounting for different sources of errors that intrinsically exist in the system, including those from sensor data, vehicle movement, and process models. Selected freeway and arterial vehicle trajectory datasets from the Next Generation Simulation (NGSIM) program and CAV traffic simulation are applied to test the proposed methodological framework. The accuracy of position and speed estimation is between 50% and 70% when the CAV market penetration rate (MPR) is 12.5%, and between 80% and 90% when the MPR is 50%. The incorporation of RSD data can further increase the accuracy by up to 10% under low CAV MPRs. The framework can also provide an accurate short-term prediction (i.e., 5 – 15 seconds) of position and speed with 60% to 90% accuracy. The proposed framework provides efficient and accurate estimations and predictions of detailed microscopic traffic states, even at low CAV MPRs, creating dynamic traffic environment world models to enable fine control and management of the connected and automated traffic systems. Xu Han 0014, Jiaqi Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Short-Term Prediction of Signal Cycle on an Arterial With Actuated-Uncoordinated Control Using Sparse Time Series ModelsabstractTraffic signals as part of intelligent transportation systems can play a significant role in making cities smart. Conventionally, most traffic lights are designed with fixed-time control, which induces a lot of slack time (unused green time). Actuated traffic lights control traffic flow in real time and are more responsive to the variation of traffic demands. For an isolated signal, a family of time series models, such as autoregressive integrated moving average (ARIMA) models, can be beneficial for predicting the next cycle length. However, when there are multiple signals placed along a corridor with different spacing and configurations, the cycle length variation of such signals is not just related to each signal’s values, but it is also affected by the platoon of vehicles coming from neighboring intersections. In this paper, a multivariate time series model is developed to analyze the behavior of signal cycle lengths of multiple intersections placed along a corridor in a fully actuated setup. Five signalized intersections have been modeled along a corridor, with different spacing among them, together with multiple levels of traffic demand. To tackle the high-dimensional nature of the problem, a penalized least-squares method is utilized in the estimation procedure to output sparse models. Two proposed sparse time series methods captured the signal data reasonably well and outperformed the conventional vector autoregressive model—in some cases up to 17%—as well as being more powerful than univariate models, such as ARIMA. Bahman Moghimi, Abolfazl Safikhani, Camille Kamga, Wei Hao 0002, Jiaqi Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2016 | How big data serves for freight safety management at highway-rail grade crossings? A spatial approach fused with path analysis
Jun Liu 0009, Xin Wang 0018, Asad J. Khattak, Jia Hu 0003, JianXun Cui, Jiaqi Ma 0003 |
Neurocomputing | 6 |