Hao Xiang 0001

dblp:156/3706-1 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0001-7591-9546ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 14 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction
abstract
Vehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores temporal cues and temporal tasks (e.g., temporal perception and prediction). In this paper, we focus on the spatio-temporal fusion in V2X scenarios and design one-step and multi-step communication strategies (when to transmit) as well as examine their integration with three fusion strategies - early, late, and intermediate (what to transmit), providing comprehensive benchmarks with 11 fusion models (how to fuse). Furthermore, we propose V2XPnP, a novel intermediate fusion framework within one-step communication for end-to-end perception and prediction. Our framework employs a unified Transformer-based architecture to effectively model complex spatio-temporal relationships across multiple agents, frames, and high-definition maps. Moreover, we introduce the V2XPnP Sequential Dataset that supports all V2X collaboration modes and addresses the limitations of existing real-world datasets, which are restricted to single-frame or single-mode cooperation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods in both perception and prediction tasks.
Zewei Zhou, Hao Xiang 0001, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Tianhui Cai, Johnson Liu, Maheswari Bajji, Xin Xia 0007, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003
ICCV2
2025 CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
abstract
Multi-agent collaborative perception enhances each agent’s perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor deficiencies, occlusions, and long-range perception. However, existing representative collaborative perception systems transmit intermediate feature maps, such as bird’s-eye view (BEV) representations, which contain a significant amount of non-critical information, leading to high communication bandwidth requirements. To enhance communication efficiency while preserving perception capability, we introduce CoCMT, an object-query-based collaboration framework that optimizes communication bandwidth by selectively extracting and transmitting essential features. Within CoCMT, we introduce the Efficient Query Transformer (EQFormer) to effectively fuse multi-agent object queries and implement a synergistic deep supervision to enhance the positive reinforcement between stages, leading to improved overall performance. Experiments on OPV2V and V2V4Real datasets show CoCMT outperforms state-of-the-art methods while drastically reducing communication needs. On V2V4Real, our model (Top-50 object queries) requires only 0.416 Mb bandwidth—83 times less than SOTA methods—while improving AP@70 by 1.1%. This efficiency breakthrough enables practical collaborative perception deployment in bandwidth-constrained environments without sacrificing detection accuracy. The code and models are open-sourced through the following link: https://github.com/taco-group/COCMT.
Rujia Wang, Xiangbo Gao, Hao Xiang 0001, Runsheng Xu, Zhengzhong Tu
IROS3
2025 CooPre: Cooperative Pretraining for V2X Cooperative Perception
abstract
Existing Vehicle-to-Everything (V2X) cooperative perception methods rely on accurate multi-agent 3D annotations. Nevertheless, it is time-consuming and expensive to collect and annotate real-world data, especially for V2X systems. In this paper, we present a self-supervised learning framwork for V2X cooperative perception, which utilizes the vast amount of unlabeled 3D V2X data to enhance the perception performance. Specifically, multi-agent sensing information is aggregated to form a holistic view and a novel proxy task is formulated to reconstruct the LiDAR point clouds across multiple connected agents to better reason multi-agent spatial correlations. Besides, we develop a V2X bird-eye-view (BEV) guided masking strategy which effectively allows the model to pay attention to 3D features across heterogeneous V2X agents (i.e., vehicles and infrastructure) in the BEV space. Noticeably, such a masking strategy effectively pretrains the 3D encoder with a multi-agent LiDAR point cloud reconstruction objective and is compatible with mainstream cooperative perception backbones. Our approach, validated through extensive experiments on representative datasets (i.e., V2X-Real, V2V4Real, and OPV2V) and multiple state-of-the-art cooperative perception methods (i.e., AttFuse, F-Cooper, and V2X-ViT), leads to a performance boost across all V2X settings. Notably, CooPre achieves a 4% mAP improvement on V2X-Real dataset and surpasses baseline performance using only 50% of the training data, highlighting its data efficiency. Additionally, we demonstrate the framework’s powerful performance in cross-domain transferability and robustness under challenging scenarios. The code will be made publicly available at https://github.com/ucla-mobility/CooPre.
Seth Z. Zhao, Hao Xiang 0001, Chenfeng Xu, Xin Xia 0007, Bolei Zhou, Jiaqi Ma 0003
IROS2
2024 V2X-Real: A Largs-Scale Dataset for Vehicle-to-Everything Cooperative Perception
Hao Xiang 0001, Zhaoliang Zheng, Xin Xia 0007, Runsheng Xu, Letian Gao, Zewei Zhou, Xu Han 0014, Xinkai Ji, Zonglin Meng, Mingyue Lei, Haoxuan Ma, Yunshuang Yuan, Yingqian Zhao, Jiaqi Ma 0003
ECCV (52)1
2024 End-to-end Cooperative Localization via Neural Feature Sharing
abstract
Cooperative driving automation attracts great attention for its potential to improve traffic safety. Knowing each vehicle’s accurate position serves as the cornerstone for the information fusion necessary in cooperative driving tasks. However, inherent errors within a vehicle’s self-localization system often necessitate correction to facilitate cooperative perception and downstream tasks. Leveraging intermediate features shared among other Connected Automated Vehicles (CAVs), we propose an end-to-end learning localization framework aimed at estimating the relative pose error between the ego vehicle and the CAV. We investigate factors that may influence learning performance and validate the algorithm’s performance using a simulation dataset. The proposed method is compared with the traditional point cloud matching-based relative localization method. Remarkably, our framework effectively corrects relative pose errors even when the vehicle exhibits significant initial localization inaccuracies, and it can be integrated into the cooperative perception system.
Letian Gao, Hao Xiang 0001, Xin Xia 0007, Jiaqi Ma 0003
IV2
2024 CooperFuse: A Real-Time Cooperative Perception Fusion Framework
abstract
Cooperative perception algorithms based on fusing sensing data across multiple connected automated vehicles (CAVs) have shown promising performance to enhance the existing individual perception algorithm in terms of object detection tracking. However, existing cooperative perception algorithms are only developed offline given the constraints from data sharing and computational resources and none of them have been verified in real-time conditions. In this work, we propose a real-time cooperative perception framework called CooperFuse, which achieves cooperative perception in a late fusion scheme. Based on object detection and tracking results from individual vehicle, the late fusion cooperative perception algorithm considers object detection confidence score, kinematics, and dynamics consistency as well as scale consistency of detected objects. The algorithm computes the kinematic and dynamic consistency of the objects by solving for the energy consumption of inter-frame trajectories, and determines scale consistency by calculating inter-frame scale changes, enabling feature-based bounding box fusion. The experimental results demonstrate the real-time performance of the proposed algorithm and reveal its effective improvements in feature fusion and object detection accuracy when dealing with heterogeneous detection models across different cooperative intelligent agents.
Zhaoliang Zheng, Xin Xia 0007, Letian Gao, Hao Xiang 0001, Jiaqi Ma 0003
IV4
2023 V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception
abstract
Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at research.seas.ucla.edu/mobility-lab/v2v4real/.
Runsheng Xu, Xin Xia 0007, Hanzhao Li, Zhengzhong Tu, Zonglin Meng, Hao Xiang 0001, Rui Song 0007, Hongkai Yu, Bolei Zhou, Jiaqi Ma 0003
CVPR8
2023 Optimizing the Placement of Roadside LiDARs for Autonomous Driving
abstract
Multi-agent cooperative perception is an increasingly popular topic in the field of autonomous driving, where roadside LiDARs play an essential role. However, how to optimize the placement of roadside LiDARs is a crucial but often overlooked problem. This paper proposes an approach to optimize the placement of roadside LiDARs by selecting optimized positions within the scene for better perception performance. To efficiently obtain the best combination of locations, a greedy algorithm based on perceptual gain is proposed, which selects the location that can maximize the perceptual gain sequentially. We define perceptual gain as the increased perceptual capability when a new LiDAR is placed. To obtain the perception capability, we propose a perception predictor that learns to evaluate LiDAR placement using only a single point cloud frame. A dataset named Roadside-Opt is created using the CARLA simulator to facilitate research on the roadside LiDAR placement problem. Extensive experiments are conducted to demonstrate the effectiveness of our proposed method.
Hao Xiang 0001, Xinyu Cai, Runsheng Xu, Jiaqi Ma 0003, Gim Hee Lee, Si Liu 0001
ICCV2
2023 HM-ViT: Hetero-modal Vehicle-to-Vehicle Cooperative Perception with Vision Transformer
abstract
Vehicle-to-Vehicle technologies have enabled autonomous vehicles to share information to see through occlusions, greatly enhancing perception performance. Nevertheless, existing works all focused on homogeneous traffic where vehicles are equipped with the same type of sensors, which significantly hampers the scale of collaboration and benefit of cross-modality interactions. In this paper, we investigate the multi-agent hetero-modal cooperative perception problem where agents may have distinct sensor modalities. We present HM-ViT, the first unified multi-agent hetero-modal cooperative perception framework that can collaboratively predict 3D objects for highly dynamic Vehicle-to-Vehicle (V2V) collaborations with varying numbers and types of agents. To effectively fuse features from multi-view images and LiDAR point clouds, we design a novel heterogeneous 3D graph transformer to jointly reason inter-agent and intra-agent interactions. The extensive experiments on the V2V perception dataset OPV2V demonstrate that the HM-ViT outperforms SOTA cooperative perception methods for V2V hetero-modal cooperative perception. Our code will be released at https://github.com/XHwind/HM-ViT.
Hao Xiang 0001, Runsheng Xu, Jiaqi Ma 0003
ICCV1
2023 V2XP-ASG: Generating Adversarial Scenes for Vehicle-to-Everything Perception
abstract
Recent advancements in Vehicle-to-Everything communication technology have enabled autonomous vehicles to share sensory information to obtain better perception performance. With the rapid growth of autonomous vehicles and intelligent infrastructure, the V2X perception systems will soon be deployed at scale, which raises a safety-critical question: how can we evaluate and improve its performance under challenging traffic scenarios before the real-world deployment? Collecting diverse large-scale real-world test scenes seems to be the most straightforward solution, but it is expensive and time-consuming, and the collections can only cover limited scenarios. To this end, we propose the first open adversarial scene generator V2XP-ASG that can produce realistic, challenging scenes for modern LiDAR-based multi-agent perception systems. V2XP-ASG learns to construct an adversarial collaboration graph and simultaneously perturb multiple agents' poses in an adversarial and plausible manner. The experiments demonstrate that V2XP-ASG can effectively identify challenging scenes for a large range of V2X perception systems. Meanwhile, by training on the limited number of generated challenging scenes, the accuracy of V2X perception systems can be further improved by 12.3% on challenging and 4% on normal scenes. Our code will be released at https://github.com/XHwind/V2XP-ASG.
Hao Xiang 0001, Runsheng Xu, Xin Xia 0007, Zhaoliang Zheng, Bolei Zhou, Jiaqi Ma 0003
ICRA1
2023 Model-Agnostic Multi-Agent Perception Framework
abstract
Existing multi-agent perception systems assume that every agent utilizes the same model with identical parameters and architecture. The performance can be degraded with different perception models due to the mismatch in their confidence scores. In this work, we propose a model-agnostic multi-agent perception framework to reduce the negative effect caused by the model discrepancies without sharing the model information. Specifically, we propose a confidence calibrator that can eliminate the prediction confidence score bias. Each agent performs such calibration independently on a standard public database to protect intellectual property. We also propose a corresponding bounding box aggregation algorithm that considers the confidence scores and the spatial agreement of neighboring boxes. Our experiments shed light on the necessity of model calibration across different agents, and the results show that the proposed framework improves the baseline 3D object detection performance of heterogeneous agents. The code can be found at this url.
Runsheng Xu, Weizhe Chen 0004, Hao Xiang 0001, Xin Xia 0007, Lantao Liu, Jiaqi Ma 0003
ICRA3
2022 V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer
Runsheng Xu, Hao Xiang 0001, Zhengzhong Tu, Xin Xia 0007, Ming-Hsuan Yang 0001, Jiaqi Ma 0003
ECCV (39)2
2022 TridentNetV2: Lightweight Graphical Global Plan Representations for Dynamic Trajectory Generation
abstract
We present a framework for dynamic trajectory generation for autonomous navigation, which does not rely on HD maps as the underlying representation. High Definition (HD) maps have become a key component in most autonomous driving frameworks, which include complete road network information annotated at a centimeter-level that include traversable waypoints, lane information, and traffic signals. Instead, the presented approach models the distributions of feasible ego-centric trajectories in real-time given a nominal graph-based global plan and a lightweight scene representation. By embedding contextual information, such as crosswalks, stop signs, and traffic signals, our approach achieves low errors across multiple urban navigation datasets that include diverse intersection maneuvers, while maintaining real-time performance and reducing network complexity. Underlying datasets introduced are available online.
David Paz, Hao Xiang 0001, Andrew Liang, Henrik I. Christensen
ICRA2
2022 OPV2V: An Open Benchmark Dataset and Fusion Pipeline for Perception with Vehicle-to-Vehicle Communication
abstract
Employing Vehicle-to-Vehicle communication to enhance perception performance in self-driving technology has attracted considerable attention recently; however, the absence of a suitable open dataset for benchmarking algorithms has made it difficult to develop and assess cooperative perception technologies. To this end, we present the first large-scale open simulated dataset for Vehicle-to-Vehicle perception. It contains over 70 interesting scenes, 11,464 frames, and 232,913 annotated 3D vehicle bounding boxes, collected from 8 towns in CARLA and a digital town of Culver City, Los Angeles. We then construct a comprehensive benchmark with a total of 16 implemented models to evaluate several information fusion strategies (i.e. early, late, and intermediate fusion) with state-of-the-art LiDAR detection algorithms. Moreover, we propose a new Attentive Intermediate Fusion pipeline to aggregate information from multiple connected vehicles. Our experiments show that the proposed pipeline can be easily integrated with existing 3D LiDAR detectors and achieve outstanding performance even with large compression rates. To encourage more researchers to investigate Vehicle-to-Vehicle perception, we will release the dataset, benchmark methods, and all related codes in https://mobility-lab.seas.ucla.edu/opv2v/.
Runsheng Xu, Hao Xiang 0001, Xin Xia 0007, Xu Han 0014, Jiaqi Ma 0003
ICRA2
2020 Probabilistic Semantic Mapping for Urban Autonomous Driving Applications
abstract
Recent advancements in statistical learning and computational abilities have enabled autonomous vehicle technology to develop at a much faster rate. While many of the architectures previously introduced are capable of operating under highly dynamic environments, many of these are constrained to smaller-scale deployments, require constant maintenance due to the associated scalability cost with highdefinition (HD) maps, and involve tedious manual labeling. As an attempt to tackle this problem, we propose to fuse image and pre-built point cloud map information to perform automatic and accurate labeling of static landmarks such as roads, sidewalks, crosswalks, and lanes. The method performs semantic segmentation on 2D images, associates the semantic labels with point cloud maps to accurately localize them in the world, and leverages the confusion matrix formulation to construct a probabilistic semantic map in bird’s eye view from semantic point clouds. Experiments from data collected in an urban environment show that this model is able to predict most road features and can be extended for automatically incorporating road features into HD maps with potential future work directions.
David Paz, Hengyuan Zhang 0001, Qinru Li, Hao Xiang 0001, Henrik I. Christensen
IROS4