EDBT 2026 Demo / reviewers in the wild / expert
Xun Huang 0003
dblp:121/2459-3
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0007-5343-7811ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR LocalizationabstractMulti-agents rely on accurate poses to share and align observations, enabling a collaborative perception of the environment. However, traditional GNSS-based localization often fails in GNSS-denied environments, making consistent feature alignment difficult in collaboration. To tackle this challenge, we propose a robust GNSS-free collaborative perception framework based on LiDAR localization. Specifically, we propose a lightweight Pose Generator with Confidence (PGC) to estimate compact pose and confidence representations. To alleviate the effects of localization errors, we further develop the Pose-Aware Spatio-Temporal Alignment Transformer (PASTAT), which performs confidence-aware spatial alignment while capturing essential temporal context. Additionally, we present a new simulation dataset, V2VLoc, which can be adapted for both LiDAR localization and collaborative detection tasks. V2VLoc comprises three subsets: Town1Loc, Town4Loc, and V2VDet. Town1Loc and Town4Loc offer multi-traversal sequences for training in localization tasks, whereas V2VDet is specifically intended for the collaborative detection task. Extensive experiments conducted on the V2VLoc dataset demonstrate that our approach achieves state-of-the-art performance under GNSS-denied conditions. We further conduct extended experiments on the real-world V2V4Real dataset to validate the effectiveness and generalizability of PASTAT. Wenkai Lin, Qiming Xia, Wen Li 0005, Xun Huang 0003, Chenglu Wen |
AAAI | 4 |
| 2026 | DOtA++: Unsupervisely and Collaboratively Detect Objects From Multi-Agent Observations With Multi-Modal Prior ConstraintsabstractEnhancing perception performance via multi-agent collaboration has gained increasing attention in the field of autonomous driving. However, as the number of agents grows, the manual annotation required for training collaborative detectors increases significantly. To tackle this problem, we introduce an unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, named DOtA, without using labels from external. DOtA first generates preliminary labels by an initial detector, which is trained by internally shared information of collaborative agents. DOtA then optimizes these preliminary labels by utilizing the physical rule constraints derived from the surrounding area of the object. Building on DOtA, we further propose DOtA++, an enhanced version that improves performance by leveraging composite prior constraints. Beyond physical rule constraints, DOtA++ further uses image data as an auxiliary modality to introduce multi-agent observation consistency constraints, boosting object classification, while also incorporating point cloud geometric distribution constraints to improve structural description. Extensive experiments on widely-used benchmarks demonstrate that DOtA and DOtA++ effectively perceive potential objects in the scene without manual annotations. In particular, DOtA++ shows 10.7% mAP improvement over traditional unsupervised methods on V2X-R dataset. Qiming Xia, Longhui Zheng, Shijia Zhao, Xun Huang 0003, Chenglu Wen, Cheng Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object DetectionabstractLiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, it faces challenges of significant differences in terms of data quality and the degree of degradation in adverse weather. To address these issues, we introduce L4DR, a weather-robust 3D object detection method that effectively achieves LiDAR and 4D Radar fusion. Our L4DR proposes Multi-Modal Encoding (MME) and Foreground-Aware Denoising (FAD) modules to reconcile sensor gaps, which is the first exploration of the complementarity of early fusion between LiDAR and 4D radar. Additionally, we design an Inter-Modal and IntraModal ({IM}2) parallel feature extraction backbone coupled with a Multi-Scale Gated Fusion (MSGF) module to counteract the varying degrees of sensor degradation under adverse weather conditions. Experimental evaluation on a VoD dataset with simulated fog proves that L4DR is more adaptable to changing weather conditions. It delivers a significant performance increase under different fog levels, improving the 3D mAP by up to 20.0% over the traditional LiDAR-only approach. Moreover, the results on the K-Radar dataset validate the consistent performance improvement of L4DR in realworld adverse weather conditions. Xun Huang 0003, Ziyu Xu 0002, Qiming Xia, Yan Xia 0003, Jonathan Li 0001, Kyle Gao, Chenglu Wen, Cheng Wang 0003 |
AAAI | 1 |
| 2025 | V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object DetectionabstractCurrent Vehicle-to-Everything (V2X) systems have significantly enhanced 3D object detection using LiDAR and camera data. However, they face performance degradation in adverse weather. Weather-robust 4D radar, with Doppler velocity and additional geometric information, offers a promising solution to this challenge. To this end, we present V2X-R, the first simulated V2X dataset incorporating LiDAR, camera, and 4D radar modalities. V2X-R contains 12,079 scenarios with 37,727 frames of LiDAR and 4D radar point clouds, 150,908 images, and 170,859 annotated 3D vehicle bounding boxes. Subsequently, we propose a novel cooperative LiDAR-4D radar fusion pipeline for 3D object detection and implement it with multiple fusion strategies. To achieve weather-robust detection, we additionally propose a Multi-modal Denoising Diffusion (MDD) module in our fusion pipeline. MDD utilizes weather-robust 4D radar feature as a condition to guide the diffusion model in denoising noisy LiDAR features. Experiments show that our LiDAR-4D radar fusion pipeline demonstrates superior performance in the V2X-R dataset. Over and above this, our MDD module further improved the foggy/snowy performance of the basic fusion model by up to 5.73%/6.70% and barely disrupting normal performance. The dataset and code will be publicly available at: https://github.com/ylwhxht/V2X-R. Xun Huang 0003, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li 0003, Cheng Wang 0003, Chenglu Wen |
CVPR | 1 |
| 2025 | Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual LabelsabstractUnsupervised 3D object detection serves as an important solution for offline 3D object annotation. However, due to the data sparsity and limited views, the clustering-based label fitting in unsupervised object detection often generates low-quality pseudo-labels. Multi-agent collaborative dataset, which involves the sharing of complementary observations among agents, holds the potential to break through this bottleneck. In this paper, we introduce a novel unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, termed DOtA, without using labels from external. DOtA first uses the internally shared ego-pose and ego-shape of collaborative agents to initialize the detector, leveraging the generalization performance of neural networks to infer preliminary labels. Subsequently, DOtA uses the complementary observations between agents to perform multi-scale encoding on preliminary labels, then decodes high-quality and low-quality labels. These labels are further used as prompts to guide a correct feature learning process, thereby enhancing the performance of the unsupervised object detection task. Extensive experiments on the V2V4Real and OPV2V datasets show that our DOtA outperforms state-of-the-art unsupervised 3D object detection methods. Additionally, we also validate the effectiveness of the DOtA labels under various collaborative perception frameworks. The code is available at https://github.com/xmuqimingxia/DOtA. Qiming Xia, Wenkai Lin, Haoen Xiang, Xun Huang 0003, Siheng Chen, Zhen Dong 0005, Cheng Wang 0003, Chenglu Wen |
CVPR | 4 |
| 2025 | Unsupervised 3D Object Detection by Commonsense ClueabstractTraditional 3D object detectors, whether fully-, semi-, or weakly-supervised, rely heavily on extensive human annotations. In contrast, this paper introduces an unsupervised 3D object detector that automatically discerns object patterns without such annotations. To achieve this, we propose a Commonsense Prototype-based Detector (CPD) for unsupervised 3D object detection. CPD first constructs Commonsense Prototypes (CProto) to represent the geometric center and size of objects. It then generates high-quality pseudo-labels and guides detector convergence using size and geometry priors from CProto. Building on CPD, we further introduce CPD++, an enhanced version that improves performance by leveraging motion cues. CPD++ learns localization from stationary objects and recognition from moving objects, facilitating the mutual transfer of localization and recognition knowledge between these two object types. Both CPD and CPD++ outperform existing state-of-the-art unsupervised 3D detectors. Furthermore, when trained on Waymo Open Dataset (WOD) and tested on KITTI, CPD++ achieves 89.25% 3D Average Precision (AP) on the moderate car class at a 0.5 IoU threshold, reaching 95.3% of the performance attained by fully supervised counterparts. These results underscore the significant advancements brought by our method. Shijia Zhao, Xun Huang 0003, Qiming Xia, Chenglu Wen, Li Jiang 0009, Xin Li 0003, Cheng Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Sunshine to Rainstorm: Cross-Weather Knowledge Distillation for Robust 3D Object DetectionabstractLiDAR-based 3D object detection models inevitably struggle under rainy conditions due to the degraded and noisy scanning signals. Previous research has attempted to address this by simulating the noise from rain to improve the robustness of detection models. However, significant disparities exist between simulated and actual rain-impacted data points. In this work, we propose a novel rain simulation method, termed DRET, that unifies Dynamics and Rainy Environment Theory to provide a cost-effective means of expanding the available realistic rain data for 3D detection training. Furthermore, we present a Sunny-to-Rainy Knowledge Distillation (SRKD) approach to enhance 3D detection under rainy conditions. Extensive experiments on the Waymo-Open-Dataset show that, when combined with the state-of-the-art DSVT model and other classical 3D detectors, our proposed framework demonstrates significant detection accuracy improvements, without losing efficiency. Remarkably, our framework also improves detection capabilities under sunny conditions, therefore offering a robust solution for 3D detection regardless of whether the weather is rainy or sunny. Xun Huang 0003, Xin Li 0003, Xiaoliang Fan, Chenglu Wen, Cheng Wang 0003 |
AAAI | 1 |
| 2024 | Commonsense Prototype for Outdoor Unsupervised 3D Object DetectionabstractThe prevalent approaches of unsupervised 3D object de-tection follow cluster-based pseudo-label generation and iterative self-training processes. However, the challenge arises due to the sparsity of LiDAR scans, which leads to pseudo-labels with erroneous size and position, resulting in subpar detection performance. To tackle this problem, this paper introduces a Commonsense Prototype-based Detector, termed CPD, for unsupervised 3D object de-tection. CPD first constructs Commonsense Prototype (CProto) characterized by high-quality bounding box and dense points, based on commonsense intuition. Subse-quently, CPD refines the low-quality pseudo-labels by lever-aging the size prior from CProto. Furthermore, CPD en-hances the detection accuracy of sparsely scanned objects by the geometric knowledge from CProto. CPD outper-forms state-of-the-art unsupervised 3D detectors on Waymo Open Dataset (WOD), PandaSet, and KITTI datasets by a large margin. Besides, by training CPD on WOD and testing on KITTI, CPD attains 90.85% and 81.01% 3D Aver-age Precision on easy and moderate car classes, respectively. These achievements position CPD in close prox-imity to fully supervised detectors, highlighting the sig-nificance of our method. The code will be available at https://github.com/hailanyi/CPD. Shijia Zhao, Xun Huang 0003, Chenglu Wen, Xin Li 0003, Cheng Wang 0003 |
CVPR | 3 |
| 2024 | HINTED: Hard Instance Enhanced Detector with Mixed-Density Feature Fusion for Sparsely-Supervised 3D Object DetectionabstractCurrent sparsely-supervised object detection methods largely depend on high threshold settings to derive high-quality pseudo labels from detector predictions. However, hard instances within point clouds frequently display incomplete structures, causing decreased confidence scores in their assigned pseudo-labels. Previous methods inevitably result in inadequate positive supervision for these instances. To address this problem, we propose a novel Hard INsTance Enhanced Detector (HINTED), for sparsely-supervised 3D object detection. Firstly, we design a self-boosting teacher (SBT) model to generate more potential pseudo-labels, enhancing the effectiveness of information transfer. Then, we introduce a mixed-density student (MDS) model to concentrate on hard instances during the training phase, thereby improving detection accuracy. Our extensive experiments on the KITTI dataset validate our method's superior performance. Compared with leading sparsely-supervised methods, HINTED significantly improves the detection performance on hard instances, no-tably outperforming fully-supervised methods in detecting challenging categories like cyclists. HINTED also significantly outperforms the state-of-the-art semi-supervised method on challenging categories. The code is available at https://github.com/xmuqimingxia/HINTED. Qiming Xia, Shijia Zhao, Leyuan Xing, Xun Huang 0003, Jinhao Deng, Xin Li 0003, Chenglu Wen, Cheng Wang 0003 |
CVPR | 6 |
| 2024 | CMD: A Cross Mechanism Domain Adaptation Dataset for 3D Object Detection
Jinhao Deng, Xun Huang 0003, Qiming Xia, Xin Li 0003, Wei Li 0111, Chenglu Wen, Cheng Wang 0003 |
ECCV (57) | 4 |