EDBT 2026 Demo / reviewers in the wild / expert
Zhe Wang 0070
dblp:75/3158-70
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3201-1903ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data DomainabstractIn autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain gaps. To bridge this gap and Improve ROAdside Monocular 3D object detection, we propose IROAM, a semantic-geometry decoupled contrastive learning framework, which takes vehicle-side and roadside data as input simultaneously. IROAM has two significant modules. In-Domain Query Interaction module utilizes a transformer to learn content and depth information for each domain and outputs object queries. Cross-Domain Query Enhancement To learn better feature representations from two domains, Cross-Domain Query Enhancement decouples queries into semantic and geometry parts and only the former is used for contrastive learning. Experiments demonstrate the effectiveness of IROAM in improving roadside detector's performance. The results validate that IROAM has the capabilities to learn cross-domain information. Zhe Wang 0070, Xiaoliang Huo, Siqi Fan 0002, Ya-Qin Zhang, Yan Wang 0105 |
ICRA | 1 |
| 2025 | CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object QueryabstractCooperative perception enhances the individual perception capabilities of autonomous vehicles (AVs) by providing a comprehensive view of the environment. However, balancing perception performance and transmission costs remains a significant challenge. Current approaches that transmit regionlevel features across agents are limited in interpretability and demand substantial bandwidth, making them unsuitable for practical applications. In this work, we propose CoopDETR, a novel cooperative perception framework that introduces objectlevel feature cooperation via object query. Our framework consists of two key modules: single-agent query generation, which efficiently encodes raw sensor data into object queries, reducing transmission cost while preserving essential information for detection; and cross-agent query fusion, which includes Spatial Query Matching (SQM) and Object Query Aggregation (OQA) to enable effective interaction between queries. Our experiments on the OPV2V and V2XSet datasets demonstrate that CoopDETR achieves state-of-the-art performance and significantly reduces transmission costs to 1/782 of previous methods. Zhe Wang 0070, Shaocong Xu, Xucai Zhuang, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
ICRA | 1 |
| 2024 | Idempotence and Perceptual Image CompressionabstractIdempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence constraint is equivalent to conditional generative codec. Based on this newfound equivalence, we propose a new paradigm of perceptual image codec by inverting unconditional generative model with idempotence constraints. Our codec is theoretically equivalent to conditional generative codec, and it does not require training new models. Instead, it only requires a pre-trained mean-square-error codec and unconditional generative model. Empirically, we show that our proposed approach outperforms state-of-the-art methods such as HiFiC and ILLM, in terms of Fréchet Inception Distance (FID). The source code is provided in https://github.com/tongdaxu/Idempotence-and-Perceptual-Image-Compression. Tongda Xu, Ziran Zhu, Dailan He, Yanghao Li, Zhe Wang 0070, Hongwei Qin, Yan Wang 0105, Ya-Qin Zhang |
ICLR | 7 |
| 2024 | EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object DetectionabstractIn autonomous driving, cooperative perception makes use of multi-view cameras from both vehicles and infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle viewpoint. Currently, two major challenges persist in vehicle-infrastructure cooperative 3D (VIC3D) object detection: 1) inherent pose errors when fusing multi-view images, caused by time asynchrony across cameras; 2) information loss in transmission process resulted from limited communication bandwidth. To address these issues, we propose a novel camera-based 3D detection framework for VIC3D task, Enhanced Multi-scale Image Feature Fusion (EMIFF). To fully exploit holistic perspectives from both vehicles and infrastructure, we propose Multi-scale Cross Attention (MCA) and Camera-aware Channel Masking (CCM) modules to enhance infrastructure and vehicle features at scale, spatial, and channel levels to correct the pose error introduced by camera asynchrony. We also introduce a Feature Compression (FC) module with channel and spatial compression blocks for transmission efficiency. Experiments show that EMIFF achieves SOTA on DAIR-V2X-C datasets, significantly outperforming previous early-fusion and late-fusion methods with comparable transmission costs. Zhe Wang 0070, Siqi Fan 0002, Xiaoliang Huo, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
ICRA | 1 |
| 2023 | Your Camera Improves Your Point Cloud CompressionabstractLiDAR point cloud compression is important for autonomous driving as it consumes a lot of storage and bandwidth. Although the fusion of camera and LiDAR for vision perception has been well studied, it remains unexplored that how we can improve the compression of LiDAR point cloud data using cross-modal information from cameras. In this paper’ we propose a multi-modality compression framework for LiDAR point cloud by exploiting the depth information predicted from its paired image. To the best of our knowledge’ our model is the first multi-modality compression framework for point cloud. Specifically’ we first represent point cloud based on octrees to reduce spatial redundancy. Then’ we propose a cross-modal fusion structure to improve the compression of these octrees’ with depth distribution extracted from the camera pixels and acts as side information. Compared to previous state-of-the-art (SOTA) method, our approach obtains up to 8.10% compression rate gain for LiDAR point cloud compression. Yuhuan Lin, Tongda Xu, Yanghao Li, Zhe Wang 0070, Yan Wang 0105 |
ICASSP | 5 |
| 2023 | Calibration-Free BEV Representation for Infrastructure PerceptionabstractEffective BEV object detection on infrastructure can greatly improve traffic scene understanding and vehicle-to-infrastructure (V2I) cooperative perception. However, cameras installed on infrastructure have various postures, and previous BEV detection methods rely on accurate calibration, which is difficult for practical applications due to inevitable natural factors (e.g., wind and snow). In this paper, we propose a Calibration-free BEV Representation (CBR) network, which achieves 3D detection based on BEV representation without calibration parameters and additional depth supervision. Specifically, we utilize two multi-layer perceptrons for decoupling the features from perspective view to front view and bird-eye view under boxes-induced foreground supervision. Then, a cross-view feature fusion module matches features from orthogonal views according to similarity and conducts BEV feature enhancement with front-view features. Experimental results on DAIR-V2X demonstrate that CBR achieves acceptable performance without any camera parameters and is naturally not affected by calibration noises. We hope CBR can serve as a baseline for future research addressing practical challenges of infrastructure perception. Siqi Fan 0002, Zhe Wang 0070, Xiaoliang Huo, Yan Wang 0105 |
IROS | 2 |