EDBT 2026 Demo / reviewers in the wild / expert
Zhenhua Xu 0003
dblp:126/7726-3
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-0700-2335ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous DrivingabstractMultimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which focused on open-loop tasks, this study explores the capabilities of LLMs in enhancing closed-loop autonomous driving. DriveGPT4-V2 processes camera images and vehicle states as input to generate low-level control signals for end-to-end vehicle operation. A multi-view visual tokenizer (MV-VT) is employed enabling DriveGPT4-V2 to perceive the environment with an extensive range while maintaining critical details. The model architecture has been refined to improve decision prediction and inference speed. To further enhance the performance, an additional expert LLM is trained for online imitation learning. The expert LLM, sharing a similar structure with DriveGPT4-V2, can access privileged information about surrounding objects for more robust and reliable predictions. Experimental results show that DriveGPT4-V2 outperforms all baselines on the challenging CARLA Longest6 benchmark. The code and data of DriveGPT4-V2 will be publicly available. Zhenhua Xu 0003, Yujia Zhang 0003, Zhuoling Li, Kwan-Yee Kenneth Wong, Hengshuang Zhao |
CVPR | 1 |
| 2025 | LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceabstractRecent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while demanding enormous computing resources. In this work, we combine their advantages while avoiding the drawbacks by conducting the proposed referee RL on our developed large auto-regressive model (LARM). Specifically, LARM is built upon a lightweight LLM (fewer than 5B parameters) and directly outputs the next action to execute rather than text. We mathematically reveal that classic RL feedbacks vanish in long-horizon embodied exploration and introduce a giant LLM based referee to handle this reward vanishment during training LARM. In this way, LARM learns to complete diverse open-world tasks without human intervention. Especially, LARM successfully harvests enchanted diamond equipment in Minecraft, which demands significantly longer decision-making chains than the highest achievements of prior best methods. Zhuoling Li, Xiaogang Xu 0002, Zhenhua Xu 0003, Ser-Nam Lim, Hengshuang Zhao |
ICML | 3 |
| 2025 | VIP: Vision Instructed Pre-training for Robotic ManipulationabstractThe effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would be confused if the task targets are not specified clearly. Existing works primarily rely on text instruction to describe targets. However, we reveal that current robotic data cannot train policies to understand text instruction effectively, and vision is much more comprehensible. Therefore, we introduce utilizing vision instruction to specify targets. A straightforward implementation is training a policy to predict the intermediate actions linking the current observation and a future image. Nevertheless, a single future image does not describe the task target in insufficient detail. To handle this problem, we propose to use sparse point flows to provide more detailed information. Extensive tasks are designed based on real and simulated environments to evaluate the effectiveness of our vision instructed pre-training (VIP) method. The results indicate VIP improves the performance on diverse tasks significantly, and the derived policy can complete competitive tasks like ``opening the lid of a tightly sealed bottle''. Zhuoling Li, Liangliang Ren, Xiaoyang Wu 0002, Zhenhua Xu 0003, Xiang Bai, Hengshuang Zhao |
ICML | 6 |
| 2024 | InsMapper: Exploring Inner-Instance Information for Vectorized HD Mapping
Zhenhua Xu 0003, Kwan-Yee Kenneth Wong, Hengshuang Zhao |
ECCV (34) | 1 |
| 2024 | PGO-IPM: Enhance IPM Accuracy with Pose-guided Optimization for Low-cost High-definition Angular Marking Map GenerationabstractHigh-definition angular marking maps (HDAM maps) are vital in large-scale environments with variable appearances. In these scenarios, unmanned ground vehicles (UGVs) can use angular markings for localization because they are easy to identify and informative for localization. However, creating such a marking map relies heavily on manual measurement and annotation, which is time-consuming and laborious. Although Inverse Perspective Mapping (IPM) offers a low-cost and automated alternative, its accuracy is compromised by vehicle motion and the arduous pre-calibration of the IPM matrix. To fill these gaps, we propose a pose-guided optimization framework for IPM. This framework enables the automated generation of HDAM maps, while concurrently refining the preliminary IPM matrix. We deployed the proposed method in two different automated ports, and the method yielded HDAM maps with near-centimeter precision. Moreover, the refined IPM matrix matched the accuracy of manual calibrations. The supplementary materials and videos are available at http://liuhongji.site/PGO-IPM/. Hongji Liu, Linwei Zheng, Xiaoyang Yan, Zhenhua Xu 0003, Bohuan Xue, Yang Yu 0028, Ming Liu 0001 |
IV | 4 |
| 2024 | FSNet: Redesign Self-Supervised MonoDepth for Full-Scale Depth Prediction for Autonomous DrivingabstractPredicting accurate depth with monocular images is important for low-cost robotic applications and autonomous driving. This study proposes a comprehensive self-supervised framework for accurate scale-aware depth prediction on autonomous driving scenes utilizing inter-frame poses obtained from inertial measurements. In particular, we introduce a Full-Scale depth prediction network named FSNet. FSNet contains four important improvements over existing self-supervised models: (1) a multichannel output representation for stable training of depth prediction in driving scenarios, (2) an optical-flow-based mask designed for dynamic object removal, (3) a self-distillation training strategy to augment the training process, and (4) an optimization-based post-processing algorithm in test time, fusing the results from visual odometry. With this framework, robots and vehicles with only one well-calibrated camera can collect sequences of training image frames and camera poses, and infer accurate 3D depths of the environment without extra labeling work or 3D data. Extensive experiments on the KITTI dataset, KITTI-360 dataset and the nuScenes dataset demonstrate the potential of FSNet. More visualizations are presented in https://sites.google.com/view/fsnet/homeNote to Practitioners—This paper was motivated by the problem of unsupervised monocular depth for robotic deployment. We notice that PoseNet is not generalizable and by nature monodepth2 only predict depths up to a scale. We believe that we should not expect PoseNet, a ResNet on a concatenation of two images, to produce more reliable poses than the localization module in a robot. So we try our best to completely avoid using PoseNet. This creates much unstability in training, but we managed to fix it in FSNet with multichannel output and self-distillation. We also believe the network should try to directly predict accurate depth with a correct scale at any cases. So our method could produce meaningful results on static frames or scenes with little/no VO points (same as the network’s direct prediction). There are images without VO points in our multi-frame experiment, but our method is robust enough to fix this problem. In future research, we will include multi-frame depth predictions for more accurate depth prediction. Yuxuan Liu 0008, Zhenhua Xu 0003, Huaiyang Huang, Lujia Wang 0001, Ming Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | CenterLineDet: CenterLine Graph Detection for Road Lanes with Vehicle-mounted Sensors by Transformer for HD Map GenerationabstractWith the fast development of autonomous driving technologies, there is an increasing demand for high-definition (HD) maps, which provide reliable and robust prior information about the static part of the traffic environments. As one of the important elements in HD maps, road lane centerline is critical for downstream tasks, such as prediction and planning. Manually annotating centerlines for road lanes in HD maps is labor-intensive, expensive and inefficient, severely restricting the wide applications of autonomous driving systems. Previous work seldom explores the lane centerline detection problem due to the complicated topology and severe overlapping issues of lane centerlines. In this paper, we propose a novel method named CenterLineDet to detect lane centerlines for automatic HD map generation. Our CenterLineDet is trained by imitation learning and can effectively detect the graph of centerlines with vehicle-mounted sensors (i.e., six cameras and one LiDAR) through iterations. Due to the use of the DETR-like transformer network, CenterLineDet can handle complicated graph topology, such as lane intersections. The proposed approach is evaluated on the large-scale public dataset NuScenes. The superiority of our CenterLineDet is demonstrated by the comparative results. Our code, supplementary materials, and video demonstrations are available at https://tonyxuqaq.github.io/projects/CenterLineDet/. Zhenhua Xu 0003, Yuxuan Liu 0008, Yuxiang Sun 0002, Ming Liu 0001, Lujia Wang 0001 |
ICRA | 1 |
| 2022 | Star-Convolution for Image-Based 3D Object Detectionabstract3D object detection with only image inputs is an interesting and important problem in computer vision and autonomous driving. Nowadays, most existing monocular 3D object detection algorithms rely solely on the approximation power of convolutional neural networks to learn a mapping from pixels to 3D predictions without knowing the projection matrix of the camera. To endow the networks with camera projection knowledge, we propose the Star-Convolution module for application to image-based 3D detection. The introduced module increases the receptive field of the detector and embeds the camera's projection geometry inside the network while keeping the network end-to-end trainable. We test the module with different baselines in both monocular and stereo 3D object detection, and we achieve significant improvements on both tasks. The code will be published at https://github.com/Owen-Liuyuxuan/visualDet3D. Yuxuan Liu 0008, Zhenhua Xu 0003, Ming Liu 0001 |
ICRA | 2 |
| 2022 | RNGDet: Road Network Graph Detection by Transformer in Aerial ImagesabstractRoad network graphs provide critical information for autonomous-vehicle applications, such as drivable areas that can be used for motion planning algorithms. To find road network graphs, manual annotation is usually inefficient and labor-intensive. Automatically detecting road network graphs could alleviate this issue, but existing works still have some limitations. For example, segmentation-based approaches could not ensure satisfactory topology correctness, and graph-based approaches could not present precise enough detection results. To provide a solution to these problems, we propose a novel approach based on transformer and imitation learning in this article. In view of that high-resolution aerial images could be easily accessed all over the world nowadays, we make use of aerial images in our approach. Taken as input an aerial image, our approach iteratively generates road network graphs vertex-by-vertex. Our approach can handle complicated intersection points with various numbers of incident road segments. We evaluate our approach on a publicly available dataset. The superiority of our approach is demonstrated through comparative experiments. Our work is accompanied by a demonstration video which is available athttps://tonyxuqaq.github.io/projects/RNGDet/. Zhenhua Xu 0003, Yuxuan Liu 0008, Lu Gan 0001, Yuxiang Sun 0002, Xinyu Wu 0001, Ming Liu 0001, Lujia Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | CP-loss: Connectivity-preserving Loss for Road Curb Detection in Autonomous Driving with Aerial ImagesabstractRoad curb detection is important for autonomous driving. It can be used to determine road boundaries to constrain vehicles on roads, so that potential accidents could be avoided. Most of the current methods detect road curbs online using vehicle-mounted sensors, such as cameras or 3-D Lidars. However, these methods usually suffer from severe occlusion issues. Especially in highly-dynamic traffic environments, most of the field of view is occupied by dynamic objects. To alleviate this issue, we detect road curbs offline using high-resolution aerial images in this paper. Moreover, the detected road curbs can be used to create high-definition (HD) maps for autonomous vehicles. Specifically, we first predict the pixel-wise segmentation map of road curbs, and then conduct a series of post-processing steps to extract the graph structure of road curbs. To tackle the disconnectivity issue in the segmentation maps, we propose an innovative connectivity-preserving loss (CP-loss) to improve the segmentation performance. The experimental results on a public dataset demonstrate the effectiveness of our proposed loss function. This paper is accompanied with a demonstration video and a supplementary document, which are available at https://sites.google.com/view/cp-loss. Zhenhua Xu 0003, Yuxiang Sun 0002, Lujia Wang 0001, Ming Liu 0001 |
IROS | 1 |
| 2021 | Visual Analysis of Discrimination in Machine LearningabstractThe growing use of automated decision-making in critical applications, such as crime prediction and college admission, has raised questions about fairness in machine learning. How can we decide whether different treatments are reasonable or discriminatory? In this paper, we investigate discrimination in machine learning from a visual analytics perspective and propose an interactive visualization tool, DiscriLens, to support a more comprehensive analysis. To reveal detailed information on algorithmic discrimination, DiscriLens identifies a collection of potentially discriminatory itemsets based on causal modeling and classification rules mining. By combining an extended Euler diagram with a matrix-based visualization, we develop a novel set visualization to facilitate the exploration and interpretation of discriminatory itemsets. A user study shows that users can interpret the visually encoded information in DiscriLens quickly and accurately. Use cases demonstrate that DiscriLens provides informative guidance in understanding and reducing algorithmic discrimination. Qianwen Wang 0001, Zhenhua Xu 0003, Chen Zhu-Tian, Yong Wang 0021, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 2 |