VLDB 2026 Research / reviewers in the wild / expert
Yupeng Zheng
dblp:339/6462
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous DrivingabstractLatent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented representation learning tangles perception with planning tasks, leading to suboptimal optimization for planning. To address this challenge, we propose WorldRFT, a planning-oriented latent world model framework that aligns scene representation learning with planning via a hierarchical planning decomposition and local-aware interactive refinement mechanism, augmented by reinforcement learning fine-tuning (RFT) to enhance safety-critical policy performance. Specifically, WorldRFT integrates a vision-geometry foundation model to improve 3D spatial awareness, employs hierarchical planning task decomposition to guide representation optimization, and utilizes local-aware iterative refinement to derive a planning-oriented driving policy. Furthermore, we introduce Group Relative Policy Optimization (GRPO), which applies trajectory Gaussianization and collision-aware rewards to fine-tune the driving policy, yielding systematic improvements in safety. WorldRFT achieves state-of-the-art (SOTA) performance on both open-loop nuScenes and closed-loop NavSim benchmarks. On nuScenes, it reduces collision rates by 83% (0.30% → 0.05%). On NavSim, using camera-only sensors input, it attains competitive performance with the LiDAR-based SOTA method DiffusionDrive (87.8 vs. 88.1 PDMS). Pengxuan Yang, Ben Lu, Zhongpu Xia, Yinfeng Gao, Kun Zhan, Xianpeng Lang, Yupeng Zheng |
AAAI | 9 |
| 2026 | SoAD: Safety-Oriented Value Estimation for Enhanced Closed-Loop End-to-End Autonomous DrivingabstractEnd-to-end (E2E) autonomous driving systems, which map sensory inputs directly to vehicle planning, have garnered attention for harnessing the potential of data-driven methodologies in motion planning. However, current methods face two limitations that undermine their safety performance in closed-loop driving tasks. First, the predominant imitation learning (IL) paradigm overlooks long-term safety beyond predefined planning horizons, potentially guiding the ego vehicle into hazardous states. Second, the lack of reliable online evaluation mechanisms limits real-time responses to safety risks. To overcome these challenges, we propose SoAD, a safety-oriented E2E framework that integrates long-term safety awareness into planning. SoAD is distinguished by a reinforcement learning (RL)-based value estimation module to quantify the safety of planned trajectories, and a vector world model (VWM) to generate interaction-aware future rollouts. During training, the system benefits from value-guided fine-tuning (VFT) that optimizes the planning distribution to favor safer trajectories. In closed-loop deployment, a planning rescoring (PRS) mechanism is designed to perform reliable online evaluation by combining ego-conditional predictions from the VWM with corresponding safety value estimates. Experimental results on the Bench2Drive closed-loop benchmark demonstrate the state-of-the-art (SoTA) performance of SoAD, achieving a 66.40% improvement in driving score (DS) compared to the vectorized scene representation for efficient autonomous driving (VAD) baseline, while zero-shot evaluation on the driving in occlusion simulation (DOS) benchmark further highlights its strong generalization ability. Yinfeng Gao, Deqing Liu, Yupeng Zheng, Dawei Ding 0001, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World ModelabstractEnd-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: constructing an informative driving world model to enable perception annotation-free, end-to-end planning via self-supervised learning. In this paper, we present World4Drive, an end-to-end autonomous driving framework that employs vision foundation models to build latent world models for generating and evaluating multi-modal planning trajectories. Specifically, World4Drive first extracts scene features, including driving intention and world latent representations enriched with spatial-semantic priors provided by vision foundation models. It then generates multi-modal planning trajectories based on current scene features and driving intentions and predicts multiple intention-driven future states within the latent space. Finally, it introduces a world model selector module to evaluate and select the best trajectory. We achieve perception annotation-free, end-to-end planning through self-supervised alignment between actual future observations and predicted observations reconstructed from the latent space. World4Drive achieves state-of-the-art performance without manual perception annotations on both the open-loop nuScenes and closed-loop NavSim benchmarks, demonstrating an 18.1\% relative reduction in L2 error, 46.7% lower collision rate, and 3.75 faster training convergence. Codes will be accessed at https://github.com/ucaszyp/World4Drive. Yupeng Zheng, Pengxuan Yang, Zebin Xing, Yuhang Zheng 0004, Yinfeng Gao, Pengfei Li 0007, Zhongpu Xia, Peng Jia 0007, Xianpeng Lang, Dongbin Zhao |
ICCV | 1 |
| 2025 | Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous DrivingabstractUnderstanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D occupancy labels are still required to produce promising results. Considering the high annotation cost for 3D outdoor scenes, we propose a semi-supervised vision-centric 3D occupancy world model, **PreWorld**, to leverage the potential of 2D labels through a novel two-stage training paradigm: the self-supervised pre-training stage and the fully-supervised fine-tuning stage. Specifically, during the pre-training stage, we utilize an attribute projection head to generate different attribute fields of a scene (e.g., RGB, density, semantic), thus enabling temporal supervision from 2D labels via volume rendering techniques. Furthermore, we introduce a simple yet effective state-conditioned forecasting module to recursively forecast future occupancy and ego trajectory in a direct manner. Extensive experiments on the nuScenes dataset validate the effectiveness and scalability of our method, and demonstrate that PreWorld achieves competitive performance across 3D occupancy prediction, 4D occupancy forecasting and motion planning tasks. Xiang Li 0205, Pengfei Li 0007, Yupeng Zheng |
ICLR | 3 |
| 2025 | UncAD: Towards Safe End-to-end Autonomous Driving via Online Map UncertaintyabstractEnd-to-end autonomous driving aims to produce planning trajectories from raw sensors directly. Currently, most approaches integrate perception, prediction, and planning modules into a fully differentiable network, promising great scalability. However, these methods typically rely on deterministic modeling of online maps in the perception module for guiding or constraining vehicle planning, which may incorporate erroneous perception information and further compromise planning safety. To address this issue, we delve into the importance of online map uncertainty for enhancing autonomous driving safety and propose a novel paradigm named UncAD. Specifically, UncAD first estimates the uncertainty of the online map in the perception module. It then leverages the uncertainty to guide motion prediction and planning modules to produce multi-modal trajectories. Finally, to achieve safer autonomous driving, UncAD proposes an uncertainty-collision-aware planning selection strategy according to the online map uncertainty to evaluate and select the best trajectory. In this study, we incorporate UncAD into various state-of-the-art (SOTA) end-to-end methods. Experiments on the nuScenes dataset show that integrating UncAD, with only a 1.9% increase in parameters, can reduce collision rates by up to 26% and drivable area conflict rate by up to 42%. Codes, pre-trained models, and demo videos can be accessed at https://github.com/pengxuanyang/UncAD. Pengxuan Yang, Yupeng Zheng, Kefei Zhu, Zebin Xing, Yun-Fu Liu, Zhiguo Su, Dongbin Zhao |
ICRA | 2 |
| 2024 | Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionabstractDepth completion is a vital taskfor autonomous driving, as it involves reconstructing the precise 3D geometry of a scene from sparse and noisy depth measurements. How-ever, most existing methods either rely only on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. To address this chal-lenge, we introduce Tri-Perspective View Decomposition (TPVD), a novel framework that can explicitly model 3D geometry. In particular, (1) TPVD ingeniously decomposes the original point cloud into three 2D views, one of which corresponds to the sparse depth input. (2) We design TPV Fusion to update the 2D TPV features through recurrent 2D-3D-2D aggregation, where a Distance-Aware Spherical Convolution (DASC) is applied. (3) By adaptively choosing TPVaffinitive neighbors, the newly proposed Geometric Spatial Propagation Network (GSPN) further improves the geometric consistency. As a result, our TPVD outperforms existing methods on KITTI, NYUv2, and SUN RGBD. Fur-thermore, we build a novel depth completion dataset named TOFDC, which is acquired by the time-of-flight (TOF) sen-sor and the color camera on smart phones. Project page. Zhiqiang Yan 0001, Yuankai Lin, Kun Wang 0042, Yupeng Zheng, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
CVPR | 4 |
| 2024 | TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Bu Jin, Yupeng Zheng, Pengfei Li 0007, Weize Li 0001, Yuhang Zheng 0004, Sujie Hu, Zhijie Yan, Kun Zhan, Peng Jia 0007, Xiaoxiao Long, Hao Zhao 0002 |
ECCV (18) | 2 |
| 2024 | MonoOcc: Digging into Monocular Semantic Occupancy PredictionabstractMonocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, particularly due to its potential to enhance the 3D perception of autonomous vehicles. However, existing methods rely on a complex cascaded framework with relatively limited information to restore 3D scenes, including a dependency on supervision solely on the whole network’s output, single-frame input, and the utilization of a small backbone. These challenges, in turn, hinder the optimization of the framework and yield inferior prediction results, particularly concerning smaller and long-tailed objects. To address these issues, we propose MonoOcc. In particular, we (i) improve the monocular occupancy prediction framework by proposing an auxiliary semantic loss as supervision to the shallow layers of the framework and an image-conditioned cross-attention module to refine voxel features with visual clues, and (ii) employ a distillation module that transfers temporal information and richer knowledge from a larger image backbone to the monocular semantic occupancy prediction framework with low cost of hardware. With these advantages, our method yields state-of-the-art performance on the camera-based SemanticKITTI Scene Completion benchmark. Codes and models can be accessed at https://github.com/ucaszyp/MonoOcc. Yupeng Zheng, Xiang Li 0205, Pengfei Li 0007, Yuhang Zheng 0004, Bu Jin, Chengliang Zhong, Xiaoxiao Long, Hao Zhao 0002 |
ICRA | 1 |
| 2024 | Flow-Audio-Synth: A Video-to-Audio Model which Captures Dynamic Features
Yupeng Zheng, Zixiang Lu, Qiguang Miao, Xiangzeng Liu |
PRCV (10) | 1 |
| 2024 | Adaptive Surface Normal Constraint for Geometric Estimation From Monocular ImagesabstractWe introduce a novel approach to learn geometries such as depth and surface normal from images while incorporating geometric context. The difficulty of reliably capturing geometric context in existing methods impedes their ability to accurately enforce the consistency between the different geometric properties, thereby leading to a bottleneck of geometric estimation quality. We therefore propose the Adaptive Surface Normal (ASN) constraint, a simple yet efficient method. Our approach extracts geometric context that encodes the geometric variations present in the input image and correlates depth estimation with geometric constraints. By dynamically determining reliable local geometry from randomly sampled candidates, we establish a surface normal constraint, where the validity of these candidates is evaluated using the geometric context. Furthermore, our normal estimation leverages the geometric context to prioritize regions that exhibit significant geometric variations, which makes the predicted normals accurately capture intricate and detailed geometric information. Through the integration of geometric context, our method unifies depth and surface normal estimations within a cohesive framework, which enables the generation of high-quality 3D geometry from images. We validate the superiority of our approach over state-of-the-art methods through extensive evaluations and comparisons on diverse indoor and outdoor datasets, showcasing its efficiency and robustness. Xiaoxiao Long, Yuhang Zheng 0004, Yupeng Zheng, Beiwen Tian, Cheng Lin 0001, Lingjie Liu, Hao Zhao 0002, Guyue Zhou, Wenping Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | DPF: Learning Dense Prediction Fields with Weak SupervisionabstractNowadays, many visual scene understanding problems are addressed by dense prediction networks. But pixel-wise dense annotations are very expensive (e.g., for scene parsing) or impossible (e.g., for intrinsic image decomposition), motivating us to leverage cheap point-level weak supervision. However, existing pointly-supervised methods still use the same architecture designed for full supervision. In stark contrast to them, we propose a new paradigm that makes predictions for point coordinate queries, as inspired by the recent success of implicit representations, like distance or radiance fields. As such, the method is named as dense prediction fields (DPFs). DPFs generate expressive intermediate features for continuous sub-pixel locations, thus allowing outputs of an arbitrary resolution. DPFs are naturally compatible with point-level supervision. We showcase the effectiveness of DPFs using two substantially different tasks: high-level semantic parsing and low-level intrinsic image decomposition. In these two cases, supervision comes in the form of single-point semantic category and two-point relative reflectance, respectively. As benchmarked by three large-scale public datasets PASCALContext, ADE20K and IIW, DPFs set new state-of-the-art performance on all of them with significant margins. Code can be accessed at https://github.com/cxx226/DPF. Xiaoxue Chen, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
CVPR | 3 |
| 2023 | 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryabstractKeypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method was introduced for 2D data, which reconstructs the target frame from the source frame to incorporate both spatial and temporal information. However, the direct application of the Transporter to 3D point clouds is infeasible due to their structural differences from 2D images. Thus, we propose the first 3D version of the Transporter, which leverages hybrid 3D representation, cross attention, and implicit reconstruction. We apply this new learning system on 3D articulated objects and non-rigid animals (humans and rodents) and show that learned keypoints are spatio-temporally consistent. Additionally, we propose a closed-loop control strategy that utilizes the learned keypoints for 3D object manipulation and demonstrate its superior performance. Codes are available at https://github.com/zhongcl-thu/3D-Implicit-Transporter. Chengliang Zhong, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Li Yi 0001, Xiaodong Mu, Ling Wang 0001, Pengfei Li 0007, Guyue Zhou, Chao Yang 0026, Jian Zhao 0006 |
ICCV | 3 |
| 2023 | ADAPT: Action-aware Driving Caption TransformerabstractEnd-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been some early attempts to use attention maps or cost volume for better model explainability which is difficult for ordinary passengers to understand. To bridge the gap, we propose an end-to-end transformer-based architecture, ADAPT (Action-aware Driving cAPtion Transformer), which provides user-friendly natural language narrations and reasoning for each decision making step of autonomous vehicular control and action. ADAPT jointly trains both the driving caption task and the vehicular control prediction task, through a shared video representation. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate state-of-the-art performance of the ADAPT framework on both automatic metrics and human evaluation. To illustrate the feasibility of the proposed framework in real-world applications, we build a novel deployable system that takes raw car videos as input and outputs the action narrations and reasoning in real time. The code, models and data are available at https://github.com/jxbbb/ADAPT. Bu Jin, Yupeng Zheng, Pengfei Li 0007, Hao Zhao 0002, Yuhang Zheng 0004, Guyue Zhou |
ICRA | 3 |
| 2023 | STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth EstimationabstractSelf-supervised depth estimation draws a lot of attention recently as it can promote the 3D sensing capa-bilities of self-driving vehicles. However, it intrinsically relies upon the photometric consistency assumption, which hardly holds during nighttime. Although various supervised night-time image enhancement methods have been proposed, their generalization performance in challenging driving scenarios is not satisfactory. To this end, we propose the first method that jointly learns a nighttime image enhancer and a depth estimator, without using ground truth for either task. Our method tightly entangles two self-supervised tasks using a newly proposed uncertain pixel masking strategy. This strategy originates from the observation that nighttime images not only suffer from underexposed regions but also from overexposed regions. By fitting a bridge-shaped curve to the illumination map distribution, both regions are suppressed and two tasks are bridged naturally. We benchmark the method on two established datasets: nuScenes and RobotCar and demonstrate state-of-the-art performance on both of them. Detailed ablations also reveal the mechanism of our proposal. Last but not least, to mitigate the problem of sparse ground truth of existing datasets, we provide a new photo-realistically enhanced nighttime dataset based upon CARLA. It brings meaningful new challenges to the community. Codes, data, and models are available at https://github.com/ucaszyp/STEPS. Yupeng Zheng, Chengliang Zhong, Pengfei Li 0007, Huan-ang Gao, Yuhang Zheng 0004, Bu Jin, Ling Wang 0001, Hao Zhao 0002, Guyue Zhou, Dongbin Zhao |
ICRA | 1 |