VLDB 2026 Research / reviewers in the wild / expert
Tuopu Wen
dblp:220/3900
· DBLP profile ↗
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-3093-765XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video GenerationabstractCorner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for generating diverse scenarios, commonly achieved through 3D Gaussian Splatting or image generative models. However, these approaches often suffer from limited visual fidelity or imprecise pose control. To address these issues, we propose G^2Editor, a framework designed for photorealistic and precise object editing in driving videos. Our method leverages a 3D Gaussian representation of the edited object as a dense prior, injected into the denoising process to ensure accurate pose control and spatial consistency. A scene-level 3D bounding box layout is employed to reconstruct occluded areas of non-target objects. Furthermore, to guide the appearance details of the edited object, we incorporate hierarchical fine-grained features as additional conditions during generation. Experiments on the Waymo Open Dataset demonstrate that G^2Editor effectively supports object repositioning, insertion, and deletion within a unified framework, outperforming existing methods in both pose controllability and visual quality, while also benefiting downstream data-driven tasks. Jiusi Li, Jackson Jiang, Jinyu Miao, Miao Long, Tuopu Wen, Peijin Jia, Shengxiang Liu, Chun-lei Yu, Maolin Liu, Yuzhan Cai, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 5 |
| 2026 | Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion PredictionabstractLiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2025 | Enhancing Autonomous Vehicle Planning With a Robust Fault-Tolerant Mechanism for Action-Induced Agent DetectionabstractIn autonomous driving, accurately identifying traffic participants that may influence vehicle behavior is crucial for effective system planning. To address this challenge, we propose a fault-tolerant mechanism for detecting action-induced objects, which significantly improves decision-making performance and system explainability. Since these objects are often linked to the vehicle’s driving intentions, we introduce a top-down attention network that adjusts attention weights for traffic participants based on navigational information. Additionally, we define potentially hazardous objects in the driving environment and employ supervised training with a classification head to detect them. To further enhance detection accuracy, we integrate a fault-tolerant process that merges attention maps with classification results, effectively reducing false positives and false negatives in identifying action-induced objects. Extensive testing validates the robustness and effectiveness of our approach, demonstrating its ability to improve both planning and interpretability in autonomous vehicles. Zheng Fu, Hezhe Lin, Kangan Qian, Tuopu Wen, Hao Gao 0005, Diange Yang |
ICASSP | 4 |
| 2025 | Efficient End-to-end Visual Localization for Autonomous Driving with Decoupled BEV Neural MatchingabstractAccurate localization plays an important role in high-level autonomous driving systems. Conventional map matching-based localization methods solve the poses by explicitly matching map elements with sensor observations, generally sensitive to perception noise, therefore requiring costly hyperparameter tuning. In this paper, we propose an end-to-end localization neural network which directly estimates vehicle poses from surrounding images, without explicitly matching perception results with HD maps. To ensure efficiency and interpretability, a decoupled BEV neural matching-based pose solver is proposed, which estimates poses in a differentiable sampling-based matching module. Moreover, the sampling space is hugely reduced by decoupling the feature representation affected by each DoF of poses. The experimental results demonstrate that the proposed network is capable of performing decimeter level localization with mean absolute errors of 0.19m, 0.13m and 0.39° in longitudinal, lateral position and yaw angle while exhibiting a 68.8% reduction in inference memory usage. Jinyu Miao, Tuopu Wen, Ziang Luo, Kangan Qian, Zheng Fu, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Jin Huang 0002, Diange Yang |
IROS | 2 |
| 2025 | EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Drivingabstract3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or transformers at the voxel level. Moreover, OccNets are supervised with expensive large-scale dense voxel labels. Model and label inefficiencies, caused by excessive network parameters and label annotation requirements, severely hinder the onboard deployment of OccNets. This paper proposes an EFFicient Occupancy learning framework, EFFOcc, that targets minimal network complexity and label requirements while achieving state-of-the-art accuracy. We first propose an efficient fusion-based OccNet that only uses simple 2D operators and improves accuracy to the state-of-the-art on three large-scale benchmarks: Occ3D-nuScenes, Occ3D-Waymo, and OpenOccupancy-nuScenes. On the Occ3D-nuScenes benchmark, the fusion-based model with ResNet-18 as the image backbone has 21.35M parameters and achieves 51.49 in terms of mean Intersection over Union (mIoU). Furthermore, we propose a multi-stage occupancy-oriented distillation to efficiently transfer knowledge to vision-only OccNet. Extensive experiments on occupancy benchmarks show state-of-the-art precision for both fusion-based and vision-based OccNets. For the demonstration of learning with limited labels, we achieve 94.38% of the performance (mIoU = 28.38) of a 100% labeled vision OccNet (mIoU = 30.07) using the same OccNet trained with only 40% labeled sequences and distillation from the fusion-based OccNet. Code is available at https://github.com/synsin0/EFFOcc. Yining Shi 0002, Kun Jiang 0002, Jinyu Miao, Ke Wang 0021, Kangan Qian, Yunlong Wang 0009, Jiusi Li, Tuopu Wen, Mengmeng Yang 0001, Yiliang Xu, Diange Yang |
IROS | 8 |
| 2025 | COME: Adding Scene-Centric Forecasting Control to Occupancy World ModelabstractWorld models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data.
Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to separate environmental changes from ego-motion by leveraging the scene-centric coordinate systems. In this paper, we introduce COME: a framework that integrates scene-centric forecasting Control into the Occupancy world ModEl. Specifically, COME first generates ego-irrelevant, spatially consistent future features through a scene-centric prediction branch, which are then converted into scene condition using a tailored ControlNet. These condition features are subsequently injected into the occupancy world model, enabling more accurate and controllable future occupancy predictions. Experimental results on the nuScenes-Occ3D dataset show that COME achieves consistent and significant improvements over state-of-the-art (SOTA) methods across diverse configurations, including different input sources (ground-truth, camera-based, fusion-based occupancy) and prediction horizons (3s and 8s). For example, under the same settings, COME achieves 26.3% better mIoU metric than DOME and 23.7% better mIoU metric than UniScene. These results highlight the efficacy of disentangled representation learning in enhancing spatio-temporal prediction fidelity for world models. Code is available at https://github.com/synsin0/COME. Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Tuopu Wen, Mengmeng Yang 0001, Diange Yang |
NeurIPS | 7 |
| 2025 | Top-Down Attention-Based Mechanisms for Interpretable Autonomous DrivingabstractDespite the remarkable advancements in autonomous driving, the challenge persists in achieving interpretable action decision-making, primarily owing to the intricate and ambiguous relationship between detected agents and driving intention. In this study, we introduce an interpretable action prediction model, denoted as the Prediction-Driven Attention Network (PDANet), designed to undertake action decisions and provide corresponding interpretations cohesively. The PDANet is inspired by the perceptual mechanisms inherent in human drivers, who allocate attention according to their driving intentions. Specifically, we elaborate a prediction module to generate vehicle prospective trajectories to characterize driving intentions. Subsequently, the features of this predicted trajectory are utilized to modulate the attention distribution among agents through the top-down attention module, yielding an attention map. Finally, two distinct task tokens are applied to aggregate agent features and generate the final output according to the derived attention map. Extensive experiments conducted on the publicly available BDD-OIA and nu-AR datasets demonstrate that our proposed method outperforms all prior works in terms of both action prediction and behavior interpretation tasks. Remarkably, our method attains a noteworthy enhancement in the behavior interpretation task, surpassing the previous state-of-the-art by a substantial margin of +10.8% in terms of F1-score on the nu-AR dataset. We also validate our algorithm on Carla Town05 long in a closed-loop decision-making scenario, highlighting the generality and robustness of our approach. Furthermore, qualitative results show that the agents selected by our model are more closely aligned with human cognitive processes. Zheng Fu, Kun Jiang 0002, Yunlong Wang 0009, Tuopu Wen, Hao Gao 0005, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Toward Democratizing High-Definition Map Update Through Consortium BlockchainabstractIn the rapidly evolving landscape of autonomous vehicles and advanced navigation systems, the accuracy of high-definition maps and real-time updating has become paramount. However, in this progression, the security of map data has not received adequate attention, although the accuracy of the data can be easily altered when the system is breached. Thus, this paper introduces a novel approach to democratizing the process of high-definition map updates by leveraging consortium blockchain technology specifically designed for Proof of Presence and Reputation (POP-R) to safeguard the update process. Our proposed system leverages the presence and reputation of vehicles through infrastructure nodes to enhance the accuracy and reliability of HD map updates. We created a trusted ecosystem for maintaining high-definition maps, marked by a superior safety score across three scenarios compared to the standard proof of reputation technique. Additionally, it demonstrates high efficiency, achieving 12,000 transactions per second (TPS) for data queries and more than 2,500 TPS for data writing in our blockchain network. This efficiency proved our prowess in the lightweight computational power required, suitable for decentralized and crowdsourced-based systems. Through our POP-R framework, we lay the foundation for a new decentralized approach to the evolution of high-definition maps in the era of autonomous mobility. Benny Wijaya, Mengmeng Yang 0001, Tuopu Wen, Kun Jiang 0002, Wei Zhang 0090, Yunlong Wang 0009, Zheng Fu, Xuewei Tang, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | LaneDAG: Automatic HD Map Topology Generator Based on Geometry and Attention Fusion MechanismabstractIn high-definition maps (HD maps), the road lane centerline and lane topology graph play essential roles in navigation, planning, and decision-making. Existing research focusing on extracting physical infrastructure, such as lane boundaries, has made significant progress. But lane centerline detection and topology reasoning still remains challenging due to the severe overlapping centerlines and complicated topology. To tackle these challenges, we introduce an automatic lane topology extraction method for HD maps, termed LaneDAG, which extracts vectorized centerlines and their topology from prebuilt lane lines and road boundaries in HD maps. It formulates centerline extraction as a set prediction problem and lane topology prediction as a directed acyclic graph (DAG) construction problem. A novel mechanism that fusing geometric and attention-based features in the DAG is proposed to model the topological relationship between centerlines. Experiments conducted on the Argoverse 2 dataset demonstrate the proposed method’s superior performance compared to existing methods, showcasing its capability to extract lane centerlines and topology in HD maps automatically. Peijin Jia, Tuopu Wen, Ziang Luo, Zheng Fu, Jiaqi Liao, Huixian Chen, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 2 |
| 2022 | Skeleton-based traffic command recognition at road intersections for intelligent vehicles
Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Tuopu Wen, Diange Yang |
Neurocomputing | 6 |
| 2022 | TM3Loc: Tightly-Coupled Monocular Map Matching for High Precision Vehicle LocalizationabstractVision-based map-matching with HD map for high precision vehicle localization has gained great attention for its low-cost and ease of deployment. However, its localization performance is still unsatisfactory in accuracy and robustness in numerous real applications due to the sparsity and noise of the perceived HD map landmarks. This article proposes the tightly-coupled monocular map-matching localization algorithm (TM3Loc) for monocular-based vehicle localization. TM3Loc introduces semantic chamfer matching (SCM) to model monocular map-matching problem and combines visual features with SCM in a tightly-coupled manner. By applying the sliding window-based optimization technique, the historical visual features and HD map constraints are also introduced, such that the vehicle poses are estimated with an abundance of visual features and multi-frame HD map landmark features, rather than with single-frame HD map observations in previous works. Experiments are conducted on large scale dataset of 15 km long in total. The results show that TM3Loc is able to achieve high precision localization performance using a low-cost monocular camera, largely exceeding the performance of the previous state-of-the-art methods, thereby promoting the development of autonomous driving. Tuopu Wen, Kun Jiang 0002, Benny Wijaya, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Fast Initialization for Monocular Map Matching Localization via Multi-lane Hypotheses in Highway ScenariosabstractMany researchers have used the map-matching algorithm to leverage inadequate traditional vehicle localization with HD maps for fast and efficient vehicle localization. The initialization process in the map-matching algorithm has always been problematic due to the nature of GNSS error, which might exceed the average lane width of approximately 3.5m. Thus, directly using the GNSS data to obtain the rough initial pose may cause the failure of the initialization process. The general solution is to randomly sample around the GNSS data and test the initialization with these hypotheses. However, this often leads to expensive computation and thus fails to run in realtime or online mode, as a dense sampling is required to achieve an acceptable level of initial estimation accuracy. As a viable alternative, we propose a multi-lane hypotheses approach to narrow down the search by limiting the sample pose to the number of lanes within the GNSS data's error radius. From these poses, we then perform an efficient map-matching to refine these pose candidates. Moreover, a novel belief function to evaluate the hypothesis is proposed to select the best hypothesis for system initialization robustly. Our evaluation result shows that we have outperformed the primary random sampling method in both accuracy and efficiency. Tuopu Wen, Benny Wijaya, Kun Jiang 0002, Dongfang Zheng, Yiliang Xu, Mengmeng Yang 0001, Diange Yang |
IV | 1 |
| 2021 | Bridging the Gap of Lane Detection Performance Between Different Datasets: Unified Viewpoint TransformationabstractConvolutional neural networks (CNNs) have shown excellent performance for vision-based lane detection. However, maintaining the performance of the trained models under new test scenarios still remains challenging due to the dataset bias between the training and test datasets; In lane detection processes, the dataset bias can be categorized into lane position bias and lane pattern bias, with the former one particularly influences the lane detection performance. To tackle this dataset bias, this article proposes aunified viewpoint transformation (UVT)method that transforms the camera viewpoints of different datasets into a common virtual world coordinate system, such that the mismatched lane position distributions can be effectively aligned. Experiments are conducted on multiple datasets including the Caltech[1], Tusimple[2], and KITTI[3]dataset. The results demonstrate the effectiveness of the UVT algorithm in improving the lane detection performance on the test datasets. Moreover, by incorporating the UVT into other techniques that tackling the dataset bias, the lane position and pattern differences are disentangled and separately minimized. As a result, the performance gap between the training data and the test scenarios can be bridged. Specifically, the model trained on the KITTI dataset have achieved high performance in the Tusimple and the Caltech dataset (F1-score: 84.8 and 87.1%). With the proposed algorithm, a lane detection model trained on one dataset can be effectively applied to datasets with different camera settings in vastly different localities, and achieve better generalization ability compared to the state of the art methods. Tuopu Wen, Diange Yang, Kun Jiang 0002, Chun-lei Yu, Benny Wijaya, Xinyu Jiao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | High Precision Vehicle Localization based on Tightly-coupled Visual Odometry and Vector HD MapabstractMatching low-cost camera and vector HD map is proven to be a practical and effective way of estimating the location and orientation of intelligent vehicles. However, map-based approach is viable only when the landmark observation is adequate and precise. In some areas with sparse and noisy observation, or even non-existent map matching features, the localization results may be unstable. In this paper, we introduce a novel algorithm by fusing visual odometry and vector HD map in a tightly-coupled optimization framework to tackle these problems. Our algorithm exploits the observation of visual feature points and vector HD map landmarks in the sliding window manner and optimize their residuals in a tightly-coupled approach. In this way, the system is more robust against the noisy HD map landmark observations. In addition, our method is able to accurately estimate vehicle pose even when landmarks are sparse. Experiments under two challenging scenarios with noisy and sparse landmark observations show that our method can achieve the Mean Absolute Error (MAE) at 0.1473m and 0.2496m respectively. Tuopu Wen, Zhongyang Xiao, Benny Wijaya, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 1 |
| 2019 | High Precision Target Positioning Method for RSU in Cooperative PerceptionabstractVehicle-road cooperative perception system can greatly improve the perception ability of intelligent vehicles by making use of perception information from road side units (RSU). This paper focuses on the target positioning of static camera for vehicle-road cooperation. A low-cost camera calibration method is proposed to complete the accurate mapping between the image plane and the 3D world space. Precise location of interested targets are achieved by an efficient tracking strategy. Real test scenarios show that our algorithm can effectively locate vehicles, pedestrians, non-motor vehicles and other targets with high accuracy. Our algorithm won the Monocular Static Camera Positioning and Ranging Competition for Autonomous Driving championship in 2019. Tuopu Wen, Zhongyang Xiao, Kun Jiang 0002, Mengmeng Yang 0001, Keqiang Li 0002, Diange Yang |
MMSP | 1 |