EDBT 2026 Demo / reviewers in the wild / expert
Mengmeng Yang 0001
dblp:121/1326-1
· DBLP profile ↗
25ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0002-3294-6437ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video GenerationabstractCorner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for generating diverse scenarios, commonly achieved through 3D Gaussian Splatting or image generative models. However, these approaches often suffer from limited visual fidelity or imprecise pose control. To address these issues, we propose G^2Editor, a framework designed for photorealistic and precise object editing in driving videos. Our method leverages a 3D Gaussian representation of the edited object as a dense prior, injected into the denoising process to ensure accurate pose control and spatial consistency. A scene-level 3D bounding box layout is employed to reconstruct occluded areas of non-target objects. Furthermore, to guide the appearance details of the edited object, we incorporate hierarchical fine-grained features as additional conditions during generation. Experiments on the Waymo Open Dataset demonstrate that G^2Editor effectively supports object repositioning, insertion, and deletion within a unified framework, outperforming existing methods in both pose controllability and visual quality, while also benefiting downstream data-driven tasks. Jiusi Li, Jackson Jiang, Jinyu Miao, Miao Long, Tuopu Wen, Peijin Jia, Shengxiang Liu, Chun-lei Yu, Maolin Liu, Yuzhan Cai, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 12 |
| 2026 | MaPLocator: Map-Prior Enhanced End-to-End Vehicle Localization with Coarse-to-Fine Pose Refinement
Muxi Tang, Jinyu Miao, Rujun Yan, Le Jia, Mengmeng Yang 0001, Diange Yang, Kun Jiang 0002 |
IV | 6 |
| 2026 | Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion PredictionabstractLiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Efficient End-to-end Visual Localization for Autonomous Driving with Decoupled BEV Neural MatchingabstractAccurate localization plays an important role in high-level autonomous driving systems. Conventional map matching-based localization methods solve the poses by explicitly matching map elements with sensor observations, generally sensitive to perception noise, therefore requiring costly hyperparameter tuning. In this paper, we propose an end-to-end localization neural network which directly estimates vehicle poses from surrounding images, without explicitly matching perception results with HD maps. To ensure efficiency and interpretability, a decoupled BEV neural matching-based pose solver is proposed, which estimates poses in a differentiable sampling-based matching module. Moreover, the sampling space is hugely reduced by decoupling the feature representation affected by each DoF of poses. The experimental results demonstrate that the proposed network is capable of performing decimeter level localization with mean absolute errors of 0.19m, 0.13m and 0.39° in longitudinal, lateral position and yaw angle while exhibiting a 68.8% reduction in inference memory usage. Jinyu Miao, Tuopu Wen, Ziang Luo, Kangan Qian, Zheng Fu, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Jin Huang 0002, Diange Yang |
IROS | 8 |
| 2025 | LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion PredictionabstractAccurate spatial and motion understanding is critical for autonomous driving systems. While object-level perception models excel in structured environments, they struggle with open-set categories and often lack precise geometric representation. Occupancy-based, class-agnostic methods offer better scene expressiveness but typically ignore inter-agent interactions and fail to ensure physical consistency in motion predictions, limiting their reliability in complex traffic scenarios. In this paper, we propose LEGO-Motion, a novel class-agnostic motion prediction framework that bridges the gap between instance-level reasoning and occupancy-based modeling. Unlike conventional grid-based methods that treat each cell independently, LEGO-Motion introduces two key components: (1) the Interaction-Augmented Instance Encoder (IaIE), which models interactions among dynamic agents via cross-attention, and (2) the Instance-Enhanced BEV Encoder (IeBE), which improves motion consistency across instances through multi-stage feature fusion. These components enable our model to learn semantically coherent and physically plausible motion fields. Extensive experiments on the nuScenes dataset show that LEGO-Motion achieves a around 6% improvement in motion prediction accuracy over the previous state-of-the-art, while maintaining real-time inference at 21ms. Moreover, our method demonstrates strong generalization on a proprietary FMCW LiDAR benchmark. These results validate LEGO-Motion's effectiveness in capturing both global scene structure and fine-grained motion dynamics, making it a promising foundation for next-generation perception systems. Kangan Qian, Jinyu Miao, Ziang Luo, Zheng Fu, Jinchen Li, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IROS | 9 |
| 2025 | EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Drivingabstract3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or transformers at the voxel level. Moreover, OccNets are supervised with expensive large-scale dense voxel labels. Model and label inefficiencies, caused by excessive network parameters and label annotation requirements, severely hinder the onboard deployment of OccNets. This paper proposes an EFFicient Occupancy learning framework, EFFOcc, that targets minimal network complexity and label requirements while achieving state-of-the-art accuracy. We first propose an efficient fusion-based OccNet that only uses simple 2D operators and improves accuracy to the state-of-the-art on three large-scale benchmarks: Occ3D-nuScenes, Occ3D-Waymo, and OpenOccupancy-nuScenes. On the Occ3D-nuScenes benchmark, the fusion-based model with ResNet-18 as the image backbone has 21.35M parameters and achieves 51.49 in terms of mean Intersection over Union (mIoU). Furthermore, we propose a multi-stage occupancy-oriented distillation to efficiently transfer knowledge to vision-only OccNet. Extensive experiments on occupancy benchmarks show state-of-the-art precision for both fusion-based and vision-based OccNets. For the demonstration of learning with limited labels, we achieve 94.38% of the performance (mIoU = 28.38) of a 100% labeled vision OccNet (mIoU = 30.07) using the same OccNet trained with only 40% labeled sequences and distillation from the fusion-based OccNet. Code is available at https://github.com/synsin0/EFFOcc. Yining Shi 0002, Kun Jiang 0002, Jinyu Miao, Ke Wang 0021, Kangan Qian, Yunlong Wang 0009, Jiusi Li, Tuopu Wen, Mengmeng Yang 0001, Yiliang Xu, Diange Yang |
IROS | 9 |
| 2025 | LDMapNet-U: An End-to-End System for City-Scale Lane-Level Map UpdatingabstractAn up-to-date city-scale lane-level map is an indispensable infrastructure and a key enabling technology for ensuring the safety and user experience of autonomous driving systems. In industrial scenarios, reliance on manual annotation for map updates creates a critical bottleneck. Lane-level updates require precise change information and must ensure consistency with adjacent data while adhering to strict standards. Traditional methods utilize a three-stage approach -- construction, change detection, and updating -- which often necessitates manual verification due to accuracy limitations. This results in labor-intensive processes and hampers timely updates. To address these challenges, we propose LDMapNet-U, which implements a new end-to-end paradigm for city-scale lane-level map updating. By reconceptualizing the update task as an end-to-end map generation process grounded in historical map data, we introduce a paradigm shift in map updating that simultaneously generates vectorized maps and change information. To achieve this, a Prior-Map Encoding (PME) module is introduced to effectively encode historical maps, serving as a critical reference for detecting changes. Additionally, we incorporate a novel Instance Change Prediction (ICP) module that learns to predict associations with historical maps. Consequently, LDMapNet-U simultaneously achieves vectorized map element generation and change detection. To demonstrate the superiority and effectiveness of LDMapNet-U, extensive experiments are conducted using large-scale real-world datasets. In addition, LDMapNet-U has been successfully deployed in production at Baidu Maps since April 2024, supporting lane-level map updating for over 360 cities and significantly shortening the update cycle from quarterly to weekly, thereby enhancing the timeliness and accuracy of lane-level map. The nationwide, high-frequency city-scale lane-level map has been instrumental in the development of the lane-level navigation product serving hundreds of millions of users, while also integrating into the autonomous driving systems of several leading vehicle companies. Deguo Xia, Weiming Zhang 0006, Xiyan Liu, Wei Zhang 0088, Chenting Gong, Xiao Tan 0001, Jizhou Huang, Mengmeng Yang 0001, Diange Yang |
KDD (1) | 8 |
| 2025 | COME: Adding Scene-Centric Forecasting Control to Occupancy World ModelabstractWorld models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data.
Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to separate environmental changes from ego-motion by leveraging the scene-centric coordinate systems. In this paper, we introduce COME: a framework that integrates scene-centric forecasting Control into the Occupancy world ModEl. Specifically, COME first generates ego-irrelevant, spatially consistent future features through a scene-centric prediction branch, which are then converted into scene condition using a tailored ControlNet. These condition features are subsequently injected into the occupancy world model, enabling more accurate and controllable future occupancy predictions. Experimental results on the nuScenes-Occ3D dataset show that COME achieves consistent and significant improvements over state-of-the-art (SOTA) methods across diverse configurations, including different input sources (ground-truth, camera-based, fusion-based occupancy) and prediction horizons (3s and 8s). For example, under the same settings, COME achieves 26.3% better mIoU metric than DOME and 23.7% better mIoU metric than UniScene. These results highlight the efficacy of disentangled representation learning in enhancing spatio-temporal prediction fidelity for world models. Code is available at https://github.com/synsin0/COME. Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Tuopu Wen, Mengmeng Yang 0001, Diange Yang |
NeurIPS | 8 |
| 2025 | Pedestrian Trajectory Prediction for Autonomous Vehicles With Multiple InteractionsabstractPedestrian trajectory prediction is significant for autonomous vehicles, but the difficulty of pedestrian trajectory prediction lies in the accurate modeling of pedestrian multiple interactions. In this paper, we attempt to explore the essential features of pedestrian interaction and propose a pedestrian trajectory prediction method based on multiple interactions. Firstly, considering that the interaction between self-driving cars and pedestrians resembles a dynamic game process involving sequential adaptation, we map them to the same feature space and design a temporal cross-attention mechanism to model the interaction between pedestrians and vehicles. Meanwhile, pedestrian-scene interaction is affected by the global environment as well as the local environment. To capture the global information while preserving the spatial location of pedestrians in the scene, we design a pedestrian-scene heatmap fusion (PSHF) framework to model the pedestrian-scene interaction features. We validate the effectiveness of our algorithm on the publicly available JAAD and PIE datasets, achieving better performance than existing representative methods in both single-trajectory and multi-trajectory prediction tasks. We conducted a thorough ablation study, cross-dataset validation, and qualitative visualization experiments, demonstrating the effectiveness and robustness of our method. Zheng Fu, Mengmeng Yang 0001, Kun Jiang 0002, Jin Huang 0002, Hao Gao 0005, Diange Yang |
IEEE Internet Things J. | 2 |
| 2025 | Lo-SLAM: Lunar Target-Oriented SLAM Using Object Identification, Relative Navigation, and Multilevel MappingabstractTo ensure long-term space missions, an autonomous localization and mapping system for lunar rovers is demanded. While the target-oriented localization and mapping problem can be solved through state-of-the-art methods, they greatly rely on human-in-loop remote operations, posing several challenges for the visual system of a rover when operating in a distant, unknown, and feature-sparse lunar environment. This article presents a segment anything model (SAM)-augmented target-oriented simultaneous localization and mapping (SLAM) framework that enables rovers to estimate the relative distance to the target on the lunar surface, thus ensuring the safety of the exploration task. Based on the proposed point-prompted object instance extraction (OIE) pipeline, object correspondences are first predicted in the middle-end of Lo-SLAM, where reliable semantic constraints are robustly associated cross image frames. We then maintain camera-object relative positioning between the camera and target in the visual odometry of Lo-SLAM. Meanwhile, a multilevel mapping and representation framework is proposed to keep the target explicit and characterized in different subtasks. Extensive experiments are conducted on our dataset, stereo planetary tracks (SePTs). Results show that the proposed Lo-SLAM is validated on challenging lunar scenarios with dramatic viewpoints and object scale changes. The average pose errors are 0.36 m in centroid and 0.27 m in scale, and the average object-centric trajectory error is 0.49% or so. An open-source dataset has been released athttps://github.com/miaTian99/SePT_Stereo-Planetary-Tracks. Yaolin Tian, Xue Wan, Shengyang Zhang, Jianhong Zuo, Yadong Shao, Baichuan Liu, Mengmeng Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | RM2Occ: Re-Projection Multi-Task Multi-Sensor Fusion for Autonomous Driving 3D Object Detection and Occupancy PerceptionabstractOccupancy prediction plays a crucial role in supporting autonomous driving planning and decision-making. Existing methods typically rely on modular stacking and fusion techniques of object detection, semantic segmentation, and depth estimation to achieve 3D occupancy. However, they fail to deeply explore the transformation relationships between 2D and 3D spaces and to efficiently fuse the different characteristics of multi-source sensors. We propose R$M^{2}$Occ, the first 3D occupancy perception network that integrates multi-sensor fusion based on different sensor principles and achieves multi-task learning. To leverage the rich 2D semantic information captured by cameras and elevate it to the 3D domain, we begin by querying and populating predefined empty voxels with multi-view image features. Subsequently, we progressively fuse 3D LiDAR point clouds with these populated voxels through an unbalanced fusion strategy that effectively supplements missing information and suppresses noise. Leveraging IMU data and calibration parameters, we then re-project the enriched voxels back onto the 2D image plane according to camera coordinates, performing a secondary query using the semantic segmentation results to recover semantic details potentially lost due to radar fusion limitations and incomplete voxel querying. Finally, supported by a multi-task detection head, R$M^{2}$Occ simultaneously accomplishes 3D object detection, semantic segmentation, Bird’s Eye View (BEV) detection, and full-scene grid occupancy prediction, enabling comprehensive multi-task output. Extensive experiments and ablation studies on the nuScenes dataset demonstrate that R$M^{2}$Occ significantly outperforms existing state-of-the-art methods, establishing a new paradigm for accurate and efficient multi-sensor fusion and multi-task perception in autonomous driving scenarios. Yilong Ren, Minda Li, Han Jiang 0003, Zhiyong Cui, Mengmeng Yang 0001, Haiyang Yu 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Toward Democratizing High-Definition Map Update Through Consortium BlockchainabstractIn the rapidly evolving landscape of autonomous vehicles and advanced navigation systems, the accuracy of high-definition maps and real-time updating has become paramount. However, in this progression, the security of map data has not received adequate attention, although the accuracy of the data can be easily altered when the system is breached. Thus, this paper introduces a novel approach to democratizing the process of high-definition map updates by leveraging consortium blockchain technology specifically designed for Proof of Presence and Reputation (POP-R) to safeguard the update process. Our proposed system leverages the presence and reputation of vehicles through infrastructure nodes to enhance the accuracy and reliability of HD map updates. We created a trusted ecosystem for maintaining high-definition maps, marked by a superior safety score across three scenarios compared to the standard proof of reputation technique. Additionally, it demonstrates high efficiency, achieving 12,000 transactions per second (TPS) for data queries and more than 2,500 TPS for data writing in our blockchain network. This efficiency proved our prowess in the lightweight computational power required, suitable for decentralized and crowdsourced-based systems. Through our POP-R framework, we lay the foundation for a new decentralized approach to the evolution of high-definition maps in the era of autonomous mobility. Benny Wijaya, Mengmeng Yang 0001, Tuopu Wen, Kun Jiang 0002, Wei Zhang 0090, Yunlong Wang 0009, Zheng Fu, Xuewei Tang, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Grid-Centric Traffic Scenario Perception for Autonomous Driving: A Comprehensive ReviewabstractThe grid-centric perception is a crucial field for mobile robot perception and navigation. Nonetheless, the grid-centric perception is less prevalent than object-centric perception as autonomous vehicles need to accurately perceive highly dynamic, large-scale traffic scenarios, and the complexity and computational costs of grid-centric perception are high. In recent years, the rapid development of deep learning techniques and hardware provides fresh insights into the evolution of grid-centric perception. The fundamental difference between grid-centric and object-centric pipeline lies in that grid-centric perception follows a geometry-first paradigm which is more robust to the open-world driving scenarios with endless long-tailed semantically unknown obstacles. Recent research demonstrates the great advantages of grid-centric perception, such as comprehensive fine-grained environmental representation, greater robustness to occlusion and irregular-shaped objects, better ground estimation, and safer planning policies. There is also a growing trend that the capacity of occupancy networks is greatly expanded to 4-D scene perception and prediction, and the latest techniques are highly related to new research topics, such as 4-D occupancy forecasting, generative artificial intelligence (GenAI), and world models in the field of autonomous driving. Given the lack of current surveys for this rapidly expanding field, we present a hierarchically structured review of grid-centric perception for autonomous vehicles. We organize previous and current knowledge of occupancy grid techniques along the main vein from 2-D bird-eye view (BEV) grids to 3-D occupancy to 4-D occupancy forecasting. We additionally summarize label-efficient occupancy learning and the role of grid-centric perception in driving systems. Finally, we present a summary of the current research trend and provide future outlooks. Yining Shi 0002, Kun Jiang 0002, Jiusi Li, Zelin Qian, Junze Wen, Mengmeng Yang 0001, Ke Wang 0021, Diange Yang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous DrivingabstractVision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect obstacles regardless of their shape and occlusion. Modern occupancy networks mainly focus on reconstructing visible voxels from object surfaces with voxel-wise semantic prediction. Usually, they suffer from inconsistent predictions of one object and mixed predictions for adjacent objects. These confusions may harm the safety of downstream planning modules. To this end, we investigate panoptic segmentation on 3D voxel scenarios and propose an instance-aware occupancy network, PanoSSC. We predict foreground objects and backgrounds separately and merge both in post-processing. For foreground instance grouping, we propose a novel 3D instance mask decoder that can efficiently extract individual objects. we unify geometric reconstruction, 3D semantic segmentation, and 3D instance segmentation into PanoSSC framework and propose new metrics for evaluating panoptic voxels. Extensive experiments show that our method achieves competitive results on SemanticKITTI semantic scene completion benchmark. Yining Shi 0002, Jiusi Li, Kun Jiang 0002, Ke Wang 0021, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang |
3DV | 6 |
| 2024 | StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential EquationabstractPredicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However, current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict uniform snapshots of future occupancy states and require strictly synchronized sensory data for sensor fusion. We propose a novel framework, StreamingFlow, to lift these strong limitations. StreamingFlow is a novel BEV occupancy predictor that ingests asynchronous multi-sensor data streams for fusion and performs streaming fore-casting of the future occupancy map at any future times-tamps. By integrating neural ordinary differential equations (N-ODE) into recurrent neural networks, StreamingFlow learns derivatives of BEV features over temporal horizons, updates the implicit sensor's BEV features as part of the fusion process, and propagates BEV states to the desired future time point. It shows good zero-shot generalization ability of prediction, reflected in the interpolation of the ob-served prediction time horizon and the reasonable inference of the unseen farther future period. Extensive experiments on two large-scale datasets, nuScenes [2] and Lyft L5 [14], demonstrate that StreamingFlow significantly outperforms previous vision-based, LiDAR-based methods, and shows superior performance compared to state-of-the-art fusion-based methods. Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Jiusi Li, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang |
CVPR | 6 |
| 2024 | LaneDAG: Automatic HD Map Topology Generator Based on Geometry and Attention Fusion MechanismabstractIn high-definition maps (HD maps), the road lane centerline and lane topology graph play essential roles in navigation, planning, and decision-making. Existing research focusing on extracting physical infrastructure, such as lane boundaries, has made significant progress. But lane centerline detection and topology reasoning still remains challenging due to the severe overlapping centerlines and complicated topology. To tackle these challenges, we introduce an automatic lane topology extraction method for HD maps, termed LaneDAG, which extracts vectorized centerlines and their topology from prebuilt lane lines and road boundaries in HD maps. It formulates centerline extraction as a set prediction problem and lane topology prediction as a directed acyclic graph (DAG) construction problem. A novel mechanism that fusing geometric and attention-based features in the DAG is proposed to model the topological relationship between centerlines. Experiments conducted on the Argoverse 2 dataset demonstrate the proposed method’s superior performance compared to existing methods, showcasing its capability to extract lane centerlines and topology in HD maps automatically. Peijin Jia, Tuopu Wen, Ziang Luo, Zheng Fu, Jiaqi Liao, Huixian Chen, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 8 |
| 2024 | DuMapNet: An End-to-End Vectorization System for City-Scale Lane-Level Map GenerationabstractGenerating city-scale lane-level maps faces significant challenges due to the intricate urban environments, such as blurred or absent lane markings. Additionally, a standard lane-level map requires a comprehensive organization of lane groupings, encompassing lane direction, style, boundary, and topology, yet has not been thoroughly examined in prior research. These obstacles result in labor-intensive human annotation and high maintenance costs. This paper overcomes these limitations and presents an industrial-grade solution named DuMapNet that outputs standardized, vectorized map elements and their topology in an end-to-end paradigm. To this end, we propose a group-wise lane prediction (GLP) system that outputs vectorized results of lane groups by meticulously tailoring a transformer-based network. Meanwhile, to enhance generalization in challenging scenarios, such as road wear and occlusions, as well as to improve global consistency, a contextual prompts encoder (CPE) module is proposed, which leverages the predicted results of spatial neighborhoods as contextual information. Extensive experiments conducted on large-scale real-world datasets demonstrate the superiority and effectiveness of DuMapNet. Additionally, DuMapNet has already been deployed in production at Baidu Maps since June 2023, supporting lane-level map generation tasks for over 360 cities while bringing a 95% reduction in costs. This demonstrates that DuMapNet serves as a practical and cost-effective industrial solution for city-scale lane-level map generation. Deguo Xia, Weiming Zhang 0006, Xiyan Liu, Wei Zhang 0114, Chenting Gong, Jizhou Huang, Mengmeng Yang 0001, Diange Yang |
KDD | 7 |
| 2023 | Traffic Police 3D Gesture Recognition Based on Spatial-Temporal Fully Adaptive Graph Convolutional NetworkabstractIt is critical for autonomous vehicles to recognize traffic police gestures timely and accurately. During the movement of the vehicle, the collected traffic police scales change all the time, in addition, the frequency and amplitude of actions of different traffic police are different. First, we use gesture normalization to fix the traffic police actions at a unified scale and remove the influence of scale changes on traffic police gesture recognition. Meanwhile, a fully adaptive spatial-temporal graph convolution network (FA-STGCN) is proposed to recognize the actions with different amplitude and frequencies. The adaptive spatial graph network can dig the latent joints connection relation of the traffic police under different gestures, which weakens the amplitude impact on the action recognition. The adaptive temporal graph network is composed of the global temporal module and the local temporal module. The global temporal module can obtain the coarse-grained features of the traffic police gestures’ speed and then naturally use the coarse-grained features to guide the local temporal module to adaptively learn the fine-grained temporal features of the traffic police action. The adaptive spatial graph network and the temporal graph network are alternately stacked to finally output accurate traffic police gestures. We thoroughly evaluated our method through intensive experiments, the result shows that our method achieved the best results on public datasets. What’s more, we proofed the effectiveness of each module and verified our methods for moving vehicles for the first time, the performance present meets the vehicle’s practical requirements. Zheng Fu, Kun Jiang 0002, Junze Wen, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Reliable Autonomous Driving Environment Model With Unified State-Extended BoundaryabstractFrom the early stage of robotic applications to current autonomous driving technologies, environment modeling has been acting as the middleware for connecting perception and decision layers. In robotic applications, space-oriented models (e.g., grid map, drivable area) are widely applied to faithfully reflect the space occupation. With the development of autonomous driving, highly dynamic and complex road environment brings rising need to understand the type and motion status of objects, thus element list has became the mainstream environment model. However, along comes the reliablity problem caused by missed detection and irregular objects, which is still inevitable despite the detection accuracy improvement. In view of this, a new view of driving environment is proposed as the unified state-extended boundary (USEB), aiming to improve the reliablity of element-oriented model. For driving decision requirements, different types of elements are consistently converted into driving constraints. Semantics and dynamics are expressed as the status of drivable area boundary, making it possible to merge space occupation to improve reliability against missed detection and irregular objects. Evaluation of USEB is carried out on the nuScenes dataset. Comparative results show that the proposed USEB could cover the required information for driving decision, whereas achieving higher reliability than the commonly applied element-oriented model. Xinyu Jiao, Kun Jiang 0002, Yunlong Wang 0009, Zhong Cao 0003, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Skeleton-based traffic command recognition at road intersections for intelligent vehicles
Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Tuopu Wen, Diange Yang |
Neurocomputing | 4 |
| 2022 | Simple But Effective: Upper-Body Geometric Features for Traffic Command Gesture RecognitionabstractRecognizing traffic command gestures with high accuracy and quick response at a low computational cost is a requisite for driver assistance or autonomous driving. However, it has been understudied for a long time. Existing research takes advantage of increasing development in human action recognition but pays little attention to onboard conditions. In this article, we propose a simple but effective recognition model based on human upper-body geometric features and a long short-term memory (LSTM) network. The handcrafted geometric features can easily be calculated with estimated 2-D human keypoints at a low computational cost but are discriminative and sufficient in classification. Offline and online inferences are implemented to comprehensively evaluate the proposed model. For the sake of robustness required in the automotive domain, dual voting is designed to filter the output in online inference. On the recently published Chinese traffic police gesture (CTPG) dataset, the presented approach is the best with a remarkable improvement of approximately 8% compared to previous LSTM-based methods with handcrafted spatial features and is competitive with advanced GCN-based deep learning methods. The tradeoff pattern is explored to demonstrate how accuracy and response time alter with different training and inference strategies so that a balanced setup can be manually chosen under various application scenarios. Field tests are also carried out with an experimental vehicle, and the results uncover the present gap between research and practical application to some extent, moving a step closer to real-life traffic command gesture recognition. Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Diange Yang |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2022 | TM3Loc: Tightly-Coupled Monocular Map Matching for High Precision Vehicle LocalizationabstractVision-based map-matching with HD map for high precision vehicle localization has gained great attention for its low-cost and ease of deployment. However, its localization performance is still unsatisfactory in accuracy and robustness in numerous real applications due to the sparsity and noise of the perceived HD map landmarks. This article proposes the tightly-coupled monocular map-matching localization algorithm (TM3Loc) for monocular-based vehicle localization. TM3Loc introduces semantic chamfer matching (SCM) to model monocular map-matching problem and combines visual features with SCM in a tightly-coupled manner. By applying the sliding window-based optimization technique, the historical visual features and HD map constraints are also introduced, such that the vehicle poses are estimated with an abundance of visual features and multi-frame HD map landmark features, rather than with single-frame HD map observations in previous works. Experiments are conducted on large scale dataset of 15 km long in total. The results show that TM3Loc is able to achieve high precision localization performance using a low-cost monocular camera, largely exceeding the performance of the previous state-of-the-art methods, thereby promoting the development of autonomous driving. Tuopu Wen, Kun Jiang 0002, Benny Wijaya, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Fast Initialization for Monocular Map Matching Localization via Multi-lane Hypotheses in Highway ScenariosabstractMany researchers have used the map-matching algorithm to leverage inadequate traditional vehicle localization with HD maps for fast and efficient vehicle localization. The initialization process in the map-matching algorithm has always been problematic due to the nature of GNSS error, which might exceed the average lane width of approximately 3.5m. Thus, directly using the GNSS data to obtain the rough initial pose may cause the failure of the initialization process. The general solution is to randomly sample around the GNSS data and test the initialization with these hypotheses. However, this often leads to expensive computation and thus fails to run in realtime or online mode, as a dense sampling is required to achieve an acceptable level of initial estimation accuracy. As a viable alternative, we propose a multi-lane hypotheses approach to narrow down the search by limiting the sample pose to the number of lanes within the GNSS data's error radius. From these poses, we then perform an efficient map-matching to refine these pose candidates. Moreover, a novel belief function to evaluate the hypothesis is proposed to select the best hypothesis for system initialization robustly. Our evaluation result shows that we have outperformed the primary random sampling method in both accuracy and efficiency. Tuopu Wen, Benny Wijaya, Kun Jiang 0002, Dongfang Zheng, Yiliang Xu, Mengmeng Yang 0001, Diange Yang |
IV | 7 |
| 2020 | High Precision Vehicle Localization based on Tightly-coupled Visual Odometry and Vector HD MapabstractMatching low-cost camera and vector HD map is proven to be a practical and effective way of estimating the location and orientation of intelligent vehicles. However, map-based approach is viable only when the landmark observation is adequate and precise. In some areas with sparse and noisy observation, or even non-existent map matching features, the localization results may be unstable. In this paper, we introduce a novel algorithm by fusing visual odometry and vector HD map in a tightly-coupled optimization framework to tackle these problems. Our algorithm exploits the observation of visual feature points and vector HD map landmarks in the sliding window manner and optimize their residuals in a tightly-coupled approach. In this way, the system is more robust against the noisy HD map landmark observations. In addition, our method is able to accurately estimate vehicle pose even when landmarks are sparse. Experiments under two challenging scenarios with noisy and sparse landmark observations show that our method can achieve the Mean Absolute Error (MAE) at 0.1473m and 0.2496m respectively. Tuopu Wen, Zhongyang Xiao, Benny Wijaya, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 5 |
| 2019 | High Precision Target Positioning Method for RSU in Cooperative PerceptionabstractVehicle-road cooperative perception system can greatly improve the perception ability of intelligent vehicles by making use of perception information from road side units (RSU). This paper focuses on the target positioning of static camera for vehicle-road cooperation. A low-cost camera calibration method is proposed to complete the accurate mapping between the image plane and the 3D world space. Precise location of interested targets are achieved by an efficient tracking strategy. Real test scenarios show that our algorithm can effectively locate vehicles, pedestrians, non-motor vehicles and other targets with high accuracy. Our algorithm won the Monocular Static Camera Positioning and Ranging Competition for Autonomous Driving championship in 2019. Tuopu Wen, Zhongyang Xiao, Kun Jiang 0002, Mengmeng Yang 0001, Keqiang Li 0002, Diange Yang |
MMSP | 4 |