Diange Yang

dblp:134/0733 · DBLP profile ↗
← Back
55ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0003-0825-5609ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LSRE: Latent Semantic Rule Encoding for Real-Time Semantic Risk Detection in Autonomous Driving
Weitao Zhou, Cheng Jing, Nanshan Deng, Junze Wen, Kun Jiang 0002, Diange Yang
IV8
2026 Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation
abstract
Corner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for generating diverse scenarios, commonly achieved through 3D Gaussian Splatting or image generative models. However, these approaches often suffer from limited visual fidelity or imprecise pose control. To address these issues, we propose G^2Editor, a framework designed for photorealistic and precise object editing in driving videos. Our method leverages a 3D Gaussian representation of the edited object as a dense prior, injected into the denoising process to ensure accurate pose control and spatial consistency. A scene-level 3D bounding box layout is employed to reconstruct occluded areas of non-target objects. Furthermore, to guide the appearance details of the edited object, we incorporate hierarchical fine-grained features as additional conditions during generation. Experiments on the Waymo Open Dataset demonstrate that G^2Editor effectively supports object repositioning, insertion, and deletion within a unified framework, outperforming existing methods in both pose controllability and visual quality, while also benefiting downstream data-driven tasks.
Jiusi Li, Jackson Jiang, Jinyu Miao, Miao Long, Tuopu Wen, Peijin Jia, Shengxiang Liu, Chun-lei Yu, Maolin Liu, Yuzhan Cai, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang
IV13
2026 MaPLocator: Map-Prior Enhanced End-to-End Vehicle Localization with Coarse-to-Fine Pose Refinement
Muxi Tang, Jinyu Miao, Rujun Yan, Le Jia, Mengmeng Yang 0001, Diange Yang, Kun Jiang 0002
IV7
2026 DTCCL: Disengagement-Triggerd Contrastive Continual Learning for Autonomous Bus Planners
Yanding Yang, Weitao Zhou, Xiaomin Guo, Junze Wen, Lang Ding, Zheng Fu, Jinyu Miao, Kun Jiang 0002, Diange Yang
IV11
2026 Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion Prediction
abstract
LiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion
Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang
IEEE Trans. Intell. Transp. Syst.9
2025 Enhancing Autonomous Vehicle Planning With a Robust Fault-Tolerant Mechanism for Action-Induced Agent Detection
abstract
In autonomous driving, accurately identifying traffic participants that may influence vehicle behavior is crucial for effective system planning. To address this challenge, we propose a fault-tolerant mechanism for detecting action-induced objects, which significantly improves decision-making performance and system explainability. Since these objects are often linked to the vehicle’s driving intentions, we introduce a top-down attention network that adjusts attention weights for traffic participants based on navigational information. Additionally, we define potentially hazardous objects in the driving environment and employ supervised training with a classification head to detect them. To further enhance detection accuracy, we integrate a fault-tolerant process that merges attention maps with classification results, effectively reducing false positives and false negatives in identifying action-induced objects. Extensive testing validates the robustness and effectiveness of our approach, demonstrating its ability to improve both planning and interpretability in autonomous vehicles.
Zheng Fu, Hezhe Lin, Kangan Qian, Tuopu Wen, Hao Gao 0005, Diange Yang
ICASSP7
2025 PriorMotion: Generative Class-Agnostic Motion Prediction with Raster-Vector Motion Field Priors
abstract
Reliable spatial and motion perception is essential for safe autonomous navigation. Recently, class-agnostic motion prediction on bird's-eye view (BEV) cell grids derived from LiDAR point clouds has gained significant attention. However, existing frameworks typically perform cell classification and motion prediction on a per-pixel basis, neglecting important motion field priors such as rigidity constraints, temporal consistency, and future interactions between agents. These limitations lead to degraded performance, particularly in sparse and distant regions. To address these challenges, we introduce \textbf{PriorMotion}, an innovative generative framework designed for class-agnostic motion prediction that integrates essential motion priors by modeling them as distributions within a structured latent space. Specifically, our method captures structured motion priors using raster-vector representations and employs a variational autoencoder with distinct dynamic and static components to learn future motion distributions in the latent space. Experiments on the nuScenes dataset demonstrate that \textbf{PriorMotion} outperforms state-of-the-art methods across both traditional metrics and our newly proposed evaluation criteria. Notably, we achieve improvements of approximately 15.24\% in accuracy for fast-moving objects, an 3.59\% increase in generalization, a reduction of 0.0163 in motion stability, and a 31.52\% reduction in prediction errors in distant regions. Further validation on FMCW LiDAR sensors confirms the robustness of our approach.
Kangan Qian, Jinyu Miao, Xinyu Jiao, Ziang Luo, Zheng Fu, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Diange Yang
ICCV9
2025 Efficient End-to-end Visual Localization for Autonomous Driving with Decoupled BEV Neural Matching
abstract
Accurate localization plays an important role in high-level autonomous driving systems. Conventional map matching-based localization methods solve the poses by explicitly matching map elements with sensor observations, generally sensitive to perception noise, therefore requiring costly hyperparameter tuning. In this paper, we propose an end-to-end localization neural network which directly estimates vehicle poses from surrounding images, without explicitly matching perception results with HD maps. To ensure efficiency and interpretability, a decoupled BEV neural matching-based pose solver is proposed, which estimates poses in a differentiable sampling-based matching module. Moreover, the sampling space is hugely reduced by decoupling the feature representation affected by each DoF of poses. The experimental results demonstrate that the proposed network is capable of performing decimeter level localization with mean absolute errors of 0.19m, 0.13m and 0.39° in longitudinal, lateral position and yaw angle while exhibiting a 68.8% reduction in inference memory usage.
Jinyu Miao, Tuopu Wen, Ziang Luo, Kangan Qian, Zheng Fu, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Jin Huang 0002, Diange Yang
IROS11
2025 LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction
abstract
Accurate spatial and motion understanding is critical for autonomous driving systems. While object-level perception models excel in structured environments, they struggle with open-set categories and often lack precise geometric representation. Occupancy-based, class-agnostic methods offer better scene expressiveness but typically ignore inter-agent interactions and fail to ensure physical consistency in motion predictions, limiting their reliability in complex traffic scenarios. In this paper, we propose LEGO-Motion, a novel class-agnostic motion prediction framework that bridges the gap between instance-level reasoning and occupancy-based modeling. Unlike conventional grid-based methods that treat each cell independently, LEGO-Motion introduces two key components: (1) the Interaction-Augmented Instance Encoder (IaIE), which models interactions among dynamic agents via cross-attention, and (2) the Instance-Enhanced BEV Encoder (IeBE), which improves motion consistency across instances through multi-stage feature fusion. These components enable our model to learn semantically coherent and physically plausible motion fields. Extensive experiments on the nuScenes dataset show that LEGO-Motion achieves a around 6% improvement in motion prediction accuracy over the previous state-of-the-art, while maintaining real-time inference at 21ms. Moreover, our method demonstrates strong generalization on a proprietary FMCW LiDAR benchmark. These results validate LEGO-Motion's effectiveness in capturing both global scene structure and fine-grained motion dynamics, making it a promising foundation for next-generation perception systems.
Kangan Qian, Jinyu Miao, Ziang Luo, Zheng Fu, Jinchen Li, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang
IROS10
2025 EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Driving
abstract
3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or transformers at the voxel level. Moreover, OccNets are supervised with expensive large-scale dense voxel labels. Model and label inefficiencies, caused by excessive network parameters and label annotation requirements, severely hinder the onboard deployment of OccNets. This paper proposes an EFFicient Occupancy learning framework, EFFOcc, that targets minimal network complexity and label requirements while achieving state-of-the-art accuracy. We first propose an efficient fusion-based OccNet that only uses simple 2D operators and improves accuracy to the state-of-the-art on three large-scale benchmarks: Occ3D-nuScenes, Occ3D-Waymo, and OpenOccupancy-nuScenes. On the Occ3D-nuScenes benchmark, the fusion-based model with ResNet-18 as the image backbone has 21.35M parameters and achieves 51.49 in terms of mean Intersection over Union (mIoU). Furthermore, we propose a multi-stage occupancy-oriented distillation to efficiently transfer knowledge to vision-only OccNet. Extensive experiments on occupancy benchmarks show state-of-the-art precision for both fusion-based and vision-based OccNets. For the demonstration of learning with limited labels, we achieve 94.38% of the performance (mIoU = 28.38) of a 100% labeled vision OccNet (mIoU = 30.07) using the same OccNet trained with only 40% labeled sequences and distillation from the fusion-based OccNet. Code is available at https://github.com/synsin0/EFFOcc.
Yining Shi 0002, Kun Jiang 0002, Jinyu Miao, Ke Wang 0021, Kangan Qian, Yunlong Wang 0009, Jiusi Li, Tuopu Wen, Mengmeng Yang 0001, Yiliang Xu, Diange Yang
IROS11
2025 DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving Policy
abstract
With the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these disengagement cases for policy improvement, the inherent scarcity of disengagement data (often occurring as a single instance) restricts training effectiveness. Furthermore, some disengagement data should be excluded since the disengagement may not always come from the failure of driving policies, e.g. the driver may casually intervene for a while. To this end, this work proposes disengagement-reason-augmented reinforcement learning (DRARL), which enhances driving policy improvement process according to the reason of disengagement cases. Specifically, the reason of disengagement is identified by an out-of-distribution (OOD) state estimation model. When the reason doesn’t exist, the case will be identified as a casual disengagement case, which doesn’t require additional policy adjustment. Otherwise, the policy can be updated under a reason-augmented imagination environment, improving the policy performance of disengagement cases with similar reasons. The method is evaluated using real-world disengagement cases collected by autonomous driving robotaxi. Experimental results demonstrate that the method accurately identifies policy-related disengagement reasons, allowing the agent to handle both original and semantically similar cases through reason-augmented training. Furthermore, the approach prevents the agent from becoming overly conservative after policy adjustments. Overall, this work provides an efficient way to improve driving policy performance with disengagement cases.
Weitao Zhou, Bo Zhang 0106, Zhong Cao 0003, Xiang Li 0001, Diange Yang
IROS8
2025 LDMapNet-U: An End-to-End System for City-Scale Lane-Level Map Updating
abstract
An up-to-date city-scale lane-level map is an indispensable infrastructure and a key enabling technology for ensuring the safety and user experience of autonomous driving systems. In industrial scenarios, reliance on manual annotation for map updates creates a critical bottleneck. Lane-level updates require precise change information and must ensure consistency with adjacent data while adhering to strict standards. Traditional methods utilize a three-stage approach -- construction, change detection, and updating -- which often necessitates manual verification due to accuracy limitations. This results in labor-intensive processes and hampers timely updates. To address these challenges, we propose LDMapNet-U, which implements a new end-to-end paradigm for city-scale lane-level map updating. By reconceptualizing the update task as an end-to-end map generation process grounded in historical map data, we introduce a paradigm shift in map updating that simultaneously generates vectorized maps and change information. To achieve this, a Prior-Map Encoding (PME) module is introduced to effectively encode historical maps, serving as a critical reference for detecting changes. Additionally, we incorporate a novel Instance Change Prediction (ICP) module that learns to predict associations with historical maps. Consequently, LDMapNet-U simultaneously achieves vectorized map element generation and change detection. To demonstrate the superiority and effectiveness of LDMapNet-U, extensive experiments are conducted using large-scale real-world datasets. In addition, LDMapNet-U has been successfully deployed in production at Baidu Maps since April 2024, supporting lane-level map updating for over 360 cities and significantly shortening the update cycle from quarterly to weekly, thereby enhancing the timeliness and accuracy of lane-level map. The nationwide, high-frequency city-scale lane-level map has been instrumental in the development of the lane-level navigation product serving hundreds of millions of users, while also integrating into the autonomous driving systems of several leading vehicle companies.
Deguo Xia, Weiming Zhang 0006, Xiyan Liu, Wei Zhang 0088, Chenting Gong, Xiao Tan 0001, Jizhou Huang, Mengmeng Yang 0001, Diange Yang
KDD (1)9
2025 COME: Adding Scene-Centric Forecasting Control to Occupancy World Model
abstract
World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to separate environmental changes from ego-motion by leveraging the scene-centric coordinate systems. In this paper, we introduce COME: a framework that integrates scene-centric forecasting Control into the Occupancy world ModEl. Specifically, COME first generates ego-irrelevant, spatially consistent future features through a scene-centric prediction branch, which are then converted into scene condition using a tailored ControlNet. These condition features are subsequently injected into the occupancy world model, enabling more accurate and controllable future occupancy predictions. Experimental results on the nuScenes-Occ3D dataset show that COME achieves consistent and significant improvements over state-of-the-art (SOTA) methods across diverse configurations, including different input sources (ground-truth, camera-based, fusion-based occupancy) and prediction horizons (3s and 8s). For example, under the same settings, COME achieves 26.3% better mIoU metric than DOME and 23.7% better mIoU metric than UniScene. These results highlight the efficacy of disentangled representation learning in enhancing spatio-temporal prediction fidelity for world models. Code is available at https://github.com/synsin0/COME.
Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Tuopu Wen, Mengmeng Yang 0001, Diange Yang
NeurIPS9
2025 Pedestrian Trajectory Prediction for Autonomous Vehicles With Multiple Interactions
abstract
Pedestrian trajectory prediction is significant for autonomous vehicles, but the difficulty of pedestrian trajectory prediction lies in the accurate modeling of pedestrian multiple interactions. In this paper, we attempt to explore the essential features of pedestrian interaction and propose a pedestrian trajectory prediction method based on multiple interactions. Firstly, considering that the interaction between self-driving cars and pedestrians resembles a dynamic game process involving sequential adaptation, we map them to the same feature space and design a temporal cross-attention mechanism to model the interaction between pedestrians and vehicles. Meanwhile, pedestrian-scene interaction is affected by the global environment as well as the local environment. To capture the global information while preserving the spatial location of pedestrians in the scene, we design a pedestrian-scene heatmap fusion (PSHF) framework to model the pedestrian-scene interaction features. We validate the effectiveness of our algorithm on the publicly available JAAD and PIE datasets, achieving better performance than existing representative methods in both single-trajectory and multi-trajectory prediction tasks. We conducted a thorough ablation study, cross-dataset validation, and qualitative visualization experiments, demonstrating the effectiveness and robustness of our method.
Zheng Fu, Mengmeng Yang 0001, Kun Jiang 0002, Jin Huang 0002, Hao Gao 0005, Diange Yang
IEEE Internet Things J.9
2025 Top-Down Attention-Based Mechanisms for Interpretable Autonomous Driving
abstract
Despite the remarkable advancements in autonomous driving, the challenge persists in achieving interpretable action decision-making, primarily owing to the intricate and ambiguous relationship between detected agents and driving intention. In this study, we introduce an interpretable action prediction model, denoted as the Prediction-Driven Attention Network (PDANet), designed to undertake action decisions and provide corresponding interpretations cohesively. The PDANet is inspired by the perceptual mechanisms inherent in human drivers, who allocate attention according to their driving intentions. Specifically, we elaborate a prediction module to generate vehicle prospective trajectories to characterize driving intentions. Subsequently, the features of this predicted trajectory are utilized to modulate the attention distribution among agents through the top-down attention module, yielding an attention map. Finally, two distinct task tokens are applied to aggregate agent features and generate the final output according to the derived attention map. Extensive experiments conducted on the publicly available BDD-OIA and nu-AR datasets demonstrate that our proposed method outperforms all prior works in terms of both action prediction and behavior interpretation tasks. Remarkably, our method attains a noteworthy enhancement in the behavior interpretation task, surpassing the previous state-of-the-art by a substantial margin of +10.8% in terms of F1-score on the nu-AR dataset. We also validate our algorithm on Carla Town05 long in a closed-loop decision-making scenario, highlighting the generality and robustness of our approach. Furthermore, qualitative results show that the agents selected by our model are more closely aligned with human cognitive processes.
Zheng Fu, Kun Jiang 0002, Yunlong Wang 0009, Tuopu Wen, Hao Gao 0005, Diange Yang
IEEE Trans. Intell. Transp. Syst.8
2025 RM2Occ: Re-Projection Multi-Task Multi-Sensor Fusion for Autonomous Driving 3D Object Detection and Occupancy Perception
abstract
Occupancy prediction plays a crucial role in supporting autonomous driving planning and decision-making. Existing methods typically rely on modular stacking and fusion techniques of object detection, semantic segmentation, and depth estimation to achieve 3D occupancy. However, they fail to deeply explore the transformation relationships between 2D and 3D spaces and to efficiently fuse the different characteristics of multi-source sensors. We propose R$M^{2}$Occ, the first 3D occupancy perception network that integrates multi-sensor fusion based on different sensor principles and achieves multi-task learning. To leverage the rich 2D semantic information captured by cameras and elevate it to the 3D domain, we begin by querying and populating predefined empty voxels with multi-view image features. Subsequently, we progressively fuse 3D LiDAR point clouds with these populated voxels through an unbalanced fusion strategy that effectively supplements missing information and suppresses noise. Leveraging IMU data and calibration parameters, we then re-project the enriched voxels back onto the 2D image plane according to camera coordinates, performing a secondary query using the semantic segmentation results to recover semantic details potentially lost due to radar fusion limitations and incomplete voxel querying. Finally, supported by a multi-task detection head, R$M^{2}$Occ simultaneously accomplishes 3D object detection, semantic segmentation, Bird’s Eye View (BEV) detection, and full-scene grid occupancy prediction, enabling comprehensive multi-task output. Extensive experiments and ablation studies on the nuScenes dataset demonstrate that R$M^{2}$Occ significantly outperforms existing state-of-the-art methods, establishing a new paradigm for accurate and efficient multi-sensor fusion and multi-task perception in autonomous driving scenarios.
Yilong Ren, Minda Li, Han Jiang 0003, Zhiyong Cui, Mengmeng Yang 0001, Haiyang Yu 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.8
2025 Toward Democratizing High-Definition Map Update Through Consortium Blockchain
abstract
In the rapidly evolving landscape of autonomous vehicles and advanced navigation systems, the accuracy of high-definition maps and real-time updating has become paramount. However, in this progression, the security of map data has not received adequate attention, although the accuracy of the data can be easily altered when the system is breached. Thus, this paper introduces a novel approach to democratizing the process of high-definition map updates by leveraging consortium blockchain technology specifically designed for Proof of Presence and Reputation (POP-R) to safeguard the update process. Our proposed system leverages the presence and reputation of vehicles through infrastructure nodes to enhance the accuracy and reliability of HD map updates. We created a trusted ecosystem for maintaining high-definition maps, marked by a superior safety score across three scenarios compared to the standard proof of reputation technique. Additionally, it demonstrates high efficiency, achieving 12,000 transactions per second (TPS) for data queries and more than 2,500 TPS for data writing in our blockchain network. This efficiency proved our prowess in the lightweight computational power required, suitable for decentralized and crowdsourced-based systems. Through our POP-R framework, we lay the foundation for a new decentralized approach to the evolution of high-definition maps in the era of autonomous mobility.
Benny Wijaya, Mengmeng Yang 0001, Tuopu Wen, Kun Jiang 0002, Wei Zhang 0090, Yunlong Wang 0009, Zheng Fu, Xuewei Tang, Diange Yang
IEEE Trans. Intell. Transp. Syst.9
2025 Grid-Centric Traffic Scenario Perception for Autonomous Driving: A Comprehensive Review
abstract
The grid-centric perception is a crucial field for mobile robot perception and navigation. Nonetheless, the grid-centric perception is less prevalent than object-centric perception as autonomous vehicles need to accurately perceive highly dynamic, large-scale traffic scenarios, and the complexity and computational costs of grid-centric perception are high. In recent years, the rapid development of deep learning techniques and hardware provides fresh insights into the evolution of grid-centric perception. The fundamental difference between grid-centric and object-centric pipeline lies in that grid-centric perception follows a geometry-first paradigm which is more robust to the open-world driving scenarios with endless long-tailed semantically unknown obstacles. Recent research demonstrates the great advantages of grid-centric perception, such as comprehensive fine-grained environmental representation, greater robustness to occlusion and irregular-shaped objects, better ground estimation, and safer planning policies. There is also a growing trend that the capacity of occupancy networks is greatly expanded to 4-D scene perception and prediction, and the latest techniques are highly related to new research topics, such as 4-D occupancy forecasting, generative artificial intelligence (GenAI), and world models in the field of autonomous driving. Given the lack of current surveys for this rapidly expanding field, we present a hierarchically structured review of grid-centric perception for autonomous vehicles. We organize previous and current knowledge of occupancy grid techniques along the main vein from 2-D bird-eye view (BEV) grids to 3-D occupancy to 4-D occupancy forecasting. We additionally summarize label-efficient occupancy learning and the role of grid-centric perception in driving systems. Finally, we present a summary of the current research trend and provide future outlooks.
Yining Shi 0002, Kun Jiang 0002, Jiusi Li, Zelin Qian, Junze Wen, Mengmeng Yang 0001, Ke Wang 0021, Diange Yang
IEEE Trans. Neural Networks Learn. Syst.8
2024 PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous Driving
abstract
Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect obstacles regardless of their shape and occlusion. Modern occupancy networks mainly focus on reconstructing visible voxels from object surfaces with voxel-wise semantic prediction. Usually, they suffer from inconsistent predictions of one object and mixed predictions for adjacent objects. These confusions may harm the safety of downstream planning modules. To this end, we investigate panoptic segmentation on 3D voxel scenarios and propose an instance-aware occupancy network, PanoSSC. We predict foreground objects and backgrounds separately and merge both in post-processing. For foreground instance grouping, we propose a novel 3D instance mask decoder that can efficiently extract individual objects. we unify geometric reconstruction, 3D semantic segmentation, and 3D instance segmentation into PanoSSC framework and propose new metrics for evaluating panoptic voxels. Extensive experiments show that our method achieves competitive results on SemanticKITTI semantic scene completion benchmark.
Yining Shi 0002, Jiusi Li, Kun Jiang 0002, Ke Wang 0021, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang
3DV7
2024 StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation
abstract
Predicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However, current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict uniform snapshots of future occupancy states and require strictly synchronized sensory data for sensor fusion. We propose a novel framework, StreamingFlow, to lift these strong limitations. StreamingFlow is a novel BEV occupancy predictor that ingests asynchronous multi-sensor data streams for fusion and performs streaming fore-casting of the future occupancy map at any future times-tamps. By integrating neural ordinary differential equations (N-ODE) into recurrent neural networks, StreamingFlow learns derivatives of BEV features over temporal horizons, updates the implicit sensor's BEV features as part of the fusion process, and propagates BEV states to the desired future time point. It shows good zero-shot generalization ability of prediction, reflected in the interpolation of the ob-served prediction time horizon and the reasonable inference of the unseen farther future period. Extensive experiments on two large-scale datasets, nuScenes [2] and Lyft L5 [14], demonstrate that StreamingFlow significantly outperforms previous vision-based, LiDAR-based methods, and shows superior performance compared to state-of-the-art fusion-based methods.
Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Jiusi Li, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang
CVPR7
2024 LaneDAG: Automatic HD Map Topology Generator Based on Geometry and Attention Fusion Mechanism
abstract
In high-definition maps (HD maps), the road lane centerline and lane topology graph play essential roles in navigation, planning, and decision-making. Existing research focusing on extracting physical infrastructure, such as lane boundaries, has made significant progress. But lane centerline detection and topology reasoning still remains challenging due to the severe overlapping centerlines and complicated topology. To tackle these challenges, we introduce an automatic lane topology extraction method for HD maps, termed LaneDAG, which extracts vectorized centerlines and their topology from prebuilt lane lines and road boundaries in HD maps. It formulates centerline extraction as a set prediction problem and lane topology prediction as a directed acyclic graph (DAG) construction problem. A novel mechanism that fusing geometric and attention-based features in the DAG is proposed to model the topological relationship between centerlines. Experiments conducted on the Argoverse 2 dataset demonstrate the proposed method’s superior performance compared to existing methods, showcasing its capability to extract lane centerlines and topology in HD maps automatically.
Peijin Jia, Tuopu Wen, Ziang Luo, Zheng Fu, Jiaqi Liao, Huixian Chen, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang
IV9
2024 DuMapNet: An End-to-End Vectorization System for City-Scale Lane-Level Map Generation
abstract
Generating city-scale lane-level maps faces significant challenges due to the intricate urban environments, such as blurred or absent lane markings. Additionally, a standard lane-level map requires a comprehensive organization of lane groupings, encompassing lane direction, style, boundary, and topology, yet has not been thoroughly examined in prior research. These obstacles result in labor-intensive human annotation and high maintenance costs. This paper overcomes these limitations and presents an industrial-grade solution named DuMapNet that outputs standardized, vectorized map elements and their topology in an end-to-end paradigm. To this end, we propose a group-wise lane prediction (GLP) system that outputs vectorized results of lane groups by meticulously tailoring a transformer-based network. Meanwhile, to enhance generalization in challenging scenarios, such as road wear and occlusions, as well as to improve global consistency, a contextual prompts encoder (CPE) module is proposed, which leverages the predicted results of spatial neighborhoods as contextual information. Extensive experiments conducted on large-scale real-world datasets demonstrate the superiority and effectiveness of DuMapNet. Additionally, DuMapNet has already been deployed in production at Baidu Maps since June 2023, supporting lane-level map generation tasks for over 360 cities while bringing a 95% reduction in costs. This demonstrates that DuMapNet serves as a practical and cost-effective industrial solution for city-scale lane-level map generation.
Deguo Xia, Weiming Zhang 0006, Xiyan Liu, Wei Zhang 0114, Chenting Gong, Jizhou Huang, Mengmeng Yang 0001, Diange Yang
KDD8
2024 Accuracy and Safety: Tracking Control of Heavy-Duty Cooperative Transportation Systems Using Constraint-Following Method
abstract
Accurate tracking control of autonomous cooperative transportation systems (CTS) remains challenging owing to the complexity of the mechanisms and the high requirements of coordination between carriers. In this paper, the trajectory tracking control of a CTS with a pair of autonomous vehicles serving as carriers is investigated. A novel constraint-oriented hierarchical modeling method is proposed to describe the dynamics of the system. By dividing the system dynamics into two portions: the lower-level individual modeling and the upper-level constraints abstraction, the modeling process is significantly simplified. Then an innovative constraint-following control law is designed to address the tracking control problem under the special system topology, based on the internal and external constraints designed in the modeling process. The asymptotic convergence of the tracking error is theoretically guaranteed. To reduce potential damage of the payload during transportation, a payload force optimization method is creatively proposed. It relies on the closed-form relationship between the control input and payload forces established by the constraint-oriented modeling. The normal and shear stress on the payload is successfully limited, without affecting the trajectory tracking performance. Simulation results show that the proposed control method and the payload force optimization strategy can help achieve accurate and safe autonomous cooperative transportation simultaneously.
Bowei Zhang 0008, Ye-Hwa Chen, Yi-fan Jia 0001, Jin Huang 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.5
2024 Distributed Collaborative Control of Multi-Vehicle Autonomous Cooperative Transportation Systems: A Hierarchical Constraint-Following Approach
abstract
In this paper, the dynamic modeling and collaborative control of a multi-vehicle cooperative transportation system for load carrying is explored. A hierarchical modeling and constraint-following control scheme is creatively proposed. In the dynamic modeling stage, the separate models of system components including the load and vehicle carriers are firstly established at the lower level. Then the internal and external constraints corresponding to the system topology and the transportation task are designed to integrate separate models at a higher level. In the system control stage, a distributed collaborative control law is proposed based on the closed-form constraint forces, with which the load can follow the external constraints actively and the carriers can maintain the internal constraints passively. In order to overcome the influence of time-varying multi-source uncertainties of the system on control effectiveness and stability, an adaptive robust control term is designed based on the Lyapunov min-max approach. Both uniform boundedness and uniform ultimate boundedness of the constraint-following error are guaranteed. Comprehensive validations show that our propose scheme can significantly reduce the modeling complexity despite the strongly coupled topology and nonlinearity of the system, as well as achieving more precise and robust trajectory following control compared with the baseline methods.
Bowei Zhang 0008, Jin Huang 0002, Yanzhao Su, Ye-Hwa Chen, Diange Yang
IEEE Trans. Intell. Transp. Syst.5
2024 Safety-Guaranteed Oversized Cargo Cooperative Transportation With Closed-Form Collision-Free Trajectory Generation and Tracking Control
abstract
In this article, the trajectory generation and motion control of autonomous driving oversized cargo cooperative transportation systems (CTS) in static but bounded environment is investigated. Different from common vehicle systems, the challenges lie on the safety-guaranteed cooperation of independently controlled carriers with inherent connections brought by the rigid payload, which results in complex system dynamics and multiple time-variant uncertainties. A constraint-oriented “leader-follower” modeling and control framework is introduced, and a trajectory generation method based on the diffeomorphism is creatively proposed to generate closed-form collision-free trajectory for the payload in the bounded environment. To achieve safety-guaranteed trajectory following under uncertainties, a transformed adaptive robust control strategy (TARC) is designed through constraint relaxation, and the coordination of the carriers is realized. An implementation with comprehensive ablation studies demonstrates the effectiveness of our trajectory generation and tracking control framework. The collision-free trajectory set is efficiently generated, and the CTS can be kept strictly inside the safe corridor with high tracking accuracy, which is extremely hard for the baseline methods.
Bowei Zhang 0008, Jin Huang 0002, Yanzhao Su, Xiangyu Wang 0005, Ye-Hwa Chen, Diange Yang
IEEE Trans. Intell. Transp. Syst.6
2023 Stochastic pedestrian avoidance for autonomous vehicles using hybrid reinforcement learning
abstract
Ensuring the safety of pedestrians is essential and challenging when autonomous vehicles are involved. Classical pedestrian avoidance strategies cannot handle uncertainty, and learning-based methods lack performance guarantees. In this paper we propose a hybrid reinforcement learning (HRL) approach for autonomous vehicles to safely interact with pedestrians behaving uncertainly. The method integrates the rule-based strategy and reinforcement learning strategy. The confidence of both strategies is evaluated using the data recorded in the training process. Then we design an activation function to select the final policy with higher confidence. In this way, we can guarantee that the final policy performance is not worse than that of the rule-based policy. To demonstrate the effectiveness of the proposed method, we validate it in simulation using an accelerated testing technique to generate stochastic pedestrians. The results indicate that it increases the success rate for pedestrian avoidance to 98.8%, compared with 94.4% of the baseline method.
Huiqian Li, Jin Huang 0002, Zhong Cao 0003, Diange Yang
Frontiers Inf. Technol. Electron. Eng.4
2023 Traffic Police 3D Gesture Recognition Based on Spatial-Temporal Fully Adaptive Graph Convolutional Network
abstract
It is critical for autonomous vehicles to recognize traffic police gestures timely and accurately. During the movement of the vehicle, the collected traffic police scales change all the time, in addition, the frequency and amplitude of actions of different traffic police are different. First, we use gesture normalization to fix the traffic police actions at a unified scale and remove the influence of scale changes on traffic police gesture recognition. Meanwhile, a fully adaptive spatial-temporal graph convolution network (FA-STGCN) is proposed to recognize the actions with different amplitude and frequencies. The adaptive spatial graph network can dig the latent joints connection relation of the traffic police under different gestures, which weakens the amplitude impact on the action recognition. The adaptive temporal graph network is composed of the global temporal module and the local temporal module. The global temporal module can obtain the coarse-grained features of the traffic police gestures’ speed and then naturally use the coarse-grained features to guide the local temporal module to adaptively learn the fine-grained temporal features of the traffic police action. The adaptive spatial graph network and the temporal graph network are alternately stacked to finally output accurate traffic police gestures. We thoroughly evaluated our method through intensive experiments, the result shows that our method achieved the best results on public datasets. What’s more, we proofed the effectiveness of each module and verified our methods for moving vehicles for the first time, the performance present meets the vehicle’s practical requirements.
Zheng Fu, Kun Jiang 0002, Junze Wen, Mengmeng Yang 0001, Diange Yang
IEEE Trans. Intell. Transp. Syst.7
2023 Reliable Autonomous Driving Environment Model With Unified State-Extended Boundary
abstract
From the early stage of robotic applications to current autonomous driving technologies, environment modeling has been acting as the middleware for connecting perception and decision layers. In robotic applications, space-oriented models (e.g., grid map, drivable area) are widely applied to faithfully reflect the space occupation. With the development of autonomous driving, highly dynamic and complex road environment brings rising need to understand the type and motion status of objects, thus element list has became the mainstream environment model. However, along comes the reliablity problem caused by missed detection and irregular objects, which is still inevitable despite the detection accuracy improvement. In view of this, a new view of driving environment is proposed as the unified state-extended boundary (USEB), aiming to improve the reliablity of element-oriented model. For driving decision requirements, different types of elements are consistently converted into driving constraints. Semantics and dynamics are expressed as the status of drivable area boundary, making it possible to merge space occupation to improve reliability against missed detection and irregular objects. Evaluation of USEB is carried out on the nuScenes dataset. Comparative results show that the proposed USEB could cover the required information for driving decision, whereas achieving higher reliability than the commonly applied element-oriented model.
Xinyu Jiao, Kun Jiang 0002, Yunlong Wang 0009, Zhong Cao 0003, Mengmeng Yang 0001, Diange Yang
IEEE Trans. Intell. Transp. Syst.7
2023 Semantic Traffic Law Adaptive Decision-Making for Self-Driving Vehicles
abstract
Facts proved that obeying traffic laws keeps the promise to promote the safety of self-driving vehicles. Current self-driving vehicles usually have fixed algorithms during autonomous driving, however the traffic laws may differ or change in different regions or times, e.g., tidal lanes. It raises a crucial requirement to make self-driving vehicles adapt to the newly received traffic laws. The challenges are that traffic laws are usually semantic and manually designed, but the original algorithms may not always contain the pre-designed interface to adapt to emerging laws. To this end, this work proposes a traffic law adaptive decision-making platform, which uses the linear temporal logic (LTL) formula to consistently describe the semantic traffic laws. Then, an LTL-based reinforcement learning framework is designed to estimate the probability of illegal behavior under different traffic laws. Finally, a law-specific backup policy is designed to maintain the performance threshold by monitoring the probability of illegal behavior. This work takes three typical scenarios where the traffic laws differ for instance to prove the effectiveness of the proposed approach, i.e., law amendment presented by the government, law difference between different regions, and temporary traffic control. The results show that the proposed method can help the original decision-making algorithms adapt to the traffic laws well without pre-defined interfaces. This method provides a way to administer on-road driving self-driving vehicles.
Hong Wang 0014, Zhong Cao 0003, Wenhao Yu 0006, Chengxiang Zhao, Ding Zhao, Diange Yang, Jun Li 0082
IEEE Trans. Intell. Transp. Syst.7
2023 Identify, Estimate and Bound the Uncertainty of Reinforcement Learning for Autonomous Driving
abstract
Deep reinforcement learning (DRL) has emerged as a promising approach for developing more intelligent autonomous vehicles (AVs). A typical DRL application on AVs is to train a neural network-based driving policy. However, the black-box nature of neural networks can result in unpredictable decision failures, making such AVs unreliable. To this end, this work proposes a method to identify and protect unreliable decisions of a DRL driving policy. The basic idea is to estimate and constrain the policy’s performance uncertainty, which quantifies potential performance drop due to insufficient training data or network fitting errors. By constraining the uncertainty, the DRL model’s performance is always greater than that of a baseline policy. The uncertainty caused by insufficient data is estimated by the bootstrapped method. Then, the uncertainty caused by the network fitting error is estimated using an ensemble network. Finally, a baseline policy is added as the performance lower bound to avoid potential decision failures. The overall framework is called uncertainty-bound reinforcement learning (UBRL). The proposed UBRL is evaluated on DRL policies with different amounts of training data, taking an unprotected left-turn driving case as an example. The result shows that the UBRL method can identify potentially unreliable decisions of DRL policy. The UBRL guarantees to outperform baseline policy even when the DRL policy is not well-trained and has high uncertainty. Meanwhile, the performance of UBRL improves with more training data. Such a method is valuable for the DRL application on real-road driving and provides a metric to evaluate a DRL policy.
Weitao Zhou, Zhong Cao 0003, Nanshan Deng, Kun Jiang 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.5
2023 Dynamically Conservative Self-Driving Planner for Long-Tail Cases
abstract
Self-driving vehicles (SDVs) are becoming reality but still suffer from “long-tail” challenges during natural driving: the SDVs will continually encounter rare, safety-critical cases that may not be included in the dataset they were trained. Some safety-assurance planners solve this problem by being conservative in all possible cases, which may significantly affect driving mobility. To this end, this work proposes a method to automatically adjust the conservative level according to each case’s “long-tail” rate, named dynamically conservative planner (DCP). We first define the “long-tail” rate as an SDV’s confidence to pass a driving case. The rate indicates the probability of safe-critical events and is estimated using the statistics bootstrapped method with historical data. Then, a reinforcement learning-based planner is designed to contain candidate policies with different conservative levels. The final policy is optimized based on the estimated “long-tail” rate. In this way, the DCP is designed to automatically adjust to be more conservative in low-confidence “long-tail” cases while keeping efficient otherwise. The DCP is evaluated in the CARLA simulator using driving cases with “long-tail” distributed training data. The results show that the DCP can accurately estimate the “long-tail” rate to identify potential risks. Based on the rate, the DCP automatically avoids potential collisions in “long-tail” cases using conservative decisions while not affecting the average velocity in other typical cases. Thus, the DCP is safer and more efficient than the baselines with fixed conservative levels, e.g., an always conservative planner. This work provides a technique to guarantee SDV’s performance in unexpected driving cases without resorting to a global conservative setting, which contributes to solving the “long-tail” problem practically.
Weitao Zhou, Zhong Cao 0003, Nanshan Deng, Kun Jiang 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.6
2022 BE-STI: Spatial-Temporal Integrated Network for Class-agnostic Motion Prediction with Bidirectional Enhancement
abstract
Determining the motion behavior of inexhaustible categories of traffic participants is critical for autonomous driving. In recent years, there has been a rising concern in performing class-agnostic motion prediction directly from the captured sensor data, like LiDAR point clouds or the combination of point clouds and images. Current motion prediction frameworks tend to perform joint semantic segmentation and motion prediction and face the trade-off between the performance of these two tasks. In this paper, we propose a novel Spatial-Temporal Integrated network with Bidirectional Enhancement, BE-STI, to improve the temporal motion prediction performance by spatial semantic features, which points out an efficient way to combine semantic segmentation and motion prediction. Specifically, we propose to enhance the spatial features of each individual point cloud with the similarity among temporal neighboring frames and enhance the global temporal features with the spatial difference among non-adjacent frames in a coarse-to-fine fashion. Extensive experiments on nuScenes and Waymo Open Dataset show that our proposed framework outperforms all state-of-the-art LiDAR-based and RGB+LiDAR-based methods with remarkable margins by using only point clouds as input.11The code will be released at https://github.com/be-sti/be-sti.
Yunlong Wang 0009, Hongyu Pan, Yu-Huan Wu, Xin Zhan, Kun Jiang 0002, Diange Yang
CVPR7
2022 Skeleton-based traffic command recognition at road intersections for intelligent vehicles
Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Tuopu Wen, Diange Yang
Neurocomputing7
2022 Temporal Point Cloud Fusion With Scene Flow for Robust 3D Object Tracking
abstract
Non-visual range sensors such as Lidar have shown the potential to detect, locate and track objects in complex dynamic scenes thanks to their higher stability in comparison with vision-based sensors like cameras. However, due to the disorder, sparsity, and irregularity of the point cloud, it is much more challenging to take advantage of the temporal information in the dynamic 3D point cloud sequences, as it has been done in the image sequences for improving detection and tracking. In this paper, we propose a novel scene-flow-based point cloud feature fusion module to tackle this challenge, based on which a 3D object tracking framework is also achieved to exploit the temporal motion information. Moreover, we carefully designed several training schemes that contribute to the success of this new module by eliminating the issues of overfitting and long-tailed distribution of object categories. Extensive experiments on the public KITTI 3D object tracking dataset demonstrate the effectiveness of the proposed method by achieving superior results to the baselines. The source code is available athttps://github.com/Tsinghua-OpenICV/SharingVan-OpenPCDet.
Yanding Yang, Kun Jiang 0002, Diange Yang, Yanqin Jiang
IEEE Signal Process. Lett.3
2022 Hybrid Car-Following Strategy Based on Deep Deterministic Policy Gradient and Cooperative Adaptive Cruise Control
abstract
Deep deterministic policy gradient (DDPG)-based car-following strategy can break through the constraints of the differential equation model due to the ability of exploration on complex environments. However, the car-following performance of DDPG is usually degraded by unreasonable reward function design, insufficient training, and low sampling efficiency. In order to solve this kind of problem, a hybrid car-following strategy based on DDPG and cooperative adaptive cruise control (CACC) is proposed. First, the car-following process is modeled as the Markov decision process to calculate CACC and DDPG simultaneously at each frame. Given a current state, two actions are obtained from CACC and DDPG, respectively. Then, an optimal action, corresponding to the one offering a larger reward, is chosen as the output of the hybrid strategy. Meanwhile, a rule is designed to ensure that the change rate of acceleration is smaller than the desired value. Therefore, the proposed strategy not only guarantees the basic performance of car-following through CACC but also makes full use of the advantages of exploration on complex environments via DDPG. Finally, simulation results show that the car-following performance of the proposed strategy is improved compared with that of DDPG and CACC. Note to Practitioners—This article presents a new car-following strategy, which avoids the impact of deep deterministic policy gradient (DDPG) performance degradation on the system. In the proposed strategy, DDPG is replaced with cooperative adaptive cruise control (CACC) when the performance of DDPG is worse than that of CACC. Meanwhile, a switching rule is designed to guarantee that the change rate of acceleration is smaller than the threshold. Simulation results show that the performance of hybrid car-following strategy has been improved compared with that of only using CACC or DDPG. Moreover, the proposed strategy has the advantages of low computational burden, high real-time performance, and good scalability.
Ruidong Yan, Rui Jiang 0008, Jin Huang 0002, Diange Yang
IEEE Trans Autom. Sci. Eng.5
2022 Design and Optimization of Robust Path Tracking Control for Autonomous Vehicles With Fuzzy Uncertainty
abstract
Uncertainty is a major concern in vehicle path tracking control design. The coefficients of the uncertainty bound are unknown. They are assumed to lie within prescribed fuzzy sets. First, based on the path tracking kinematic model, this article innovatively formulates the vehicle path tracking task as a constraint-following problem. Second, we put forward a deterministic adaptive robust control law with a tunable parameter to ensure the uniform boundedness and ultimate uniform boundedness of the closed-loop system. Third, an optimal scheme for the tunable parameter is proposed based on the fuzzy uncertainty. The resulting optimal robust control (ORC) minimizes a comprehensive fuzzy performance index that involves the fuzzy system performance and the control cost. The results of the CarSim-Simulink cosimulation and the hardware-in-loop experiment together show that the proposed ORC exhibits a superior path tracking performance.
Zeyu Yang 0002, Jin Huang 0002, Diange Yang
IEEE Trans. Fuzzy Syst.3
2022 Simple But Effective: Upper-Body Geometric Features for Traffic Command Gesture Recognition
abstract
Recognizing traffic command gestures with high accuracy and quick response at a low computational cost is a requisite for driver assistance or autonomous driving. However, it has been understudied for a long time. Existing research takes advantage of increasing development in human action recognition but pays little attention to onboard conditions. In this article, we propose a simple but effective recognition model based on human upper-body geometric features and a long short-term memory (LSTM) network. The handcrafted geometric features can easily be calculated with estimated 2-D human keypoints at a low computational cost but are discriminative and sufficient in classification. Offline and online inferences are implemented to comprehensively evaluate the proposed model. For the sake of robustness required in the automotive domain, dual voting is designed to filter the output in online inference. On the recently published Chinese traffic police gesture (CTPG) dataset, the presented approach is the best with a remarkable improvement of approximately 8% compared to previous LSTM-based methods with handcrafted spatial features and is competitive with advanced GCN-based deep learning methods. The tradeoff pattern is explored to demonstrate how accuracy and response time alter with different training and inference strategies so that a balanced setup can be manually chosen under various application scenarios. Field tests are also carried out with an experimental vehicle, and the results uncover the present gap between research and practical application to some extent, moving a step closer to real-life traffic command gesture recognition.
Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Diange Yang
IEEE Trans. Hum. Mach. Syst.6
2022 Confidence-Aware Reinforcement Learning for Self-Driving Cars
abstract
Reinforcement learning (RL) can be used to design smart driving policies in complex situations where traditional methods cannot. However, they are frequently black-box in nature, and the resulting policy may perform poorly, including in scenarios where few training cases are available. In this paper, we propose a method to use RL under two conditions: (i) RL works together with a baseline rule-based driving policy; and (ii) the RL intervenes only when the rule-based method seems to have difficulty handling and when the confidence of the RL policy is high. Our motivation is to use a not-well trained RL policy to reliably improve AV performance. The confidence of the policy is evaluated by Lindeberg-Levy Theorem using the recorded data distribution in the training process. The overall framework is named “confidence-aware reinforcement learning” (CARL). The condition to switch between the RL policy and the baseline policy is analyzed and presented. Driving in a two-lane roundabout scenario is used as the application case study. Simulation results show the proposed method outperforms the pure RL policy and the baseline rule-based policy.
Zhong Cao 0003, Shaobing Xu, Huei Peng, Diange Yang, Robert Zidek
IEEE Trans. Intell. Transp. Syst.4
2022 A General Autonomous Driving Planner Adaptive to Scenario Characteristics
abstract
Autonomous vehicle requires a general planner for all possible scenarios. Existing researches design such a planner by a unified scenario description. However, it may significantly increase the planner complexity even in some simple tasks, e.g., car following, further resulting in unsatisfactory driving performance. This work aims to design a general planner which can 1) drive in all possible scenarios and 2) have lower complexity in some common scenarios. To this end, this work proposes a pertinent boundary for multi-scenario driving planning. The total approach is named as Pertinent Boundary-based Unified Decision system. Based on the original drivable area, the pertinent boundary can further support motion status and semantics of the traffic elements, which provides the potential of pertinent performance for given scenarios. The pertinent boundary can support unified driving with the drivable area, in the meantime, can be pertinently modified to support the pertinent driving decisions for identified driving scenarios (e.g., car-following, junction left turning). It will further avoid the bump between the connections of the scenarios due to the continuity of space boundary. Thus, the planner is suitable for the fully autonomous driving. The proposed method is validated in different classical driving decision scenarios. Results show that the proposed method can support pertinent driving decisions in identified scenarios, in the meantime, assure generalized cross-scenario planning when no scenario information is available. Such a method shed light on fully autonomous driving by pertinence improvement of multi-scenario decision in the complex real world.
Xinyu Jiao, Zhong Cao 0003, Kun Jiang 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.5
2022 PNNUAD: Perception Neural Networks Uncertainty Aware Decision-Making for Autonomous Vehicle
abstract
Most environment perception methods in autonomous vehicles rely on deep neural networks because of their impressive performance. However, neural networks have black-box characteristics in nature, which may lead to perception uncertainty and untrustworthy autonomous vehicles. Thus, this work proposes a decision-making method to adapt the potential perception uncertainty due to the sensor noises, fuzzy features, and unfamiliar inputs. The whole method is named as Perception Neural Networks Uncertainty Aware Decision-Making (PNNUAD) method. PNNUAD first uses the Monte Carlo dropout method to estimate the perception neural network uncertainty into a distribution around the original output. Then, the perception uncertainty will be considered in a designed reinforcement learning-based planner using a distributed value function. Finally, a backup policy will maintain the vehicle’s performance to avoid disastrous perception uncertainty. The evaluation section uses an augmented reality urban driving scenario; namely, the scenario builds in the CARLA simulator while the perception uncertainty comes from the real dataset. This case study focuses on the object class uncertainty of a widely used neural network, i.e., YOLO-V3. The results indicate that the proposed method can maintain AV safety even with poor perception performance. Meanwhile, the AV has not become too conservative by defending the perception uncertainty. This work is necessary for applying the statistics neural networks to safety-critical autonomous vehicles, and the source code will be open-source in this work.
Hong Wang 0014, Zhong Cao 0003, Diange Yang, Jun Li 0082
IEEE Trans. Intell. Transp. Syst.5
2022 TM3Loc: Tightly-Coupled Monocular Map Matching for High Precision Vehicle Localization
abstract
Vision-based map-matching with HD map for high precision vehicle localization has gained great attention for its low-cost and ease of deployment. However, its localization performance is still unsatisfactory in accuracy and robustness in numerous real applications due to the sparsity and noise of the perceived HD map landmarks. This article proposes the tightly-coupled monocular map-matching localization algorithm (TM3Loc) for monocular-based vehicle localization. TM3Loc introduces semantic chamfer matching (SCM) to model monocular map-matching problem and combines visual features with SCM in a tightly-coupled manner. By applying the sliding window-based optimization technique, the historical visual features and HD map constraints are also introduced, such that the vehicle poses are estimated with an abundance of visual features and multi-frame HD map landmark features, rather than with single-frame HD map observations in previous works. Experiments are conducted on large scale dataset of 15 km long in total. The results show that TM3Loc is able to achieve high precision localization performance using a low-cost monocular camera, largely exceeding the performance of the previous state-of-the-art methods, thereby promoting the development of autonomous driving.
Tuopu Wen, Kun Jiang 0002, Benny Wijaya, Mengmeng Yang 0001, Diange Yang
IEEE Trans. Intell. Transp. Syst.6
2022 Distributed Car-Following Control for Intelligent Connected Vehicle Using Improved Super-Twisting Compensator Subject to Sudden Velocity Changes of Leading Vehicle
abstract
The optimal velocity-based model has been successfully applied to distributed car-following systems. However, the car-following performance is inevitably affected by a series of disturbances, particularly, sudden velocity changes of leading vehicle. To improve accuracy and response rate of car-following control in the presence of such disturbances, an improved super-twisting compensator (ISTC) is proposed and a composite controller is designed by combining ISTC with a finite- time controller. A second-order nominal system is constructed by using a virtual measurement signal along with its integration to facilitate the design of ISTC. By introducing the feedback of high-order estimation error, the accuracy and response rate of ISTC are increased significantly as compared with the conventional one under same gains. Such improvement further enhances the disturbance rejection ability of the composite controller. Both Lyapunov approach and numerical simulations are carried out to verify the effectiveness of the proposed method.
Ruidong Yan, Diange Yang, Jin Huang 0002, Kun Jiang 0002, Xinyu Jiao
IEEE Trans. Intell. Transp. Syst.2
2022 Path Tracking Control for Underactuated Vehicles With Matched-Mismatched Uncertainties: An Uncertainty Decomposition Based Constraint-Following Approach
abstract
This paper presents a robust path tracking control method by utilizing the ideology of constraint-following approach for uncertain underactuated autonomous vehicles. The uncertainties are bounded with an unknown boundary. They do not all fall within the range space of the input matrix. Based on kinematic relations between the desired path and vehicle, the path tracking task is transformed into an equality constraint of the vehicle lateral dynamics states. The control goal is to make the underactuated vehicle follow the constraint, thus realizing the desired tracking performance. The constraint-following robust control (CFRC) is designed in two steps. First, a servo control design for the nominal system is devised without considering uncertainty and initial constraint deviation. Second, the uncertainty is meticulously decomposed into matched and mismatched portions based on the geometric structural characteristics of the constraint dynamics system. As a result, since the mismatched uncertainty is orthogonal to the constraint-following geometric space, it “disappears” in the stability analysis. The matched uncertainty is estimated by a self-adjusting leakage type adaptive law. On this basis, a robust control is designed based on the estimated matched uncertainty. Through Lyapunov minimax analysis, the proposed control method guarantees the approximate constraint-following performance. Finally, the TruckSim–Simulink co-simulations and real vehicle experiments are presented. The results show that the proposed control can robustly realize excellent tracking performance in the presence of time-vary uncertainties.
Zeyu Yang 0002, Jin Huang 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.4
2021 Efficient Computing Platform Design for Autonomous Driving Systems
abstract
Autonomous driving is becoming a hot topic in both academic and industrial communities. Traditional algorithms can hardly achieve the complex tasks and meet the high safety criteria. Recent research on deep learning shows significant performance improvement over traditional algorithms and is believed to be a strong candidate in autonomous driving system. Despite the attractive performance, deep learning does not solve the problem totally. The application scenario requires that an autonomous driving system must work in real-time to keep safety. But the high computation complexity of neural network model, together with complicated pre-process and post-process, brings great challenges. System designers need to do dedicated optimizations to make a practical computing platform for autonomous driving. In this paper, we introduce our work on efficient computing platform design for autonomous driving systems. In the software level, we introduce neural network compression and hardware-aware architecture search to reduce the workload. In the hardware level, we propose customized hardware accelerators for pre- and post-process of deep learning algorithms. Finally, we introduce the hardware platform design, NOVA-30, and our on-vehicle evaluation project.
Shuang Liang 0010, Changcheng Tang, Xuefei Ning, Shulin Zeng, Yu Wang 0002, Kaiyuan Guo, Diange Yang, Huazhong Yang
ASP-DAC8
2021 LiDAR-based Object Detection Failure Tolerated Autonomous Driving Planning System
abstract
A typical autonomous driving system usually relies on the detected objects from an environment perception module. Current research still cannot guarantee a perfect perception, and failure detections may cause collisions, leading to untrustworthy autonomous vehicles. This work proposes a trajectory planner to tolerate the detection failure of the LiDAR sensors. This method will plan the path relying on the detected objects as well as the raw sensor data. The overlapping and contradiction of both perception routes will be carefully addressed for safe and efficient driving. The object detector in this work uses a deep learning-based method, i.e., CNN-Segmentation neural network. The designed trajectory planner has multi-layers to handle the multi-resolution environment formed by different perception routes. The final system will dynamically adjust its attention to the detected objects or the point cloud to avoid collision due to detection failures. This method is implemented on a real autonomous vehicle to drive in an open urban area. The results show that when the autonomous vehicle fails to detect a surrounding object, e.g., vehicles or some undefined objects, the autonomous vehicles still can plan an efficient and safe trajectory. In the meantime, when the perception system works well, the A V will not be affected by the point clouds. This technology can make the autonomous vehicle trustworthy even with the black-box neural networks. The codes are open-source with our autonomous driving platform to help other researchers for A V development.
Zhong Cao 0003, Weitao Zhou, Xinyu Jiao, Diange Yang
IV5
2021 UrbanPose: A New Benchmark for VRU Pose Estimation in Urban Traffic Scenes
abstract
Human pose, serving as a robust appearance-invariant mid-level feature, has proven to be effective and efficient for human action recognition and intention estimation. Pose features also have a great potential to improve trajectory prediction for the Vulnerable Road User (VRU) in ADAS or automated driving applications. However, the lack of highly diverse and large VRU pose datasets makes a transfer and application to the VRU rather difficult. This paper introduces the Tsinghua-Daimler Urban Pose dataset (TDUP), a large-scale 2D VRU pose image dataset collected in Chinese urban traffic environments from on-board a moving vehicle. The TDUP dataset contains 21k images with more than 90k high-quality, manually labeled VRU bounding boxes with pose keypoint annotations and additional tags. We optimize four state-of-the-art deep learning approaches (AlphaPose, Mask R-CNN, Pose-SSD and PitPaf) to serve as baselines for the new pose estimation benchmark. We further analyze the effect of using large pre-training datasets and different data proportions as well as optional labeled information during training. Our new benchmark is expected to lay the foundation for further VRU pose studies and to empower the development of accurate VRU trajectory prediction methods in complex urban traffic scenes. The dataset (including an evaluation server) is available on www.urbanpose-dataset.com for non-commercial scientific use.
Diange Yang, Baofeng Wang, Zijie Guo, Rishabh Verma, Jayanth Ramesh, Christoph Weinrich, Ulrich Kressel, Fabian Flohr
IV2
2021 Fast Initialization for Monocular Map Matching Localization via Multi-lane Hypotheses in Highway Scenarios
abstract
Many researchers have used the map-matching algorithm to leverage inadequate traditional vehicle localization with HD maps for fast and efficient vehicle localization. The initialization process in the map-matching algorithm has always been problematic due to the nature of GNSS error, which might exceed the average lane width of approximately 3.5m. Thus, directly using the GNSS data to obtain the rough initial pose may cause the failure of the initialization process. The general solution is to randomly sample around the GNSS data and test the initialization with these hypotheses. However, this often leads to expensive computation and thus fails to run in realtime or online mode, as a dense sampling is required to achieve an acceptable level of initial estimation accuracy. As a viable alternative, we propose a multi-lane hypotheses approach to narrow down the search by limiting the sample pose to the number of lanes within the GNSS data's error radius. From these poses, we then perform an efficient map-matching to refine these pose candidates. Moreover, a novel belief function to evaluate the hypothesis is proposed to select the best hypothesis for system initialization robustly. Our evaluation result shows that we have outperformed the primary random sampling method in both accuracy and efficiency.
Tuopu Wen, Benny Wijaya, Kun Jiang 0002, Dongfang Zheng, Yiliang Xu, Mengmeng Yang 0001, Diange Yang
IV8
2021 Highway Exiting Planner for Automated Vehicles Using Reinforcement Learning
abstract
Exiting from highways in crowded dynamic traffic is an important path planning task for autonomous vehicles (AVs). This task can be challenging because of the uncertain motion of surrounding vehicles and limited sensing/observing window. Conventional path planning methods usually compute a mandatory lane change (MLC) command, but the lane change behavior (e.g., vehicle speed and gap acceptance) should also adapt to traffic conditions and the urgency for exiting. In this paper, we propose a reinforcement learning-enhanced highway-exit planner. The learning-based strategy learns from past failures and adjusts the vehicle motion when the AV fails to exit. The reinforcement learning is based on the Monte Carlo tree search (MCTS) approach. The proposed learning-enhanced highway-exit planner is tested 6000 times in stochastic simulations. The results indicate that the proposed planner achieves a higher probability of successful highway exiting than a benchmark MLC planner.
Zhong Cao 0003, Diange Yang, Shaobing Xu, Huei Peng, Boqi Li 0001, Shuo Feng 0002, Ding Zhao
IEEE Trans. Intell. Transp. Syst.2
2021 Bridging the Gap of Lane Detection Performance Between Different Datasets: Unified Viewpoint Transformation
abstract
Convolutional neural networks (CNNs) have shown excellent performance for vision-based lane detection. However, maintaining the performance of the trained models under new test scenarios still remains challenging due to the dataset bias between the training and test datasets; In lane detection processes, the dataset bias can be categorized into lane position bias and lane pattern bias, with the former one particularly influences the lane detection performance. To tackle this dataset bias, this article proposes aunified viewpoint transformation (UVT)method that transforms the camera viewpoints of different datasets into a common virtual world coordinate system, such that the mismatched lane position distributions can be effectively aligned. Experiments are conducted on multiple datasets including the Caltech[1], Tusimple[2], and KITTI[3]dataset. The results demonstrate the effectiveness of the UVT algorithm in improving the lane detection performance on the test datasets. Moreover, by incorporating the UVT into other techniques that tackling the dataset bias, the lane position and pattern differences are disentangled and separately minimized. As a result, the performance gap between the training data and the test scenarios can be bridged. Specifically, the model trained on the KITTI dataset have achieved high performance in the Tusimple and the Caltech dataset (F1-score: 84.8 and 87.1%). With the proposed algorithm, a lane detection model trained on one dataset can be effectively applied to datasets with different camera settings in vastly different localities, and achieve better generalization ability compared to the state of the art methods.
Tuopu Wen, Diange Yang, Kun Jiang 0002, Chun-lei Yu, Benny Wijaya, Xinyu Jiao
IEEE Trans. Intell. Transp. Syst.2
2020 High Precision Vehicle Localization based on Tightly-coupled Visual Odometry and Vector HD Map
abstract
Matching low-cost camera and vector HD map is proven to be a practical and effective way of estimating the location and orientation of intelligent vehicles. However, map-based approach is viable only when the landmark observation is adequate and precise. In some areas with sparse and noisy observation, or even non-existent map matching features, the localization results may be unstable. In this paper, we introduce a novel algorithm by fusing visual odometry and vector HD map in a tightly-coupled optimization framework to tackle these problems. Our algorithm exploits the observation of visual feature points and vector HD map landmarks in the sliding window manner and optimize their residuals in a tightly-coupled approach. In this way, the system is more robust against the noisy HD map landmark observations. In addition, our method is able to accurately estimate vehicle pose even when landmarks are sparse. Experiments under two challenging scenarios with noisy and sparse landmark observations show that our method can achieve the Mean Absolute Error (MAE) at 0.1473m and 0.2496m respectively.
Tuopu Wen, Zhongyang Xiao, Benny Wijaya, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang
IV6
2020 CLAP: Cloud-and-Learning-compatible Autonomous driving Platform
abstract
Autonomous driving (AV) has been intensively researched over the last decade. In this paper, we introduce an open autonomous driving software stack to enable faster design, development, and testing of algorithms on simulated or experimental vehicles, which we hope will become a useful tool for AV researchers.
Yuanxin Zhong, Zhong Cao 0003, Minghan Zhu, Xinpeng Wang 0002, Diange Yang, Huei Peng
IV5
2020 Feedforward Compensation-Based Finite-Time Traffic Flow Controller for Intelligent Connected Vehicle Subject to Sudden Velocity Changes of Leading Vehicle
abstract
Optimal velocity (OV)-based car-following model can be easily applied to the transportation system composed of intelligent connected vehicle (ICV) via vehicle-to-vehicle (V2V) communication, since this model only requires space headway and velocity differences of preceding vehicles. However, the sudden velocity changes of the leading vehicle will decrease the control performance of following vehicles. Particularly, the farther the distance between the following vehicle and the leading vehicle is, the worse the control performance of the following vehicle is. Besides this problem, the nonlinear information of OV-based car-following models is often not fully utilized due to the linearization for the convenience of stability analysis and controller design. To address these problems, the factors related to velocity sudden changes are taken into account for each ICV simultaneously and compensated by the proposed feedforward compensator. Based on the compensator, a finite-time traffic flow controller is designed for each ICV to smooth the space headway in traffic flow subject to sudden velocity changes and make full use of the nonlinear information of OV-based model. Finally, the theoretical analysis via Lyapunov approach and the numerical simulations are carried out to verify the effectiveness of the proposed method.
Ruidong Yan, Diange Yang, Benny Wijaya, Chun-lei Yu
IEEE Trans. Intell. Transp. Syst.2
2019 Real-time Adaptive UWB Positioning System Enhanced by Sensor Fusion for Multiple Targets Detection
abstract
Ultra-wideband (UWB) as a state-of-the-art Real-Time Localization System (RTLS) has shown outstanding performance in tackling difficult positioning task. However, the implementation of this technology remains a challenge as several problems such as clock synchronization and line-of-sight (LOS) problem often occurs during integration. Moreover, when this technology faced with real-time multi targets detection, this technology still does not produce a stable result. This paper addresses these two problems by introduces a sound approach to tackle clock synchronization, LOS problem, and create a stable multi positioning system. We managed to secure 8.48 cm of RMS error for NLOS condition and 7.29 cm of RMS error for LOS condition. Besides, we also enhance the system by adding sensor fusion in order to create more effective multi targets localization in real-time condition. This enhancement derives from support by map information and speed sensor as a support system. Finally, this system is tested to support a realtime application of the model cars, and it can handle the task and obtains 11.27 cm of RMS error for dynamic positioning result.
Benny Wijaya, Nanshan Deng, Kun Jiang 0002, Ruidong Yan, Diange Yang
IV5
2019 High Precision Target Positioning Method for RSU in Cooperative Perception
abstract
Vehicle-road cooperative perception system can greatly improve the perception ability of intelligent vehicles by making use of perception information from road side units (RSU). This paper focuses on the target positioning of static camera for vehicle-road cooperation. A low-cost camera calibration method is proposed to complete the accurate mapping between the image plane and the 3D world space. Precise location of interested targets are achieved by an efficient tracking strategy. Real test scenarios show that our algorithm can effectively locate vehicles, pedestrians, non-motor vehicles and other targets with high accuracy. Our algorithm won the Monocular Static Camera Positioning and Ranging Competition for Autonomous Driving championship in 2019.
Tuopu Wen, Zhongyang Xiao, Kun Jiang 0002, Mengmeng Yang 0001, Keqiang Li 0002, Diange Yang
MMSP6
2013 A Study on the Method for Cleaning and Repairing the Probe Vehicle Data
abstract
Probe vehicle data are being increasingly applied in urban dynamic traffic data collection. However, the mobility and scale limit of probe vehicles may lead to incomplete or inaccurate data and thus influence the measurement of the state of traffic. At present, probe vehicle data are usually repaired by linear interpolation or a historical average method, but the repair accuracy is relatively low. To address the given problems, the multithreshold control repair method (MTCRM) was proposed to clean and repair the probe vehicle data. The MTCRM adopts threshold control and a rule based on the approximate normalization transform to clean abnormal traffic data and to fill in the missing data by a weighted average method and an exponential smoothing method. In this approach, we combine topological road network characteristics to fill in the missing data from data for neighboring road sections and repair noisy data by reconstructing the principal components. This paper mainly focuses on analyzing the component of the recurring pattern of probe vehicle data, which can provide guidelines for the subsequent traffic forecasts. The findings of data repair for different grades of road in Beijing, China, demonstrate that the mean repair error may meet the requirements of traffic-state measurement, demonstrating that MTCRM can effectively clean probe vehicle data.
Zhaosheng Zhang, Diange Yang, Tao Zhang 0061, Qiaochu He, Xiaomin Lian
IEEE Trans. Intell. Transp. Syst.2