EDBT 2026 Demo / reviewers in the wild / expert
Kun Jiang 0002
dblp:03/4406-2
· DBLP profile ↗
37ranked-venue papers
1as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LSRE: Latent Semantic Rule Encoding for Real-Time Semantic Risk Detection in Autonomous Driving
Weitao Zhou, Cheng Jing, Nanshan Deng, Junze Wen, Kun Jiang 0002, Diange Yang |
IV | 7 |
| 2026 | Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video GenerationabstractCorner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for generating diverse scenarios, commonly achieved through 3D Gaussian Splatting or image generative models. However, these approaches often suffer from limited visual fidelity or imprecise pose control. To address these issues, we propose G^2Editor, a framework designed for photorealistic and precise object editing in driving videos. Our method leverages a 3D Gaussian representation of the edited object as a dense prior, injected into the denoising process to ensure accurate pose control and spatial consistency. A scene-level 3D bounding box layout is employed to reconstruct occluded areas of non-target objects. Furthermore, to guide the appearance details of the edited object, we incorporate hierarchical fine-grained features as additional conditions during generation. Experiments on the Waymo Open Dataset demonstrate that G^2Editor effectively supports object repositioning, insertion, and deletion within a unified framework, outperforming existing methods in both pose controllability and visual quality, while also benefiting downstream data-driven tasks. Jiusi Li, Jackson Jiang, Jinyu Miao, Miao Long, Tuopu Wen, Peijin Jia, Shengxiang Liu, Chun-lei Yu, Maolin Liu, Yuzhan Cai, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 11 |
| 2026 | MaPLocator: Map-Prior Enhanced End-to-End Vehicle Localization with Coarse-to-Fine Pose Refinement
Muxi Tang, Jinyu Miao, Rujun Yan, Le Jia, Mengmeng Yang 0001, Diange Yang, Kun Jiang 0002 |
IV | 8 |
| 2026 | DTCCL: Disengagement-Triggerd Contrastive Continual Learning for Autonomous Bus Planners
Yanding Yang, Weitao Zhou, Xiaomin Guo, Junze Wen, Lang Ding, Zheng Fu, Jinyu Miao, Kun Jiang 0002, Diange Yang |
IV | 10 |
| 2026 | Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion PredictionabstractLiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | PriorMotion: Generative Class-Agnostic Motion Prediction with Raster-Vector Motion Field PriorsabstractReliable spatial and motion perception is essential for safe autonomous navigation. Recently, class-agnostic motion prediction on bird's-eye view (BEV) cell grids derived from LiDAR point clouds has gained significant attention. However, existing frameworks typically perform cell classification and motion prediction on a per-pixel basis, neglecting important motion field priors such as rigidity constraints, temporal consistency, and future interactions between agents. These limitations lead to degraded performance, particularly in sparse and distant regions. To address these challenges, we introduce \textbf{PriorMotion}, an innovative generative framework designed for class-agnostic motion prediction that integrates essential motion priors by modeling them as distributions within a structured latent space. Specifically, our method captures structured motion priors using raster-vector representations and employs a variational autoencoder with distinct dynamic and static components to learn future motion distributions in the latent space. Experiments on the nuScenes dataset demonstrate that \textbf{PriorMotion} outperforms state-of-the-art methods across both traditional metrics and our newly proposed evaluation criteria. Notably, we achieve improvements of approximately 15.24\% in accuracy for fast-moving objects, an 3.59\% increase in generalization, a reduction of 0.0163 in motion stability, and a 31.52\% reduction in prediction errors in distant regions. Further validation on FMCW LiDAR sensors confirms the robustness of our approach. Kangan Qian, Jinyu Miao, Xinyu Jiao, Ziang Luo, Zheng Fu, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Diange Yang |
ICCV | 8 |
| 2025 | Efficient End-to-end Visual Localization for Autonomous Driving with Decoupled BEV Neural MatchingabstractAccurate localization plays an important role in high-level autonomous driving systems. Conventional map matching-based localization methods solve the poses by explicitly matching map elements with sensor observations, generally sensitive to perception noise, therefore requiring costly hyperparameter tuning. In this paper, we propose an end-to-end localization neural network which directly estimates vehicle poses from surrounding images, without explicitly matching perception results with HD maps. To ensure efficiency and interpretability, a decoupled BEV neural matching-based pose solver is proposed, which estimates poses in a differentiable sampling-based matching module. Moreover, the sampling space is hugely reduced by decoupling the feature representation affected by each DoF of poses. The experimental results demonstrate that the proposed network is capable of performing decimeter level localization with mean absolute errors of 0.19m, 0.13m and 0.39° in longitudinal, lateral position and yaw angle while exhibiting a 68.8% reduction in inference memory usage. Jinyu Miao, Tuopu Wen, Ziang Luo, Kangan Qian, Zheng Fu, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Jin Huang 0002, Diange Yang |
IROS | 7 |
| 2025 | LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion PredictionabstractAccurate spatial and motion understanding is critical for autonomous driving systems. While object-level perception models excel in structured environments, they struggle with open-set categories and often lack precise geometric representation. Occupancy-based, class-agnostic methods offer better scene expressiveness but typically ignore inter-agent interactions and fail to ensure physical consistency in motion predictions, limiting their reliability in complex traffic scenarios. In this paper, we propose LEGO-Motion, a novel class-agnostic motion prediction framework that bridges the gap between instance-level reasoning and occupancy-based modeling. Unlike conventional grid-based methods that treat each cell independently, LEGO-Motion introduces two key components: (1) the Interaction-Augmented Instance Encoder (IaIE), which models interactions among dynamic agents via cross-attention, and (2) the Instance-Enhanced BEV Encoder (IeBE), which improves motion consistency across instances through multi-stage feature fusion. These components enable our model to learn semantically coherent and physically plausible motion fields. Extensive experiments on the nuScenes dataset show that LEGO-Motion achieves a around 6% improvement in motion prediction accuracy over the previous state-of-the-art, while maintaining real-time inference at 21ms. Moreover, our method demonstrates strong generalization on a proprietary FMCW LiDAR benchmark. These results validate LEGO-Motion's effectiveness in capturing both global scene structure and fine-grained motion dynamics, making it a promising foundation for next-generation perception systems. Kangan Qian, Jinyu Miao, Ziang Luo, Zheng Fu, Jinchen Li, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IROS | 8 |
| 2025 | EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Drivingabstract3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or transformers at the voxel level. Moreover, OccNets are supervised with expensive large-scale dense voxel labels. Model and label inefficiencies, caused by excessive network parameters and label annotation requirements, severely hinder the onboard deployment of OccNets. This paper proposes an EFFicient Occupancy learning framework, EFFOcc, that targets minimal network complexity and label requirements while achieving state-of-the-art accuracy. We first propose an efficient fusion-based OccNet that only uses simple 2D operators and improves accuracy to the state-of-the-art on three large-scale benchmarks: Occ3D-nuScenes, Occ3D-Waymo, and OpenOccupancy-nuScenes. On the Occ3D-nuScenes benchmark, the fusion-based model with ResNet-18 as the image backbone has 21.35M parameters and achieves 51.49 in terms of mean Intersection over Union (mIoU). Furthermore, we propose a multi-stage occupancy-oriented distillation to efficiently transfer knowledge to vision-only OccNet. Extensive experiments on occupancy benchmarks show state-of-the-art precision for both fusion-based and vision-based OccNets. For the demonstration of learning with limited labels, we achieve 94.38% of the performance (mIoU = 28.38) of a 100% labeled vision OccNet (mIoU = 30.07) using the same OccNet trained with only 40% labeled sequences and distillation from the fusion-based OccNet. Code is available at https://github.com/synsin0/EFFOcc. Yining Shi 0002, Kun Jiang 0002, Jinyu Miao, Ke Wang 0021, Kangan Qian, Yunlong Wang 0009, Jiusi Li, Tuopu Wen, Mengmeng Yang 0001, Yiliang Xu, Diange Yang |
IROS | 2 |
| 2025 | A Benchmark for Vision-Centric HD Mapping by V2I SystemsabstractAutonomous driving faces safety challenges due to a lack of global perspective and the semantic information of vectorized high-definition (HD) maps. Information from roadside cameras can greatly expand the map perception range through vehicle-to-infrastructure (V2I) communications. However, there is still no dataset from the real world available for the study on map vectorization onboard under the scenario of vehicle-infrastructure cooperation. To prosper the research on online HD mapping for Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD), we release a real-world dataset, which contains collaborative camera frames from both vehicles and roadside infrastructures, and provides human annotations of HD map elements. We also present an end-to-end neural framework (i.e., V2I-HD) leveraging vision-centric V2I systems to construct vectorized maps. To reduce computation costs and further deploy V2I-HD on autonomous vehicles, we introduce a directionally decoupled self-attention mechanism to V2I-HD. Extensive experiments show that V2I-HD has superior performance in real-time inference speed, as tested by our real-world dataset. Abundant qualitative results also demonstrate stable and robust map construction quality with low cost in complex and various driving scenes. As a benchmark, both source codes and the dataset have been released at OneDrive11https://ldrv.ms/f/c/76645c25a8914a0b/EgWy5XCUk6pKgvE9vB-HbVEBCdCQjJvgxlKKjeKF7hPdZw for the purpose of further study. Shengtong Xu, Kun Jiang 0002, Haoyi Xiong, Xiangzeng Liu |
IV | 4 |
| 2025 | COME: Adding Scene-Centric Forecasting Control to Occupancy World ModelabstractWorld models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data.
Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to separate environmental changes from ego-motion by leveraging the scene-centric coordinate systems. In this paper, we introduce COME: a framework that integrates scene-centric forecasting Control into the Occupancy world ModEl. Specifically, COME first generates ego-irrelevant, spatially consistent future features through a scene-centric prediction branch, which are then converted into scene condition using a tailored ControlNet. These condition features are subsequently injected into the occupancy world model, enabling more accurate and controllable future occupancy predictions. Experimental results on the nuScenes-Occ3D dataset show that COME achieves consistent and significant improvements over state-of-the-art (SOTA) methods across diverse configurations, including different input sources (ground-truth, camera-based, fusion-based occupancy) and prediction horizons (3s and 8s). For example, under the same settings, COME achieves 26.3% better mIoU metric than DOME and 23.7% better mIoU metric than UniScene. These results highlight the efficacy of disentangled representation learning in enhancing spatio-temporal prediction fidelity for world models. Code is available at https://github.com/synsin0/COME. Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Tuopu Wen, Mengmeng Yang 0001, Diange Yang |
NeurIPS | 2 |
| 2025 | Pedestrian Trajectory Prediction for Autonomous Vehicles With Multiple InteractionsabstractPedestrian trajectory prediction is significant for autonomous vehicles, but the difficulty of pedestrian trajectory prediction lies in the accurate modeling of pedestrian multiple interactions. In this paper, we attempt to explore the essential features of pedestrian interaction and propose a pedestrian trajectory prediction method based on multiple interactions. Firstly, considering that the interaction between self-driving cars and pedestrians resembles a dynamic game process involving sequential adaptation, we map them to the same feature space and design a temporal cross-attention mechanism to model the interaction between pedestrians and vehicles. Meanwhile, pedestrian-scene interaction is affected by the global environment as well as the local environment. To capture the global information while preserving the spatial location of pedestrians in the scene, we design a pedestrian-scene heatmap fusion (PSHF) framework to model the pedestrian-scene interaction features. We validate the effectiveness of our algorithm on the publicly available JAAD and PIE datasets, achieving better performance than existing representative methods in both single-trajectory and multi-trajectory prediction tasks. We conducted a thorough ablation study, cross-dataset validation, and qualitative visualization experiments, demonstrating the effectiveness and robustness of our method. Zheng Fu, Mengmeng Yang 0001, Kun Jiang 0002, Jin Huang 0002, Hao Gao 0005, Diange Yang |
IEEE Internet Things J. | 4 |
| 2025 | Top-Down Attention-Based Mechanisms for Interpretable Autonomous DrivingabstractDespite the remarkable advancements in autonomous driving, the challenge persists in achieving interpretable action decision-making, primarily owing to the intricate and ambiguous relationship between detected agents and driving intention. In this study, we introduce an interpretable action prediction model, denoted as the Prediction-Driven Attention Network (PDANet), designed to undertake action decisions and provide corresponding interpretations cohesively. The PDANet is inspired by the perceptual mechanisms inherent in human drivers, who allocate attention according to their driving intentions. Specifically, we elaborate a prediction module to generate vehicle prospective trajectories to characterize driving intentions. Subsequently, the features of this predicted trajectory are utilized to modulate the attention distribution among agents through the top-down attention module, yielding an attention map. Finally, two distinct task tokens are applied to aggregate agent features and generate the final output according to the derived attention map. Extensive experiments conducted on the publicly available BDD-OIA and nu-AR datasets demonstrate that our proposed method outperforms all prior works in terms of both action prediction and behavior interpretation tasks. Remarkably, our method attains a noteworthy enhancement in the behavior interpretation task, surpassing the previous state-of-the-art by a substantial margin of +10.8% in terms of F1-score on the nu-AR dataset. We also validate our algorithm on Carla Town05 long in a closed-loop decision-making scenario, highlighting the generality and robustness of our approach. Furthermore, qualitative results show that the agents selected by our model are more closely aligned with human cognitive processes. Zheng Fu, Kun Jiang 0002, Yunlong Wang 0009, Tuopu Wen, Hao Gao 0005, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Toward Democratizing High-Definition Map Update Through Consortium BlockchainabstractIn the rapidly evolving landscape of autonomous vehicles and advanced navigation systems, the accuracy of high-definition maps and real-time updating has become paramount. However, in this progression, the security of map data has not received adequate attention, although the accuracy of the data can be easily altered when the system is breached. Thus, this paper introduces a novel approach to democratizing the process of high-definition map updates by leveraging consortium blockchain technology specifically designed for Proof of Presence and Reputation (POP-R) to safeguard the update process. Our proposed system leverages the presence and reputation of vehicles through infrastructure nodes to enhance the accuracy and reliability of HD map updates. We created a trusted ecosystem for maintaining high-definition maps, marked by a superior safety score across three scenarios compared to the standard proof of reputation technique. Additionally, it demonstrates high efficiency, achieving 12,000 transactions per second (TPS) for data queries and more than 2,500 TPS for data writing in our blockchain network. This efficiency proved our prowess in the lightweight computational power required, suitable for decentralized and crowdsourced-based systems. Through our POP-R framework, we lay the foundation for a new decentralized approach to the evolution of high-definition maps in the era of autonomous mobility. Benny Wijaya, Mengmeng Yang 0001, Tuopu Wen, Kun Jiang 0002, Wei Zhang 0090, Yunlong Wang 0009, Zheng Fu, Xuewei Tang, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Grid-Centric Traffic Scenario Perception for Autonomous Driving: A Comprehensive ReviewabstractThe grid-centric perception is a crucial field for mobile robot perception and navigation. Nonetheless, the grid-centric perception is less prevalent than object-centric perception as autonomous vehicles need to accurately perceive highly dynamic, large-scale traffic scenarios, and the complexity and computational costs of grid-centric perception are high. In recent years, the rapid development of deep learning techniques and hardware provides fresh insights into the evolution of grid-centric perception. The fundamental difference between grid-centric and object-centric pipeline lies in that grid-centric perception follows a geometry-first paradigm which is more robust to the open-world driving scenarios with endless long-tailed semantically unknown obstacles. Recent research demonstrates the great advantages of grid-centric perception, such as comprehensive fine-grained environmental representation, greater robustness to occlusion and irregular-shaped objects, better ground estimation, and safer planning policies. There is also a growing trend that the capacity of occupancy networks is greatly expanded to 4-D scene perception and prediction, and the latest techniques are highly related to new research topics, such as 4-D occupancy forecasting, generative artificial intelligence (GenAI), and world models in the field of autonomous driving. Given the lack of current surveys for this rapidly expanding field, we present a hierarchically structured review of grid-centric perception for autonomous vehicles. We organize previous and current knowledge of occupancy grid techniques along the main vein from 2-D bird-eye view (BEV) grids to 3-D occupancy to 4-D occupancy forecasting. We additionally summarize label-efficient occupancy learning and the role of grid-centric perception in driving systems. Finally, we present a summary of the current research trend and provide future outlooks. Yining Shi 0002, Kun Jiang 0002, Jiusi Li, Zelin Qian, Junze Wen, Mengmeng Yang 0001, Ke Wang 0021, Diange Yang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous DrivingabstractVision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect obstacles regardless of their shape and occlusion. Modern occupancy networks mainly focus on reconstructing visible voxels from object surfaces with voxel-wise semantic prediction. Usually, they suffer from inconsistent predictions of one object and mixed predictions for adjacent objects. These confusions may harm the safety of downstream planning modules. To this end, we investigate panoptic segmentation on 3D voxel scenarios and propose an instance-aware occupancy network, PanoSSC. We predict foreground objects and backgrounds separately and merge both in post-processing. For foreground instance grouping, we propose a novel 3D instance mask decoder that can efficiently extract individual objects. we unify geometric reconstruction, 3D semantic segmentation, and 3D instance segmentation into PanoSSC framework and propose new metrics for evaluating panoptic voxels. Extensive experiments show that our method achieves competitive results on SemanticKITTI semantic scene completion benchmark. Yining Shi 0002, Jiusi Li, Kun Jiang 0002, Ke Wang 0021, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang |
3DV | 3 |
| 2024 | StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential EquationabstractPredicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However, current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict uniform snapshots of future occupancy states and require strictly synchronized sensory data for sensor fusion. We propose a novel framework, StreamingFlow, to lift these strong limitations. StreamingFlow is a novel BEV occupancy predictor that ingests asynchronous multi-sensor data streams for fusion and performs streaming fore-casting of the future occupancy map at any future times-tamps. By integrating neural ordinary differential equations (N-ODE) into recurrent neural networks, StreamingFlow learns derivatives of BEV features over temporal horizons, updates the implicit sensor's BEV features as part of the fusion process, and propagates BEV states to the desired future time point. It shows good zero-shot generalization ability of prediction, reflected in the interpolation of the ob-served prediction time horizon and the reasonable inference of the unseen farther future period. Extensive experiments on two large-scale datasets, nuScenes [2] and Lyft L5 [14], demonstrate that StreamingFlow significantly outperforms previous vision-based, LiDAR-based methods, and shows superior performance compared to state-of-the-art fusion-based methods. Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Jiusi Li, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang |
CVPR | 2 |
| 2024 | LaneSegNet: Map Learning with Lane Segment Perception for Autonomous DrivingabstractA map, as crucial information for downstream applications of an autonomous driving system, is usually represented in lanelines or centerlines. However, existing literature on map learning primarily focuses on either detecting geometry-based lanelines or perceiving topology relationships of centerlines. Both of these methods ignore the intrinsic relationship of lanelines and centerlines, that lanelines bind centerlines. While simply predicting both types of lane in one model is mutually excluded in learning objective, we advocate lane segment as a new representation that seamlessly incorporates both geometry and topology information. Thus, we introduce LaneSegNet, the first end-to-end mapping network generating lane segments to obtain a complete representation of the road structure. Our algorithm features two key modifications. One is a lane attention module to capture pivotal region details within the long-range feature space. Another is an identical initialization strategy for reference points, which enhances the learning of positional priors for lane attention. On the OpenLane-V2 dataset, LaneSegNet outperforms previous counterparts by a substantial gain across three tasks, i.e., map element detection (+4.8 mAP), centerline perception (+6.9 DET$_l$), and the newly defined one, lane segment perception (+5.6 mAP). Furthermore, it obtains a real-time inference speed of 14.7 FPS. Code is accessible at https://github.com/OpenDriveLab/LaneSegNet. Tianyu Li 0004, Peijin Jia, Bangjun Wang, Li Chen 0008, Kun Jiang 0002, Junchi Yan, Hongyang Li 0001 |
ICLR | 5 |
| 2024 | LaneDAG: Automatic HD Map Topology Generator Based on Geometry and Attention Fusion MechanismabstractIn high-definition maps (HD maps), the road lane centerline and lane topology graph play essential roles in navigation, planning, and decision-making. Existing research focusing on extracting physical infrastructure, such as lane boundaries, has made significant progress. But lane centerline detection and topology reasoning still remains challenging due to the severe overlapping centerlines and complicated topology. To tackle these challenges, we introduce an automatic lane topology extraction method for HD maps, termed LaneDAG, which extracts vectorized centerlines and their topology from prebuilt lane lines and road boundaries in HD maps. It formulates centerline extraction as a set prediction problem and lane topology prediction as a directed acyclic graph (DAG) construction problem. A novel mechanism that fusing geometric and attention-based features in the DAG is proposed to model the topological relationship between centerlines. Experiments conducted on the Argoverse 2 dataset demonstrate the proposed method’s superior performance compared to existing methods, showcasing its capability to extract lane centerlines and topology in HD maps automatically. Peijin Jia, Tuopu Wen, Ziang Luo, Zheng Fu, Jiaqi Liao, Huixian Chen, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 7 |
| 2023 | Traffic Police 3D Gesture Recognition Based on Spatial-Temporal Fully Adaptive Graph Convolutional NetworkabstractIt is critical for autonomous vehicles to recognize traffic police gestures timely and accurately. During the movement of the vehicle, the collected traffic police scales change all the time, in addition, the frequency and amplitude of actions of different traffic police are different. First, we use gesture normalization to fix the traffic police actions at a unified scale and remove the influence of scale changes on traffic police gesture recognition. Meanwhile, a fully adaptive spatial-temporal graph convolution network (FA-STGCN) is proposed to recognize the actions with different amplitude and frequencies. The adaptive spatial graph network can dig the latent joints connection relation of the traffic police under different gestures, which weakens the amplitude impact on the action recognition. The adaptive temporal graph network is composed of the global temporal module and the local temporal module. The global temporal module can obtain the coarse-grained features of the traffic police gestures’ speed and then naturally use the coarse-grained features to guide the local temporal module to adaptively learn the fine-grained temporal features of the traffic police action. The adaptive spatial graph network and the temporal graph network are alternately stacked to finally output accurate traffic police gestures. We thoroughly evaluated our method through intensive experiments, the result shows that our method achieved the best results on public datasets. What’s more, we proofed the effectiveness of each module and verified our methods for moving vehicles for the first time, the performance present meets the vehicle’s practical requirements. Zheng Fu, Kun Jiang 0002, Junze Wen, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Reliable Autonomous Driving Environment Model With Unified State-Extended BoundaryabstractFrom the early stage of robotic applications to current autonomous driving technologies, environment modeling has been acting as the middleware for connecting perception and decision layers. In robotic applications, space-oriented models (e.g., grid map, drivable area) are widely applied to faithfully reflect the space occupation. With the development of autonomous driving, highly dynamic and complex road environment brings rising need to understand the type and motion status of objects, thus element list has became the mainstream environment model. However, along comes the reliablity problem caused by missed detection and irregular objects, which is still inevitable despite the detection accuracy improvement. In view of this, a new view of driving environment is proposed as the unified state-extended boundary (USEB), aiming to improve the reliablity of element-oriented model. For driving decision requirements, different types of elements are consistently converted into driving constraints. Semantics and dynamics are expressed as the status of drivable area boundary, making it possible to merge space occupation to improve reliability against missed detection and irregular objects. Evaluation of USEB is carried out on the nuScenes dataset. Comparative results show that the proposed USEB could cover the required information for driving decision, whereas achieving higher reliability than the commonly applied element-oriented model. Xinyu Jiao, Kun Jiang 0002, Yunlong Wang 0009, Zhong Cao 0003, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Identify, Estimate and Bound the Uncertainty of Reinforcement Learning for Autonomous DrivingabstractDeep reinforcement learning (DRL) has emerged as a promising approach for developing more intelligent autonomous vehicles (AVs). A typical DRL application on AVs is to train a neural network-based driving policy. However, the black-box nature of neural networks can result in unpredictable decision failures, making such AVs unreliable. To this end, this work proposes a method to identify and protect unreliable decisions of a DRL driving policy. The basic idea is to estimate and constrain the policy’s performance uncertainty, which quantifies potential performance drop due to insufficient training data or network fitting errors. By constraining the uncertainty, the DRL model’s performance is always greater than that of a baseline policy. The uncertainty caused by insufficient data is estimated by the bootstrapped method. Then, the uncertainty caused by the network fitting error is estimated using an ensemble network. Finally, a baseline policy is added as the performance lower bound to avoid potential decision failures. The overall framework is called uncertainty-bound reinforcement learning (UBRL). The proposed UBRL is evaluated on DRL policies with different amounts of training data, taking an unprotected left-turn driving case as an example. The result shows that the UBRL method can identify potentially unreliable decisions of DRL policy. The UBRL guarantees to outperform baseline policy even when the DRL policy is not well-trained and has high uncertainty. Meanwhile, the performance of UBRL improves with more training data. Such a method is valuable for the DRL application on real-road driving and provides a metric to evaluate a DRL policy. Weitao Zhou, Zhong Cao 0003, Nanshan Deng, Kun Jiang 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Dynamically Conservative Self-Driving Planner for Long-Tail CasesabstractSelf-driving vehicles (SDVs) are becoming reality but still suffer from “long-tail” challenges during natural driving: the SDVs will continually encounter rare, safety-critical cases that may not be included in the dataset they were trained. Some safety-assurance planners solve this problem by being conservative in all possible cases, which may significantly affect driving mobility. To this end, this work proposes a method to automatically adjust the conservative level according to each case’s “long-tail” rate, named dynamically conservative planner (DCP). We first define the “long-tail” rate as an SDV’s confidence to pass a driving case. The rate indicates the probability of safe-critical events and is estimated using the statistics bootstrapped method with historical data. Then, a reinforcement learning-based planner is designed to contain candidate policies with different conservative levels. The final policy is optimized based on the estimated “long-tail” rate. In this way, the DCP is designed to automatically adjust to be more conservative in low-confidence “long-tail” cases while keeping efficient otherwise. The DCP is evaluated in the CARLA simulator using driving cases with “long-tail” distributed training data. The results show that the DCP can accurately estimate the “long-tail” rate to identify potential risks. Based on the rate, the DCP automatically avoids potential collisions in “long-tail” cases using conservative decisions while not affecting the average velocity in other typical cases. Thus, the DCP is safer and more efficient than the baselines with fixed conservative levels, e.g., an always conservative planner. This work provides a technique to guarantee SDV’s performance in unexpected driving cases without resorting to a global conservative setting, which contributes to solving the “long-tail” problem practically. Weitao Zhou, Zhong Cao 0003, Nanshan Deng, Kun Jiang 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | BE-STI: Spatial-Temporal Integrated Network for Class-agnostic Motion Prediction with Bidirectional EnhancementabstractDetermining the motion behavior of inexhaustible categories of traffic participants is critical for autonomous driving. In recent years, there has been a rising concern in performing class-agnostic motion prediction directly from the captured sensor data, like LiDAR point clouds or the combination of point clouds and images. Current motion prediction frameworks tend to perform joint semantic segmentation and motion prediction and face the trade-off between the performance of these two tasks. In this paper, we propose a novel Spatial-Temporal Integrated network with Bidirectional Enhancement, BE-STI, to improve the temporal motion prediction performance by spatial semantic features, which points out an efficient way to combine semantic segmentation and motion prediction. Specifically, we propose to enhance the spatial features of each individual point cloud with the similarity among temporal neighboring frames and enhance the global temporal features with the spatial difference among non-adjacent frames in a coarse-to-fine fashion. Extensive experiments on nuScenes and Waymo Open Dataset show that our proposed framework outperforms all state-of-the-art LiDAR-based and RGB+LiDAR-based methods with remarkable margins by using only point clouds as input.11The code will be released at https://github.com/be-sti/be-sti. Yunlong Wang 0009, Hongyu Pan, Yu-Huan Wu, Xin Zhan, Kun Jiang 0002, Diange Yang |
CVPR | 6 |
| 2022 | Skeleton-based traffic command recognition at road intersections for intelligent vehicles
Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Tuopu Wen, Diange Yang |
Neurocomputing | 2 |
| 2022 | Temporal Point Cloud Fusion With Scene Flow for Robust 3D Object TrackingabstractNon-visual range sensors such as Lidar have shown the potential to detect, locate and track objects in complex dynamic scenes thanks to their higher stability in comparison with vision-based sensors like cameras. However, due to the disorder, sparsity, and irregularity of the point cloud, it is much more challenging to take advantage of the temporal information in the dynamic 3D point cloud sequences, as it has been done in the image sequences for improving detection and tracking. In this paper, we propose a novel scene-flow-based point cloud feature fusion module to tackle this challenge, based on which a 3D object tracking framework is also achieved to exploit the temporal motion information. Moreover, we carefully designed several training schemes that contribute to the success of this new module by eliminating the issues of overfitting and long-tailed distribution of object categories. Extensive experiments on the public KITTI 3D object tracking dataset demonstrate the effectiveness of the proposed method by achieving superior results to the baselines. The source code is available athttps://github.com/Tsinghua-OpenICV/SharingVan-OpenPCDet. Yanding Yang, Kun Jiang 0002, Diange Yang, Yanqin Jiang |
IEEE Signal Process. Lett. | 2 |
| 2022 | Simple But Effective: Upper-Body Geometric Features for Traffic Command Gesture RecognitionabstractRecognizing traffic command gestures with high accuracy and quick response at a low computational cost is a requisite for driver assistance or autonomous driving. However, it has been understudied for a long time. Existing research takes advantage of increasing development in human action recognition but pays little attention to onboard conditions. In this article, we propose a simple but effective recognition model based on human upper-body geometric features and a long short-term memory (LSTM) network. The handcrafted geometric features can easily be calculated with estimated 2-D human keypoints at a low computational cost but are discriminative and sufficient in classification. Offline and online inferences are implemented to comprehensively evaluate the proposed model. For the sake of robustness required in the automotive domain, dual voting is designed to filter the output in online inference. On the recently published Chinese traffic police gesture (CTPG) dataset, the presented approach is the best with a remarkable improvement of approximately 8% compared to previous LSTM-based methods with handcrafted spatial features and is competitive with advanced GCN-based deep learning methods. The tradeoff pattern is explored to demonstrate how accuracy and response time alter with different training and inference strategies so that a balanced setup can be manually chosen under various application scenarios. Field tests are also carried out with an experimental vehicle, and the results uncover the present gap between research and practical application to some extent, moving a step closer to real-life traffic command gesture recognition. Kun Jiang 0002, Mengmeng Yang 0001, Zheng Fu, Diange Yang |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2022 | A General Autonomous Driving Planner Adaptive to Scenario CharacteristicsabstractAutonomous vehicle requires a general planner for all possible scenarios. Existing researches design such a planner by a unified scenario description. However, it may significantly increase the planner complexity even in some simple tasks, e.g., car following, further resulting in unsatisfactory driving performance. This work aims to design a general planner which can 1) drive in all possible scenarios and 2) have lower complexity in some common scenarios. To this end, this work proposes a pertinent boundary for multi-scenario driving planning. The total approach is named as Pertinent Boundary-based Unified Decision system. Based on the original drivable area, the pertinent boundary can further support motion status and semantics of the traffic elements, which provides the potential of pertinent performance for given scenarios. The pertinent boundary can support unified driving with the drivable area, in the meantime, can be pertinently modified to support the pertinent driving decisions for identified driving scenarios (e.g., car-following, junction left turning). It will further avoid the bump between the connections of the scenarios due to the continuity of space boundary. Thus, the planner is suitable for the fully autonomous driving. The proposed method is validated in different classical driving decision scenarios. Results show that the proposed method can support pertinent driving decisions in identified scenarios, in the meantime, assure generalized cross-scenario planning when no scenario information is available. Such a method shed light on fully autonomous driving by pertinence improvement of multi-scenario decision in the complex real world. Xinyu Jiao, Zhong Cao 0003, Kun Jiang 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Interactive Trajectory Prediction Using a Driving Risk Map-Integrated Deep Learning Method for Surrounding Vehicles on HighwaysabstractAccurate trajectory prediction of surrounding vehicles is vital for automated vehicles to achieve high-level driving safety in complex situations. However, most state-of-the-art approaches for multi-vehicle trajectory prediction ignore vehicle motion uncertainty caused by different driving styles. Moreover, the interrelationship between the vehicle and the environment is seldom considered. To address the above problems, this paper proposes a driving risk map-integrated deep learning (DRM-DL) method for interactive trajectory prediction of surrounding vehicles, which comprehensively considers the motion uncertainty, trajectory intention uncertainty and interactions among vehicles, lane lines and road boundaries. Specifically, we adopt a conditional variational autoencoder (CVAE) to generate the candidate trajectories, in which the motion uncertainty is considered using a conditional Gaussian distribution. Furthermore, a driving risk map is constructed to realize a unified and interpretable representation of vehicle-vehicle and vehicle-environment interactions. The probability of each candidate trajectory is assigned using a trajectory probability model and a random selection is adopted to select a guided trajectory, which simulates the driver’s trajectory intention uncertainty. Finally, a relearning module is designed to obtain the precise trajectory prediction for surrounding vehicles. The proposed method is evaluated on the HighD dataset, and the results demonstrate a more accurate and reliable trajectory prediction for surrounding vehicles compared with state-of-the-art methods. Xulei Liu, Yafei Wang 0001, Kun Jiang 0002, Zhisong Zhou, Kanghyun Nam, Chengliang Yin |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | TM3Loc: Tightly-Coupled Monocular Map Matching for High Precision Vehicle LocalizationabstractVision-based map-matching with HD map for high precision vehicle localization has gained great attention for its low-cost and ease of deployment. However, its localization performance is still unsatisfactory in accuracy and robustness in numerous real applications due to the sparsity and noise of the perceived HD map landmarks. This article proposes the tightly-coupled monocular map-matching localization algorithm (TM3Loc) for monocular-based vehicle localization. TM3Loc introduces semantic chamfer matching (SCM) to model monocular map-matching problem and combines visual features with SCM in a tightly-coupled manner. By applying the sliding window-based optimization technique, the historical visual features and HD map constraints are also introduced, such that the vehicle poses are estimated with an abundance of visual features and multi-frame HD map landmark features, rather than with single-frame HD map observations in previous works. Experiments are conducted on large scale dataset of 15 km long in total. The results show that TM3Loc is able to achieve high precision localization performance using a low-cost monocular camera, largely exceeding the performance of the previous state-of-the-art methods, thereby promoting the development of autonomous driving. Tuopu Wen, Kun Jiang 0002, Benny Wijaya, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Distributed Car-Following Control for Intelligent Connected Vehicle Using Improved Super-Twisting Compensator Subject to Sudden Velocity Changes of Leading VehicleabstractThe optimal velocity-based model has been successfully applied to distributed car-following systems. However, the car-following performance is inevitably affected by a series of disturbances, particularly, sudden velocity changes of leading vehicle. To improve accuracy and response rate of car-following control in the presence of such disturbances, an improved super-twisting compensator (ISTC) is proposed and a composite controller is designed by combining ISTC with a finite- time controller. A second-order nominal system is constructed by using a virtual measurement signal along with its integration to facilitate the design of ISTC. By introducing the feedback of high-order estimation error, the accuracy and response rate of ISTC are increased significantly as compared with the conventional one under same gains. Such improvement further enhances the disturbance rejection ability of the composite controller. Both Lyapunov approach and numerical simulations are carried out to verify the effectiveness of the proposed method. Ruidong Yan, Diange Yang, Jin Huang 0002, Kun Jiang 0002, Xinyu Jiao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Fast Initialization for Monocular Map Matching Localization via Multi-lane Hypotheses in Highway ScenariosabstractMany researchers have used the map-matching algorithm to leverage inadequate traditional vehicle localization with HD maps for fast and efficient vehicle localization. The initialization process in the map-matching algorithm has always been problematic due to the nature of GNSS error, which might exceed the average lane width of approximately 3.5m. Thus, directly using the GNSS data to obtain the rough initial pose may cause the failure of the initialization process. The general solution is to randomly sample around the GNSS data and test the initialization with these hypotheses. However, this often leads to expensive computation and thus fails to run in realtime or online mode, as a dense sampling is required to achieve an acceptable level of initial estimation accuracy. As a viable alternative, we propose a multi-lane hypotheses approach to narrow down the search by limiting the sample pose to the number of lanes within the GNSS data's error radius. From these poses, we then perform an efficient map-matching to refine these pose candidates. Moreover, a novel belief function to evaluate the hypothesis is proposed to select the best hypothesis for system initialization robustly. Our evaluation result shows that we have outperformed the primary random sampling method in both accuracy and efficiency. Tuopu Wen, Benny Wijaya, Kun Jiang 0002, Dongfang Zheng, Yiliang Xu, Mengmeng Yang 0001, Diange Yang |
IV | 3 |
| 2021 | Bridging the Gap of Lane Detection Performance Between Different Datasets: Unified Viewpoint TransformationabstractConvolutional neural networks (CNNs) have shown excellent performance for vision-based lane detection. However, maintaining the performance of the trained models under new test scenarios still remains challenging due to the dataset bias between the training and test datasets; In lane detection processes, the dataset bias can be categorized into lane position bias and lane pattern bias, with the former one particularly influences the lane detection performance. To tackle this dataset bias, this article proposes aunified viewpoint transformation (UVT)method that transforms the camera viewpoints of different datasets into a common virtual world coordinate system, such that the mismatched lane position distributions can be effectively aligned. Experiments are conducted on multiple datasets including the Caltech[1], Tusimple[2], and KITTI[3]dataset. The results demonstrate the effectiveness of the UVT algorithm in improving the lane detection performance on the test datasets. Moreover, by incorporating the UVT into other techniques that tackling the dataset bias, the lane position and pattern differences are disentangled and separately minimized. As a result, the performance gap between the training data and the test scenarios can be bridged. Specifically, the model trained on the KITTI dataset have achieved high performance in the Tusimple and the Caltech dataset (F1-score: 84.8 and 87.1%). With the proposed algorithm, a lane detection model trained on one dataset can be effectively applied to datasets with different camera settings in vastly different localities, and achieve better generalization ability compared to the state of the art methods. Tuopu Wen, Diange Yang, Kun Jiang 0002, Chun-lei Yu, Benny Wijaya, Xinyu Jiao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | High Precision Vehicle Localization based on Tightly-coupled Visual Odometry and Vector HD MapabstractMatching low-cost camera and vector HD map is proven to be a practical and effective way of estimating the location and orientation of intelligent vehicles. However, map-based approach is viable only when the landmark observation is adequate and precise. In some areas with sparse and noisy observation, or even non-existent map matching features, the localization results may be unstable. In this paper, we introduce a novel algorithm by fusing visual odometry and vector HD map in a tightly-coupled optimization framework to tackle these problems. Our algorithm exploits the observation of visual feature points and vector HD map landmarks in the sliding window manner and optimize their residuals in a tightly-coupled approach. In this way, the system is more robust against the noisy HD map landmark observations. In addition, our method is able to accurately estimate vehicle pose even when landmarks are sparse. Experiments under two challenging scenarios with noisy and sparse landmark observations show that our method can achieve the Mean Absolute Error (MAE) at 0.1473m and 0.2496m respectively. Tuopu Wen, Zhongyang Xiao, Benny Wijaya, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 4 |
| 2019 | Real-time Adaptive UWB Positioning System Enhanced by Sensor Fusion for Multiple Targets DetectionabstractUltra-wideband (UWB) as a state-of-the-art Real-Time Localization System (RTLS) has shown outstanding performance in tackling difficult positioning task. However, the implementation of this technology remains a challenge as several problems such as clock synchronization and line-of-sight (LOS) problem often occurs during integration. Moreover, when this technology faced with real-time multi targets detection, this technology still does not produce a stable result. This paper addresses these two problems by introduces a sound approach to tackle clock synchronization, LOS problem, and create a stable multi positioning system. We managed to secure 8.48 cm of RMS error for NLOS condition and 7.29 cm of RMS error for LOS condition. Besides, we also enhance the system by adding sensor fusion in order to create more effective multi targets localization in real-time condition. This enhancement derives from support by map information and speed sensor as a support system. Finally, this system is tested to support a realtime application of the model cars, and it can handle the task and obtains 11.27 cm of RMS error for dynamic positioning result. Benny Wijaya, Nanshan Deng, Kun Jiang 0002, Ruidong Yan, Diange Yang |
IV | 3 |
| 2019 | High Precision Target Positioning Method for RSU in Cooperative PerceptionabstractVehicle-road cooperative perception system can greatly improve the perception ability of intelligent vehicles by making use of perception information from road side units (RSU). This paper focuses on the target positioning of static camera for vehicle-road cooperation. A low-cost camera calibration method is proposed to complete the accurate mapping between the image plane and the 3D world space. Precise location of interested targets are achieved by an efficient tracking strategy. Real test scenarios show that our algorithm can effectively locate vehicles, pedestrians, non-motor vehicles and other targets with high accuracy. Our algorithm won the Monocular Static Camera Positioning and Ranging Competition for Autonomous Driving championship in 2019. Tuopu Wen, Zhongyang Xiao, Kun Jiang 0002, Mengmeng Yang 0001, Keqiang Li 0002, Diange Yang |
MMSP | 3 |
| 2016 | Estimation and prediction of vehicle dynamics states based on fusion of OpenStreetMap and vehicle dynamics modelsabstractThis paper presents a novel approach for estimation and prediction of vehicle dynamics states by incorporating digital road map and vehicle dynamics models. Precise information about vehicle dynamics states is essential for the safety and stability of vehicle. In particular, the tire-road contact forces and vehicle side slip angle are the most important parameters for evaluating the safety of vehicle. Nevertheless, these dynamics states are immeasurable with low cost sensors. Therefore, different observers, or the so-called virtual sensors are developed to estimate vehicle dynamics states. However, the existing observers are only capable in estimating vehicle dynamics states at a current instant but not to predict the potential dangers in a future instant. In order to make time for correcting drive behaviors, especially when driving at high speed, it seems very appealing for us to predict an impending dangerous event and react before the danger occurs. In this paper, the estimation of vehicle dynamics states is based on the fusion of information from inertial sensors, GPS and OpenStreetMap. The geometry of the upcoming path ahead of vehicle is provided by the digital map and is employed to predict the future dynamics states. Kun Jiang 0002, Alessandro Corrêa Victorino, Ali Charara 0002 |
Intelligent Vehicles Symposium | 1 |