Jun Li 0082

dblp:116/1011-82 · DBLP profile ↗
← Back
43ranked-venue papers
0as first author
42since 2021 · last 2026
0000-0002-0437-5112ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 15 since 2021Systems, architecture and hardware · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal Large Language Models for Perception in Autonomous Driving: Architecture, Taxonomy, and Challenges
abstract
Autonomous vehicles rely on continuous environmental perception to assess obstacle distribution to ensure safe driving. However, current perception technologies face substantial challenges, particularly under adverse weather conditions and when encountering long-tail scenarios. With the advent of transformer-based attention mechanisms, large language models (LLMs), exemplified by GPT, have exhibited emergent intelligence, offering new possibilities for achieving high-performance perception. This technological advancement has led to the development of multi-modal large language models (MLLMs), which incorporate multi-modal encoders. These models enable a single LLM to process multi-source data while performing advanced understanding and reasoning tasks, enhancing complex environmental perception capabilities. Despite significant progress in MLLMs, there remains a notable gap in systematic research on their optimal application to perception tasks. Therefore, this paper presents a comprehensive survey of recent advancements in MLLM-based perception. First, we introduce the mainstream vision–language perception tasks, widely adopted evaluation metrics, and existing language-enhanced autonomous driving datasets. Next, we outline the general architectural design principles of MLLMs. Subsequently, we provide a taxonomy and indepth analysis of current MLLMs, focusing on three dimensions: input modality, alignment technique, and scene representation, elucidating their underlying implementation paradigms. Finally, we summarize the key challenges and emerging research directions in MLLM-driven perception. This survey aims to facilitate further progress in MLLMs by synthesizing these insights.
Xinyu Zhang 0001, Yuchuan Ji, Yanchao Ding, Jialun Yin, Ruizhi Jia, Yijin Xiong, Jun Li 0082, Huaping Liu 0001
IEEE Internet Things J.11
2026 An Empirical Analysis of Cooperative Perception for Occlusion Risk Mitigation
abstract
Occlusions present a significant challenge for connected and automated vehicles, as they can obscure critical road users from perception systems. Traditional risk metrics often fail to capture the cumulative nature of these threats over time adequately. In this paper, we propose a novel and universal risk assessment metric, the Risk of Tracking Loss (RTL), which aggregates instantaneous risk intensity throughout occluded periods. This provides a holistic risk profile that encompasses both high-intensity, short-term threats and prolonged exposure. Utilizing diverse and high-fidelity real-world datasets, a large-scale statistical analysis is conducted to characterize occlusion risk and validate the effectiveness of the proposed metric. The metric is applied to evaluate different vehicle-to-everything (V2X) deployment strategies. Our study shows that full V2X penetration theoretically eliminates this risk, the reduction is highly nonlinear; a substantial statistical benefit requires a high penetration threshold of 75–90%. To overcome this limitation, we propose a novel asymmetric communication framework that allows even non-connected vehicles to receive warnings. Experimental results demonstrate that this paradigm achieves better risk mitigation performance. We found that our approach at 25% penetration outperforms the traditional symmetric model at 75%, and benefits saturate at only 50% penetration. This work provides a crucial risk assessment metric and a cost-effective, strategic roadmap for accelerating the safety benefits of V2X deployment.
Aihong Wang, Tenghui Xie, Fuxi Wen, Jun Li 0082
IEEE Internet Things J.4
2026 Diffusion Models for Autonomous Driving in Smart Parking: A Data Synthesis Framework
abstract
The accelerating urbanization process and the increase in vehicle ownership have increasingly highlighted the issue of parking difficulties. Intelligent parking systems, as an emerging solution to this problem, are centered around the perception tasks of autonomous driving technology within parking lots. However, the complexity of parking environments poses significant challenges to the perception systems of autonomous vehicles. To address this, this paper proposes a data synthesis framework based on diffusion models, aimed at generating high-quality synthetic parking lot datasets to support the development of intelligent parking systems. By constructing standardized parking lot scenarios in simulators, simulated images containing original images and real labels are automatically generated to address the high cost and low efficiency of real data collection. The innovative application of diffusion models to the synthesis of intelligent parking perception data results in the generation of highly realistic images that closely resemble real parking lot scenes, while preserving accurate labeling data. This approach effectively reduces the discrepancy between traditional simulation data and real-world data. Experimental evaluations have demonstrated the superior performance of this approach in terms of visual quality, generation diversity, and object detection performance. This research provides a substantial resource for training autonomous driving perception algorithms and offers an effective means to verify and improve system safety and reliability.
Xinyu Zhang 0001, Jialun Yin, Jun Li 0082
IEEE Trans Autom. Sci. Eng.6
2026 RFFRDet: A Refined Feature Fusion Rotation Detector for Prohibited Item Recognition in X-Ray Images
abstract
With increasing security threats in critical infrastructure, intelligent X-ray inspection systems have become essential for modern security frameworks. Contraband detection faces dual challenges: semantic dilution across cross-scale features and degraded directional sensitivity. Existing detection paradigms rely on Euclidean spatial assumptions while ignoring X-ray projection geometry characteristics, leading to feature representation ambiguity and insufficient spatial relationship modeling in complex occlusion scenarios. To address these challenges, we propose RFFRDet, a rotation detector based on refined feature decoupling. First, an Integrated Local Attention (ILA) module is constructed that enables material-aware and geometry-aware feature enhancement through channel-spatial decoupling. Second, a Multi-Resolution Feature Fusion (MRFF) network is designed that achieves optimal coupling between fine-grained spatial positioning and high-level semantic understanding through parallel multi-scale aggregation. Finally, we construct the first multi-directional X-ray security benchmark dataset with rotation annotations for 27,000 images. Extensive experiments demonstrate that RFFRDet achieves 97.4% mAR and 96.8% mAP, representing improvements of 1.6% and 1.5% over state-of-the-art methods, respectively.
Mingxun Wang 0003, Wenfeng Guo, Jun Li 0082
IEEE Trans. Inf. Forensics Secur.5
2026 Brain-in-the-Loop Learning for Intelligent Vehicle Decision-Making
abstract
The inflexible human-autonomy relationship within autonomous driving scenarios still has not realized synergetic intelligence, therefore unable to provide adaptive and context-sensitive decision-making and sometimes leading to violation of human pReferences or even hazards. In this paper, we utilize functional near-infrared spectroscopy (fNIRS) signals as real-time human risk-perception feedback to establish a brain-in-the-loop (BiTL) trained artificial intelligence algorithm for decision-making. The proposed algorithm uses the result of driving risk reasoning as one input of reinforcement learning combining fNIRS-based risk and driving safety field model-based risk, realizing integrating human brain activity into the reinforcement learning scheme, then overcoming the disadvantage of machine-oriented intelligence that could violate human intentions. To achieve policy learning within limited BiTL training periods, we add two modification features to the proposed algorithm based on TD3. The experiment involving twenty participants has been conducted, and the results show that in continuously high-risk driving scenarios, compared to traditional reinforcement learning algorithms without human participation, the proposed algorithm can maintain a cautious driving policy and avoid potential collisions, validated with both proximal surrogate indicators and success rates.
Haoyi Zheng, Jun Li 0082, Chaosheng Huang, Hong Wang 0014
IEEE Trans. Intell. Transp. Syst.3
2025 TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception
abstract
Cooperative perception presents significant potential for enhancing the sensing capabilities of individual vehicles, however, inter-agent latency remains a critical challenge. Latencies cause misalignments in both spatial and semantic features, complicating the fusion of real-time observations from the ego vehicle with delayed data from others. To address these issues, we propose TraF-Align, a novel framework that learns the flow path of features by predicting the feature-level trajectory of objects from past observations up to the ego vehicle’s current time. By generating temporally ordered sampling points along these paths, TraF-Align directs attention from the current-time query to relevant historical features along each trajectory, supporting the reconstruction of current-time features and promoting semantic interaction across multiple frames. This approach corrects spatial misalignment and ensures semantic consistency across agents, effectively compensating for motion and achieving coherent feature fusion. Experiments on two real-world datasets, V2V4Real and DAIR-V2X-Seq, show that TraF-Align sets a new benchmark for asynchronous cooperative perception. The code is available at https://github.com/zhyingS/TraF-Align.
Zhiying Song, Lei Yang 0060, Fuxi Wen, Jun Li 0082
CVPR4
2025 Learning Based MPC for Autonomous Driving Using a Low Dimensional Residual Model
abstract
In this paper, a learning based Model Predictive Control (MPC) using a low dimensional residual model is proposed for autonomous driving. One of the critical challenge in autonomous driving is the complexity of vehicle dynamics, which impedes the formulation of accurate vehicle model. Inaccurate vehicle model can significantly impact the performance of MPC controller. To address this issue, this paper decomposes the nominal vehicle model into invariable and variable elements. The accuracy of invariable elements are ensured by calibration, while the deviations in the variable elements are learned by a low-dimensional residual model. The features of residual model are selected as the physical variables most correlated with nominal model errors. Physical constraints among these features are formulated to explicitly define the valid region within the feature space. The formulated model and constraints are incorporated into the MPC framework and validated through both simulation and real vehicle experiments. The results indicate that the proposed method significantly enhances the model accuracy and controller performance.
Chaosheng Huang, Jun Li 0082
ICRA5
2025 Autonomous Vehicle Path Planning and Path Tracking Control
abstract
Path planning and path tracking play an important role in the development of the auto drive system. Based on the advantages of Bezier Curves, such as smoothness and ease of calculation, this paper proposes a path planning method for two segment Bezier Curves, and establishes the relationship between the control points of Bezier Curves and the driving status of vehicles, so as to ensure the feasibility of the planned curves. Based on the planned path and the vehicle’s own state, a model predictive control method is adopted to develop a path tracking control algorithm. In order to fully ensure the driving stability of the vehicle tracking path, this paper developed a vehicle stability control algorithm based on model predictive control to ensure the stability of the vehicle during the path tracking process. Finally, the corresponding verification test is designed, and the test results show that the path planning and path tracking control algorithm considering stability developed in this paper have good control effects. It can not only plan the driving path according to the obstacle information and the current driving state of the vehicle, but also ensure the driving stability of the vehicle in the path tracking process, and improve the driving safety of autonomous vehicle.
Shubin Lin, Chaosheng Huang, Jun Li 0082
INDIN5
2025 V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception
abstract
Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In recent years, a series of cooperative perception datasets have emerged; however, these datasets primarily focus on cameras and LiDAR, neglecting 4D Radar—a sensor used in single-vehicle autonomous driving to provide robust perception in adverse weather conditions. In this paper, to bridge the gap created by the absence of 4D Radar datasets in cooperative perception, we present V2X-Radar, the first large-scale, real-world multi-modal dataset featuring 4D Radar. V2X-Radar dataset is collected using a connected vehicle platform and an intelligent roadside unit equipped with 4D Radar, LiDAR, and multi-view cameras. The collected data encompasses sunny and rainy weather conditions, spanning daytime, dusk, and nighttime, as well as various typical challenging scenarios. The dataset consists of 20K LiDAR frames, 40K camera images, and 20K 4D Radar data, including 350K annotated boxes across five categories. To support various research domains, we have established V2X-Radar-C for cooperative perception, V2X-Radar-I for roadside perception, and V2X-Radar-V for single-vehicle perception. Furthermore, we provide comprehensive benchmarks across these three sub-datasets.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Jiaqi Ma 0003, Zhiying Song, Ziying Song, Li Wang 0092, Yang Shen 0005, Chen Lv 0001
NeurIPS3
2025 Cooperative Camera-LiDAR Extrinsic Calibration for Vehicle-Infrastructure Systems in Urban Intersections
abstract
Cameras and LiDARs are fundamental sensors for autonomous driving, with wide applicability. This paper comprehensively studies and analyzes the calibration of the multi-terminal camera-LiDAR devices from the perspectives of vehicle, roadside, and road cooperation, and outlines its related applications and far-reaching significance. Different from previous studies that focused on sensor calibration of a single platform, ignoring the differences between vehicle-end and roadside, this paper particularly emphasizes the difference between vehicle-end and roadside, as well as heterogeneous calibration challenges emerging in vehicle-road cooperation scenarios. These challenges include cross-sensor temporal synchronization and spatial alignment in dynamic environments. It reveals the advantages and differences of different methods in dealing with constraints such as the fixed installation postures of vehicle-mounted and roadside devices and the large-scale observation domain. This paper further points out key research gaps, such as the adaptability of online calibration and the maintenance of calibration consistency among multiagents, providing important references for the technological evolution from isolated calibration to a sensor-collaborative calibration framework in traffic scenarios.
Yijin Xiong, Xinyu Zhang 0001, Xin Gao 0028, Qianxin Qu, Chun Duan, Jun Li 0082
IEEE Internet Things J.8
2025 BEVHeight++: Toward Robust Visual Centric 3D Object Detection
abstract
While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric detection methods perform poorly on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight++, to address this issue. In essence, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. By incorporating both height and depth encoding techniques, we achieve a more accurate and robust projection from 2D to BEV spaces. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. In terms of the ego-vehicle scenario, BEVHeight++ surpasses depth-only methods with increases of +2.8% NDS and +1.7% mAP on the nuScenes test set, and even higher gains of +9.3% NDS and +8.8% mAP on the nuScenes-C benchmark with object-level distortion. Consistent and substantial performance improvements are achieved across the KITTI, KITTI-360, and Waymo datasets as well.
Lei Yang 0060, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Yi Huang 0038, Xinyu Zhang 0001, Kaicheng Yu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 FMRT: Learning Accurate Feature Matching With Reconciliatory Transformer
abstract
Local Feature Matching, a pivotal component of numerous computer vision tasks (e.g., structure from motion and visual localization), has been effectively addressed by Transformer-based methods. Nevertheless, these methods solely incorporate long-range context information among keypoints with a fixed receptive field, which constrains the network from appropriately reconciling the importance of features with diverse receptive fields to realize complete image perception, hence limiting feature matching accuracy. In addition, these methods employ a conventional handcrafted encoding approach to incorporate positional information of keypoints into visual descriptors, which limits the capability of networks to extract effective positional encoding message. In this study, we propose FMRT, a novel detector-free method that reconciles local features with diverse receptive fields adaptively and utilizes parallel networks to realize reliable positional encoding. Specifically, FMRT proposes a dedicated reconciliatory transformer (RecFormer) that contains a global perception attention layer to identify visual descriptors with different receptive fields and integrate global context information under various scales, a perception weight layer to measure the importance of various receptive fields adaptively, and a local perception feed-forward network to extract deep aggregated multi-scale local feature representation. Moreover, we introduce a novel axis-wise position encoder (AWPE) that views positional encoding as two keypoints encoding tasks along the row and column dimensions, decouples the x- and y-coordinates of keypoints into two independent 1D vectors, and designs two parallel network branches to explicitly encodes geometric correlations among keypoints, hence realizing reliable positional encoding. Extensive experiments indicate that FMRT yields impressive performance on multiple tasks, including relative pose estimation, visual localization, homography estimation, and image matching. Besides, we integrate FMRT into a localization framework and conduct a visual localization experiment in a real scene, which further demonstrate the superiority of FMRT. Note to Practitioners—This paper presents a novel approach to enhancing the performance of local feature matching in computer vision tasks. Traditional methods often rely on fixed receptive fields for integrating context among keypoints, which can limit the perception of the complete image and, consequently, the precision of feature matching. Our work introduces a Reconciliatory Transformer that not only addresses these limitations by effectively reconciling the importance of features across varying receptive fields but also improves the integration of positional information into visual descriptors. The techniques developed here can be adapted to a wide range of systems, e.g., image matching for computer vision and visual localization for autonomous driving, offering practitioners a tool to significantly improve the fidelity of feature matching, which is foundational for accurate interaction with the surrounding environment.
Li Wang 0092, Xinyu Zhang 0001, Tao Xie 0010, Lei Yang 0060, Wenhao Yu 0006, Yang Shen 0005, Bin Xu 0003, Jun Li 0082
IEEE Trans Autom. Sci. Eng.10
2025 GF-SLAM: A Novel Hybrid Localization Method Incorporating Global and Arc Features
abstract
A global and feature-based hybrid algorithm, which integrates global information and feature-based simultaneous localization and mapping (GF-SLAM). This system can operate adaptively when external signals are unstable, thereby avoiding cumulative errors produced by local methods. In agricultural planting bases with abundant circular arc features, the focus is on efficiently exploring the correlations between these features to optimize the robust real-time positioning system for work vehicles. In this process, feature-based SLAM (F-SLAM) is applied to partial positioning using a particle filter. Available global information is then fused using an extended Kalman filter (EKF) for precise positioning and deviation correction in the mapping process, thereby achieving an effective combination of two positioning modes. The proposed model was evaluated using two simulation environments and a comparison with representative techniques. Results showed that GF-SLAM was competitive in normal conditions while requiring fewer computations and significantly reducing the drift in F-SLAM for stable global signals. Switching between these two algorithms eliminated positioning errors to within 1 cm for a test case in which global localization was lost in 54.7% of the route, producing an error within 3 cm. The code will be open source. Note to Practitioners—Our adaptive fusion strategy aims to address challenges in real-world agricultural scenarios where global information may be lacking or unstable. This approach enhances the robustness and reliability of robot positioning in practical applications. We invite practitioners to consider the adaptability of our system to diverse environments, particularly those with limited or fluctuating global information. In this paper, we first introduce the EKF module for global localization, then illustrate the establishment of feature maps and particle filter positioning in F-SLAM. We conduct comparison experiments in two virtual environments and real agricultural planting bases, where approximately half of the route lacks global information. The results demonstrate the benefits of our adaptive fusion strategy in practical positioning applications. We also plan to explore the integration of additional sensors, such as cameras, combined with deep learning, to further improve the efficiency and quality of feature extraction. This extension is aimed at mitigating the issue of positioning failure caused by crop and equipment occlusion in agricultural scenarios.
Yijin Xiong, Xinyu Zhang 0001, Wenju Gao, Qianxin Qu, Shichun Guo, Yang Shen 0005, Jun Li 0082
IEEE Trans Autom. Sci. Eng.9
2025 PDDepth: Pose Decoupled Monocular Depth Estimation for Roadside Perception System
abstract
Accurate depth information is crucial for roadside perception in Cooperative Vehicle Infrastructure Systems. Beyond existing radar and LiDAR solutions, monocular depth estimation using surveillance cameras is emerging as a superior approach due to its cost-effectiveness and dense depth output. Unlike onboard cameras, roadside cameras are relatively fixed in position. Many existing monocular depth estimation methods, which do not independently model camera pose, tend to overfit to training data and produce suboptimal results when confronted with slight variations in camera poses, which may be caused by external forces within the same camera or across different cameras. To address this issue, a pose decoupled monocular depth estimation method specifically designed for roadside perception systems is proposed. This method separates depth estimation into two components: a pose-dependent modeling portion that recovers ground depth based on the current camera pose, and a pose-agnostic portion that estimates pixel height relative to the ground plane. Additionally, a knowledge distillation framework is introduced to improve the robustness of the proposed method against variations in roadside cameras. To validate the method, we propose the first open source dataset for roadside monocular depth estimation, DAIR-MDE, and a roadside instance segmentation dataset, DAIR-Ins, both derived from the DAIR dataset. The proposed method demonstrates significant advances over the state-of-the-art methods on DAIR-MDE. The proposed dataset and source code are publicly available athttps://github.com/441599828/PDDepth.
Huanan Wang, Xinyu Zhang 0001, Zhengxian Chen, Jun Li 0082, Huaping Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 From Prediction to Planning: Comprehensive Uncertainty Management in Autonomous Driving
Wenbo Shao, Zhong Cao 0003, Hong Wang 0014, Jun Li 0082
IEEE Trans. Intell. Transp. Syst.5
2025 SGV3D: Toward Scenario Generalization for Vision-Based Roadside 3D Object Detection
abstract
Roadside perception can significantly enhance the safety of autonomous vehicles by extending their perceptual capabilities beyond the visual range and addressing occluded regions. However, current state-of-the-art vision-based roadside detection methods exhibit high accuracy on labeled scenes but perform poorly on new scenes. This limitation arises because roadside cameras remain stationary after installation and can only gather data from a single scene, leading the algorithm to overfit these roadside backgrounds and camera positions. To tackle this issue, we propose an innovativeScenarioGeneralization Framework forVision-based Roadside3DObject Detection, calledSGV3D. Specifically, we utilize a Background-suppressed Module (BSM) to reduce background overfitting in vision-centric pipelines by diminishing background features during the 2D to bird’s-eye-view projection. Furthermore, by introducing the Semi-supervised Data Generation Pipeline (SSDG) that employs unlabeled images from new scenes, we generate diverse foreground instances with varying camera poses, mitigating the risk of overfitting to specific camera positions. Experiments conducted on two large-scale roadside benchmarks demonstrate that SGV3D, with only a minimal increase in latency, effectively improves the scenario generalization capabilities of vision-based roadside 3D object detectors. The code is available here (https://github.com/yanglei18/SGV3D).
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Zhiwei Li 0011, Yang Shen 0005, Chen Lv 0001, Hong Wang 0014
IEEE Trans. Intell. Transp. Syst.3
2025 V2X-Reg++: A Real-Time Global Registration Method for Multi-End Sensing System in Urban Intersections
abstract
Urban intersections, dense with pedestrian and vehicular traffic and compounded by positioning signal obstructions, are among the most challenging areas in urban traffic systems. Traditional single-vehicle intelligence systems often perform poorly in such environments due to a lack of global scene observations and the inherent uncertainty in predicting other agents’ intentions. Vehicle-to-Everything (V2X) technology, through real-time communication between vehicles (V2V) and vehicles to infrastructure (V2I), offers a robust solution. However, practical applications still face numerous challenges. Spatial registration among vehicle and infrastructure endpoints with different configurations in multi-end sensing systems is crucial for ensuring the accuracy of perception system data. Most existing multi-end spatial registration methods rely on initial extrinsic values provided by positioning systems, but the instability of GNSS signals due to high buildings in urban canyons poses severe challenges to these methods. To address this issue, this paper proposes a novel multi-end spatial registration method that does not require positioning priors to determine initial external parameters and meets real-time requirements. Our method introduces an innovative multi-end perception object association technique that leverages a newOverall Distance(oDist) metric to measure the spatial association between perception objects, subsequently using this metric as the foundation for an optimal transport formulation. By this means, we can extract co-observed targets from object association results for further external parameter computation and optimization. Extensive comparative and ablation experiments conducted on the simulated dataset V2X-Sim and the real dataset DAIR-V2X confirm the effectiveness and efficiency of our method. The code for this method can be accessed at:https://github.com/MassimoQu/v2i-calib.
Xinyu Zhang 0001, Qianxin Qu, Yijin Xiong, Chen Xia, Ziqiang Song, Kang Liu 0008, Jun Li 0082, Keqiang Li 0002
IEEE Trans. Intell. Transp. Syst.8
2024 Towards Safe and Reliable Autonomous Driving: Dynamic Occupancy Set Prediction
abstract
In the rapidly evolving field of autonomous driving, reliable prediction is pivotal for vehicular safety. However, trajectory predictions often deviate from actual paths, particularly in complex and challenging environments, leading to significant errors. To address this issue, our study introduces a novel method for Dynamic Occupancy Set (DOS) prediction, it effectively combines advanced trajectory prediction networks with a DOS prediction module, overcoming the shortcomings of existing models. It provides a comprehensive and adaptable framework for predicting the potential occupancy sets of traffic participants. The innovative contributions of this study include the development of a novel DOS prediction model specifically tailored for navigating complex scenarios, the introduction of precise DOS mathematical representations, and the formulation of optimized loss functions that collectively advance the safety and efficiency of autonomous systems. Through rigorous validation, our method demonstrates marked improvements over traditional models, establishing a new benchmark for safety and operational efficiency in intelligent transportation systems.
Wenbo Shao, Wenhao Yu 0006, Jun Li 0082, Hong Wang 0014
IV4
2024 Si-GAIS: Siamese Generalizable-Attention Instance Segmentation for Intersection Perception System
abstract
Instance segmentation of traffic participants using vision-based techniques serves as a cornerstone for numerous intelligent transportation systems. Although existing deep learning-based methods have made significant advancements in this field, these algorithms still present considerable challenges with regard to generalizability for commercial deployment. Specifically, the mean average perception (mAP) of these algorithms degrades rapidly when dealing with intersections that are not included in the training set. To address this limitation, a novel instance segmentation approach named Si-GAIS is proposed, which incorporates a siamese structure for the first time with the proposed Generalizable-Attention Encoder (GA). Through the proposed Foreground-Background Fusion Unit (FBF) within GA, efficient feature-level fusion for the foreground and background images is achieved. Additionally, the Interpretable Attention Neck (IA) in GA enables the feature encoder to focus exclusively on the foreground traffic participants while ignoring various backgrounds. To utilize Si-GAIS, an unsupervised method named P-DBSCAN is proposed to obtain high-quality background image for each intersection with slow-moving traffic and camera jitters. Finally, the first multi-intersection multi-category instance segmentation datasets named RopeIns is proposed for validation. Si-GAIS achieves a 7.7% mAP (All APs used in this paper are abbreviations of AP$_{\textit {50}}$the same with PASCAL VOC.) accuracy improvement compared to the state-of-the-art (SOTA) methods while using fewer parameters, with only a 6.4% decline in AP for car segmentation in unseen intersections and weather conditions, whereas all other SOTA methods decline more than 10%. The proposed dataset and source code are publicly available athttps://github.com/441599828/SiGAISand we hope Si-GAIS will be a new baseline for IPS instance segmentation research.
Huanan Wang, Xinyu Zhang 0001, Hong Wang 0014, Jun Li 0082
IEEE Trans. Intell. Transp. Syst.4
2024 MonoGAE: Roadside Monocular 3D Object Detection With Ground-Aware Embeddings
abstract
Although the majority of recent autonomous driving systems concentrate on developing perception methods based on ego-vehicle sensors, there is an overlooked alternative approach that involves leveraging intelligent roadside cameras to help extend the ego-vehicle perception ability beyond the visual range. We discover that most existing monocular 3D object detectors rely on the ego-vehicle prior assumption that the optical axis of the camera is parallel to the ground. However, the roadside camera is installed on a pole with a pitched angle, which makes the existing methods not optimal for roadside scenes. In this paper, we introduce a novel framework for Roadside Monocular 3D object detection with ground-aware embeddings, named MonoGAE. Specifically, the ground plane is a stable and strong prior knowledge due to the fixed installation of cameras in roadside scenarios. In order to reduce the domain gap between the ground geometry information and high-dimensional image features, we employ a supervised training paradigm with a ground plane to predict high-dimensional ground-aware embeddings. These embeddings are subsequently integrated with image features through cross-attention mechanisms. Furthermore, to improve the detector’s robustness to the divergences in cameras’ installation poses, we replace the ground plane depth map with a novel pixel-level refined ground plane equation map. Our approach demonstrates a substantial performance advantage over all previous monocular 3D object detectors on widely recognized 3D detection benchmarks for roadside cameras. The code and pre-trained models will be released soon.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Yi Huang 0038, Hong Wang 0014
IEEE Trans. Intell. Transp. Syst.4
2024 Auto-Points: Automatic Learning for Point Cloud Analysis With Neural Architecture Search
abstract
Pure point-based neural networks have recently shown tremendous promise for point cloud tasks, including 3D object classification, 3D object part segmentation, 3D semantic segmentation, and 3D object detection. Nevertheless, it is a laborious process to construct a network for each task due to the artificial parameters and hyperparameters involved, e.g., the depths and widths of the network and the number of sampled points at each stage. In this work, we propose Auto-Points, a novel one-shot search framework that automatically seeks the optimal architecture configuration for point cloud tasks. Technically, we introduce a set abstraction mixer (SAM) layer that is capable of scaling up flexibly along the depth and width of the network. Each SAM layer consists of numerous child candidates, which simplifies architecture search and enables us to discover the optimum design for each point cloud task pursuant to resource constraint from an enormous search space. To fully optimize the child candidates, we develop a weight-entwinement neural architecture search (NAS) technique that entwines the weights of different candidates in the same layer during supernet training such that all candidates can be extremely optimized. Benefiting from the proposed techniques, the trained supernet allows the searched subnets to be exceptionally well-optimized without further retraining or finetuning. In particular, the searched models deliver superior performances on multiple extensively employed benchmarks, 93.9% overall accuracy (OA) on ModelNet40, 89.1% OA on ScanObjectNN, 87.1% instance average IoU on ShapeNetPart, 69.1% mIoU on S3DIS, 70.4% [email protected] on ScanNet V2, and 64.4% [email protected] on SUN RGB-D.
Li Wang 0092, Tao Xie 0010, Xinyu Zhang 0001, Linqi Yang, Yilong Ren, Haiyang Yu 0002, Jun Li 0082, Huaping Liu 0001
IEEE Trans. Multim.10
2024 FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object Detection
abstract
In this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance.
Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082
IEEE Trans. Multim.9
2024 Informative Data Selection With Uncertainty for Multimodal Object Detection
abstract
Noise has always been nonnegligible trouble in object detection by creating confusion in model reasoning, thereby reducing the informativeness of the data. It can lead to inaccurate recognition due to the shift in the observed pattern, that requires a robust generalization of the models. To implement a general vision model, we need to develop deep learning models that can adaptively select valid information from multimodal data. This is mainly based on two reasons. Multimodal learning can break through the inherent defects of single-modal data, and adaptive information selection can reduce chaos in multimodal data. To tackle this problem, we propose a universal uncertainty-aware multimodal fusion model. It adopts a multipipeline loosely coupled architecture to combine the features and results from point clouds and images. To quantify the correlation in multimodal information, we model the uncertainty, as the inverse of data information, in different modalities and embed it in the bounding box generation. In this way, our model reduces the randomness in fusion and generates reliable output. Moreover, we conducted a completed investigation on the KITTI 2-D object detection dataset and its derived dirty data. Our fusion model is proven to resist severe noise interference like Gaussian, motion blur, and frost, with only slight degradation. The experiment results demonstrate the benefits of our adaptive fusion. Our analysis on the robustness of multimodal fusion will provide further insights for future research.
Xinyu Zhang 0001, Zhiwei Li 0011, Zhenhong Zou, Xin Gao 0028, Yijin Xiong, Dafeng Jin, Jun Li 0082, Huaping Liu 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection
abstract
While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric bird's eye view detection methods have inferior performances on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight, to address this issue. In essence, instead of predicting the pixel-wise depth, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. The code is available at https://github.com/ADLab-AutoDrive/BEVHeight.
Lei Yang 0060, Kaicheng Yu, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Xinyu Zhang 0001
CVPR4
2023 Failure Detection for Motion Prediction of Autonomous Driving: An Uncertainty Perspective
abstract
Motion prediction is essential for safe and efficient autonomous driving. However, the inexplicability and uncertainty of complex artificial intelligence models may lead to unpredictable failures of the motion prediction module, which may mislead the system to make unsafe decisions. Therefore, it is necessary to develop methods to guarantee reliable autonomous driving, where failure detection is a potential direction. Uncertainty estimates can be used to quantify the degree of confidence a model has in its predictions and may be valuable for failure detection. We propose a framework of failure detection for motion prediction from the uncertainty perspective, considering both motion uncertainty and model uncertainty, and formulate various uncertainty scores according to different prediction stages. The proposed approach is evaluated based on different motion prediction algorithms, uncertainty estimation methods, uncertainty scores, etc., and the results show that uncertainty is promising for failure detection for motion prediction but should be used with caution.
Wenbo Shao, Yanchao Xu, Jun Li 0082, Hong Wang 0014
ICRA4
2023 PeSOTIF: a Challenging Visual Dataset for Perception SOTIF Problems in Long-tail Traffic Scenarios
abstract
Perception algorithms in autonomous driving systems confront great challenges in long-tail traffic scenarios, where the problems of Safety of the Intended Functionality (SOTIF) could be triggered by the algorithm performance insufficiency and dynamic operational environment. However, such scenarios are not systematically included in current open-source datasets, and this paper fills the gap accordingly. Based on the analysis and enumeration of trigger conditions, a high-quality diverse dataset is released, including various long-tail traffic scenarios collected from multiple resources. Considering the development of probabilistic object detection (POD), this dataset marks trigger sources that may cause perception SOTIF problems in the scenarios as key objects. In addition, an evaluation protocol is suggested to verify the effectiveness of POD algorithms in identifying the key objects via uncertainty. The dataset never stops expanding, and the first batch of open-source data includes 1126 frames with an average of 2.27 key objects and 2.47 normal objects in each frame. To demonstrate how to use this dataset for SOTIF research, this paper further quantifies the perception SOTIF entropy to confirm whether a scenario is unknown and unsafe for a perception system. The experimental results show that the quantified entropy can effectively and efficiently reflect the failure of the perception algorithm.
Jun Li 0082, Wenbo Shao, Hong Wang 0014
IV2
2023 Self-Aware Trajectory Prediction for Safe Autonomous Driving
abstract
Trajectory prediction is one of the key components of the autonomous driving software stack. Accurate prediction for the future movement of surrounding traffic participants is an important prerequisite for ensuring the driving efficiency and safety of intelligent vehicles. Trajectory prediction algorithms based on artificial intelligence have been widely studied and applied in recent years and have achieved remarkable results. However, complex artificial intelligence models are uncertain and difficult to explain, so they may face unintended failures when applied in the real world. In this paper, a self-aware trajectory prediction method is proposed. By introducing a self-awareness module and a two-stage training process, the original trajectory prediction module's performance is estimated online, to facilitate the system to deal with the possible scenario of insufficient prediction function in time, and create conditions for the realization of safe and reliable autonomous driving. Comprehensive experiments and analysis are performed, and the proposed method performed well in terms of self-awareness, memory footprint, and real-time performance, showing that it may serve as a promising paradigm for safe autonomous driving.
Wenbo Shao, Jun Li 0082, Hong Wang 0014
IV2
2023 SAT-GCN: Self-attention graph convolutional network-based 3D object detection for autonomous driving
Li Wang 0092, Ziying Song, Xinyu Zhang 0001, Jun Li 0082, Huaping Liu 0001
Knowl. Based Syst.7
2023 Lite-FPN for keypoint-based monocular 3D object detection
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu
Knowl. Based Syst.3
2023 Mix-Teaching: A Simple, Unified and Effective Semi-Supervised Learning Framework for Monocular 3D Object Detection
abstract
Semi-supervised learning (SSL) has promising potential for improving model performance using both labelled and unlabelled data. Since recovering 3D information from 2D images is an ill-posed problem, the current state-of-the-art methods of monocular 3D object detection (Mono3D) have relatively low precision and recall, making semi-supervised learning for Mono3D tasks challenging and understudied. In this work, we propose a unified and effective semi-supervised learning framework called Mix-Teaching that can be applied to most monocular 3D object detectors. Based on the idea of decomposition and recombination, unlabelled samples are firstly decomposed into collections of image patches with high-quality predictions and collections of background images containing no objects. The student model is then trained on the mixed images containing dense instances with high-quality pseudo-labels generated by the recombination operation. In addition, we propose an uncertainty-based filter to distinguish high-quality pseudo-labels from noisy predictions during the decomposition process. As results in KITTI and nuScenes benchmarks, Mix-Teaching consistently improves MonoFlex and GUPNet by significant margins under various labeling ratios. Our method achieves around +6.34%$AP_{3D}$improvement against the GUPNet on the validation set when using only 10% labelled data. Using the full training set and the additional 38K raw images from KITTI, it can further improve the MonoFlex by +4.65% absolute improvement on$AP_{3D}$for car detection, reaching 18.54%$AP_{3D}$, which ranks the 1st place among all monocular based methods on the KITTI test leaderboard.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu, Huaping Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Semantic Traffic Law Adaptive Decision-Making for Self-Driving Vehicles
abstract
Facts proved that obeying traffic laws keeps the promise to promote the safety of self-driving vehicles. Current self-driving vehicles usually have fixed algorithms during autonomous driving, however the traffic laws may differ or change in different regions or times, e.g., tidal lanes. It raises a crucial requirement to make self-driving vehicles adapt to the newly received traffic laws. The challenges are that traffic laws are usually semantic and manually designed, but the original algorithms may not always contain the pre-designed interface to adapt to emerging laws. To this end, this work proposes a traffic law adaptive decision-making platform, which uses the linear temporal logic (LTL) formula to consistently describe the semantic traffic laws. Then, an LTL-based reinforcement learning framework is designed to estimate the probability of illegal behavior under different traffic laws. Finally, a law-specific backup policy is designed to maintain the performance threshold by monitoring the probability of illegal behavior. This work takes three typical scenarios where the traffic laws differ for instance to prove the effectiveness of the proposed approach, i.e., law amendment presented by the government, law difference between different regions, and temporary traffic control. The results show that the proposed method can help the original decision-making algorithms adapt to the traffic laws well without pre-defined interfaces. This method provides a way to administer on-road driving self-driving vehicles.
Hong Wang 0014, Zhong Cao 0003, Wenhao Yu 0006, Chengxiang Zhao, Ding Zhao, Diange Yang, Jun Li 0082
IEEE Trans. Intell. Transp. Syst.8
2023 How Does Traffic Environment Quantitatively Affect the Autonomous Driving Prediction?
abstract
Accurate trajectory prediction is essential for safe and efficient autonomous driving in complex traffic environments. While artificial intelligence has shown great promise in improving prediction accuracy, its inherent uncertainty and lack of explainability may lead to unpredictable failures, creating challenges for safety-critical decision-making. This study aims to address these challenges by exploring the impact of traffic environment on prediction algorithms. The study proposes a trajectory prediction framework with epistemic uncertainty estimation ability to output high uncertainty when facing unforeseeable or unknown scenarios. The framework analyzes the environmental effect on the trajectory prediction by considering scenario features and shifts. Features are divided into kinematic features of a target agent, features of surrounding traffic participants, and other scenario features. Feature correlation and importance analyses are performed to study their influence on prediction error and epistemic uncertainty. The impact of unavoidable distributional shifts in the real world on trajectory predictions is investigated using multiple intersection datasets. The results indicate that deep ensemble-based methods have advantages in improving robustness while estimating epistemic uncertainty. Consistent conclusions were obtained from the correlation and importance analyses, indicating that kinematic features of the target agent have relatively strong effects on both prediction error and epistemic uncertainty. Finally, the study analyzes the accuracy deterioration caused by distributional shifts and the potential of the deep ensemble-based method. Through deep ensemble, the errors of the prediction methods based on GRIP++ and Trajectron++ have been improved by 6.4% and 10.8% in the same-dataset test, and 6.3% and 10.8% in the cross-dataset test.
Wenbo Shao, Yanchao Xu, Jun Li 0082, Chen Lv 0001, Weida Wang, Hong Wang 0014
IEEE Trans. Intell. Transp. Syst.3
2023 CAMO-MOT: Combined Appearance-Motion Optimization for 3D Multi-Object Tracking With Camera-LiDAR Fusion
abstract
3D Multi-object tracking (MOT) ensures consistency during continuous dynamic detection, conducive to subsequent motion planning and navigation tasks in autonomous driving. However, camera-based methods suffer in the case of occlusions and it can be challenging to track the irregular motion of objects for LiDAR-based methods accurately. Some fusion methods work well but do not consider the untrustworthy issue of appearance features under occlusion. At the same time, the false detection problem also significantly affects tracking. As such, we propose a novel camera-LiDAR fusion 3D MOT framework based on Combined Appearance-Motion Optimization (CAMO-MOT), which uses both camera and LiDAR data and significantly reduces tracking failures caused by occlusion and false detection. For occlusion problems, we are the first to propose an occlusion head to select the best object appearance features multiple times effectively, reducing the influence of occlusions. To decrease the impact of false detection in tracking, we design a motion cost matrix based on confidence scores which improve the positioning and object prediction accuracy in 3D space. As existing multi-object tracking methods always evaluate each category separately and do not consider the mismatch between objects of different categories, we also propose to build a multi-category cost to implement multi-object tracking in multi-category scenes. A series of validation experiments are conducted on the KITTI and nuScenes tracking benchmarks. Our proposed method achieves state-of-the-art performance with 79.99% HOTA and the lowest identity switches (IDS) value (23 for Car and 137 for Pedestrian) among all multi-modal MOT methods on the KITTI test dataset. And our method achieves state-of-the-art performance among all algorithms on the nuScenes test dataset with 75.3% AMOTA.
Li Wang 0092, Xinyu Zhang 0001, Wenyuan Qin, Jinghan Gao, Lei Yang 0060, Zhiwei Li 0011, Jun Li 0082, Hong Wang 0014, Huaping Liu 0001
IEEE Trans. Intell. Transp. Syst.8
2023 Uncertainties in Onboard Algorithms for Autonomous Vehicles: Challenges, Mitigation, and Perspectives
abstract
Autonomous driving is considered one of the revolutionary technologies shaping humanity’s future mobility and quality of life. However, safety remains a critical hurdle in the way of commercialization and widespread deployment of autonomous vehicles on public roads. Safety concerns require the autonomous driving system to handle uncertainties from multiple sources that are either preexisting, e.g., the stochastic behavior of traffic participants or scenario occlusion, or introduced as a result of processing, e.g., the application of neural networks. Thus, it is crucial to analyze the sources of uncertainties and quantify the risks associated with them, including the propagated risks that accumulate in the decision-making system. In this context, this paper provides an overview of uncertainty challenges and state-of-the-art techniques for mitigating these challenges. We argue that the uncertainties mainly originate from two aspects: 1) the external traffic environment, and 2) the internal autonomous driving system. Specifically, this paper first analyzes the safety challenges caused by the uncertainties and summarizes their sources. In addition, the corresponding techniques that mitigate and quantify the risk of uncertainties are presented. Finally, research perspectives are highlighted to facilitate future studies for guaranteeing the safety of autonomous vehicles.
Kai Yang 0032, Xiaolin Tang, Jun Li 0082, Hong Wang 0014, Guichuan Zhong, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.3
2022 IPS300+: a Challenging multi-modal data sets for Intersection Perception System
abstract
Due to high complexity and occlusion, insufficient perception in the crowded urban intersection can be a serious safety risk for both human drivers and autonomous algorithms, whereas CVIS (Cooperative Vehicle Infrastructure System) is a proposed solution for full-participants perception under this scenario. However, the research on roadside multi-modal perception is still in its infancy, and there is no open-source data sets for such scene. Accordingly, this paper fills the gap. Through an IPS (Intersection Perception System) installed at the diagonal of the intersection, this paper proposes a high-quality multi-modal data sets for the intersection perception task. The center of the experimental intersection covers an area of 3000m2, and the extended distance reaches 300m, which is typical for CVIS. The first batch of open-source data includes 14198 frames, and each frame has an average of 319.84 labels, which is 9.6 times larger than the most crowded data sets (H3D data sets in 2019) by now. Our data sets is available at: http://www.openmpd.com/column/IPS300.
Huanan Wang, Xinyu Zhang 0001, Zhiwei Li 0011, Jun Li 0082, Zhu Lei, Haibing Ren
ICRA4
2022 InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection
abstract
Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggested. In recent years, the fusion of LiDAR and Radar has gained ever-increasing attention, especially 4D Radar, which can adapt to bad weather conditions due to its penetrability. Although features have been fused from multiple sensing modalities, most methods cannot learn interactions from different modalities, which does not make for their best use. Inspired by the self-attention mechanism, we present InterFusion, an interaction-based fusion framework, to fuse 16-line LiDAR with 4D Radar. It aggregates features from two modalities and identifies cross-modal relations between Radar and LiDAR features. In experimental evaluations on the Astyx HiRes 2019 dataset, our method outperformed the baseline by 4.20% mAP in 3D and 10.76% BEV mAP for the car class at the moderate level.
Li Wang 0092, Xinyu Zhang 0001, Baowei Xv, Jinzhao Zhang, Haibing Ren, Pingping Lu, Jun Li 0082, Huaping Liu 0001
IROS10
2022 Safety Decision of Running Speed Based on Real-time Weather
abstract
The safety of autonomous vehicles is hard to ensure in adverse weather since the sensors will degrade drastically. Setting a variable speed limit based on real-time weather condition is the most efficient method to make the vehicle safe. But most current speed limit methods are based on human visibility rather than the sensor, which is not suitable for autonomous vehicles. Thus, it is necessary to explore the performance of sensors in different weathers and propose a speed limit method based on sensor performance. Safety decisions will be made based on the calculated speed limit to ensure safety.This paper describes how to make safety decisions based on sensor performance and road conditions in real-time. The experiment explores the degradation of different sensors, and variable speed limit methods are proposed for rainy and foggy days. MPC controller is used to generate safety decisions.
Hong Wang 0014, Jun Li 0082, Wenhao Yu 0008
IV3
2022 Risk Assessment and Mitigation in Local Path Planning for Autonomous Vehicles With LSTM Based Predictive Model
abstract
Accurate trajectory prediction of surrounding vehicles enables lower risk path planning in advance for autonomous vehicles, thus promising the safety of automated driving. A low-risk and high-efficiency path planning approach is proposed for autonomous driving based on the high-performance and practical trajectory prediction method. A long short-term memory (LSTM) network is trained and tested using the highD dataset, and the validated LSTM is used to predict the trajectories of surrounding vehicles combining the information extracted from vehicle-to-vehicle (V2V) technology. A risk assessment and mitigation-based local path planning algorithm is proposed according to the information of predicted trajectories of surrounding vehicles. Two driving scenarios are extracted and reconstructed from the highD dataset for validation and evaluation, i.e., an active lane-change scenario and a longitudinal collision-avoidance scenario. The results illustrate that the risk is mitigated and the driving efficiency is improved with the proposed path planning algorithm comparing to the constant-velocity prediction and the prediction method of the nonlinear input–output (NIO) network, especially when the velocity and trajectory with sudden changes. Note to Practitioners—This article was motivated by the problem of promising the safety decision-making and path planning through accurate environment prediction. There are two main parts included in this article. First, this article proposed one pragmatic approach to predict the environment movement correctly based on the long short-term memory (LSTM) approach. The prediction performance of LSTM was compared with nonlinear input–output (NIO). The results showed that the LSTM approach has a significant advantage in motivation prediction of the surrounded vehicles during path planning. The second part of this article is to make the decision and realize local path planning based on the risk assessment. The potential field-based approach is implemented on the risk assessment based on these accurate predictions. Some primary results demonstrate that the decision-making algorithm performs better under the accurate prediction model. The results also show that the safety and driving efficiency of the ego vehicle were improved by tracking the trajectory, which was planned based on the risk assessment. The only concern for the real-time application is the computation time; in future, we will figure it out how to further reduce the computation time.
Hong Wang 0014, Bing Lu 0005, Jun Li 0082, Yang Xing 0002, Chen Lv 0001, Dongpu Cao, Ehsan Hashemi
IEEE Trans Autom. Sci. Eng.3
2022 PNNUAD: Perception Neural Networks Uncertainty Aware Decision-Making for Autonomous Vehicle
abstract
Most environment perception methods in autonomous vehicles rely on deep neural networks because of their impressive performance. However, neural networks have black-box characteristics in nature, which may lead to perception uncertainty and untrustworthy autonomous vehicles. Thus, this work proposes a decision-making method to adapt the potential perception uncertainty due to the sensor noises, fuzzy features, and unfamiliar inputs. The whole method is named as Perception Neural Networks Uncertainty Aware Decision-Making (PNNUAD) method. PNNUAD first uses the Monte Carlo dropout method to estimate the perception neural network uncertainty into a distribution around the original output. Then, the perception uncertainty will be considered in a designed reinforcement learning-based planner using a distributed value function. Finally, a backup policy will maintain the vehicle’s performance to avoid disastrous perception uncertainty. The evaluation section uses an augmented reality urban driving scenario; namely, the scenario builds in the CARLA simulator while the perception uncertainty comes from the real dataset. This case study focuses on the object class uncertainty of a widely used neural network, i.e., YOLO-V3. The results indicate that the proposed method can maintain AV safety even with poor perception performance. Meanwhile, the AV has not become too conservative by defending the perception uncertainty. This work is necessary for applying the statistics neural networks to safety-critical autonomous vehicles, and the source code will be open-source in this work.
Hong Wang 0014, Zhong Cao 0003, Diange Yang, Jun Li 0082
IEEE Trans. Intell. Transp. Syst.6
2021 Line-based Automatic Extrinsic Calibration of LiDAR and Camera
abstract
Reliable real-time extrinsic parameters of 3D Light Detection and Ranging (LiDAR) and camera are a key component of multi-modal perception systems. However, extrinsic transformation may drift gradually during operation, which can result in decreased accuracy of perception system. To solve this problem, we propose a line-based method that enables automatic online extrinsic calibration of LiDAR and camera in real-world scenes. Herein, the line feature is selected to constrain the extrinsic parameters for its ubiquity. Initially, the line features are extracted and filtered from point clouds and images. Afterwards, an adaptive optimization is utilized to provide accurate extrinsic parameters. We demonstrate that line features are robust geometric features that can be extracted from point clouds and images, thus contributing to the extrinsic calibration. To demonstrate the benefits of this method, we evaluate it on KITTI benchmark with ground truth value. The experiments verify the accuracy of the calibration approach. In online experiments on hundreds of frames, our approach automatically corrects miscalibration errors and achieves an accuracy of 0.2 degrees, which verifies its applicability in various scenarios. This work can provide basis for perception systems and further improve the performance of other algorithms that utilize these sensors.
Xinyu Zhang 0001, Shifan Zhu, Shichun Guo, Jun Li 0082, Huaping Liu 0001
ICRA4
2021 Lifelong Localization in Semi-Dynamic Environment
abstract
Mapping and localization in non-static environments are fundamental problems in robotics. Most of previous methods mainly focus on static and highly dynamic objects in the environment, which may suffer from localization failure in semi-dynamic scenarios without considering objects with lower dynamics, such as parked cars and stopped pedestrians. In this paper, we introduce semantic mapping and lifelong localization approaches to recognize semi-dynamic objects in non-static environments. We also propose a generic framework that can integrate mainstream object detection algorithms with mapping and localization algorithms. The mapping method combines an object detection algorithm and a SLAM algorithm to detect semi-dynamic objects and constructs a semantic map that only contains semi-dynamic objects in the environment. During navigation, the localization method can classify observation corresponding to static and non-static objects respectively and evaluate whether those semi-dynamic objects have moved, to reduce the weight of invalid observation and localization fluctuation. Real-world experiments show that the proposed method can improve the localization accuracy of mobile robots in non-static scenarios.
Shifan Zhu, Xinyu Zhang 0001, Shichun Guo, Jun Li 0082, Huaping Liu 0001
ICRA4
2021 Channel Attention in LiDAR-camera Fusion for Lane Line Segmentation
Xinyu Zhang 0001, Zhiwei Li 0011, Xin Gao 0028, Dafeng Jin, Jun Li 0082
Pattern Recognit.5
2019 The Architecture of the Intended Safety System for Intelligent Driving
abstract
As the development direction of intelligent driving in the future, the research of key technologies has made significant progress. However, due to the recent unmanned accidents, there are concerns about safety performance. To solve the safety problem, an intended safety systems for intelligent driving was proposed. This system provides security analysis and monitoring services in real time for intended problems with smart car perception, decision and control modules. Based on the concept of safety of the intended functionality, the driving scene and system safety are analyzed and evaluated to improve the safety of intelligent driving, which may help the development of intelligent driving.
Xinyu Zhang 0001, Wenbo Shao, Jun Li 0082
ISCAS5