Yingfeng Cai

dblp:81/8456 · DBLP profile ↗
← Back
62ranked-venue papers
7as first author
53since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 2 first-author · 33 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 9 since 2021Computer networks · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Electro-optical and infrared multi-sensor fusion based airborne target perception: A unified framework
Zhouyu Zhang 0002, Chenyuan He, Yingfeng Cai, Long Chen 0003, Hai Wang 0003, Can Zhong
Eng. Appl. Artif. Intell.3
2026 Semi-supervised driving style recognition via deep metric learning and liquid time-constant networks
Shangwu Jiang, Renkai Ding, Yingfeng Cai
Neurocomputing5
2026 End-to-End Autonomous Driving: From Classic Paradigm to Large Model Empowerment - A Comprehensive Survey
abstract
In recent years, autonomous driving technology has witnessed rapid global development, offering effective solutions to increasingly severe challenges related to traffic congestion and road safety. Among various paradigms, end-to-end autonomous driving has emerged as a promising alternative to traditional modular systems, owing to its streamlined architecture, enhanced decision consistency, and superior generalization capabilities. This survey provides a comprehensive review of the evolution and core technologies of end-to-end autonomous driving. It emphasizes the applications and developments of imitation learning (IL), reinforcement learning (RL), and other paradigms. Furthermore, it highlights the emerging paradigm empowered by foundation models, such as large language models (LLMs) and vision-language models (VLMs), and systematically categorizes recent advances in planning, reasoning, data generation, and scene understanding. In light of persistent challenges such as multi-modal fusion complexity, low sample efficiency, and safety risks, this survey summarizes representative research efforts and corresponding solutions. Finally, it outlines future directions, including the integration of world models (WMs) for unified data generation and inference optimization, the advancement of foundation model architectures toward modular design, sparse activation mechanisms, and knowledge distillation, and the realization of deployable, transferable, and unified multi-modal frameworks. This survey aims to serve as a comprehensive theoretical and technical reference for researchers in end-to-end autonomous driving, facilitating its evolution toward higher performance, stronger generalization, and enhanced safety assurance.
Sikai Lu, Xinhe Chen, Shunyao Zhang 0001, Qingchao Liu, Long Chen 0003, Hai Wang 0003, Yingfeng Cai
IEEE Internet Things J.9
2026 V2I-GridCalib: A Real-Time Online Calibration for Vehicle-to-Infrastructure Systems
abstract
Real-time spatial alignment between vehicles and roadside sensors is essential for unlocking collaborative perception. However, existing methods often struggle in dynamic environments due to observational scale mismatch and high latency. This paper proposes a novel edge-assisted framework for robust, real-time, and target-free online V2I calibration. First, a Directional Handshake mechanism pre-selects task-relevant point cloud subsets to reduce computational and transmission overhead. Second, a distance-adaptive rendering strategy generates scale-optimized rasterized views to explicitly address scale variation based on the vehicle's distance. Finally, a particle filter estimates the 6-DoF pose by fusing complementary 2D-3D reprojection and 2D-2D epipolar geometric constraints. Extensive experiments in real-world intersections and high-fidelity simulations validate the system. The framework maintains trajectory consistency errors below 5.5% for paths up to 400 meters and achieves centimeter-level accuracy with 5.2 cm translation and 0.3° rotation errors in static tests. Crucially, the system consistently meets stringent real-time constraints of less than 66 ms for V2X applications.
Zhuoyi Jiang, Yingfeng Cai, Yicheng Li 0001, Qingchao Liu, Long Chen 0003, Hai Wang 0003
IEEE Internet Things J.2
2026 ProtoAug: Prototype-Guided Uncertainty-Aware Augmentation for Long-Tail Motion Prediction
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yang Wang 0003, Xinxin Zuo
IEEE Internet Things J.2
2026 Learning discrete latent representations for scene-guided multi-modal motion prediction
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Long Chen 0003
Pattern Recognit.2
2026 CAMAv2: A Vision-Centric Approach for Static Map Element Annotation
abstract
The recent advancement of Bird’s Eye View (BEV) perception algorithms requires extensive, high-quality annotated map data for effective training and deployment in real-world scenarios. However, existing HD map-based auto-labeling methods, like those found in public datasets, face significant issues related to efficiency and accuracy. For instance, the nuScenes dataset reveals considerable misalignment and inconsistency between images and their annotations, with an average reprojection error of approximately 8.03 pixels. Additionally, the dependence on HD maps limits the applicability of these methods for large-scale, real-world auto-labeling. To tackle these challenges, we introduce CAMAv2: a vision-centric approach for Consistent and Accurate Map Annotation. This pipeline primarily utilizes camera inputs to generate precise 3D annotations of static map elements, achieving high reprojection accuracy across all surrounding cameras and maintaining spatiotemporal consistency throughout the entire sequence. Importantly, CAMAv2 annotations show lower reprojection errors compared to the original nuScenes map elements, with errors of 4.96 pixels versus 8.03 pixels. Comprehensive evaluations across various public datasets confirm the feasibility and generalizability of our pipeline in diverse environments, as well as under different weather and lighting conditions.
Shiyuan Chen, Jiaxin Zhang 0014, Ruohong Mei, Yingfeng Cai, Wei Sui
IEEE Trans. Intell. Transp. Syst.4
2026 Low-Cost Monocular End-to-End Autonomous Driving With Enhanced Interpretability
abstract
In recent years, end-to-end autonomous driving has achieved remarkable progress, yet challenges remain in achieving efficient deployment. Existing methods often rely on multi-modal sensor fusion, which incurs high costs and computational demands, limiting their large-scale adoption in low-cost scenarios. Reducing system cost while maintaining performance has consistently been a key focus of the industry. Moreover, the opaque nature of end-to-end models, due to the lack of intermediate representations, limits interpretability and hampers root-cause analysis when misjudgments occur in complex environments, thereby posing safety risks. To address these challenges, this paper proposes a lightweight and interpretable end-to-end autonomous driving architecture that utilizes only monocular camera images as input. The aim is to enhance decision-making efficiency and trajectory planning accuracy while ensuring system safety. The proposed architecture comprises two core components. First, in the scene perception part, a depth information supervision module is introduced to enable hierarchical geometric perception, and an online vector map construction module is designed to build structured environmental representations at low cost. This replaces high-definition maps with explicit constraint expressions, thereby improving model interpretability and decision-making reliability. Second, in the planning part, a dual-layer prediction framework is developed to decouple speed and waypoint prediction. It independently forecasts vehicle velocity and waypoints, while an optimized decoder structure produces fine-grained spatiotemporal trajectories, enhancing planning capabilities in complex traffic scenarios. Experimental results on multiple benchmarks from the CARLA Leaderboard 1.0 demonstrate that the proposed method significantly improves system efficiency while maintaining advanced decision-making performance. It outperforms existing multi-sensor fusion approaches, showcasing strong potential for practical deployment.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Jianmei Lei
IEEE Trans. Intell. Transp. Syst.2
2026 Scenario-Driven Federated Reinforcement Learning Autonomous Driving System With High Generalizability
abstract
Data-driven autonomous driving (AD) systems typically use incremental optimization based on various scenarios to improve driving safety and overall performance. However, existing algorithms lack the capability to integrate information across scenarios and optimize synchronously. Incremental optimization for different scenarios often results in catastrophic forgetting during continual learning, reducing the system’s generalizability and reliability in complex driving scenarios. In this paper, we propose a scenario-driven federated reinforcement learning (FRL) AD system with broad applicability. Our proposed cross-scenario optimization uses a federated architecture to replace the data-driven incremental learning process, allowing specific experiences to be shared across different expert data distributions. This approach develops highly generalizable imitation learning (IL) experts and effectively addresses the problem of catastrophic forgetting in continual learning. Furthermore, our proposed dynamic driving suggestions provide multifaceted guidance to reinforcement learning (RL) students, such as feature extraction, reward function modeling, loss function construction, and population-based optimization. This guidance specifically targets the challenges of high-dimensional state space representation and objective alignment in RL. Experimental results on various benchmarks demonstrate that our proposed framework outperforms current state-of-the-art (SOTA) multimodal fusion algorithms. Notably, our system achieves highly reliable control performance in complex traffic environments with only a single monocular camera.
Sikai Lu, Yingfeng Cai, Hai Wang 0003, Qingchao Liu
IEEE Trans. Intell. Transp. Syst.2
2026 RAMACoDrive: A Real-Time Asynchronous Framework for Cooperative Perception With Realistic V2X Communication
abstract
Recent advances in cooperative driving have garnered significant attention, as it is believed to have greater potential than single-vehicle systems in complex environments. However, gaps in current research frameworks impede its further development. Dataset-based frameworks provide a limited perspective, failing to capture the dynamic, asynchronous, and interactive complexities of the real world. Simulation-based frameworks often employ synchronous implementation and oversimplified communication models, impeding meaningful real-time performance evaluation in asynchronous environments. To this end, we propose RAMACoDrive, a novel real-time asynchronous framework for cooperative perception research. RAMACoDrive implements complete process separation, allowing agents and their internal components to operate asynchronously and independently. It uniquely integrates hardware-in-the-loop V2X communication using an external device for realistic data forwarding. Furthermore, RAMACoDrive incorporates an asynchronous online evaluation method with real-time performance metrics, enabling comprehensive evaluation beyond mere accuracy. Through experiments centered on cooperative 3D object detection, we demonstrate its effectiveness in revealing the impact of real-world asynchronous effects and communication constraints. RAMACoDrive provides a valuable and realistic testbed for cooperative perception algorithms and helps to advance the deployment of practical applications. The code will be released on our GitHub repository.
Shunyao Zhang 0001, Yingfeng Cai, Chengzhi Gao, Caizhen He, Hai Wang 0003, Long Chen 0003
IEEE Trans. Mob. Comput.2
2025 Dual-perspective safety driver secondary task detection method based on swin-transformer and cross-attention
Qingchao Liu, Lie Yang, Yingfeng Cai, Long Chen 0003
Adv. Eng. Informatics6
2025 Uncertainty-Aware End-to-End Autonomous Driving Modeling
abstract
End-to-end autonomous driving has become a prominent research paradigm. However, these methods typically rely on deterministic static and dynamic representations, leading to compounded errors that accumulate over time across modules and become increasingly amplified, ultimately degrading planning performance. Meanwhile, the deterministic sequential arrangement of modules leads to information loss and error propagation, and makes it difficult to model the interactions between the ego and surrounding agents, thereby compromising the reliability of the planning outcomes. Moreover, these methods exhibit limited capacity in perceiving and prioritizing critical features, as they assume a deterministic and uniform importance across all feature dimensions. To overcome these limitations, we introduce an uncertainty-aware end-to-end driving model to enhance planning performance and overall safety. First, we jointly model both static and dynamic uncertainties and incorporate them into the model, enabling the model to perceive and respond to environment uncertainties. Static uncertainty guides the extraction of correct map structures and driving rules, while dynamic uncertainty enables the model to capture the stochastic interactions between traffic participants and environment and provides probabilistic scores for trajectory selection, helping to avoid collisions. Second, we propose a distributed query mechanism that leverages planning-specific queries to decode raw features, thereby removing the reliance of planning on scene understanding. In addition, we jointly model prediction and planning, enabling rich interactions between ego and surrounding agents. Third, we introduce a multi-scale dilated aggregation that adjusts the attention to informative features, highlighting critical feature contributions and enlarges the receptive field. We conduct extensive evaluations on nuScenes dataset and Bench2Drive benchmark. Experimental results demonstrate that our approach achieves the lowest planning collision rate and L2 error, and enhances inference speed and efficiency, supporting the deployment of high-performance, low-cost.
Yingfeng Cai, Xinhe Chen, Hai Wang 0003, Long Chen 0003
IEEE Internet Things J.1
2025 An Autonomous Driving Vehicle-Road Collaborative Heterogeneous Fusion Localization Method Based on the Epipolar Plane Model and Graph Optimization
abstract
High-precision localization is central to the safety of high-level autonomous driving. However, existing SLAM and map-based localization methods suffer from challenges, such as cumulative errors and the need for frequent map updates. To address these issues, this article proposes a vehicle–road collaborative localization fusion approach designed to achieve high-accuracy localization without cumulative errors. The proposed method employs visual-inertial odometry (VIO) on the vehicle side, where gyroscope bias is optimized through image observations and an epipolar plane model to ensure robust pose initialization. On the roadside, 3-D object detection is utilized to provide drift-free vehicle position and orientation information. By leveraging factor graph optimization and an edge-based marginalization mechanism, this approach efficiently integrates local constraints from vehicle-side VIO data with global constraints from roadside localization data, thereby reducing computational complexity. Experimental results demonstrate that incorporating roadside infrastructure significantly reduces localization errors, effectively mitigating long-term drift. Evaluations conducted in both road and indoor environments at Jiangsu University indicate that the proposed system improves localization accuracy by 30% compared to visual-inertial navigation system (VINS)-Mono. The findings validate the superiority of vehicle–road collaborative localization in both accuracy and stability, offering a promising solution for advancing intelligent driving technologies.
Yicheng Li 0001, Yunqi Xia, Yingfeng Cai, Hai Wang 0003, Long Chen 0003
IEEE Internet Things J.3
2025 Beyond classification and regression: A novel multi-task deep learning framework for driver state understanding in human-machine Co-driving system
Yingfeng Cai
Knowl. Based Syst.4
2025 Multibranch Attentive Transformer With Joint Temporal and Social Correlations for Traffic Agents Trajectory Prediction
abstract
Accurately predicting the future trajectories of traffic agents is paramount for autonomous unmanned systems, such as self-driving cars and mobile robotics. Extracting abundant temporal and social features from trajectory data and integrating the resulting features effectively pose great challenges for predictive models. To address these issues, this article proposes a novel multibranch attentive transformer (MBAT) trajectory prediction network for traffic agents. Specifically, to explore and reveal diverse correlations of agents, we propose a decoupled temporal and spatial feature learning module with multibranch to extract temporal, spatial, as well as spatiotemporal features. Such design ensures each branch can be specifically tailored for different types of correlations, thus enhancing the flexibility and representation ability of features. Besides, we put forward an attentive transformer architecture that simultaneously models the complex correlations possibly occurring in historical and future timesteps. Moreover, the temporal, spatial, and spatiotemporal features can be effectively integrated based on different types of attention mechanisms. Empirical results demonstrate that our model achieves outstanding performance on public ETH, UCY, SDD, and INTERACTION datasets. Detailed ablation studies are conducted to verify the effectiveness of the model components.
Xiaobo Chen 0001, Yuwen Liang, Qiaolin Ye, Yingfeng Cai
IEEE Trans. Comput. Soc. Syst.5
2025 MliG: Scene-Level Multimodal Motion Prediction Based on Multi-Layer Interaction Graph
abstract
The representation form of multimodal motion prediction results is critical to the efficiency of the prediction-planning workflow. However, most existing methods generate multiple sets of trajectories for each independent agent to cover diverse modalities, overlooking interactions among agents during the prediction period, making it challenging to achieve scene-level consistency in multi-agent predictions. This paper introduces MliG, a scene-level multimodal conditional motion prediction framework based on Multi-Layer Interaction Graph. The interaction graph designed explicitly represent the interaction behaviors of multi-agents over future periods, guiding conditional trajectory prediction to achieve planning-friendly prediction result representation. First, an Interaction-Mask-based determination strategy is introduced to obtain ground truth interaction labels, enabling data-driven implicit relationship learning, and the generated training labels can support the accurate construction of the interaction graph. Second, scene simplification and modality decoupling using the interaction graph make it easier to fuse information and represent driving scenarios efficiently. Based on the multi-layer design, the complex interactions between agent pairs are decomposed into different group and modality layers, ensuring the diversity of interactions while improving the efficiency of conditional predictions. Extensive experiments on the Argoverse 2, Argoverse 1, and INTERACTION Datasets demonstrate that the proposed MliG achieves high prediction accuracy and significantly improves the scene-level consistency of trajectory prediction modalities for multi-agents in the same driving scenario.
Yingfeng Cai, Ziheng Lu, Hai Wang 0003, Yubo Lian, Long Chen 0003, Qingchao Liu
IEEE Trans. Intell. Transp. Syst.1
2025 Pedestrian Crossing Intention Prediction via Progressive Multimodal Token Fusion for Autonomous Driving
abstract
Pedestrians’ intention to cross the street exercises a substantial influence on the decision-making process of autonomous vehicles in urban traffic environments. However, accurately predicting pedestrian crossing intention is non-trivial due to the interweaving of pedestrian personalities and traffic scene elements. Despite the significant achievement of previous studies, challenges remain in effectively extracting and integrating diverse features from different modalities of observation data. In response, this paper proposes a novel model leveraging pedestrian bounding boxes, poses, and ego-vehicle speed to predict crossing intention. We introduce mixture expert feature embedding (MEFE) to project raw data into high-dimensional space based on different types of inputs. A multi-branch spatial and temporal graph convolutional network (MB-STGCN) is applied to capture multi-scale spatial and temporal features of pedestrian pose skeleton joints. A multi-token temporal aggregation (MTTA) method is devised to preserve abundant temporal information of observation data. Additionally, a progressive multimodal feature fusion (PMFF) method based on symmetric channel split attention (SCSA) is employed to enhance the interaction between different modalities when integrating different features. Extensive comparison experiments and ablation studies on public benchmark datasets substantiate the effectiveness of our approach, showing significant improvements over existing models.
Xiaobo Chen 0001, Wei Xu 0052, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.4
2025 Distributed Modeling and Scenario-Driven Extension Hybrid-DMPC Coordinated Control of Autonomous Vehicle Chassis
abstract
As a critical technology to improve vehicle safety and handling stability, chassis coordinated control technology faces challenges in multi-system coordination, multi-variable solutions, and multi-task allocation. This paper combines the advantages of decentralized and distributed architectures and proposes a scenario-driven extension Hybrid-DMPC algorithm with variable topology and distributed modeling. Firstly, a two-dimensional extension coordinate is constructed with the$\beta - \dot {\beta } $phase plane and LTR as feature states. The correlation function is then solved to divide the vehicle states into classical domain, extension domain, and non-domain, forming a mapping relationship with the control architecture. Secondly, the distributed state space equations with state coupling and input coupling are constructed. The proposed algorithm is designed to employ decentralized, distributed, and hybrid architectures in the classical domain, non-domain, and extension domain, respectively. The correlation function determines the fusion rules for multiple control inputs in the extension domain. Thirdly, a cost-coupled weighted optimization function is designed, enabling local agents to coordinate global performance goals through iterative optimization. Finally, co-simulation and HIL test results verified that the proposed algorithm effectively coordinates multiple control objectives, achieving high-precision trajectory tracking, excellent handling stability, and anti-roll performance.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yubo Lian, Zhaozhi Dong
IEEE Trans. Intell. Transp. Syst.2
2025 Takeover Time Prediction for Conditionally Automated Driving Vehicles: Considering Mixed Traffic Flow Environment
abstract
The length of time for drivers to take over conditionally automated driving vehicles (CADV) is often influenced by numerous factors, especially the mixed traffic flow environment. To analyze the influencing factors of driver’s takeover time in the mixed traffic flow environment, we study a real-vehicle takeover experiment with the CADV and radar video integrated machine provided by Jiangsu University. DeepGBM algorithm and Shapley additive explanation (SHAP) are utilized to predict and analyze the takeover time based on the data obtained from the real-vehicle experiment. The results show that the DeepGBM algorithm performs a better accuracy in predicting the takeover time compared with the other algorithms. Moreover, the predicting effect of DeepGBM in the scene of average takeover time is better than the shorter and longer ones. On the other hand, when taking over CADV in the intersection area, the driver pays more attention to the vehicle’s speed. While in the non-intersection area, the driver is more attentive to the longitudinal distance difference with the vehicle ahead. This study further found that CADVs are subject to significant lateral interference for the takeover process in intersection areas, significantly impacting the driver’s takeover time and vehicle safety. This study can provide a theoretical reference for automated driving companies to design takeover times.
Qingchao Liu, Jingya Zhao, Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.4
2025 A Preference-Based Multi-Agent Federated Reinforcement Learning Algorithm Framework for Trustworthy Interactive Urban Autonomous Driving
abstract
In autonomous driving (AD) tasks, data-driven deep reinforcement learning (DRL) outperforms rule-based methods in terms of continuous decision-making and adaptability. However, traditional DRL relies on hand-crafted reward functions, which introduce objective alignment challenges and reward loopholes. Moreover, the black-box structure makes it difficult to explain the decision-making process, which has a direct impact on DRL performance in complex driving situations. To address these shortcomings, a preference-based decomposable proximal policy optimization algorithm (PDPPO) is proposed for reliable interactive urban AD. The framework deconstructs the federated reinforcement learning (FRL) algorithm from various perspectives using a rule-based preference model, resulting in high-availability algorithmic performance for AD. PDPPO employs a data-rule fusion-driven hybrid vision transformer to overcome the objective alignment and high-dimensional state-space representation challenges of traditional DRL in complex urban traffic environments. Furthermore, to address the issue of algorithmic trustworthiness, PDPPO models the multi-agent FRL co-optimization process as an interpretable self-organized group collaboration process. This approach enables the algorithm to strike a balance between model robustness and sample efficiency using preference-heuristic parameter aggregation. The simulation results demonstrate that the proposed PDPPO algorithmic framework can implement interpretable single-agent decision control and multi-agent co-optimization processes. Furthermore, it exhibits competitive performance on various benchmark tests.
Sikai Lu, Yingfeng Cai, Yubo Lian, Long Chen 0003, Hai Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2025 Self-Distillation Attention for Efficient and Accurate Motion Prediction in Autonomous Driving
abstract
The accuracy and stability of motion prediction are crucial for the safe planning of autonomous driving systems. The widely used attention mechanisms effectively improve prediction accuracy. However, their computational cost grows quadratically with sequence length, presenting challenges for handling complex, large-scale scenarios. The attention patterns of motion prediction tasks exhibit significant data-related sparsity, indicating that not all scene elements are worth interaction. Therefore, this paper proposes a novel motion prediction framework that enhances the efficiency and stability of interaction fusion while achieving promising prediction accuracy. Firstly, an adaptive self-distillation attention module based on sparse interaction graph is designed. This module adaptively filters high-value sequences for each target to achieve sparse attention calculation and balance scene scale. Due to the lack of interaction labels in the original dataset, a self-distillation strategy is employed for model training. Secondly, a multi-stage dynamic anchor decoder is introduced that leverages the information filtering and aggregation capabilities of the sparse attention mechanism to improve prediction accuracy. The decoder dynamically updates the target anchor representations as the prediction progresses, ensuring consistency between interaction states and the prediction process in long-term forecasting. This approach effectively focuses attention calculation on the most relevant scene context fusion and trajectory decoding at each prediction stage. Validation results on Argoverse 1 and Argoverse 2 demonstrate that the proposed method achieves competitive accuracy while effectively reducing computational resource consumption. The proposed method can also serve as a plug-in that can be seamlessly incorporated into standard motion prediction pipelines to optimize scene interaction, making it more friendly to low-cost devices.
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Yubo Lian, Long Chen 0003
IEEE Trans. Intell. Transp. Syst.2
2025 A Multi-Agent Federated Reinforcement Learning-Based Cooperative Vehicle-Infrastructure Control Approach Framework for Roundabouts
abstract
In contrast to autonomous driving (AD) systems based on a single vehicle, the intelligent vehicle infrastructure cooperative system enables intelligent connected vehicles (ICVs) to perform active safety control and road collaboration management through vehicle-Infrastructure dynamic information interaction. However, complex intersections involve a wide range of interactions between agents, traffic participants and road infrastructures. The presence of time-varying traffic conditions and redundant traffic information may impede the efficiency and effectiveness of the vehicle-infrastructure system. Meanwhile, existing cooperative systems typically assume voluntary client access and data sharing, overlooking the information asymmetry challenges stemming from privacy awareness and the model interpretability challenges associated with data-driven approaches. This paper proposes a federated reinforcement learning (FRL) architecture for vehicle-infrastructure cooperative control is predicated on data and rule fusion, with the objective of considering privacy awareness and sample efficiency by employing a screening-based dual aggregation federated learning (FL) method. The proposed architecture enhances the interpretability of the algorithmic framework at four perspectives: modeling, training, optimization, and application. In this paper, we initially undertake training and performance comparisons on the OpenDD real traffic roundabout road driving dataset, and then extend the proposed method to more complex Nocrash benchmarks and real-world campus traffic environment for algorithm scalability validation.
Sikai Lu, Hai Wang 0003, Long Chen 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.4
2025 RTMDet-R: A Robust Instance Segmentation Network for Complex Traffic Scenarios
abstract
In complex traffic scenarios, several factors including lighting, weather, the size of the traffic participants, the distance between the traffic participants and the camera, and occlusions impact the features of the traffic participants. The impact of these factors is a huge challenge, especially for vision-based instance segmentation networks. To this end, this paper proposes an enhanced version of the RTMDet to promote the overall performance of instance segmentation in complex traffic scenarios. Firstly, an extended CSP-style backbone with large kernel convolutions of different kernel sizes is used to enhance the robustness of feature extraction capability, which contributes to obtaining more information about traffic objects of different scales. Secondly, a plugin pre-fusion module is designed to enhance the network’s robustness to multi-scale changes caused by distance changes. Additionally, instance kernel distinguish module is proposed to further highlight and distinguish different instance objects under poor lighting or weather and occlusion situations. Finally, the existing advanced image generation technology is used to expand the BDD100k dataset, enriching the dataset with severe scenarios. With an input resolution of$\mathbf {1280}\times \mathbf {720}$on the expanded BDD100K dataset, the proposed RTMDet-R achieves an accuracy of 25.4% mAP on the instance mask and 27.8% mAP on the instance box. This surpasses other similar models in terms of accuracy. Additionally, it maintains a good inference speed of 23.1 FPS, achieving the trade-off between accuracy and speed. Code and models are released athttps://github.com/GTrui6/RTMDet-R.git.
Hai Wang 0003, Qirui Qin, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.5
2025 MixedFusion: An Efficient Multimodal Data Fusion Framework for 3-D Object Detection and Tracking
abstract
The performance of environmental perception is critical for the safe driving of intelligent connected vehicles (ICVs). Currently, the most prevalent technical solutions are based on multimodal data fusion to achieve a comprehensive perception of the surrounding environment. However, existing fusion perception methods suffer from issues such as low sensor data utilization and unreasonable fusion strategies, which severely limit their performance in adverse weather conditions. To address these issues, this article proposes a novel multimodal data fusion framework called MixedFusion. In this framework, we introduce two innovative fusion strategies for the data characteristics of each sensor: high-level semantic guidance (HLSG) and multipriority matching (MPM). It not only realizes the efficient utilization of the multimodal data but also further realizes the complementary fusion between the multimodal data. We perform extensive experiments on the nuScenes and K-radar datasets. The experimental results demonstrate that the fusion framework proposed in this article significantly improves the performance of 3-D object detection and tracking in severe weather conditions.
Cheng Zhang 0031, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Neural Networks Learn. Syst.5
2024 Robust Control of Electric Power Steering System Based on Aligning Torque Model
abstract
The majority of previous studies on steering system control have concentrated on the torque associated with the lateral tire force. However, during parking maneuvers, the tire load and the front axle vertical displacement constitute the primary contributors to the torque that is responsible for aligning. Additionally, the nonlinearity imposed on the electric power steering (EPS) system arising from the required steering portability during parking maneuvers is another control issue that needs to be considered. Therefore, this paper focuses on the aligning torque model and the EPS system control during parking maneuvers. Firstly, an aligning torque model for parking maneuvers is developed. Then, an EPS system control strategy is presented, in which an observer and an integral sliding mode controller are developed with the objective of tracking desired steering torque and achieving robustness to disturbances. The simulations indicate that this developed model is capable of accurately predicting the aligning torque, and that this control strategy can well realize the desired steering feel during parking maneuvers.
Xing Xu 0002, Long Chen 0003, Yingfeng Cai
INDIN4
2024 VRSO: Visual-Centric Reconstruction for Static Object Annotation
abstract
As a part of the perception results of intelligent driving systems, static object detection (SOD) in 3D space provides crucial cues for driving environment understanding. With the rapid deployment of deep neural networks for SOD tasks, the demand for high-quality training samples soars. The traditional, also reliable, way is manual labelling over the dense LiDAR point clouds and reference images. Though most public driving datasets adopt this strategy to provide SOD ground truth (GT), it is still expensive and time-consuming in practice. This paper introduces VRSO, a visual-centric approach for static object annotation. Experiments on the Waymo Open Dataset show that the mean reprojection error from VRSO annotation is only 2.6 pixels, around four times lower than the Waymo Open Dataset labels (10.6 pixels). VRSO is distinguished in low cost, high efficiency, and high quality: (1) It recovers static objects in 3D space with only camera images as input, and (2) manual annotation is barely involved since GT for SOD tasks is generated based on an automatic reconstruction and annotation pipeline.
Chenyao Yu, Yingfeng Cai, Jiaxin Zhang 0014, Hui Kong 0001, Wei Sui
IROS2
2024 Unsupervised Cross-Scenario Abnormal Driving Behavior Recognition Using Smartphone Sensor Data
abstract
Accurately recognizing abnormal behavior of drivers (e.g., aggressive driving and fatigued driving) based on multivariate sensor data is vital for human-centric assistive driving systems. Existing data-driven deep learning models for abnormal driving behavior recognition (ADBR) achieve promising performance under specific driving scenes with sufficient labeled data. However, in the real world, dynamic driving scenes and unlabeled data pose a great challenge to the adaptability of models. In light of this, we put forward a novel unsupervised cross-scenario ADBR approach that can transfer domain knowledge in the source scenario with labeled data to the target scenario with only unlabeled data, thus considerably enhancing the adaptability of our model. Specifically, we first propose a feature extraction module that can obtain domain-shared and domain-specific features from raw sensor data derived from different driving scenes. Then, adversarial learning is presented to align the feature distribution of source and target domains to reduce the domain shift. A self-training strategy is further developed to boost the target domain classification performance by iteratively using the pseudo labels. Moreover, prediction uncertainty and ensemble classification are proposed to enhance the quality of pseudo labels. Extensive experiments on cross-scenario ADBR are conducted to evaluate the effectiveness of our model. The results manifest that our model significantly improves the recognition performance for the target domain and outperforms the competing algorithms.
Xiaobo Chen 0001, Yong Wang 0059, Xiaodong Sun 0001, Yingfeng Cai
IEEE Internet Things J.4
2024 Localization for Intelligent Vehicles in Underground Car Parks Based on Semantic Information
abstract
Global navigation satellite system (GNSS) signals cannot be received indoors, thus to deploy intelligent vehicles in underground car parks other localization methods are needed. In this paper, we use various carpark signs that are widely and uniformly distributed in underground parking lots as localization references. We propose a coarse-to-fine multiscale localization method that relies solely on vision sensors for underground parking lot localization based on a preconstructed lightweight node map. In coarse localization, we propose a semantic keyframe topological localization method to predict the localization range (candidate set of nodes). In node-level localization, we extract features by learning-based neural networks and construct a hybrid k-nearest neighbor (H-KNN) model to search for the closest node within the coarse localization results. In metric localization, we construct plane homography and perspective-n-point (PnP) models, allowing the vehicle’s pose (rotation and translation relative to the closest node) to be computed for refined localization. The proposed method has been tested in two underground parking lots of an office building and a shopping mall with different characteristics. Experimental results demonstrate that root mean square error (RMSE) is 0.38 m and the proposed method exhibits strong robustness in various scenarios.
Yicheng Li 0001, Yingfeng Cai, Zhixiong Li 0001, Miguel Ángel Sotelo
IEEE Trans. Intell. Transp. Syst.3
2024 DAIR-V2XReid: A New Real-World Vehicle-Infrastructure Cooperative Re-ID Dataset and Cross-Shot Feature Aggregation Network Perception Method
abstract
As an emerging research field, vehicle re-identification (Re-ID) can realize identity search between the vehicles, which plays an important role in the over-the-horizon perception of Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD). At present, due to the lack of data sets, the relevant research on Vehicle-Infrastructure Cooperative (VIC) Re-ID can only be evaluated in the cross-view monitoring test set which leads to the lack of persuasion of the research. Therefore, based on the DAID-V2X dataset of Tsinghua University, this paper constructs a VIC Re-ID dataset “DAIR-V2XReid” from real vehicle scenarios through vehicle-road end target tag association, thereby making it better applicable to the research of VIC Re-ID. Owing to different task scenarios, existing algorithms trained on monitoring test sets are unable to effectively complete the Re-ID task in this new dataset. Therefore, Cross-shot Feature Aggregation Network (CFA-Net) is also proposed in this paper, to tackle the case where a vehicle becomes unrecognizable due to a large change in its visual appearance across different cameras. Firstly, we put forward a camera embedding module and add it to the Backbone, to group different cameras and solve the problem of cross-shot perspective mutation. Secondly, in order to address the situation where background and vehicle division are not distinguishable, we propose a cross-stage feature fusion module, which integrates low-order semantics with high-order semantics. Finally, we use multi-directional attention network to achieve the final feature extraction. The experimental results show that our proposed CFA-Net method achieves new state-of-the-art in DAIR-V2XReid, with mAP of 58.47%.
Hai Wang 0003, Yaqing Niu, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.7
2024 RTMDet-MGG: A Multi-Task Model With Global Guidance
abstract
In this paper, we propose the concept of global guidance, design a global guidance structure based on a dual-scale global feature enhancement module, and construct a multi-task network (RTMDet-MGG) for road scene instance segmentation and drivable area segmentation in order to overcome the limitation of current multi-task networks in sharing features to different task branches. The proposed global guidance structure enhances the connection between different task branches in the multi-task network by sharing the enhanced global feature with different task branches through various fusion methods, thereby enhancing the multi-task network’s overall performance. With an input resolution of$\mathbf {640}\times \mathbf {360}$on the BDD100K dataset, RTMDet-MGG achieves an accuracy of 18.7% mAP (mean average precision) on the instance mask and 82.9% mIoU (mean intersection over union) on the drivable area, with an inference speed of 32.1 FPS (frames per second), which satisfies the requirements of real-time tasks. In addition, the algorithm has excellent scene generalization capabilities, and the mIoU of drivable area segmentation on our custom-built dataset for unstructured road drivable area segmentation reaches 93.3%.
Hai Wang 0003, Qirui Qin, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.5
2024 Real-Time Pedestrian Crossing Anticipation Based on an Action-Interaction Dual-Branch Network
abstract
Accurate anticipation of pedestrian crossing intentions is critical for preventing pedestrian-vehicle conflicts and ensuring road safety. This issue is a significant focus in intelligent transportation systems and autonomous driving. However, current approaches often face challenges, such as high computational costs due to complex scene understanding and inadequate consideration of spatiotemporal dependencies in pedestrian actions. To handle these challenges, we propose RAIDN (real-time action-interaction dual-branch network), which comprises the pedestrian action encoding and traffic-object interaction modules to anticipate pedestrian crossing intentions in real time. The pedestrian action encoding module employs a multi-scale graph transformer, efficiently extracting the intrinsic topology of long- and short-term action variations. This module effectively addresses issues of information redundancy in multi-channel graph convolution networks and local limitations in multi-scale temporal convolutions. Subsequently, the traffic-object interaction module introduces an interaction relation graph convolution network to excavate relevant traffic-object interactions, thereby shortening the prolonged scene semantic inference. Finally, global average pooling and attention layers fuse the action and interaction cues for real-time intention anticipation. The effectiveness of RAIDN has been validated on public datasets JAAD and PIE, achieving competitive metrics with Accuracy, ROC-AUC, F1-Score, Precision, and Recall rates of 0.89/0.92, 0.80/0.89, 0.66/0.85, 0.65/0.82, and 0.72/0.89 respectively. Notably, RAIDN demonstrates a remarkable inference time of just 0.28ms, outperforming other state-of-the-art methods and establishing its suitability for real-time applications in intelligent transportation and autonomous driving. Code is available athttps://github.com/wrysmile99/RAIG.
Zhiwen Wei, Chuan Hu 0003, Yingfeng Cai, Hai Wang 0003, Hongyu Hu
IEEE Trans. Intell. Transp. Syst.4
2024 Small Fixed-Wing Unmanned Aerial Vehicle Path Following Under Low Altitude Wind Shear Disturbance
abstract
The wind shear in low altitude can bring about time-varying and unknown disturbances, which makes small fixed-wing Unmanned Aerial Vehicle (UAV) hard to accurately follow desired curved paths. In this paper, a Vector Field (VF) based curved path following algorithm is designed for UAV to overcome the above difficulties. Firstly, the path following problem is mathematically formulated to give the kinematics model of UAV with constant airspeed. Secondly, the curved path following control laws under both constant and time-varying unknown wind disturbances are designed with VF. A ground velocity estimator is additionally designed to overcome the difficulties of unmeasurable ground velocity, and the Lyapunov stability of curved path following is analyzed. Finally, both simulations and field tests on a real small fixed-wing UAV are carried out to evaluate the algorithm performance in the presence of time-varying and unknown wind disturbances. Verification results demonstrate that the algorithm proposed in this paper outperforms other widely used path following algorithms and can effectively follow arbitrary curved paths.
Zhouyu Zhang 0002, Chenyuan He, Hongshen Chen, Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Haitong Li, Tongwei Lu
IEEE Trans. Intell. Transp. Syst.6
2024 TransFusion: Multi-Modal Robust Fusion for 3D Object Detection in Foggy Weather Based on Spatial Vision Transformer
abstract
A practical approach to realizing the comprehensive perception of the surrounding environment is to use a multi-modal fusion method based on various types of vehicular sensors. In clear weather, the camera and LiDAR can provide high-resolution images and point clouds that can be utilized for 3D object detection. However, in foggy weather, the propagation of light is affected by the fog in the air. Consequently, both images and point clouds become distorted to varying degrees. Thus, it is challenging to implement accurate detection in adverse weather conditions. Compared to cameras and LiDAR, Radar possesses strong penetrating power and is not affected by fog. Therefore, this paper proposes a novel two-stage detection framework called “TransFusion”, which leverages LiDAR and Radar fusion to solve the problem of environment perception in foggy weather. The proposed framework is composed of Multi-modal Rotate Region Proposal Network (MM-RRPN) and Multi-modal Refine Network (MM-RFN). Specifically, Spatial Vision Transformer (SVT) and Cross-Modal Attention Mechanism (CMAM) are introduced in the MM-RRPN to improve the robustness of the algorithm in foggy weather. Furthermore, Temporal-Spatial Memory Fusion (TSMF) module in MM-RFN is employed to fuse the spatial-temporal prior information. In addition, the Multi-branches Combination Loss function (MC-Loss) is designed to efficiently supervise the learning of the network. Extensive experiments were conducted on Oxford Radar RobotCar (ORR) dataset. The experimental results show that the proposed algorithm has excellent performance in both foggy and clear weather. Especially in foggy weather, the proposed TransFusion achieves 85.31mAP, outperforming all other competing approaches. The demo is available at:https://youtu.be/ugjIYHLgn98.
Cheng Zhang 0031, Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Robust Dimensionality Reduction via Low-rank Laplacian Graph Learning
abstract
Manifold learning is a widely used technique for dimensionality reduction as it can reveal the intrinsic geometric structure of data. However, its performance decreases drastically when data samples are contaminated by heavy noise or occlusions, which leads to unsatisfying data processing performance. We propose a novel robust dimensionality reduction method via low-rank Laplacian graph learning for classification and clustering tasks to solve the above problem. First, we construct a low-rank Laplacian graph by combining manifold learning and subspace learning. This graph can capture both global and local structural information of the data. And we introduce rank constraints for the Laplacian graph to make it more discriminative. Second, we put the learning of projection matrix and sample affinity graph into a unified framework. The projection matrix is embedded into a robust low-rank Laplacian graph so that the low-dimensional mapping of data can maintain the structural information in the graph well. Finally, we add a regularization term to the projection matrix to make it have the ability of both feature extraction and feature selection. Therefore, the proposed model can resist the interference of noise or data damage to learn the optimal projection to achieve better performance in dimensionality reduction through such a data dimensionality reduction joint framework. Comprehensive experiments on various benchmark datasets with varying degrees of occlusions or corruptions are carried out to evaluate the performance of the proposed method. Compared with the state-of-the-art dimensionality reduction methods in the literature, the experimental results are inspiring, showing our method’s effectiveness and robustness in classification and clustering, especially in object recognition scenarios with noise or occlusions.
Mingjian Cai, Xiangjun Shen, Stanley Ebhohimhen Abhadiomhen, Yingfeng Cai, Sirui Tian
ACM Trans. Intell. Syst. Technol.4
2023 Trajectory and Velocity Planning Method of Emergency Rescue Vehicle Based on Segmented Three-Dimensional Quartic Bezier Curve
abstract
In order to ensure the driving safety, comfort, stability and high mobility of emergency rescue vehicle, and considering the actual characteristics of emergency rescue vehicles like large vehicle width, high center of gravity and large turning radius, a trajectory and velocity planning method for collision avoidance based on segmented three-dimensional quartic Bezier curve is proposed. The velocity information is regarded as the vertical coordinates of corresponding trajectory coordinates, and the three-dimensional Bezier curve is used for vehicle trajectory planning, so as to make the velocity information and vehicle trajectory coordinates correspond one by one in planning results. In addition, in order to realize the effective fit between the multiple optimization objectives and the actual needs in intelligent driving, the segmented Bezier curve is used to design the vehicle trajectory, and the vehicle trajectory is divided into collision-avoidance process and lane-alignment process, so as to realize the dynamic optimization of trajectory design objectives. The results show that the proposed trajectory and velocity planning method based on segmented three-dimensional quartic Bezier curve can effectively improve the mobility of trajectory planning results, improve lane-changing efficiency and reduce lane-changing time while ensuring the design objectives of vehicle safety, comfort and stability.
Te Chen, Yingfeng Cai, Long Chen 0003, Xing Xu 0002
IEEE Trans. Intell. Transp. Syst.2
2023 Driving Style Feature Extraction and Recognition Based on Hyperdimensional Computing and Semi-Supervised Twin Projection Vector Machine
abstract
Driving style recognition is one of the most crucial requirements for human-centric autonomous and assistive driving systems. Existing studies either require a large amount of labeled data or suffer from hand-crafted features, thus hindering the practical application of these methods. To address the above challenges, this paper proposes a driving style recognition approach from the perspective of brain-inspired hyperdimensional computing and semi-supervised learning. Specifically, raw sensor signals are treated as multivariate time series which are encoded by hyperdimensional computing, thus generating a large holistic feature representation. Then, considering the high dimension of the resulting representation, we put forward a semi-supervised twin projection vector machine (SSTPVM) model that can take full advantage of unlabeled data and jointly optimize multiple projection directions tailored for dimension reduction. In addition, a heuristic method based on particle swarm optimization is developed for the parameter selection of SSTPVM. Finally, extensive experiment comparisons with other related methods are performed on naturalistic driving style data. The results show that hyperdimensional computing is rather suitable for semi-supervised learning while our proposed SSTPVM can significantly improve recognition performance even with a small portion of labeled data.
Xiaobo Chen 0001, Yuxiang Gao, Haoze Yu, Hai Wang 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.5
2023 Monocular Road Scene Bird's Eye View Prediction via Big Kernel-Size Encoder and Spatial-Channel Transform Module
abstract
A detailed representation of the surrounding road scene is crucial for an autonomous driving system. yellow The camera-based Bird’s Eye View map has been a popular solution to present the surrounding information, due to its low cost and rich spatial context information. Most of the existing methods predict the BEV map based on the depth-estimation or the trivial homography method, which may cause the error propagation and the absence of content. To overcome these drawbacks, we propose a novel end-to-end framework that employs the front monocular image to predict the road layout and vehicle occupancy. In particular, to capture the long-range feature, we redesign a CNN encoder with a large kernel size to extract the image features. For reducing the big difference between the front image features and the top-down features, we propose a novel Spatial-Channel projection module to convert the front map into the top-down space. Additionally, concerning the correlation between front view and top-down view, we propose the Dual Cross-view Transformer module to refine the top-down view feature maps and strengthen the transformation. Extensive evaluations on the KITTI and Argoverse datasets present that the proposed model achieves the state-of-the-art results for both datasets. Furthermore, the proposed model runs in 37 FPS on a single GPU, demonstrating the generation of a real-time BEV map. The code will be published at https://github.com/raozhongyu/BEV_LKA.
Zhongyu Rao, Hai Wang 0003, Long Chen 0003, Yubo Lian, Yilin Zhong, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.7
2023 DYC Design for Autonomous Distributed Drive Electric Vehicle Considering Tire Nonlinear Mechanical Characteristics in the PWA Form
abstract
This paper presents a novel direct yaw moment control (DYC) strategy for autonomous distributed drive electric vehicle (DDEV) to improve path-following driving stability under critical maneuvers. Due to the fact that the tire mechanical characteristics show highly nonlinear under those critical conditions, the piecewise affine (PWA) identification method is chosen in this work to model the tire nonlinear mechanical characteristics since it shows best model accuracy/simplicity compromise for system control design. On this basis, to properly reflect the coupling behaviors between longitudinal motion and lateral motion, the vehicle dynamics model which considers the tire nonlinear mechanical characteristics in the PWA form is established. By analyzing the phase plane of vehicle nonlinear model, a novel external yaw moment calculation method based on global fast terminal sliding mode (GFTSM) control method is applied to the upper layer of DYC architecture for eliminating the errors of yaw rate and sideslip angle. Then, different weight coefficients are allocated to obtain the final external yaw moment. In the lower controller, the optimal four-wheel torque is distributed according to a concise and accurate algorithm in which adaptive weight coefficients are considered. The simulation results verify the effectiveness of the proposed control strategy in improving the driving stability performance of autonomous DDEV against two other control strategies under three extreme conditions.
Zhenqiang Quan, Yingfeng Cai, Long Chen 0003, Shaoyi Bei
IEEE Trans. Intell. Transp. Syst.4
2023 Sparse U-PDP: A Unified Multi-Task Framework for Panoptic Driving Perception
abstract
This paper proposes a multi-task unified decoding framework for the panoptic driving perception, Sparse U-PDP. This framework’s primary objective is to combine vehicle object detection, lane detection and drivable area segmentation tasks. This paper mainly builds a multi-task unified decoder and further explores whether the potential connection between multi-tasks can improve the robustness of the model. Experiments demonstrate that the proposed Sparse U-PDP outperforms the present state-of-the-art multi-task model in terms of accuracy. The primary contributions of this work are as follows: First, we present the approach to unified multi-task representation. We abstract the multi-task into the “dynamic convolution kernels” representation form to build a highly unified multi-task decoder. Second, we use the proposed dynamic interaction module to establish different feature sampling pipelines for various task features. Lastly, our model is verified on the BDD100K dataset, where we achieve an AP50 of 84.1 in the “vehicle” category, 32.0 IoU in the lane detection, and 93.0 mIoU in the drivable area segmentation with the helper of CSP-Darknet. That is to say; Sparse U-PDP verifies that a more unified task representation form can implicitly increase the mutual help between different task branches.
Hai Wang 0003, Meng Qiu, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.3
2023 CenterPoint-SE: A Single-Stage Anchor-Free 3-D Object Detection Algorithm With Spatial Awareness Enhancement
abstract
Real-time and accurate 3-D object detection is one of the foundational technologies for environmental perception in autonomous vehicles. However, the existing second-stage anchor-based 3-D object detection algorithms have high accuracy, but they are challenging in terms of computation complexity and latency. Due to poor perception of spatial features, the accuracy of the existing single-stage anchor-free detection algorithms with low latency are difficult to be implemented into autonomous vehicles. Therefore, we focus on enhancing the spatial perception ability of the anchor-free detection network based on CenterPoints. In this paper, we propose a single-stage anchor-free 3-D object detector CenterPoint-Space-Enhancement (CenterPoint-SE) algorithm and construct an efficient 3-D backbone network to extract fine-grained spatial geometric features by introducing a spatial attention mechanism and residual structure. At the same time, a powerful spatial semantic feature fusion module, the enhancement of feature fusion (EF-Fusion), is designed. In addition, we add a lightweight IoU prediction branch to improve the algorithm’s perception of various object sizes. Finally, we add a foreground point segmentation auxiliary training branch to enable the 3-D backbone to obtain object boundary features. We use the ONCE dataset to train and validate the proposed model, and the results showed that the proposed CenterPoint-SE achieves 70.33 mAP and an inference speed of 17.15 FPS, outperforming other methods.
Hai Wang 0003, Le Tao, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Map-based localization for intelligent vehicles from bi-sensor data fusion
Yicheng Li 0001, Yingfeng Cai, Zhixiong Li 0001, Shizhe Feng, Hai Wang 0003, Miguel Ángel Sotelo
Expert Syst. Appl.2
2022 Pedestrian Motion Trajectory Prediction in Intelligent Driving from Far Shot First-Person Perspective Video
abstract
Pedestrian motion trajectory prediction is an important task in intelligent driving, and it can provide a valuable reference for the subsequent path decision of intelligent driving. However, so far, there are only a few models in the field of specific pedestrian motion track prediction in intelligent driving from far shot first-person perspective video. To accomplish this task, we proposed a deep learning model for pedestrian motion trajectory prediction from far shot first-person perspective video with four key innovations: a) A macroscopic pedestrian trajectory prediction module is established under the close correlation between neighboring frames to estimate the pedestrian motion track on the whole; b) A relative motion transformation module of vehicle-mounted camera is designed to consider the effect of vehicle-mounted camera’s ego-motion on the pedestrian motion track; c) We set up a circular training module to maintain the number of parameters in our model to simplify and reduce the size of model; d) A new far shot first-person pedestrian motion dataset under intelligent driving is specifically established to train and test the proposed model. The above four modules are integrated into the proposed deep learning model, which achieves state-of-the-art results for predicting pedestrian motion trajectory from both far and close shot first-person perspective video.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.1
2022 An Automatic Vehicle Avoidance Control Model for Dangerous Lane-Changing Behavior
abstract
This paper proposes a new avoidance control model for automatic vehicle in facing dangerous lane-changing behavior. Firstly, the new lane-changing probability factor based on Gaussian-mixture-based hidden Markov model is constructed to predict the lateral-vehicle lane-changing probability and output the pre-control parameters. Secondly, the back propagation neural network avoidance model, which combined with driver’s avoidance behavior, is developed for achieving the instantaneous collision avoidance control. Moreover, the optimal solution between control parameters and vehicle stability is obtained by using linear quadratic regulator. Finally, the accuracy of the avoidance model is verified by the semi-physical driver-in-the-loop simulation based on PreScan/Simulink. Results show that the automatic vehicle with the proposed avoidance model can accurately and effectively take pre-braking and micro-steering behavior. The proposed model can greatly reduce vehicle collision probability and effectively take both safety and comfort of collision avoidance into account. In addition, the robustness of the control model under different network penetration is discussed.
Sensen Cong, Wensa Wang, Jun Liang 0004, Long Chen 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.5
2022 Robust Target Recognition and Tracking of Self-Driving Cars With Radar and Camera Information Fusion Under Severe Weather Conditions
abstract
Radar and camera information fusion sensing methods are used to solve the inherent shortcomings of the single sensor in severe weather. Our fusion scheme uses radar as the main hardware and camera as the auxiliary hardware framework. At the same time, the Mahalanobis distance is used to match the observed values of the target sequence. Data fusion based on the joint probability function method. Moreover, the algorithm was tested using actual sensor data collected from a vehicle, performing real-time environment perception. The test results show that radar and camera fusion algorithms perform better than single sensor environmental perception in severe weather, which can effectively reduce the missed detection rate of autonomous vehicle environment perception in severe weather. The fusion algorithm improves the robustness of the environment perception system and provides accurate environment perception information for the decision-making system and control system of autonomous vehicles.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Hongbo Gao 0001, Yunyi Jia, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Exploring the Impact of the Takeover Time for Conditionally Automated Driving Vehicles on Traffic Flow in Highway Merging Area
abstract
Human-machine handover of conditionally automated driving vehicles (CADVs) significantly affects traffic safety. Therefore, a simulation modeling was conducted for the traffic flow mixed with manual driving vehicles, fully automated driving vehicles (FADVs), and CADVs under different takeover times to unravel the impact of CADVs on traffic flow. The results showed that different takeover times significantly affected traffic flow stability. A moderate takeover time allows the driver to complete the takeover quickly with sufficient observation of the surrounding traffic conditions and mitigate the adverse effects of CADVs takeover transition on traffic flow and improve traffic flow safety. Taking moderate takeover time (7s) as the given takeover time, we developed a traffic flow model, and it is found that increasing the total penetration rates of CADVs and FADVs or that of CADVs alone will expand the traffic flow stability area. Moreover, the improving effects of traffic flow stability increase with the value of both penetration rates. This research can be a reference for the safety analysis of heterogeneous traffic flow mixed with CADVs.
Qingchao Liu, Yingfeng Cai, Long Chen 0003
IEEE Trans. Intell. Transp. Syst.3
2022 SFNet-N: An Improved SFNet Algorithm for Semantic Segmentation of Low-Light Autonomous Driving Road Scenes
abstract
In recent years, considerable progress has been made in semantic segmentation of images with favorable environments. However, the environmental perception of autonomous driving under adverse weather conditions is still very challenging. In particular, the low visibility at nighttime greatly affects driving safety. In this paper, we aim to explore image segmentation in low-light scenarios, thereby expanding the application range of autonomous vehicles. The segmentation algorithms for road scenes based on deep learning are highly dependent on the volume of images with pixel-level annotations. Considering the scarcity of labeled large-scale nighttime data, we performed synthetic data collection and data style transfer using images acquired in daytime based on the autonomous driving simulation platform and generative adversarial network, respectively. In addition, we also proposed a novel nighttime segmentation framework (SFNET-N) to effectively recognize objects in dark environments, aiming at the boundary blurring caused by low semantic contrast in low-illumination images. Specifically, the framework comprises a light enhancement network which introduces semantic information for the first time and a segmentation network with strong feature extraction capability. Extensive experiments with Dark Zurich-test and Nighttime Driving-test datasets show the effectiveness of our method compared with existing state-of-the art approaches, with 56.9% and 57.4% mIoU (mean of category-wise intersection-over-union) respectively. Finally, we also performed real-vehicle verification of the proposed models in road scenes of Zhenjiang city with poor lighting. The datasets are available athttps://github.com/pupu-chenyanyan/semantic-segmentation-on-nightime.
Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.3
2022 NLS Based Hierarchical Anti-Disturbance Controller for Vehicle Platoons With Time-Varying Parameter Uncertainties
abstract
Cooperative adaptive cruise control (CACC) is a promising technology for vehicle platoons to increase roadway capacity. This paper proposes a networked Lagrange system (NLS) based CACC dynamic model based on which a hierarchical anti-disturbance controller is developed to solve the stability problem of CACC in vehicle platoons with time-varying parameter uncertainties, external disturbances, and directional dynamic communication topology. First, a hierarchical anti-disturbance controller for NLS is constructed which comprises three layers, these are, the adaptive smooth estimator-control layer, the distributed classifying amplitude-related layer, and the anti-disturbance control layer, in which the parameter uncertainties are estimated in the adaptive smooth estimator-control layer, the disturbance classification and distribute control are executed in the distributed classifying amplitude-related layer and the anti-disturbance control layer respectively. In addition, the proposed controller is extended to address the stability problem of CACC in vehicle platoons with time-varying parameter uncertainties, external disturbances, and directional dynamic communication topology conditions, which shows the versatility of the controller. Finally, comparison studies and simulation results are provided to demonstrate the effectiveness, significance, and advantages of the presented controllers.
Wensa Wang, Jun Liang 0004, Chaofeng Pan, Yingfeng Cai, Long Chen 0003
IEEE Trans. Intell. Transp. Syst.4
2022 Crossing or Not? Context-Based Recognition of Pedestrian Crossing Intention in the Urban Environment
abstract
The booming self-driving cars need to understand the behaviors of other road users for better performance. Recognizing pedestrians’ crossing intentions is one of the most critical capabilities of self-driving cars to guarantee safe operations in the urban environment. Some researchers try to predict the future trajectories of pedestrians to avoid a potential collision. Others extract skeleton features from pedestrians to detect specific actions related to the crossing intention. All these methods either neglect the abundant appearance information of pedestrians, or work poorly under severe conditions, e.g., pedestrians standing far away, dim light, and occlusions. In this work, a pedestrian crossing intention recognition (PCIR) framework is proposed to recognize pedestrians’ crossing intention. A module for searching targets of interest is introduced to find the pedestrians who are likely to cross the streets, and perform scene perception simultaneously. An action recognition module utilizes a 3D convolutional neural network (CNN) to automatically extract spatiotemporal features that imply early actions before crossing, like limb movements. A distance encoding module makes full use of the contextual cues, e.g., distances between pedestrians and the oncoming vehicle, local traffic scenes around pedestrians, and vehicle speeds, to improve the recognition accuracy obtained from the action recognition module only. Finally, the PCIR framework fuses these factors to predict pedestrians’ crossing intentions by minimizing a focal loss. Experimental evaluations verify the effectiveness of the proposed method in recognizing pedestrians’ crossing intentions. Comparisons with the skeleton-based methods reveal the robustness of the PCIR framework in the urban environment.
Weiqin Zhan, Ching-Yao Chan, Yingfeng Cai, Nan Wang 0013
IEEE Trans. Intell. Transp. Syst.5
2022 DLnet With Training Task Conversion Stream for Precise Semantic Segmentation in Actual Traffic Scene
abstract
Many successful semantic segmentation models trained on certain datasets experience a performance gap when they are applied to the actual scene images, expressing weak robustness of these models in the actual scene. The training task conversion (TTC) and domain adaption field have been originally proposed to solve the performance gap problem. Unfortunately, many existing models for TTC and domain adaptation have defects, and even if the TTC is completed, the performance is far from the original task model. Thus, how to maintain excellent performance while completing TTC is the main challenge. In order to address this challenge, a deep learning model named DLnet is proposed for TTC from the existing image dataset-based training task to the actual scene image-based training task. The proposed network, named the DLnet, contains three main innovations. The proposed network is verified by experiments. The experimental results show that the proposed DLnet not only can achieve state-of-the-art quantitative performance on four popular datasets but also can obtain outstanding qualitative performance in four actual urban scenes, which demonstrates the robustness and performance of the proposed DLnet. In addition, although the proposed DLnet cannot achieve outstanding performance in real time, it can still achieve a moderate performance in real time, which is within an acceptable range.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Trajectory prediction of cyclist based on dynamic Bayesian network and long short-term memory model at unsignalized intersections
Hongbo Gao 0001, Hang Su 0001, Yingfeng Cai, Renfei Wu, Zhengyuan Hao, Yongneng Xu, Jianqing Wang, Zhijun Li 0001, Zhen Kan
Sci. China Inf. Sci.3
2021 Creating navigation map in semi-open scenarios for intelligent vehicle localization using multi-sensor fusion
Yicheng Li 0001, Yingfeng Cai, Reza Malekian, Hai Wang 0003, Miguel Ángel Sotelo, Zhixiong Li 0001
Expert Syst. Appl.2
2021 Multi-Target Pan-Class Intrinsic Relevance Driven Model for Improving Semantic Segmentation in Autonomous Driving
abstract
At present, most semantic segmentation models rely on the excellent feature extraction capabilities of a deep learning network structure. Although these models can achieve excellent performance on multiple datasets, ways of refining the target main body segmentation and overcoming the performance limitation of deep learning networks are still a research focus. We discovered a pan-class intrinsic relevance phenomenon among targets that can link the targets cross-class. This cross-class strategy is different from the latest semantic segmentation model via context where targets are divided into an intra-class and inter-class. This paper proposes a model for refining the target main body segmentation using multi-target pan-class intrinsic relevance. The main contributions of the proposed model can be summarized as follows: a) The multi-target pan-class intrinsic relevance prior knowledge establishment (RPK-Est) module builds the prior knowledge of the intrinsic relevance to lay the foundation for the following extraction of the pan-class intrinsic relevance feature. b) The multi-target pan-class intrinsic relevance feature extraction (RF-Ext) module is designed to extract the pan-class intrinsic relevance feature based on the proposed multi-target node graph and graph convolution network. c) The multi-target pan-class intrinsic relevance feature integration (RF-Int) module is proposed to integrate the intrinsic relevance features and semantic features by a generative adversarial learning strategy at the gradient level, which can make intrinsic relevance features play a role in semantic segmentation. The proposed model achieved outstanding performance in semantic segmentation testing on four authoritative datasets compared to other state-of-the-art models.
Yingfeng Cai, Hai Wang 0003, Zhixiong Li 0001
IEEE Trans. Image Process.1
2021 Visual Map-Based Localization for Intelligent Vehicles From Multi-View Site Matching
abstract
Accurate localization is a crucial step for intelligent vehicles (IVs). And vision-based localization methods are promising due to its good accuracy and low cost. However, vision-based methods are usually not robust enough due to the errors of matching similar road scenarios. In this paper, we proposed a visual map-based localization method, called multi-view site matching (MVSM). We proposed using two camera views (i.e., downward-view and front-view) to construct visual map. The visual map consists of a serial of nodes. Each node encodes the features of the road, the 2D structure, and the poses of the vehicle. Based on the constructed visual map, we proposed a multi-scale method for accurate vehicle localization. In coarse localization, we adopt a topological model to obtain a set of candidate nodes from visual map. Furthermore, holistic features from front view are matched within the candidates such that the best matched node is determined for image-level localization. In metric localization, the best matched is first verified with the local features from downward view. And the vehicle pose is finally computed by utilizing the 2D structure from the verified nodes in the map. In the experiment, the proposed MVSM method has been tested with actual field data covering different pavement types in different seasons. The proposed MVSM method can achieve less than 0.20m mean localization errors. Compared to existing vision-based methods, the proposed method utilizes two views to enhance image-level localization and 2D pavement structure to improve metric localization so as to greatly improve the overall localization performance.
Yicheng Li 0001, Zhaozheng Hu, Yingfeng Cai, Huawei Wu, Zhixiong Li 0001, Miguel Ángel Sotelo
IEEE Trans. Intell. Transp. Syst.3
2020 Robust crowd counting based on refined density map
Jinmeng Cao, Nan Wang 0013, Hai Wang 0003, Yingfeng Cai
Multim. Tools Appl.5
2020 A Novel Saliency Detection Algorithm Based on Adversarial Learning Model
abstract
The traditional salient object detection models can be divided into several classes based on the low-level features of images and contrast between the pixels. This paper proposes an adversarial learning model (ALM) that includes the generative model and discriminative model. The ALM uses the original image as an input of the generative model to extract the high-level features and forms an initial salient map. Then, the discriminative model is utilized to compare differences in the features between the initial salient map and the ground truth, and the obtained differences are sent to the convolutional layers of the generative model to adjust the parameters for the generative model updating. Due to the serial-iterative adjustment, the salient map of the generative model becomes more similar to the ground truth. Lastly, the ALM forms the salient map fused with the super-pixels by enhancing the color and texture features, so the final salient map is obtained. The ALM is not limited to the color and texture features; on the contrary, it fuses multiple features and achieves good results in the salient target extraction. The experimental results show that ALM performs better than the other ten state-of-the-art models on three different datasets. Thus, the proposed ALM is widely applicable to the salient target extraction.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Image Process.1
2018 Graph regularized local self-representation for missing value imputation with applications to on-road traffic sensor data
Xiaobo Chen 0001, Yingfeng Cai, Qiaolin Ye, Lei Chen 0011
Neurocomputing2
2018 Salient object detection based on multi-scale contrast
Hai Wang 0003, Yingfeng Cai, Long Chen 0003
Neural Networks3
2017 Ensemble correlation-based low-rank matrix completion with applications to traffic data imputation
Xiaobo Chen 0001, Zhongjie Wei, Jun Liang 0004, Yingfeng Cai, Bob Zhang 0001
Knowl. Based Syst.5
2016 Discriminative sparsity preserving graph embedding
abstract
In this paper, we propose a new dimensionality reduction method called discriminative sparsity preserving graph embedding (DSPGE). Unlike many existing graph embedding methods such as locality preserving projections (LPP) and sparsity preserving projections (SPP), the aim of DSPGE is to preserve the sparse reconstructive relationships of data while simultaneously capture the geometric and discriminant structure of data in the embedding space. Through the sparse reconstruction and class-specific adjacent graphs, DSPGE characterizes the intra-class and inter-class sparsity preserving scatters, seeking to achieve the optimal projections that simultaneously maximize the inter-class sparsity preserving scatter and minimize intra-class sparsity preserving scatter. The effectiveness of the proposed DSPGE is demonstrated on two popular face databases, compared to up-to-date methods. The experimental results show that DSPGE outperforms the competing methods with the satisfactory classification performance.
Jianping Gou, Lan Du 0002, Keyang Cheng, Yingfeng Cai
CEC4
2016 Multilevel framework to handle object occlusions for real-time tracking
abstract
This study proposes an efficient method to handle the object occlusions seen in monocular traffic image sequences. The motivation of this study is different methods perform differently in occlusion segmentation and the authors’ idea is to use a situation‐driven approach to aggregate different methods in order to get a good performance. This study classifies occlusion into four categories according to the foreground situation and a multilevel occlusion handling framework is utilised. First, the image segmentation algorithm based on convex hull analysis is utilised for intra‐frame level occlusion segmentation. The segmentation algorithm is established by the compactness ratio and interior distance ratio of the foreground. Second, an online sample‐based classification algorithm is utilised for tracking level occlusion segmentation. Training samples are extracted from the historical frames before occlusion and testing samples are extracted from the current frame by an adaptive searching strategy. The segmentation of occlusion is transferred into the online classification of testing samples. Such algorithm is established by the similarity and coherence of target's property between continuous frames. Experiments on video sequences illustrate the good performance of the proposed method under different conditions with low computational cost.
Yingfeng Cai, Hai Wang 0003, Xiaobo Chen 0001, Long Chen 0003
IET Image Process.1
2016 Occluded vehicle detection with local connected deep model
Hai Wang 0003, Yingfeng Cai, Xiaobo Chen 0001, Long Chen 0003
Multim. Tools Appl.2
2015 Discriminant feature extraction for image recognition using complete robust maximum margin criterion
Xiaobo Chen 0001, Yingfeng Cai, Long Chen 0003
Mach. Vis. Appl.2