Hai Wang 0003

dblp:59/3767-3 · DBLP profile ↗
← Back
43ranked-venue papers
8as first author
38since 2021 · last 2026
0000-0002-9136-8091ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 6 first-author · 24 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Computer networks · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Electro-optical and infrared multi-sensor fusion based airborne target perception: A unified framework
Zhouyu Zhang 0002, Chenyuan He, Yingfeng Cai, Long Chen 0003, Hai Wang 0003, Can Zhong
Eng. Appl. Artif. Intell.5
2026 End-to-End Autonomous Driving: From Classic Paradigm to Large Model Empowerment - A Comprehensive Survey
abstract
In recent years, autonomous driving technology has witnessed rapid global development, offering effective solutions to increasingly severe challenges related to traffic congestion and road safety. Among various paradigms, end-to-end autonomous driving has emerged as a promising alternative to traditional modular systems, owing to its streamlined architecture, enhanced decision consistency, and superior generalization capabilities. This survey provides a comprehensive review of the evolution and core technologies of end-to-end autonomous driving. It emphasizes the applications and developments of imitation learning (IL), reinforcement learning (RL), and other paradigms. Furthermore, it highlights the emerging paradigm empowered by foundation models, such as large language models (LLMs) and vision-language models (VLMs), and systematically categorizes recent advances in planning, reasoning, data generation, and scene understanding. In light of persistent challenges such as multi-modal fusion complexity, low sample efficiency, and safety risks, this survey summarizes representative research efforts and corresponding solutions. Finally, it outlines future directions, including the integration of world models (WMs) for unified data generation and inference optimization, the advancement of foundation model architectures toward modular design, sparse activation mechanisms, and knowledge distillation, and the realization of deployable, transferable, and unified multi-modal frameworks. This survey aims to serve as a comprehensive theoretical and technical reference for researchers in end-to-end autonomous driving, facilitating its evolution toward higher performance, stronger generalization, and enhanced safety assurance.
Sikai Lu, Xinhe Chen, Shunyao Zhang 0001, Qingchao Liu, Long Chen 0003, Hai Wang 0003, Yingfeng Cai
IEEE Internet Things J.8
2026 V2I-GridCalib: A Real-Time Online Calibration for Vehicle-to-Infrastructure Systems
abstract
Real-time spatial alignment between vehicles and roadside sensors is essential for unlocking collaborative perception. However, existing methods often struggle in dynamic environments due to observational scale mismatch and high latency. This paper proposes a novel edge-assisted framework for robust, real-time, and target-free online V2I calibration. First, a Directional Handshake mechanism pre-selects task-relevant point cloud subsets to reduce computational and transmission overhead. Second, a distance-adaptive rendering strategy generates scale-optimized rasterized views to explicitly address scale variation based on the vehicle's distance. Finally, a particle filter estimates the 6-DoF pose by fusing complementary 2D-3D reprojection and 2D-2D epipolar geometric constraints. Extensive experiments in real-world intersections and high-fidelity simulations validate the system. The framework maintains trajectory consistency errors below 5.5% for paths up to 400 meters and achieves centimeter-level accuracy with 5.2 cm translation and 0.3° rotation errors in static tests. Crucially, the system consistently meets stringent real-time constraints of less than 66 ms for V2X applications.
Zhuoyi Jiang, Yingfeng Cai, Yicheng Li 0001, Qingchao Liu, Long Chen 0003, Hai Wang 0003
IEEE Internet Things J.7
2026 ProtoAug: Prototype-Guided Uncertainty-Aware Augmentation for Long-Tail Motion Prediction
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yang Wang 0003, Xinxin Zuo
IEEE Internet Things J.3
2026 Mul-VMamba: Multimodal semantic segmentation using selection-fusion-based vision-Mamba
Yuanhui Guo, Yi Liu 0038, Hai Wang 0003, Chuan Hu 0003
Knowl. Based Syst.5
2026 Learning discrete latent representations for scene-guided multi-modal motion prediction
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Long Chen 0003
Pattern Recognit.3
2026 Low-Cost Monocular End-to-End Autonomous Driving With Enhanced Interpretability
abstract
In recent years, end-to-end autonomous driving has achieved remarkable progress, yet challenges remain in achieving efficient deployment. Existing methods often rely on multi-modal sensor fusion, which incurs high costs and computational demands, limiting their large-scale adoption in low-cost scenarios. Reducing system cost while maintaining performance has consistently been a key focus of the industry. Moreover, the opaque nature of end-to-end models, due to the lack of intermediate representations, limits interpretability and hampers root-cause analysis when misjudgments occur in complex environments, thereby posing safety risks. To address these challenges, this paper proposes a lightweight and interpretable end-to-end autonomous driving architecture that utilizes only monocular camera images as input. The aim is to enhance decision-making efficiency and trajectory planning accuracy while ensuring system safety. The proposed architecture comprises two core components. First, in the scene perception part, a depth information supervision module is introduced to enable hierarchical geometric perception, and an online vector map construction module is designed to build structured environmental representations at low cost. This replaces high-definition maps with explicit constraint expressions, thereby improving model interpretability and decision-making reliability. Second, in the planning part, a dual-layer prediction framework is developed to decouple speed and waypoint prediction. It independently forecasts vehicle velocity and waypoints, while an optimized decoder structure produces fine-grained spatiotemporal trajectories, enhancing planning capabilities in complex traffic scenarios. Experimental results on multiple benchmarks from the CARLA Leaderboard 1.0 demonstrate that the proposed method significantly improves system efficiency while maintaining advanced decision-making performance. It outperforms existing multi-sensor fusion approaches, showcasing strong potential for practical deployment.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Jianmei Lei
IEEE Trans. Intell. Transp. Syst.3
2026 Scenario-Driven Federated Reinforcement Learning Autonomous Driving System With High Generalizability
abstract
Data-driven autonomous driving (AD) systems typically use incremental optimization based on various scenarios to improve driving safety and overall performance. However, existing algorithms lack the capability to integrate information across scenarios and optimize synchronously. Incremental optimization for different scenarios often results in catastrophic forgetting during continual learning, reducing the system’s generalizability and reliability in complex driving scenarios. In this paper, we propose a scenario-driven federated reinforcement learning (FRL) AD system with broad applicability. Our proposed cross-scenario optimization uses a federated architecture to replace the data-driven incremental learning process, allowing specific experiences to be shared across different expert data distributions. This approach develops highly generalizable imitation learning (IL) experts and effectively addresses the problem of catastrophic forgetting in continual learning. Furthermore, our proposed dynamic driving suggestions provide multifaceted guidance to reinforcement learning (RL) students, such as feature extraction, reward function modeling, loss function construction, and population-based optimization. This guidance specifically targets the challenges of high-dimensional state space representation and objective alignment in RL. Experimental results on various benchmarks demonstrate that our proposed framework outperforms current state-of-the-art (SOTA) multimodal fusion algorithms. Notably, our system achieves highly reliable control performance in complex traffic environments with only a single monocular camera.
Sikai Lu, Yingfeng Cai, Hai Wang 0003, Qingchao Liu
IEEE Trans. Intell. Transp. Syst.3
2026 PCIP: Toward Real-Time Pedestrian Crossing Intention Prediction With Adaptive Fusion and Temporal Dynamics Encoding
abstract
Pedestrian crossing intention prediction is critical for autonomous vehicles. Analyzing pedestrian motion characteristics provides a preemptive decision-making basis for the planning module. The prediction accuracy directly affects the timeliness of collision avoidance mechanisms in vehicles, especially in mixed traffic scenarios, which is crucial for the safety and reliability of autonomous driving systems. Existing methods struggle to jointly achieve accuracy and real-time performance due to motion-unaware representations and inflexible fusion, which introduce redundancy and latency on embedded platforms. To tackle this problem, we propose PCIP, which leverages adaptive multimodal fusion to achieve efficient and robust crossing intention prediction. PCIP proposes three key contributions: 1) The dynamic selection of the quasi-polar coordinate system accurately characterizes the relative motion trends between pedestrians and vehicles through adaptive multi-pole selection. 2) The Koopman-based CI-TDE module linearizes the spatiotemporal dynamics of skeletal actions, significantly enhancing the expressive capability of temporal features that describe pedestrian motions. 3) A multi-scale dynamic weighting strategy based on distance and angle effectively suppresses noisy modalities and enhances pivotal features. Experiments show that PCIP achieves intention classification accuracies of 89% and 93% on the JAAD and PIE datasets, respectively, and achieves a real-time inference speed of 0.5ms per frame on an embedded platform (NVIDIA Jetson AGX). This study provides a lightweight and robust solution for pedestrian intention prediction in complex traffic scenarios.
Chuan Hu 0003, Hai Wang 0003
IEEE Trans. Intell. Transp. Syst.4
2026 RAMACoDrive: A Real-Time Asynchronous Framework for Cooperative Perception With Realistic V2X Communication
abstract
Recent advances in cooperative driving have garnered significant attention, as it is believed to have greater potential than single-vehicle systems in complex environments. However, gaps in current research frameworks impede its further development. Dataset-based frameworks provide a limited perspective, failing to capture the dynamic, asynchronous, and interactive complexities of the real world. Simulation-based frameworks often employ synchronous implementation and oversimplified communication models, impeding meaningful real-time performance evaluation in asynchronous environments. To this end, we propose RAMACoDrive, a novel real-time asynchronous framework for cooperative perception research. RAMACoDrive implements complete process separation, allowing agents and their internal components to operate asynchronously and independently. It uniquely integrates hardware-in-the-loop V2X communication using an external device for realistic data forwarding. Furthermore, RAMACoDrive incorporates an asynchronous online evaluation method with real-time performance metrics, enabling comprehensive evaluation beyond mere accuracy. Through experiments centered on cooperative 3D object detection, we demonstrate its effectiveness in revealing the impact of real-world asynchronous effects and communication constraints. RAMACoDrive provides a valuable and realistic testbed for cooperative perception algorithms and helps to advance the deployment of practical applications. The code will be released on our GitHub repository.
Shunyao Zhang 0001, Yingfeng Cai, Chengzhi Gao, Caizhen He, Hai Wang 0003, Long Chen 0003
IEEE Trans. Mob. Comput.5
2025 Uncertainty-Aware End-to-End Autonomous Driving Modeling
abstract
End-to-end autonomous driving has become a prominent research paradigm. However, these methods typically rely on deterministic static and dynamic representations, leading to compounded errors that accumulate over time across modules and become increasingly amplified, ultimately degrading planning performance. Meanwhile, the deterministic sequential arrangement of modules leads to information loss and error propagation, and makes it difficult to model the interactions between the ego and surrounding agents, thereby compromising the reliability of the planning outcomes. Moreover, these methods exhibit limited capacity in perceiving and prioritizing critical features, as they assume a deterministic and uniform importance across all feature dimensions. To overcome these limitations, we introduce an uncertainty-aware end-to-end driving model to enhance planning performance and overall safety. First, we jointly model both static and dynamic uncertainties and incorporate them into the model, enabling the model to perceive and respond to environment uncertainties. Static uncertainty guides the extraction of correct map structures and driving rules, while dynamic uncertainty enables the model to capture the stochastic interactions between traffic participants and environment and provides probabilistic scores for trajectory selection, helping to avoid collisions. Second, we propose a distributed query mechanism that leverages planning-specific queries to decode raw features, thereby removing the reliance of planning on scene understanding. In addition, we jointly model prediction and planning, enabling rich interactions between ego and surrounding agents. Third, we introduce a multi-scale dilated aggregation that adjusts the attention to informative features, highlighting critical feature contributions and enlarges the receptive field. We conduct extensive evaluations on nuScenes dataset and Bench2Drive benchmark. Experimental results demonstrate that our approach achieves the lowest planning collision rate and L2 error, and enhances inference speed and efficiency, supporting the deployment of high-performance, low-cost.
Yingfeng Cai, Xinhe Chen, Hai Wang 0003, Long Chen 0003
IEEE Internet Things J.3
2025 An Autonomous Driving Vehicle-Road Collaborative Heterogeneous Fusion Localization Method Based on the Epipolar Plane Model and Graph Optimization
abstract
High-precision localization is central to the safety of high-level autonomous driving. However, existing SLAM and map-based localization methods suffer from challenges, such as cumulative errors and the need for frequent map updates. To address these issues, this article proposes a vehicle–road collaborative localization fusion approach designed to achieve high-accuracy localization without cumulative errors. The proposed method employs visual-inertial odometry (VIO) on the vehicle side, where gyroscope bias is optimized through image observations and an epipolar plane model to ensure robust pose initialization. On the roadside, 3-D object detection is utilized to provide drift-free vehicle position and orientation information. By leveraging factor graph optimization and an edge-based marginalization mechanism, this approach efficiently integrates local constraints from vehicle-side VIO data with global constraints from roadside localization data, thereby reducing computational complexity. Experimental results demonstrate that incorporating roadside infrastructure significantly reduces localization errors, effectively mitigating long-term drift. Evaluations conducted in both road and indoor environments at Jiangsu University indicate that the proposed system improves localization accuracy by 30% compared to visual-inertial navigation system (VINS)-Mono. The findings validate the superiority of vehicle–road collaborative localization in both accuracy and stability, offering a promising solution for advancing intelligent driving technologies.
Yicheng Li 0001, Yunqi Xia, Yingfeng Cai, Hai Wang 0003, Long Chen 0003
IEEE Internet Things J.4
2025 Traffic Agents Trajectory Prediction Based on Enhanced Bidirectional Recurrent Network and Adaptive Social Interaction Model
abstract
Accurate prediction of the future trajectory of traffic agents is imperative to the effective motion planning of autonomous vehicles and mobile robots. Despite enormous progress that has been made toward trajectory prediction, dynamic and crowded traffic scenarios pose major challenges to the understanding and forecasting of traffic agents’ motion behavior. In this paper, we propose a novel trajectory prediction method from the perspective of temporal modeling and social interaction. Specifically, we first put forward a recurrent modeling approach to learn temporal features in favor of capturing long-range and short-range temporal dependencies of individual agents. Then, we construct a social feature learning module to capture the sparse and directional interactions among agents while suppressing the spurious connections. Finally, to reduce the accumulated error during prediction, a coordinated bidirectional decoding module is developed where temporal and social features can be properly integrated into the forward and backward prediction processes. Extensive experiments are performed on four real-world trajectory prediction benchmarks, and the results demonstrate the superiority of our method compared with other competing approaches. Detailed ablation studies are also performed to evaluate the effectiveness of each model component. Note to Practitioners—Motion planning is one of the crucial components of autonomous systems, such as intelligent vehicles and mobile robots. For example, the safety and efficiency of motion planning can be drastically improved if the future trajectories of surrounding agents, e.g., pedestrians, bicyclists, cars, etc., can be accurately forecasted. Motivated by the above requirements, this article develops an advanced deep learning model that can learn temporal and social features from trajectory data and perform accuracy prediction. This work aims to enhance the bidirectional recurrent network for dealing with trajectory data with evident temporal characteristics. In addition, this work introduces a novel adaptive social interaction modeling approach that overcomes the inherent defect of the fixed threshold method. This work also addresses the error accumulation problem in the prediction. The proposed model is evaluated on several datasets, and the results demonstrate its effectiveness. Our approach has broad application prospects in autonomous driving and mobile robots.
Xiaobo Chen 0001, Yuwen Liang, Chuan Hu 0003, Hai Wang 0003, Qiaolin Ye
IEEE Trans Autom. Sci. Eng.4
2025 MliG: Scene-Level Multimodal Motion Prediction Based on Multi-Layer Interaction Graph
abstract
The representation form of multimodal motion prediction results is critical to the efficiency of the prediction-planning workflow. However, most existing methods generate multiple sets of trajectories for each independent agent to cover diverse modalities, overlooking interactions among agents during the prediction period, making it challenging to achieve scene-level consistency in multi-agent predictions. This paper introduces MliG, a scene-level multimodal conditional motion prediction framework based on Multi-Layer Interaction Graph. The interaction graph designed explicitly represent the interaction behaviors of multi-agents over future periods, guiding conditional trajectory prediction to achieve planning-friendly prediction result representation. First, an Interaction-Mask-based determination strategy is introduced to obtain ground truth interaction labels, enabling data-driven implicit relationship learning, and the generated training labels can support the accurate construction of the interaction graph. Second, scene simplification and modality decoupling using the interaction graph make it easier to fuse information and represent driving scenarios efficiently. Based on the multi-layer design, the complex interactions between agent pairs are decomposed into different group and modality layers, ensuring the diversity of interactions while improving the efficiency of conditional predictions. Extensive experiments on the Argoverse 2, Argoverse 1, and INTERACTION Datasets demonstrate that the proposed MliG achieves high prediction accuracy and significantly improves the scene-level consistency of trajectory prediction modalities for multi-agents in the same driving scenario.
Yingfeng Cai, Ziheng Lu, Hai Wang 0003, Yubo Lian, Long Chen 0003, Qingchao Liu
IEEE Trans. Intell. Transp. Syst.3
2025 Distributed Modeling and Scenario-Driven Extension Hybrid-DMPC Coordinated Control of Autonomous Vehicle Chassis
abstract
As a critical technology to improve vehicle safety and handling stability, chassis coordinated control technology faces challenges in multi-system coordination, multi-variable solutions, and multi-task allocation. This paper combines the advantages of decentralized and distributed architectures and proposes a scenario-driven extension Hybrid-DMPC algorithm with variable topology and distributed modeling. Firstly, a two-dimensional extension coordinate is constructed with the$\beta - \dot {\beta } $phase plane and LTR as feature states. The correlation function is then solved to divide the vehicle states into classical domain, extension domain, and non-domain, forming a mapping relationship with the control architecture. Secondly, the distributed state space equations with state coupling and input coupling are constructed. The proposed algorithm is designed to employ decentralized, distributed, and hybrid architectures in the classical domain, non-domain, and extension domain, respectively. The correlation function determines the fusion rules for multiple control inputs in the extension domain. Thirdly, a cost-coupled weighted optimization function is designed, enabling local agents to coordinate global performance goals through iterative optimization. Finally, co-simulation and HIL test results verified that the proposed algorithm effectively coordinates multiple control objectives, achieving high-precision trajectory tracking, excellent handling stability, and anti-roll performance.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yubo Lian, Zhaozhi Dong
IEEE Trans. Intell. Transp. Syst.4
2025 Takeover Time Prediction for Conditionally Automated Driving Vehicles: Considering Mixed Traffic Flow Environment
abstract
The length of time for drivers to take over conditionally automated driving vehicles (CADV) is often influenced by numerous factors, especially the mixed traffic flow environment. To analyze the influencing factors of driver’s takeover time in the mixed traffic flow environment, we study a real-vehicle takeover experiment with the CADV and radar video integrated machine provided by Jiangsu University. DeepGBM algorithm and Shapley additive explanation (SHAP) are utilized to predict and analyze the takeover time based on the data obtained from the real-vehicle experiment. The results show that the DeepGBM algorithm performs a better accuracy in predicting the takeover time compared with the other algorithms. Moreover, the predicting effect of DeepGBM in the scene of average takeover time is better than the shorter and longer ones. On the other hand, when taking over CADV in the intersection area, the driver pays more attention to the vehicle’s speed. While in the non-intersection area, the driver is more attentive to the longitudinal distance difference with the vehicle ahead. This study further found that CADVs are subject to significant lateral interference for the takeover process in intersection areas, significantly impacting the driver’s takeover time and vehicle safety. This study can provide a theoretical reference for automated driving companies to design takeover times.
Qingchao Liu, Jingya Zhao, Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.5
2025 A Preference-Based Multi-Agent Federated Reinforcement Learning Algorithm Framework for Trustworthy Interactive Urban Autonomous Driving
abstract
In autonomous driving (AD) tasks, data-driven deep reinforcement learning (DRL) outperforms rule-based methods in terms of continuous decision-making and adaptability. However, traditional DRL relies on hand-crafted reward functions, which introduce objective alignment challenges and reward loopholes. Moreover, the black-box structure makes it difficult to explain the decision-making process, which has a direct impact on DRL performance in complex driving situations. To address these shortcomings, a preference-based decomposable proximal policy optimization algorithm (PDPPO) is proposed for reliable interactive urban AD. The framework deconstructs the federated reinforcement learning (FRL) algorithm from various perspectives using a rule-based preference model, resulting in high-availability algorithmic performance for AD. PDPPO employs a data-rule fusion-driven hybrid vision transformer to overcome the objective alignment and high-dimensional state-space representation challenges of traditional DRL in complex urban traffic environments. Furthermore, to address the issue of algorithmic trustworthiness, PDPPO models the multi-agent FRL co-optimization process as an interpretable self-organized group collaboration process. This approach enables the algorithm to strike a balance between model robustness and sample efficiency using preference-heuristic parameter aggregation. The simulation results demonstrate that the proposed PDPPO algorithmic framework can implement interpretable single-agent decision control and multi-agent co-optimization processes. Furthermore, it exhibits competitive performance on various benchmark tests.
Sikai Lu, Yingfeng Cai, Yubo Lian, Long Chen 0003, Hai Wang 0003
IEEE Trans. Intell. Transp. Syst.6
2025 Self-Distillation Attention for Efficient and Accurate Motion Prediction in Autonomous Driving
abstract
The accuracy and stability of motion prediction are crucial for the safe planning of autonomous driving systems. The widely used attention mechanisms effectively improve prediction accuracy. However, their computational cost grows quadratically with sequence length, presenting challenges for handling complex, large-scale scenarios. The attention patterns of motion prediction tasks exhibit significant data-related sparsity, indicating that not all scene elements are worth interaction. Therefore, this paper proposes a novel motion prediction framework that enhances the efficiency and stability of interaction fusion while achieving promising prediction accuracy. Firstly, an adaptive self-distillation attention module based on sparse interaction graph is designed. This module adaptively filters high-value sequences for each target to achieve sparse attention calculation and balance scene scale. Due to the lack of interaction labels in the original dataset, a self-distillation strategy is employed for model training. Secondly, a multi-stage dynamic anchor decoder is introduced that leverages the information filtering and aggregation capabilities of the sparse attention mechanism to improve prediction accuracy. The decoder dynamically updates the target anchor representations as the prediction progresses, ensuring consistency between interaction states and the prediction process in long-term forecasting. This approach effectively focuses attention calculation on the most relevant scene context fusion and trajectory decoding at each prediction stage. Validation results on Argoverse 1 and Argoverse 2 demonstrate that the proposed method achieves competitive accuracy while effectively reducing computational resource consumption. The proposed method can also serve as a plug-in that can be seamlessly incorporated into standard motion prediction pipelines to optimize scene interaction, making it more friendly to low-cost devices.
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Yubo Lian, Long Chen 0003
IEEE Trans. Intell. Transp. Syst.4
2025 A Multi-Agent Federated Reinforcement Learning-Based Cooperative Vehicle-Infrastructure Control Approach Framework for Roundabouts
abstract
In contrast to autonomous driving (AD) systems based on a single vehicle, the intelligent vehicle infrastructure cooperative system enables intelligent connected vehicles (ICVs) to perform active safety control and road collaboration management through vehicle-Infrastructure dynamic information interaction. However, complex intersections involve a wide range of interactions between agents, traffic participants and road infrastructures. The presence of time-varying traffic conditions and redundant traffic information may impede the efficiency and effectiveness of the vehicle-infrastructure system. Meanwhile, existing cooperative systems typically assume voluntary client access and data sharing, overlooking the information asymmetry challenges stemming from privacy awareness and the model interpretability challenges associated with data-driven approaches. This paper proposes a federated reinforcement learning (FRL) architecture for vehicle-infrastructure cooperative control is predicated on data and rule fusion, with the objective of considering privacy awareness and sample efficiency by employing a screening-based dual aggregation federated learning (FL) method. The proposed architecture enhances the interpretability of the algorithmic framework at four perspectives: modeling, training, optimization, and application. In this paper, we initially undertake training and performance comparisons on the OpenDD real traffic roundabout road driving dataset, and then extend the proposed method to more complex Nocrash benchmarks and real-world campus traffic environment for algorithm scalability validation.
Sikai Lu, Hai Wang 0003, Long Chen 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.2
2025 RTMDet-R: A Robust Instance Segmentation Network for Complex Traffic Scenarios
abstract
In complex traffic scenarios, several factors including lighting, weather, the size of the traffic participants, the distance between the traffic participants and the camera, and occlusions impact the features of the traffic participants. The impact of these factors is a huge challenge, especially for vision-based instance segmentation networks. To this end, this paper proposes an enhanced version of the RTMDet to promote the overall performance of instance segmentation in complex traffic scenarios. Firstly, an extended CSP-style backbone with large kernel convolutions of different kernel sizes is used to enhance the robustness of feature extraction capability, which contributes to obtaining more information about traffic objects of different scales. Secondly, a plugin pre-fusion module is designed to enhance the network’s robustness to multi-scale changes caused by distance changes. Additionally, instance kernel distinguish module is proposed to further highlight and distinguish different instance objects under poor lighting or weather and occlusion situations. Finally, the existing advanced image generation technology is used to expand the BDD100k dataset, enriching the dataset with severe scenarios. With an input resolution of$\mathbf {1280}\times \mathbf {720}$on the expanded BDD100K dataset, the proposed RTMDet-R achieves an accuracy of 25.4% mAP on the instance mask and 27.8% mAP on the instance box. This surpasses other similar models in terms of accuracy. Additionally, it maintains a good inference speed of 23.1 FPS, achieving the trade-off between accuracy and speed. Code and models are released athttps://github.com/GTrui6/RTMDet-R.git.
Hai Wang 0003, Qirui Qin, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.1
2025 MixedFusion: An Efficient Multimodal Data Fusion Framework for 3-D Object Detection and Tracking
abstract
The performance of environmental perception is critical for the safe driving of intelligent connected vehicles (ICVs). Currently, the most prevalent technical solutions are based on multimodal data fusion to achieve a comprehensive perception of the surrounding environment. However, existing fusion perception methods suffer from issues such as low sensor data utilization and unreasonable fusion strategies, which severely limit their performance in adverse weather conditions. To address these issues, this article proposes a novel multimodal data fusion framework called MixedFusion. In this framework, we introduce two innovative fusion strategies for the data characteristics of each sensor: high-level semantic guidance (HLSG) and multipriority matching (MPM). It not only realizes the efficient utilization of the multimodal data but also further realizes the complementary fusion between the multimodal data. We perform extensive experiments on the nuScenes and K-radar datasets. The experimental results demonstrate that the fusion framework proposed in this article significantly improves the performance of 3-D object detection and tracking in severe weather conditions.
Cheng Zhang 0031, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Neural Networks Learn. Syst.2
2024 DAIR-V2XReid: A New Real-World Vehicle-Infrastructure Cooperative Re-ID Dataset and Cross-Shot Feature Aggregation Network Perception Method
abstract
As an emerging research field, vehicle re-identification (Re-ID) can realize identity search between the vehicles, which plays an important role in the over-the-horizon perception of Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD). At present, due to the lack of data sets, the relevant research on Vehicle-Infrastructure Cooperative (VIC) Re-ID can only be evaluated in the cross-view monitoring test set which leads to the lack of persuasion of the research. Therefore, based on the DAID-V2X dataset of Tsinghua University, this paper constructs a VIC Re-ID dataset “DAIR-V2XReid” from real vehicle scenarios through vehicle-road end target tag association, thereby making it better applicable to the research of VIC Re-ID. Owing to different task scenarios, existing algorithms trained on monitoring test sets are unable to effectively complete the Re-ID task in this new dataset. Therefore, Cross-shot Feature Aggregation Network (CFA-Net) is also proposed in this paper, to tackle the case where a vehicle becomes unrecognizable due to a large change in its visual appearance across different cameras. Firstly, we put forward a camera embedding module and add it to the Backbone, to group different cameras and solve the problem of cross-shot perspective mutation. Secondly, in order to address the situation where background and vehicle division are not distinguishable, we propose a cross-stage feature fusion module, which integrates low-order semantics with high-order semantics. Finally, we use multi-directional attention network to achieve the final feature extraction. The experimental results show that our proposed CFA-Net method achieves new state-of-the-art in DAIR-V2XReid, with mAP of 58.47%.
Hai Wang 0003, Yaqing Niu, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.1
2024 RTMDet-MGG: A Multi-Task Model With Global Guidance
abstract
In this paper, we propose the concept of global guidance, design a global guidance structure based on a dual-scale global feature enhancement module, and construct a multi-task network (RTMDet-MGG) for road scene instance segmentation and drivable area segmentation in order to overcome the limitation of current multi-task networks in sharing features to different task branches. The proposed global guidance structure enhances the connection between different task branches in the multi-task network by sharing the enhanced global feature with different task branches through various fusion methods, thereby enhancing the multi-task network’s overall performance. With an input resolution of$\mathbf {640}\times \mathbf {360}$on the BDD100K dataset, RTMDet-MGG achieves an accuracy of 18.7% mAP (mean average precision) on the instance mask and 82.9% mIoU (mean intersection over union) on the drivable area, with an inference speed of 32.1 FPS (frames per second), which satisfies the requirements of real-time tasks. In addition, the algorithm has excellent scene generalization capabilities, and the mIoU of drivable area segmentation on our custom-built dataset for unstructured road drivable area segmentation reaches 93.3%.
Hai Wang 0003, Qirui Qin, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.1
2024 A Multi-Task Learning Network With a Collision-Aware Graph Transformer for Traffic-Agents Trajectory Prediction
abstract
It is critical for autonomous vehicles to accurately forecast the future trajectories of surrounding agents to avoid collisions. However, capturing the complex interactions between agents in complex urban scenes is challenging. As a result, complex interactions may impair trajectory prediction accuracy. A trajectory prediction network with an enhanced Graph Transformer (TP-EGT) is proposed to forecast the future trajectories of traffic-agents. A collision-aware Graph Transformer is introduced to capture the complex social interactions between traffic-agents. Following that, an additional interaction prediction task that could predict the interaction probabilities between agents is proposed to mitigate the over-smoothing issue of the Graph Transformer via a multi-task learning strategy. Afterward, the trajectory prediction performance is improved with additional interaction probabilities, which are beneficial for the decision-making and planning modules of autonomous vehicles. Quantitative and qualitative evaluations of TP-EGT on the ETH/UCY and ApolloScape databases demonstrate that the trajectory prediction accuracy of TP-EGT is comparable to the state-of-the-art baseline methods, and the predicted interaction probabilities can help autonomous vehicles comprehend the complex traffic scenes.
Fucheng Fan, Hai Wang 0003, Ammar Jafaripournimchahi, Hongyu Hu
IEEE Trans. Intell. Transp. Syst.4
2024 Real-Time Pedestrian Crossing Anticipation Based on an Action-Interaction Dual-Branch Network
abstract
Accurate anticipation of pedestrian crossing intentions is critical for preventing pedestrian-vehicle conflicts and ensuring road safety. This issue is a significant focus in intelligent transportation systems and autonomous driving. However, current approaches often face challenges, such as high computational costs due to complex scene understanding and inadequate consideration of spatiotemporal dependencies in pedestrian actions. To handle these challenges, we propose RAIDN (real-time action-interaction dual-branch network), which comprises the pedestrian action encoding and traffic-object interaction modules to anticipate pedestrian crossing intentions in real time. The pedestrian action encoding module employs a multi-scale graph transformer, efficiently extracting the intrinsic topology of long- and short-term action variations. This module effectively addresses issues of information redundancy in multi-channel graph convolution networks and local limitations in multi-scale temporal convolutions. Subsequently, the traffic-object interaction module introduces an interaction relation graph convolution network to excavate relevant traffic-object interactions, thereby shortening the prolonged scene semantic inference. Finally, global average pooling and attention layers fuse the action and interaction cues for real-time intention anticipation. The effectiveness of RAIDN has been validated on public datasets JAAD and PIE, achieving competitive metrics with Accuracy, ROC-AUC, F1-Score, Precision, and Recall rates of 0.89/0.92, 0.80/0.89, 0.66/0.85, 0.65/0.82, and 0.72/0.89 respectively. Notably, RAIDN demonstrates a remarkable inference time of just 0.28ms, outperforming other state-of-the-art methods and establishing its suitability for real-time applications in intelligent transportation and autonomous driving. Code is available athttps://github.com/wrysmile99/RAIG.
Zhiwen Wei, Chuan Hu 0003, Yingfeng Cai, Hai Wang 0003, Hongyu Hu
IEEE Trans. Intell. Transp. Syst.5
2024 Small Fixed-Wing Unmanned Aerial Vehicle Path Following Under Low Altitude Wind Shear Disturbance
abstract
The wind shear in low altitude can bring about time-varying and unknown disturbances, which makes small fixed-wing Unmanned Aerial Vehicle (UAV) hard to accurately follow desired curved paths. In this paper, a Vector Field (VF) based curved path following algorithm is designed for UAV to overcome the above difficulties. Firstly, the path following problem is mathematically formulated to give the kinematics model of UAV with constant airspeed. Secondly, the curved path following control laws under both constant and time-varying unknown wind disturbances are designed with VF. A ground velocity estimator is additionally designed to overcome the difficulties of unmeasurable ground velocity, and the Lyapunov stability of curved path following is analyzed. Finally, both simulations and field tests on a real small fixed-wing UAV are carried out to evaluate the algorithm performance in the presence of time-varying and unknown wind disturbances. Verification results demonstrate that the algorithm proposed in this paper outperforms other widely used path following algorithms and can effectively follow arbitrary curved paths.
Zhouyu Zhang 0002, Chenyuan He, Hongshen Chen, Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Haitong Li, Tongwei Lu
IEEE Trans. Intell. Transp. Syst.5
2024 TransFusion: Multi-Modal Robust Fusion for 3D Object Detection in Foggy Weather Based on Spatial Vision Transformer
abstract
A practical approach to realizing the comprehensive perception of the surrounding environment is to use a multi-modal fusion method based on various types of vehicular sensors. In clear weather, the camera and LiDAR can provide high-resolution images and point clouds that can be utilized for 3D object detection. However, in foggy weather, the propagation of light is affected by the fog in the air. Consequently, both images and point clouds become distorted to varying degrees. Thus, it is challenging to implement accurate detection in adverse weather conditions. Compared to cameras and LiDAR, Radar possesses strong penetrating power and is not affected by fog. Therefore, this paper proposes a novel two-stage detection framework called “TransFusion”, which leverages LiDAR and Radar fusion to solve the problem of environment perception in foggy weather. The proposed framework is composed of Multi-modal Rotate Region Proposal Network (MM-RRPN) and Multi-modal Refine Network (MM-RFN). Specifically, Spatial Vision Transformer (SVT) and Cross-Modal Attention Mechanism (CMAM) are introduced in the MM-RRPN to improve the robustness of the algorithm in foggy weather. Furthermore, Temporal-Spatial Memory Fusion (TSMF) module in MM-RFN is employed to fuse the spatial-temporal prior information. In addition, the Multi-branches Combination Loss function (MC-Loss) is designed to efficiently supervise the learning of the network. Extensive experiments were conducted on Oxford Radar RobotCar (ORR) dataset. The experimental results show that the proposed algorithm has excellent performance in both foggy and clear weather. Especially in foggy weather, the proposed TransFusion achieves 85.31mAP, outperforming all other competing approaches. The demo is available at:https://youtu.be/ugjIYHLgn98.
Cheng Zhang 0031, Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.2
2023 Driving Style Feature Extraction and Recognition Based on Hyperdimensional Computing and Semi-Supervised Twin Projection Vector Machine
abstract
Driving style recognition is one of the most crucial requirements for human-centric autonomous and assistive driving systems. Existing studies either require a large amount of labeled data or suffer from hand-crafted features, thus hindering the practical application of these methods. To address the above challenges, this paper proposes a driving style recognition approach from the perspective of brain-inspired hyperdimensional computing and semi-supervised learning. Specifically, raw sensor signals are treated as multivariate time series which are encoded by hyperdimensional computing, thus generating a large holistic feature representation. Then, considering the high dimension of the resulting representation, we put forward a semi-supervised twin projection vector machine (SSTPVM) model that can take full advantage of unlabeled data and jointly optimize multiple projection directions tailored for dimension reduction. In addition, a heuristic method based on particle swarm optimization is developed for the parameter selection of SSTPVM. Finally, extensive experiment comparisons with other related methods are performed on naturalistic driving style data. The results show that hyperdimensional computing is rather suitable for semi-supervised learning while our proposed SSTPVM can significantly improve recognition performance even with a small portion of labeled data.
Xiaobo Chen 0001, Yuxiang Gao, Haoze Yu, Hai Wang 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.4
2023 Monocular Road Scene Bird's Eye View Prediction via Big Kernel-Size Encoder and Spatial-Channel Transform Module
abstract
A detailed representation of the surrounding road scene is crucial for an autonomous driving system. yellow The camera-based Bird’s Eye View map has been a popular solution to present the surrounding information, due to its low cost and rich spatial context information. Most of the existing methods predict the BEV map based on the depth-estimation or the trivial homography method, which may cause the error propagation and the absence of content. To overcome these drawbacks, we propose a novel end-to-end framework that employs the front monocular image to predict the road layout and vehicle occupancy. In particular, to capture the long-range feature, we redesign a CNN encoder with a large kernel size to extract the image features. For reducing the big difference between the front image features and the top-down features, we propose a novel Spatial-Channel projection module to convert the front map into the top-down space. Additionally, concerning the correlation between front view and top-down view, we propose the Dual Cross-view Transformer module to refine the top-down view feature maps and strengthen the transformation. Extensive evaluations on the KITTI and Argoverse datasets present that the proposed model achieves the state-of-the-art results for both datasets. Furthermore, the proposed model runs in 37 FPS on a single GPU, demonstrating the generation of a real-time BEV map. The code will be published at https://github.com/raozhongyu/BEV_LKA.
Zhongyu Rao, Hai Wang 0003, Long Chen 0003, Yubo Lian, Yilin Zhong, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.2
2023 Sparse U-PDP: A Unified Multi-Task Framework for Panoptic Driving Perception
abstract
This paper proposes a multi-task unified decoding framework for the panoptic driving perception, Sparse U-PDP. This framework’s primary objective is to combine vehicle object detection, lane detection and drivable area segmentation tasks. This paper mainly builds a multi-task unified decoder and further explores whether the potential connection between multi-tasks can improve the robustness of the model. Experiments demonstrate that the proposed Sparse U-PDP outperforms the present state-of-the-art multi-task model in terms of accuracy. The primary contributions of this work are as follows: First, we present the approach to unified multi-task representation. We abstract the multi-task into the “dynamic convolution kernels” representation form to build a highly unified multi-task decoder. Second, we use the proposed dynamic interaction module to establish different feature sampling pipelines for various task features. Lastly, our model is verified on the BDD100K dataset, where we achieve an AP50 of 84.1 in the “vehicle” category, 32.0 IoU in the lane detection, and 93.0 mIoU in the drivable area segmentation with the helper of CSP-Darknet. That is to say; Sparse U-PDP verifies that a more unified task representation form can implicitly increase the mutual help between different task branches.
Hai Wang 0003, Meng Qiu, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.1
2023 CenterPoint-SE: A Single-Stage Anchor-Free 3-D Object Detection Algorithm With Spatial Awareness Enhancement
abstract
Real-time and accurate 3-D object detection is one of the foundational technologies for environmental perception in autonomous vehicles. However, the existing second-stage anchor-based 3-D object detection algorithms have high accuracy, but they are challenging in terms of computation complexity and latency. Due to poor perception of spatial features, the accuracy of the existing single-stage anchor-free detection algorithms with low latency are difficult to be implemented into autonomous vehicles. Therefore, we focus on enhancing the spatial perception ability of the anchor-free detection network based on CenterPoints. In this paper, we propose a single-stage anchor-free 3-D object detector CenterPoint-Space-Enhancement (CenterPoint-SE) algorithm and construct an efficient 3-D backbone network to extract fine-grained spatial geometric features by introducing a spatial attention mechanism and residual structure. At the same time, a powerful spatial semantic feature fusion module, the enhancement of feature fusion (EF-Fusion), is designed. In addition, we add a lightweight IoU prediction branch to improve the algorithm’s perception of various object sizes. Finally, we add a foreground point segmentation auxiliary training branch to enable the 3-D backbone to obtain object boundary features. We use the ONCE dataset to train and validate the proposed model, and the results showed that the proposed CenterPoint-SE achieves 70.33 mAP and an inference speed of 17.15 FPS, outperforming other methods.
Hai Wang 0003, Le Tao, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Map-based localization for intelligent vehicles from bi-sensor data fusion
Yicheng Li 0001, Yingfeng Cai, Zhixiong Li 0001, Shizhe Feng, Hai Wang 0003, Miguel Ángel Sotelo
Expert Syst. Appl.5
2022 Pedestrian Motion Trajectory Prediction in Intelligent Driving from Far Shot First-Person Perspective Video
abstract
Pedestrian motion trajectory prediction is an important task in intelligent driving, and it can provide a valuable reference for the subsequent path decision of intelligent driving. However, so far, there are only a few models in the field of specific pedestrian motion track prediction in intelligent driving from far shot first-person perspective video. To accomplish this task, we proposed a deep learning model for pedestrian motion trajectory prediction from far shot first-person perspective video with four key innovations: a) A macroscopic pedestrian trajectory prediction module is established under the close correlation between neighboring frames to estimate the pedestrian motion track on the whole; b) A relative motion transformation module of vehicle-mounted camera is designed to consider the effect of vehicle-mounted camera’s ego-motion on the pedestrian motion track; c) We set up a circular training module to maintain the number of parameters in our model to simplify and reduce the size of model; d) A new far shot first-person pedestrian motion dataset under intelligent driving is specifically established to train and test the proposed model. The above four modules are integrated into the proposed deep learning model, which achieves state-of-the-art results for predicting pedestrian motion trajectory from both far and close shot first-person perspective video.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Robust Target Recognition and Tracking of Self-Driving Cars With Radar and Camera Information Fusion Under Severe Weather Conditions
abstract
Radar and camera information fusion sensing methods are used to solve the inherent shortcomings of the single sensor in severe weather. Our fusion scheme uses radar as the main hardware and camera as the auxiliary hardware framework. At the same time, the Mahalanobis distance is used to match the observed values of the target sequence. Data fusion based on the joint probability function method. Moreover, the algorithm was tested using actual sensor data collected from a vehicle, performing real-time environment perception. The test results show that radar and camera fusion algorithms perform better than single sensor environmental perception in severe weather, which can effectively reduce the missed detection rate of autonomous vehicle environment perception in severe weather. The fusion algorithm improves the robustness of the environment perception system and provides accurate environment perception information for the decision-making system and control system of autonomous vehicles.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Hongbo Gao 0001, Yunyi Jia, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.3
2022 SFNet-N: An Improved SFNet Algorithm for Semantic Segmentation of Low-Light Autonomous Driving Road Scenes
abstract
In recent years, considerable progress has been made in semantic segmentation of images with favorable environments. However, the environmental perception of autonomous driving under adverse weather conditions is still very challenging. In particular, the low visibility at nighttime greatly affects driving safety. In this paper, we aim to explore image segmentation in low-light scenarios, thereby expanding the application range of autonomous vehicles. The segmentation algorithms for road scenes based on deep learning are highly dependent on the volume of images with pixel-level annotations. Considering the scarcity of labeled large-scale nighttime data, we performed synthetic data collection and data style transfer using images acquired in daytime based on the autonomous driving simulation platform and generative adversarial network, respectively. In addition, we also proposed a novel nighttime segmentation framework (SFNET-N) to effectively recognize objects in dark environments, aiming at the boundary blurring caused by low semantic contrast in low-illumination images. Specifically, the framework comprises a light enhancement network which introduces semantic information for the first time and a segmentation network with strong feature extraction capability. Extensive experiments with Dark Zurich-test and Nighttime Driving-test datasets show the effectiveness of our method compared with existing state-of-the art approaches, with 56.9% and 57.4% mIoU (mean of category-wise intersection-over-union) respectively. Finally, we also performed real-vehicle verification of the proposed models in road scenes of Zhenjiang city with poor lighting. The datasets are available athttps://github.com/pupu-chenyanyan/semantic-segmentation-on-nightime.
Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.1
2022 DLnet With Training Task Conversion Stream for Precise Semantic Segmentation in Actual Traffic Scene
abstract
Many successful semantic segmentation models trained on certain datasets experience a performance gap when they are applied to the actual scene images, expressing weak robustness of these models in the actual scene. The training task conversion (TTC) and domain adaption field have been originally proposed to solve the performance gap problem. Unfortunately, many existing models for TTC and domain adaptation have defects, and even if the TTC is completed, the performance is far from the original task model. Thus, how to maintain excellent performance while completing TTC is the main challenge. In order to address this challenge, a deep learning model named DLnet is proposed for TTC from the existing image dataset-based training task to the actual scene image-based training task. The proposed network, named the DLnet, contains three main innovations. The proposed network is verified by experiments. The experimental results show that the proposed DLnet not only can achieve state-of-the-art quantitative performance on four popular datasets but also can obtain outstanding qualitative performance in four actual urban scenes, which demonstrates the robustness and performance of the proposed DLnet. In addition, although the proposed DLnet cannot achieve outstanding performance in real time, it can still achieve a moderate performance in real time, which is within an acceptable range.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Creating navigation map in semi-open scenarios for intelligent vehicle localization using multi-sensor fusion
Yicheng Li 0001, Yingfeng Cai, Reza Malekian, Hai Wang 0003, Miguel Ángel Sotelo, Zhixiong Li 0001
Expert Syst. Appl.4
2021 Multi-Target Pan-Class Intrinsic Relevance Driven Model for Improving Semantic Segmentation in Autonomous Driving
abstract
At present, most semantic segmentation models rely on the excellent feature extraction capabilities of a deep learning network structure. Although these models can achieve excellent performance on multiple datasets, ways of refining the target main body segmentation and overcoming the performance limitation of deep learning networks are still a research focus. We discovered a pan-class intrinsic relevance phenomenon among targets that can link the targets cross-class. This cross-class strategy is different from the latest semantic segmentation model via context where targets are divided into an intra-class and inter-class. This paper proposes a model for refining the target main body segmentation using multi-target pan-class intrinsic relevance. The main contributions of the proposed model can be summarized as follows: a) The multi-target pan-class intrinsic relevance prior knowledge establishment (RPK-Est) module builds the prior knowledge of the intrinsic relevance to lay the foundation for the following extraction of the pan-class intrinsic relevance feature. b) The multi-target pan-class intrinsic relevance feature extraction (RF-Ext) module is designed to extract the pan-class intrinsic relevance feature based on the proposed multi-target node graph and graph convolution network. c) The multi-target pan-class intrinsic relevance feature integration (RF-Int) module is proposed to integrate the intrinsic relevance features and semantic features by a generative adversarial learning strategy at the gradient level, which can make intrinsic relevance features play a role in semantic segmentation. The proposed model achieved outstanding performance in semantic segmentation testing on four authoritative datasets compared to other state-of-the-art models.
Yingfeng Cai, Hai Wang 0003, Zhixiong Li 0001
IEEE Trans. Image Process.3
2020 Robust crowd counting based on refined density map
Jinmeng Cao, Nan Wang 0013, Hai Wang 0003, Yingfeng Cai
Multim. Tools Appl.4
2020 A Novel Saliency Detection Algorithm Based on Adversarial Learning Model
abstract
The traditional salient object detection models can be divided into several classes based on the low-level features of images and contrast between the pixels. This paper proposes an adversarial learning model (ALM) that includes the generative model and discriminative model. The ALM uses the original image as an input of the generative model to extract the high-level features and forms an initial salient map. Then, the discriminative model is utilized to compare differences in the features between the initial salient map and the ground truth, and the obtained differences are sent to the convolutional layers of the generative model to adjust the parameters for the generative model updating. Due to the serial-iterative adjustment, the salient map of the generative model becomes more similar to the ground truth. Lastly, the ALM forms the salient map fused with the super-pixels by enhancing the color and texture features, so the final salient map is obtained. The ALM is not limited to the color and texture features; on the contrary, it fuses multiple features and achieves good results in the salient target extraction. The experimental results show that ALM performs better than the other ten state-of-the-art models on three different datasets. Thus, the proposed ALM is widely applicable to the salient target extraction.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Image Process.3
2018 Salient object detection based on multi-scale contrast
Hai Wang 0003, Yingfeng Cai, Long Chen 0003
Neural Networks1
2016 Multilevel framework to handle object occlusions for real-time tracking
abstract
This study proposes an efficient method to handle the object occlusions seen in monocular traffic image sequences. The motivation of this study is different methods perform differently in occlusion segmentation and the authors’ idea is to use a situation‐driven approach to aggregate different methods in order to get a good performance. This study classifies occlusion into four categories according to the foreground situation and a multilevel occlusion handling framework is utilised. First, the image segmentation algorithm based on convex hull analysis is utilised for intra‐frame level occlusion segmentation. The segmentation algorithm is established by the compactness ratio and interior distance ratio of the foreground. Second, an online sample‐based classification algorithm is utilised for tracking level occlusion segmentation. Training samples are extracted from the historical frames before occlusion and testing samples are extracted from the current frame by an adaptive searching strategy. The segmentation of occlusion is transferred into the online classification of testing samples. Such algorithm is established by the similarity and coherence of target's property between continuous frames. Experiments on video sequences illustrate the good performance of the proposed method under different conditions with low computational cost.
Yingfeng Cai, Hai Wang 0003, Xiaobo Chen 0001, Long Chen 0003
IET Image Process.2
2016 Occluded vehicle detection with local connected deep model
Hai Wang 0003, Yingfeng Cai, Xiaobo Chen 0001, Long Chen 0003
Multim. Tools Appl.1