Bolin Gao

dblp:120/1968 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ROGAD: Risk-Aware Occupancy Guided End-to-End Autonomous Driving
Hengduo Zou, Kunsong Shi, Zhongke Wu, Bolin Gao
IV7
2026 VRU Trajectory Prediction Based on SA-TF-LSTM Considering Traffic-Actor Interaction for Automated Vehicles
abstract
This paper systematically investigates the prediction of Vulnerable Road Users’ (VRU, including pedestrian, cyclist, and electric cyclist) trajectories by leveraging the Action Intention Model (AIM), Revised Social Force Model (RSFM) and Social Attention-Transformer-Long Short Term Memory Network (SA-TF-LSTM) for automated vehicles. Firstly, an AIM based on the Transformer is developed for predicting the VRUs’ crossing/waiting intention. VRU type and heterogeneity (age and gender), distance between vehicle and VRU, speed of vehicle and VRU are considered. Secondly, a micro-dynamic RSFM is used to model the trajectories of VRUs for generating initially hypothetical future trajectories, which are then merged the historical trajectories with observed time as a new feature input. Furthermore, traffic data gathered by an unmanned aerial vehicle (UAV) is acquired and examined, and the Maximum Likelihood Estimation (MLE) is utilized to adjust the parameters of the RSFM. Finally, a data driven SA-TF-LSTM is proposed for VRU trajectory prediction, and VRU crossing intention, traffic-actor interaction, VRU type and heterogeneity are considered. Social Attention is employed to ascertain the attention coefficients of the aforementioned factors. The results demonstrate that the data driven SA-TF-LSTM surpasses the existing methods, with a prediction accuracy enhancement of over 9% utilizing the collected traffic data. This significant improvement grants us substantial confidence in employing the SA-TF-LSTM within automated vehicles to bolster the safety of VRUs.
Tianshu Pang, Bolin Gao, Hao Chen 0074, Xi Zhang 0016
IEEE Internet Things J.2
2026 Cloud-Vehicle Collaborative Stability Control via Learning-Based Variable-Matrix Model Predictive Control With Region-of-Attraction Constraints
Shengru Chen, Lin Zhang 0035, Bolin Gao, Yingen Ge, Lu Xiong 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Vision-Driven 2D Supervised Fine-Tuning Framework for Bird's Eye View Perception
abstract
Visual bird’s eye view (BEV) perception, dute to its excellent perceptual capabilities, is progressively replacing costly LiDAR-based perception systems, especially in the realm of urban intelligent driving. However, this type of perception still relies on LiDAR data to construct ground truth databases, a process that is both cumbersome and time-consuming. Additionally, most mass-produced autonomous driving systems are equipped solely with surround camera sensors and lack the LiDAR data necessary for precise annotation. To tackle this challenge, we propose a fine-tuning method for BEV perception network based on visual 2D semantic perception, aimed at enhancing the model’s generalization capabilities in new scene data. Leveraging the maturity of 2D perception technologies, our method utilizes only 2D semantic segmentation labels and monocular depth estimations, thereby significantly reducing the dependence on expensive BEV ground truths and offering strong potential for industrial deployment. Extensive experiments and comparative analyses on the nuScenes and Waymo datasets demonstrate the effectiveness of our method. Specifically, it improves mAP and NDS by 2.51% and 1.93% on nuScenes, and by 1.21% and 0.78% on Waymo, respectively, validating its practical utility and robustness across diverse domains.
Qiaoyi Wang, Honglin Sun, Qing Xu 0010, Bolin Gao, Shengbo Eben Li, Jianqiang Wang 0003, Keqiang Li 0002
IROS5
2025 Minds on the Move: Decoding Trajectory Prediction in Autonomous Driving With Cognitive Insights
abstract
In mixed autonomous driving environments, accurately predicting the future trajectories of surrounding vehicles is crucial for the safe operation of autonomous vehicles (AVs). In driving scenarios, a vehicle’s trajectory is determined by the decision-making process of human drivers. However, existing models primarily focus on the inherent statistical patterns in the data, often neglecting the critical aspect of understanding the decision-making processes of human drivers. This oversight results in models that fail to capture the true intentions of human drivers, leading to suboptimal performance in long-term trajectory prediction. To address this limitation, we introduce a Cognitive-Informed Transformer (CITF) that incorporates a cognitive concept, Perceived Safety, to interpret drivers’ decision-making mechanisms. Perceived Safety encapsulates the varying risk tolerances across drivers with different driving behaviors. Specifically, we develop a Perceived Safety-aware Module that includes a Quantitative Safety Assessment for measuring the subject risk levels within scenarios, and Driver Behavior Profiling for characterizing driver behaviors. Furthermore, we present a novel module, Leanformer, designed to capture social interactions among vehicles. CITF demonstrates significant performance improvements on three well-established datasets. In terms of long-term prediction, it surpasses existing benchmarks by 12.0% on the NGSIM, 28.2% on the HighD, and 20.8% on the MoCAD dataset. Additionally, its robustness in scenarios with limited or missing data is evident, surpassing most state-of-the-art (SOTA) baselines, and paving the way for real-world applications.
Haicheng Liao, Chengyue Wang 0001, Kaiqun Zhu, Yilong Ren, Bolin Gao, Shengbo Eben Li, Cheng-Zhong Xu 0001, Zhenning Li 0001
IEEE Trans. Intell. Transp. Syst.5
2025 HSTI: A Light Hierarchical Spatial-Temporal Interaction Model for Map-Free Trajectory Prediction
abstract
Trajectory prediction is a crucial task for autonomous driving, but current models’ reliance on high-definition (HD) maps limits their broader applicability. To cope with this challenge, we propose a novel map-free trajectory prediction method that leverages spatiotemporal attention mechanisms. The method consists of three key stages: 1) we first encode spatial and temporal features separately using spatial and temporal attention mechanisms, 2) we then model spatial and temporal interactions through Crystal Graph Convolutional Networks (CGCN) and Multi-Head Attention (MHA), 3) finally, an adaptive anchor generation technique is introduced to tackle the multimodal trajectory prediction challenge. This self-adaptive technique generates context-specific anchors, enabling accurate prediction of multiple possible future vehicle trajectories. Extensive experiments on the Argoverse1 and V2X-Seq datasets validate the effectiveness of our approach. On the Argoverse1 dataset, our method outperforms CRAT-Pred by 5.8% in minADE and 6.25% in minFDE. On the V2X-Seq dataset, it achieves improvements of 82.6%, 85.1%, and 44.0% in minADE, minFDE, and MR, respectively, compared to the baseline model.
Xiaoyang Luo, Shuaiqi Fu, Bolin Gao, Yanan Zhao 0004, Huachun Tan, Zeye Song
IEEE Trans. Intell. Transp. Syst.3
2025 Anti-Rollover Path Tracking Control for an Autonomous Semi-Trailer Tank Truck
abstract
This study focuses on the anti-rollover control problem for autonomous semi-trailer tank trucks and proposes an anti-rollover path tracking control algorithm suitable for autonomous driving scenarios. A simplified semi-trailer tank truck model is established in the controller, modeling liquid as a single pendulum considering both lateral and roll inputs, and an anti-rollover path tracking algorithm that utilizes it is developed based on multi-constraint model predictive control (MPC). The equivalent lateral load transfer rate (LTR), liquid sloshing angle, and angular velocity are used as constraints to achieve multi-objective optimization for path tracking, sloshing suppression, and rollover prevention. A vehicle-fluid coupling co-simulation platform based on computational fluid dynamics (CFD) is built to verify the control performance under extreme scenarios of left turning, high-speed single-lane change (SLC), and short-distance double-lane change (DLC). Additionally, the similarity principle for experimental validation using a down-scale model tank truck is derived, and experiments are conducted on the model semi-trailer tank truck under the DLC scenario. Through simulation and experimental verification, the proposed anti-rollover path tracking algorithm demonstrates acceptable tracking performance while ensuring that no rollover is carried out under extreme conditions, limiting the angle of$|LTR|<0.75$and sloshing within -20∘-20∘, reducing the angle of sloshing by a maximum of 42% in the experiment. Moreover, the real-time capability of the proposed algorithm meets the requirements for practical applications with a peak time consumption of one control step less than 20msin the controller tested.
Xiaojing Qi, Yugong Luo, Bolin Gao
IEEE Trans. Intell. Transp. Syst.3
2025 Multi-Option Hierarchical Reinforcement Learning Framework With State Segmentation for Mixed On-Ramp Merging
abstract
Deep Reinforcement Learning (DRL) has achieved significant advancements in the transportation domain, effectively enhancing traffic network efficiency, reducing pollutant emissions, and improving driving safety. A prominent approach within DRL, Hierarchical Reinforcement Learning (HRL), simplifies complex tasks by grouping states and decomposing the Markov Decision Process (MDP), facilitating the exploration in multi-dimensional state spaces. These concepts of state abstraction and temporal abstraction prove to be particularly beneficial in complex, high-risk transportation scenarios, such as on-ramp merging. In this context, this paper introduces a novel Multi-Option HRL (MO-HRL) framework with state segmentation. Unlike traditional option-based HRL, the proposed framework enables the simultaneous activation of multiple options, with each option observing diverse states. After carefully defining and justifying the framework, we apply MO-HRL to a simplified on-ramp merging scenario. To enhance training, curriculum learning is incorporated into the MO-HRL framework. Extensive experiments involve discussions of different training modes, the “shared critic” problem, and comparisons with state-of-the-art baselines. Additionally, a six-lane mainline on-ramp merging scenario, based on the NGSIM I-80 dataset, is constructed. Simulation results from both scenarios show that the proposed approach outperforms existing methods and maintains a balance between the mainline and on-ramp traffic.
Zoutao Wen, Huachun Tan, Yanan Zhao 0004, Hailong Zhang 0018, Peifeng Li 0003, Xinguo Chen, Bolin Gao
IEEE Trans. Intell. Transp. Syst.7
2025 AGSENet: A Robust Road Ponding Detection Method for Proactive Traffic Safety
abstract
Road ponding, a prevalent traffic hazard, poses a serious threat to road safety by causing vehicles to lose control and leading to accidents ranging from minor fender benders to severe collisions. Existing technologies struggle to accurately identify road ponding due to complex road textures and variable ponding coloration influenced by reflection characteristics. To address this challenge, we propose a novel approach called Self-Attention-based Global Saliency-Enhanced Network (AGSENet) for proactive road ponding detection and traffic safety improvement. AGSENet incorporates saliency detection techniques through the Channel Saliency Information Focus (CSIF) and Spatial Saliency Information Enhancement (SSIE) modules. The CSIF module, integrated into the encoder, employs self-attention to highlight similar features by fusing spatial and channel information. The SSIE module, embedded in the decoder, refines edge features and reduces noise by leveraging correlations across different feature levels. To ensure accurate and reliable evaluation, we corrected significant mislabeling and missing annotations in the Puddle-1000 dataset. Additionally, we constructed the Foggy-Puddle and Night-Puddle datasets for road ponding detection in low-light and foggy conditions, respectively. Experimental results demonstrate that AGSENet outperforms existing methods, achieving IoU improvements of 2.03%, 0.62%, and 1.06% on the Puddle-1000, Foggy-Puddle, and Night-Puddle datasets, respectively, setting a new state-of-the-art in this field. Finally, we verified the algorithm’s reliability on edge computing devices. This work provides a valuable reference for proactive warning research in road traffic safety. The source code and datasets are placed in thehttps://github.com/Lyu-Dakang/AGSENet.
Shangyu Yang, Dakang Lyu, Junzhou Chen 0001, Yilong Ren, Bolin Gao, Zhihan Lyu
IEEE Trans. Intell. Transp. Syst.7
2024 Vehicle-Road-Cloud Collaborative Perception Framework and Key Technologies: A Review
abstract
Over recent years, the Vehicle-Road-Cloud Integration System (VRCIS) and Intelligent and Connected Vehicles (ICVs) have gained significant attention in the realm of autonomous driving. By sharing data across diverse traffic participants and coordinating with VRCIS, ICVs can achieve enhanced perception accuracy and superior driving decisions, surpassing autonomous vehicles that rely solely on onboard sensors. Existing literature explores VRCIS’ overall architecture, applications, and deployment status. However, there is a lack of a comprehensive review focusing on the overarching architecture of ICV’s perception and its associated technologies, which are fundamental to VRCIS from an information integration perspective. This gap hinders the development of a robust perception framework for VRCIS, including the crucial perception technologies specific to it. This survey seeks to bridge this gap by offering an exhaustive review of the designed VRCIS perception framework and its specific perception technologies. Firstly, an overview of VRCIS’ perception architecture is provided, and the application relationships among various perception technologies are elucidated. Then, single-node, multi-node, and vehicle-road-cloud collaborative perception technologies are explored in sequence. Finally, the survey concludes with a discussion of insights and prospective future directions for VRCIS.
Bolin Gao, Hengduo Zou, Keqiang Li 0002
IEEE Trans. Intell. Transp. Syst.1
2020 Taking a Look at Small-Scale Pedestrians and Occluded Pedestrians
abstract
Small-scale pedestrian detection and occluded pedestrian detection are two challenging tasks. However, most state-of-the-art methods merely handle one single task each time, thus giving rise to relatively poor performance when the two tasks, in practice, are required simultaneously. In this paper, it is found that small-scale pedestrian detection and occluded pedestrian detection actually have a common problem, i.e., an inaccurate location problem. Therefore, solving this problem enables to improve the performance of both tasks. To this end, we pay more attention to the predicted bounding box with worse location precision and extract more contextual information around objects, where two modules (i.e., location bootstrap and semantic transition) are proposed. The location bootstrap is used to reweight regression loss, where the loss of the predicted bounding box far from the corresponding ground-truth is upweighted and the loss of the predicted bounding box near the corresponding ground-truth is downweighted. Additionally, the semantic transition adds more contextual information and relieves semantic inconsistency of the skip-layer fusion. Since the location bootstrap is not used at the test stage and the semantic transition is lightweight, the proposed method does not add many extra computational costs during inference. Experiments on the challenging CityPersons and Caltech datasets show that the proposed method outperforms the state-of-the-art methods on the small-scale pedestrians and occluded pedestrians (e.g., 5.20% and 4.73% improvements on the Caltech).
Jiale Cao, Yanwei Pang, Jungong Han, Bolin Gao, Xuelong Li 0001
IEEE Trans. Image Process.4
2019 Complementary Features with Reasonable Receptive Field for Road Scene 3D Object Detection
abstract
Accurate and efficient 3D object detection is of great importance for autonomous driving and robot perception. There are two problems in lidar based 3D object detection networks. (1) Semantic information (e.g., class label of each point) and spatial details are not fully explored for feature extraction. (2) The variance of object sizes represented by point cloud is much smaller than those represented by 2D images. But existing methods do not make use of this property and the receptive field sizes generally mismatch the physical sizes of road scene objects. Based on these two aspects, we propose Complementary Features with Reasonable receptive field networks (CFRNet). CFRNet first exploits Complementary Feature Extractor to learn semantic and positional features, then utilizes a RPN (Region Proposal Networks) with reasonable receptive field to collect correlated context in road scene. Experimental results on KITTI benchmark show the effectiveness of our method. Moreover, our method achieves state-of-the-art performance at a high inference speed.
Yuefeng Wu, Yanwei Pang, Bolin Gao, Jungong Han
ICIP3
2013 Computer-aided 3D human modeling for portrait-based product development using point- and curve-based deformation
Li Li 0081, Fei Liu 0012, Chao-hua Peng, Bolin Gao
Comput. Aided Des.4