Yicheng Li 0001

dblp:59/8179-1 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 V2I-GridCalib: A Real-Time Online Calibration for Vehicle-to-Infrastructure Systems
abstract
Real-time spatial alignment between vehicles and roadside sensors is essential for unlocking collaborative perception. However, existing methods often struggle in dynamic environments due to observational scale mismatch and high latency. This paper proposes a novel edge-assisted framework for robust, real-time, and target-free online V2I calibration. First, a Directional Handshake mechanism pre-selects task-relevant point cloud subsets to reduce computational and transmission overhead. Second, a distance-adaptive rendering strategy generates scale-optimized rasterized views to explicitly address scale variation based on the vehicle's distance. Finally, a particle filter estimates the 6-DoF pose by fusing complementary 2D-3D reprojection and 2D-2D epipolar geometric constraints. Extensive experiments in real-world intersections and high-fidelity simulations validate the system. The framework maintains trajectory consistency errors below 5.5% for paths up to 400 meters and achieves centimeter-level accuracy with 5.2 cm translation and 0.3° rotation errors in static tests. Crucially, the system consistently meets stringent real-time constraints of less than 66 ms for V2X applications.
Zhuoyi Jiang, Yingfeng Cai, Yicheng Li 0001, Qingchao Liu, Long Chen 0003, Hai Wang 0003
IEEE Internet Things J.3
2025 An Autonomous Driving Vehicle-Road Collaborative Heterogeneous Fusion Localization Method Based on the Epipolar Plane Model and Graph Optimization
abstract
High-precision localization is central to the safety of high-level autonomous driving. However, existing SLAM and map-based localization methods suffer from challenges, such as cumulative errors and the need for frequent map updates. To address these issues, this article proposes a vehicle–road collaborative localization fusion approach designed to achieve high-accuracy localization without cumulative errors. The proposed method employs visual-inertial odometry (VIO) on the vehicle side, where gyroscope bias is optimized through image observations and an epipolar plane model to ensure robust pose initialization. On the roadside, 3-D object detection is utilized to provide drift-free vehicle position and orientation information. By leveraging factor graph optimization and an edge-based marginalization mechanism, this approach efficiently integrates local constraints from vehicle-side VIO data with global constraints from roadside localization data, thereby reducing computational complexity. Experimental results demonstrate that incorporating roadside infrastructure significantly reduces localization errors, effectively mitigating long-term drift. Evaluations conducted in both road and indoor environments at Jiangsu University indicate that the proposed system improves localization accuracy by 30% compared to visual-inertial navigation system (VINS)-Mono. The findings validate the superiority of vehicle–road collaborative localization in both accuracy and stability, offering a promising solution for advancing intelligent driving technologies.
Yicheng Li 0001, Yunqi Xia, Yingfeng Cai, Hai Wang 0003, Long Chen 0003
IEEE Internet Things J.1
2025 RTMDet-R: A Robust Instance Segmentation Network for Complex Traffic Scenarios
abstract
In complex traffic scenarios, several factors including lighting, weather, the size of the traffic participants, the distance between the traffic participants and the camera, and occlusions impact the features of the traffic participants. The impact of these factors is a huge challenge, especially for vision-based instance segmentation networks. To this end, this paper proposes an enhanced version of the RTMDet to promote the overall performance of instance segmentation in complex traffic scenarios. Firstly, an extended CSP-style backbone with large kernel convolutions of different kernel sizes is used to enhance the robustness of feature extraction capability, which contributes to obtaining more information about traffic objects of different scales. Secondly, a plugin pre-fusion module is designed to enhance the network’s robustness to multi-scale changes caused by distance changes. Additionally, instance kernel distinguish module is proposed to further highlight and distinguish different instance objects under poor lighting or weather and occlusion situations. Finally, the existing advanced image generation technology is used to expand the BDD100k dataset, enriching the dataset with severe scenarios. With an input resolution of$\mathbf {1280}\times \mathbf {720}$on the expanded BDD100K dataset, the proposed RTMDet-R achieves an accuracy of 25.4% mAP on the instance mask and 27.8% mAP on the instance box. This surpasses other similar models in terms of accuracy. Additionally, it maintains a good inference speed of 23.1 FPS, achieving the trade-off between accuracy and speed. Code and models are released athttps://github.com/GTrui6/RTMDet-R.git.
Hai Wang 0003, Qirui Qin, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.4
2025 MixedFusion: An Efficient Multimodal Data Fusion Framework for 3-D Object Detection and Tracking
abstract
The performance of environmental perception is critical for the safe driving of intelligent connected vehicles (ICVs). Currently, the most prevalent technical solutions are based on multimodal data fusion to achieve a comprehensive perception of the surrounding environment. However, existing fusion perception methods suffer from issues such as low sensor data utilization and unreasonable fusion strategies, which severely limit their performance in adverse weather conditions. To address these issues, this article proposes a novel multimodal data fusion framework called MixedFusion. In this framework, we introduce two innovative fusion strategies for the data characteristics of each sensor: high-level semantic guidance (HLSG) and multipriority matching (MPM). It not only realizes the efficient utilization of the multimodal data but also further realizes the complementary fusion between the multimodal data. We perform extensive experiments on the nuScenes and K-radar datasets. The experimental results demonstrate that the fusion framework proposed in this article significantly improves the performance of 3-D object detection and tracking in severe weather conditions.
Cheng Zhang 0031, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Neural Networks Learn. Syst.4
2024 Localization for Intelligent Vehicles in Underground Car Parks Based on Semantic Information
abstract
Global navigation satellite system (GNSS) signals cannot be received indoors, thus to deploy intelligent vehicles in underground car parks other localization methods are needed. In this paper, we use various carpark signs that are widely and uniformly distributed in underground parking lots as localization references. We propose a coarse-to-fine multiscale localization method that relies solely on vision sensors for underground parking lot localization based on a preconstructed lightweight node map. In coarse localization, we propose a semantic keyframe topological localization method to predict the localization range (candidate set of nodes). In node-level localization, we extract features by learning-based neural networks and construct a hybrid k-nearest neighbor (H-KNN) model to search for the closest node within the coarse localization results. In metric localization, we construct plane homography and perspective-n-point (PnP) models, allowing the vehicle’s pose (rotation and translation relative to the closest node) to be computed for refined localization. The proposed method has been tested in two underground parking lots of an office building and a shopping mall with different characteristics. Experimental results demonstrate that root mean square error (RMSE) is 0.38 m and the proposed method exhibits strong robustness in various scenarios.
Yicheng Li 0001, Yingfeng Cai, Zhixiong Li 0001, Miguel Ángel Sotelo
IEEE Trans. Intell. Transp. Syst.1
2024 DAIR-V2XReid: A New Real-World Vehicle-Infrastructure Cooperative Re-ID Dataset and Cross-Shot Feature Aggregation Network Perception Method
abstract
As an emerging research field, vehicle re-identification (Re-ID) can realize identity search between the vehicles, which plays an important role in the over-the-horizon perception of Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD). At present, due to the lack of data sets, the relevant research on Vehicle-Infrastructure Cooperative (VIC) Re-ID can only be evaluated in the cross-view monitoring test set which leads to the lack of persuasion of the research. Therefore, based on the DAID-V2X dataset of Tsinghua University, this paper constructs a VIC Re-ID dataset “DAIR-V2XReid” from real vehicle scenarios through vehicle-road end target tag association, thereby making it better applicable to the research of VIC Re-ID. Owing to different task scenarios, existing algorithms trained on monitoring test sets are unable to effectively complete the Re-ID task in this new dataset. Therefore, Cross-shot Feature Aggregation Network (CFA-Net) is also proposed in this paper, to tackle the case where a vehicle becomes unrecognizable due to a large change in its visual appearance across different cameras. Firstly, we put forward a camera embedding module and add it to the Backbone, to group different cameras and solve the problem of cross-shot perspective mutation. Secondly, in order to address the situation where background and vehicle division are not distinguishable, we propose a cross-stage feature fusion module, which integrates low-order semantics with high-order semantics. Finally, we use multi-directional attention network to achieve the final feature extraction. The experimental results show that our proposed CFA-Net method achieves new state-of-the-art in DAIR-V2XReid, with mAP of 58.47%.
Hai Wang 0003, Yaqing Niu, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.4
2024 RTMDet-MGG: A Multi-Task Model With Global Guidance
abstract
In this paper, we propose the concept of global guidance, design a global guidance structure based on a dual-scale global feature enhancement module, and construct a multi-task network (RTMDet-MGG) for road scene instance segmentation and drivable area segmentation in order to overcome the limitation of current multi-task networks in sharing features to different task branches. The proposed global guidance structure enhances the connection between different task branches in the multi-task network by sharing the enhanced global feature with different task branches through various fusion methods, thereby enhancing the multi-task network’s overall performance. With an input resolution of$\mathbf {640}\times \mathbf {360}$on the BDD100K dataset, RTMDet-MGG achieves an accuracy of 18.7% mAP (mean average precision) on the instance mask and 82.9% mIoU (mean intersection over union) on the drivable area, with an inference speed of 32.1 FPS (frames per second), which satisfies the requirements of real-time tasks. In addition, the algorithm has excellent scene generalization capabilities, and the mIoU of drivable area segmentation on our custom-built dataset for unstructured road drivable area segmentation reaches 93.3%.
Hai Wang 0003, Qirui Qin, Long Chen 0003, Yicheng Li 0001, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.4
2024 TransFusion: Multi-Modal Robust Fusion for 3D Object Detection in Foggy Weather Based on Spatial Vision Transformer
abstract
A practical approach to realizing the comprehensive perception of the surrounding environment is to use a multi-modal fusion method based on various types of vehicular sensors. In clear weather, the camera and LiDAR can provide high-resolution images and point clouds that can be utilized for 3D object detection. However, in foggy weather, the propagation of light is affected by the fog in the air. Consequently, both images and point clouds become distorted to varying degrees. Thus, it is challenging to implement accurate detection in adverse weather conditions. Compared to cameras and LiDAR, Radar possesses strong penetrating power and is not affected by fog. Therefore, this paper proposes a novel two-stage detection framework called “TransFusion”, which leverages LiDAR and Radar fusion to solve the problem of environment perception in foggy weather. The proposed framework is composed of Multi-modal Rotate Region Proposal Network (MM-RRPN) and Multi-modal Refine Network (MM-RFN). Specifically, Spatial Vision Transformer (SVT) and Cross-Modal Attention Mechanism (CMAM) are introduced in the MM-RRPN to improve the robustness of the algorithm in foggy weather. Furthermore, Temporal-Spatial Memory Fusion (TSMF) module in MM-RFN is employed to fuse the spatial-temporal prior information. In addition, the Multi-branches Combination Loss function (MC-Loss) is designed to efficiently supervise the learning of the network. Extensive experiments were conducted on Oxford Radar RobotCar (ORR) dataset. The experimental results show that the proposed algorithm has excellent performance in both foggy and clear weather. Especially in foggy weather, the proposed TransFusion achieves 85.31mAP, outperforming all other competing approaches. The demo is available at:https://youtu.be/ugjIYHLgn98.
Cheng Zhang 0031, Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.5
2023 Sparse U-PDP: A Unified Multi-Task Framework for Panoptic Driving Perception
abstract
This paper proposes a multi-task unified decoding framework for the panoptic driving perception, Sparse U-PDP. This framework’s primary objective is to combine vehicle object detection, lane detection and drivable area segmentation tasks. This paper mainly builds a multi-task unified decoder and further explores whether the potential connection between multi-tasks can improve the robustness of the model. Experiments demonstrate that the proposed Sparse U-PDP outperforms the present state-of-the-art multi-task model in terms of accuracy. The primary contributions of this work are as follows: First, we present the approach to unified multi-task representation. We abstract the multi-task into the “dynamic convolution kernels” representation form to build a highly unified multi-task decoder. Second, we use the proposed dynamic interaction module to establish different feature sampling pipelines for various task features. Lastly, our model is verified on the BDD100K dataset, where we achieve an AP50 of 84.1 in the “vehicle” category, 32.0 IoU in the lane detection, and 93.0 mIoU in the drivable area segmentation with the helper of CSP-Darknet. That is to say; Sparse U-PDP verifies that a more unified task representation form can implicitly increase the mutual help between different task branches.
Hai Wang 0003, Meng Qiu, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.5
2023 CenterPoint-SE: A Single-Stage Anchor-Free 3-D Object Detection Algorithm With Spatial Awareness Enhancement
abstract
Real-time and accurate 3-D object detection is one of the foundational technologies for environmental perception in autonomous vehicles. However, the existing second-stage anchor-based 3-D object detection algorithms have high accuracy, but they are challenging in terms of computation complexity and latency. Due to poor perception of spatial features, the accuracy of the existing single-stage anchor-free detection algorithms with low latency are difficult to be implemented into autonomous vehicles. Therefore, we focus on enhancing the spatial perception ability of the anchor-free detection network based on CenterPoints. In this paper, we propose a single-stage anchor-free 3-D object detector CenterPoint-Space-Enhancement (CenterPoint-SE) algorithm and construct an efficient 3-D backbone network to extract fine-grained spatial geometric features by introducing a spatial attention mechanism and residual structure. At the same time, a powerful spatial semantic feature fusion module, the enhancement of feature fusion (EF-Fusion), is designed. In addition, we add a lightweight IoU prediction branch to improve the algorithm’s perception of various object sizes. Finally, we add a foreground point segmentation auxiliary training branch to enable the 3-D backbone to obtain object boundary features. We use the ONCE dataset to train and validate the proposed model, and the results showed that the proposed CenterPoint-SE achieves 70.33 mAP and an inference speed of 17.15 FPS, outperforming other methods.
Hai Wang 0003, Le Tao, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.5
2022 Map-based localization for intelligent vehicles from bi-sensor data fusion
Yicheng Li 0001, Yingfeng Cai, Zhixiong Li 0001, Shizhe Feng, Hai Wang 0003, Miguel Ángel Sotelo
Expert Syst. Appl.1
2022 Pedestrian Motion Trajectory Prediction in Intelligent Driving from Far Shot First-Person Perspective Video
abstract
Pedestrian motion trajectory prediction is an important task in intelligent driving, and it can provide a valuable reference for the subsequent path decision of intelligent driving. However, so far, there are only a few models in the field of specific pedestrian motion track prediction in intelligent driving from far shot first-person perspective video. To accomplish this task, we proposed a deep learning model for pedestrian motion trajectory prediction from far shot first-person perspective video with four key innovations: a) A macroscopic pedestrian trajectory prediction module is established under the close correlation between neighboring frames to estimate the pedestrian motion track on the whole; b) A relative motion transformation module of vehicle-mounted camera is designed to consider the effect of vehicle-mounted camera’s ego-motion on the pedestrian motion track; c) We set up a circular training module to maintain the number of parameters in our model to simplify and reduce the size of model; d) A new far shot first-person pedestrian motion dataset under intelligent driving is specifically established to train and test the proposed model. The above four modules are integrated into the proposed deep learning model, which achieves state-of-the-art results for predicting pedestrian motion trajectory from both far and close shot first-person perspective video.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.5
2022 Robust Target Recognition and Tracking of Self-Driving Cars With Radar and Camera Information Fusion Under Severe Weather Conditions
abstract
Radar and camera information fusion sensing methods are used to solve the inherent shortcomings of the single sensor in severe weather. Our fusion scheme uses radar as the main hardware and camera as the auxiliary hardware framework. At the same time, the Mahalanobis distance is used to match the observed values of the target sequence. Data fusion based on the joint probability function method. Moreover, the algorithm was tested using actual sensor data collected from a vehicle, performing real-time environment perception. The test results show that radar and camera fusion algorithms perform better than single sensor environmental perception in severe weather, which can effectively reduce the missed detection rate of autonomous vehicle environment perception in severe weather. The fusion algorithm improves the robustness of the environment perception system and provides accurate environment perception information for the decision-making system and control system of autonomous vehicles.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Hongbo Gao 0001, Yunyi Jia, Yicheng Li 0001
IEEE Trans. Intell. Transp. Syst.7
2022 SFNet-N: An Improved SFNet Algorithm for Semantic Segmentation of Low-Light Autonomous Driving Road Scenes
abstract
In recent years, considerable progress has been made in semantic segmentation of images with favorable environments. However, the environmental perception of autonomous driving under adverse weather conditions is still very challenging. In particular, the low visibility at nighttime greatly affects driving safety. In this paper, we aim to explore image segmentation in low-light scenarios, thereby expanding the application range of autonomous vehicles. The segmentation algorithms for road scenes based on deep learning are highly dependent on the volume of images with pixel-level annotations. Considering the scarcity of labeled large-scale nighttime data, we performed synthetic data collection and data style transfer using images acquired in daytime based on the autonomous driving simulation platform and generative adversarial network, respectively. In addition, we also proposed a novel nighttime segmentation framework (SFNET-N) to effectively recognize objects in dark environments, aiming at the boundary blurring caused by low semantic contrast in low-illumination images. Specifically, the framework comprises a light enhancement network which introduces semantic information for the first time and a segmentation network with strong feature extraction capability. Extensive experiments with Dark Zurich-test and Nighttime Driving-test datasets show the effectiveness of our method compared with existing state-of-the art approaches, with 56.9% and 57.4% mIoU (mean of category-wise intersection-over-union) respectively. Finally, we also performed real-vehicle verification of the proposed models in road scenes of Zhenjiang city with poor lighting. The datasets are available athttps://github.com/pupu-chenyanyan/semantic-segmentation-on-nightime.
Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Yicheng Li 0001, Miguel Ángel Sotelo, Zhixiong Li 0001
IEEE Trans. Intell. Transp. Syst.5
2022 DLnet With Training Task Conversion Stream for Precise Semantic Segmentation in Actual Traffic Scene
abstract
Many successful semantic segmentation models trained on certain datasets experience a performance gap when they are applied to the actual scene images, expressing weak robustness of these models in the actual scene. The training task conversion (TTC) and domain adaption field have been originally proposed to solve the performance gap problem. Unfortunately, many existing models for TTC and domain adaptation have defects, and even if the TTC is completed, the performance is far from the original task model. Thus, how to maintain excellent performance while completing TTC is the main challenge. In order to address this challenge, a deep learning model named DLnet is proposed for TTC from the existing image dataset-based training task to the actual scene image-based training task. The proposed network, named the DLnet, contains three main innovations. The proposed network is verified by experiments. The experimental results show that the proposed DLnet not only can achieve state-of-the-art quantitative performance on four popular datasets but also can obtain outstanding qualitative performance in four actual urban scenes, which demonstrates the robustness and performance of the proposed DLnet. In addition, although the proposed DLnet cannot achieve outstanding performance in real time, it can still achieve a moderate performance in real time, which is within an acceptable range.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2021 Creating navigation map in semi-open scenarios for intelligent vehicle localization using multi-sensor fusion
Yicheng Li 0001, Yingfeng Cai, Reza Malekian, Hai Wang 0003, Miguel Ángel Sotelo, Zhixiong Li 0001
Expert Syst. Appl.1
2021 Visual Map-Based Localization for Intelligent Vehicles From Multi-View Site Matching
abstract
Accurate localization is a crucial step for intelligent vehicles (IVs). And vision-based localization methods are promising due to its good accuracy and low cost. However, vision-based methods are usually not robust enough due to the errors of matching similar road scenarios. In this paper, we proposed a visual map-based localization method, called multi-view site matching (MVSM). We proposed using two camera views (i.e., downward-view and front-view) to construct visual map. The visual map consists of a serial of nodes. Each node encodes the features of the road, the 2D structure, and the poses of the vehicle. Based on the constructed visual map, we proposed a multi-scale method for accurate vehicle localization. In coarse localization, we adopt a topological model to obtain a set of candidate nodes from visual map. Furthermore, holistic features from front view are matched within the candidates such that the best matched node is determined for image-level localization. In metric localization, the best matched is first verified with the local features from downward view. And the vehicle pose is finally computed by utilizing the 2D structure from the verified nodes in the map. In the experiment, the proposed MVSM method has been tested with actual field data covering different pavement types in different seasons. The proposed MVSM method can achieve less than 0.20m mean localization errors. Compared to existing vision-based methods, the proposed method utilizes two views to enhance image-level localization and 2D pavement structure to improve metric localization so as to greatly improve the overall localization performance.
Yicheng Li 0001, Zhaozheng Hu, Yingfeng Cai, Huawei Wu, Zhixiong Li 0001, Miguel Ángel Sotelo
IEEE Trans. Intell. Transp. Syst.1
2020 A Novel Saliency Detection Algorithm Based on Adversarial Learning Model
abstract
The traditional salient object detection models can be divided into several classes based on the low-level features of images and contrast between the pixels. This paper proposes an adversarial learning model (ALM) that includes the generative model and discriminative model. The ALM uses the original image as an input of the generative model to extract the high-level features and forms an initial salient map. Then, the discriminative model is utilized to compare differences in the features between the initial salient map and the ground truth, and the obtained differences are sent to the convolutional layers of the generative model to adjust the parameters for the generative model updating. Due to the serial-iterative adjustment, the salient map of the generative model becomes more similar to the ground truth. Lastly, the ALM forms the salient map fused with the super-pixels by enhancing the color and texture features, so the final salient map is obtained. The ALM is not limited to the color and texture features; on the contrary, it fuses multiple features and achieves good results in the salient target extraction. The experimental results show that ALM performs better than the other ten state-of-the-art models on three different datasets. Thus, the proposed ALM is widely applicable to the salient target extraction.
Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yicheng Li 0001
IEEE Trans. Image Process.5