Xu Li 0004

dblp:25/3528-4 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-2772-7114ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VICooper: Communication-Efficient Vehicle-Infrastructure Cooperative 3-D Object Detection Leveraging Roadside HD Point Cloud Background Map Priors
abstract
Recently, LiDAR-based Vehicle-Infrastructure Cooperative (VIC) perception has shown an advantage in expanding the horizon of Connected Autonomous Vehicles (CAVs), enabling occlusion-aware 3D scene understanding. However, limited communication bandwidth hampers multi-agent cooperation in urban Internet of Things (IoT) environments. Existing solutions often compress or implicitly filter high-resolution features, resulting in semantic redundancy or information loss, which degrades overall performance. To tackle this, we propose VICooper, a communication-efficient VIC perception framework. Driven by the stable context of static infrastructure LiDARs, VICooper employs offline Background Mapping (BgM) to extract foreground points of interest, thereby offering explicit guidance for communication reduction. For the sparse yet critical foreground point clouds, we introduce the Multi-dimensional Foreground Backbone (MdFB) that incorporates geometric cues, including height, scale, and spatial density, to enrich feature encoding. To address the fusion imbalance between dense vehicle-side and sparse roadside features, we customize a Progressive Bilateral Feature Aggregation (PbFA) using a deformable transformer to capture inter-agent mutual correlations, thereby enabling the deep coupling of heterogeneous agents under asymmetric information. Extensive evaluations on real-world VIC benchmark validate that VICooper achieves superior performance with substantially lower bandwidth, demonstrating its potential in intelligent transportation IoT ecosystems.
Benwu Wang, Xu Li 0004, Qimin Xu, Wenkai Zhu, Yinan Du, Haoyang Che, Baidan Li
IEEE Internet Things J.2
2026 VI_MCPR: Viewpoint Invariant Place Recognition Driven by Multicamera for Large-Scale Environments
abstract
Place recognition (PR) is a critical component of simultaneous localization and mapping in the fields of autonomous driving and robotics. In outdoor large-scale and complex environments, existing vision-based place recognition (VPR) methods typically rely on single-camera input, which is inherently limited by its restricted field of view and, thus, vulnerable to viewpoint variations. To effectively fill the aforementioned drawbacks, we propose VI_MCPR, a novel method that supports input from any number of cameras. This method utilizes a multibranch, weight-sharing encoder structure to encode image features from multiperspective simultaneously. The robust feature attention pooling block is then utilized to learn high-order nonlinear features and latent correlations between features, effectively mitigating the loss of key features during down-sampling. To generate a discriminative global descriptor representing the image, we designed a geometry and spatial relationship enhanced block, named graph-SE-transform (GSET), which captures the overall shape of objects in a manner similar to the human visual system. Extensive comparative experiments on the NuScenes, Argoverse 2 Sensors, and real-vehicle datasets demonstrate that VI_MCPR outperforms state-of-the-art VPR methods. Compared to the strongest representative baselines, our approach increases PR performance by approximately 6% under viewpoint variations, by approximately 6% in dynamic environments, and by approximately 8% in extreme scenarios such as adverse weather, illumination changes, and low-texture conditions.
Xu Li 0004, Qimin Xu, Dong Kong
IEEE Trans. Ind. Informatics2
2025 Radial awareness with adaptive hybrid CNN-Transformer range-view representation for outdoor LiDAR point cloud semantic segmentation
Xu Li 0004, Qimin Xu, Yue Hu 0009, Zhengliang Sun
Expert Syst. Appl.2
2025 Scenario potentiality-constrain network for RGB-D salient object detection
Guanyu Zong, Xu Li 0004, Qimin Xu
Knowl. Based Syst.2
2025 3D multi-object tracking based on parallel multimodal data association
Shiyu Tan, Xu Li 0004, Qimin Xu, Jianxiao Zhu
Mach. Vis. Appl.2
2025 Accurate Representation Modeling and Interindividual Constraint Learning for Roadside Three-Dimensional Object Detection
abstract
Roadside three-dimensional (3-D) object detection is essential for enhancing blind-less perception performance in cooperative autonomous driving. Current research has preliminarily explored the representations from different perspectives and constraints inside individuals. However, the accuracy of existing representations is hardly guaranteed under the wide-range conditions on roadside applications, and the constraints involving multiple targets are less considered. To address these, an accurate representation modeling and interindividual constraint-learning method is proposed. In representation modeling, the deficiencies of existing distance representations are systematically analyzed, and the relative depth is developed with the consideration of numerical distribution and error tendency under wide-range conditions. Besides, the limitation of existing reference-pixel representation is addressed by introducing Affine–Gaussian heatmaps, which accurately selects reference pixels and depresses extra responses based on geometric discrepancies. In constraint learning, the interindividual constraints are effectively extracted in the proposed graph-based module, which introduces powerful graph attention operations and a special designed implicit gradient flow to induce constraints into image feature maps. Extensive experiments on DAIR-V2X-I and Rope3-D demonstrate that significant improvements are achieved compared with concurrent state-of-arts on both the trained scenarios and unseen scenarios.
Jianxiao Zhu, Xu Li 0004, Qimin Xu, Benwu Wang
IEEE Trans. Ind. Informatics2
2024 Radial Transformer for Large-Scale Outdoor LiDAR Point Cloud Semantic Segmentation
abstract
Semantic segmentation of large-scale outdoor point cloud captured by light detection and ranging (LiDAR) sensors can provide fine-grain and stereoscopic comprehension for the surrounding environment. However, limited by the receptive field of convolution kernel and ignoration of specific spatial properties inherent to the large-scale outdoor point cloud, the existing advanced LiDAR semantic segmentation methods inevitably abandon the unique radial long-range topological relationships. To this end, from the LiDAR perspective, we propose a novel Radial Transformer that can naturally and efficiently exploit the radial long-range dependencies exclusive to the outdoor point cloud for accurate LiDAR semantic segmentation. Specifically, we first develop a radial window partition to generate a series of candidate point sequences and then construct the long-range interactions among the densely continuous point sequences by the self-attention mechanism. Moreover, considering the varying-distance distribution of point cloud in 3-D space, a spatial-adaptive position encoding is particularly designed to elaborate the relative position. Furthermore, we fusion radial balanced attention for a better structure representation of real-world scenes and distant points. Extensive experiments demonstrate the effectiveness and superiority of our method, which achieves 67.5% and 77.7% mean intersection-over-union (mIoU) on two recognized large-scale outdoor LiDAR point cloud datasets SemanticKITTI and nuScenes, respectively.
Xu Li 0004, Peizhou Ni, Qimin Xu, Xixiang Liu
IEEE Trans. Geosci. Remote. Sens.2
2024 An Enhanced-LiDAR/UWB/INS Integrated Positioning Methodology for Unmanned Ground Vehicle in Sparse Environments
abstract
Light detection and ranging (LiDAR) positioning has received great attention especially when satellites fail. However, the positioning accuracy is still subjected to the following. First, the LiDAR positioning accuracy is affected by the sparsity due to LiDAR beams and environment. Second, the existing methods for suppressing cumulative errors are based on ego-vehicle sensing that requires the motion of repeated paths. Moreover, the output frequency of LiDAR is low. To solve the above problems, an enhanced-LiDAR/ultrawideband (UWB)/inertial navigation system integrated positioning methodology is proposed. First, the enhanced LiDAR odometry module is designed to improve the resolution of LiDAR beams for more accurate odometry. Then, the cooperative optimization module is proposed to introduce UWB observation to suppress the accumulated error without relying on ego-vehicle sensing. Finally, the factor graph fusion module is used to fuse multisensor information dynamically and improve the output frequency. Experimental results prove the effectiveness of our methodology.
Yue Hu 0009, Xu Li 0004, Dong Kong, Peizhou Ni, Weiming Hu 0002, Xiang Song 0004
IEEE Trans. Ind. Informatics2
2024 SC_LPR: Semantically Consistent LiDAR Place Recognition Based on Chained Cascade Network in Long-Term Dynamic Environments
abstract
In large-scale long-term dynamic environments, high-frequency dynamic objects inevitably lead to significant changes in the appearance of the scene at the same location at different times, which is catastrophic for place recognition (PR). Therefore, how to eliminate the influence of dynamic objects to achieve robust PR has universal practical value for mobile robots and autonomous vehicles. To this end, we suggest a novel semantically consistent LiDAR PR method based on chained cascade network, called SC_LPR, which mainly consists of a LiDAR semantic image inpainting network (LSI-Net) and a semantic pyramid Transformer-based PR network (SPT-Net). Specifically, LSI-Net is a coarse-to-fine generative adversarial network (GAN) with a gated convolutional autoencoder as the backbone. To effectively address the challenges posed by variable-scale dynamic object masks, we integrate the updated Transformer block with mask attention and gated trident block into LSI-Net. Sequentially, in order to generate a discriminative global descriptor representing the point cloud, we design an encoder with pyramid Transformer block to efficiently encode long-range dependencies and global contexts between different categories in the inpainted semantic image, followed by an augmented NetVALD, a generalized VLAD (Vector of Locally Aggregated Descriptors) layer that adaptively aggregates salient local features. Last but not least, we first attempt to create a LiDAR semantic inpainting dataset, called LSI-Dataset, to effectively validate the proposed method. Experimental comparisons show that our method not only improves semantic inpainting performance by about 6%, but also improves PR performance in dynamic environments by about 8% compared to the representative optimal baseline. LSI-Dataset will be publicly available at https://github.KD.LPR.com/.
Dong Kong, Xu Li 0004, Qimin Xu, Yue Hu 0009, Peizhou Ni
IEEE Trans. Image Process.2
2024 A Cooperative Control Methodology Considering Dynamic Interaction for Multiple Connected and Automated Vehicles in the Merging Zone
abstract
The dynamic interaction among Connected and Automated Vehicles (CAVs) is becoming increasingly complex, encompassing factors such as dynamic topology and the dynamic states of multiple CAVs. Existing cooperative control methods struggle to explicitly represent dynamic interaction, which can lead to dangerous behavior, severe congestion, and even accidents. In this paper, we propose a cooperative control methodology that aims to improve safety and efficiency in the merging zone by deeply representing dynamic interaction among multiple CAVs. Our proposed methodology, named GMA-DRL, utilizes a spatial graph convolutional encoder with a multi-head attention mechanism to explicitly represent dynamic interaction among vehicles. Furthermore, deep reinforcement learning based on Actor-Critic with temporal relation regularization is utilized to ensure the consistency of dynamic interaction and generate cooperative driving actions for multiple CAVs. The GMA-DRL is tested in a series of typical merging scenarios with dynamic interaction. Extensive experimental results show that the GMA-DRL outperforms the existing cooperative control models in term of headway, average speed and acceleration. It demonstrates that the GMA-DRL with explicitly represent dynamic interaction can improve safety and efficiency of multiple CAVs in the merging zone.
Jinchao Hu, Xu Li 0004, Weiming Hu 0002, Qimin Xu, Dong Kong
IEEE Trans. Intell. Transp. Syst.2
2023 Explicit Points-of-Interest Driven Siamese Transformer for 3D LiDAR Place Recognition in Outdoor Challenging Environments
abstract
Place recognition plays a crucial role in simultaneous localization and mapping. Unfortunately, however, changes in viewpoints and conditions in large-scale environments impose tricky challenges for PR. To this end, this article specifically proposes an explicit points-of-interest driven PR method, which consists of a road segmentation module based on grid-wise patch U-transformer and a PR module based on regions of interest siamese transformer NetVLAD (RI_STV). Especially for RI_STV, in the individual dimension, it is dedicated to exploring the local topological features of nonroad regions of interest. In the spatial dimension, an improved Transformer is introduced to capture the global interactions between features of interest. In the cluster dimension, NetVLAD embedded with weighted pooling is created to perform weighted aggregation of feature clusters to generate discriminative and general descriptors. Evaluation on various datasets shows that our customized method is not only impressively competitive, but also strikes the best balance between accuracy and real-time performance.
Dong Kong, Xu Li 0004, Weiming Hu 0002, Jinchao Hu, Yue Hu 0009, Qimin Xu, Xiang Song 0004
IEEE Trans. Ind. Informatics2
2023 Highly Robust Vehicle Lateral Localization Using Multilevel Robust Network
abstract
Vision-based vehicle lateral localization has been extensively studied in the literature. However, it faces great challenges when dealing with occlusion situations where the road is frequently occluded by moving/static objects. To address the occlusion problem, we propose a highly robust lateral localization framework called multilevel robust network (MLRN) in this article. MLRN utilizes three deep neural networks (DNNs) to reduce the impact of occluding objects on localization performance from the object, feature, and decision levels, respectively, which shows strong robustness to varying degrees of road occlusion. At the object level, an attention-guided network (AGNet) is designed to achieve accurate road detection by paying more attention to the interested road area. Then, at the feature level, a lateral-connection fully convolutional denoising autoencoder (LC-FCDAE) is proposed to learn robust location features from the road area. Finally, at the decision level, a long short-term memory (LSTM) network is used to enhance the prediction accuracy of lateral position by establishing the temporal correlations of positioning decisions. Experimental results validate the effectiveness of the proposed framework in improving the reliability and accuracy of vehicle lateral localization.
Zhiyong Zheng, Xu Li 0004, Jianxiao Zhu, Jianhua Yuan, Linqi Wu
IEEE Trans. Neural Networks Learn. Syst.2
2022 A Roadside Decision-Making Methodology Based on Deep Reinforcement Learning to Simultaneously Improve the Safety and Efficiency of Merging Zone
abstract
The safety and efficiency of the merging zone is particularly important for traffic networks. Although autonomous vehicle improves the safety and efficiency from vehicle view, traffic controlling in merging zone mostly focus on improving efficiency from roadside view. Lacking of detailed driving recommendation, it ignores the safety of merging zone where commercial vehicle pose a high collision risk in real traffic. This paper proposes a roadside decision-making methodology to simultaneously improve the safety and efficiency of merging zone. We have built two modules, namely assessment and decision-making. Assessment module takes advantage of Bayesian inference to evaluate dynamic collision risk. Decision-making module based on deep reinforcement learning recommends the actions to commercial vehicles by roadside unit. A series of typical simulation tests show that our method increases the TTC of commercial vehicles by an average of 62.7%. In the free flow, the overall travel time of vehicles in the merging zone is reduced by 11.68%. Most notably, when congestion occurred, the average jam length is reduced by 59.68% on the premise of safety. Moreover, the average accuracy of the roadside decision-making method on the evaluation metrics of TTC, travel time, and jam length are 93.73%, 91.65%, and 94.45%, respectively. The experimental results show that the roadside decision-making methodology simultaneously improves safety and efficiency, and it dynamically adapts free and congested traffic flow.
Jinchao Hu, Xu Li 0004, Yanqing Cen, Qimin Xu, Weiming Hu 0002
IEEE Trans. Intell. Transp. Syst.2
2021 A Novel Visual Measurement Framework for Land Vehicle Positioning Based on Multimodule Cascaded Deep Neural Network
abstract
This article proposes a novel visual measurement framework, multimodule cascaded deep neural network (MMC-DNN), to achieve accurate, reliable, and cost-effective vehicle positioning in complex urban environments. The MMC-DNN is inspired by the mechanism of the human eyes' lateral positioning, which consists of three modules called siamesed fully convolutional network (S-FCN), skip-connection fully convolutional autoencoder (SC-FCAE), and multitask neural network regressor (MT-NNR), respectively. The S-FCN is first designed to accurately detect the road area. Then, the segmented road was executed inverse perspective mapping and the result is fed to the developed SC-FCAE for extracting equivalent positioning features. Furthermore, the MT-NNR is proposed to efficiently estimate lateral position and yaw angle with the help of a road map. Based on the estimation results, the MEMS INS/GPS integration is significantly augmented by extended Kalman filter. Experimental results validate the effectiveness of the proposed framework in enhancing positioning performance.
Zhiyong Zheng, Xu Li 0004, Zhengliang Sun, Xiang Song 0004
IEEE Trans. Ind. Informatics2
2021 Deep Inference Networks for Reliable Vehicle Lateral Position Estimation in Congested Urban Environments
abstract
Reliable estimation of vehicle lateral position plays an essential role in enhancing the safety of autonomous vehicles. However, it remains a challenging problem due to the frequently occurred road occlusion and the unreliability of employed reference objects (e.g., lane markings, curbs, etc.). Most existing works can only solve part of the problem, resulting in unsatisfactory performance. This paper proposes a novel deep inference network (DINet) to estimate vehicle lateral position, which can adequately address the challenges. DINet integrates three deep neural network (DNN)-based components in a human-like manner. A road area detection and occluding object segmentation (RADOOS) model focuses on detecting road areas and segmenting occluding objects on the road. A road area reconstruction (RAR) model tries to reconstruct the corrupted road area to a complete one as realistic as possible, by inferring missing road regions conditioned on the occluding objects segmented before. A lateral position estimator (LPE) model estimates the position from the reconstructed road area. To verify the effectiveness of DINet, road-test experiments were carried out in the scenarios with different degrees of occlusion. The experimental results demonstrate that DINet can obtain reliable and accurate (centimeter-level) lateral position even in severe road occlusion.
Zhiyong Zheng, Xu Li 0004, Qimin Xu, Xiang Song 0004
IEEE Trans. Image Process.2
2020 A novel vehicle lateral positioning methodology based on the integrated deep neural network
Zhiyong Zheng, Xu Li 0004
Expert Syst. Appl.2
2018 Modeling the Special Intersection for Enhanced Digital Map
abstract
The enhanced digital map is of great significance for various Intelligent Transportation System (ITS) applications and services, especially at the lane-level. This requirement motives the development of road modeling in the enhanced digital map at the lane-level. However, previous enhanced digital maps do not provide detailed modeling of special intersections which are covered by vegetation in the central region. In this paper we propose a novel lane-level road model for this special intersection scenario. The proposed intersection model can be considered into two levels: topological structure and geometrical structure. Topological structure of this model helps describe the connectivity, turn restrictions and other attributes of the special intersection in the real world. Geometrical structure of this model helps describe the virtual lanes of the internal part of the special intersection using cardinal spline, which better approximates the real vehicle trajectory at the intersection. The proposed intersection model has been verified and evaluated through experiments. The results demonstrate the effectiveness of the proposed intersection model in representing the lane-level topological and geometrical details of special intersection which is covered by vegetation in its central region.
Xu Li 0004, Xianghui Song, Qimin Xu
Intelligent Vehicles Symposium1
2016 A Reliable Hybrid Positioning Methodology for Land Vehicles Using Low-Cost Sensors
abstract
In this paper, we propose a reliable hybrid positioning methodology by combining the advantages of H∞filter and extreme learning machine (ELM), which addresses GPS outages and uncertain nonlinear drift of MEMS INS simultaneously. A novel parallel-dual-H∞filtering (PDHF) mechanism is proposed to prevent the H∞filter from diverging during GPS outages and to make full use of supplementary observations. The PDHF is composed of an enhanced H∞filter and an auxiliary H∞filter. The enhanced H∞filter is developed by fusing not only GPS information but also supplementary observations, which include yaw angle provided by electronic compass, longitudinal velocity derived from wheel speed sensor, and lateral velocity constrained by assumptions, whereas the auxiliary H∞filter only fuses the supplementary observations. Furthermore, an ELM module with good generalization ability is designed and augmented with the auxiliary H∞filter to constitute a “virtual enhanced H∞filter.” In the case of GPS outages, the “virtual enhanced H∞filter” provides accurate corrections for stand-alone INS. Due to the characteristics of the Hoc filter, the proposed methodology is immune to uncertain nonlinear drift of MEMS INS in nature. To verify the effectiveness of the proposed methodology, road-test experiments with various scenarios were performed. The experimental results indicate that the proposed methodology outperform all the compared counterparts.
Qimin Xu, Xu Li 0004, Xianghui Song, Zhixiang Cai
IEEE Trans. Intell. Transp. Syst.2