Xiaojun Tan

dblp:20/5390 · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Confidence-V2X: Confidence-driven sparse communication for efficient V2X cooperative perception
Xiaojun Tan, Dongsheng Wu
Adv. Eng. Informatics1
2026 Solving the vehicle coordination problem towards unsignalized intersections using risk-aware graph attention and deep reinforcement learning approach
Yuhao Ding, Xiaojun Tan
Eng. Appl. Artif. Intell.3
2025 Goal-guided multi-agent motion prediction with interactive state refinement
Xiangyi Qin, Xiaojun Tan
Adv. Eng. Informatics5
2025 RRT*ASV: Improved RRT* path planning method for Ackermann steering vehicles
Chaofan Lei, Yushan Deng, Xiaojun Tan
Expert Syst. Appl.4
2025 Multimodal end-to-end autonomous driving via bilateral modality interaction
Jun Li 0105, Zesong Chen, Xiaojun Tan
Expert Syst. Appl.6
2025 Solving the Vehicle Cooperation Problem at Signal-Free Intersection via an Asynchronous Deep Reinforcement Learning Approach
abstract
With the rapid development of modern intelligent transportation systems, connected and automated vehicles (CAVs) have garnered significant attention due to their advanced communication and decision-making capabilities. Intelligent collaborative decision-making in dynamic traffic scenarios poses significant challenges in the research of CAV technology. Strategies and methods have been developed to tackle the collaboration problem toward different scenarios. However, most existing studies have not been fully investigated the performance optimization and practicability of the vehicle collaboration at signal-free intersections. Meanwhile, the urban signal-free intersections represent a critical application scenario for vehicle cooperation, where a thorough study has not been given yet. Therefore, this study intends to solve the vehicle collaboration problem utilizing the deep reinforcement learning approach. Initially, the problem is formulated as an elaborated Markov decision process, comprising the state space, the action space, and the reward function. Then, a shared advantage actor-critic (A2C) model is proposed to effectively extract temporal and spatial features through the shared network, thereby improving the consistency of feature learning processes between the Actor and Critic networks. Furthermore, the asynchronous training strategy employed in this study involves multiple training processes concurrently, thereby enhancing the model’s convergence speed and stability. Finally, the effectiveness of the proposed method is verified in two typical intersection scenarios. Simulation results reveal that our method exhibits competitive performance compared with existing approaches, and an improvement up to 30% and 40% can be achieved in the retreat time and the averaged delay time. Additionally, the field experiments have been conducted on a miniaturized autonomous driving platform, verifying the considerable potential of the proposed method for real-world applications.
Shuai Wang 0009, Yuhao Ding, Xiaoqi Ding, Xiaojun Tan
IEEE Internet Things J.4
2025 MAEMOT: Pretrained MAE-Based Antiocclusion 3-D Multiobject Tracking for Autonomous Driving
abstract
The existing 3-D multiobject tracking (MOT) methods suffer from object occlusion in real-world traffic scenes. However, previous works have faced challenges in providing a reasonable solution to the fundamental question: "How can the interference of the perception data loss caused by occlusion be overcome?" Therefore, this article attempts to provide a reasonable solution by developing a novel pretrained movement-constrained masked autoencoder (M-MAE) for an antiocclusion 3-D MOT called MAEMOT. Specifically, for the pretrained M-MAE, this article adopts an efficient multistage transformer (MST) encoder and a spatiotemporal-based motion decoder to predict and reconstruct occluded point cloud data, following the properties of object motion. Afterward, the well-trained M-MAE model extracts the global features of occluded objects, ensuring that the features of the intraobjects between interframes are as consistent as possible throughout the spatiotemporal sequence. Next, a proposal-based geometric graph aggregation (PG2A) module is utilized to extract and fuse the spatial features of each proposal, producing refined region-of-interest (RoI) components. Finally, this article designs an object association module that combines geometric and corner affinities, which helps to match the predicted occlusion objects more robustly. According to an extensive evaluation, the proposed MAEMOT method can effectively overcome the interference of occlusion and achieve improved 3-D MOT performance under challenging conditions.
Zhengping Fan, Ying Shen 0001, Yasong An, Xiaojun Tan
IEEE Trans. Neural Networks Learn. Syst.6
2024 BEVLOC: End-to-End 6-DoF Localization Via Cross-Modality Correlation Under Bird's Eye View
abstract
Accurate ego-centric localization assumes a paramount significance in the domain of autonomous driving. However, traditional methods for camera-LiDAR map localization rely on perspective projection to create a unified representation, which often falls short due to challenges such as occlusion and the sparse nature of point cloud data. Despite the recent surge in popularity of the Bird’s-Eye-View (BEV) paradigm within autonomous driving, its potential applications in localization tasks have remained relatively underexplored. In response to this concern, this paper presents a pioneering end-to-end approach called the BEV Localization Network via LiDAR Map (BEVLoc). By fusing the image and LiDAR map in the BEV space via the concept of optical flow-based correlation, the BEVLoc framework can leverage the synergistic power of cross-modalities in localizing the vehicle. Experimental results conducted on the KITTI dataset highlight the efficacy and performance of BEVLoc in the realm of autonomous vehicle localization.
Nanjie Chen, Shuai Wang 0009, Xiaojun Tan
ICASSP6
2024 DualAT: Dual Attention Transformer for End-to-End Autonomous Driving
abstract
The effective reasoning of integrated multimodal perception information is crucial for achieving enhanced end-to-end autonomous driving performance. In this paper, we introduce a novel multitask imitation learning framework for end-to-end autonomous driving that leverages a dual attention transformer (DualAT) to enhance the multimodal fusion and waypoint prediction processes. A self-attention mechanism captures global context information and models the long-term temporal dependencies of waypoints for multiple time steps. On the other hand, a cross-attention mechanism implicitly associates the latent feature representations derived from different modalities through a learnable geometrically linked positional embedding. Specifically, the DualAT excels at processing and fusing information from multiple camera views and LiDAR sensors, enabling comprehensive scene understanding for multitask learning. Furthermore, the DualAT introduces a novel waypoint prediction architecture that combines the temporal relationships between waypoints with the spatial features extracted from sensor inputs. We evaluate our approach on both the Town05 and Longest6 benchmarks using the closed-loop CARLA urban driving simulator and provide extensive ablation studies. The experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods.
Zesong Chen, Jun Li 0075, Linlin You, Xiaojun Tan
ICRA5
2024 Sensor Fusion and Motion Planning with Unified Bird's-Eye View Representation for End-to-end Autonomous Driving
abstract
End-to-end autonomous driving has made significant advancements in these years. Efficiently fusing multi-modal sensor information to enhance the scene understanding capability and motion planning performance of end-to-end models is currently a prominent research topic. Existing methods fuse multimodal information from different views, which lack efficiency and restrict the sensor extension of the models. In this work, we propose an end-to-end autonomous driving method with unified multi-modal feature representation. Construction of the entire end-to-end autonomous driving model, including sensor fusion and motion planning, is conducted in a unified Bird’s-Eye View(BEV) representation space. We summarize our proposed method into three stages: multi-modal BEV feature construction, multi-modal BEV feature fusion and motion planning in BEV space. This construction method significantly improves the environmental perception capability and planning performance of the end-to-end model. We conduct closed-loop experiments in the Carla simulator and our method achieves superior performance compared with other methods.
Yuandong Lyu, Xiaojun Tan, Zhengping Fan
IJCNN2
2024 OATracker: Object-aware anti-occlusion 3D multiobject tracking for autonomous driving
Xiaojun Tan, Yasong An, Zhengping Fan
Expert Syst. Appl.2
2024 PDLC-LIO: A Precise and Direct SLAM System Toward Large-Scale Environments With Loop Closures
abstract
As a key technology of autonomous driving, the high-precision vehicle localization is the prerequisite for automatic steer. With the advantage of accuracy and robustness, the Light Detection and Ranging (LiDAR) is widely employed in Simultaneous Localization and Mapping (SLAM), which has been utilized to provide reliable positioning information. Compared with the experimental vehicles in the restricted laboratory driving circumstances, autonomous vehicles are devoted to actual roads and contain complicated activities, where closed loops can be noticed. In general, the accumulated error may be accumulated with the increasing of maneuver scope of SLAM. A distinct drift phenomenon may thus be incurred in the completed point cloud map and threaten the accuracy of positioning results. Focusing on this deficiency, a tightly-coupled LiDAR inertial odometry, named PDLC-LIO, is developed in this paper towards large-scale environments with close loops. A direct point cloud registration approach without extracting features has been introduced into this framework. Strategies including the pre-integration on the Inertial Measurement Unit (IMU), the direct scan match in local-scale, and an efficient fusion loop closure detection method with conditional detection are included to guarantee the precision. Also, the factor graph optimization is considered, and measurements from LiDAR odometry, IMU and loop closures can thus be integrated into the back-end optimization. The proposed method has been verified on both the public datasets and the test cases collected by our own autonomous driving experiment platform. Accurate experimental results can be obtained. Such results show that the performance including accumulated errors and the drift phenomenon outperforms the state-of-the-art algorithms LeGO-LOAM/ LIO-SAM /FAST-LIO2, and the loop closure detection rate can be promoted up to 75%.
Shuai Wang 0009, Jiazhong Zhang, Xiaojun Tan
IEEE Trans. Intell. Transp. Syst.3
2023 LMBAO: A Landmark Map for Bundle Adjustment Odometry in LiDAR SLAM
abstract
Existing LiDAR odometry strategies match a new scan iteratively with previous fixed-pose scans, gradually accumulating errors. Furthermore, as an effective joint optimization mechanism, bundle adjustment (BA) cannot be directly introduced into odometry due to the intensive computation of global landmarks. Therefore, this paper designs a landmark map for bundle adjustment odometry (LMBAO) in LiDAR SLAM. First, an active landmark maintenance strategy is developed to obtain a local map of limited size that enables real-time BA. Specifically, this paper keeps entire stable landmarks on the map instead of just their feature points in the sliding window and timely deletes inactive landmarks. Next, unlike visual marginalization to approximate the Gaussian distribution, and a direct and efficient marginalization strategy is performed to retain the scans outside the window to greatly simplifying the computation. Experiments show the effectiveness of LMBAO in outdoor driving.
Lu Jie 0003, Nanjie Chen, Xiaojun Tan, Zhifei Duan
ICASSP5
2023 Short-Term EV Charging Load Predicting Based on Adaptive VMD and LSTM Methods
abstract
The uncoordinated charging of large-scale electric vehicles (EVs) generally deteriorates the peak-valley difference of daily electric demands. To facilitate the operation of charging stations and electric power distributers, this work proposes a charging load prediction algorithm by combining the Variational Mode Decomposition (VMD) and the Long Short-Term Memory (LSTM) methods. The VMD is adopted to extract the EV charging load features at different time scales, obtaining multiple intrinsic mode functions (IMFs). Then the LSTM establishes the dependencies between these IMFs of historical data and the predicted load. To trade-off between the prediction accuracy and the computation overhead, an additional Snake Optimization (SO) technique is applied to adaptively optimize the VMD parameters. Experimental results show that the proposed algorithm outperforms the traditional LSTM alone and the Gate Recurrent Unit alone neural networks in terms of the overall prediction accuracy. The proposed parallel LSTM structure with the optimized VMD further reduces the Root Mean Square Error (RMSE) and the Mean Absolute Error (MAE) significantly by 55.1% and 55.9% with respect to the LSTM method with non-optimized VMD.
Quanxue Guan, Qinhe Liu, Yunjian Xu, Xiaojun Tan
IECON5
2023 A Memetic algorithm for determining robust and influential seeds against structural perturbances in competitive networks
Shuai Wang 0009, Xiaojun Tan
Inf. Sci.2
2023 Spatiotemporal adaptive attention 3D multiobject tracking for autonomous driving
Zhengping Fan, Xiaojun Tan, Qunming Liu, Yanli Shi
Knowl. Based Syst.3
2023 Mutually Beneficial Transformer for Multimodal Data Fusion
abstract
Multimodal feature fusion representation, e.g., hyperspectral image and light detection and ranging (HSI-LiDAR) fusion, is an essential topic for fusion perception. However, existing networks tend to employ mandatory feature stacking or local context fusion strategies between multiple modalities, ignoring the power of globally mutual-guided feature transmission. Therefore, this paper develops a mutually beneficial transformer method for multimodal data fusion (MBFormer), which contains the following steps. First, a spatial constraint-based self-attention (SCS) module. In this module, spectralwise attention and a spatialwise convolution are applied to HSI and LiDAR data individually, and then a spatial guide mask generated from LiDAR elevation information is used as an agent to bridge with HSI for spatial feature constraints. Second, a channel diversity-based transformer (CDT) module. On the basis of local spectral embedding explorations, an adaptive token-mixer mechanism is conducted on the groupwise classification token of HSI and individual LiDAR data for global information connectivity and transitivity. At last, the selected features are embedded into a classification layer for the final result calculation. Experimental results show that the proposed MBFormer can obtain 97.76% and 98.62% classification accuracies on Houston and Trento datasets, respectively, indicating the advantages and competitiveness of the MBFormer over the compared state-of-the-art methods.
Xiaojun Tan
IEEE Trans. Circuits Syst. Video Technol.2
2023 Point-Guided Contrastive Learning for Monocular 3-D Object Detection
abstract
3-D object detection is a fundamental task in the context of autonomous driving. In the literature, cheap monocular image-based methods show a significant performance drop compared to the expensive LiDAR and stereo-images-based algorithms. In this article, we aim to close this performance gap by bridging the representation capability between 2-D and 3-D domains. We propose a novel monocular 3-D object detection model using self-supervised learning and auxiliary learning, resorting to mimicking the representations over 3-D point clouds. Specifically, given a 2-D region proposal and the corresponding instance point cloud, we supervise the feature activation from our image-based convolution network to mimic the latent feature of a point-based neural network at the training stage. While state-of-the-art (SOTA) monocular 3-D detection algorithms typically convert images to pseudo-LiDAR with depth estimation and regress 3-D detection with LiDAR-based methods, our approach seeks the power of the 2-D neural network straightforwardly and essentially enhances the 2-D module capability with latent spatial-aware representations by contrastive learning. We empirically validate the performance improvement from the feature mimicking the KITTI and ApolloScape datasets and achieve the SOTA performance on the KITTI and ApolloScape leaderboard.
Dapeng Feng, Songfang Han, Hang Xu 0004, Xiaodan Liang, Xiaojun Tan
IEEE Trans. Cybern.5
2023 Estimating Human Weight From a Single Image
abstract
Body weight, as one of the biometric traits, has been studied in both the forensic and medical domains. However, estimating weight directly from 2-D images is particularly challenging since visual inspection is rather sensitive to the distance between the subject and camera, even for frontal view images. In this case, the widely used body mass index (BMI), which is associated with body height and weight, can be employed as a measure of weight to indicate health conditions. Previous works on the estimation of BMI have predominantly focused on using multiple 2-D images, 3-D images, or facial images; however, these cues are not always available. To address this issue, we explore the feasibility of obtaining BMI from a single 2-D body image with the dual-branch regression framework proposed in this work. More specifically, the framework comprises an anthropometric feature computation branch and a deep learning-based feature extraction branch. One aggregation layer maps all the features to an estimated BMI value. In addition, a new public 2-D image-to-BMI dataset, which contains 4189 images (1477 males and 2712 females) from approximately 3000 subjects with attributes including gender, age, height, and weight, was collected and released to facilitate the study. Extensive experiments confirm that the proposed framework combining anthropometric features and deep features outperforms the single-type feature approaches to BMI estimation in most cases.
Zhi Jin 0002, Junjia Huang, Wenjin Wang 0002, Aolin Xiong, Xiaojun Tan
IEEE Trans. Multim.5
2022 Spectral-Spatial Symmetrical Aggregation Cross-Linking Multi-Modal Data Fusion Network
abstract
In this paper, a spectral-spatial symmetrical aggregation cross-linking network (SACLNet) is developed for multi-modal data classification, which contains three modules as follows. First, the Spectro-Spatial Feature Learning Module is proposed, using the involution operation sliding over the spectral channels of hyperspectral image (HSI) and fused-sharing weight obtained from HSI and light detection and ranging (HSI-LiDAR) data for spectral and spatial information representation. Second, the pyramid feature fusion module interacts among spectral and spatial information to further share and guide each other. In this step, multistage features, including low-level, middle-level, and high-level, are fused and adjusted in a pyramidal and mutually guided learning process. Third, the fused Spectro-Spatial features are embedded into the Multimodal Data Fusion Module, obtaining the final classification results. Experimental results show that the proposed SACLNet has a satisfactory classification performance than the state-of-the-art methods.
Jun Li 0075, Xiaojun Tan
ICASSP3
2022 Radio Resource Selection in C-V2X Mode 4: A Multiagent Deep Reinforcement Learning Approach
abstract
The Third Generation Partnership Project has standardized cellular vehicle-to-everything (C-V2X) sidelink Mode 4 communication to support vehicle-to-vehicle safety applications. In Mode 4, the sensing-based semi-persistent scheduling (SPS) scheme allows vehicles to select radio resources autonomously. In particular, SPS has three steps to generate available resource lists (ARLs) for resource selection. However, the overlapping of ARLs is inevitable when radio resources are insufficient. In this case, randomly selecting from ARLs according to SPS is likely to make adjacent vehicles select the same resources, resulting in packet collisions. Unlike SPS, this paper proposes a multiagent deep reinforcement learning-based SPS (RL-SPS) algorithm to help vehicles select appropriate radio resources. As a consequence, packet collisions can be largely reduced. Furthermore, a centralized multi-head attention mechanism is adopted to improve the efficiency of the training process of RL-SPS. Simulation results demonstrate the reliability, scalability and robustness of the RL-SPS in a dynamic vehicular network.
Weixiang Chen, Bo Gu 0003, Xiaojun Tan, Chenhua Wei
ICCCN3
2022 Local-guided Global Collaborative Learning Transformer for Vehicle Reidentification
abstract
Vehicle reidentification(ReID) has attracted much attention and is significant for traffic security surveillance. Due to the variety of views of the same vehicle captured by different camera and the great similarity in the visual appearance of different vehicles, it is necessary to explore how to effectively utilize local detail information to achieve collaborative perception to highlight discriminative appearance features. Different from existing local feature exploration methods that focus on using extra part or keypoint information, we propose a global collaborative learning Transformer guided by local abstract features, named LG-CoT, which aims to highlight the highest-attention regions of vehicle images. We adopt Vision Transformer(ViT) as our backbone to extract global features and obtain all local tokens. To reduce the distribution from the background and drive the network to focus more on details, all attention maps containing low-level texture information and high-level semantic information are multiplied to obtain the local regions with highest-attention. Finally, we design a local-attention-guided pose-optimization feature encoding module, which can help the global features focus on local regions adaptively. Extensive experiments on two popular datasets and a dataset we built in a T-junction traffic scene suggest that our method can achieve comparable performance.
Yanli Shi, Xiaojun Tan
ICTAI3
2022 Deep Reinforcement Learning Based Radio Resource Selection Approach for C- V2X Mode 4 in Cooperative Perception Scenario
abstract
In recent years, vehicles have been equipped with multiple sensors to enable assisted driving and even autonomous driving. However, due to the physical characteristics of the sensors, there are numerous shortcomings in the perception of the surrounding environment by a single vehicle. The development of vehicle-to-everything technology enables vehicles to extend their sensing range or enhance the reliability of perception by exchanging sensor data via vehicle-to-vehicle communication, which is called cooperative perception. In cellular vehicle-to-everything Mode 4, vehicles use the sensing-based semi-persistent scheduling scheme to select radio resource autonomously before transmission. But this scheme is hardly adaptable to cooperative perception scenario due to the time-sensitive of cooperative perception and the impact caused by the position of the per-ception information. In this paper, we modeled the cooperative perception scenario and the communication between vehicles, and then we formulated the optimization objective considering the characteristics of cooperative perception. Finally, we propose a multi-agent deep reinforcement learning based resource selection algorithm to tackle this problem and demonstrate its effectiveness through simulations.
Chenhua Wei, Xiaojun Tan
MSN2
2022 ASPCNet: Deep adaptive spatial pattern capsule network for hyperspectral image classification
Xiaojun Tan, Jian-Huang Lai, Jun Li 0075
Neurocomputing2
2022 VAERHNN: Voting-averaged ensemble regression and hybrid neural network to investigate potent leads against colorectal cancer
Guanxing Chen, Xuefei Jiang, Qiujie Lv, Xiaojun Tan, Zihuan Yang, Calvin Yu-Chian Chen
Knowl. Based Syst.4
2022 Determining seeds with robust influential ability from multi-layer networks: A multi-factorial evolutionary approach
Shuai Wang 0009, Xiaojun Tan
Knowl. Based Syst.2
2022 AM³Net: Adaptive Mutual-Learning-Based Multimodal Data Fusion Network
abstract
Multimodal data fusion, e.g., hyperspectral image (HSI) and light detection and ranging (LiDAR) data fusion, plays an important role in object recognition and classification tasks. However, existing methods pay little attention to the specificity of HSI spectral channels and the complementarity of HSI and LiDAR spatial information. In addition, the utilized feature extraction modules tend to consider the feature transmission processes among different modalities independently. Therefore, a new data fusion network named AM3Net is proposed for multimodal data classification; it includes three parts. First, an involution operator slides over the input HSI’s spectral channels, which can independently measure the contribution rate of the spectral channel of each pixel to the spectral feature tensor construction. Furthermore, the spatial information of HSI and LiDAR data is integrated and excavated in an adaptively fused, modality-oriented manner. Second, aspectral-spatial mutual-guided moduleis designed for the feature collaborative transmission among spectral features and spatial information, which can increase the semantic relatedness connection through adaptive, multiscale, and mutual-learning transmission. Finally, the fused spatial-spectral features are embedded into a classification module to obtain the final results, which determines whether to continue updating the network weights. Experimental evaluations on HSI-LiDAR datasets indicate that AM3Net possesses a better feature representation ability than the state-of-the-art methods. Additionally, AM3Net still maintains considerable performance when its input is replaced with multispectral and synthetic aperture radar data. The result indicates that the proposed data fusion framework is compatible with diversified data types.
Jun Li 0075, Yanli Shi, Jian-Huang Lai, Xiaojun Tan
IEEE Trans. Circuits Syst. Video Technol.5
2020 Towards Lighter and Faster: Learning Wavelets Progressively for Image Super-Resolution
abstract
Due to the significant development of deep learning (DL) techniques, recent advances in the super-resolution (SR) field have achieved a great performance. While seeking for better performance, the later proposed networks prone to be deeper and heavier, which limits the applications of SR algorithms in the resource-constrain devices. Some advances rely on recurrent/recursive learning to reduce the number of network parameters, however, they ignore the caused long inference time, since the more recurrences/recursions are involved, the longer inference time the network needs. To address this trade-off issue between reconstruction performance, the number of network parameters, and inference time, we propose a lightweight and fast network (WSR) to learn wavelet coefficients of the target image progressively for single image super-resolution. More specifically, the network comprises two main branches. One is used for predicting the second level low-frequency wavelet coefficients, and the other one is designed in a recurrent way for predicting the rest wavelet coefficients at the first and second levels. Finally, an inverse wavelet transformation is adopted to reconstruct the SR images from these coefficients. In addition, we propose a deformable convolution kernel (side window) to construct the side-information multi-distillation block (S-IMDB), which is the basic unit of the recurrent blocks (RBs). We train the WSR with loss constraints at wavelet and spatial domains. Comprehensive experiments demonstrate that our WSR achieves a better trade-off than most of the state-of-the-art approaches. Code is available at https://github.com/FVL2020/WSR.
Huanrong Zhang, Zhi Jin 0002, Xiaojun Tan
ACM Multimedia3