Xiaolei Li 0003

dblp:77/709-3 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0001-9848-9200ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 36% Face, body and person analysis · 18% Representation and self-supervised learning · 15%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Smart cities and intelligent transportation · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › actor-critic methods
asymmetric actor-critic
0.912025
Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning · IEEE Trans. Robotics 2025
Machine learning › Reinforcement learning › exploration › autonomous exploration › mobile robot exploration
cooperative exploration
0.912025
Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning · IEEE Trans. Robotics 2025
Machine learning › Reinforcement learning
deep reinforcement learning
0.912025
Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning · IEEE Trans. Robotics 2025
Machine learning › Reinforcement learning › exploration
multi-robot exploration
0.912025
Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning · IEEE Trans. Robotics 2025
Computer vision › Face, body and person analysis
person re-identification
0.912025
ViV-ReID: Bidirectional Structural-Aware Spatial-Temporal Graph Networks on Large-Scale Video-Based Vessel Re-Identification Dataset · IEEE Trans. Image Process. 2025
Machine learning › Graph learning › spatio-temporal graph learning
spatio-temporal graph network
0.912025
ViV-ReID: Bidirectional Structural-Aware Spatial-Temporal Graph Networks on Large-Scale Video-Based Vessel Re-Identification Dataset · IEEE Trans. Image Process. 2025
Computer vision › Face, body and person analysis › person re-identification
video-based person re-identification
0.912025
ViV-ReID: Bidirectional Structural-Aware Spatial-Temporal Graph Networks on Large-Scale Video-Based Vessel Re-Identification Dataset · IEEE Trans. Image Process. 2025
Machine learning › Representation and self-supervised learning
mutual information maximization
0.812024
Neighborhood-Aware Mutual Information Maximization for Source-Free Domain Adaptation · IEEE Trans. Multim. 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
source-free domain adaptation
0.812024
Neighborhood-Aware Mutual Information Maximization for Source-Free Domain Adaptation · IEEE Trans. Multim. 2024
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
unsupervised embedding learning
0.712023
Unsupervised Embedding Learning With Mutual-Information Graph Convolutional Networks · IEEE Trans. Multim. 2023
Machine learning › Generative modeling
generative adversarial network
0.512021
PoT-GAN: Pose Transform GAN for Person Image Synthesis · IEEE Trans. Image Process. 2021
Machine learning › Generative modeling › image generation
person image synthesis
0.512021
PoT-GAN: Pose Transform GAN for Person Image Synthesis · IEEE Trans. Image Process. 2021
Robotics › Robot navigation and mapping › robot mapping
topological mapping
0.312025
Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning · IEEE Trans. Robotics 2025
Machine learning › Graph learning › graph neural network
graph convolutional network
0.212023
Unsupervised Embedding Learning With Mutual-Information Graph Convolutional Networks · IEEE Trans. Multim. 2023

Methods — techniques the papers use, named apart from their topics

spatial-temporal feature alignment · 1.7graph neural network · 1.7mutual information · 1.5data augmentation · 1.4topological graph matching · 0.9deep reinforcement learning · 0.9self-supervised learning · 0.8contrastive learning · 0.8graph convolutional network · 0.7multi-scale feature map manipulation · 0.5
YearPublicationVenuePosition
2026 RDPrompter: Reference-Defect Prompt Learning for Few-Shot Defect Segmentation Based on Visual Foundation Model
abstract
Few-shot industrial defect segmentation (FIDS) is an extremely challenging task in industrial inspection, which focuses on segmenting unseen defect categories with only a few samples. Existing few-shot segmentation methods are commonly developed with constrained feature extraction capabilities, making it difficult to accurately segment diverse unseen defects from distinctive industrial scenarios. To this end, we aim to leverage strong generalization strengths of the large-scale visual foundation model, segment anything model (SAM), to handle FIDS task. However, simply incorporating SAM into FIDS is insufficient for automatically categorizing defects, as the SAM heavily relies on manual prompts for segmentation. To address this, we propose a novel reference-defect prompt learning method RDPrompter, which learns adaptive prompts from defect similarities based on SAM for guiding automated FIDS. Specifically, to obtain adaptive prompts, we introduce a multiscale feature extraction strategy to mine rich defect features by the image encoder of SAM. Then, self-correcting probability prototype and cosine similarity attention are proposed to form a mixed similarity aggregation module, which obtains multiscale feature similarities for defect localization. Based on the similarities, a prompt embedding construction module is designed to further extract fine defect location information and generate prompt embeddings for FIDS. Sufficient experiments on MVTec-Unseen, SDD, FSSD-12 and the collected CID datasets demonstrate the superiority of our method.
Tiyu Fang, Lin Zhang 0041, Ran Song 0001, Xiaolei Li 0003, Wei Zhang 0021
IEEE Trans. Ind. Informatics4
2025 Transformer-Driven Semantic-Spatial Adaptive Fusion Representation for Object-Goal Navigation
abstract
Visual object-goal navigation requires an agent to make decisions to search for specified target objects within an unknown environment. While learning-based approaches have achieved progress, they still face two limitations: (1) visual representation: the lack of a low-dimensional representation that balances object priors with current observations, and the absence of effective modeling of scene spatial structures, which restricts the agent’s ability to perceive its surroundings. (2) policy learning: existing reward functions often fail to consider the visual essence of object goal-driven tasks, leading to inefficient learning and suboptimal performance. To address these issues, this paper introduces a semantic-spatial fusion representation framework that incorporates both object semantics and the intrinsic spatial structure of the scene. Specifically, a new object context matrix captures semantic relationships between objects while providing distinct low-dimensional representations for different observations, and an episode memory graph is also constructed and dynamically updated based on observation similarity and geodesic distance to represent the spatial structure of the agent’s environment in real-time. Then, a transformer-based adaptive multimodal feature fusion module is proposed to integrate these dual representations. Moreover, a sparse reward function is designed based on the target’s bounding box to guide the agent to learn correctly. The proposed method is the first to be evaluated on both the simulation platform AI2-THOR and the real-world dataset AVD, demonstrating its generalization in unseen environments. The method is also deployed on a physical mobile robot and tested in real-world scenarios, further validating its practical effectiveness.
Fei Lu 0005, Xiaolei Li 0003, Guohui Tian, Tengfan Fu
IEEE Trans Autom. Sci. Eng.3
2025 ViV-ReID: Bidirectional Structural-Aware Spatial-Temporal Graph Networks on Large-Scale Video-Based Vessel Re-Identification Dataset
abstract
Vessel re-identification (ReID) serves as a foundational task for intelligent maritime transportation systems. To enhance maritime surveillance capabilities, this study investigates video-based vessel ReID, a critical yet underexplored task in intelligent transportation systems. The lack of relevant datasets has limited the progress of Video-based vessel ReID research work. We established ViV-ReID, the first publicly available large-scale video-based vessel ReID dataset, comprising 480 vessel identities captured from 20 cross-port camera views (7,165 tracklets and 1.14 million frames), establishing a benchmark for advancing vessel ReID from image to video processing. Videos offer significantly richer information than single-frame images. The dynamic nature of video often leads to fragmented spatio-temporal features causing disrupted contextual understanding, and to address this problem, we further propose a Bidirectional Structural-Aware Spatial-Temporal Graph Network (Bi-SSTN) that explicitly aligns spatio-temporal features using vessel structural priors. Extensive experiments on the ViV-ReID dataset demonstrate that image-based ReID methods often show suboptimal performance when applied to video data. Meanwhile, it is crucial to validate the effectiveness of spatio-temporal information and establish performance benchmarks for different methods. The Bidirectional Structural-Aware Spatial-Temporal Graph Network (Bi-SSTN) significantly outperforms state-of-the-art methods on ViV-ReID, confirming its efficacy in modeling vessel-specific spatio-temporal patterns. Project web page: https://vsislab.github.io/ViV_ReID/.
Mingxin Zhang 0006, Fuxiang Feng, Lin Zhang 0041, Youmei Zhang, Xiaolei Li 0003, Wei Zhang 0021
IEEE Trans. Image Process.6
2025 Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning
abstract
Despite significant advancements in multirobot technologies, efficiently and collaboratively exploring an unknown environment remains a major challenge. In this paper, we propose AIM-Mapping, an Asymmetric InforMation enhanced Mapping framework based on deep reinforcement learning. The framework fully leverages the privileged information to help construct the environmental representation as well as the supervised signal in an asymmetric actor-critic training framework. Specifically, privileged information is used to evaluate exploration performance through an asymmetric feature representation module and a mutual information evaluation module. The decision-making network employs the trained feature encoder to extract structural information of the environment and integrates it with a topological map constructed based on geometric distance. By leveraging this topological map representation, we apply topological graph matching to assign corresponding boundary points to each robot as long-term goal points. We conduct experiments in both iGibson simulation environments and real-world scenarios. The results demonstrate that the proposed method achieves significant performance improvements compared to existing approaches.
Jiyu Cheng, Junhui Fan, Xiaolei Li 0003, Paul L. Rosin, Yibin Li 0001, Wei Zhang 0021
IEEE Trans. Robotics3
2024 MPGNet: Learning Move-Push-Grasping Synergy for Target-Oriented Grasping in Occluded Scenes
abstract
This paper focuses on target-oriented grasping in occluded scenes, where the target object is specified by a binary mask and the goal is to grasp the target object with as few robotic manipulations as possible. Most existing methods rely on a push-grasping synergy to complete this task. To deliver a more powerful target-oriented grasping pipeline, we present MPGNet, a three-branch network for learning a synergy between moving, pushing, and grasping actions. We also propose a multi-stage training strategy to train the MPGNet which contains three policy networks corresponding to the three actions. The effectiveness of our method is demonstrated via both simulated and real-world experiments. Video of the real-world experiments is at https://youtu.be/S_QKZqkh0w8.
Dayou Li, Chenkun Zhao, Ran Song 0001, Xiaolei Li 0003, Wei Zhang 0021
IROS5
2024 Neighborhood-Aware Mutual Information Maximization for Source-Free Domain Adaptation
abstract
Recently, the source-free domain adaptation (SFDA) problem has attracted much attention, where the pre-trained model for the source domain is adapted to the target domain in the absence of source data. However, due to domain shift, the negative alignment usually exists between samples from the same class, which may lower intra-class feature similarity. To address this issue, we present a self-supervised representation learning strategy for SFDA, named as neighborhood-aware mutual information (NAMI), which maximizes the mutual information (MI) between the representations of target samples and their corresponding neighbors. Moreover, we theoretically demonstrate that NAMI can be decomposed into a weighted sum of local MI, which suggests that the weighted terms can better estimate NAMI. To this end, we introduce neighborhood consensus score over the set of weakly and strongly augmented views and point-wise density based on neighborhood, both of which determine the weights of local MI for NAMI by leveraging the neighborhood information of samples. The proposed method can significantly handle domain shift and adaptively reduce the noise in the neighborhood of each target sample. In combination with the consistency loss over views, NAMI leads to consistent improvement over existing state-of-the-art methods on three popular SFDA benchmarks.
Lin Zhang 0041, Yifan Wang 0020, Ran Song 0001, Mingxin Zhang 0006, Xiaolei Li 0003, Wei Zhang 0021
IEEE Trans. Multim.5
2023 Unsupervised Embedding Learning With Mutual-Information Graph Convolutional Networks
abstract
Recently, methods for unsupervised embedding learning have exhibited promising results for extracting desirable representations from unlabeled samples. In general, most methods learn the feature embeddings by handling each sample individually while the structural and semantic relationships between samples are not fully exploited. As a result, the learned embeddings are not sufficiently discriminative. To make use of such inter-sample information for deep embedding learning, this paper proposes an unsupervised method based on the graph convolutional network (GCN). On one hand, our method encodes structural information between the samples corresponding to the nodes in a local neighbourhood of the GCN graph. On the other hand, it leverages the mutual information between the original samples and the augmented ones to ensure that they are globally consistent with each other. Extensive experiments show that our method is not just robust to augmentation perturbations, but also learns discriminative embeddings. Consequently, it achieves the state-of-the-art performance on several challenging datasets.
Lin Zhang 0041, Mingxin Zhang 0006, Ran Song 0001, Ziying Zhao, Xiaolei Li 0003
IEEE Trans. Multim.5
2022 LiTMNet: A deep CNN for efficient HDR image reconstruction from a single LDR image
Guotao Wu, Ran Song 0001, Mingxin Zhang 0006, Xiaolei Li 0003, Paul L. Rosin
Pattern Recognit.4
2021 EFNet: Enhancement-Fusion Network for Semantic Segmentation
Zhijie Wang 0010, Ran Song 0001, Peng Duan 0002, Xiaolei Li 0003
Pattern Recognit.4
2021 Triple-Input-Unsupervised neural Networks for deformable image registration
Wan Wan, Kaixuan Guo, Jun Tang 0007, Xiaolei Li 0003, Jun Wu 0024
Pattern Recognit. Lett.6
2021 PoT-GAN: Pose Transform GAN for Person Image Synthesis
abstract
Pose-based person image synthesis aims to generate a new image containing a person with a target pose conditioned on a source image containing a person with a specified pose. It is challenging as the target pose is arbitrary and often significantly differs from the specified source pose, which leads to large appearance discrepancy between the source and the target images. This paper presents the Pose Transform Generative Adversarial Network (PoT-GAN) for person image synthesis where the generator explicitly learns the transform between the two poses by manipulating the corresponding multi-scale feature maps. By incorporating the learned pose transform information into the multi-scale feature maps of the source image in a GAN architecture, our method reliably transfers the appearance of the person in the source image to the target pose with no need for any hard-coded spatial information depicting the change of pose. According to both qualitative and quantitative results, the proposed PoT-GAN demonstrates a state-of-the-art performance on three publicly available datasets for person image synthesis.
Wei Zhang 0021, Ran Song 0001, Zhiheng Li 0005, Jun Liu 0036, Xiaolei Li 0003, Shijian Lu
IEEE Trans. Image Process.6
2006 The Application of Robot Formation Approach in the Control of Subway Train
abstract
In order to increase the carrying capacity of subway train, it is necessary to adjust the running of the train for the realization of high operation frequency and safety. After an analysis of the subway train's running under moving block system, a dynamic mathematical model of subway train is constructed. The control among multiple trains is discussed, the method of decreasing the headway and increasing the carrying capacity is discussed by using the event-based technology and formation approach to ensure the safety and to avoid the stop out of station, this approach can reconstruct the system easily and harmonize the cooperation between subsystems. Take account of three trains, the control strategy of the following train is given, the velocity, acceleration of the train can be adjusted according to the safety distance between the successive trains and the distance which the train had traveled
Fei Lu 0005, Mumin Song, Guohui Tian, Xiaolei Li 0003
IROS4