VLDB 2026 Research / reviewers in the wild / expert
Xiaohong Li 0002
dblp:08/2489-2 · also Xiao-Hong Li 0002
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-1187-2042ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HMR-net: Hierarchical multi-scale relation network for visible-infrared person re-identification
Ruixiang Quan, Xiaohong Li 0002, Shuo Zhuang, Xueliang Liu |
Neurocomputing | 2 |
| 2025 | Contrastive learning-based joint pre-training for unsupervised domain adaptive person re-identification
Xiaohong Li 0002, Xuesong Dai, Shuo Zhuang, Meibin Qi |
Multim. Syst. | 2 |
| 2024 | Asymmetric Deformable Spatio-temporal Framework for Infrared Object TrackingabstractThe Infrared Object Tracking (IOT) task aims to locate objects in infrared sequences. Since color and texture information is unavailable in infrared modality, most existing infrared trackers merely rely on capturing spatial contexts from the image to enhance feature representation, where other complementary information is rarely deployed. To fill this gap, we in this article propose a novel Asymmetric Deformable Spatio-Temporal Framework (ADSF) to fully exploit collaborative shape and temporal clues in terms of the objects. Firstly, an asymmetric deformable cross-attention module is designed to extract shape information, which attends to the deformable correlations between distinct frames in an asymmetric manner. Secondly, a spatio-temporal tracking framework is coined to learn the temporal variance trend of the object during the training process and store the template information closest to the tracking frame when testing. Comprehensive experiments demonstrate that ADSF outperforms state-of-the-art methods on three public datasets. Extensive ablation experiments further confirm the effectiveness of each component in ADSF. Furthermore, we conduct generalization validation to demonstrate that the proposed method also achieves performance gains in RGB-based tracking scenarios. Jingjing Wu 0001, Xi Zhou 0004, Xiaohong Li 0002, Hao Liu 0003, Meibin Qi, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Camera Proxy based Contrastive Learning with Hard Sampling for Unsupervised Person Re-identificationabstractBecause of the advantages of dealing with large-scale unlabelled data, unsupervised learning has recently attracted more attention for person re-identification. Particularly, the combination of the unsupervised learning paradigm with contrastive learning shows promising efficiency in network optimization. This work adopts the successful camera-aware contrastive learning approach and further explores its capability on the camera proxy level to improve the data pair consistency. Thus, it is more robust to the camera change, which still challenges the unsupervised person re-identification. This work proposed a Camera Proxy-based Contrastive Learning framework, which explicitly considers inter-camera scenario and intra-camera scenario. Moreover, this work is motivated by the strategy of selecting a hard negative sample in triplet loss learning and further extends it to contrastive learning for both negative and positive pair creation on the camera proxy level. Extensive experiments demonstrate the superiority of the proposed framework over state-of-the-art approaches on purely unsupervised re-identification. Yimin Liu 0001, Meibin Qi, Qiang Wu 0001, Yanfang Yang, Xiaohong Li 0002, Jian Zhang 0002 |
ICME | 5 |
| 2022 | A Local-Global Self-attention Interaction Network for RGB-D Cross-Modal Person Re-identification
Chuanlei Zhu, Xiaohong Li 0002, Meibin Qi, Yimin Liu 0001 |
PRCV (4) | 2 |
| 2021 | Towards accurate estimation for visual object tracking with multi-hierarchy feature aggregation
Jingjing Wu 0001, Meibin Qi, Xiaohong Li 0002 |
Neurocomputing | 4 |
| 2021 | Global-Local Graph Convolutional Network for cross-modality person re-identification
Xiaohong Li 0002, Cuiqun Chen, Meibin Qi, Jingjing Wu 0001 |
Neurocomputing | 2 |
| 2021 | Learning discriminative features with a dual-constrained guided network for video-based person re-identification
Cuiqun Chen, Meibin Qi, Guanghong Huang, Jingjing Wu 0001, Xiaohong Li 0002 |
Multim. Tools Appl. | 6 |
| 2020 | Person Attribute Recognition by Sequence Contextual Relation LearningabstractPerson attribute recognition aims to identify the attribute labels from the pedestrian images. Extracting contextual relation from the images and attributes, including the spatial-semantic relations, the spatial context and the semantic correlation, is beneficial to enhance the discrimination of the features for recognizing the attributes. Thus, this work proposes a sequence contextual relation learning (SCRL) method to capture these relations. It first embeds the images and attributes into sequences in two branches. Then SCRL flexibly learns the contextual relation from the sequences with the parallel attention model structure, which integrates the inter-attention and intra-attention models. The inter-attention module is utilized to extract the spatial-semantic relations, while the intra-attention is designed to gain the spatial context and the semantic correlation. Both attention modules are comprised of several parallel attention units and each unit can obtain the pairwise relations in one subspace. Therefore, they obtain the relations in multiple subspaces, which can improve the comprehensiveness of the relation learning. Additionally, for the sake of better extraction of spatial-semantic relations, this paper employs connectionist temporal classification (CTC) loss which is capable of driving the network to enforce monotonic alignment between the image and attribute. It can also accelerate the convergence of the network by the algorithm in it. Extensive experiments on five public datasets, i.e., Market-1501 attribute, Duke attribute, PETA, RAP and PA-100K datasets, demonstrate the effectiveness of the proposed method. Jingjing Wu 0001, Hao Liu 0003, Meibin Qi, Bo Ren 0002, Xiaohong Li 0002, Yashen Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Robust face detection using local CNN and SVM based on kernel combination
Qinqin Tao, Shu Zhan, Xiaohong Li 0002, Toru Kurihara |
Neurocomputing | 3 |
| 2016 | Face detection using representation learning
Shu Zhan, Qinqin Tao, Xiaohong Li 0002 |
Neurocomputing | 3 |