EDBT 2026 Demo / reviewers in the wild / expert
Xuehu Liu
dblp:287/4752
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-IdentificationabstractLarge-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Identification (VVI-ReID) remains largely unexplored. The primary challenges are narrowing the modality gap and leveraging spatiotemporal information in video sequences. To address the above issues, in this paper, we propose a novel cross-modality feature learning framework named X-ReID for VVI-ReID. Specifically, we first propose a Cross-modality Prototype Collaboration (CPC) to align and integrate features from different modalities, guiding the network to reduce the modality discrepancy. Then, a Multi-granularity Information Interaction (MII) is designed, incorporating short-term interactions from adjacent frames, long-term cross-frame information fusion, and cross-modality feature alignment to enhance temporal modeling and further reduce modality gaps. Finally, by integrating multi-granularity information, a robust sequence-level representation is achieved. Extensive experiments on two large-scale VVI-ReID benchmarks (i.e., HITSZ-VCM and BUPTCampus) demonstrate the superiority of our method over state-of-the-art methods. Chenyang Yu, Xuehu Liu, Huchuan Lu |
AAAI | 2 |
| 2026 | Low-frequency constrained generative adversarial network: An attack framework for remote sensing image scene classification
Huixiao Meng, Yuhang Hong, Xuehu Liu, Zhixi Feng, Zhihao Chang, Shuyuan Yang 0001 |
Neurocomputing | 3 |
| 2026 | Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image FusionabstractMulti-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently, some MMIF methods incorporate frequency domain information to enhance spatial features. However, these methods typically rely on simple serial or parallel spatial-frequency fusion without interaction. In this paper, we propose a novel Interactive Spatial-Frequency Fusion Mamba (ISFM) framework for MMIF. Specifically, we begin with a Modality-Specific Extractor (MSE) to extract features from different modalities. It models long-range dependencies across the image with linear computational complexity. To effectively leverage frequency information, we then propose a Multi-scale Frequency Fusion (MFF). It adaptively integrates low-frequency and high-frequency components across multiple scales, enabling robust representations of frequency features. More importantly, we further propose an Interactive Spatial-Frequency Fusion (ISF). It incorporates frequency features to guide spatial features across modalities, enhancing complementary representations. Extensive experiments are conducted on six MMIF datasets. The experimental results demonstrate that our ISFM can achieve better performances than other state-of-the-art methods. The source code is available at https://github.com/Namn23/ISFM. Long Lv, Xuehu Liu, Tongdan Tang, Feng Tian 0001, Weibing Sun, Huchuan Lu |
IEEE Trans. Image Process. | 4 |
| 2025 | MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptabstractMulti-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary image information from different modalities. Recently, large-scale pre-trained models like CLIP have demonstrated impressive performance in traditional single-modal ReID tasks. However, they remain unexplored for multi-modal object ReID. Furthermore, current multi-modal aggregation methods have obvious limitations in dealing with long sequences from different modalities. To address above issues, we introduce a novel framework called MambaPro for multi-modal object ReID. To be specific, we first employ a Parallel Feed-Forward Adapter (PFA) for adapting CLIP to multi-modal object ReID. Then, we propose the Synergistic Residual Prompt (SRP) to guide the joint learning of multi-modal features. Finally, leveraging Mamba's superior scalability for long sequences, we introduce Mamba Aggregation (MA) to efficiently model interactions between different modalities. As a result, MambaPro could extract more robust features with lower complexity. Extensive experiments on three multi-modal object ReID benchmarks (i.e., RGBNT201, RGBNT100 and MSVR310) validate the effectiveness of our proposed methods. Xuehu Liu, Tianyu Yan, Aihua Zheng, Huchuan Lu |
AAAI | 2 |
| 2025 | CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationabstractPerson Re-IDentification (ReID) aims to identify specific persons from non-overlapping cameras. Recently, some works have suggested using large-scale pre-trained vision-language models like CLIP to boost ReID performance. Unfortunately, existing methods still struggle to address two key issues simultaneously: efficiently transferring the knowledge learned from CLIP and comprehensively extracting the context information from images or videos. To address these issues, we introduce CLIMB-ReID, a pioneering hybrid framework that synergizes the impressive power of CLIP with the remarkable computational efficiency of Mamba. Specifically, we first propose a novel Multi-Memory Collaboration (MMC) strategy to transfer CLIP's knowledge in a parameter-free and prompt-free form. Then, we design a Multi-Temporal Mamba (MTM) to capture multi-granular spatiotemporal information in videos. Finally, with Importance-aware Reorder Mamba (IRM), information from various scales is combined to produce robust sequence features. Extensive experiments show that our proposed method outperforms other state-of-the-art methods on both image and video person ReID benchmarks. Chenyang Yu, Xuehu Liu, Jiawen Zhu 0003, Huchuan Lu |
AAAI | 2 |
| 2025 | Hierarchical Proxy Learning for Cloth-Changing Person Re-IdentificationabstractCloth-Changing person Re-Identification (CC-ReID) depends significantly on learning discriminative features under the cloth-changing scenario. It is quite challenging due to the large intra-person variance and small inter-person variance caused by clothes changing. To address these issues, in this work we propose a Hierarchical Proxy Learning (HPL) framework to extract clothes-irrelevant and person-invariant features. Specifically, we employ person labels as the main proxy. Instead of leveraging clothing labels as sub proxy, we further propose a clustering-based automatic sub-proxy mining scheme. More specifically, we first construct a person-aware Main Proxy Learning (MPL) to improve the separability of different persons. Then, a Sub Proxy Learning (SPL) is constructed to enhance the intra-person compactness. Finally, a Sub-to-Main Proxy Learning (S2MPL) is proposed to promote the cooperation between the main proxies and sub proxies. In addition, to weed out the negative effect of clothes, we propose a Sample Balance and Diversity (SBD) module, which balances the number of sub proxies in a mini-batch and utilizes semantic guidance to enrich the diversity of clothes, simultaneously. Extensive experiments on two public CC-ReID datasets demonstrate the superiority of our proposed method over most state-of-the-art methods. Chenyang Yu, Xuehu Liu, Ju Dai, Huchuan Lu |
ICASSP | 2 |
| 2025 | AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge ResultsabstractPerson re-identification (ReID) across aerial and ground vantage points has become crucial for large-scale surveillance and public safety applications. Although significant progress has been made in ground-only scenarios, bridging the aerial-ground domain gap remains a formidable challenge due to extreme viewpoint differences, scale variations, and occlusions. Building upon the achievements of the AG-ReID 2023 Challenge, this paper introduces the AG-VPReID 2025 Challenge—the first large-scale video-based competition focused on high-altitude (80–120 m) aerial-ground person ReID. Constructed on the new AG-VPReID dataset with 3,027 identities, over 13,500 tracklets, and approximately 3.7 million frames captured from UAVs, CCTV, and wearable cameras, the challenge featured four international teams. These teams developed solutions ranging from multi-stream architectures to transformer-based temporal reasoning and physics-informed modeling. The leading approach, X-TFCLIP from UAM, attained 72.28% Rank-1 accuracy in the aerial-to-ground ReID setting and 70.77% in the ground-to-aerial ReID setting, surpassing existing baselines while highlighting the dataset’s complexity. For additional details, please refer to the official website at https://agvpreid25.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Tamás Endrei, Ivan DeAndres-Tame, Ruben Tolosana, Rubén Vera-Rodríguez, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Zijing Gong, Xuehu Liu, Md. Rashidunnabi, Hugo Proença 0001, Kailash A. Hambarde, Saeid Rezaei |
IJCB | 17 |
| 2025 | Spatial Invariant Hash Based on Self-Attention Mechanism for Remote Sensing Ship Image RetrievalabstractIn the task of remote sensing ship image retrieval, due to the significant Angle changes and multi-scale characteristics of ships in the image, it is difficult to extract advanced features and the efficiency of feature descriptors is low. Therefore, this paper proposes a spatial invariant hash based on self-attention mechanism for remote sensing ship image retrieval (SIHS). The algorithm consists of two core modules: First, a module based on spatial invariance is designed, which uses deep convolutional neural network to extract continuous real-valued descriptors, and introduces the spatial transformation attention mechanism, and enhances the adaptation ability and learning efficiency of the model to the spatial invariance features through self-learning affine transformation and attention calculation; Secondly, a module based on self-attention hashing is proposed, which improves the efficiency of image representation by multi-scale image embedding, optimizes the attention regularization in the visual encoder, and effectively solves the problem of quantization loss in hash mapping. The experimental results show that the retrieval performance of SIHS algorithm on GGWS, DSCR, FGSC-23 and FGSCR-42 datasets is superior to the existing methods based on deep features. Fuwei Huang, Yaxiong Chen, Kai Yan 0001, Yin Ye, Xuehu Liu, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Unity Is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-IdentificationabstractPerson Re-identification (ReID) aims to retrieve the specific person across non-overlapping cameras, which greatly helps intelligent transportation systems. As we all know, Convolutional Neural Networks (CNNs) and Transformers have the unique strengths to extract local and global features, respectively. Considering this fact, we focus on the mutual fusion between them to learn more comprehensive representations for persons. In particular, we utilize the complementary integration of deep features from different model structures. We propose a novel fusion framework called FusionReID to unify the strengths of CNNs and Transformers for image-based person ReID. More specifically, we first deploy a Dual-branch Feature Extraction (DFE) to extract features through CNNs and Transformers from a single image. Moreover, we design a novel Dual-attention Mutual Fusion (DMF) to achieve sufficient feature fusions. The DMF comprises Local Refinement Units (LRU) and Heterogenous Transmission Modules (HTM). LRU utilizes depth-separable convolutions to align deep features in channel dimensions and spatial sizes. HTM consists of a Shared Encoding Unit (SEU) and two Mutual Fusion Units (MFU). Through the continuous stacking of HTM, deep features after LRU are repeatedly utilized to generate more discriminative features. Extensive experiments on three public ReID benchmarks demonstrate that our method can attain superior performances than most state-of-the-arts. The source code is available athttps://github.com/924973292/FusionReID. Xuehu Liu, Zhengzheng Tu, Huchuan Lu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | TOP-ReID: Multi-Spectral Object Re-identification with Token PermutationabstractMulti-spectral object Re-identification (ReID) aims to retrieve specific objects by leveraging complementary information from different image spectra. It delivers great advantages over traditional single-spectral ReID in complex visual environment. However, the significant distribution gap among different image spectra poses great challenges for effective multi-spectral feature representations. In addition, most of current Transformer-based ReID methods only utilize the global feature of class tokens to achieve the holistic retrieval, ignoring the local discriminative ones. To address the above issues, we step further to utilize all the tokens of Transformers and propose a cyclic token permutation framework for multi-spectral object ReID, dubbled TOP-ReID. More specifically, we first deploy a multi-stream deep network based on vision Transformers to preserve distinct information from different image spectra. Then, we propose a Token Permutation Module (TPM) for cyclic multi-spectral feature aggregation. It not only facilitates the spatial feature alignment across different image spectra, but also allows the class token of each spectrum to perceive the local details of other spectra. Meanwhile, we propose a Complementary Reconstruction Module (CRM), which introduces dense token-level reconstruction constraints to reduce the distribution gap across different image spectra. With the above modules, our proposed framework can generate more discriminative multi-spectral features for robust object ReID. Extensive experiments on three ReID benchmarks (i.e., RGBNT201, RGBNT100 and MSVR310) verify the effectiveness of our methods. The code is available at https://github.com/924973292/TOP-ReID. Xuehu Liu, Hu Lu, Zhengzheng Tu, Huchuan Lu |
AAAI | 2 |
| 2024 | TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identificationabstractLarge-scale language-image pre-trained models (e.g., CLIP) have shown superior performances on many cross-modal retrieval tasks. However, the problem of transferring the knowledge learned from such models to video-based person re-identification (ReID) has barely been explored. In addition, there is a lack of decent text descriptions in current ReID benchmarks. To address these issues, in this work, we propose a novel one-stage text-free CLIP-based learning framework named TF-CLIP for video-based person ReID. More specifically, we extract the identity-specific sequence feature as the CLIP-Memory to replace the text feature. Meanwhile, we design a Sequence-Specific Prompt (SSP) module to update the CLIP-Memory online. To capture temporal information, we further propose a Temporal Memory Diffusion (TMD) module, which consists of two key components: Temporal Memory Construction (TMC) and Memory Diffusion (MD). Technically, TMC allows the frame-level memories in a sequence to communicate with each other, and to extract temporal information based on the relations within the sequence. MD further diffuses the temporal memories to each token in the original features to obtain more robust sequence features. Extensive experiments demonstrate that our proposed method shows much better results than other state-of-the-art methods on MARS, LS-VID and iLIDS-VID. Chenyang Yu, Xuehu Liu, Yingquan Wang, Huchuan Lu |
AAAI | 2 |
| 2024 | D3R-Net: Denoising Diffusion-Based Defense Restore Network for Adversarial Defense in Remote Sensing Scene ClassificationabstractDeep learning models (algorithms) have demonstrated their superior performance in interpreting Earth science and remote sensing data. However, adversarial examples generated with perturbations imperceptible to humans could render deep learning algorithms ineffective. This significant vulnerability of deep learning models, thus, inspires the exploration of defense methods resistible to adversarial examples. Although numerous countermeasures against adversarial examples have been proposed, the design of a universally applicable defense method across multiple scenarios still remains to be explored. In this study, we propose an effective denoising diffusion-based defense restore network (D3R-Net) based on the denoising diffusion model from the perspective of adversarial restoration, which transforms the adversarial examples into clean samples. Utilizing a highly effective denoising diffusion probabilistic model (DDPM), our D3R-Net transforms input adversarial examples into a state of noise, where diverse forms of adversarial noise transition into Gaussian noise. Subsequently, it captures semantic information through a series of iterative denoising steps. The pixel distribution of adversarial examples is restored in the proposed network to match the original distribution, enabling the classifier to identify adversarial examples correctly. Furthermore, we introduce a combined filtering module to preserve the semantic information of the original image, thereby further enhancing the defensive performance. Instead of modifying the model structure or excluding suspected samples, the proposed method restores the adversarial examples, making it simple yet effective and applicable to a broader range of scenarios. Extensive experiments are conducted on four benchmark datasets, and the results demonstrate that D3R-Net has significant defense capabilities against known and unknown attacks. Our source code is available athttps://github.com/SIM-xidian/D3R-Net. Xuehu Liu, Zhixi Feng, Yue Ma 0008, Shuyuan Yang 0001, Zhihao Chang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Video Is Worth Three Views: Trigeminal Transformers for Video-Based Person Re-IdentificationabstractVideo-based person Re-Identification (Re-ID) is a hot research topic in intelligent transportation systems, which aims to retrieve video sequences of the same person under non-overlapping surveillance cameras. Compared with static images, video sequences contain more visual information from multiple views, such as spatial and temporal views. However, previous Re-ID methods usually focus on single limited views, lacking diverse observations from different views. To capture richer perceptions and extract more comprehensive representations, we propose a novel learning framework namedTrigeminal Transformers (TMT)to tackle video-based person Re-ID. More specifically, we first design aView-wise Projector (VP)to jointly transform raw videos from spatial, temporal and spatial-temporal views. In addition, inspired by the great success of Vision Transformers (ViT), we introduce the Transformer structure for information enhancement and aggregation. In our work, threeSelf-view Transformers (ST)are proposed to exploit the relationships of local features for information enhancement in spatial, temporal and spatial-temporal. Moreover, aCross-view Transformer (CT)is proposed to aggregate the multi-view features for comprehensive representations. Experimental results indicate that our approach can obtain better performance than some other state-of-the-art approaches on four public Re-ID benchmarks. Xuehu Liu, Chenyang Yu, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Deeply Coupled Convolution-Transformer With Spatial-Temporal Complementary Learning for Video-Based Person Re-IdentificationabstractAdvanced deep convolutional neural networks (CNNs) have shown great success in video-based person re-identification (Re-ID). However, they usually focus on the most obvious regions of persons with a limited global representation ability. Recently, it witnesses that Transformers explore the interpatch relationships with global observations for performance improvements. In this work, we take both the sides and propose a novel spatial-temporal complementary learning framework named deeply coupled convolution-transformer (DCCT) for high-performance video-based person Re-ID. First, we couple CNNs and Transformers to extract two kinds of visual features and experimentally verify their complementarity. Furthermore, in spatial, we propose a complementary content attention (CCA) to take advantages of the coupled structure and guide independent features for spatial complementary learning. In temporal, a hierarchical temporal aggregation (HTA) is proposed to progressively capture the interframe dependencies and encode temporal information. Besides, a gated attention (GA) is used to deliver aggregated temporal information into the CNN and Transformer branches for temporal complementary learning. Finally, we introduce a self-distillation training strategy to transfer the superior spatial-temporal knowledge to backbone networks for higher accuracy and more efficiency. In this way, two kinds of typical features from same videos are integrated mechanically for more informative representations. Extensive experiments on four public Re-ID benchmarks demonstrate that our framework could attain better performances than most state-of-the-art methods. Xuehu Liu, Chenyang Yu, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Video-Based Person Re-Identification with Long Short-Term Representation Learning
Xuehu Liu, Huchuan Lu |
ICIG (1) | 1 |
| 2023 | Hierarchical Feature Fusion and Selection for Hyperspectral Image ClassificationabstractMost existing classification methods design complicated and large deep neural network (DNN) model to deal with the ubiquitous spectral variability and nonlinearity of hyperspectral images (HSIs). However, their application is blocked by limited training samples and considerable computational costs in real scenes. To solve these problems, we propose a simple spectral hierarchical feature fusion and selection network (HFFSNet). Specifically, we apply 1-D grouped convolution for dimensionality reduction and multilevel feature extraction, then the multilevel features are fused to assist the adaptive feature selection of different layer features via the soft attention mechanism, and finally the selected features are fused to further enhance the feature representation. Extensive experimental results on three hyperspectral datasets demonstrate the effectiveness of the proposed network. Zhixi Feng, Xuehu Liu, Shuyuan Yang 0001, Kai Zhang 0010, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Watching You: Global-Guided Reciprocal Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to automatically retrieve video sequences of the same person under non-overlapping cameras. To achieve this goal, it is the key to fully utilize abundant spatial and temporal cues in videos. Existing methods usually focus on the most conspicuous image regions, thus they may easily miss out fine-grained clues due to the person varieties in image sequences. To address above issues, in this paper, we propose a novel Global-guided Reciprocal Learning (GRL) framework for video-based person Re-ID. Specifically, we first propose a Global-guided Correlation Estimation (GCE) to generate feature correlation maps of local features and global features, which help to localize the high- and low-correlation regions for identifying the same person. After that, the discriminative features are disentangled into high-correlation features and low-correlation features under the guidance of the global representations. Moreover, a novel Temporal Reciprocal Learning (TRL) mechanism is designed to sequentially enhance the high-correlation semantic information and accumulate the low-correlation sub-critical clues. Extensive experiments are conducted on three public benchmarks. The experimental results indicate that our approach can achieve better performance than other state-of-the-art approaches. The code is released at https://github.com/flysnowtiger/GRL. Xuehu Liu, Chenyang Yu, Huchuan Lu, Xiaoyun Yang |
CVPR | 1 |