VLDB 2026 Research / reviewers in the wild / expert
Haijun Xiong
dblp:41/553
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0856-8250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gait Recognition via Collaborating Discriminative and Generative Diffusion ModelsabstractGait recognition offers a non-intrusive biometric solution by identifying individuals through their walking patterns. Although discriminative models have achieved notable success in this domain, the full potential of generative models remains largely unexplored. In this paper, we introduce CoD², a novel framework that combines the data distribution modeling capabilities of diffusion models with the semantic representation learning strengths of discriminative models to extract robust gait features. We propose a Multi-level Conditional Control strategy that integrates both high-level identity-aware semantic conditions and low-level visual details. Specifically, the high-level condition, extracted by the discriminative extractor, guides the generation of identity-consistent gait sequences, while low-level visual details, such as appearance and motion, are preserved to enhance consistency. Moreover, the generated sequences facilitate the discriminative extractor's learning, enabling it to capture more comprehensive high-level semantic features. Extensive experiments on four datasets (SUSTech1K, CCPG, GREW, and Gait3D) demonstrate that CoD² achieves state-of-the-art performance and can be seamlessly integrated with existing discriminative methods, yielding consistent improvements. Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001 |
AAAI | 1 |
| 2026 | Fourier-based adaptive counterfactual intervention for object re-identification
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001 |
Neural Networks | 1 |
| 2025 | Collaborative Spatial and Channel Attention for Structural Vibration-based Gait RecognitionabstractStructural vibration-based gait recognition aims to identify pedestrians through their unique vibration patterns and has gained significant attention for its non-invasive nature. However, existing methods primarily rely on manually crafted features and convolutional neural networks (CNNs) to extract local details, without considering global gait features in vibration signals. To address this limitation, we propose CSCA, a novel framework designed to capture discriminative global gait features through collaborative spatial and channel attention. Specifically, CSCA consists of two key components: Mamba-inspired Linear Channel Attention (MLCA) and Strip Pooling-guided Spatial Attention (SPSA). MLCA integrates Mamba with linear attention to enhance discriminative attention in the channel dimension by capturing global gait features. Additionally, SPSA integrates axial strip pooling with multikernel depth-wise convolutions to extract multi-level spatial semantic features. Experimental results on the VIBEID dataset demonstrate that CSCA achieves state-of-the-art performance across diverse environmental conditions, with identity recognition accuracy exceeding 96%. Code is available at https://github.com/xiaomush/CSCA. Junhao Lu, Haijun Xiong, Ziyu Lin, Bin Feng 0001 |
IJCB | 2 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 26 |
| 2025 | CGTGait: Collaborative Graph and Transformer for Gait Emotion RecognitionabstractSkeleton-based gait emotion recognition has received significant attention due to its wide-ranging applications. However, existing methods primarily focus on extracting spatial and local temporal motion information, failing to capture long-range temporal representations. In this paper, we propose CGTGait, a novel framework that collaboratively integrates graph convolution and transformers to extract discriminative spatiotemporal features for gait emotion recognition. Specifically, CGTGait consists of multiple CGT blocks, where each block employs graph convolution to capture frame-level spatial topology and the transformer to model global temporal dependencies. Additionally, we introduce a Bidirectional Cross-Stream Fusion (BCSF) module to effectively aggregate posture and motion spatiotemporal features, facilitating the exchange of complementary information between the two streams. We evaluate our method on two widely used datasets, Emotion-Gait and ELMD, demonstrating that our CGTGait achieves state-of-the-art or at least competitive performance while reducing computational complexity by approximately 82.2% (only requiring 0.34G FLOPs) during testing. Code is available at https://github.com/githubzjj1/CGTGait. Haijun Xiong, Junhao Lu, Ziyu Lin, Bin Feng 0001 |
IJCB | 2 |
| 2025 | STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-To-4D Gaussian SplattingabstractText-To-4D generation is rapidly developing and widely applied in various scenarios. However, existing methods often fail to incorporate adequate spatio-temporal modeling and prompt alignment within a unified framework, resulting in temporal inconsistencies, geometric distortions, or low-quality 4D content that deviates from the provided texts. Therefore, we propose STP4D, a novel approach that aims to integrate comprehensive spatio-temporal-prompt consistency modeling for high-quality text-to-4D generation. Specifically, STP4D employs three carefully designed modules: Time-Varying Prompt Embedding, Geometric Information Enhancement, and Temporal Extension Deformation, which collaborate to accomplish this goal. Furthermore, STP4D is among the first methods to exploit the Diffusion model to generate 4D Gaussians, combining the fine-grained modeling capabilities and the real-time rendering process of 4DGS with the rapid inference speed of the Diffusion model. Extensive experiments demonstrate that STP4D excels in generating high-fidelity 4D content with exceptional efficiency (approximately 4.6s per asset), surpassing existing methods in both quality and speed. Yunze Deng, Haijun Xiong, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
ICME | 2 |
| 2025 | MambaGait: Gait recognition approach combining explicit representation and implicit state space model
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001 |
Image Vis. Comput. | 1 |
| 2024 | Causality-Inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
Haijun Xiong, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
ECCV (54) | 1 |
| 2024 | Licaf: Lidar-Camera Asymmetric Fusion For Gait RecognitionabstractGait recognition is a biometric technology that identifies individuals by using walking patterns. Due to the significant achievements of multimodal fusion in gait recognition, we consider employing LiDAR-camera fusion to obtain robust gait representations. However, existing methods often overlook intrinsic characteristics of modalities, and lack fine-grained fusion and temporal modeling. In this paper, we introduce a novel modality-sensitive network LiCAF for LiDAR-camera fusion, which employs an asymmetric modeling strategy. Specifically, we propose Asymmetric Cross-modal Channel Attention (ACCA) and Interlaced Crossmodal Temporal Modeling (ICTM) for cross-modal valuable channel information selection and powerful temporal modeling. Our method achieves state-of-the-art performance ($93.9 \%$ in Rank-1 and $98.8 \%$ in Rank-5) on the SUSTech1K dataset, demonstrating its effectiveness. Yunze Deng, Haijun Xiong, Bin Feng 0001 |
ICIP | 2 |
| 2024 | Gaitgs: Temporal Feature Learning in Granularity And Span Dimension for Gait RecognitionabstractGait recognition, a growing field in biological recognition technology, utilizes distinct walking patterns for accurate individual identification. However, existing methods lack the incorporation of temporal information. To reach the full potential of gait recognition, we advocate for the consideration of temporal features at varying granularities and spans. This paper introduces a novel framework, GaitGS, which aggregates temporal features simultaneously in both granularity and span dimensions. Specifically, the Multi-Granularity Feature Extractor (MGFE) is designed to capture micro-motion and macro-motion information at fine and coarse levels respectively, while the Multi-Span Feature Extractor (MSFE) generates local and global temporal representations. Through extensive experiments on two datasets, our method demonstrates state-of-the-art performance, achieving Rank-1 accuracy of 98.2%, 96.5%, and 89.7% on CASIA-B under different conditions, and 97.6% on OU-MVLP. The source code will be available at https://github.com/Haijun-Xiong/GaitGS. Haijun Xiong, Yunze Deng, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
ICIP | 1 |
| 2021 | HID 2021: Competition on Human Identification at a Distance 2021abstractThe Competition on Human Identification at a Distance 2021 (HID 2021) is to promote the research in human identification at a distance and to provide a benchmark to evaluate different methods. HID 2021 is the second follow-up from the first one, HID 2020. The dataset size and the evaluation protocal are the same with the previous competition, but the data in the test set has been changed. The paper firstly introduces the dataset and the evaluation protocol, then describes the methods from the top teams and their results. The methods show how to achieve state-of-the-art performance on gait recognition. The results in HID 2021 are better than those in HID 2020. From the comparisons and analysis, some useful conclusions can be drawn. We hope more improvements can be achieved by better followup competitions. Shiqi Yu 0001, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Edel B. García Reyes, Feng Zheng 0001, Md. Atiqur Rahman Ahad, Beibei Lin, Haijun Xiong, Binyuan Huang |
IJCB | 10 |