Chunshui Cao

dblp:176/1432 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0001-6634-1682ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 15 since 2021Security and privacy · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MPANet: Motion Pattern Aggregation Network for Gait Recognition
Chunshui Cao, Guoren Wang, Dezhi Zheng
Int. J. Comput. Vis.3
2025 Bridging Gait Recognition and Large Language Models Sequence Modeling
abstract
Gait sequences exhibit sequential structures and contextual relationships similar to those in natural language, where each element—whether a word or a gait step—is connected to its predecessors and successors. This similarity enables the transformation of gait sequences into "texts" containing identity-related information. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized for gait sequence modeling to enhance gait recognition performance. Leveraging these insights, we make a pioneering effort to apply LLMs to gait recognition, which we refer to as GaitLLM. Specifically, we propose the Gait-to-Language (G2L) module, which converts gait sequences into a textual format suitable for LLMs, and the Language-to-Gait (L2G) module, which maps the LLM’s output back to the gait feature space, thereby bridging the gap between LLM outputs and gait recognition. Notably, GaitLLM leverages the powerful modeling capabilities of LLMs without relying on complex architectural designs, improving gait recognition performance with only a small number of trainable parameters. Our method achieves state-of-the-art results on four popular gait datasets—SUSTech1K, CCPG, Gait3D, and GREW—demonstrating the effectiveness of applying LLMs in this domain. This work highlights the potential of LLMs to significantly enhance gait recognition, paving the way for future research and practical applications.
Shaopeng Yang, Jilong Wang 0010, Saihui Hou, Xu Liu 0008, Chunshui Cao, Liang Wang 0001, Yongzhen Huang
CVPR5
2025 Gait: Exploring X Modality for Generalized Gait Recognition
Zengbin Wang, Saihui Hou, Junjie Li 0002, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Siye Wang, Man Zhang 0005
ICCV5
2025 Vocabulary-Guided Gait Recognition
abstract
What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points. However, the considerations remain vague for humans to comprehend truly. In this work, we introduce a novel paradigm Vocabulary-Guided Gait Recognition, dubbed Gait-World, which attempts to explore gait concepts through human vocabularies with Vision-Language Models (VLMs). Despite VLMs have achieved the remarkable progress in various vision tasks, the cognitive capability regarding gait modalities remains limited. The success element in Gait-World is the proper vocabulary prompt where this paradigm carefully selects gait cycle actions as Vocabulary Base, bridging the gait and vocabulary feature spaces and further promoting human understanding for the gait. How to extract gait features? Although previous gait networks have made significant progress, learning solely from gait modalities on limited gait databases makes it difficult to learn robust gait features for practicality. Therefore, we propose the first Gait-World model, dubbed $\alpha$-Gait, which guides the gait network learning with universal vocabulary knowledge from VLMs. However, due to the heterogeneity of the modalities, directly integrating vocabulary and gait features is highly challenging as they reside in different embedding spaces. To address the issues, $\alpha$-Gait designs Vocabulary Relation Mapper and Gait Fine-grained Detector to map and establish vocabulary relations in the gait space for detecting corresponding gait features. Extensive experiments on CASIA-B, CCPG, SUSTech1K, Gait3D and GREW reveal the potential value and research directions of vocabulary information from VLMs in the gait field.
Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
NeurIPS3
2025 Edge-Oriented Adversarial Attack for Deep Gait Recognition
Saihui Hou, Zengbin Wang, Man Zhang 0005, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
Int. J. Comput. Vis.4
2025 From FastPoseGait to GPGait++: Bridging the Past and Future for Pose-Based Gait Recognition
abstract
Recent studies on pose-based gait recognition have underscored the potential of utilizing such fundamental data to achieve superior outcomes. Nonetheless, the development of current pose-based methods faces significant obstacles due to several critical issues: (1) Misaligned Settings, which results in a lack of thorough and unbiased comparative analysis. (2) Inferior Performance, which causes diminished focus on pose-based gait representations. (3) Limited Generalization, which hinders the effective application in real-world scenarios. Focused on tackling the aforementioned challenges, our study introduces a comprehensive benchmark and a versatile approach to bridge the past and future for pose-based gait recognition. First, we revisit previous pose-based methods and make great efforts to establish a unified framework, FastPoseGait, aiming at a fair and comprehensive comparison investigation with consistent experimental settings and a more stable training process. Then, within this framework, we propose GPGait++, a generalized pose-based gait recognition method featuring a human-oriented input and part-aware modeling, intended to enhance the generalization ability and discriminative power across diverse environments and camera viewpoints. Experiments on six public gait recognition datasets reveal that our unified framework significantly enhances the performance of previous approaches, and GPGait++ exhibits state-of-the-art cross-domain capabilities compared to existing pose-based methods, marking a significant advancement in the field of pose-based gait recognition.
Shibei Meng, Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Multimodal Mutual Learning for Unsupervised Gait Recognition
abstract
The primary challenge in unsupervised gait recognition lies in generating meaningful and diverse supervisory signals to guide representation learning. The effectiveness of such methods largely depends on the richness of the supervisory signals. Unlike previous methods that construct supervisory signals solely from a single modality, we propose a novel framework, named Multimodal Mutual Learning (M3L), that leverages the identity consistency and complementary nature of both silhouette and skeleton modalities to generate richer and more informative supervisory signals. To fully leverage the richer supervisory signals, M3L encourages mutual prediction between the silhouette and skeleton modalities, guiding the network toward modality-invariant representations. However, mutual prediction alone is hindered by the inherent modality gap, so we introduce a Multimodal Collaborative Module to explicitly bridge this gap and promote cross-modal knowledge transfer. Moreover, to make the framework practical when only one modality is available at inference, we introduce a Multimodal Disentanglement Module. Multimodal Disentanglement Module decouples the two branches and distills a shared representation, preserving the gains of multimodal training while allowing the model to maintain robust performance under single-modality conditions. Extensive experiments on four widely used gait datasets—Gait3D, GREW, CASIA-B, and SUSTech1K—demonstrate the effectiveness of our approach and highlight its potential to advance unsupervised gait recognition.
Shaopeng Yang, Saihui Hou, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.4
2024 QAGait: Revisit Gait Recognition from a Quality Perspective
abstract
Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research. However, as gait recognition rapidly advances from in-the-lab to in-the-wild scenarios, various conditions raise significant challenges for silhouette modality, including 1) unidentifiable low-quality silhouettes (abnormal segmentation, severe occlusion, or even non-human shape), and 2) identifiable but challenging silhouettes (background noise, non-standard posture, slight occlusion). To address these challenges, we revisit gait recognition pipeline and approach gait recognition from a quality perspective, namely QAGait. Specifically, we propose a series of cost-effective quality assessment strategies, including Maxmial Connect Area and Template Match to eliminate background noises and unidentifiable silhouettes, Alignment strategy to handle non-standard postures. We also propose two quality-aware loss functions to integrate silhouette quality into optimization within the embedding space. Extensive experiments demonstrate our QAGait can guarantee both gait reliability and performance enhancement. Furthermore, our quality assessment strategies can seamlessly integrate with existing gait datasets, showcasing our superiority. Code is available at https://github.com/wzb-bupt/QAGait.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu
AAAI5
2024 Learning Visual Prompt for Gait Recognition
abstract
Gait, a prevalent and complex form of human motion, plays a significant role in the field of long-range pedestrian retrieval due to the unique characteristics inherent in individual motion patterns. However, gait recognition in real-world scenarios is challenging due to the limitations of capturing comprehensive cross-viewing and crossclothing data. Additionally, distractors such as occlusions, directional changes, and lingering movements further complicate the problem. The widespread application of deep learning techniques has led to the development of various potential gait recognition methods. However, these methods utilize convolutional networks to extract shared information across different views and attire conditions. Once trained, the parameters and non-linear function become constrained to fixed patterns, limiting their adaptability to various distractors in real-world scenarios. In this paper, we present a unified gait recognition framework to extract global motion patterns and develop a novel dynamic transformer to generate representative gait features. Specifically, we develop a trainable part-based prompt pool with numerous key-value pairs that can dynamically select prompt templates to incorporate into the gait sequence, thereby providing task-relevant shared knowledge information. Furthermore, we specifically design dynamic attention to extract robust motion patterns and address the length generalization issue. Extensive experiments on four widely recognized gait datasets, i.e., Gait3D, GREW, OUMVLP, and CASIA-B, reveal that the proposed method yields substantial improvements compared to current state-of-the-art approaches.
Ying Fu 0001, Chunshui Cao, Saihui Hou, Yongzhen Huang, Dezhi Zheng
CVPR3
2024 Cut Out the Middleman: Revisiting Pose-Based Gait Recognition
Saihui Hou, Shibei Meng, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
ECCV (31)5
2024 Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective
Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu 0008, Zhiqiang He 0002, Yongzhen Huang
ECCV (6)4
2024 Free Lunch for Gait Recognition: A Novel Relation Descriptor
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001
ECCV (38)4
2024 Cloth-Imbalanced Gait Recognition via Hallucination
abstract
The study in the gait field has rarely paid attention to the class-imbalanced learning, while the realistic data always exhibits an imbalanced distribution. The main reason lies in the difficulty of collecting the cross-clothes sequences, since the collection is usually aided by person re-identification and it is more likely to obtain the sequences for a subject wearing the same clothes. In this work, we formulate a new problem to tackle the task-specific cloth-imbalanced issue, dubbed as Cloth-Imbalanced Gait Recognition, and the training data consists of two parts denoted as head set and tail set. The sequences for a subject in head set cover the cross-clothes variation which is scarce in tail set to mimic the collection difficulty. Along with the problem formulation, we design a new method to deal with the inherent challenges, called Cross-Clothes Hallucination or CCH for short. Our method is inspired by the observation that certain directions in deep feature space correspond to meaningful semantic transformations, and it tries to generate the cross-clothes sequences for tail set referring to the cloth-changing transformation in head set. To evaluate CCH, we build two cloth-imbalanced benchmarks based on the widely-used CASIA-B and Outdoor-Gait. Extensive experiments demonstrate that CCH brings significant improvements over the baselines.
Saihui Hou, Panjian Huang, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Adaptive Knowledge Transfer for Weak-Shot Gait Recognition
abstract
Most works for cloth-changing gait recognition assume that the sequences of different clothes are accessible for each subject in the training set, which, however, is almost impossible for real applications. In practice, the collection of gait sequences is usually aided by person re-identification which is more likely to cluster the cloth-consistent sequences for a subject, and it is laborious to merge the cloth-changing clusters with the same identity. As a result, the training set is usually comprised of two subsets, i.e., a fully-annotated base set where the cloth-changing sequences are available for each subject, and a weakly-annotated wild set where the sequences of different clothes for a subject are assigned diverse labels. In this work, we formulate a problem named Weak-Shot Gait Recognition which seeks to learn discriminative features from the mixture of base set and wild set. Furthermore, we propose an effective method called Adaptive Knowledge Transfer to deal with the weak-shot issue. In particular, we define the knowledge as the ability to judge whether two sequences come from the same subject or not, and we take an adaptive way to mine the useful information from wild set. For the experimental study, we build three weak-shot benchmarks based on CASIA-B, Outdoor-Gait, and CASIA-E respectively. Extensive experiments show that our method can bring consistent improvement. For example, under the cloth-changing condition of on the weak-shot CASIA-B, our method exceeds a naïve baseline by 7.10%.
Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.3
2024 Integral Pose Learning via Appearance Transfer for Gait Recognition
abstract
Gait recognition plays an important role in video surveillance and security by identifying humans based on their unique walking patterns. The existing gait recognition methods have achieved competitive accuracy with shape and motion patterns under limited-covariate conditions. However, when extreme appearance changes distort discriminative features, gait recognition yields unsatisfactory results under cross-covariate conditions. In this work, we first indicate that the integral pose in each silhouette maintains an appearance-unrelated discriminative identity. However, the monotonous appearance variables in a gait database cause gait models to have difficulty extracting integral poses. Therefore, we propose an Appearance-transferable Disentangling and Generative Network (GaitApp) to generate gait silhouettes with rich appearances and invariant poses. Specifically, GaitApp leverages multi-branch cooperation to disentangle pose features and appearance features, and transfers the appearance information from one subject to another. By simulating a person constantly changing appearances under limited-covariate conditions, downstream models enable to extract integral discriminative pose features. Extensive experiments demonstrate that our method allows representative gait models to stand at a new altitude, further promoting the exploration to cross-covariate gait recognition. All the code is available at https://github.com/Hpjhpjhs/GaitApp.git.
Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Xuecai Hu, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.3
2024 Gait Attribute Recognition: A New Benchmark for Learning Richer Attributes From Human Gait Patterns
abstract
Compared to gait recognition, Gait Attribute Recognition (GAR) is a seldom-investigated problem. However, since gait attribute recognition can provide richer and finer semantic descriptions, it is an indispensable part of building intelligent gait analysis systems. Nonetheless, the types of attributes considered in the existing datasets are very limited. This paper contributes a new benchmark dataset for gait attribute recognition named Multi-Attribute Gait (MA-Gait). Our MA-Gait contains 95 subjects recorded from 12 camera views, resulting in more than 13000 sequences, with 16 attributes labeled, including six attributes that have never been considered in the literature. Moreover, we propose a Multi-Scale Motion Encoder (MSME) to extract robust motion features, and an Attribute-Guided Feature Selection Module (AGFSM) to adaptively capture the most discriminative attribute features from static appearance features and dynamic motion features for different attributes. Our method achieves the best GAR accuracy on the new dataset. Comprehensive experiments show the effectiveness of the proposed method through both quantitative and qualitative evaluations.
Xu Song, Saihui Hou, Yan Huang 0023, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Caifeng Shan
IEEE Trans. Inf. Forensics Secur.4
2024 GaitParsing: Human Semantic Parsing for Gait Recognition
abstract
Gait recognition is a soft biotechnology to identify pedestrians observed from different camera views based on specific walking patterns. However, various dressing and wearing conditions bring great challenges to realistic gait recognition. Most existing methods take holistic gait silhouette as input and focus on local areas through horizontal strip division or attention map. We consider that this processing may contain mixed or incomplete information about multiple body parts so that gait information is misused or underutilized. In this paper, we propose a parsing-guided framework for gait recognition, namedGaitParsing, which explores human semantic parsing to dissect human body into a set of specific and complete body parts. Correspondingly, a simple yet effective dual-branch feature extraction network is adopted to process holistic gait and distinct body parts. To maximize the use of highly discriminated gait frames, we propose a self-occlusion frame assessment to measure the self-occlusion in a gait sequence. Since there is no human parsing modality in current gait datasets, we further develop a general human parsing pipeline specifically tailored for gait datasets. This single training enables widespread application across various gait datasets. Extensive experiments with ablation analyses demonstrate competitive performance even in the most challenging conditions, e.g., Cloth-Changing (CC+5.9%). Especially, It is gratifying to see that our model can be easily applied to existing methods and significantly outperform the original architecture, even without much modification.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Multim.5
2024 Deep Learning Based Occluded Person Re-Identification: A Survey
abstract
Occluded person re-identification (Re-ID) focuses on addressing the occlusion problem when retrieving the person of interest across non-overlapping cameras. With the increasing demand for intelligent video surveillance and the application of person Re-ID technology, the real-world occlusion problem draws considerable interest from researchers. Although a large number of occluded person Re-ID methods have been proposed, there are few surveys that focus on occlusion. To fill this gap and help boost future research, this article provides a systematic survey of occluded person Re-ID. In this work, we review recent deep learning based occluded person Re-ID research. First, we summarize the main issues caused by occlusion as four groups: position misalignment, scale misalignment, noisy information, and missing information. Second, we categorize existing methods into six solution groups: matching, image transformation, multi-scale features, attention mechanism, auxiliary information, and contextual recovery. We also discuss the characteristics of each approach, as well as the issues they address. Furthermore, we present the performance comparison of recent occluded person Re-ID methods on four public datasets: Partial-ReID, Partial-iLIDS, Occluded-ReID, and Occluded-DukeMTMC. We conclude the study with thoughts on promising future research directions.
Yunjie Peng, Jinlin Wu, Boqiang Xu, Chunshui Cao, Xu Liu 0008, Zhenan Sun, Zhiqiang He 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2023 An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions
abstract
The target of person re-identification (ReID) and gait recognition is consistent, that is to match the target pedestrian under surveillance cameras. For the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is often researched under controlled conditions. To tackle this problem, we propose a Cloth-Changing benchmark for Person re-identification and Gait recognition (CCPG). It is a cloth-changing dataset, and there are several highlights in CCPG, (1) it provides 200 identities and over 16K sequences are captured indoors and outdoors, (2) each identity has seven different cloth-changing statuses, which is hardly seen in previous datasets, (3) RGB and silhouettes version data are both available for research purposes. Moreover, aiming to investigate the cloth-changing problem systematically, comprehensive experiments are conducted on video-based ReID and gait recognition methods. The experimental results demonstrate the superiority of ReID and gait recognition separately in different cloth-changing conditions and suggest that gait recognition is a potential solution for addressing the cloth-changing problem. Our dataset will be available at https://github.com/BNU-IVC/CCPG.
Saihui Hou, Chunjie Zhang 0001, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Yao Zhao 0001
CVPR4
2023 Dynamic Aggregated Network for Gait Recognition
abstract
Gait recognition is beneficial for a variety of applications, including video surveillance, crime scene investigation, and social security, to mention a few. However, gait recognition often suffers from multiple exterior factors in real scenes, such as carrying conditions, wearing overcoats, and diverse viewing angles. Recently, various deep learning-based gait recognition methods have achieved promising results, but they tend to extract one of the salient features using fixed-weighted convolutional networks, do not well consider the relationship within gait features in key regions, and ignore the aggregation of complete motion patterns. In this paper, we propose a new perspective that actual gait features include global motion patterns in multiple key regions, and each global motion pattern is composed of a series of local motion patterns. To this end, we propose a Dynamic Aggregation Network (DANet) to learn more discriminative gait features. Specifically, we create a dynamic attention mechanism between the features of neighboring pixels that not only adaptively focuses on key regions but also generates more expressive local motion patterns. In addition, we develop a selfattention mechanism to select representative local motion patterns and further learn robust global motion patterns. Extensive experiments on three popular public gait datasets, i.e., CASIA-B, OUMVLP, and Gait3D, demonstrate that the proposed method can provide substantial improvements over the current state-of-the-art methods.1
Ying Fu 0001, Dezhi Zheng, Chunshui Cao, Xuecai Hu, Yongzhen Huang
CVPR4
2023 ChatEdit: Towards Multi-turn Interactive Facial Image Editing via Dialogue
abstract
Single-turn Multi-turn Input Lipstick Pale skin Smiling Black Hair Figure 2: Comparison of previous repeated singleturn editing approaches and our proposed multiturn editing approach.The cascaded errors in the single-turn approach lead to unintended changes in gender and eye makeup.
Xing Cui, Zekun Li 0001, Yibo Hu 0001, Hailin Shi, Chunshui Cao, Zhaofeng He 0001
EMNLP6
2023 Fine-grained Unsupervised Domain Adaptation for Gait Recognition
abstract
Gait recognition has emerged as a promising technique for the long-range retrieval of pedestrians, providing numerous advantages such as accurate identification in challenging conditions and non-intrusiveness, making it highly desirable for improving public safety and security. However, the high cost of labeling datasets, which is a prerequisite for most existing fully supervised approaches, poses a significant obstacle to the development of gait recognition. Recently, some unsupervised methods for gait recognition have shown promising results. However, these methods mainly rely on a fine-tuning approach that does not sufficiently consider the relationship between source and target domains, leading to the catastrophic forgetting of source domain knowledge. This paper presents a novel perspective that adjacent-view sequences exhibit overlapping views, which can be leveraged by the network to gradually attain cross-view and cross-dressing capabilities without pre-training on the labeled source domain. Specifically, we propose a fine-grained Unsupervised Domain Adaptation (UDA) framework that iteratively alternates between two stages. The initial stage involves offline clustering, which transfers knowledge from the labeled source domain to the unlabeled target domain and adaptively generates pseudo-labels according to the expressiveness of each part. Subsequently, the second stage encompasses online training, which further achieves cross-dressing capabilities by continuously learning to distinguish numerous features of source and target domains. The effectiveness of the proposed method is demonstrated through extensive experiments conducted on widely-used public gait datasets.
Ying Fu 0001, Dezhi Zheng, Yunjie Peng, Chunshui Cao, Yongzhen Huang
ICCV5
2023 Occluded Gait Recognition
abstract
Gait recognition suffers from common occlusions in real-world applications. However, academic research on gait recognition usually assumes access to full-body input data. For bridging the gap to practical applications, we propose to identify people when given an occluded gait sequence, namely occluded gait recognition. Since publicly available datasets do not meet the requirements of the intended research, we design a new framework named OccSilGait to generate realistic occluded gait silhouette sequences based on the principle of perspective transformation. Specifically, OccSilGait considers various occlusion scenarios including non-occlusion, crowd occlusion, static occlusion, and detection occlusion. And we employ OccSilGait to build the occluded gait dataset OccCASIA-B for further research. To address challenges brought by occlusion for gait recognition, we propose a novel SpaAlignTemOccRecover network consisting of 1) a Spatial auto-Align module that transforms silhouettes into spatially aligned ones with well-designed self-supervision; 2) a Spatial-Temporal Backbone that alternatively extracts spatial and temporal features to avoid the diffusion of occlusion; 3) a Temporal Occlusion Recovery module that reconstructs the current frame based on time index and temporal context, exploiting gait periodicity for occlusion recovery. Experiments on the newly built occluded dataset show the superiority of the proposed method. Both the OccSilGait framework and the code are available at https://github.com/YunjiePeng/OccludedGaitRecognition.
Yunjie Peng, Chunshui Cao, Zhiqiang He 0002
IJCNN2
2023 Causal Intervention for Sparse-View Gait Recognition
abstract
Gait recognition aims at identifying individuals by unique walking patterns at a long distance. However, prevailing methods suffer from a large degradation when applied to large-scale surveillance systems. We find a significant cause of this issue is that previous methods heavily rely on full-view person annotations to reduce view differences by pulling closer the anchor to positive samples from different viewpoints. But, subjects under in-the-wild scenarios usually have only a limited number of sequences from different viewpoints. As a result, the available viewpoints of each subject are sparse compared to the whole dataset, and simply minimizing intra-identity differences cannot well reducing the view differences in the whole dataset. In this work, we formulate this overlooked problem as Sparse-View Gait Recognition and provide a comprehensive analysis of it by a Structural Causal Model for causalities among latent features, view distribution, and labels. Based on our analysis, we propose a simple yet effective method that enables networks to learn a more robust representation among different views. Specifically, our method consists of two parts: 1) an effective metric learning algorithmic implementation based on the backdoor adjustment, which improves the consistency of representations among different views; 2) an unsupervised view cluster algorithm to discover and identify the most influential view contexts. We evaluate the effectiveness of our method on popular GREW, Gait3D, CASIA-B, and OU-MVLP, showing that our method consistently outperforms baselines and achieves state-of-the-art performance. The code will be available at https://github.com/wj1tr0y/GaitCSV.
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Liang Wang 0001
ACM Multimedia4
2023 LandmarkGait: Intrinsic Human Parsing for Gait Recognition
abstract
Gait recognition is an emerging biometric technology for identifying pedestrians based on their unique walking patterns. In past gait recognition, global-based methods are inadequate to meet the growing demand for accuracy, while commonly used part-based methods provided coarse and inaccurate feature representation for specific body parts. Human parsing appears to be a better option for accurately representing specific and complete body parts in gait recognition. However, its practical application in gait recognition is often hindered by missing RGB modality, lack of annotated body parts, and difficulty in balancing parsing quantity and quality. To address this issue, we propose LandmarkGait, an accessible and alternative parsing-based solution for gait recognition. LandmarkGait introduces an unsupervised landmark discovery network to transform the dense silhouette into a finite set of landmarks with remarkable consistency across various conditions. By grouping landmarks subsets corresponding to distinct body part regions, following a reconstruction task and further refinement from high-quality input silhouettes, we can directly obtain fine-grained parsing results from original binary silhouettes in an unsupervised manner. Moreover, we also develop a multi-scale feature extractor that simultaneously captures global and parsing feature representations based on the integrity and flexibility of specific body parts. Extensive experiments demonstrate that our LandmarkGait can extract more stable features and exhibit significant performance improvement under all conditions, especially in various dressing conditions. Code is available at https://github.com/wzb-bupt/LandmarkGait.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu
ACM Multimedia5
2023 Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification
abstract
We present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a vision-language alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance.
Rui Wang 0124, Peipei Li 0002, Huaibo Huang, Chunshui Cao, Ran He 0001, Zhaofeng He 0001
NeurIPS4
2023 Gait Quality Aware Network: Toward the Interpretability of Silhouette-Based Gait Recognition
abstract
Gait recognition receives increasing attention since it can be conducted at a long distance in a nonintrusive way and applied to the condition of changing clothes. Most existing methods take the silhouettes of gait sequences as the input and learn a unified representation from multiple silhouettes to match probe and gallery. However, these models are all faced with the lack of interpretability, e.g., it is not clear which silhouette in a gait sequence and which part in the human body are relatively more important for recognition. In this work, we propose a gait quality aware network (GQAN) for gait recognition which explicitly assesses the quality of each silhouette and each part via two blocks: frame quality block (FQBlock) and part quality block (PQBlock). Specifically, FQBlock works in a squeeze-and-excitation style to recalibrate the features for each silhouette, and the scores of all the channels are added as frame quality indicator. PQBlock predicts a score for each part which is used to compute the weighted distance between the probe and gallery. Particularly, we propose a part quality loss (PQLoss) which enables GQAN to be trained in an end-to-end manner with only sequence-level identity annotations. This work is meaningful by moving toward the interpretability of silhouette-based gait recognition, and our method also achieves very competitive performance on CASIA-B and OUMVLP.
Saihui Hou, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Neural Networks Learn. Syst.3
2020 GaitPart: Temporal Part-Based Model for Gait Recognition
abstract
Gait recognition, applied to identify individual walking patterns in a long-distance, is one of the most promising video-based biometric technologies. At present, most gait recognition methods take the whole human body as a unit to establish the spatio-temporal representations. However, we have observed that different parts of human body possess evidently various visual appearances and movement patterns during walking. In the latest literature, employing partial features for human body description has been verified being beneficial to individual recognition. Taken above insights together, we assume that each part of human body needs its own spatio-temporal expression. Then, we propose a novel part-based model GaitPart and get two aspects effect of boosting the performance: On the one hand, Focal Convolution Layer, a new applying of convolution, is presented to enhance the fine-grained learning of the part-level spatial features. On the other hand, the Micro-motion Capture Module (MCM) is proposed and there are several parallel MCMs in the GaitPart corresponding to the pre-defined parts of the human body, respectively. It is worth mentioning that the MCM is a novel way of temporal modeling for gait task, which focuses on the short-range temporal features rather than the redundant long-range features for cycle gait. Experiments on two of the most popular public datasets, CASIA-B and OU-MVLP, richly exemplified that our method meets a new state-of-the-art on multiple standard benchmarks. The source code will be available on https://github.com/ChaoFan96/GaitPart.
Chao Fan 0001, Yunjie Peng, Chunshui Cao, Xu Liu 0008, Saihui Hou, Jiannan Chi, Yongzhen Huang, Qing Li 0015, Zhiqiang He 0002
CVPR3
2020 Gait Lateral Network: Learning Discriminative and Compact Representations for Gait Recognition
Saihui Hou, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
ECCV (9)2
2019 Feedback Convolutional Neural Network for Visual Localization and Segmentation
abstract
Feedback is a fundamental mechanism existing in the human visual system, but has not been explored deeply in designing computer vision algorithms. In this paper, we claim that feedback plays a critical role in understanding convolutional neural networks (CNNs), e.g., how a neuron in CNNs describes an object's pattern, and how a collection of neurons form comprehensive perception to an object. To model the feedback in CNNs, we propose a novel model named Feedback CNN and develop two new processing algorithms, i.e., neural pathway pruning and pattern recovering. We mathematically prove that the proposed method can reach local optimum. Note that Feedback CNN belongs to weakly supervised methods and can be trained only using category-level labels. But it possesses a powerful capability to accurately localize and segment category-specific objects. We conduct extensive visualization analysis, and the results reveal the close relationship between neurons and object parts in Feedback CNN. Finally, we evaluate the proposed Feedback CNN over the tasks of weakly supervised object localization and segmentation, and the experimental results on ImageNet and Pascal VOC show that our method remarkably outperforms the state-of-the-art ones.
Chunshui Cao, Yongzhen Huang, Yi Yang 0007, Liang Wang 0001, Zilei Wang, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Lateral Inhibition-Inspired Convolutional Neural Network for Visual Attention and Saliency Detection
abstract
Lateral inhibition in top-down feedback is widely existing in visual neurobiology, but such an important mechanism has not be well explored yet in computer vision. In our recent research, we find that modeling lateral inhibition in convolutional neural network (LICNN) is very useful for visual attention and saliency detection. In this paper, we propose to formulate lateral inhibition inspired by the related studies from neurobiology, and embed it into the top-down gradient computation of a general CNN for classification, i.e. only category-level information is used. After this operation (only conducted once), the network has the ability to generate accurate category-specific attention maps. Further, we apply LICNN for weakly-supervised salient object detection.Extensive experimental studies on a set of databases, e.g., ECSSD, HKU-IS, PASCAL-S and DUT-OMRON, demonstrate the great advantage of LICNN which achieves the state-of-the-art performance. It is especially impressive that LICNN with only category-level supervised information even outperforms some recent methods with segmentation-level supervised learning.
Chunshui Cao, Yongzhen Huang, Zilei Wang, Liang Wang 0001, Ninglong Xu, Tieniu Tan
AAAI1
2015 Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural Networks
abstract
While feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to note that the human visual cortex generally contains more feedback than feedforward connections. In this paper, we will briefly introduce the background of feedbacks in the human visual cortex, which motivates us to develop a computational feedback mechanism in deep neural networks. In addition to the feedforward inference in traditional neural networks, a feedback loop is introduced to infer the activation status of hidden layer neurons according to the "goal" of the network, e.g., high-level semantic labels. We analogize this mechanism as "Look and Think Twice." The feedback networks help better visualize and understand how deep neural networks work, and capture visual attention on expected objects, even in images with cluttered background and multiple objects. Experiments on ImageNet dataset demonstrate its effectiveness in solving tasks such as image classification and object localization.
Chunshui Cao, Xianming Liu 0005, Yi Yang 0007, Yinan Yu, Jiang Wang 0001, Zilei Wang, Yongzhen Huang, Liang Wang 0001, Chang Huang, Wei Xu 0017, Deva Ramanan, Thomas S. Huang
ICCV1