Saihui Hou

dblp:180/5646 · DBLP profile ↗
← Back
51ranked-venue papers
14as first author
41since 2021 · last 2026
0000-0003-4689-2860ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 first-author · 28 since 2021Artificial intelligence and machine learning · 34 · 11 first-author · 25 since 2021Security and privacy · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Gait Transformer: End-to-End Transformer Backbone for Gait Recognition
abstract
Gait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modality, spatial-temporal dynamics, fine-grained motion cues, and limited training data. In this paper, we propose Gait Transformer (GaT), an end-to-end Transformer backbone specifically tailored for silhouette-based gait recognition. GaT introduces three key components: (1) a hybrid patch embedding module that combines convolutional stems with group-batch normalization to enhance structural preservation; (2) a decomposed token mixer that explicitly models both short-range and long-range dependencies across spatial-temporal dimensions; and (3) a hybrid positional encoding strategy that integrates absolute, relative, and rotary embeddings to support efficient training under data scarcity. Without relying on any pretraining, GaT achieves state-of-the-art performance on Gait3D, GREW, and CCGR-MINI.
Saihui Hou, Wenpeng Lang, Jilong Wang 0010, Yan Huang 0008, Liang Wang 0001, Yongzhen Huang
AAAI1
2026 ForgeryMoE: Mixture of Experts for Image Forgery Detection under JPEG Compression
abstract
AI-generated imagery threatens information integrity, making reliable forgery detection crucial. However, the pervasive use of JPEG compression throughout image sharing pipelines critically undermines the robustness and generalization capability of existing detectors. To address this, we propose ForgeryMoE, a new Mixture of Experts (MoE) framework designed for universal and robust image forgery detection. Our framework strategically integrates three complementary experts: a Frequency Domain Expert that analyzes wavelet-based artifacts, a Pre-trained model expert that leverages CLIP and DINOv2 for semantic inconsistencies, and a Visual Domain Expert that captures pixel-level features and reconstruction anomalies. A dual gating mechanism dynamically estimates expert reliability and adaptively fuses their decisions. Extensive experiments on the GenImage benchmark show our method achieves state-of-the-art performance in both pristine and JPEG compressed settings, demonstrating strong generalization across diverse generative models and compression qualities.
Saihui Hou, Jian Zhao 0006, Zhaofeng He 0001
ICMR2
2025 RA-GAR: A Richly Annotated Benchmark for Gait Attribute Recognition
abstract
Gait attracts growing interest from researchers due to its advantages as a non-invasive and non-cooperative biometric feature. Current gait-based attribute recognition methods primarily focus on estimating attributes such as gender, age, and emotions. However, there is insufficient attention to diverse gait attributes in various covariate scenarios. In this paper, we design and collect a Richly Annotated benchmark for 15 gait attributes, named RA-GAR, comprising data from 533 individuals with over 120,000 sequences. To our knowledge, RA-GAR represents the largest and most diverse benchmark of gait attributes currently available. Furthermore, to fully leverage the semantic information and enhance attribute-specific local perception, we propose a two-stage CLIP-based method for Gait Attribute Recognition, named CLIP-GAR. Experiments on the RA-GAR and MA-Gait datasets demonstrate the effectiveness of CLIP-GAR, showing significant improvements in mean accuracy and F1 score.
Chenye Wang, Saihui Hou, Aoqi Li, Qingyuan Cai, Yongzhen Huang
AAAI2
2025 Bridging Gait Recognition and Large Language Models Sequence Modeling
abstract
Gait sequences exhibit sequential structures and contextual relationships similar to those in natural language, where each element—whether a word or a gait step—is connected to its predecessors and successors. This similarity enables the transformation of gait sequences into "texts" containing identity-related information. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized for gait sequence modeling to enhance gait recognition performance. Leveraging these insights, we make a pioneering effort to apply LLMs to gait recognition, which we refer to as GaitLLM. Specifically, we propose the Gait-to-Language (G2L) module, which converts gait sequences into a textual format suitable for LLMs, and the Language-to-Gait (L2G) module, which maps the LLM’s output back to the gait feature space, thereby bridging the gap between LLM outputs and gait recognition. Notably, GaitLLM leverages the powerful modeling capabilities of LLMs without relying on complex architectural designs, improving gait recognition performance with only a small number of trainable parameters. Our method achieves state-of-the-art results on four popular gait datasets—SUSTech1K, CCPG, Gait3D, and GREW—demonstrating the effectiveness of applying LLMs in this domain. This work highlights the potential of LLMs to significantly enhance gait recognition, paving the way for future research and practical applications.
Shaopeng Yang, Jilong Wang 0010, Saihui Hou, Xu Liu 0008, Chunshui Cao, Liang Wang 0001, Yongzhen Huang
CVPR3
2025 OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better Generalization
abstract
This paper addresses the challenge of animal re-identification, an emerging field that shares similarities with person re-identification but presents unique complexities due to the diverse species, environments and poses. To facilitate research in this domain, we introduce OpenAnimals, a flexible and extensible codebase designed specifically for animal re-identification. We conduct a comprehensive study by revisiting several state-of-the-art person re-identification methods, including BoT, AGW, SBS, and MGN, and evaluate their effectiveness on animal re-identification benchmarks such as HyenaID, LeopardID, SeaTurtleID, and WhaleSharkID. Our findings reveal that while some techniques generalize well, many do not, underscoring the significant differences between the two tasks. To bridge this gap, we propose ARBase, a strong \textbf{Base} model tailored for \textbf{A}nimal \textbf{R}e-identification, which incorporates insights from extensive experiments and introduces simple yet effective animal-oriented designs. Experiments demonstrate that ARBase consistently outperforms existing baselines, achieving state-of-the-art performance across various benchmarks.
Saihui Hou, Panjian Huang, Zengbin Wang, Yuan Liu 0043, Man Zhang 0005, Yongzhen Huang
ICCV1
2025 Learning a Unified Template for Gait Recognition
abstract
"What I cannot create, I do not understand."Human wisdom reveals that creation is one of the highest forms of learning. For example, Diffusion Models have demonstrated remarkable semantic structure and memory in image generation, understanding, and restoration, which intuitively benefits representation learning. However, current gait networks rarely embrace this perspective, relying primarily on learning by contrasting gait samples under varying complex conditions, leading to semantic inconsistency and uniformity issues. To address these issues, we propose Origins with generative capabilities whose underlying philosophy is that different entities are generated from a unified template, inherently regularizing gait representations within a consistent and diverse semantic space to capture accurate gait differences. Admittedly, learning this unified template is exceedingly challenging, as it requires the comprehensiveness of the template to encompass gait representations with various conditions. Inspired by Diffusion Models, Origins diffuses the unified template into timestep templates for gait generative learning, and meanwhile transfers the unified template for gait representation learning. Especially, gait generative and representation learning serve as a unified framework for end-to-end joint training. Extensive experiments on CASIA-B, CCPG,SUSTech1K, Gait3D, GREW and CCGR-MINI demonstrate that Origins performs unified generative and representation learning, achieving superior performance.
Panjian Huang, Saihui Hou, Junzhou Huang, Yongzhen Huang
ICCV2
2025 OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference Captioning
abstract
Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements, existing datasets often lack breadth and depth, limiting their applicability in complex and dynamic environments: (1) from a breadth perspective, current datasets are constrained to limited variations of objects in specific scenes, and (2) from a depth perspective, prior benchmarks often provide overly simplistic descriptions. To address these challenges, we introduce OmniDiff, a comprehensive dataset comprising 324 diverse scenarios-spanning real-world complex environments and 3D synthetic settings-with fine-grained human annotations averaging 60 words in length and covering 12 distinct change types. Building on this foundation, we propose M$^3$Diff, a MultiModal large language model enhanced by a plug-and-play Multi-scale Differential Perception (MDP) module. This module improves the model's ability to accurately identify and describe inter-image differences while maintaining the foundational model's generalization capabilities. With the addition of the OmniDiff dataset, M$^3$Diff achieves state-of-the-art performance across multiple benchmarks, including Spot-the-Diff, IEdit, CLEVR-Change, CLEVR-DC, and OmniDiff, demonstrating significant improvements in cross-scenario difference recognition accuracy compared to existing methods. The dataset, code, and models will be made publicly available to support further research.
Yuan Liu 0043, Saihui Hou, Saijie Hou, Jiabao Du, Shibei Meng, Yongzhen Huang
ICCV2
2025 Gait: Exploring X Modality for Generalized Gait Recognition
Zengbin Wang, Saihui Hou, Junjie Li 0002, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Siye Wang, Man Zhang 0005
ICCV2
2025 Beyond Sparse Keypoints: Dense Pose Modeling for Robust Gait Recognition
abstract
Gait recognition has emerged as a promising biometric technology due to its ability to operate at a distance without subject cooperation. While pose-based methods offer advantages over appearance-based approaches in robustness and interpretability, their performance has been limited by the sparse keypoint representations of current pose estimation frameworks. We identify two critical limitations: (1) incomplete motion representation due to insufficient keypoints for dynamic body parts, and (2) lack of shape information from minimal skeleton points. This paper presents DPGait, a novel framework that addresses these challenges through innovations in both upstream processing and downstream modeling. First, we enhance pose estimation by extending the standard COCO keypoint format with additional motion-sensitive points and shape-descriptive keypoints inspired by human mesh estimation. Second, we propose a divide-and-conquer modeling strategy that processes dense keypoints through group convolution with cross-group attention, coupled with multi-granularity supervision for improved training. Our comprehensive experiments demonstrate state-of-the-art performance in pose-based gait recognition, achieving 85.8% rank-1 accuracy on SUSTech1K-surpassing leading silhouette-based methods for the first time. The results validate that dense pose representation combined with our novel modeling approach significantly advances the field of gait recognition.
Wenpeng Lang, Saihui Hou, Yongzhen Huang
ACM Multimedia2
2025 Seeing from Magic Mirror: Contrastive Learning from Reconstruction for Pose-based Gait Recognition
abstract
While recent advancements in supervised gait recognition have yielded promising results, these approaches rely heavily on annotated walking data, limiting their generalizability to complex environments. This paper presents a self-supervised gait recognition framework using human poses as input to address this challenge, focusing on high-quality pretrained data and self-supervised learning strategies. We first introduce StreamGait, a large-scale, unlabelled dataset that captures in-the-wild distributions of walking sequences. This dataset is curated from Internet livestreams across diverse geographic and environmental scenarios, reflecting variations in real-world camera angles, weather, and pedestrian behavior. Our framework, MirrorGait, conducts self-supervised learning by integration with 2D-to-3D pose reconstruction to synthesize multi-view perspectives for effective 3D-aware contrastive learning. With specific designs of temporal position embedding and gait partition head on a Transformer backbone, the encoder can readily adapt to the periodic and fine-grained nature of gait. Extensive experiments on three widely used gait datasets, Gait3D, GREW, and OUMVLP-Pose, demonstrate that our method, with minimal fine-tuning on the pretrained model, achieves state-of-the-art performance among pose-based gait recognition approaches. The dataset, code, and models are available at https://github.com/BNU-IVC/StreamGait.
Shibei Meng, Saihui Hou, Xuecai Hu, Junzhou Huang, Yongzhen Huang
ACM Multimedia2
2025 Vocabulary-Guided Gait Recognition
abstract
What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points. However, the considerations remain vague for humans to comprehend truly. In this work, we introduce a novel paradigm Vocabulary-Guided Gait Recognition, dubbed Gait-World, which attempts to explore gait concepts through human vocabularies with Vision-Language Models (VLMs). Despite VLMs have achieved the remarkable progress in various vision tasks, the cognitive capability regarding gait modalities remains limited. The success element in Gait-World is the proper vocabulary prompt where this paradigm carefully selects gait cycle actions as Vocabulary Base, bridging the gait and vocabulary feature spaces and further promoting human understanding for the gait. How to extract gait features? Although previous gait networks have made significant progress, learning solely from gait modalities on limited gait databases makes it difficult to learn robust gait features for practicality. Therefore, we propose the first Gait-World model, dubbed $\alpha$-Gait, which guides the gait network learning with universal vocabulary knowledge from VLMs. However, due to the heterogeneity of the modalities, directly integrating vocabulary and gait features is highly challenging as they reside in different embedding spaces. To address the issues, $\alpha$-Gait designs Vocabulary Relation Mapper and Gait Fine-grained Detector to map and establish vocabulary relations in the gait space for detecting corresponding gait features. Extensive experiments on CASIA-B, CCPG, SUSTech1K, Gait3D and GREW reveal the potential value and research directions of vocabulary information from VLMs in the gait field.
Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
NeurIPS2
2025 Edge-Oriented Adversarial Attack for Deep Gait Recognition
Saihui Hou, Zengbin Wang, Man Zhang 0005, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
Int. J. Comput. Vis.1
2025 OpenGait: A Comprehensive Benchmark Study for Gait Recognition Toward Better Practicality
abstract
Gait recognition, a rapidly advancing vision technology for person identification from a distance, has made significant strides in indoor settings. However, evidence suggests that existing methods often yield unsatisfactory results when applied to newly released real-world gait datasets. Furthermore, conclusions drawn from indoor gait datasets may not easily generalize to outdoor ones. Therefore, the primary goal of this paper is to present a comprehensive benchmark study aimed at improving practicality rather than solely focusing on enhancing performance. To this end, we developed OpenGait, a flexible and efficient gait recognition platform. Using OpenGait, we conducted in-depth ablation experiments to revisit recent developments in gait recognition. Surprisingly, we detected some imperfect parts of some prior methods and thereby uncovered several critical yet previously neglected insights. These findings led us to develop three structurally simple yet empirically powerful and practically robust baseline models: DeepGaitV2, SkeletonGait, and SkeletonGait++, which represent the appearance-based, model-based, and multi-modal methodologies for gait pattern description, respectively. In addition to achieving state-of-the-art performance, our careful exploration provides new perspectives on the modeling experience of deep gait models and the representational capacity of typical gait modalities. In the end, we discuss the key trends and challenges in current gait recognition, aiming to inspire further advancements towards better practicality.
Chao Fan 0001, Saihui Hou, Chuanfu Shen, Jingzhe Ma, Dongyang Jin, Yongzhen Huang, Shiqi Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 From FastPoseGait to GPGait++: Bridging the Past and Future for Pose-Based Gait Recognition
abstract
Recent studies on pose-based gait recognition have underscored the potential of utilizing such fundamental data to achieve superior outcomes. Nonetheless, the development of current pose-based methods faces significant obstacles due to several critical issues: (1) Misaligned Settings, which results in a lack of thorough and unbiased comparative analysis. (2) Inferior Performance, which causes diminished focus on pose-based gait representations. (3) Limited Generalization, which hinders the effective application in real-world scenarios. Focused on tackling the aforementioned challenges, our study introduces a comprehensive benchmark and a versatile approach to bridge the past and future for pose-based gait recognition. First, we revisit previous pose-based methods and make great efforts to establish a unified framework, FastPoseGait, aiming at a fair and comprehensive comparison investigation with consistent experimental settings and a more stable training process. Then, within this framework, we propose GPGait++, a generalized pose-based gait recognition method featuring a human-oriented input and part-aware modeling, intended to enhance the generalization ability and discriminative power across diverse environments and camera viewpoints. Experiments on six public gait recognition datasets reveal that our unified framework significantly enhances the performance of previous approaches, and GPGait++ exhibits state-of-the-art cross-domain capabilities compared to existing pose-based methods, marking a significant advancement in the field of pose-based gait recognition.
Shibei Meng, Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 GaitC3I: Robust Cross-Covariate Gait Recognition via Causal Intervention
abstract
Cross-covariate gait recognition aims to analyze a pedestrian’s gait to extract an identity representation that is invariant across varying covariates. However, prevailing methods that have achieved good results on controlled in-the-lab datasets often perform poorly on realistic datasets. In this work, we find a significant cause is that the widely used pairwise metric learning paradigm cannot correctly handle the relationship between samples from different covariate conditions. Even worse, it may yield harmful signals that inadvertently mislead models to focus on covariate-related features, particularly when covariate distributions vary across subjects. To address this issue, we propose a Cross-Covariate Causal Intervention (GaitC3I) framework, a unified causality-inspired approach aimed at enhancing the robustness of gait recognition across diverse conditions. Specifically, our method consists of two parts: 1) an effective causal intervention metric learning paradigm based on backdoor adjustment, which strategically mitigates spurious correlations induced by covariates, thus ensuring a more invariant gait representation; and 2) an annotation-free selection strategy that progressively matches each positive sample with negative samples from similar covariate conditions at various granularities. We demonstrate the effectiveness of our GaitC3I through extensive evaluation on six popular gait datasets-Gait3D, GREW, OUMVLP, CASIA-B, CCPG, and CCGR-achieving substantial improvements. Our method not only outperforms existing state-of-the-art models but also provides a systematic solution to remove the spurious correlations in gait recognition.
Jilong Wang 0010, Saihui Hou, Xianda Guo, Yan Huang 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 GaitAsset: In Defense of Regarding Gait as a Set
abstract
In the field of gait recognition, regarding gait as a set has emerged as a seminal approach, notably eliminating the dependence on template-based input. Although set-based methods offer notable advantages, such as insensitivity to frame order permutations and robustness to varying frame counts, their performance has consistently lagged behind that of sequence-based methods in subsequent studies. In this work, we advocate for treating gait as an unordered set and argue thatthe lack of set context aggregation in frame-level feature extraction is the primary limitation hindering the full potential of set-based gait recognition. To substantiate this claim, we develop a gait-oriented self-attention module and introduce a Gating Mechanism that facilitates set context awareness for each silhouette whilepreserving the permutation-invariant property. Specifically, the context aggregation operates on diverse bins of feature maps, interleaving fine-grained shape and motion details in an almost parameter-free manner. The Gating Mechanism is employed to ensure that frame-level features are not overwhelmed by the aggregated context. Furthermore, the sampling strategy is carefully enhanced to better support set context modeling. Our research demonstrates that set-based gait recognition can achieve state-of-the-art accuracy on in-the-wild benchmarks (77.6% on Gait3D and 81.1% on GREW) while retaining its inherent advantages.
Saihui Hou, Chenye Wang, Aoqi Li, Jilong Wang 0010, Liang Wang 0001, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.1
2025 Multimodal Mutual Learning for Unsupervised Gait Recognition
abstract
The primary challenge in unsupervised gait recognition lies in generating meaningful and diverse supervisory signals to guide representation learning. The effectiveness of such methods largely depends on the richness of the supervisory signals. Unlike previous methods that construct supervisory signals solely from a single modality, we propose a novel framework, named Multimodal Mutual Learning (M3L), that leverages the identity consistency and complementary nature of both silhouette and skeleton modalities to generate richer and more informative supervisory signals. To fully leverage the richer supervisory signals, M3L encourages mutual prediction between the silhouette and skeleton modalities, guiding the network toward modality-invariant representations. However, mutual prediction alone is hindered by the inherent modality gap, so we introduce a Multimodal Collaborative Module to explicitly bridge this gap and promote cross-modal knowledge transfer. Moreover, to make the framework practical when only one modality is available at inference, we introduce a Multimodal Disentanglement Module. Multimodal Disentanglement Module decouples the two branches and distills a shared representation, preserving the gains of multimodal training while allowing the model to maintain robust performance under single-modality conditions. Extensive experiments on four widely used gait datasets—Gait3D, GREW, CASIA-B, and SUSTech1K—demonstrate the effectiveness of our approach and highlight its potential to advance unsupervised gait recognition.
Shaopeng Yang, Saihui Hou, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.2
2024 Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser
abstract
Recently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods decompose the task into bone length and bone direction prediction based on the human anatomical skeleton to explicitly incorporate more human body prior constraints, the performance of these methods is significantly lower than that of the SOTA diffusion-based methods. This can be attributed to the tree structure of the human skeleton. Direct application of the disentangled method could amplify the accumulation of hierarchical errors, propagating through each hierarchy. Meanwhile, the hierarchical information has not been fully explored by the previous methods. To address these problems, a Disentangled Diffusion-based 3D human Pose Estimation method with Hierarchical Spatial and Temporal Denoiser is proposed, termed DDHPose. In our approach: (1) We disentangle the 3d pose and diffuse the bone length and bone direction during the forward process of the diffusion model to effectively model the human pose prior. A disentanglement loss is proposed to supervise diffusion model learning. (2) For the reverse process, we propose Hierarchical Spatial and Temporal Denoiser (HSTDenoiser) to improve the hierarchical modelling of each joint. Our HSTDenoiser comprises two components: the Hierarchical-Related Spatial Transformer (HRST) and the Hierarchical-Related Temporal Transformer (HRTT). HRST exploits joint spatial information and the influence of the parent joint on each joint for spatial modeling, while HRTT utilizes information from both the joint and its hierarchical adjacent joints to explore the hierarchical temporal correlations among joints. Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets show that our method outperforms the SOTA disentangled-based, non-disentangled based, and probabilistic approaches by 10.0%, 2.0%, and 1.3%, respectively.
Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 0002, Yongzhen Huang
AAAI3
2024 QAGait: Revisit Gait Recognition from a Quality Perspective
abstract
Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research. However, as gait recognition rapidly advances from in-the-lab to in-the-wild scenarios, various conditions raise significant challenges for silhouette modality, including 1) unidentifiable low-quality silhouettes (abnormal segmentation, severe occlusion, or even non-human shape), and 2) identifiable but challenging silhouettes (background noise, non-standard posture, slight occlusion). To address these challenges, we revisit gait recognition pipeline and approach gait recognition from a quality perspective, namely QAGait. Specifically, we propose a series of cost-effective quality assessment strategies, including Maxmial Connect Area and Template Match to eliminate background noises and unidentifiable silhouettes, Alignment strategy to handle non-standard postures. We also propose two quality-aware loss functions to integrate silhouette quality into optimization within the embedding space. Extensive experiments demonstrate our QAGait can guarantee both gait reliability and performance enhancement. Furthermore, our quality assessment strategies can seamlessly integrate with existing gait datasets, showcasing our superiority. Code is available at https://github.com/wzb-bupt/QAGait.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu
AAAI2
2024 Learning Visual Prompt for Gait Recognition
abstract
Gait, a prevalent and complex form of human motion, plays a significant role in the field of long-range pedestrian retrieval due to the unique characteristics inherent in individual motion patterns. However, gait recognition in real-world scenarios is challenging due to the limitations of capturing comprehensive cross-viewing and crossclothing data. Additionally, distractors such as occlusions, directional changes, and lingering movements further complicate the problem. The widespread application of deep learning techniques has led to the development of various potential gait recognition methods. However, these methods utilize convolutional networks to extract shared information across different views and attire conditions. Once trained, the parameters and non-linear function become constrained to fixed patterns, limiting their adaptability to various distractors in real-world scenarios. In this paper, we present a unified gait recognition framework to extract global motion patterns and develop a novel dynamic transformer to generate representative gait features. Specifically, we develop a trainable part-based prompt pool with numerous key-value pairs that can dynamically select prompt templates to incorporate into the gait sequence, thereby providing task-relevant shared knowledge information. Furthermore, we specifically design dynamic attention to extract robust motion patterns and address the length generalization issue. Extensive experiments on four widely recognized gait datasets, i.e., Gait3D, GREW, OUMVLP, and CASIA-B, reveal that the proposed method yields substantial improvements compared to current state-of-the-art approaches.
Ying Fu 0001, Chunshui Cao, Saihui Hou, Yongzhen Huang, Dezhi Zheng
CVPR4
2024 Cut Out the Middleman: Revisiting Pose-Based Gait Recognition
Saihui Hou, Shibei Meng, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
ECCV (31)2
2024 Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective
Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu 0008, Zhiqiang He 0002, Yongzhen Huang
ECCV (6)3
2024 Free Lunch for Gait Recognition: A Novel Relation Descriptor
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001
ECCV (38)2
2024 Probabilistic Contrastive Learning for Domain Adaptation
Junjie Li 0002, Yixin Zhang 0007, Zilei Wang, Saihui Hou, Keyu Tu, Man Zhang 0005
IJCAI4
2024 AerialGait: Bridging Aerial and Ground Views for Gait Recognition
abstract
In this work, we present AerialGait, a comprehensive dataset for aerial-ground gait recognition. This dataset comprises 82,454 sequences totaling over 10 million frames from 533 subjects, captured from both aerial and ground perspectives. To align with real-life scenarios of aerial and ground surveillance, we utilize a drone and a ground surveillance camera for data acquisition. The drone is operated at various speeds, directions, and altitudes. Meanwhile, we conduct data collection across five diverse surveillance sites to ensure a comprehensive simulation of real-world settings. AerialGait has several unique features: 1) The gait sequences exhibit significant variations in views, resolutions, and illumination across five distinct scenes. 2) It incorporates challenges of motion blur and frame discontinuity due to drone mobility. 3) The dataset reflects the domain gap caused by the view disparity between aerial and ground views, presenting a realistic challenge for drone-based gait recognition. Moreover, we perform a comprehensive analysis of existing gait recognition methods on AerialGait dataset and propose the Aerial-Ground Gait Network (AGG-Net). AGG-Net effectively learns discriminative features from aerial views by uncertainty learning and clusters features across aerial and ground views through prototype learning. Our model achieves state-of-the-art performance on both AerialGait and DroneGait datasets.
Aoqi Li, Saihui Hou, Chenye Wang, Qingyuan Cai, Yongzhen Huang
ACM Multimedia2
2024 Depression risk recognition based on gait: A benchmark
Saihui Hou, Xuecai Hu, Yongzhen Huang
Neurocomputing3
2024 Cloth-Imbalanced Gait Recognition via Hallucination
abstract
The study in the gait field has rarely paid attention to the class-imbalanced learning, while the realistic data always exhibits an imbalanced distribution. The main reason lies in the difficulty of collecting the cross-clothes sequences, since the collection is usually aided by person re-identification and it is more likely to obtain the sequences for a subject wearing the same clothes. In this work, we formulate a new problem to tackle the task-specific cloth-imbalanced issue, dubbed as Cloth-Imbalanced Gait Recognition, and the training data consists of two parts denoted as head set and tail set. The sequences for a subject in head set cover the cross-clothes variation which is scarce in tail set to mimic the collection difficulty. Along with the problem formulation, we design a new method to deal with the inherent challenges, called Cross-Clothes Hallucination or CCH for short. Our method is inspired by the observation that certain directions in deep feature space correspond to meaningful semantic transformations, and it tries to generate the cross-clothes sequences for tail set referring to the cloth-changing transformation in head set. To evaluate CCH, we build two cloth-imbalanced benchmarks based on the widely-used CASIA-B and Outdoor-Gait. Extensive experiments demonstrate that CCH brings significant improvements over the baselines.
Saihui Hou, Panjian Huang, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Verifying in the Dark: Verifiable Machine Unlearning by Using Invisible Backdoor Triggers
abstract
Machine unlearning as a fundamental requirement in Machine-Learning-as-a-Service (MLaaS) has been extensively studied with increasing concerns about data privacy. It requires MLaaS providers should delete training data upon user requests. Unfortunately, none of the existing studies can efficiently achieve machine unlearning validation while preserving the retraining efficiency and the service quality after data deletion. Besides, how to craft the validation scheme to prevent providers from spoofing validation by forging proofs remains under-explored. In this paper, we introduce a backdoor-assisted validation scheme for machine unlearning. The proposed design is built from the ingenious combination of backdoor triggers and incremental learning to assist users in verifying proofs of machine unlearning without compromising performance and service quality. We propose to embed invisible markers based on backdoor triggers into privacy-sensitive data to prevent MLaaS providers from distinguishing poisoned data for validation spoofing. Users can use prediction results to determine whether providers comply with data deletion requests. Besides, we incorporate our validation scheme into an efficient incremental learning approach via our index structure to further facilitate the performance of retraining after data deletion. Evaluation results on real-world datasets confirm the efficiency and effectiveness of our proposed verifiable machine unlearning scheme.
Yu Guo 0003, Saihui Hou, Cong Wang 0001, Xiaohua Jia
IEEE Trans. Inf. Forensics Secur.3
2024 Adaptive Knowledge Transfer for Weak-Shot Gait Recognition
abstract
Most works for cloth-changing gait recognition assume that the sequences of different clothes are accessible for each subject in the training set, which, however, is almost impossible for real applications. In practice, the collection of gait sequences is usually aided by person re-identification which is more likely to cluster the cloth-consistent sequences for a subject, and it is laborious to merge the cloth-changing clusters with the same identity. As a result, the training set is usually comprised of two subsets, i.e., a fully-annotated base set where the cloth-changing sequences are available for each subject, and a weakly-annotated wild set where the sequences of different clothes for a subject are assigned diverse labels. In this work, we formulate a problem named Weak-Shot Gait Recognition which seeks to learn discriminative features from the mixture of base set and wild set. Furthermore, we propose an effective method called Adaptive Knowledge Transfer to deal with the weak-shot issue. In particular, we define the knowledge as the ability to judge whether two sequences come from the same subject or not, and we take an adaptive way to mine the useful information from wild set. For the experimental study, we build three weak-shot benchmarks based on CASIA-B, Outdoor-Gait, and CASIA-E respectively. Extensive experiments show that our method can bring consistent improvement. For example, under the cloth-changing condition of on the weak-shot CASIA-B, our method exceeds a naïve baseline by 7.10%.
Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.1
2024 Integral Pose Learning via Appearance Transfer for Gait Recognition
abstract
Gait recognition plays an important role in video surveillance and security by identifying humans based on their unique walking patterns. The existing gait recognition methods have achieved competitive accuracy with shape and motion patterns under limited-covariate conditions. However, when extreme appearance changes distort discriminative features, gait recognition yields unsatisfactory results under cross-covariate conditions. In this work, we first indicate that the integral pose in each silhouette maintains an appearance-unrelated discriminative identity. However, the monotonous appearance variables in a gait database cause gait models to have difficulty extracting integral poses. Therefore, we propose an Appearance-transferable Disentangling and Generative Network (GaitApp) to generate gait silhouettes with rich appearances and invariant poses. Specifically, GaitApp leverages multi-branch cooperation to disentangle pose features and appearance features, and transfers the appearance information from one subject to another. By simulating a person constantly changing appearances under limited-covariate conditions, downstream models enable to extract integral discriminative pose features. Extensive experiments demonstrate that our method allows representative gait models to stand at a new altitude, further promoting the exploration to cross-covariate gait recognition. All the code is available at https://github.com/Hpjhpjhs/GaitApp.git.
Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Xuecai Hu, Yongzhen Huang
IEEE Trans. Inf. Forensics Secur.2
2024 Gait Attribute Recognition: A New Benchmark for Learning Richer Attributes From Human Gait Patterns
abstract
Compared to gait recognition, Gait Attribute Recognition (GAR) is a seldom-investigated problem. However, since gait attribute recognition can provide richer and finer semantic descriptions, it is an indispensable part of building intelligent gait analysis systems. Nonetheless, the types of attributes considered in the existing datasets are very limited. This paper contributes a new benchmark dataset for gait attribute recognition named Multi-Attribute Gait (MA-Gait). Our MA-Gait contains 95 subjects recorded from 12 camera views, resulting in more than 13000 sequences, with 16 attributes labeled, including six attributes that have never been considered in the literature. Moreover, we propose a Multi-Scale Motion Encoder (MSME) to extract robust motion features, and an Attribute-Guided Feature Selection Module (AGFSM) to adaptively capture the most discriminative attribute features from static appearance features and dynamic motion features for different attributes. Our method achieves the best GAR accuracy on the new dataset. Comprehensive experiments show the effectiveness of the proposed method through both quantitative and qualitative evaluations.
Xu Song, Saihui Hou, Yan Huang 0023, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Caifeng Shan
IEEE Trans. Inf. Forensics Secur.2
2024 Gait Recognition With Drones: A Benchmark
abstract
Gait recognition aims to obtain people's identity through body shape and walking posture. Existing gait recognition studies focus on low vertical view recognition, in which the person and the camera are nearly at the same height. Differently, in this work, we focus on gait recognition at high vertical views. To facilitate the research, we propose a new dataset named DroneGait, where the drones are used to collect the gait data. This dataset contains 22 k sequences of 96 subjects taken at different vertical views, varying from about 0$^{\circ }$to 80$^{\circ }$. Furthermore, we evaluate the effectiveness of several state-of-the-art appearance-based and skeleton-based models using our dataset and establish comprehensive baselines. Our results demonstrate that the dataset is challenging and presents significant opportunities to improve existing gait recognition methods. Moreover, we propose a new method called Vertical Distillation, which is based on the feature distillation across different vertical views. Our proposed method substantially outperforms the state-of-the-art models on DroneGait at high vertical views. Cross-vertical-view and cross-domain experiments are also made to explain the importance of gait recognition at high vertical views. Furthermore, we analyze the differences between gait recognition at different vertical views using heatmap visualization techniques. We will make our dataset and code publicly available upon acceptance.
Aoqi Li, Saihui Hou, Qingyuan Cai, Yongzhen Huang
IEEE Trans. Multim.2
2024 GaitParsing: Human Semantic Parsing for Gait Recognition
abstract
Gait recognition is a soft biotechnology to identify pedestrians observed from different camera views based on specific walking patterns. However, various dressing and wearing conditions bring great challenges to realistic gait recognition. Most existing methods take holistic gait silhouette as input and focus on local areas through horizontal strip division or attention map. We consider that this processing may contain mixed or incomplete information about multiple body parts so that gait information is misused or underutilized. In this paper, we propose a parsing-guided framework for gait recognition, namedGaitParsing, which explores human semantic parsing to dissect human body into a set of specific and complete body parts. Correspondingly, a simple yet effective dual-branch feature extraction network is adopted to process holistic gait and distinct body parts. To maximize the use of highly discriminated gait frames, we propose a self-occlusion frame assessment to measure the self-occlusion in a gait sequence. Since there is no human parsing modality in current gait datasets, we further develop a general human parsing pipeline specifically tailored for gait datasets. This single training enables widespread application across various gait datasets. Extensive experiments with ablation analyses demonstrate competitive performance even in the most challenging conditions, e.g., Cloth-Changing (CC+5.9%). Especially, It is gratifying to see that our model can be easily applied to existing methods and significantly outperform the original architecture, even without much modification.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Multim.2
2023 OpenGait: Revisiting Gait Recognition Toward Better Practicality
abstract
Gait recognition is one of the most critical long-distance identification technologies and increasingly gains popularity in both research and industry communities. Despite the significant progress made in indoor datasets, much evidence shows that gait recognition techniques perform poorly in the wild. More importantly, we also find that some conclusions drawn from indoor datasets cannot be generalized to real applications. Therefore, the primary goal of this paper is to present a comprehensive benchmark study for better practicality rather than only a particular model for better performance. To this end, we first develop a flexible and efficient gait recognition codebase named OpenGait. Based on OpenGait, we deeply revisit the recent development of gait recognition by re-conducting the ablative experiments. Encouragingly, we detect some unperfect parts of certain prior woks, as well as new insights. Inspired by these discoveries, we develop a structurally simple, empirically powerful, and practically robust baseline model, Gait-Base. Experimentally, we comprehensively compare Gait-Base with many current gait recognition methods on multiple public datasets, and the results reflect that GaitBase achieves significantly strong performance in most cases regardless of indoor or outdoor situations. Code is available at https://github.com/ShiqiYu/OpenGait.
Chao Fan 0001, Chuanfu Shen, Saihui Hou, Yongzhen Huang, Shiqi Yu 0001
CVPR4
2023 An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions
abstract
The target of person re-identification (ReID) and gait recognition is consistent, that is to match the target pedestrian under surveillance cameras. For the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is often researched under controlled conditions. To tackle this problem, we propose a Cloth-Changing benchmark for Person re-identification and Gait recognition (CCPG). It is a cloth-changing dataset, and there are several highlights in CCPG, (1) it provides 200 identities and over 16K sequences are captured indoors and outdoors, (2) each identity has seven different cloth-changing statuses, which is hardly seen in previous datasets, (3) RGB and silhouettes version data are both available for research purposes. Moreover, aiming to investigate the cloth-changing problem systematically, comprehensive experiments are conducted on video-based ReID and gait recognition methods. The experimental results demonstrate the superiority of ReID and gait recognition separately in different cloth-changing conditions and suggest that gait recognition is a potential solution for addressing the cloth-changing problem. Our dataset will be available at https://github.com/BNU-IVC/CCPG.
Saihui Hou, Chunjie Zhang 0001, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Yao Zhao 0001
CVPR2
2023 GPGait: Generalized Pose-based Gait Recognition
abstract
Recent works on pose-based gait recognition have demonstrated the potential of using such simple information to achieve results comparable to silhouette-based methods. However, the generalization ability of pose-based methods on different datasets is undesirably inferior to that of silhouette-based ones, which has received little attention but hinders the application of these methods in real-world scenarios. To improve the generalization ability of pose-based methods across datasets, we propose a Generalized Pose-based Gait recognition (GPGait) framework. First, a Human-Oriented Transformation (HOT) and a series of Human-Oriented Descriptors (HOD) are proposed to obtain a unified pose representation with discriminative multi-features. Then, given the slight variations in the unified representation after HOT and HOD, it becomes crucial for the network to extract local-global relationships between the keypoints. To this end, a Part-Aware Graph Convolutional Network (PAGCN) is proposed to enable efficient graph partition and local-global spatial feature extraction. Experiments on four public gait recognition datasets, CASIA-B, OUMVLP-Pose, Gait3D and GREW, show that our model demonstrates better and more stable cross-domain capabilities compared to existing skeleton-based methods, achieving comparable recognition results to silhouette-based ones. Code is available at https://github.com/BNU-IVC/FastPoseGait.
Shibei Meng, Saihui Hou, Xuecai Hu, Yongzhen Huang
ICCV3
2023 Causal Intervention for Sparse-View Gait Recognition
abstract
Gait recognition aims at identifying individuals by unique walking patterns at a long distance. However, prevailing methods suffer from a large degradation when applied to large-scale surveillance systems. We find a significant cause of this issue is that previous methods heavily rely on full-view person annotations to reduce view differences by pulling closer the anchor to positive samples from different viewpoints. But, subjects under in-the-wild scenarios usually have only a limited number of sequences from different viewpoints. As a result, the available viewpoints of each subject are sparse compared to the whole dataset, and simply minimizing intra-identity differences cannot well reducing the view differences in the whole dataset. In this work, we formulate this overlooked problem as Sparse-View Gait Recognition and provide a comprehensive analysis of it by a Structural Causal Model for causalities among latent features, view distribution, and labels. Based on our analysis, we propose a simple yet effective method that enables networks to learn a more robust representation among different views. Specifically, our method consists of two parts: 1) an effective metric learning algorithmic implementation based on the backdoor adjustment, which improves the consistency of representations among different views; 2) an unsupervised view cluster algorithm to discover and identify the most influential view contexts. We evaluate the effectiveness of our method on popular GREW, Gait3D, CASIA-B, and OU-MVLP, showing that our method consistently outperforms baselines and achieves state-of-the-art performance. The code will be available at https://github.com/wj1tr0y/GaitCSV.
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Liang Wang 0001
ACM Multimedia2
2023 LandmarkGait: Intrinsic Human Parsing for Gait Recognition
abstract
Gait recognition is an emerging biometric technology for identifying pedestrians based on their unique walking patterns. In past gait recognition, global-based methods are inadequate to meet the growing demand for accuracy, while commonly used part-based methods provided coarse and inaccurate feature representation for specific body parts. Human parsing appears to be a better option for accurately representing specific and complete body parts in gait recognition. However, its practical application in gait recognition is often hindered by missing RGB modality, lack of annotated body parts, and difficulty in balancing parsing quantity and quality. To address this issue, we propose LandmarkGait, an accessible and alternative parsing-based solution for gait recognition. LandmarkGait introduces an unsupervised landmark discovery network to transform the dense silhouette into a finite set of landmarks with remarkable consistency across various conditions. By grouping landmarks subsets corresponding to distinct body part regions, following a reconstruction task and further refinement from high-quality input silhouettes, we can directly obtain fine-grained parsing results from original binary silhouettes in an unsupervised manner. Moreover, we also develop a multi-scale feature extractor that simultaneously captures global and parsing feature representations based on the integrity and flexibility of specific body parts. Extensive experiments demonstrate that our LandmarkGait can extract more stable features and exhibit significant performance improvement under all conditions, especially in various dressing conditions. Code is available at https://github.com/wzb-bupt/LandmarkGait.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu
ACM Multimedia2
2023 Learning Gait Representation From Massive Unlabelled Walking Videos: A Benchmark
abstract
Gait depicts individuals' unique and distinguishing walking patterns and has become one of the most promising biometric features for human identification. As a fine-grained recognition task, gait recognition is easily affected by many factors and usually requires a large amount of completely annotated data that is costly and insatiable. This paper proposes a large-scale self-supervised benchmark for gait recognition with contrastive learning, aiming to learn the general gait representation from massive unlabelled walking videos for practical applications via offering informative walking priors and diverse real-world variations. Specifically, we collect a large-scale unlabelled gait dataset GaitLU-1M consisting of 1.02M walking sequences and propose a conceptually simple yet empirically powerful baseline model GaitSSB. Experimentally, we evaluate the pre-trained model on four widely-used gait benchmarks, CASIA-B, OU-MVLP, GREW and Gait3D with or without transfer learning. The unsupervised results are comparable to or even better than the early model-based and GEI-based methods. After transfer learning, GaitSSB outperforms existing methods by a large margin in most cases, and also showcases the superior generalization capacity. Further experiments indicate that the pre-training can save about 50% and 80% annotation costs of GREW and Gait3D. Theoretically, we discuss the critical issues for gait-specific contrastive framework and present some insights for further study. As far as we know, GaitLU-1M is the first large-scale unlabelled gait dataset, and GaitSSB is the first method that achieves remarkable unsupervised results on the aforementioned benchmarks.
Chao Fan 0001, Saihui Hou, Jilong Wang 0010, Yongzhen Huang, Shiqi Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Gait Quality Aware Network: Toward the Interpretability of Silhouette-Based Gait Recognition
abstract
Gait recognition receives increasing attention since it can be conducted at a long distance in a nonintrusive way and applied to the condition of changing clothes. Most existing methods take the silhouettes of gait sequences as the input and learn a unified representation from multiple silhouettes to match probe and gallery. However, these models are all faced with the lack of interpretability, e.g., it is not clear which silhouette in a gait sequence and which part in the human body are relatively more important for recognition. In this work, we propose a gait quality aware network (GQAN) for gait recognition which explicitly assesses the quality of each silhouette and each part via two blocks: frame quality block (FQBlock) and part quality block (PQBlock). Specifically, FQBlock works in a squeeze-and-excitation style to recalibrate the features for each silhouette, and the scores of all the channels are added as frame quality indicator. PQBlock predicts a score for each part which is used to compute the weighted distance between the probe and gallery. Particularly, we propose a part quality loss (PQLoss) which enables GQAN to be trained in an end-to-end manner with only sequence-level identity annotations. This work is meaningful by moving toward the interpretability of silhouette-based gait recognition, and our method also achieves very competitive performance on CASIA-B and OUMVLP.
Saihui Hou, Xu Liu 0008, Chunshui Cao, Yongzhen Huang
IEEE Trans. Neural Networks Learn. Syst.1
2022 GaitEdge: Beyond Plain End-to-End Gait Recognition for Better Practicality
Chao Fan 0001, Saihui Hou, Chuanfu Shen, Yongzhen Huang, Shiqi Yu 0001
ECCV (5)3
2020 GaitPart: Temporal Part-Based Model for Gait Recognition
abstract
Gait recognition, applied to identify individual walking patterns in a long-distance, is one of the most promising video-based biometric technologies. At present, most gait recognition methods take the whole human body as a unit to establish the spatio-temporal representations. However, we have observed that different parts of human body possess evidently various visual appearances and movement patterns during walking. In the latest literature, employing partial features for human body description has been verified being beneficial to individual recognition. Taken above insights together, we assume that each part of human body needs its own spatio-temporal expression. Then, we propose a novel part-based model GaitPart and get two aspects effect of boosting the performance: On the one hand, Focal Convolution Layer, a new applying of convolution, is presented to enhance the fine-grained learning of the part-level spatial features. On the other hand, the Micro-motion Capture Module (MCM) is proposed and there are several parallel MCMs in the GaitPart corresponding to the pre-defined parts of the human body, respectively. It is worth mentioning that the MCM is a novel way of temporal modeling for gait task, which focuses on the short-range temporal features rather than the redundant long-range features for cycle gait. Experiments on two of the most popular public datasets, CASIA-B and OU-MVLP, richly exemplified that our method meets a new state-of-the-art on multiple standard benchmarks. The source code will be available on https://github.com/ChaoFan96/GaitPart.
Chao Fan 0001, Yunjie Peng, Chunshui Cao, Xu Liu 0008, Saihui Hou, Jiannan Chi, Yongzhen Huang, Qing Li 0015, Zhiqiang He 0002
CVPR5
2020 Gait Lateral Network: Learning Discriminative and Compact Representations for Gait Recognition
Saihui Hou, Chunshui Cao, Xu Liu 0008, Yongzhen Huang
ECCV (9)1
2019 Weighted Channel Dropout for Regularization of Deep Convolutional Neural Network
abstract
In this work, we propose a novel method named Weighted Channel Dropout (WCD) for the regularization of deep Convolutional Neural Network (CNN). Different from Dropout which randomly selects the neurons to set to zero in the fully-connected layers, WCD operates on the channels in the stack of convolutional layers. Specifically, WCD consists of two steps, i.e., Rating Channels and Selecting Channels, and three modules, i.e., Global Average Pooling, Weighted Random Selection and Random Number Generator. It filters the channels according to their activation status and can be plugged into any two consecutive layers, which unifies the original Dropout and Channel-Wise Dropout. WCD is totally parameter-free and deployed only in training phase with very slight computation cost. The network in test phase remains unchanged and thus the inference cost is not added at all. Besides, when combining with the existing networks, it requires no re-pretraining on ImageNet and thus is well-suited for the application on small datasets. Finally, WCD with VGGNet-16, ResNet-101, Inception-V3 are experimentally evaluated on multiple datasets. The extensive results demonstrate that WCD can bring consistent improvements over the baselines.
Saihui Hou, Zilei Wang
AAAI1
2019 Learning a Unified Classifier Incrementally via Rebalancing
abstract
Conventionally, deep neural networks are trained offline, relying on a large dataset prepared in advance. This paradigm is often challenged in real-world applications, e.g. online services that involve continuous streams of incoming data. Recently, incremental learning receives increasing attention, and is considered as a promising solution to the practical challenges mentioned above. However, it has been observed that incremental learning is subject to a fundamental difficulty -- catastrophic forgetting, namely adapting a model to new data often results in severe performance degradation on previous tasks or classes. Our study reveals that the imbalance between previous and new data is a crucial cause to this problem. In this work, we develop a new framework for incrementally learning a unified classifier, e.g. a classifier that treats both old and new classes uniformly. Specifically, we incorporate three components, cosine normalization, less-forget constraint, and inter-class separation, to mitigate the adverse effects of the imbalance. Experiments show that the proposed method can effectively rebalance the training process, thus obtaining superior performance compared to the existing methods. On CIFAR-100 and ImageNet, our method can reduce the classification errors by more than 6% and 13% respectively, under the incremental setting of 10 phases.
Saihui Hou, Chen Change Loy, Zilei Wang, Dahua Lin
CVPR1
2018 Lifelong Learning via Progressive Distillation and Retrospection
Saihui Hou, Chen Change Loy, Zilei Wang, Dahua Lin
ECCV (3)1
2018 Towards Human-Level License Plate Recognition
Jiafan Zhuang, Saihui Hou, Zilei Wang, Zhengjun Zha
ECCV (3)2
2018 Object detection via deeply exploiting depth information
Saihui Hou, Zilei Wang, Feng Wu 0001
Neurocomputing1
2017 VegFru: A Domain-Specific Dataset for Fine-Grained Visual Categorization
abstract
In this paper, we propose a novel domain-specific dataset named VegFru for fine-grained visual categorization (FGVC). While the existing datasets for FGVC are mainly focused on animal breeds or man-made objects with limited labelled data, VegFru is a larger dataset consisting of vegetables and fruits which are closely associated with the daily life of everyone. Aiming at domestic cooking and food management, VegFru categorizes vegetables and fruits according to their eating characteristics, and each image contains at least one edible part of vegetables or fruits with the same cooking usage. Particularly, all the images are labelled hierarchically. The current version covers vegetables and fruits of 25 upper-level categories and 292 subordinate classes. And it contains more than 160,000 images in total and at least 200 images for each subordinate class. Accompanying the dataset, we also propose an effective framework called HybridNet to exploit the label hierarchy for FGVC. Specifically, multiple granularity features are first extracted by dealing with the hierarchical labels separately. And then they are fused through explicit operation, e.g., Compact Bilinear Pooling, to form a unified representation for the ultimate recognition. The experimental results on the novel VegFru, the public FGVC-Aircraft and CUB-200-2011 indicate that HybridNet achieves one of the top performance on these datasets. The dataset and code are available at https://github.com/ustc-vim/vegfru.
Saihui Hou, Yushan Feng, Zilei Wang
ICCV1
2017 DualNet: Learn Complementary Features for Image Recognition
abstract
In this work we propose a novel framework named Dual-Net aiming at learning more accurate representation for image recognition. Here two parallel neural networks are coordinated to learn complementary features and thus a wider network is constructed. Specifically, we logically divide an end-to-end deep convolutional neural network into two functional parts, i.e., feature extractor and image classifier. The extractors of two subnetworks are placed side by side, which exactly form the feature extractor of DualNet. Then the two-stream features are aggregated to the final classifier for overall classification, while two auxiliary classifiers are appended behind the feature extractor of each subnetwork to make the separately learned features discriminative alone. The complementary constraint is imposed by weighting the three classifiers, which is indeed the key of DualNet. The corresponding training strategy is also proposed, consisting of iterative training and joint fine tuning, to make the two subnetworks cooperate well with each other. Finally, DualNet based on the well-known CaffeNet, VGGNet, NIN and ResNet are thoroughly investigated and experimentally evaluated on multiple datasets including CIFAR-100, Stanford Dogs and UEC FOOD-100. The results demonstrate that DualNet can really help learn more accurate image representation, and thus result in higher accuracy for recognition. In particular, the performance on CIFAR-100 is state-of-the-art compared to the recent works.
Saihui Hou, Xu Liu 0008, Zilei Wang
ICCV1
2017 Background-Driven Salient Object Detection
abstract
The background information is a significant prior for salient object detection, especially when images contain cluttered background and diverse object parts. In this paper, we propose a background-driven salient object detection (BD-SOD) method to more comprehensively exploit the background prior, aiming at generating more accurate and robust salient maps. To be specific, we first exploit the background prior to conduct the saliency estimation, i.e., computing the regional saliency values. In this stage, the background prior is utilized in threefold: restricting the reference regions to only the background regions, weighting the contribution of reference regions, and leveraging the importance of different features. Benefiting from such an explicit utilization, the proposed model can greatly mitigate the negative interference of the cluttered background and diverse object parts. We then embed the background prior into the optimization graph for saliency refinement. Specifically, two virtual supernodes (representing the background and foreground, respectively) are introduced with extra connections, and the nonlocal feature connections between similar regions are also set up. These connections enhance the power of optimization graph to alleviate the perturbations from diverse parts, and thus help to achieve the uniformity of saliency values. Finally, we provide systematical studies to investigate the effectiveness of the proposed BD-SOD in exploiting the valuable background prior. Experimental results on multiple public benchmark datasets, including MSRA-1000, THUS-10000, PASCAL-S, and ECSSD, clearly show that BD-SOD consistently outperforms the well-established baselines and achieves state-of-the-art performance.
Zilei Wang, Dao Xiang, Saihui Hou, Feng Wu 0001
IEEE Trans. Multim.3