EDBT 2026 Demo / reviewers in the wild / expert
Yongzhen Huang
dblp:11/753
· DBLP profile ↗
115ranked-venue papers
10as first author
56since 2021 · last 2026
0000-0003-4389-9805ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 7 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 6 first-author · 35 since 2021Security and privacy · 12 · 11 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gait Transformer: End-to-End Transformer Backbone for Gait RecognitionabstractGait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modality, spatial-temporal dynamics, fine-grained motion cues, and limited training data. In this paper, we propose Gait Transformer (GaT), an end-to-end Transformer backbone specifically tailored for silhouette-based gait recognition. GaT introduces three key components: (1) a hybrid patch embedding module that combines convolutional stems with group-batch normalization to enhance structural preservation; (2) a decomposed token mixer that explicitly models both short-range and long-range dependencies across spatial-temporal dimensions; and (3) a hybrid positional encoding strategy that integrates absolute, relative, and rotary embeddings to support efficient training under data scarcity. Without relying on any pretraining, GaT achieves state-of-the-art performance on Gait3D, GREW, and CCGR-MINI. Saihui Hou, Wenpeng Lang, Jilong Wang 0010, Yan Huang 0008, Liang Wang 0001, Yongzhen Huang |
AAAI | 6 |
| 2026 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks (Abstract Reprint)abstractDeep learning-based face recognition models are vulnerable to adversarial attacks. In contrast to general noises, the presence of imperceptible adversarial noises can lead to catastrophic errors in deep face recognition models. The primary difference between adversarial noise and general noise lies in its specificity. Adversarial attack methods give rise to noises tailored to the characteristics of the individual image and recognition model at hand. Diverse samples and recognition models can engender specific adversarial noise patterns, which pose significant challenges for adversarial defense. Addressing this challenge in the realm of face recognition presents a more formidable endeavor due to the inherent nature of face recognition as an open set task. In order to tackle this challenge, it is imperative to employ customized processing for each individual input sample. Drawing inspiration from the biological immune system, which can identify and respond to various threats, this paper aims to create an artificial immune system to provide adversarial defense for face recognition. The proposed defense model incorporates the principles of antibody cloning, mutation, selection, and memory mechanisms to generate a distinct antibody for each input sample, wherein the term antibody refers to a specialized noise removal manner. Furthermore, we introduce a self-supervised adversarial training mechanism that serves as a simulated rehearsal of immune system invasions. Extensive experimental results demonstrate the efficacy of the proposed method, surpassing state-of-the-art adversarial defense methods. Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
AAAI | 4 |
| 2026 | Audio-Visual Feature Disentanglement and Fusion Network for Automatic Depression Severity PredictionabstractIn order to achieve early screening and assist clinical decision-making, automatic depression assessment based on multimodal data are highly anticipated. However, the existed methods often suffer from semantic gap and information redundancy due to heterogeneity among modalities. To address this challenge, this paper investigates a novel Feature Disentanglement and Fusion Network (FDFNet) for predicting depression severity from audio-visual cues. Firstly, we design the shared and private encoders to disentangle modality-shared and modalityprivate representations. The former representation that acquires joint information is subjected by similarity constraints between modalities to ensure their distributions as close as possible. The latter that can capture unique features of each modality is restrained by independence constraints for keeping their distributions distinct. The decoder is then developed to reconstruct unimodal representation with constraints to minimize information loss. Finally, an efficient fusion strategy through addition and concatenation is ultilized for aggregating information. Experimental results on four benchmark datasets demonstrate that the proposed FDFNet consistently outperforms several stateof-the-art methods, with the competitive MAE/RMSE values of 6.22/7.58 on AVEC2013, 5.21/6.49 on AVEC2014, 4.25/5.34 on DAIC-WOZ, and 4.41/5.10 on E-DAIC, indicating that multimodal deep learning based on audio-visual is an attractive solution for objectively evaluating the depression severity. Zhuhong Shao, Rongyin Qin, Yongzhen Huang, Peipeng Liang, Yinan Jiang, Yanhe Deng, Xiaohui Tan |
IEEE Trans. Affect. Comput. | 4 |
| 2026 | Spatio-Temporal Multi-Granularity for Skeleton-Based Depression Risk RecognitionabstractAs the prevalence of depression continues to rise, the timely and accurate recognition of its early signs is crucial for effective prevention and intervention. However, current clinical diagnostic methods are limited by the absence of objective biomarkers and inefficiencies in early recognition. Recent research has revealed a significant correlation between gait patterns and depression risk, suggesting that gait analysis could serve as a promising tool for early diagnosis. Depression-associated gait characteristics are defined by two key aspects: (1) they are dynamic, reflecting temporal abnormalities in movement, and (2) they manifest across both localized body regions and broader global movement patterns of the body. Based on these insights, we propose a novel Spatio-temporal Multi-granularity Network (STM-Net) for depression risk recognition. In the temporal domain, we present a Multi-grain Temporal Focus (MTF) module, designed to capture the rich dynamic temporal information embedded in the gait cycle of individuals with depression. In the spatial domain, we introduce a Multi-grain Spatial Focus (MSF) module, which effectively captures spatial features and their interactions in depression-related body regions through joint-level and part-level attention mechanisms. Extensive experimental results demonstrate that STM-Net achieves state-of-the-art performance on a large open-source dataset. Xuecai Hu, Li Yao 0002, Yongzhen Huang |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | RA-GAR: A Richly Annotated Benchmark for Gait Attribute RecognitionabstractGait attracts growing interest from researchers due to its advantages as a non-invasive and non-cooperative biometric feature. Current gait-based attribute recognition methods primarily focus on estimating attributes such as gender, age, and emotions. However, there is insufficient attention to diverse gait attributes in various covariate scenarios. In this paper, we design and collect a Richly Annotated benchmark for 15 gait attributes, named RA-GAR, comprising data from 533 individuals with over 120,000 sequences. To our knowledge, RA-GAR represents the largest and most diverse benchmark of gait attributes currently available. Furthermore, to fully leverage the semantic information and enhance attribute-specific local perception, we propose a two-stage CLIP-based method for Gait Attribute Recognition, named CLIP-GAR. Experiments on the RA-GAR and MA-Gait datasets demonstrate the effectiveness of CLIP-GAR, showing significant improvements in mean accuracy and F1 score. Chenye Wang, Saihui Hou, Aoqi Li, Qingyuan Cai, Yongzhen Huang |
AAAI | 5 |
| 2025 | Bridging Gait Recognition and Large Language Models Sequence ModelingabstractGait sequences exhibit sequential structures and contextual relationships similar to those in natural language, where each element—whether a word or a gait step—is connected to its predecessors and successors. This similarity enables the transformation of gait sequences into "texts" containing identity-related information. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized for gait sequence modeling to enhance gait recognition performance. Leveraging these insights, we make a pioneering effort to apply LLMs to gait recognition, which we refer to as GaitLLM. Specifically, we propose the Gait-to-Language (G2L) module, which converts gait sequences into a textual format suitable for LLMs, and the Language-to-Gait (L2G) module, which maps the LLM’s output back to the gait feature space, thereby bridging the gap between LLM outputs and gait recognition. Notably, GaitLLM leverages the powerful modeling capabilities of LLMs without relying on complex architectural designs, improving gait recognition performance with only a small number of trainable parameters. Our method achieves state-of-the-art results on four popular gait datasets—SUSTech1K, CCPG, Gait3D, and GREW—demonstrating the effectiveness of applying LLMs in this domain. This work highlights the potential of LLMs to significantly enhance gait recognition, paving the way for future research and practical applications. Shaopeng Yang, Jilong Wang 0010, Saihui Hou, Xu Liu 0008, Chunshui Cao, Liang Wang 0001, Yongzhen Huang |
CVPR | 7 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 33 |
| 2025 | OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better GeneralizationabstractThis paper addresses the challenge of animal re-identification, an emerging field that shares similarities with person re-identification but presents unique complexities due to the diverse species, environments and poses. To facilitate research in this domain, we introduce OpenAnimals, a flexible and extensible codebase designed specifically for animal re-identification. We conduct a comprehensive study by revisiting several state-of-the-art person re-identification methods, including BoT, AGW, SBS, and MGN, and evaluate their effectiveness on animal re-identification benchmarks such as HyenaID, LeopardID, SeaTurtleID, and WhaleSharkID. Our findings reveal that while some techniques generalize well, many do not, underscoring the significant differences between the two tasks. To bridge this gap, we propose ARBase, a strong \textbf{Base} model tailored for \textbf{A}nimal \textbf{R}e-identification, which incorporates insights from extensive experiments and introduces simple yet effective animal-oriented designs. Experiments demonstrate that ARBase consistently outperforms existing baselines, achieving state-of-the-art performance across various benchmarks. Saihui Hou, Panjian Huang, Zengbin Wang, Yuan Liu 0043, Man Zhang 0005, Yongzhen Huang |
ICCV | 7 |
| 2025 | Learning a Unified Template for Gait Recognitionabstract"What I cannot create, I do not understand."Human wisdom reveals that creation is one of the highest forms of learning. For example, Diffusion Models have demonstrated remarkable semantic structure and memory in image generation, understanding, and restoration, which intuitively benefits representation learning. However, current gait networks rarely embrace this perspective, relying primarily on learning by contrasting gait samples under varying complex conditions, leading to semantic inconsistency and uniformity issues. To address these issues, we propose Origins with generative capabilities whose underlying philosophy is that different entities are generated from a unified template, inherently regularizing gait representations within a consistent and diverse semantic space to capture accurate gait differences. Admittedly, learning this unified template is exceedingly challenging, as it requires the comprehensiveness of the template to encompass gait representations with various conditions. Inspired by Diffusion Models, Origins diffuses the unified template into timestep templates for gait generative learning, and meanwhile transfers the unified template for gait representation learning. Especially, gait generative and representation learning serve as a unified framework for end-to-end joint training. Extensive experiments on CASIA-B, CCPG,SUSTech1K, Gait3D, GREW and CCGR-MINI demonstrate that Origins performs unified generative and representation learning, achieving superior performance. Panjian Huang, Saihui Hou, Junzhou Huang, Yongzhen Huang |
ICCV | 4 |
| 2025 | OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningabstractImage Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements, existing datasets often lack breadth and depth, limiting their applicability in complex and dynamic environments: (1) from a breadth perspective, current datasets are constrained to limited variations of objects in specific scenes, and (2) from a depth perspective, prior benchmarks often provide overly simplistic descriptions. To address these challenges, we introduce OmniDiff, a comprehensive dataset comprising 324 diverse scenarios-spanning real-world complex environments and 3D synthetic settings-with fine-grained human annotations averaging 60 words in length and covering 12 distinct change types. Building on this foundation, we propose M$^3$Diff, a MultiModal large language model enhanced by a plug-and-play Multi-scale Differential Perception (MDP) module. This module improves the model's ability to accurately identify and describe inter-image differences while maintaining the foundational model's generalization capabilities. With the addition of the OmniDiff dataset, M$^3$Diff achieves state-of-the-art performance across multiple benchmarks, including Spot-the-Diff, IEdit, CLEVR-Change, CLEVR-DC, and OmniDiff, demonstrating significant improvements in cross-scenario difference recognition accuracy compared to existing methods. The dataset, code, and models will be made publicly available to support further research. Yuan Liu 0043, Saihui Hou, Saijie Hou, Jiabao Du, Shibei Meng, Yongzhen Huang |
ICCV | 6 |
| 2025 | Gait: Exploring X Modality for Generalized Gait Recognition
Zengbin Wang, Saihui Hou, Junjie Li 0002, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Siye Wang, Man Zhang 0005 |
ICCV | 6 |
| 2025 | Beyond Sparse Keypoints: Dense Pose Modeling for Robust Gait RecognitionabstractGait recognition has emerged as a promising biometric technology due to its ability to operate at a distance without subject cooperation. While pose-based methods offer advantages over appearance-based approaches in robustness and interpretability, their performance has been limited by the sparse keypoint representations of current pose estimation frameworks. We identify two critical limitations: (1) incomplete motion representation due to insufficient keypoints for dynamic body parts, and (2) lack of shape information from minimal skeleton points. This paper presents DPGait, a novel framework that addresses these challenges through innovations in both upstream processing and downstream modeling. First, we enhance pose estimation by extending the standard COCO keypoint format with additional motion-sensitive points and shape-descriptive keypoints inspired by human mesh estimation. Second, we propose a divide-and-conquer modeling strategy that processes dense keypoints through group convolution with cross-group attention, coupled with multi-granularity supervision for improved training. Our comprehensive experiments demonstrate state-of-the-art performance in pose-based gait recognition, achieving 85.8% rank-1 accuracy on SUSTech1K-surpassing leading silhouette-based methods for the first time. The results validate that dense pose representation combined with our novel modeling approach significantly advances the field of gait recognition. Wenpeng Lang, Saihui Hou, Yongzhen Huang |
ACM Multimedia | 3 |
| 2025 | Seeing from Magic Mirror: Contrastive Learning from Reconstruction for Pose-based Gait RecognitionabstractWhile recent advancements in supervised gait recognition have yielded promising results, these approaches rely heavily on annotated walking data, limiting their generalizability to complex environments. This paper presents a self-supervised gait recognition framework using human poses as input to address this challenge, focusing on high-quality pretrained data and self-supervised learning strategies. We first introduce StreamGait, a large-scale, unlabelled dataset that captures in-the-wild distributions of walking sequences. This dataset is curated from Internet livestreams across diverse geographic and environmental scenarios, reflecting variations in real-world camera angles, weather, and pedestrian behavior. Our framework, MirrorGait, conducts self-supervised learning by integration with 2D-to-3D pose reconstruction to synthesize multi-view perspectives for effective 3D-aware contrastive learning. With specific designs of temporal position embedding and gait partition head on a Transformer backbone, the encoder can readily adapt to the periodic and fine-grained nature of gait. Extensive experiments on three widely used gait datasets, Gait3D, GREW, and OUMVLP-Pose, demonstrate that our method, with minimal fine-tuning on the pretrained model, achieves state-of-the-art performance among pose-based gait recognition approaches. The dataset, code, and models are available at https://github.com/BNU-IVC/StreamGait. Shibei Meng, Saihui Hou, Xuecai Hu, Junzhou Huang, Yongzhen Huang |
ACM Multimedia | 6 |
| 2025 | Vocabulary-Guided Gait RecognitionabstractWhat is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points. However, the considerations remain vague for humans to comprehend truly. In this work, we introduce a novel paradigm Vocabulary-Guided Gait Recognition, dubbed Gait-World, which attempts to explore gait concepts through human vocabularies with Vision-Language Models (VLMs). Despite VLMs have achieved the remarkable progress in various vision tasks, the cognitive capability regarding gait modalities remains limited. The success element in Gait-World is the proper vocabulary prompt where this paradigm carefully selects gait cycle actions as Vocabulary Base, bridging the gait and vocabulary feature spaces and further promoting human understanding for the gait. How to extract gait features? Although previous gait networks have made significant progress, learning solely from gait modalities on limited gait databases makes it difficult to learn robust gait features for practicality. Therefore, we propose the first Gait-World model, dubbed $\alpha$-Gait, which guides the gait network learning with universal vocabulary knowledge from VLMs. However, due to the heterogeneity of the modalities, directly integrating vocabulary and gait features is highly challenging as they reside in different embedding spaces. To address the issues, $\alpha$-Gait designs Vocabulary Relation Mapper and Gait Fine-grained Detector to map and establish vocabulary relations in the gait space for detecting corresponding gait features. Extensive experiments on CASIA-B, CCPG, SUSTech1K, Gait3D and GREW reveal the potential value and research directions of vocabulary information from VLMs in the gait field. Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
NeurIPS | 5 |
| 2025 | Multimodal depression recognition based on gait and rating scale
Xuecai Hu, Yongzhen Huang |
Expert Syst. Appl. | 5 |
| 2025 | Edge-Oriented Adversarial Attack for Deep Gait Recognition
Saihui Hou, Zengbin Wang, Man Zhang 0005, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
Int. J. Comput. Vis. | 6 |
| 2025 | OpenGait: A Comprehensive Benchmark Study for Gait Recognition Toward Better PracticalityabstractGait recognition, a rapidly advancing vision technology for person identification from a distance, has made significant strides in indoor settings. However, evidence suggests that existing methods often yield unsatisfactory results when applied to newly released real-world gait datasets. Furthermore, conclusions drawn from indoor gait datasets may not easily generalize to outdoor ones. Therefore, the primary goal of this paper is to present a comprehensive benchmark study aimed at improving practicality rather than solely focusing on enhancing performance. To this end, we developed OpenGait, a flexible and efficient gait recognition platform. Using OpenGait, we conducted in-depth ablation experiments to revisit recent developments in gait recognition. Surprisingly, we detected some imperfect parts of some prior methods and thereby uncovered several critical yet previously neglected insights. These findings led us to develop three structurally simple yet empirically powerful and practically robust baseline models: DeepGaitV2, SkeletonGait, and SkeletonGait++, which represent the appearance-based, model-based, and multi-modal methodologies for gait pattern description, respectively. In addition to achieving state-of-the-art performance, our careful exploration provides new perspectives on the modeling experience of deep gait models and the representational capacity of typical gait modalities. In the end, we discuss the key trends and challenges in current gait recognition, aiming to inspire further advancements towards better practicality. Chao Fan 0001, Saihui Hou, Chuanfu Shen, Jingzhe Ma, Dongyang Jin, Yongzhen Huang, Shiqi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | From FastPoseGait to GPGait++: Bridging the Past and Future for Pose-Based Gait RecognitionabstractRecent studies on pose-based gait recognition have underscored the potential of utilizing such fundamental data to achieve superior outcomes. Nonetheless, the development of current pose-based methods faces significant obstacles due to several critical issues: (1) Misaligned Settings, which results in a lack of thorough and unbiased comparative analysis. (2) Inferior Performance, which causes diminished focus on pose-based gait representations. (3) Limited Generalization, which hinders the effective application in real-world scenarios. Focused on tackling the aforementioned challenges, our study introduces a comprehensive benchmark and a versatile approach to bridge the past and future for pose-based gait recognition. First, we revisit previous pose-based methods and make great efforts to establish a unified framework, FastPoseGait, aiming at a fair and comprehensive comparison investigation with consistent experimental settings and a more stable training process. Then, within this framework, we propose GPGait++, a generalized pose-based gait recognition method featuring a human-oriented input and part-aware modeling, intended to enhance the generalization ability and discriminative power across diverse environments and camera viewpoints. Experiments on six public gait recognition datasets reveal that our unified framework significantly enhances the performance of previous approaches, and GPGait++ exhibits state-of-the-art cross-domain capabilities compared to existing pose-based methods, marking a significant advancement in the field of pose-based gait recognition. Shibei Meng, Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | GaitC3I: Robust Cross-Covariate Gait Recognition via Causal InterventionabstractCross-covariate gait recognition aims to analyze a pedestrian’s gait to extract an identity representation that is invariant across varying covariates. However, prevailing methods that have achieved good results on controlled in-the-lab datasets often perform poorly on realistic datasets. In this work, we find a significant cause is that the widely used pairwise metric learning paradigm cannot correctly handle the relationship between samples from different covariate conditions. Even worse, it may yield harmful signals that inadvertently mislead models to focus on covariate-related features, particularly when covariate distributions vary across subjects. To address this issue, we propose a Cross-Covariate Causal Intervention (GaitC3I) framework, a unified causality-inspired approach aimed at enhancing the robustness of gait recognition across diverse conditions. Specifically, our method consists of two parts: 1) an effective causal intervention metric learning paradigm based on backdoor adjustment, which strategically mitigates spurious correlations induced by covariates, thus ensuring a more invariant gait representation; and 2) an annotation-free selection strategy that progressively matches each positive sample with negative samples from similar covariate conditions at various granularities. We demonstrate the effectiveness of our GaitC3I through extensive evaluation on six popular gait datasets-Gait3D, GREW, OUMVLP, CASIA-B, CCPG, and CCGR-achieving substantial improvements. Our method not only outperforms existing state-of-the-art models but also provides a systematic solution to remove the spurious correlations in gait recognition. Jilong Wang 0010, Saihui Hou, Xianda Guo, Yan Huang 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | GaitAsset: In Defense of Regarding Gait as a SetabstractIn the field of gait recognition, regarding gait as a set has emerged as a seminal approach, notably eliminating the dependence on template-based input. Although set-based methods offer notable advantages, such as insensitivity to frame order permutations and robustness to varying frame counts, their performance has consistently lagged behind that of sequence-based methods in subsequent studies. In this work, we advocate for treating gait as an unordered set and argue thatthe lack of set context aggregation in frame-level feature extraction is the primary limitation hindering the full potential of set-based gait recognition. To substantiate this claim, we develop a gait-oriented self-attention module and introduce a Gating Mechanism that facilitates set context awareness for each silhouette whilepreserving the permutation-invariant property. Specifically, the context aggregation operates on diverse bins of feature maps, interleaving fine-grained shape and motion details in an almost parameter-free manner. The Gating Mechanism is employed to ensure that frame-level features are not overwhelmed by the aggregated context. Furthermore, the sampling strategy is carefully enhanced to better support set context modeling. Our research demonstrates that set-based gait recognition can achieve state-of-the-art accuracy on in-the-wild benchmarks (77.6% on Gait3D and 81.1% on GREW) while retaining its inherent advantages. Saihui Hou, Chenye Wang, Aoqi Li, Jilong Wang 0010, Liang Wang 0001, Yongzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Multimodal Mutual Learning for Unsupervised Gait RecognitionabstractThe primary challenge in unsupervised gait recognition lies in generating meaningful and diverse supervisory signals to guide representation learning. The effectiveness of such methods largely depends on the richness of the supervisory signals. Unlike previous methods that construct supervisory signals solely from a single modality, we propose a novel framework, named Multimodal Mutual Learning (M3L), that leverages the identity consistency and complementary nature of both silhouette and skeleton modalities to generate richer and more informative supervisory signals. To fully leverage the richer supervisory signals, M3L encourages mutual prediction between the silhouette and skeleton modalities, guiding the network toward modality-invariant representations. However, mutual prediction alone is hindered by the inherent modality gap, so we introduce a Multimodal Collaborative Module to explicitly bridge this gap and promote cross-modal knowledge transfer. Moreover, to make the framework practical when only one modality is available at inference, we introduce a Multimodal Disentanglement Module. Multimodal Disentanglement Module decouples the two branches and distills a shared representation, preserving the gains of multimodal training while allowing the model to maintain robust performance under single-modality conditions. Extensive experiments on four widely used gait datasets—Gait3D, GREW, CASIA-B, and SUSTech1K—demonstrate the effectiveness of our approach and highlight its potential to advance unsupervised gait recognition. Shaopeng Yang, Saihui Hou, Xu Liu 0008, Chunshui Cao, Yongzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal DenoiserabstractRecently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods decompose the task into bone length and bone direction prediction based on the human anatomical skeleton to explicitly incorporate more human body prior constraints, the performance of these methods is significantly lower than that of the SOTA diffusion-based methods. This can be attributed to the tree structure of the human skeleton. Direct application of the disentangled method could amplify the accumulation of hierarchical errors, propagating through each hierarchy. Meanwhile, the hierarchical information has not been fully explored by the previous methods. To address these problems, a Disentangled Diffusion-based 3D human Pose Estimation method with Hierarchical Spatial and Temporal Denoiser is proposed, termed DDHPose. In our approach: (1) We disentangle the 3d pose and diffuse the bone length and bone direction during the forward process of the diffusion model to effectively model the human pose prior. A disentanglement loss is proposed to supervise diffusion model learning. (2) For the reverse process, we propose Hierarchical Spatial and Temporal Denoiser (HSTDenoiser) to improve the hierarchical modelling of each joint. Our HSTDenoiser comprises two components: the Hierarchical-Related Spatial Transformer (HRST) and the Hierarchical-Related Temporal Transformer (HRTT). HRST exploits joint spatial information and the influence of the parent joint on each joint for spatial modeling, while HRTT utilizes information from both the joint and its hierarchical adjacent joints to explore the hierarchical temporal correlations among joints. Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets show that our method outperforms the SOTA disentangled-based, non-disentangled based, and probabilistic approaches by 10.0%, 2.0%, and 1.3%, respectively. Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 0002, Yongzhen Huang |
AAAI | 5 |
| 2024 | QAGait: Revisit Gait Recognition from a Quality PerspectiveabstractGait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research. However, as gait recognition rapidly advances from in-the-lab to in-the-wild scenarios, various conditions raise significant challenges for silhouette modality, including 1) unidentifiable low-quality silhouettes (abnormal segmentation, severe occlusion, or even non-human shape), and 2) identifiable but challenging silhouettes (background noise, non-standard posture, slight occlusion). To address these challenges, we revisit gait recognition pipeline and approach gait recognition from a quality perspective, namely QAGait. Specifically, we propose a series of cost-effective quality assessment strategies, including Maxmial Connect Area and Template Match to eliminate background noises and unidentifiable silhouettes, Alignment strategy to handle non-standard postures. We also propose two quality-aware loss functions to integrate silhouette quality into optimization within the embedding space. Extensive experiments demonstrate our QAGait can guarantee both gait reliability and performance enhancement. Furthermore, our quality assessment strategies can seamlessly integrate with existing gait datasets, showcasing our superiority. Code is available at https://github.com/wzb-bupt/QAGait. Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu |
AAAI | 6 |
| 2024 | POPDG: Popular 3D Dance Generation with PopDanceSetabstractGenerating dances that are both lifelike and well-aligned with music continues to be a challenging task in the cross-modal domain. This paper introduces PopDanceSet, the first dataset tailored to the preferences of young audiences, enabling the generation of aesthetically oriented dances. And it surpasses the${\it AIST}++{\it dataset}$in music genre di-versity and the intricacy and depth of dance movements. Moreover, the proposed POPDG model within the iD-DPMframework enhances dance diversity and, through the Space Augmentation Algorithm, strengthens spatial physi-cal connections between human body joints, ensuring that increased diversity does not compromise generation qual-ity. A streamlined Alignment Module is also designed to improve the temporal alignment between dance and mu-sic. Extensive experiments show that POPDG achieves SOTA results on two datasets. Furthermore, the paper also expands on current evaluation metrics. The dataset and code are available at https://github.com/Luke-Luol/POPDG. Zhenye Luo, Xuecai Hu, Yongzhen Huang, Li Yao 0002 |
CVPR | 4 |
| 2024 | Learning Visual Prompt for Gait RecognitionabstractGait, a prevalent and complex form of human motion, plays a significant role in the field of long-range pedestrian retrieval due to the unique characteristics inherent in individual motion patterns. However, gait recognition in real-world scenarios is challenging due to the limitations of capturing comprehensive cross-viewing and crossclothing data. Additionally, distractors such as occlusions, directional changes, and lingering movements further complicate the problem. The widespread application of deep learning techniques has led to the development of various potential gait recognition methods. However, these methods utilize convolutional networks to extract shared information across different views and attire conditions. Once trained, the parameters and non-linear function become constrained to fixed patterns, limiting their adaptability to various distractors in real-world scenarios. In this paper, we present a unified gait recognition framework to extract global motion patterns and develop a novel dynamic transformer to generate representative gait features. Specifically, we develop a trainable part-based prompt pool with numerous key-value pairs that can dynamically select prompt templates to incorporate into the gait sequence, thereby providing task-relevant shared knowledge information. Furthermore, we specifically design dynamic attention to extract robust motion patterns and address the length generalization issue. Extensive experiments on four widely recognized gait datasets, i.e., Gait3D, GREW, OUMVLP, and CASIA-B, reveal that the proposed method yields substantial improvements compared to current state-of-the-art approaches. Ying Fu 0001, Chunshui Cao, Saihui Hou, Yongzhen Huang, Dezhi Zheng |
CVPR | 5 |
| 2024 | Cut Out the Middleman: Revisiting Pose-Based Gait Recognition
Saihui Hou, Shibei Meng, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
ECCV (31) | 7 |
| 2024 | Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective
Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu 0008, Zhiqiang He 0002, Yongzhen Huang |
ECCV (6) | 7 |
| 2024 | Free Lunch for Gait Recognition: A Novel Relation Descriptor
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001 |
ECCV (38) | 6 |
| 2024 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2024abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To advance the algorithm development and provide fair evaluations, the International Competition on Human Identification at a Distance (HID) has been held annually since 2020, with HID 2024 marking the fifth edition. Despite increased difficulty, participants demonstrated remarkable capabilities, surpassing previous accuracy levels. This paper, co-authored by competition organizers and top participants, provides a comprehensive summary of HID 2024, including an overview of the competition, and insights into the methods employed by the top teams. Specifically, inspired by the achievements of the 5 competitions of HID, we also provide the insights for the future directions on gait recognition. Shiqi Yu 0001, Weiming Wu, Jiacong Hu, Zepeng Wang 0002, Runsheng Wang, Yunfei Ni, Yongzhen Huang, Liang Wang 0001, Md. Atiqur Rahman Ahad |
IJCB | 9 |
| 2024 | AerialGait: Bridging Aerial and Ground Views for Gait RecognitionabstractIn this work, we present AerialGait, a comprehensive dataset for aerial-ground gait recognition. This dataset comprises 82,454 sequences totaling over 10 million frames from 533 subjects, captured from both aerial and ground perspectives. To align with real-life scenarios of aerial and ground surveillance, we utilize a drone and a ground surveillance camera for data acquisition. The drone is operated at various speeds, directions, and altitudes. Meanwhile, we conduct data collection across five diverse surveillance sites to ensure a comprehensive simulation of real-world settings. AerialGait has several unique features: 1) The gait sequences exhibit significant variations in views, resolutions, and illumination across five distinct scenes. 2) It incorporates challenges of motion blur and frame discontinuity due to drone mobility. 3) The dataset reflects the domain gap caused by the view disparity between aerial and ground views, presenting a realistic challenge for drone-based gait recognition. Moreover, we perform a comprehensive analysis of existing gait recognition methods on AerialGait dataset and propose the Aerial-Ground Gait Network (AGG-Net). AGG-Net effectively learns discriminative features from aerial views by uncertainty learning and clusters features across aerial and ground views through prototype learning. Our model achieves state-of-the-art performance on both AerialGait and DroneGait datasets. Aoqi Li, Saihui Hou, Chenye Wang, Qingyuan Cai, Yongzhen Huang |
ACM Multimedia | 5 |
| 2024 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
Int. J. Comput. Vis. | 4 |
| 2024 | CrowdTrans: Learning top-down visual perception for crowd counting by transformer
Weiyu Guo, Shaopeng Yang, Yuheng Ren, Yongzhen Huang |
Neurocomputing | 4 |
| 2024 | Depression risk recognition based on gait: A benchmark
Saihui Hou, Xuecai Hu, Yongzhen Huang |
Neurocomputing | 6 |
| 2024 | Cloth-Imbalanced Gait Recognition via HallucinationabstractThe study in the gait field has rarely paid attention to the class-imbalanced learning, while the realistic data always exhibits an imbalanced distribution. The main reason lies in the difficulty of collecting the cross-clothes sequences, since the collection is usually aided by person re-identification and it is more likely to obtain the sequences for a subject wearing the same clothes. In this work, we formulate a new problem to tackle the task-specific cloth-imbalanced issue, dubbed as Cloth-Imbalanced Gait Recognition, and the training data consists of two parts denoted as head set and tail set. The sequences for a subject in head set cover the cross-clothes variation which is scarce in tail set to mimic the collection difficulty. Along with the problem formulation, we design a new method to deal with the inherent challenges, called Cross-Clothes Hallucination or CCH for short. Our method is inspired by the observation that certain directions in deep feature space correspond to meaningful semantic transformations, and it tries to generate the cross-clothes sequences for tail set referring to the cloth-changing transformation in head set. To evaluate CCH, we build two cloth-imbalanced benchmarks based on the widely-used CASIA-B and Outdoor-Gait. Extensive experiments demonstrate that CCH brings significant improvements over the baselines. Saihui Hou, Panjian Huang, Xu Liu 0008, Chunshui Cao, Yongzhen Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Adaptive Knowledge Transfer for Weak-Shot Gait RecognitionabstractMost works for cloth-changing gait recognition assume that the sequences of different clothes are accessible for each subject in the training set, which, however, is almost impossible for real applications. In practice, the collection of gait sequences is usually aided by person re-identification which is more likely to cluster the cloth-consistent sequences for a subject, and it is laborious to merge the cloth-changing clusters with the same identity. As a result, the training set is usually comprised of two subsets, i.e., a fully-annotated base set where the cloth-changing sequences are available for each subject, and a weakly-annotated wild set where the sequences of different clothes for a subject are assigned diverse labels. In this work, we formulate a problem named Weak-Shot Gait Recognition which seeks to learn discriminative features from the mixture of base set and wild set. Furthermore, we propose an effective method called Adaptive Knowledge Transfer to deal with the weak-shot issue. In particular, we define the knowledge as the ability to judge whether two sequences come from the same subject or not, and we take an adaptive way to mine the useful information from wild set. For the experimental study, we build three weak-shot benchmarks based on CASIA-B, Outdoor-Gait, and CASIA-E respectively. Extensive experiments show that our method can bring consistent improvement. For example, under the cloth-changing condition of on the weak-shot CASIA-B, our method exceeds a naïve baseline by 7.10%. Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Integral Pose Learning via Appearance Transfer for Gait RecognitionabstractGait recognition plays an important role in video surveillance and security by identifying humans based on their unique walking patterns. The existing gait recognition methods have achieved competitive accuracy with shape and motion patterns under limited-covariate conditions. However, when extreme appearance changes distort discriminative features, gait recognition yields unsatisfactory results under cross-covariate conditions. In this work, we first indicate that the integral pose in each silhouette maintains an appearance-unrelated discriminative identity. However, the monotonous appearance variables in a gait database cause gait models to have difficulty extracting integral poses. Therefore, we propose an Appearance-transferable Disentangling and Generative Network (GaitApp) to generate gait silhouettes with rich appearances and invariant poses. Specifically, GaitApp leverages multi-branch cooperation to disentangle pose features and appearance features, and transfers the appearance information from one subject to another. By simulating a person constantly changing appearances under limited-covariate conditions, downstream models enable to extract integral discriminative pose features. Extensive experiments demonstrate that our method allows representative gait models to stand at a new altitude, further promoting the exploration to cross-covariate gait recognition. All the code is available at https://github.com/Hpjhpjhs/GaitApp.git. Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Xuecai Hu, Yongzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Understanding Deep Face Representation via Attribute RecoveryabstractDeep neural networks have proven to be highly effective in the face recognition task, as they can map raw samples into a discriminative high-dimensional representation space. However, understanding this complex space proves to be challenging for human observers. In this paper, we propose a novel approach that interprets deep face recognition models via facial attributes. To achieve this, we introduce a two-stage framework that recovers attributes from the deep face representations. This framework allows us to quantitatively measure the significance of facial attributes in relation to the recognition model. Moreover, this framework enables us to generate sample-specific explanations through counterfactual methodology. These explanations are not only understandable but also quantitative. Through the proposed approach, we are able to acquire a deeper understanding of how the recognition model conceptualizes the notion of “identity” and understand the reasons behind the error decisions made by the deep models. By utilizing attributes as an interpretable interface, the proposed method marks a paradigm shift in our comprehension of deep face recognition models. It allows a complex model, obtained through gradient backpropagation, to effectively “communicate” with humans. The source code is available here, or you can visit this website:https://github.com/RenMin1991/Facial-Attribute-Recovery. Yuhao Zhu 0003, Yunlong Wang 0003, Yongzhen Huang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Gait Attribute Recognition: A New Benchmark for Learning Richer Attributes From Human Gait PatternsabstractCompared to gait recognition, Gait Attribute Recognition (GAR) is a seldom-investigated problem. However, since gait attribute recognition can provide richer and finer semantic descriptions, it is an indispensable part of building intelligent gait analysis systems. Nonetheless, the types of attributes considered in the existing datasets are very limited. This paper contributes a new benchmark dataset for gait attribute recognition named Multi-Attribute Gait (MA-Gait). Our MA-Gait contains 95 subjects recorded from 12 camera views, resulting in more than 13000 sequences, with 16 attributes labeled, including six attributes that have never been considered in the literature. Moreover, we propose a Multi-Scale Motion Encoder (MSME) to extract robust motion features, and an Attribute-Guided Feature Selection Module (AGFSM) to adaptively capture the most discriminative attribute features from static appearance features and dynamic motion features for different attributes. Our method achieves the best GAR accuracy on the new dataset. Comprehensive experiments show the effectiveness of the proposed method through both quantitative and qualitative evaluations. Xu Song, Saihui Hou, Yan Huang 0023, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Caifeng Shan |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Gait Recognition With Drones: A BenchmarkabstractGait recognition aims to obtain people's identity through body shape and walking posture. Existing gait recognition studies focus on low vertical view recognition, in which the person and the camera are nearly at the same height. Differently, in this work, we focus on gait recognition at high vertical views. To facilitate the research, we propose a new dataset named DroneGait, where the drones are used to collect the gait data. This dataset contains 22 k sequences of 96 subjects taken at different vertical views, varying from about 0$^{\circ }$to 80$^{\circ }$. Furthermore, we evaluate the effectiveness of several state-of-the-art appearance-based and skeleton-based models using our dataset and establish comprehensive baselines. Our results demonstrate that the dataset is challenging and presents significant opportunities to improve existing gait recognition methods. Moreover, we propose a new method called Vertical Distillation, which is based on the feature distillation across different vertical views. Our proposed method substantially outperforms the state-of-the-art models on DroneGait at high vertical views. Cross-vertical-view and cross-domain experiments are also made to explain the importance of gait recognition at high vertical views. Furthermore, we analyze the differences between gait recognition at different vertical views using heatmap visualization techniques. We will make our dataset and code publicly available upon acceptance. Aoqi Li, Saihui Hou, Qingyuan Cai, Yongzhen Huang |
IEEE Trans. Multim. | 5 |
| 2024 | GaitParsing: Human Semantic Parsing for Gait RecognitionabstractGait recognition is a soft biotechnology to identify pedestrians observed from different camera views based on specific walking patterns. However, various dressing and wearing conditions bring great challenges to realistic gait recognition. Most existing methods take holistic gait silhouette as input and focus on local areas through horizontal strip division or attention map. We consider that this processing may contain mixed or incomplete information about multiple body parts so that gait information is misused or underutilized. In this paper, we propose a parsing-guided framework for gait recognition, namedGaitParsing, which explores human semantic parsing to dissect human body into a set of specific and complete body parts. Correspondingly, a simple yet effective dual-branch feature extraction network is adopted to process holistic gait and distinct body parts. To maximize the use of highly discriminated gait frames, we propose a self-occlusion frame assessment to measure the self-occlusion in a gait sequence. Since there is no human parsing modality in current gait datasets, we further develop a general human parsing pipeline specifically tailored for gait datasets. This single training enables widespread application across various gait datasets. Extensive experiments with ablation analyses demonstrate competitive performance even in the most challenging conditions, e.g., Cloth-Changing (CC+5.9%). Especially, It is gratifying to see that our model can be easily applied to existing methods and significantly outperform the original architecture, even without much modification. Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang |
IEEE Trans. Multim. | 6 |
| 2023 | OpenGait: Revisiting Gait Recognition Toward Better PracticalityabstractGait recognition is one of the most critical long-distance identification technologies and increasingly gains popularity in both research and industry communities. Despite the significant progress made in indoor datasets, much evidence shows that gait recognition techniques perform poorly in the wild. More importantly, we also find that some conclusions drawn from indoor datasets cannot be generalized to real applications. Therefore, the primary goal of this paper is to present a comprehensive benchmark study for better practicality rather than only a particular model for better performance. To this end, we first develop a flexible and efficient gait recognition codebase named OpenGait. Based on OpenGait, we deeply revisit the recent development of gait recognition by re-conducting the ablative experiments. Encouragingly, we detect some unperfect parts of certain prior woks, as well as new insights. Inspired by these discoveries, we develop a structurally simple, empirically powerful, and practically robust baseline model, Gait-Base. Experimentally, we comprehensively compare Gait-Base with many current gait recognition methods on multiple public datasets, and the results reflect that GaitBase achieves significantly strong performance in most cases regardless of indoor or outdoor situations. Code is available at https://github.com/ShiqiYu/OpenGait. Chao Fan 0001, Chuanfu Shen, Saihui Hou, Yongzhen Huang, Shiqi Yu 0001 |
CVPR | 5 |
| 2023 | An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing ConditionsabstractThe target of person re-identification (ReID) and gait recognition is consistent, that is to match the target pedestrian under surveillance cameras. For the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is often researched under controlled conditions. To tackle this problem, we propose a Cloth-Changing benchmark for Person re-identification and Gait recognition (CCPG). It is a cloth-changing dataset, and there are several highlights in CCPG, (1) it provides 200 identities and over 16K sequences are captured indoors and outdoors, (2) each identity has seven different cloth-changing statuses, which is hardly seen in previous datasets, (3) RGB and silhouettes version data are both available for research purposes. Moreover, aiming to investigate the cloth-changing problem systematically, comprehensive experiments are conducted on video-based ReID and gait recognition methods. The experimental results demonstrate the superiority of ReID and gait recognition separately in different cloth-changing conditions and suggest that gait recognition is a potential solution for addressing the cloth-changing problem. Our dataset will be available at https://github.com/BNU-IVC/CCPG. Saihui Hou, Chunjie Zhang 0001, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Yao Zhao 0001 |
CVPR | 6 |
| 2023 | Dynamic Aggregated Network for Gait RecognitionabstractGait recognition is beneficial for a variety of applications, including video surveillance, crime scene investigation, and social security, to mention a few. However, gait recognition often suffers from multiple exterior factors in real scenes, such as carrying conditions, wearing overcoats, and diverse viewing angles. Recently, various deep learning-based gait recognition methods have achieved promising results, but they tend to extract one of the salient features using fixed-weighted convolutional networks, do not well consider the relationship within gait features in key regions, and ignore the aggregation of complete motion patterns. In this paper, we propose a new perspective that actual gait features include global motion patterns in multiple key regions, and each global motion pattern is composed of a series of local motion patterns. To this end, we propose a Dynamic Aggregation Network (DANet) to learn more discriminative gait features. Specifically, we create a dynamic attention mechanism between the features of neighboring pixels that not only adaptively focuses on key regions but also generates more expressive local motion patterns. In addition, we develop a selfattention mechanism to select representative local motion patterns and further learn robust global motion patterns. Extensive experiments on three popular public gait datasets, i.e., CASIA-B, OUMVLP, and Gait3D, demonstrate that the proposed method can provide substantial improvements over the current state-of-the-art methods.1 Ying Fu 0001, Dezhi Zheng, Chunshui Cao, Xuecai Hu, Yongzhen Huang |
CVPR | 6 |
| 2023 | Human Identification at a Distance: Challenges, Methods and Results on HID 2023abstractHuman Identification at a Distance (HID) is an important research area due to its importance (especially in biometrics) and inherent challenges within this domain. To mitigate some of the constraints, we have introduced the HID challenge. This paper presents an overview of the 4th International Competition on Human Identification at a Distance (HID 2023), which serves as a benchmark for evaluating various methods in the field of human identification at a distance. We have introduced a new dataset, SUSTech-Competition, engulfing a cross-domain challenge. This dataset has 859 subjects, having various variations of clothing, carrying conditions, occlusions, and view angles. With a substantial participation of 254 registered teams, HID 2023 has attracted considerable attention and yielded highly encouraging results. Notably, the top-performing teams achieved significantly good accuracies. In this paper, we provide an introduction to the competition, encompassing the dataset, experimental settings, and competition organization, as well as an analysis of the results obtained by the top teams. Additionally, we delve into the methodologies employed by these leading teams. The progress demonstrated in this competition offers an optimistic outlook on the advancements in gait recognition, highlighting its potential for robust real applications. Shiqi Yu 0001, Chenye Wang, Li Wang 0033, Qing Li 0015, Runsheng Wang, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Md. Atiqur Rahman Ahad |
IJCB | 9 |
| 2023 | GPGait: Generalized Pose-based Gait RecognitionabstractRecent works on pose-based gait recognition have demonstrated the potential of using such simple information to achieve results comparable to silhouette-based methods. However, the generalization ability of pose-based methods on different datasets is undesirably inferior to that of silhouette-based ones, which has received little attention but hinders the application of these methods in real-world scenarios. To improve the generalization ability of pose-based methods across datasets, we propose a Generalized Pose-based Gait recognition (GPGait) framework. First, a Human-Oriented Transformation (HOT) and a series of Human-Oriented Descriptors (HOD) are proposed to obtain a unified pose representation with discriminative multi-features. Then, given the slight variations in the unified representation after HOT and HOD, it becomes crucial for the network to extract local-global relationships between the keypoints. To this end, a Part-Aware Graph Convolutional Network (PAGCN) is proposed to enable efficient graph partition and local-global spatial feature extraction. Experiments on four public gait recognition datasets, CASIA-B, OUMVLP-Pose, Gait3D and GREW, show that our model demonstrates better and more stable cross-domain capabilities compared to existing skeleton-based methods, achieving comparable recognition results to silhouette-based ones. Code is available at https://github.com/BNU-IVC/FastPoseGait. Shibei Meng, Saihui Hou, Xuecai Hu, Yongzhen Huang |
ICCV | 5 |
| 2023 | Fine-grained Unsupervised Domain Adaptation for Gait RecognitionabstractGait recognition has emerged as a promising technique for the long-range retrieval of pedestrians, providing numerous advantages such as accurate identification in challenging conditions and non-intrusiveness, making it highly desirable for improving public safety and security. However, the high cost of labeling datasets, which is a prerequisite for most existing fully supervised approaches, poses a significant obstacle to the development of gait recognition. Recently, some unsupervised methods for gait recognition have shown promising results. However, these methods mainly rely on a fine-tuning approach that does not sufficiently consider the relationship between source and target domains, leading to the catastrophic forgetting of source domain knowledge. This paper presents a novel perspective that adjacent-view sequences exhibit overlapping views, which can be leveraged by the network to gradually attain cross-view and cross-dressing capabilities without pre-training on the labeled source domain. Specifically, we propose a fine-grained Unsupervised Domain Adaptation (UDA) framework that iteratively alternates between two stages. The initial stage involves offline clustering, which transfers knowledge from the labeled source domain to the unlabeled target domain and adaptively generates pseudo-labels according to the expressiveness of each part. Subsequently, the second stage encompasses online training, which further achieves cross-dressing capabilities by continuously learning to distinguish numerous features of source and target domains. The effectiveness of the proposed method is demonstrated through extensive experiments conducted on widely-used public gait datasets. Ying Fu 0001, Dezhi Zheng, Yunjie Peng, Chunshui Cao, Yongzhen Huang |
ICCV | 6 |
| 2023 | Causal Intervention for Sparse-View Gait RecognitionabstractGait recognition aims at identifying individuals by unique walking patterns at a long distance. However, prevailing methods suffer from a large degradation when applied to large-scale surveillance systems. We find a significant cause of this issue is that previous methods heavily rely on full-view person annotations to reduce view differences by pulling closer the anchor to positive samples from different viewpoints. But, subjects under in-the-wild scenarios usually have only a limited number of sequences from different viewpoints. As a result, the available viewpoints of each subject are sparse compared to the whole dataset, and simply minimizing intra-identity differences cannot well reducing the view differences in the whole dataset. In this work, we formulate this overlooked problem as Sparse-View Gait Recognition and provide a comprehensive analysis of it by a Structural Causal Model for causalities among latent features, view distribution, and labels. Based on our analysis, we propose a simple yet effective method that enables networks to learn a more robust representation among different views. Specifically, our method consists of two parts: 1) an effective metric learning algorithmic implementation based on the backdoor adjustment, which improves the consistency of representations among different views; 2) an unsupervised view cluster algorithm to discover and identify the most influential view contexts. We evaluate the effectiveness of our method on popular GREW, Gait3D, CASIA-B, and OU-MVLP, showing that our method consistently outperforms baselines and achieves state-of-the-art performance. The code will be available at https://github.com/wj1tr0y/GaitCSV. Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Liang Wang 0001 |
ACM Multimedia | 6 |
| 2023 | LandmarkGait: Intrinsic Human Parsing for Gait RecognitionabstractGait recognition is an emerging biometric technology for identifying pedestrians based on their unique walking patterns. In past gait recognition, global-based methods are inadequate to meet the growing demand for accuracy, while commonly used part-based methods provided coarse and inaccurate feature representation for specific body parts. Human parsing appears to be a better option for accurately representing specific and complete body parts in gait recognition. However, its practical application in gait recognition is often hindered by missing RGB modality, lack of annotated body parts, and difficulty in balancing parsing quantity and quality. To address this issue, we propose LandmarkGait, an accessible and alternative parsing-based solution for gait recognition. LandmarkGait introduces an unsupervised landmark discovery network to transform the dense silhouette into a finite set of landmarks with remarkable consistency across various conditions. By grouping landmarks subsets corresponding to distinct body part regions, following a reconstruction task and further refinement from high-quality input silhouettes, we can directly obtain fine-grained parsing results from original binary silhouettes in an unsupervised manner. Moreover, we also develop a multi-scale feature extractor that simultaneously captures global and parsing feature representations based on the integrity and flexibility of specific body parts. Extensive experiments demonstrate that our LandmarkGait can extract more stable features and exhibit significant performance improvement under all conditions, especially in various dressing conditions. Code is available at https://github.com/wzb-bupt/LandmarkGait. Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu |
ACM Multimedia | 6 |
| 2023 | Learning Gait Representation From Massive Unlabelled Walking Videos: A BenchmarkabstractGait depicts individuals' unique and distinguishing walking patterns and has become one of the most promising biometric features for human identification. As a fine-grained recognition task, gait recognition is easily affected by many factors and usually requires a large amount of completely annotated data that is costly and insatiable. This paper proposes a large-scale self-supervised benchmark for gait recognition with contrastive learning, aiming to learn the general gait representation from massive unlabelled walking videos for practical applications via offering informative walking priors and diverse real-world variations. Specifically, we collect a large-scale unlabelled gait dataset GaitLU-1M consisting of 1.02M walking sequences and propose a conceptually simple yet empirically powerful baseline model GaitSSB. Experimentally, we evaluate the pre-trained model on four widely-used gait benchmarks, CASIA-B, OU-MVLP, GREW and Gait3D with or without transfer learning. The unsupervised results are comparable to or even better than the early model-based and GEI-based methods. After transfer learning, GaitSSB outperforms existing methods by a large margin in most cases, and also showcases the superior generalization capacity. Further experiments indicate that the pre-training can save about 50% and 80% annotation costs of GREW and Gait3D. Theoretically, we discuss the critical issues for gait-specific contrastive framework and present some insights for further study. As far as we know, GaitLU-1M is the first large-scale unlabelled gait dataset, and GaitSSB is the first method that achieves remarkable unsupervised results on the aforementioned benchmarks. Chao Fan 0001, Saihui Hou, Jilong Wang 0010, Yongzhen Huang, Shiqi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | CASIA-E: A Large Comprehensive Dataset for Gait RecognitionabstractGait recognition plays a special role in visual surveillance due to its unique advantage, e.g., long-distance, cross-view and non-cooperative recognition. However, it has not yet been widely applied. One reason for this awkwardness is the lack of a truly big dataset captured in practical outdoor scenarios. Here, the "big" at least means: (1) huge amount of gait videos; (2) sufficient subjects; (3) rich attributes; and (4) spatial and temporal variations. Moreover, most existing large-scale gait datasets are collected indoors, which have few challenges from real scenes, such as the dynamic and complex background clutters, illumination variations, vertical view variations, etc. In this article, we introduce a newly built big outdoor gait dataset, called CASIA-E. It contains more than one thousand people distributed over near one million videos. Each person involves 26 view angles and varied appearances caused by changes of bag carrying, dressing and walking styles. The videos are captured across five months and across three kinds of outdoor scenes. Soft biometric features are also recorded for all subjects including age, gender, height, weight, and nationality. Besides, we report an experimental benchmark and examine some meaningful problems that have not been well studied previously, e.g., the influence of million-level training videos, vertical view angles, walking styles, and the thermal infrared modality. We believe that such a big outdoor dataset and the experimental benchmark will promote the development of gait recognition in both academic research and industrial applications. Chunfeng Song, Yongzhen Huang, Weining Wang 0001, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Gait Quality Aware Network: Toward the Interpretability of Silhouette-Based Gait RecognitionabstractGait recognition receives increasing attention since it can be conducted at a long distance in a nonintrusive way and applied to the condition of changing clothes. Most existing methods take the silhouettes of gait sequences as the input and learn a unified representation from multiple silhouettes to match probe and gallery. However, these models are all faced with the lack of interpretability, e.g., it is not clear which silhouette in a gait sequence and which part in the human body are relatively more important for recognition. In this work, we propose a gait quality aware network (GQAN) for gait recognition which explicitly assesses the quality of each silhouette and each part via two blocks: frame quality block (FQBlock) and part quality block (PQBlock). Specifically, FQBlock works in a squeeze-and-excitation style to recalibrate the features for each silhouette, and the scores of all the channels are added as frame quality indicator. PQBlock predicts a score for each part which is used to compute the weighted distance between the probe and gallery. Particularly, we propose a part quality loss (PQLoss) which enables GQAN to be trained in an end-to-end manner with only sequence-level identity annotations. This work is meaningful by moving toward the interpretability of silhouette-based gait recognition, and our method also achieves very competitive performance on CASIA-B and OUMVLP. Saihui Hou, Xu Liu 0008, Chunshui Cao, Yongzhen Huang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | GaitEdge: Beyond Plain End-to-End Gait Recognition for Better Practicality
Chao Fan 0001, Saihui Hou, Chuanfu Shen, Yongzhen Huang, Shiqi Yu 0001 |
ECCV (5) | 5 |
| 2022 | HID 2022: The 3rd International Competition on Human Identification at a DistanceabstractThe paper provides a summary of the Competition on Human Identification at a Distance 2022 (HID 2022), which is the third one in a series of competitions. HID 2022 is for promoting the research in human identification at a distance by providing a benchmark to evaluate different methods. The competition attracted 112 valid registered teams. 71 teams and 51 teams submitted their results in the first phase and the second phase, respectively. Very encouraging results have been achieved, and the accuracies of the top teams are much higher than those achieved in the previous two competitions. In this paper, we introduce the competition including the dataset, experimental settings, competition organization, results from the top teams and their analysis. The methods used by the top teams are also presented in the paper. The progress of this competition can give us an optimistic view on gait recognition. Shiqi Yu 0001, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Shengjin Wang, Md. Atiqur Rahman Ahad, Mark S. Nixon |
IJCB | 2 |
| 2022 | Global modes and coupled modes for integrated twin circular-side octagon microlasers
Yuede Yang, Youzeng Hao, Jiliang Wu, Yongtao Huang, Jinlong Xiao, Yongzhen Huang |
Sci. China Inf. Sci. | 8 |
| 2022 | Efficient convolutional networks learning through irregular convolutional kernels
Weiyu Guo, Jiabin Ma, Yidong Ouyang, Liang Wang 0001, Yongzhen Huang |
Neurocomputing | 5 |
| 2021 | HID 2021: Competition on Human Identification at a Distance 2021abstractThe Competition on Human Identification at a Distance 2021 (HID 2021) is to promote the research in human identification at a distance and to provide a benchmark to evaluate different methods. HID 2021 is the second follow-up from the first one, HID 2020. The dataset size and the evaluation protocal are the same with the previous competition, but the data in the test set has been changed. The paper firstly introduces the dataset and the evaluation protocol, then describes the methods from the top teams and their results. The methods show how to achieve state-of-the-art performance on gait recognition. The results in HID 2021 are better than those in HID 2020. From the comparisons and analysis, some useful conclusions can be drawn. We hope more improvements can be achieved by better followup competitions. Shiqi Yu 0001, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Edel B. García Reyes, Feng Zheng 0001, Md. Atiqur Rahman Ahad, Beibei Lin, Haijun Xiong, Binyuan Huang |
IJCB | 2 |
| 2020 | GaitPart: Temporal Part-Based Model for Gait RecognitionabstractGait recognition, applied to identify individual walking patterns in a long-distance, is one of the most promising video-based biometric technologies. At present, most gait recognition methods take the whole human body as a unit to establish the spatio-temporal representations. However, we have observed that different parts of human body possess evidently various visual appearances and movement patterns during walking. In the latest literature, employing partial features for human body description has been verified being beneficial to individual recognition. Taken above insights together, we assume that each part of human body needs its own spatio-temporal expression. Then, we propose a novel part-based model GaitPart and get two aspects effect of boosting the performance: On the one hand, Focal Convolution Layer, a new applying of convolution, is presented to enhance the fine-grained learning of the part-level spatial features. On the other hand, the Micro-motion Capture Module (MCM) is proposed and there are several parallel MCMs in the GaitPart corresponding to the pre-defined parts of the human body, respectively. It is worth mentioning that the MCM is a novel way of temporal modeling for gait task, which focuses on the short-range temporal features rather than the redundant long-range features for cycle gait. Experiments on two of the most popular public datasets, CASIA-B and OU-MVLP, richly exemplified that our method meets a new state-of-the-art on multiple standard benchmarks. The source code will be available on https://github.com/ChaoFan96/GaitPart. Chao Fan 0001, Yunjie Peng, Chunshui Cao, Xu Liu 0008, Saihui Hou, Jiannan Chi, Yongzhen Huang, Qing Li 0015, Zhiqiang He 0002 |
CVPR | 7 |
| 2020 | Gait Lateral Network: Learning Discriminative and Compact Representations for Gait Recognition
Saihui Hou, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
ECCV (9) | 4 |
| 2020 | Dense-View GEIs Set: View Space Covering for Gait Recognition based on Dense-View GANabstractGait recognition has proven to be effective for long-distance human recognition. But view variance of gait features would change human appearance greatly and reduce its performance. Most existing gait datasets usually collect data with a dozen different angles, or even more few. Limited view angles would prevent learning better view invariant feature. It can further improve robustness of gait recognition if we collect data with various angles at 1° interval. But it is time consuming and labor consuming to collect this kind of dataset. In this paper, we, therefore, introduce a Dense-View GEIs Set (DV-GEIs) to deal with the challenge of limited view angles. This set can cover the whole view space, view angle from 0° to 180° with 1° interval. In addition, Dense-View GAN (DV-GAN) is proposed to synthesize this dense view set. DV-GAN consists of Generator, Discriminator and Monitor, where Monitor is designed to preserve human identification and view information. The proposed method is evaluated on the CASIA-B and OU-ISIR dataset. The experimental results show that DV-GEIs synthesized by DV-GAN is an effective way to learn better view invariant feature. We believe the idea of dense view generated samples will further improve the development of gait recognition. Rijun Liao, Weizhi An, Shiqi Yu 0001, Zhu Li 0001, Yongzhen Huang |
IJCB | 5 |
| 2020 | A model-based gait recognition method with body pose and human prior knowledge
Rijun Liao, Shiqi Yu 0001, Weizhi An, Yongzhen Huang |
Pattern Recognit. | 4 |
| 2020 | Cross-View Gait Recognition by Discriminative Feature LearningabstractRecently, deep learning based cross-view gait recognition becomes popular owing to the strong capacity of convolutional neural networks (CNNs). Current deep learning methods often rely on loss functions used widely in the task of face recognition, e.g., contrastive loss and triplet loss. These loss functions have the problem of hard negative mining. In this paper, a robust, effective and gait-related loss function, called angle center loss (ACL), is proposed to learn discriminative gait features. The proposed loss function is robust to different local parts and temporal window sizes. Different from center loss which learns a center for each identity, the proposed loss function learns multiple sub-centers for each angle of the same identity. Only the largest distance between the anchor feature and the corresponding crossview sub-centers is penalized, which achieves better intra-subject compactness. We also propose to extract discriminative spatialtemporal features by local feature extractors and a temporal attention model. A simplified spatial transformer network is proposed to localize the suitable horizontal parts of the human body. Local gait features for each horizontal part are extracted and then concatenated as the descriptor. We introduce long-short term memory (LSTM) units as the temporal attention model to learn the attention score for each frame, e.g., focusing more on discriminative frames and less on frames with bad quality. The temporal attention model shows better performance than the temporal average pooling or gait energy images (GEI). By combing the three aspects, we achieve the state-of-the-art results on several cross-view gait recognition benchmarks. Yuqi Zhang 0001, Yongzhen Huang, Shiqi Yu 0001, Liang Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Feedback Convolutional Neural Network for Visual Localization and SegmentationabstractFeedback is a fundamental mechanism existing in the human visual system, but has not been explored deeply in designing computer vision algorithms. In this paper, we claim that feedback plays a critical role in understanding convolutional neural networks (CNNs), e.g., how a neuron in CNNs describes an object's pattern, and how a collection of neurons form comprehensive perception to an object. To model the feedback in CNNs, we propose a novel model named Feedback CNN and develop two new processing algorithms, i.e., neural pathway pruning and pattern recovering. We mathematically prove that the proposed method can reach local optimum. Note that Feedback CNN belongs to weakly supervised methods and can be trained only using category-level labels. But it possesses a powerful capability to accurately localize and segment category-specific objects. We conduct extensive visualization analysis, and the results reveal the close relationship between neurons and object parts in Feedback CNN. Finally, we evaluate the proposed Feedback CNN over the tasks of weakly supervised object localization and segmentation, and the experimental results on ImageNet and Pascal VOC show that our method remarkably outperforms the state-of-the-art ones. Chunshui Cao, Yongzhen Huang, Yi Yang 0007, Liang Wang 0001, Zilei Wang, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | GaitNet: An end-to-end network for gait based human identification
Chunfeng Song, Yongzhen Huang, Yan Huang 0008, Liang Wang 0001 |
Pattern Recognit. | 2 |
| 2019 | GaitGANv2: Invariant gait feature extraction using generative adversarial networks
Shiqi Yu 0001, Rijun Liao, Weizhi An, Edel B. García Reyes, Yongzhen Huang, Norman Poh |
Pattern Recognit. | 6 |
| 2019 | A comprehensive study on gait biometrics using a joint CNN-based method
Yuqi Zhang 0001, Yongzhen Huang, Liang Wang 0001, Shiqi Yu 0001 |
Pattern Recognit. | 2 |
| 2018 | Lateral Inhibition-Inspired Convolutional Neural Network for Visual Attention and Saliency DetectionabstractLateral inhibition in top-down feedback is widely existing in visual neurobiology, but such an important mechanism has not be well explored yet in computer vision. In our recent research, we find that modeling lateral inhibition in convolutional neural network (LICNN) is very useful for visual attention and saliency detection. In this paper, we propose to formulate lateral inhibition inspired by the related studies from neurobiology, and embed it into the top-down gradient computation of a general CNN for classification, i.e. only category-level information is used. After this operation (only conducted once), the network has the ability to generate accurate category-specific attention maps. Further, we apply LICNN for weakly-supervised salient object detection.Extensive experimental studies on a set of databases, e.g., ECSSD, HKU-IS, PASCAL-S and DUT-OMRON, demonstrate the great advantage of LICNN which achieves the state-of-the-art performance. It is especially impressive that LICNN with only category-level supervised information even outperforms some recent methods with segmentation-level supervised learning. Chunshui Cao, Yongzhen Huang, Zilei Wang, Liang Wang 0001, Ninglong Xu, Tieniu Tan |
AAAI | 2 |
| 2018 | Hybrid-cavity semiconductor lasers with a whispering-gallery cavity for controlling Q factor
Yongzhen Huang, Xiuwen Ma, Yuede Yang, Jinlong Xiao |
Sci. China Inf. Sci. | 1 |
| 2018 | Kernel-Based Semantic Hashing for Gait RetrievalabstractIt is very important to retrieve a specific person in locating and tracking the missing people as well as the suspects quickly. However, the well-studied face-based and appearance-based individual retrieval methods are ineffective in the surveillance scenarios because of the far photograph distances, the low camera resolutions, the long time intervals, and the complex lighting conditions. To avoid the disadvantages of face-based and appearance-based methods, we propose to retrieve individuals from the surveillance videos with the gait biometric, which has been proved to be beneficial to remote person recognition and robust to lighting variations. What's more, the gait biometric can be collected without conscious cooperation, making the data collection much easier. But it varies greatly with the view angles, the clothing style, and the carrying conditions. Therefore, the videos of the target person from a similar view angle with the same clothing style and carrying conditions should rank higher than the others. To achieve this purpose and improve the efficiency, this paper proposes a kernel-based semantic hashing (KSH) model, which is learnt by optimizing a semantic triplet ranking loss. Specifically, in the training phase, a semantic similarity score, which depends on the view angles, the clothing style, and the carrying conditions, is calculated for each training pair. Then, a weighted triplet loss considering these semantic scores is designed, which encourages videos with a higher score to stay closer to the gallery in the binary Hamming space. To evaluate the performance of the proposed method, we compare it with several methods on the CASIA Gait Database B and the OU-ISIR Gait Database. The experimental results demonstrate that the KSH is effective and efficient. Yucan Zhou, Yongzhen Huang, Qinghua Hu, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic MonitoringabstractThe rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu. Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001 |
AVSS | 39 |
| 2017 | Invariant feature extraction for gait recognition using only one uniform model
Shiqi Yu 0001, LinLin Shen, Yongzhen Huang |
Neurocomputing | 5 |
| 2017 | A Comprehensive Study on Cross-View Gait Based Human Identification with Deep CNNsabstractThis paper studies an approach to gait based human identification via similarity learning by deep convolutional neural networks (CNNs). With a pretty small group of labeled multi-view human walking videos, we can train deep networks to recognize the most discriminative changes of gait patterns which suggest the change of human identity. To the best of our knowledge, this is the first work based on deep CNNs for gait recognition in the literature. Here, we provide an extensive empirical evaluation in terms of various scenarios, namely, cross-view and cross-walking-condition, with different preprocessing approaches and network architectures. The method is first evaluated on the challenging CASIA-B dataset in terms of cross-view gait recognition. Experimental results show that it outperforms the previous state-of-the-art methods by a significant margin. In particular, our method shows advantages when the cross-view angle is large, i.e., no less than 36 degree. And the average recognition rate can reach 94 percent, much better than the previous best result (less than 65 percent). The method is further evaluated on the OU-ISIR gait dataset to test its generalization ability to larger data. OU-ISIR is currently the largest dataset available in the literature for gait recognition, with 4,007 subjects. On this dataset, the average accuracy of our method under identical view conditions is above 98 percent, and the one for cross-view scenarios is above 91 percent. Finally, the method also performs the best on the USF gait dataset, whose gait sequences are imaged in a real outdoor scene. These results show great potential of this method for practical applications. Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Xiaogang Wang 0001, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Exploring generalized shape analysis by topological representations
Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. Lett. | 2 |
| 2017 | Facial Expression Recognition Based on Deep Evolutional Spatial-Temporal NetworksabstractOne key challenging issue of facial expression recognition is to capture the dynamic variation of facial physical structure from videos. In this paper, we propose a part-based hierarchical bidirectional recurrent neural network (PHRNN) to analyze the facial expression information of temporal sequences. Our PHRNN models facial morphological variations and dynamical evolution of expressions, which is effective to extract "temporal features" based on facial landmarks (geometry information) from consecutive frames. Meanwhile, in order to complement the still appearance information, a multi-signal convolutional neural network (MSCNN) is proposed to extract "spatial features" from still frames. We use both recognition and verification signals as supervision to calculate different loss functions, which are helpful to increase the variations of different expressions and reduce the differences among identical expressions. This deep evolutional spatial-temporal network (composed of PHRNN and MSCNN) extracts the partial-whole, geometry-appearance, and dynamic-still information, effectively boosting the performance of facial expression recognition. Experimental results show that this method largely outperforms the state-of-the-art ones. On three widely used facial expression databases (CK+, Oulu-CASIA, and MMI), our method reduces the error rates of the previous best ones by 45.5%, 25.8%, and 24.4%, respectively. Kaihao Zhang, Yongzhen Huang, Liang Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Localize heavily occluded human faces via deep segmentationabstractLocalizing heavily occluded human faces is a challenging problem in facial detection. Previous methods mainly employ sliding windows by determining whether windows include human faces. In this paper, we provide a novel segmentation-based perspective for heavily occluded face localization with deep convolutional neural networks (CNN). Our model takes an image as input without complicated pre-processing. After several convolutional layers, fully-connected layers and a softmax classifier, we can predict the labels of all pixels in an image, which is the key to localize heavily occluded human faces. Finally, we search a minimal rectangle to localize the human face. Our detector needs neither complex pre-processing nor the time-consuming sliding window. Besides, we use a single model to localize faces to further alleviate computational complexity. Experimental results show that our proposed method is a very effective way to localize heavily occluded human face. Kaihao Zhang, Yongzhen Huang, Ran He 0001, Liang Wang 0001 |
ICIP | 2 |
| 2016 | View invariant gait recognition using only one uniform modelabstractGait recognition has been proved useful in human identification at a distance. But view variance of gait feature is always a great challenge because of the difference in appearance. If the view of the probe is different from that of the gallery, one view transformation model can be employed to convert the gait feature from one view to another. But most existing models need to estimate the view angle first, and can work for only one view pair. They can not convert multi-view data to one specific view efficiently. We employ one deep model based on auto-encoder for view invariant gait extraction. The model can synthesize gait feature in a progressive way by stacked multi-layer auto-encoders. The unique advantage is that it can extract view invariant feature from any view using only one model, and view estimation is not needed. The proposed method is evaluated on a large dataset, CASIA Gait Dataset B. The experimental results show that it can achieve state-of-the-art performance, and the improvement is more obvious when the view variance is larger. Shiqi Yu 0001, LinLin Shen, Yongzhen Huang |
ICPR | 4 |
| 2016 | Learning Relevance Restricted Boltzmann Machine for Unstructured Group Activity and Event Understanding
Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tao Xiang 0002, Tieniu Tan |
Int. J. Comput. Vis. | 2 |
| 2016 | Robust tracking with adaptive appearance learning and occlusion detection
Jianwei Ding, Yunqi Tang, Huawei Tian, Yongzhen Huang |
Multim. Syst. | 5 |
| 2016 | Severely Blurred Object Tracking by Learning Deep Image RepresentationsabstractAn implicit assumption in many generic object trackers is that the videos are blur free. However, motion blur is very common in real videos. The performance of a generic object tracker may drop significantly when it is applied to videos with severe motion blur. In this paper, we propose a new Tracking-Learning-Data approach to transfer a generic object tracker to a blur-invariant object tracker without deblurring image sequences. Before object tracking, a large set of unlabeled images is used to learn objects' visual prior knowledge, which is then transferred to the appearance model of a specific target. During object tracking, online training samples are collected from the tracking results and the context information. Different blur kernels are involved with the training samples to increase the robustness of the appearance model to severe blur, and the motion parameters of the object are estimated in the particle filter framework. Extensive experimental results demonstrate that the proposed algorithm can robustly track objects not only in severely blurred videos but also in other challenging scenes. Jianwei Ding, Yongzhen Huang, Wei Liu 0023, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Coupled Topic Model for Collaborative Filtering With User-Generated ContentabstractThe user-generated content (UGC) is a type of dyadic information that provides description of the interaction between users and items (such as rating, purchasing, etc.). Most conventional methods incorporate either a user profile or the item description, which cannot well utilize this kind of content information. Some other works jointly consider user ratings and reviews, but they are based on the factorization technique and have difficulty in providing explanations on generated recommendations. In this study, a coupled topic model (CoTM) for recommendation with UGC is developed. By combining UGC and ratings, the method discussed in this study captures both the content-based preferences and collaborative preferences and, thus, can explain both the user and item latent spaces using the topics discovered from the UGC. The learned topics in CoTM can also serve as proper explanations for the generated recommendations. Experimental results show that the proposed CoTM model yields significant improvements over the compared competitive methods on two typical datasets, that is, MovieLens-10M and Citation-network V1. The topics discovered by CoTM can be used not only to illustrate the topic distributions of users and items, but also to explain the generated user-item recommendations. Weiyu Guo, Song Xu 0002, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2015 | Kinship Verification with Deep Convolutional Neural Networks
Kaihao Zhang, Yongzhen Huang, Chunfeng Song, Liang Wang 0001 |
BMVC | 2 |
| 2015 | Deep semantic ranking based hashing for multi-label image retrievalabstractWith the rapid growth of web images, hashing has received increasing interests in large scale image retrieval. Research efforts have been devoted to learning compact binary codes that preserve semantic similarity based on labels. However, most of these hashing methods are designed to handle simple binary similarity. The complex multi-level semantic structure of images associated with multiple labels have not yet been well explored. Here we propose a deep semantic ranking based method for learning hash functions that preserve multilevel semantic similarity between multi-label images. In our approach, deep convolutional neural network is incorporated into hash functions to jointly learn feature representations and mappings from them to hash codes, which avoids the limitation of semantic representation power of hand-crafted features. Meanwhile, a ranking list that encodes the multilevel similarity information is employed to guide the learning of such deep hash functions. An effective scheme based on surrogate loss is used to solve the intractable optimization problem of nonsmooth and multivariate ranking measures involved in the learning procedure. Experimental results show the superiority of our proposed approach over several state-of-the-art hashing methods in term of ranking evaluation metrics when tested on multi-label image datasets. Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
CVPR | 2 |
| 2015 | Rebooting Computing and Low-Power Image Recognition Challengeabstract“Rebooting Computing” (RC) is an effort in the IEEE to rethink future computers. RC started in 2012 by the co-chairs, Elie Track (IEEE Council on Superconductivity) and Tom Conte (Computer Society). RC takes a holistic approach, considering revolutionary as well as evolutionary solutions needed to advance computer technologies. Three summits have been held in 2013 and 2014, discussing different technologies, from emerging devices to user interface, from security to energy efficiency, from neuromorphic to reversible computing. The first part of this paper introduces RC to the design automation community and solicits revolutionary ideas from the community for the directions of future computer research. Energy efficiency is identified as one of the most important challenges in future computer technologies. The importance of energy efficiency spans from miniature embedded sensors to wearable computers, from individual desktops to data centers. To gauge the state of the art, the RC Committee organized the first Low Power Image Recognition Challenge (LPIRC). Each image contains one or multiple objects, among 200 categories. A contestant has to provide a working system that can recognize the objects and report the bounding boxes of the objects. The second part of this paper explains LPIRC and the solutions from the top two winners. Yung-Hsiang Lu, Alan M. Kadin, Alexander C. Berg, Thomas M. Conte, Erik DeBenedictis, Ganesh Gingade, Bichlien Hoang, Yongzhen Huang, Boxun Li, Jingyu Liu 0004, Wei Liu 0015, Huizi Mao, Junran Peng, Tianqi Tang 0001, Elie K. Track, Jingqiu Wang, Tao Wang 0004, Yu Wang 0002 |
ICCAD | 9 |
| 2015 | Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural NetworksabstractWhile feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to note that the human visual cortex generally contains more feedback than feedforward connections. In this paper, we will briefly introduce the background of feedbacks in the human visual cortex, which motivates us to develop a computational feedback mechanism in deep neural networks. In addition to the feedforward inference in traditional neural networks, a feedback loop is introduced to infer the activation status of hidden layer neurons according to the "goal" of the network, e.g., high-level semantic labels. We analogize this mechanism as "Look and Think Twice." The feedback networks help better visualize and understand how deep neural networks work, and capture visual attention on expected objects, even in images with cluttered background and multiple objects. Experiments on ImageNet dataset demonstrate its effectiveness in solving tasks such as image classification and object localization. Chunshui Cao, Xianming Liu 0005, Yi Yang 0007, Yinan Yu, Jiang Wang 0001, Zilei Wang, Yongzhen Huang, Liang Wang 0001, Chang Huang, Wei Xu 0017, Deva Ramanan, Thomas S. Huang |
ICCV | 7 |
| 2015 | Multi-view descriptor mining via codeword net for action recognitionabstractAction recognition is an important yet challenging task in computer vision. A successful and widely used framework in this field is the Bag of Visual Words (BoVW), wherein the first step is to extract local features. One critical property of local features is that they are often multi-view, e.g., dense trajectory feature includes both appearance and motion properties. Different types of features are aligned together in coding and pooling thus leading the process to be heavily entangled. Our motivation is to disentangle each sub-descriptor and let them contribute to the maximum extent. To achieve this, a codeword net is constructed via exploiting the relation between features and codewords. Based on the codeword net, features from the same viewpoint are pooled together. Experiments on two large scale action recognition datasets, UCF50 and HMDB51, demonstrate that our approach can enhance the state-of-the-art algorithms. Jingyu Liu 0004, Yongzhen Huang, Xiaojiang Peng, Liang Wang 0001 |
ICIP | 2 |
| 2015 | Tracking by local structural manifold learning in a new SSIR particle filter
Jianwei Ding, Yunqi Tang, Wei Liu 0023, Yongzhen Huang, Kaiqi Huang |
Neurocomputing | 4 |
| 2015 | Learning Representative Deep Features for Image Set AnalysisabstractThis paper proposes to learn features from sets of labeled raw images. With this method, the problem of over-fitting can be effectively suppressed, so that deep CNNs can be trained from scratch with a small number of training data, i.e., 420 labeled albums with about 30 000 photos. This method can effectively deal with sets of images, no matter if the sets bear temporal structures. A typical approach to sequential image analysis usually leverages motions between adjacent frames, while the proposed method focuses on capturing the co-occurrences and frequencies of features. Nevertheless, our method outperforms previous best performers in terms of album classification, and achieves comparable or even better performances in terms of gait based human identification. These results demonstrate its effectiveness and good adaptivity to different kinds of set data. Zifeng Wu, Yongzhen Huang, Liang Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2014 | Fusion of Multibiometrics Based on a New Robust Linear ProgrammingabstractMultibiometrics provides a reliable method for identity authentication and has the potential to be widely applied. The success of a multibiometrics method depends critically on its ability to fuse complementary information supplied by different modalities, where the most challenging problem is to evaluate the importance of different modalities. In addition, identity authentication at a distance has become a development trend of multibiometrics. In this paper, we propose a new robust linear programming method to fuse multibiometrics by combining the modalities optimally. The proposed method can provide a reasonable trade off between conservatism and robustness. Experimental results on CASIA-Iris-Distance, a public and challenging multibiometric database, demonstrate the effectiveness and robustness of this method. Di Miao, Zhenan Sun, Yongzhen Huang |
ICPR | 3 |
| 2014 | Early Hierarchical Contexts Learned by Convolutional Networks for Image SegmentationabstractWe propose a foreground segmentation method based on convolutional networks. To predict the label of a pixel in an image, the model takes a hierarchical context as the input, which is obtained by combining multiple context patches on different scales. Short range contexts depict the local details, while long range contexts capture the object-scene relationships in an image. Early means that we combine the context patches of a pixel into a hierarchical one before any trainable layers are learned, i.e., early-combing. In contrast, late-combing means that the combination occurs later, e.g., when the convolutional feature extractor in a network has already been learned. We find that it is vital for the whole model to jointly learn the patterns of contexts on different scales in our task. Experiments show that early-combing performs better than late-combing. On the dataset1 built up by Baidu IDL2 for a latest person segmentation contest, our method beats all the competitors with a considerable margin. Qualitative results also show that the proposed method is almost ready for practical application. Zifeng Wu, Yongzhen Huang, Yinan Yu, Liang Wang 0001, Tieniu Tan |
ICPR | 2 |
| 2014 | Deep auto-encoder based clusteringabstractFor unsupervised problems like clustering, linear or non-linear data transformations are widely used techniques. Generally, they are beneficial to data representation. However, if data have a complicated structure, these techniques would be unsatisfying for clustering. In this paper, we propose a new clustering method based on the deep auto-encoder network, which can learn a highly non-linear mapping function. Via simultaneously considering data reconstruction and compactness, our method can obtain stable and effective clustering. Experimental results on four databases demonstrate that the proposed model can achieve promising performance in terms of normalized mutual information, cluster purity and accuracy. Chunfeng Song, Yongzhen Huang, Feng Liu 0036, Zhenyu Wang 0012, Liang Wang 0001 |
Intell. Data Anal. | 2 |
| 2014 | Multiple spatial pooling for visual object recognition
Yongzhen Huang, Zifeng Wu, Liang Wang 0001, Chunfeng Song |
Neurocomputing | 1 |
| 2014 | Hierarchical feature coding for image classification
Jingyu Liu 0004, Yongzhen Huang, Liang Wang 0001 |
Neurocomputing | 2 |
| 2014 | Spatial modeling via feature co-pooling and SG grafting
Feng Liu 0036, Yongzhen Huang, Liang Wang 0001, Wankou Yang, Changyin Sun 0001 |
Neurocomputing | 2 |
| 2014 | Feature Coding in Image Classification: A Comprehensive StudyabstractImage classification is a hot topic in computer vision and pattern recognition. Feature coding, as a key component of image classification, has been widely studied over the past several years, and a number of coding algorithms have been proposed. However, there is no comprehensive study concerning the connections between different coding methods, especially how they have evolved. In this paper, we first make a survey on various feature coding methods, including their motivations and mathematical representations, and then exploit their relations, based on which a taxonomy is proposed to reveal their evolution. Further, we summarize the main characteristics of current algorithms, each of which is shared by several coding strategies. Finally, we choose several representatives from different kinds of coding approaches and empirically evaluate them with respect to the size of the codebook and the number of training samples on several widely used databases (15-Scenes, Caltech-256, PASCAL VOC07, and SUN397). Experimental findings firmly justify our theoretical analysis, which is expected to benefit both practical applications and future research. Yongzhen Huang, Zifeng Wu, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Auto-encoder Based Data Clustering
Chunfeng Song, Feng Liu 0036, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
CIARP (1) | 3 |
| 2013 | Recent Progress on Object Classification and Detection
Tieniu Tan, Yongzhen Huang, Junge Zhang |
CIARP (2) | 2 |
| 2013 | Discovering compact topical descriptors for web video retrievalabstractDescribing videos efficiently is an important task for content based web video retrieval. To solve this problem, we propose an unsupervised approach based on an undirected topic model to learn a compact topical descriptor upon the bag-of-words (BoW) video representation. In our method, words in a BoW are assumed to have different topic features, and the topical descriptor of an entire video is obtained by aggregating those features, which makes the descriptor contain information about relative strength of topics. To improve the descriptor interpretability, an L1penalty is used to control the topical sparsity. Furthermore, efficient learning and inference algorithms are presented. We evaluate the proposed descriptor on the Columbia Consumer Video dataset. Experimental results demonstrate that compared with the BoW and other topical representations, the proposed compact descriptor has better performance in web video retrieval. Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICIP | 2 |
| 2013 | Depth-embedded multiple pooling for image classificationabstractMost existing methods of image classification ignore the role of depth information hidden in 2-D images. However, the depth information is important for visual perception, especially when the appearance information does not perform well. In this paper, we propose to embed depth information within multiple pooling into the classic platform of image classification, namely bag-of-features. The proposed method quantifies depth diversity by projecting objects to their nearby depth planes, resulting pooling features in the 3-D space indirectly. Experimental results on the MIT Indoor Scene database demonstrate that our proposed depth-embedded multiple pooling is effective to enhance the accuracy of image classification, especially when the appearance features alone are not so discriminative. Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICIP | 2 |
| 2013 | Relevance Topic Model for Unstructured Social Group Activity RecognitionabstractUnstructured social group activity recognition in web videos is a challenging task due to 1) the semantic gap between class labels and low-level visual features and 2) the lack of labeled training data. To tackle this problem, we propose a relevance topic model" for jointly learning meaningful mid-level representations upon bag-of-words (BoW) video representations and a classifier with sparse weights. In our approach, sparse Bayesian learning is incorporated into an undirected topic model (i.e., Replicated Softmax) to discover topics which are relevant to video classes and suitable for prediction. Rectified linear units are utilized to increase the expressive power of topics so as to explain better video data containing complex contents and make variational inference tractable for the proposed model. An efficient variational EM algorithm is presented for model parameter estimation and inference. Experimental results on the Unstructured Social Activity Attribute dataset show that our model achieves state of the art performance and outperforms other supervised topic model in terms of classification accuracy, particularly in the case of a very small number of labeled training videos." Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
NIPS | 2 |
| 2012 | Local Hypersphere Coding Based on Edges between Visual Words
Weiqiang Ren, Yongzhen Huang, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
ACCV (1) | 2 |
| 2012 | Contextual Pooling in Image Classification
Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ACCV (1) | 2 |
| 2012 | Spatial Graph for Image Classification
Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ACCV (1) | 2 |
| 2012 | Data Decomposition and Spatial Mixture Modeling for Part Based Model
Junge Zhang, Yongzhen Huang, Kaiqi Huang, Zifeng Wu, Tieniu Tan |
ACCV (1) | 2 |
| 2012 | CLUMOC: Multiple Motion Estimation by Cluster Motion ConsensusabstractIn this paper, we present techniques for robust multiple motions estimation based on dual consensus via clustering in both the image spatial space and the motion parameter space. Starting from traditional Random Samples Consensus algorithm, we novelly propose the CLUster MOtion Consensus (CLUMOC) to extract robust motions. The proposed algorithm has two advantages: (1), instead of random samples, the CLUMOC employs clustering in initial sample selection, which can remove outliers from correct pairs of motion, (2), CLUMOC automatically decides the number of motions, by employing competition among motion and samples, that each motion needs to compete for matching pairs and each pair of matching competes for motions. The experimental results show that the proposed method is effective and efficient under various situations. Yinan Yu, Weiqiang Ren, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
AVSS | 3 |
| 2012 | Feature coding via vector difference for image classificationabstractAn effective image representation is important to an image classification task. The most popular image representation framework utilizes a feature coding algorithm to encode the extracted low-level feature descriptors into a vector representation. In this paper, we analyze the recently developed feature coding methods in a general way. According to their common characteristics, we propose a new coding scheme to perform feature coding based on the vector difference in a high-dimensional space which is obtained by explicit feature maps. As we illustrate, our method has promising results with small codebook sizes and generalizes most existing coding methods in a unified form. Xin Zhao 0012, Yinan Yu, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2012 | Group encoding of local features in image classification
Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICPR | 2 |
| 2012 | Semantic windows mining in sliding window based object detection
Junge Zhang, Xin Zhao 0012, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICPR | 3 |
| 2011 | Exploring relations of visual codes for image classificationabstractThe classic Bag-of-Features (BOF) model and its extensional work use a single value to represent a visual code. This strategy ignores the relation of visual codes. In this paper, we explore this relation and propose a new algorithm for image classification. It consists of two main parts: 1) construct the codebook graph wherein a visual code is linked with other codes; 2) describe each local feature using a pair of related codes, corresponding to an edge of the graph. Our approach contains richer information than previous BOF models. Moreover, we demonstrate that these models are special cases of ours. Various coding and pooling algorithms can be embedded into our framework to obtain better performance. Experiments on different kinds of image classification databases demonstrate that our approach can stably achieve excellent performance compared with various BOF models. Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
CVPR | 1 |
| 2011 | Salient coding for image classificationabstractThe codebook based (bag-of-words) model is a widely applied model for image classification. We analyze recent coding strategies in this model, and find that saliency is the fundamental characteristic of coding. The saliency in coding means that if a visual code is much closer to a descriptor than other codes, it will obtain a very strong response. The salient representation under maximum pooling operation leads to the state-of-the-art performance on many databases and competitions. However, most current coding schemes do not recognize the role of salient representation, so that they may lead to large deviations in representing local descriptors. In this paper, we propose “salient coding”, which employs the ratio between descriptors' nearest code and other codes to describe descriptors. This approach can guarantee salient representation without deviations. We study salient coding on two sets of image classification databases (15-Scenes and PASCAL VOC2007). The experimental results demonstrate that our approach outperforms all other coding methods in image classification. Yongzhen Huang, Kaiqi Huang, Yinan Yu, Tieniu Tan |
CVPR | 1 |
| 2011 | Enhanced Biologically Inspired Model for Object RecognitionabstractThe biologically inspired model (BIM) proposed by Serre presents a promising solution to object categorization. It emulates the process of object recognition in primates' visual cortex by constructing a set of scale- and position-tolerant features whose properties are similar to those of the cells along the ventral stream of visual cortex. However, BIM has potential to be further improved in two aspects: mismatch by dense input and randomly feature selection due to the feedforward framework. To solve or alleviate these limitations, we develop an enhanced BIM (EBIM) in terms of the following two aspects: 1) removing uninformative inputs by imposing sparsity constraints, 2) apply a feedback loop to middle level feature selection. Each aspect is motivated by relevant psychophysical research findings. To show the effectiveness of the EBIM, we apply it to object categorization and conduct empirical studies on four computer vision data sets. Experimental results demonstrate that the EBIM outperforms the BIM and is comparable to state-of-the-art approaches in terms of accuracy. Moreover, the new system is about 20 times faster than the BIM. Yongzhen Huang, Kaiqi Huang, Dacheng Tao, Tieniu Tan, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | A Heuristic Deformable Pedestrian Detection Method
Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ACCV (2) | 1 |
| 2009 | A Novel Visual Organization Based on Topological Perception
Yongzhen Huang, Kaiqi Huang, Tieniu Tan, Dacheng Tao |
ACCV (1) | 1 |
| 2009 | Computational primitives of visual perceptionabstractGreat stride has been made in psychological research about primitives of visual perception, which is important to computer vision and image processing. In this paper, we propose a computational model to imitate the primitives of visual perception based on the pyschological theory of topological perceptual organization. First, we adopt geodesic distance based descriptor to describe an independent topological structure. Then, we consider the spatial relationship of two independent structures. Experiments on structures classification demonstrates that the propose model is consistent with the psychological theory. Further experiments on patches clustering prove that our approach can be used to enhance other algorithms. Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICIP | 1 |
| 2009 | Object detection and tracking for night surveillance based on salient contrast analysisabstractNight surveillance is a challenging task because of low brightness, low contrast, low Signal to Noise Ratio (SNR) and low appearance information. Most existing models for night surveillance share the following problems: a lack of adaptability for different scenes and separation between detection and tracking. To solve these problems we propose a model based on Salient Contrast Change (SCC) feature, which applies learning process to enhance adaptability and analyzes trajectories to improve the effectiveness of detection. Empirical studies on several real night videos show that the proposed model is more effective than the original CC model and other traditional models. Liangsheng Wang, Kaiqi Huang, Yongzhen Huang, Tieniu Tan |
ICIP | 3 |
| 2009 | View-invariant action recognition using cross ratios across framesabstractWe present a new method of computing invariants in videos captured from different views to achieve view-invariant action recognition. To avoid the constraints of collinearity or coplanarity of image points for constructing invariants, we consider several neighboring frames to compute cross ratios, namely cross ratios across frames (CRAF), as our invariant representation of action. For every five points sampled with different intervals from the trajectories of action, we construct a pair of cross ratios (CRs). Afterwards, we transform the CRs to histograms as the feature vectors for classification. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods in effectiveness and stability. Yeying Zhang, Kaiqi Huang, Yongzhen Huang, Tieniu Tan |
ICIP | 3 |
| 2008 | Enhanced biologically inspired modelabstractIt has been demonstrated by Serre et al. that the biologically inspired model (BIM) is effective for object recognition. It outperforms many state-of-the-art methods in challenging databases. However, BIM has the following three problems: a very heavy computational cost due to dense input, a disputable pooling operation in modeling relations of the visual cortex, and blind feature selection in a feed-forward framework. To solve these problems, we develop an enhanced BIM (EBIM), which removes uninformative input by imposing sparsity constraints, utilizes a novel local weighted pooling operation with stronger physiological motivations, and applies a feedback procedure that selects effective features for combination. Empirical studies on the CalTech5 database and CalTech101 database show that EBIM is more effective and efficient than BIM. We also apply EBIM to the MIT-CBCL street scene database to show it achieves comparable performance in comparison with the current best performance. Moreover, the new system can process images with resolution 128 times 128 at a rate of 50 frames per second and enhances the speed 20 times at least in comparison with BIM in common applications. Yongzhen Huang, Kaiqi Huang, Liangsheng Wang, Dacheng Tao, Tieniu Tan, Xuelong Li 0001 |
CVPR | 1 |