Yan Huang 0008

dblp:75/6434-8 · DBLP profile ↗
← Back
110ranked-venue papers
19as first author
62since 2021 · last 2026
0000-0002-8239-7229ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 71 · 15 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 11 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Gait Transformer: End-to-End Transformer Backbone for Gait Recognition
abstract
Gait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modality, spatial-temporal dynamics, fine-grained motion cues, and limited training data. In this paper, we propose Gait Transformer (GaT), an end-to-end Transformer backbone specifically tailored for silhouette-based gait recognition. GaT introduces three key components: (1) a hybrid patch embedding module that combines convolutional stems with group-batch normalization to enhance structural preservation; (2) a decomposed token mixer that explicitly models both short-range and long-range dependencies across spatial-temporal dimensions; and (3) a hybrid positional encoding strategy that integrates absolute, relative, and rotary embeddings to support efficient training under data scarcity. Without relying on any pretraining, GaT achieves state-of-the-art performance on Gait3D, GREW, and CCGR-MINI.
Saihui Hou, Wenpeng Lang, Jilong Wang 0010, Yan Huang 0008, Liang Wang 0001, Yongzhen Huang
AAAI4
2026 HyCTAS: Multi-objective hybrid convolution-transformer architecture search for real-time image segmentation
Hongyuan Yu, Cheng Wan 0006, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Bin Xiao 0004, Yan Huang 0008, Liang Wang 0001
Neurocomputing7
2026 Fine-Grained Alignment Supervision Matters in Vision-and-Language Navigation
abstract
The Vision-and-Language Navigation (VLN) task involves an agent navigating within 3D indoor environments based on provided instructions. Achieving cross-modal alignment presents one of the most critical challenges in VLN, as the predicted trajectory needs to precisely align with the given instruction. This paper focuses on addressing cross-modal alignment in VLN from a fine-grained perspective. Firstly, to address the issue of weak cross-modal alignment supervision arising from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset called Landmark-RxR. This dataset aims to offer precise, fine-grained supervision for VLN. Secondly, in order to comprehensively demonstrate the potential and advantage of the fine-grained data from Landmark-RxR, we explore the core components of the training process that depend on the characteristics of the training data. These components include data augmentation, training paradigm, reward shaping, and navigation loss design. Leveraging our fine-grained data, we carefully design methods for handling them and introduce a novel evaluation mechanism. The experimental results demonstrate that the fine-grained data can effectively improve the agent's cross-modal alignment ability.
Keji He, Yan Huang 0008, Ya Jing, Qi Wu 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation
abstract
Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks in top-down views. To achieve this, we first construct a Fine-Grained AVDN (FG-AVDN) dataset via a semi-automatic annotation pipeline, providing diverse multimodal annotations at the entity-landmark level. Based on this, a novel Fine-grained Entity-Landmark Alignment (FELA) method is proposed to learn the cross-modal alignment explicitly. Concretely, FELA first boosts the drone's visual understanding with a precise semantic grid representation, which captures the environmental semantics and spatial structure simultaneously. Subsequently, to learn the entity-landmark alignment, we devise cross-modal auxiliary tasks from three perspectives, including grounding, captioning, and contrastive learning. Extensive experiments demonstrate that our explicit entity-landmark alignment learning is beneficial for AVDN. As a result, FELA achieves leading performance with 3.2% SR and 4.9% GP improvements over prior arts. Code and dataset will be publicly available.
Yifei Su, Dong An 0002, Weichen Yu, Baiyang Ning, Yonggen Ling, Yan Huang 0008, Liang Wang 0001
AAAI7
2025 EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
abstract
Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate this issue, they remain limited to scenarios involving rigid objects with clear displacement and minimal occlusion. In this work, we present Embodiment-Centric Flow (EC-Flow), a framework that directly learns manipulation from action-unlabeled videos by predicting embodiment-centric flow. Our key insight is that incorporating the embodiment's inherent kinematics significantly enhances generalization to versatile manipulation scenarios, including deformable object handling, occlusions, and non-object-displacement tasks. To connect the EC-Flow with language instructions and object interactions, we further introduce a goal-alignment module by jointly optimizing movement consistency and goal-image prediction. Moreover, translating EC-Flow to executable robot actions only requires a standard robot URDF (Unified Robot Description Format) file to specify kinematic constraints across joints, which makes it easy to use in practice. We validate EC-Flow on both simulation (Meta-World) and real-world tasks, demonstrating its state-of-the-art performance in occluded object handling (62% improvement), deformable object manipulation (45% improvement), and non-object-displacement tasks (80% improvement) than prior state-of-the-art object-centric flow methods. For more information, see our project website at https://ec-flow1.github.io .
Peiyan Li 0001, Yan Huang 0008, Jiabing Yang, Liang Wang 0001
ICCV3
2025 Learning Supplementary Information for First-Person Perception Referring Expression Comprehension
Zetao Du, Yan Huang 0008, Liang Wang 0001, Feng Chen 0044, Zhepeng Wang 0002
ICIG (2)3
2025 CAPNet: Context-Aware Prompt Network for Weakly-Supervised Open-World Phrase-Grounding
Naigong Yu, Zhaoxuan Lu, Yan Huang 0008, Jinhan Yan, Liang Wang 0001
ICIG (3)4
2025 SPENet: Self-guided Prototype Enhancement Network for Few-Shot Medical Image Segmentation
Chao Fan 0001, Xibin Jia, Anqi Xiao, Hongyuan Yu, Zhenghan Yang, Yan Huang 0008, Liang Wang 0001
MICCAI (5)8
2025 Frequency-Domain Multi-modal Fusion for Language-Guided Medical Image Segmentation
Zetao Du, Yan Huang 0008, Chenglong Li 0002, Liang Wang 0001
MICCAI (9)4
2025 VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
abstract
Most of existing weakly supervised GAR methods are typically bottom-up, automatically mining key areas by the attention mechanism. Due to the lack of a semantic connection to individual actions, some regions associated with these actions may be omitted, potentially impacting performance. In fact, a group activity is a combination of multiple individual actions, and the prototype of a specific action can be obtained from visual representations of individuals performing it, denoted as visual conceptual knowledge. In this paper, we propose a Visual Conceptual Knowledge Guided Action Map framework. It uses prototypes to produce individual action maps that indicate the likelihood of actions occurring at different locations. In some scenarios, the spatial distribution of actions shows strong regularity, which we compile as A-A Maps to enhance individual action maps. The action maps are integrated with action semantic representations for group activity recognition. Extensive experiments on two public benchmarks, the Volleyball and the NBA datasets, demonstrate the effectiveness of our proposed method, even in cases of limited training data.
Zhuming Wang, Yihao Zheng 0002, Jiarui Li 0002, Yaofei Wu, Yan Huang 0008, Zun Li 0001, Lifang Wu, Liang Wang 0001
ACM Multimedia5
2025 Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression
abstract
Vision-language models heavily rely on visual representations, yet ensuring its efficiency remains a critical challenge. Most existing approaches focus on reducing visual tokens either at the visual encoder phase or during the LLM decoder stage. Inspired by human visual cognition, where an initial global glance precedes focused attention on semantically salient regions, we introduce Glance2Gaze, a cognitively inspired framework that mimics the human two-stage attention process. The framework consists of two key components: the Glance Fusion module, which integrates multi-layer vision transformer features with text-aware attention to generate a semantically enriched global representation, and the Gaze Compression module, which utilizes a novel query-guided mechanism to selectively compress visual tokens based on their semantic relevance. Experimental results on widely adopted benchmarks demonstrate that Glance2Gaze outperforms existing methods, achieving superior performance with equal or lower computational cost. Furthermore, it generalizes well to high-resolution and video scenarios, showcasing robust and scalable efficiency improvements in VLMs.
Yingying Ao, Ting Zhang 0006, Yan Huang 0008, Jintao Fang
NeurIPS5
2025 BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
abstract
Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods incorporate 3D signals into VLMs for action prediction, and they do not fully leverage the spatial structure inherent in 3D data, leading to low data efficiency. In this paper, we introduce a new paradigm for constructing 3D VLAs. Specifically, we first pre-train the VLM backbone to take 2D images as input and produce 2D heatmaps as output. Using this pre-trained VLM as the backbone, we then fine-tune the entire VLA model while maintaining alignment between inputs and outputs by: (1) projecting raw point cloud inputs into multi-view images, and (2) predicting heatmaps before generating the final action. Extensive experiments show that the resulting model, BridgeVLA, can learn 3D manipulation both efficiently and effectively. BridgeVLA outperforms state-of-the-art baselines across three simulation benchmarks. In RLBench, it improves the average success rate from 81.4\% to 88.2\%. In COLOSSEUM, it demonstrates significantly better performance in challenging generalization settings, boosting the average success rate from 56.7\% to 64.0\%. In GemBench, it surpasses all the comparing baseline methods in terms of average success rate. In real-robot experiments, BridgeVLA outperforms a state-of-the-art baseline method by 32\% on average. It generalizes robustly in multiple out-of-distribution settings, including visual disturbances and unseen instructions. Remarkably, it is able to achieve a success rate of 95.4\% on 10+ tasks with only 3 trajectories per task, while other VLA methods such as $\pi_{0}$ fail completely. Project Website: https://bridgevla.github.io/.
Peiyan Li 0001, Xiangnan Wu, Yan Huang 0008, Liang Wang 0001, Tao Kong, Tieniu Tan
NeurIPS6
2025 CSFRNet: Integrating Clothing Status Awareness for Long-Term Person Re-identification
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
Int. J. Comput. Vis.1
2025 ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
abstract
Vision-language navigation is a task that requires an agent to follow instructions to navigate in environments. It becomes increasingly crucial in the field of embodied AI, with potential applications in autonomous navigation, search and rescue, and human-robot interaction. In this paper, we propose to address a more practical yet challenging counterpart setting - vision-language navigation in continuous environments (VLN-CE). To develop a robust VLN-CE agent, we propose a new navigation framework, ETPNav, which focuses on two critical skills: 1) the capability to abstract environments and generate long-range navigation plans, and 2) the ability of obstacle-avoiding control in continuous environments. ETPNav performs online topological mapping of environments by self-organizing predicted waypoints along a traversed path, without prior environmental experience. It privileges the agent to break down the navigation procedure into high-level planning and low-level control. Concurrently, ETPNav utilizes a transformer-based cross-modal planner to generate navigation plans based on topological maps and instructions. The plan is then performed through an obstacle-avoiding controller that leverages a trial-and-error heuristic to prevent navigation from getting stuck in obstacles. Experimental results demonstrate the effectiveness of the proposed method. ETPNav yields more than 10% and 20% improvements over prior state-of-the-art on R2R-CE and RxR-CE datasets, respectively.
Dong An 0002, Hanqing Wang 0001, Wenguan Wang, Zun Wang 0001, Yan Huang 0008, Keji He, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
abstract
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence of expert demonstrations for training and minimal environment structural prior to guide navigation. To confront these challenges, we propose a Constraint-Aware Navigator (CA-Nav), which reframes zero-shot VLN-CE as a sequential, constraint-aware sub-instruction completion process. CA-Nav continuously translates sub-instructions into navigation plans using two core modules: the Constraint-Aware Sub-instruction Manager (CSM) and the Constraint-Aware Value Mapper (CVM). CSM defines the completion criteria for decomposed sub-instructions as constraints and tracks navigation progress by switching sub-instructions in a constraint-aware manner. CVM, guided by CSM's constraints, generates a value map on the fly and refines it using superpixel clustering to improve navigation stability. CA-Nav achieves the state-of-the-art performance on two VLN-CE benchmarks, surpassing the previous best method by 12% and 13% in Success Rate on the validation unseen splits of R2R-CE and RxR-CE, respectively. Moreover, CA-Nav demonstrates its effectiveness in real-world robot deployments across various indoor scenes and instructions.
Dong An 0002, Yan Huang 0008, Rongtao Xu, Yifei Su, Yonggen Ling, Ian D. Reid 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Unpaired Image-Text Matching via Multimodal Aligned Conceptual Knowledge
abstract
Recently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which use millions or billions of paired images and texts for supervised model learning. Different from them, human brains can well match images with texts using their stored multimodal knowledge. Inspired by that, this paper studies a new scenario as unpaired image-text matching, in which paired images and texts are assumed to be unavailable during model learning. To deal with it, we accordingly propose a simple yet effective method namely Multimodal Aligned Conceptual Knowledge (MACK). First, we collect a set of words and their related image regions from publicly available datasets, and compute prototypical region representations to obtain pretrained general knowledge. To make the obtained knowledge better suit for certain datasets, we refine it using unpaired images and texts in a self-supervised learning manner to obtain fine-tuned domain knowledge. Then, to match given images with texts based on the knowledge, we represent parsed words in the texts by prototypical region representations, and compute region-word similarity scores. At last, the scores are aggregated based on bidirectional similarity pooling into an image-text similarity score, which can be directly used for unpaired image-text matching. The proposed MACK is complementary with existing models, which can be easily extended as a re-ranking method to substantially improve their performance of zero-shot and cross-dataset image-text matching.
Yan Huang 0008, Yunan Zeng, Junshi Huang, Zhenhua Chai, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 GaitC3I: Robust Cross-Covariate Gait Recognition via Causal Intervention
abstract
Cross-covariate gait recognition aims to analyze a pedestrian’s gait to extract an identity representation that is invariant across varying covariates. However, prevailing methods that have achieved good results on controlled in-the-lab datasets often perform poorly on realistic datasets. In this work, we find a significant cause is that the widely used pairwise metric learning paradigm cannot correctly handle the relationship between samples from different covariate conditions. Even worse, it may yield harmful signals that inadvertently mislead models to focus on covariate-related features, particularly when covariate distributions vary across subjects. To address this issue, we propose a Cross-Covariate Causal Intervention (GaitC3I) framework, a unified causality-inspired approach aimed at enhancing the robustness of gait recognition across diverse conditions. Specifically, our method consists of two parts: 1) an effective causal intervention metric learning paradigm based on backdoor adjustment, which strategically mitigates spurious correlations induced by covariates, thus ensuring a more invariant gait representation; and 2) an annotation-free selection strategy that progressively matches each positive sample with negative samples from similar covariate conditions at various granularities. We demonstrate the effectiveness of our GaitC3I through extensive evaluation on six popular gait datasets-Gait3D, GREW, OUMVLP, CASIA-B, CCPG, and CCGR-achieving substantial improvements. Our method not only outperforms existing state-of-the-art models but also provides a systematic solution to remove the spurious correlations in gait recognition.
Jilong Wang 0010, Saihui Hou, Xianda Guo, Yan Huang 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 SliceMamba With Neural Architecture Search for Medical Image Segmentation
abstract
Despite the progress made in Mamba-based medical image segmentation models, existing methods utilizing unidirectional or multi-directional feature scanning mechanisms struggle to effectively capture dependencies between neighboring positions, limiting the discriminant representation learning of local features. These local features are crucial for medical image segmentation as they provide critical structural information about lesions and organs. To address this limitation, we propose SliceMamba, a simple yet effective locally sensitive Mamba-based medical image segmentation model. SliceMamba features an efficient Bidirectional Slicing and Scanning (BSS) module, which performs bidirectional feature slicing and employs varied scanning mechanisms for sliced features with distinct shapes. This design keeps spatially adjacent features close in the scan sequence, preserving the local structure of the image and enhancing segmentation performance. Additionally, to fit the varying sizes and shapes of lesions and organs, we introduce an Adaptive Slicing Search method that automatically identifies the optimal feature slicing method based on the characteristics of the target data. Extensive experiments on two skin lesion datasets (ISIC2017 and ISIC2018), two polyp segmentation datasets (Kvasir and ClinicDB), one ultra-wide field retinal hemorrhage segmentation dataset (UWF-RHS), and one multi-organ segmentation dataset (Synapse) demonstrate the effectiveness of our method.
Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Liang Wang 0001, Zhenghan Yang, Xibin Jia
IEEE J. Biomed. Health Informatics3
2024 Investigating Compositional Challenges in Vision-Language Models for Visual Grounding
abstract
Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks, which have been widely used for visual grounding tasks in a weakly supervised manner. However, despite the per-formance gains contributed by large vision and language pre-training, we find that state-of-the-art VLMs struggle with compositional reasoning on grounding tasks. To demonstrate this, we propose Attribute, Relation, and Pri-ority grounding (ARPGrounding) benchmark to test VLMs' compositional reasoning ability on visual grounding tasks. ARPGrounding contains 11,425 samples and evaluates the compositional understanding of VLMs in three dimensions: 1) attribute, denoting comprehension of objects' properties; 2) relation, indicating an understanding of relation between objects; 3) priority, reflecting an awareness of the part of speech associated with nouns. Using the ARPGrounding benchmark, we evaluate several mainstream VLMs. We empirically find that these models perform quite well on conventional visual grounding datasets, achieving performance comparable to or surpassing state-of-the-art methods but showing strong deficiencies in compositional reasoning. Furthermore, we propose a composition-aware fine-tuning pipeline, demonstrating the potential to lever-age cost-effective image-text annotations for enhancing the compositional understanding of VLMs in grounding tasks. Code is available at link.
Yunan Zeng, Yan Huang 0008, Zequn Jie, Zhenhua Chai, Liang Wang 0001
CVPR2
2024 Free Lunch for Gait Recognition: A Novel Relation Descriptor
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Tianzhu Zhang 0001, Liang Wang 0001
ECCV (38)3
2024 SPK: Semantic and Positional Knowledge for Zero-Shot Referring Expression Comprehension
Zetao Du, Junbo Wang 0003, Yan Huang 0008, Liang Wang 0001
ICPR (30)4
2024 Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor
abstract
Vision-and-Language Navigation (VLN) requires an agent to dynamically explore environments following natural language. The VLN agent, closely integrated into daily lives, poses a substantial threat to the security of privacy and property upon the occurrence of malicious behavior. However, this serious issue has long been overlooked. In this paper, we pioneer the exploration of an object-aware backdoored VLN, achieved by implanting object-aware backdoors during the training phase. Tailored to the unique VLN nature of cross-modality and continuous decision-making, we propose a novel backdoored VLN paradigm: IPR Backdoor. This enables the agent to act in abnormal behavior once encountering the object triggers during language-guided navigation in unseen environments, thereby executing an attack on the target scene. Experiments demonstrate the effectiveness of our method in both physical and digital spaces across different VLN agents, as well as its robustness to various visual and textual variations. Additionally, our method also well ensures navigation performance in normal scenarios with remarkable stealthiness.
Keji He, Jiawang Bai, Yan Huang 0008, Qi Wu 0001, Shutao Xia, Liang Wang 0001
NeurIPS4
2024 Memory-Adaptive Vision-and-Language Navigation
Keji He, Ya Jing, Yan Huang 0008, Zhihe Lu, Dong An 0002, Liang Wang 0001
Pattern Recognit.3
2024 Customized meta-dataset for automatic classifier accuracy evaluation
Yan Huang 0023, Zhang Zhang 0001, Yan Huang 0008, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
Pattern Recognit.3
2024 An Overview of Text-Based Person Search: Recent Advances and Future Directions
abstract
Due to the practical significance in smart video surveillance systems, Text-Based Person Search (TBPS) has been one of the research hotspots recently, which refers to searching for the interested pedestrian images given natural language sentences. To help researchers quickly grasp the developments of this important task, we comprehensively summarize the recent research advances of TBPS from two perspectives,i.e., Feature Extraction (FE) and Semantic Alignments (SA). Specifically, the FE mainly consists of pre-processing approaches and end-to-end frameworks, and the SA could be briefly divided into cross-modal attention mechanism, non-attention alignments, training objectives, and generative approaches. Afterwards, we elaborate four widely-used benchmarks and also the evaluation criterion for TBPS. And comparisons and analyses among the state-of-the-art (SOTA) solutions are provided based on these large-scale benchmarks. At last, we point out some future research directions that need to be further addressed, which will greatly facilitate the practical applications of TBPS.
Kai Niu 0002, Yanyi Liu, Yuzhou Long, Yan Huang 0008, Liang Wang 0001, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Self-Supervised Recovery and Guide for Low-Resolution Person Re-Identification
abstract
Low-resolution person re-identification is a challenging task to match low-resolution (LR) probes with high-resolution (HR) gallery images. To address the resolution gap, existing methods typically recover missing details for LR probes by super-resolution, and then match the recovered HR images (instead of the original LR probes) with gallery images. However, they usually pre-specify fixed scale factors for all LR images, and ignore that choosing a preferable scale factor for each image can recover more discriminative content and accordingly benefit the re-id performance. Moreover, these methods do not focus on learning LR representations themselves and always resort to extra recovery to handle LR probes, which is quite time-consuming during inference. To tackle these problems, we propose a Self-supervised Recovery and Guide (SRG) re-id model in this paper. Given LR images during training, our model firstrecoversmore discriminative HR images by finding out preferable scale factors, and further leverages them asguideto improve original LR representations. Through enforcing LR representations to approach the self-recovered HR guide in a self-supervised manner, our model can learn more discriminative representations for LR images. As a result, our model is able to directly handle LR probes without requiring recovery during inference, thereby reducing inference time significantly. Extensive experiments demonstrate the effectiveness of our method on four datasets.
Yan Huang 0008, Liang Wang 0001, Zikun Liu 0001
IEEE Trans. Inf. Forensics Secur.2
2024 Enhancing Person Re-Identification Performance Through In Vivo Learning
abstract
This research investigates the potential of in vivo learning to enhance visual representation learning for image-based person re-identification (re-ID). Compared to traditional self-supervised learning (which require external data), the introduced in vivo learning utilizes supervisory labels generated from pedestrian images to improve re-ID accuracy without relying on external data sources. Three carefully designed in vivo learning tasks, leveraging statistical regularities within images, are proposed without the need for laborious manual annotations. These tasks enable feature extractors to learn more comprehensive and discriminative person representations by jointly modeling various aspects of human biological structure information, contributing to enhanced re-ID performance. Notably, the method seamlessly integrates with existing re-ID frameworks, requiring minimal modifications and no additional data beyond the existing training set. Extensive experiments on diverse datasets, including Market1501, CUHK03-NP, Celeb-reID, Celeb-reid-light, PRCC, and LTCC, demonstrate substantial enhancements in rank-1 precision compared to state-of-the-art methods.
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
IEEE Trans. Image Process.1
2024 Meta Clothing Status Calibration for Long-Term Person Re-Identification
abstract
Recent studies have seen significant advancements in the field of long-term person re-identification (LT-reID) through the use of clothing-irrelevant or insensitive features. This work takes the field a step further by addressing a previously unexplored issue, the Clothing Status Distribution Shift (CSDS). CSDS refers to the differing ratios of samples with clothing changes to those without clothing changes between the training and test sets, leading to a decline in LT-reID performance. We establish a connection between the performance of LT-reID and CSDS, and argue that addressing CSDS can improve LT-reID performance. To that end, we propose a novel framework called Meta Clothing Status Calibration (MCSC), which uses meta-learning to optimize the LT-reID model. Specifically, MCSC simulates CSDS between meta-train and meta-test with meta-optimization objectives, optimizing the LT-reID model and making it robust to CSDS. This framework is designed to prevent overfitting and improve the generalization ability of the LT-reID model in the presence of CSDS. Comprehensive evaluations on seven datasets demonstrate that the proposed MCSC framework effectively handles CSDS and improves current state-of-the-art LT-reID methods on several LT-reID benchmarks.
Yan Huang 0023, Qiang Wu 0001, Zhang Zhang 0001, Caifeng Shan, Yan Huang 0008, Yi Zhong 0002, Liang Wang 0001
IEEE Trans. Image Process.5
2024 Comprehensive Attribute Prediction Learning for Person Search by Language
abstract
Person search by language refers to searching for the interested pedestrian images given natural language sentences, which requires capturing fine-grained differences to accurately distinguish different pedestrians, while still far from being well addressed by most of the current solutions. In this paper, we propose the Comprehensive Attribute Prediction Learning (CAPL) method, which explicitly carries out attribute prediction learning, for improving the modeling capabilities of fine-grained semantic attributes and obtaining more discriminative visual and textual representations. First, we construct the semantic ATTribute Vocabulary (ATT-Vocab) based on sentence analysis. Second, the complementary context-wise and attribute-wise attribute predictions are simultaneously conducted to better model the high-frequency in-vocab attributes in our In-vocab Attribute Prediction (IAP) module. Third, to additionally consider the out-of-vocab semantics, we present the Attribute Completeness Learning (ACL) module for better capturing the low-frequency attributes outside the ATT-Vocab, obtaining more comprehensive representations. Combining the IAP and ACL modules together, our CAPL method has obtained the currently state-of-the-art retrieval performance on two widely-used benchmarks, i.e., CUHK-PEDES and ICFG-PEDES datasets. Extensive experiments and analyses have been carried out to validate the effectiveness and generalization capacities of our CAPL method.
Kai Niu 0002, Linjiang Huang, Yuzhou Long, Yan Huang 0008, Liang Wang 0001, Yanning Zhang 0001
IEEE Trans. Image Process.4
2024 Text-to-Image Vehicle Re-Identification: Multi-Scale Multi-View Cross-Modal Alignment Network and a Unified Benchmark
abstract
Vehicle Re-IDentification (Re-ID) aims to retrieve the most similar images with a given query vehicle image from a set of images captured by non-overlapping cameras, and plays a crucial role in intelligent transportation systems and has made impressive advancements in recent years. In real-world scenarios, we can often acquire the text descriptions of target vehicle through witness accounts, and then manually search the image queries for vehicle Re-ID, which is time-consuming and labor-intensive. To solve this problem, this paper introduces a new fine-grained cross-modal retrieval task called text-to-image vehicle re-identification, which seeks to retrieve target vehicle images based on the given text descriptions. To bridge the significant gap between language and visual modalities, we propose a novel Multi-scale multi-view Cross-modal Alignment Network (MCANet). In particular, we incorporate view masks and multi-scale features to align image and text features in a progressive way. In addition, we design the Masked Bidirectional InfoNCE (MB-InfoNCE) loss to enhance the training stability and make the best use of negative samples. To provide an evaluation platform for text-to-image vehicle re-identification, we create a Text-to-Image Vehicle Re-Identification dataset (T2I VeRi), which contains 2465 image-text pairs from 776 vehicles with an average sentence length of 26.8 words. Extensive experiments conducted on T2I VeRi demonstrate MCANet outperforms the current state-of-art (SOTA) method by 2.2% in rank-1 accuracy.
Leqi Ding, Lei Liu 0049, Yan Huang 0008, Chenglong Li 0002, Cheng Zhang 0010, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Towards Few-shot Image Captioning with Cycle-based Compositional Semantic Enhancement Framework
abstract
Many efforts paid attention to the multi-modal task, of which image captioning is a classic work. Especially the Clip model improves the performance of image captioning; meantime, its few-shot and zero-shot problems have become a significant research project. In this work, aiming at the image captioning task, we design the new few-shot and zero-shot settings different from popular directions. The direction focuses on the impact of the exited dataset for captioning model ability. According to analysis, we discover the frequency of the word combination can directly influence the performance of the captioning model. Based on this, we define the new few-shot and zero-shot settings. In terms of this, a Cycle-based captioning framework based on data augmentation is proposed to overcome this problem, of which the novelty switcher module is the critical component. Finally, experiments demonstrate that our framework can achieve state-of-the-art performance on both traditional, few-shot and zero-shot settings.
Peng Zhang 0058, Yang Bai 0011, Jie Su 0001, Yan Huang 0008, Yang Long 0001
IJCNN4
2023 Causal Intervention for Sparse-View Gait Recognition
abstract
Gait recognition aims at identifying individuals by unique walking patterns at a long distance. However, prevailing methods suffer from a large degradation when applied to large-scale surveillance systems. We find a significant cause of this issue is that previous methods heavily rely on full-view person annotations to reduce view differences by pulling closer the anchor to positive samples from different viewpoints. But, subjects under in-the-wild scenarios usually have only a limited number of sequences from different viewpoints. As a result, the available viewpoints of each subject are sparse compared to the whole dataset, and simply minimizing intra-identity differences cannot well reducing the view differences in the whole dataset. In this work, we formulate this overlooked problem as Sparse-View Gait Recognition and provide a comprehensive analysis of it by a Structural Causal Model for causalities among latent features, view distribution, and labels. Based on our analysis, we propose a simple yet effective method that enables networks to learn a more robust representation among different views. Specifically, our method consists of two parts: 1) an effective metric learning algorithmic implementation based on the backdoor adjustment, which improves the consistency of representations among different views; 2) an unsupervised view cluster algorithm to discover and identify the most influential view contexts. We evaluate the effectiveness of our method on popular GREW, Gait3D, CASIA-B, and OU-MVLP, showing that our method consistently outperforms baselines and achieves state-of-the-art performance. The code will be available at https://github.com/wj1tr0y/GaitCSV.
Jilong Wang 0010, Saihui Hou, Yan Huang 0008, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Liang Wang 0001
ACM Multimedia3
2023 Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through complex environments based on natural language instructions. In contrast to conventional approaches, which primarily focus on the spatial domain exploration, we propose a paradigm shift toward the Fourier domain. This alternative perspective aims to enhance visual-textual matching, ultimately improving the agent's ability to understand and execute navigation tasks based on the given instructions. In this study, we first explore the significance of high-frequency information in VLN and provide evidence that it is instrumental in bolstering visual-textual matching processes. Building upon this insight, we further propose a sophisticated and versatile Frequency-enhanced Data Augmentation (FDA) technique to improve the VLN model's capability of capturing critical high-frequency information. Specifically, this approach requires the agent to navigate in environments where only a subset of high-frequency visual information corresponds with the provided textual instructions, ultimately fostering the agent's ability to selectively discern and capture pertinent high-frequency features according to the given instructions. Promising results on R2R, RxR, CVDN and REVERIE demonstrate that our FDA can be readily integrated with existing VLN approaches, improving performance without adding extra parameters, and keeping models simple and efficient. The code is available at https://github.com/hekj/FDA.
Keji He, Chenyang Si, Zhihe Lu, Yan Huang 0008, Liang Wang 0001, Xinchao Wang
NeurIPS4
2023 End-to-End Alternating Optimization for Real-World Blind Super Resolution
Zhengxiong Luo 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
Int. J. Comput. Vis.2
2023 Efficient Image and Sentence Matching
abstract
Recently, the accuracy of image and sentence matching has been continuously improved by larger and larger models. However, such large models not only need huge storage space but also slow down inference speed, which are not very suitable for low-cost devices in real-world applications. To our knowledge, this work makes the first attempt to improve the model efficiency in the context of image and sentence matching, and accordingly proposes a simple yet effective Whitened Similarity Distillation (WSD) method, which can distill cross-modal knowledge from a large teacher model to a small student model of both high efficiency and accuracy. The high efficiency is achieved by performing: 1) feature representation based on efficient backbone networks; and 2) similarity measurement in a fast N-to-N manner. However, the accuracy of such a student model is much worse than that of teacher model, because there exists very large variation inconsistency between two cross-modal similarity matrices of teacher and student models, which is hard to reduce during the similarity distillation. By performing two whitening-like transformations in the orthogonal space, the proposed WSD can reduce the large variation inconsistency more isotropically and is able to improve the accuracy of student model. We perform extensive experiments on two benchmark datasets and demonstrate the effectiveness of the proposed WSD. Compared with the teacher model, our distilled student model is 7× smaller (in model size) and 9× faster (in testing speed), only at the cost of 2% accuracy decrease.
Yan Huang 0008, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Cyclic Differentiable Architecture Search
abstract
Differentiable ARchiTecture Search, i.e., DARTS, has drawn great attention in neural architecture search. It tries to find the optimal architecture in a shallow search network and then measures its performance in a deep evaluation network. The independent optimization of the search and evaluation networks, however, leaves a room for potential improvement by allowing interaction between the two networks. To address the problematic optimization issue, we propose new joint optimization objectives and a novel Cyclic Differentiable ARchiTecture Search framework, dubbed CDARTS. Considering the structure difference, CDARTS builds a cyclic feedback mechanism between the search and evaluation networks with introspective distillation. First, the search network generates an initial architecture for evaluation, and the weights of the evaluation network are optimized. Second, the architecture weights in the search network are further optimized by the label supervision in classification, as well as the regularization from the evaluation network through feature distillation. Repeating the above cycle results in a joint optimization of the search and evaluation networks and thus enables the evolution of the architecture to fit the final evaluation network. The experiments and analysis on CIFAR, ImageNet and NATS-Bench [95] demonstrate the effectiveness of the proposed approach over the state-of-the-art ones. Specifically, in the DARTS search space, we achieve 97.52% top-1 accuracy on CIFAR10 and 76.3% top-1 accuracy on ImageNet. In the chain-structured search space, we achieve 78.2% top-1 accuracy on ImageNet, which is 1.1% higher than EfficientNet-B0. Our code and models are publicly available at https://github.com/microsoft/Cream.
Hongyuan Yu, Houwen Peng, Yan Huang 0008, Jianlong Fu, Hao Du 0006, Liang Wang 0001, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 SiamON: Siamese Occlusion-Aware Network for Visual Tracking
abstract
Occlusion has been proven to be one of the most challenging factors faced by most visual trackers. There are mainly two difficulties, the first one is that the number of occlusion samples are very limited even though collecting a large-scale training data set, and another one is how to correctly learn the features of the target when comes to occlusion situations. In this paper, we tried to solve these two problems together in our proposed model. To this end, we propose a novel Siamese Occlusion-aware Network (SiamON) for high-performance visual tracking. In particular, we predefine some soft-masks to solve the problem of fewer occlusion samples, which perceive patterns of occlusion contents at different locations and take these masks as the conditions to guide occlusion-aware feature learning. Meanwhile, we propose a target-aware attention mechanism allows the model to pay more attention to the target and further weaken the impact of occlusion. Extensive experiments on several popular benchmarks show that our tracking method exceeds many state-of-the-art trackers especially in the presence of occlusion and meets the requirements of real-time.
Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Caifeng Shan, Liang Wang 0001, Chenglong Li 0002
IEEE Trans. Circuits Syst. Video Technol.3
2023 Efficient Token-Guided Image-Text Retrieval With Consistent Multimodal Contrastive Training
abstract
Image-text retrieval is a central problem for understanding the semantic relationship between vision and language, and serves as the basis for various visual and language tasks. Most previous works either simply learn coarse-grained representations of the overall image and text, or elaborately establish the correspondence between image regions or pixels and text words. However, the close relations between coarse- and fine-grained representations for each modality are important for image-text retrieval but almost neglected. As a result, such previous works inevitably suffer from low retrieval accuracy or heavy computational cost. In this work, we address image-text retrieval from a novel perspective by combining coarse- and fine-grained representation learning into a unified framework. This framework is consistent with human cognition, as humans simultaneously pay attention to the entire sample and regional elements to understand the semantic content. To this end, a Token-Guided Dual Transformer (TGDT) architecture which consists of two homogeneous branches for image and text modalities, respectively, is proposed for image-text retrieval. The TGDT incorporates both coarse- and fine-grained retrievals into a unified framework and beneficially leverages the advantages of both retrieval approaches. A novel training objective called Consistent Multimodal Contrastive (CMC) loss is proposed accordingly to ensure the intra- and inter-modal semantic consistencies between images and texts in the common embedding space. Equipped with a two-stage inference method based on the mixed global and local cross-modal similarity, the proposed method achieves state-of-the-art retrieval performances with extremely low inference time when compared with representative recent approaches. Code is publicly available: github.com/LCFractal/TGDT.
Chong Liu 0002, Yuqi Zhang 0001, Hongsong Wang 0001, Fan Wang 0019, Yan Huang 0008, Yidong Shen, Liang Wang 0001
IEEE Trans. Image Process.6
2023 A Reconstruction-Based Visual-Acoustic-Semantic Embedding Method for Speech-Image Retrieval
abstract
Speech-image retrieval aims at learning the relevance between image and speech.Prior approaches are mainly based on bi-modal contrastive learning, which can not alleviate the cross-modal heterogeneous issue between visual and acoustic modalities well. To address this issue, we propose a visual-acoustic-semantic embedding (VASE) method. First, we propose a tri-modal ranking loss by taking advantage of semantic information corresponding to the acoustic data, which introduces the auxiliary alignment to enhance the alignment between image and speech. Second, we introduce a cycle-consistency loss based on feature reconstruction. It can further alleviate the heterogeneous issue between different data modalities (e.g., visual-acoustic, visual-textual and acoustic-textual). Extensive experiments have demonstrated the effectiveness of our proposed method. In addition, our VASE model achieves state-of-the-art performance on the speech-image retrieval task on the Flickr8K [Harwath and Glass, 2015]s and Places [Harwathet al., 2018] datasets.
Wei Tang 0016, Yan Huang 0008, Yiwen Luo, Liang Wang 0001
IEEE Trans. Multim.3
2022 Generalizable Person Re-identification via Self-Supervised Batch Norm Test-Time Adaption
abstract
In this paper, we investigate the generalization problem of person re-identification (re-id), whose major challenge is the distribution shift on an unseen domain. As an important tool of regularizing the distribution, batch normalization (BN) has been widely used in existing methods. However, they neglect that BN is severely biased to the training domain and inevitably suffers the performance drop if directly generalized without being updated. To tackle this issue, we propose Batch Norm Test-time Adaption (BNTA), a novel re-id framework that applies the self-supervised strategy to update BN parameters adaptively. Specifically, BNTA quickly explores the domain-aware information within unlabeled target data before inference, and accordingly modulates the feature distribution normalized by BN to adapt to the target domain. This is accomplished by two designed self-supervised auxiliary tasks, namely part positioning and part nearest neighbor matching, which help the model mine the domain-aware information with respect to the structure and identity of body parts, respectively. To demonstrate the effectiveness of our method, we conduct extensive experiments on three re-id datasets and confirm the superior performance to the state-of-the-art methods.
Chenyang Si, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
AAAI3
2022 3D Shape Temporal Aggregation for Video-Based Clothing-Change Person Re-identification
Yan Huang 0008, Shaogang Gong, Liang Wang 0001, Tieniu Tan
ACCV (5)2
2022 Regularized Graph Structure Learning with Semantic Knowledge for Multi-variates Time-Series Forecasting
abstract
Multivariate time-series forecasting is a critical task for many applications, and graph time-series network is widely studied due to its capability to capture the spatial-temporal correlation simultaneously. However, most existing works focus more on learning with the explicit prior graph structure, while ignoring potential information from the implicit graph structure, yielding incomplete structure modeling. Some recent works attempts to learn the intrinsic or implicit graph structure directly, while lacking a way to combine explicit prior structure with implicit structure together. In this paper, we propose Regularized Graph Structure Learning (RGSL) model to incorporate both explicit prior structure and implicit structure together, and learn the forecasting deep networks along with the graph structure. RGSL consists of two innovative modules. First, we derive an implicit dense similarity matrix through node embedding, and learn the sparse graph structure using the Regularized Graph Generation (RGG) based on the Gumbel Softmax trick. Second, we propose a Laplacian Matrix Mixed-up Module (LM3) to fuse the explicit graph and implicit graph together. We conduct experiments on three real-word datasets. Results show that the proposed RGSL model outperforms existing graph forecasting algorithms with a notable margin, while learning meaningful graph structure simultaneously. Our code and models are made publicly available at https://github.com/alipay/RGSL.git.
Hongyuan Yu, Weichen Yu, Yan Huang 0008, Liang Wang 0001, Alex X. Liu
IJCAI5
2022 Cross-modal Co-occurrence Attributes Alignments for Person Search by Language
abstract
Person search by language refers to retrieving the interested pedestrian images based on a free-form natural language description, which has important applications in smart video surveillance. Although great efforts have been made to align images with sentences, the challenge of reporting bias, i.e., attributes are only partially matched across modalities, still incurs large noise and influences the accurate retrieval seriously. To address this challenge, we propose a novel cross-modal matching method named Cross-modal Co-occurrence Attributes Alignments (C2A2), which can better deal with noise and obtain significant improvements in retrieval performance for person search by language. First, we construct visual and textual attribute dictionaries relying on matrix decomposition, and carry out cross-modal alignments using denoising reconstruction features to address the noise from pedestrian-unrelated elements. Second, we re-gather pixels of image and words of sentence under the guidance of learned attribute dictionaries, to adaptively constitute more discriminative co-occurrence attributes in both modalities. And the re-gathered co-occurrence attributes are carefully captured by imposing explicit cross-modal one-to-one alignments which consider relations across modalities, better alleviating the noise from non-correspondence attributes. The whole C_2A_2 method can be trained end-to-end without any pre-processing, i.e., requiring negligible additional computation overheads. It significantly outperforms the existing solutions, and finally achieves the new state-of-the-art retrieval performance on two large-scale benchmarks, CUHK-PEDES and RSTPReid datasets.
Kai Niu 0002, Linjiang Huang, Yan Huang 0008, Peng Wang 0015, Liang Wang 0001, Yanning Zhang 0001
ACM Multimedia3
2022 Generalized Inter-class Loss for Gait Recognition
abstract
Gait recognition is a unique biometric technique that can be performed at a long distance non-cooperatively and has broad applications in public safety and intelligent traffic systems. The previous gait works focus more on minimizing the intra-class variance while ignoring the significance of constraining inter-class variance. To this end, we propose a generalized inter-class loss that resolves the inter-class variance from both sample-level feature distribution and class-level feature distribution. Instead of equal penalty strength on pair scores, the proposed loss optimizes sample-level inter-class feature distribution by dynamically adjusting the pairwise weight. Further, in class-level distribution, the proposed loss adds a constraint on the uniformity of inter-class feature distribution, which forces the feature representations to approximate a hypersphere and keep maximal inter-class variance. In addition, the proposed method automatically adjusts the margin between classes which enables the inter-class feature distribution to be more flexible. The proposed method can be generalized to different gait recognition networks and achieves significant improvements. We conduct a series of experiments on CASIA-B and OUMVLP, and the experimental results show that the proposed loss can significantly improve the performance and achieves the state-of-the-art performances.
Weichen Yu, Hongyuan Yu, Yan Huang 0008, Liang Wang 0001
ACM Multimedia3
2022 Part Based Interaction Learning for Group Activity Recognition
abstract
Group activity recognition is a subject with broad applications, and its main challenge is to model the interactions between individuals. Existing algorithms mostly model the interactions merely based on holistic features of persons, which completely ignore the local details and local interactions that could be significant for recognition. In this paper, we propose a novel part based interaction learning algorithm for group activity recognition. Our proposed algorithm introduces both the physical structural information and fine-grained contextual information into representations, through exploring the intraand inter-actor part interactions. Specifically, a dual-branch framework is adopted to extract the appearance and motion features respectively. For each branch, we utilize the key point detection technique for proper part division and then extract the part features. The part features are further enhanced by the transformers for intra- and inter-actor part interactions, and are lastly used for group activity recognition. Comparison with the state-of-the-arts on two public datasets demonstrate the effectiveness of our proposed algorithm.
Xianglong Lang, Ye Xiang, Yan Huang 0008, Lifang Wu
MMSP4
2022 MACK: Multimodal Aligned Conceptual Knowledge for Unpaired Image-text Matching
abstract
Recently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which are trained on millions or billions of paired images and texts. Different from them, this paper studies a new scenario as unpaired image-text matching, in which paired images and texts are assumed to be unavailable during model training. To deal with this, we propose a simple yet effective method namely Multimodal Aligned Conceptual Knowledge (MACK), which is inspired by the knowledge use in human brain. It can be directly used as general knowledge to correlate images and texts even without model training, or further fine-tuned based on unpaired images and texts to better generalize to certain datasets. In addition, we extend it as a re-ranking method, which can be easily combined with existing image-text matching models to substantially improve their performance.
Yan Huang 0008, Yunan Zeng, Liang Wang 0001
NeurIPS1
2022 Learning a Robust Part-Aware Monocular 3D Human Pose Estimator via Neural Architecture Search
Zerui Chen, Yan Huang 0008, Hongyuan Yu, Liang Wang 0001
Int. J. Comput. Vis.2
2022 Two-Branch Relational Prototypical Network for Weakly Supervised Temporal Action Localization
abstract
As a challenging task of high-level video understanding, weakly supervised temporal action localization has attracted more attention recently. With only video-level category labels, this task should indistinguishably identify the background and action categories frame by frame. However, it is non-trivial to achieve this in untrimmed videos, due to the unconstrained background, complex and multi-label actions. With the observation that these difficulties are mainly brought by the large variations within background and actions, we propose to address these challenges from the perspective of modeling variations. Moreover, it is desired to further reduce the variations, or learn compact features, so as to cast the problem of background identification as rejecting background and alleviate the contradiction between classification and detection. Accordingly, in this paper, we propose a two-branch relational prototypical network. The first branch, namely action-branch, adopts class-wise prototypes and mainly acts as an auxiliary to introduce priori knowledge about label dependencies and be a guide for the second branch. Meanwhile, the second branch, namely sub-branch, starts with multiple prototypes, namely sub-prototypes, to enable a powerful ability of modeling variations. As a further benefit, we elaborately design a multi-label clustering loss based on the sub-prototypes to learn compact features under the multi-label setting. The two branches are associated using the correspondences between two types of prototypes, leading to a special two-stage classifier in the s-branch, on the other hand, the two branches serve as regularization terms to each other, improving the final performance. Ablation studies find that the proposed model is capable of modeling classes with large variations and learning compact features. Extensive experimental evaluations on Thumos14, MultiThumos and ActivityNet datasets demonstrate the effectiveness of the proposed method and superior performance over state-of-the-art approaches.
Linjiang Huang, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Few-Shot Image and Sentence Matching via Aligned Cross-Modal Memory
abstract
Image and sentence matching has attracted much attention recently, and many effective methods have been proposed to deal with it. But even the current state-of-the-arts still cannot well associate those challenging pairs of images and sentences containing few-shot content in their regions and words. In fact, such a few-shot matching problem is seldom studied and has become a bottleneck for further performance improvement in real-world applications. In this work, we formulate this challenging problem as few-shot image and sentence matching, and accordingly propose an Aligned Cross-Modal Memory (ACMM) model to deal with it. The model can not only softly align few-shot regions and words in a weakly-supervised manner, but also persistently store and update cross-modal prototypical representations of few-shot classes as references, without using any groundtruth region-word correspondence. The model can also adaptively balance the relative importance between few-shot and common content in the image and sentence, which leads to better measurement of overall similarity. We perform extensive experiments in terms of both few-shot and conventional image and sentence matching, and demonstrate the effectiveness of the proposed model by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Jingdong Wang 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Distilled light GaitSet: Towards scalable gait recognition
Xu Song, Yan Huang 0008, Caifeng Shan, Jilong Wang 0010
Pattern Recognit. Lett.2
2022 Joint Token and Feature Alignment Framework for Text-Based Person Search
abstract
Text-based person search is a challenging crossmodal retrieval task. Existing works reduce the inter-modality and intra-class gaps by aligning local features extracted from image and text modalities, which easily lead to mismatching problems due to the lack of annotation information. Besides, it is sub-optimal to reduce two gaps simultaneously in the same feature space. This work proposes a novel joint token and feature alignment framework to reduce the inter-modality and intraclass gaps progressively. Specifically, we first build a dual-path feature learning network to extract features and conduct feature alignment to reduce the inter-modality gap. Second, we design a text generation module to generate token sequences using visual features, and then token alignment is performed to reduce the intra-class gap. Last, a fusion interaction module is introduced to further eliminate the modality heterogeneity using the strategy of multi-stage feature fusion. Extensive experiments on the CUHKPEDES dataset demonstrate the effectiveness of our model, which significantly outperforms previous state-of-the-art methods.
Shangze Li, Andong Lu, Yan Huang 0008, Chenglong Li 0002, Liang Wang 0001
IEEE Signal Process. Lett.3
2022 Actor and Action Modular Network for Text-Based Video Segmentation
abstract
Text-based video segmentation aims to segment an actor in video sequences by specifying the actor and its performing action with a textual query. Previous methods fail to explicitly align the video content with the textual query in a fine-grained manner according to the actor and its action, due to the problem of semantic asymmetry. The semantic asymmetry implies that two modalities contain different amounts of semantic information during the multi-modal fusion process. To alleviate this problem, we propose a novel actor and action modular network that individually localizes the actor and its action in two separate modules. Specifically, we first learn the actor-/action-related content from the video and textual query, and then match them in a symmetrical manner to localize the target tube. The target tube contains the desired actor and action which is then fed into a fully convolutional network to predict segmentation masks of the actor. Our method also establishes the association of objects cross multiple frames with the proposed temporal proposal aggregation mechanism. This enables our method to segment the video effectively and keep the temporal consistency of predictions. The whole model is allowed for joint learning of the actor-action matching and segmentation, as well as achieves the state-of-the-art performance for both single-frame segmentation and full video segmentation on A2D Sentences and J-HMDB Sentences datasets.
Yan Huang 0008, Kai Niu 0002, Linjiang Huang, Zhanyu Ma, Liang Wang 0001
IEEE Trans. Image Process.2
2021 Rethinking the Heatmap Regression for Bottom-Up Human Pose Estimation
abstract
Heatmap regression has become the most prevalent choice for nowadays human pose estimation methods. The ground-truth heatmaps are usually constructed via covering all skeletal keypoints by 2D gaussian kernels. The standard deviations of these kernels are fixed. However, for bottom-up methods, which need to handle a large variance of human scales and labeling ambiguities, the current practice seems unreasonable. To better cope with these problems, we propose the scale-adaptive heatmap regression (SAHR) method, which can adaptively adjust the standard deviation for each keypoint. In this way, SAHR is more tolerant of various human scales and labeling ambiguities. However, SAHR may aggravate the imbalance between fore-background samples, which potentially hurts the improvement of SAHR. Thus, we further introduce the weight-adaptive heatmap regression (WAHR) to help balance the fore-background samples. Extensive experiments show that SAHR together with WAHR largely improves the accuracy of bottom-up human pose estimation. As a result, we finally outperform the state-of-the-art model by +1.5AP and achieve 72.0AP on COCO test-dev2017, which is comparable with the performances of most top-down methods. Source codes are available at https://github.com/greatlog/SWAHR-HumanPose.
Zhengxiong Luo 0001, Zhicheng Wang 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan, Erjin Zhou
CVPR3
2021 CMF: Cascaded Multi-Model Fusion For Referring Image Segmentation
abstract
In this work, we address the task of referring image segmentation (RIS), which aims at predicting a segmentation mask for the object described by a natural language expression. Most existing methods focus on establishing unidirectional or directional relationships between visual and linguistic features to associate two modalities together, while the multi-scale context is ignored or insufficiently modeled. Multi-scale context is crucial to localize and segment those objects that have large scale variations during the multi-modal fusion process. To solve this problem, we propose a simple yet effective Cascaded Multi-modal Fusion (CMF) module, which stacks multiple atrous convolutional layers in parallel and further introduces a cascaded branch to fuse visual and linguistic features. The cascaded branch can progressively integrate multi-scale contextual information and facilitate the alignment of two modalities during the multi-modal fusion process. Experimental results on four benchmark datasets demonstrate that our method outperforms most state-of-the-art methods. Code is available at https://github.com/jianhua2022/CMF-Refseg.
Yan Huang 0008, Zhanyu Ma, Liang Wang 0001
ICIP2
2021 Efficient Human Pose Estimation by Learning Deeply Aggregated Representations
abstract
In this paper, we propose an efficient human pose estimation network (DANet) by learning deeply aggregated representations. Most existing models explore multi-scale infonnation mainly from features with different spatial sizes. Powerful multi-scale representations usually rely on the cascaded pyramid framework. This framework largely boosts the performance but in the meanwhile makes networks very deep and complex. Instead, we focus on exploiting multi-scale information from layers with different receptive-field sizes and then making full of use this infonnation by improving the fusion method. Specifically, we propose an orthogonal attention block (OAB) and a second-order fusion unit (SFU). The OAB learns multi-scale infonnation from different layers and enhances them by encouraging them to be diverse. The SFU adaptively selects and fuses diverse multi-scale infonnation and suppress the redundant ones. With the help of OAB and SFU, our networks could achieve comparable or even better accuracy with much smaller model complexity. Specifically, our DANet-72 achieves 71.0 in AP score on COCO val2017 with only 1.0G FLOPS. Its speed on a CPU platfonn achieves 58 Persons-Per-Second (PPS).
Zhengxiong Luo 0001, Zhicheng Wang 0001, Yuanhao Cai, Guan'an Wang, Liang Wang 0001, Yan Huang 0008, Erjin Zhou, Tieniu Tan, Jian Sun 0001
ICME6
2021 Neighbor-view Enhanced Model for Vision and Language Navigation
abstract
Vision and Language Navigation (VLN) requires an agent to navigate to a target location by following natural language instructions. Most of existing works represent a navigation candidate by the feature of the corresponding single view where the candidate lies in. However, an instruction may mention landmarks out of the single view as references, which might lead to failures of textual-visual matching of existing methods. In this work, we propose a multi-module Neighbor-View Enhanced Model (NvEM) to adaptively incorporate visual contexts from neighbor views for better textual-visual matching. Specifically, our NvEM utilizes a subject module and a reference module to collect contexts from neighbor views. The subject module fuses neighbor views at a global level, and the reference module fuses neighbor objects at a local level. Subjects and references are adaptively determined via attention mechanisms. Our model also includes an action module to utilize the strong orientation guidance (e.g., "turn left'') in instructions. Each module predicts navigation action separately and their weighted sum is used for predicting the final action. Extensive experimental results demonstrate the effectiveness of the proposed method on the R2R and R4R benchmarks against several state-of-the-art navigators, and NvEM even beats some pre-training ones. Our code is available at https://github.com/MarSaKi/NvEM.
Dong An 0002, Yuankai Qi, Yan Huang 0008, Qi Wu 0001, Liang Wang 0001, Tieniu Tan
ACM Multimedia3
2021 Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision
abstract
In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper, we address the cross-modal alignment challenge from the perspective of fine-grain. Firstly, to alleviate weak cross-modal alignment supervision from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset, namely Landmark-RxR. Secondly, to further enhance local cross-modal alignment under fine-grained supervision, we investigate the focal-oriented rewards with soft and hard forms, by focusing on the critical points sampled from fine-grained Landmark-RxR. Moreover, to fully evaluate the navigation process, we also propose a re-initialization mechanism that makes metrics insensitive to difficult points, which can cause the agent to deviate from the correct trajectories. Experimental results show that our agent has superior navigation performance on Landmark-RxR, en-RxR and R2R. Our dataset and code are available at https://github.com/hekj/Landmark-RxR.
Keji He, Yan Huang 0008, Qi Wu 0001, Dong An 0002, Shuanglin Sima, Liang Wang 0001
NeurIPS2
2021 Mask-guided contrastive attention and two-stream metric co-learning for person Re-identification
Chunfeng Song, Caifeng Shan, Yan Huang 0008, Liang Wang 0001
Neurocomputing3
2021 Adaptive super-resolution for person re-identification with low-resolution images
Yan Huang 0008, Chunfeng Song, Liang Wang 0001, Tieniu Tan
Pattern Recognit.2
2021 End-to-end video text detection with online tracking
Hongyuan Yu, Yan Huang 0008, Lihong Pi, Chengquan Zhang, Liang Wang 0001
Pattern Recognit.2
2021 Attribute Prototype Learning for Interactive Face Retrieval
abstract
Interactive face retrieval aims at finding target subjects in face databases through human and machine interaction, which involves user feedback based on human perception and machine similarity measure in feature spaces. In this article, we propose an attribute prototype learning method to tackle the semantic gap between human and machine in face perception for fast interactive face retrieval. We reformulate the theoretical explanation of the interactive retrieval model and develop the algorithm of the heuristic solution of the model. Each module of the prototype model is learned with a set of identity-related facial attributes. The outputs of the prototype modules form the semantic representation. To adapt the prototype models across different databases, we propose a transfer selection algorithm based on the coherence measurements in interactive face retrieval. Coherence analysis proves that the proposed attribute prototype representation can effectively narrow down the semantic gap even in the case of cross-database transfer learning. The prototype representation can effectively reduce the feature dimension in the retrieval process. Real user retrieval with the Bayesian relevance feedback model shows that attribute prototype space is superior to low-level feature space and proves that interactive retrieval with attribute prototype representation can converge fast in large face databases.
Yuchun Fang, Zhengye Xiao, Yan Huang 0008, Liang Wang 0001, Nozha Boujemaa, Donald Geman
IEEE Trans. Inf. Forensics Secur.4
2021 Modeling Sub-Actions for Weakly Supervised Temporal Action Localization
abstract
As a challenging task of high-level video understanding, weakly supervised temporal action localization has attracted more attention recently. Due to the usage of video-level category labels, this task is usually formulated as the task of classification, which always suffers from the contradiction between classification and detection. In this paper, we describe a novel approach to alleviate the contradiction for detecting more complete action instances by explicitly modeling sub-actions. Our method makes use of three innovations to model the latent sub-actions. First, our framework uses prototypes to represent sub-actions, which can be automatically learned in an end-to-end way. Second, we regard the relations among sub-actions as a graph, and construct the correspondences between sub-actions and actions by the graph pooling operation. Doing so not only makes the sub-actions inter-dependent to facilitate the multi-label setting, but also naturally use the video-level labels as weak supervision. Third, we devise three complementary loss functions, namely, representation loss, balance loss and relation loss to ensure the learned sub-actions are diverse and have clear semantic meanings. Experimental results on THUMOS14 and ActivityNet1.3 datasets demonstrate the effectiveness of our method and superior performance over state-of-the-art approaches.
Linjiang Huang, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
IEEE Trans. Image Process.2
2020 Part-Level Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Recently, graph convolutional networks have achieved remarkable performance for skeleton-based action recognition. In this work, we identify a problem posed by the GCNs for skeleton-based action recognition, namely part-level action modeling. To address this problem, a novel Part-Level Graph Convolutional Network (PL-GCN) is proposed to capture part-level information of skeletons. Different from previous methods, the partition of body parts is learnable rather than manually defined. We propose two part-level blocks, namely Part Relation block (PR block) and Part Attention block (PA block), which are achieved by two differentiable operations, namely graph pooling operation and graph unpooling operation. The PR block aims at learning high-level relations between body parts while the PA block aims at highlighting the important body parts in the action. Integrating the original GCN with the two blocks, the PL-GCN can learn both part-level and joint-level information of the action. Extensive experiments on two benchmark datasets show the state-of-the-art performance on skeleton-based action recognition and demonstrate the effectiveness of the proposed method.
Linjiang Huang, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
AAAI2
2020 Relational Prototypical Network for Weakly Supervised Temporal Action Localization
abstract
In this paper, we propose a weakly supervised temporal action localization method on untrimmed videos based on prototypical networks. We observe two challenges posed by weakly supervision, namely action-background separation and action relation construction. Unlike the previous method, we propose to achieve action-background separation only by the original videos. To achieve this, a clustering loss is adopted to separate actions from backgrounds and learn intra-compact features, which helps in detecting complete action instances. Besides, a similarity weighting module is devised to further separate actions from backgrounds. To effectively identify actions, we propose to construct relations among actions for prototype learning. A GCN-based prototype embedding module is introduced to generate relational prototypes. Experiments on THUMOS14 and ActivityNet1.2 datasets show that our method outperforms the state-of-the-art methods.
Linjiang Huang, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
AAAI2
2020 Towards Part-Aware Monocular 3D Human Pose Estimation: An Architecture Search Approach
Zerui Chen, Yan Huang 0008, Hongyuan Yu, Yiru Guo, Liang Wang 0001
ECCV (3)2
2020 Prediction and Recovery for Adaptive Low-Resolution Person Re-Identification
Yan Huang 0008, Zerui Chen, Liang Wang 0001, Tieniu Tan
ECCV (26)2
2020 On the Robustness of 3D Human Pose Estimation
abstract
It is widely shown that Convolutional Neural Networks (CNNs) are vulnerable to adversarial examples on most recognition tasks, such as image classification and segmentation. However, few work studies the more complicated task - 3D human pose estimation. This task often requires large-scale datasets, specialized network architectures, and it can be solved either from single-view RGB images or from multi-view RGB images. In this paper, we make the first attempt to investigate the robustness of current state-of-the-art 3D human pose estimation methods. To this end, we build four representative baseline models, where most of the current methods can be generally classified as one of them. Furthermore, we design targeted adversarial attacks to detect whether 3D pose estimators are robust to different camera parameters. For different types of methods, we present a comprehensive study of their robustness on the large-scale Human3.6M benchmark. Our work shows that different methods vary significantly in their resistance to adversarial attacks. Through extensive experiments, we show that multi-view 3D pose estimators can be more vulnerable to adversarial examples. We believe that our efforts can shed light on future works to design more robust 3D human pose estimators.
Zerui Chen, Yan Huang 0008, Liang Wang 0001
ICPR2
2020 Efficient Super Resolution by Recursive Aggregation
abstract
Deep neural networks have achieved remarkable results on image super resolution (SR), but the efficiency problem of deep SR networks is rarely studied. We experimentally find that many sequentially stacked convolutional blocks in nowadays SR networks are far from being fully optimized, which largely damages their overall efficiency. It indicates that comparable or even better results could be achieved with less but sufficiently optimized blocks. In this paper, we try to construct more efficient SR model via the proposed recursive aggregation network (RAN). It recursively aggregates convolutional blocks in different orders, and avoids too many sequentially stacked blocks. In this way, multiple shortcuts are introduced in RAN, and help gradients easier flow to all inner layers, even for very deep SR networks. As a result, all blocks in RAN can be better optimized, thus RAN can achieve better performance with smaller model size than existing methods.
Zhengxiong Luo 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
ICPR2
2020 VSR++: Improving Visual Semantic Reasoning for Fine-Grained Image-Text Matching
abstract
Image-text matching has made great progresses recently, but there still remains challenges in fine-grained matching. To deal with this problem, we propose an Improved Visual Semantic Reasoning model (VSR++), which jointly models 1) global alignment between images and texts and 2) local correspondence between regions and words in a unified framework. To exploit their complementary advantages, we also develop a suitable learning strategy to balance their relative importance. As a result, our model can distinguish image regions and text words in a fine-grained level, and thus achieves the current state-of-the-art performance on two benchmark datasets.
Yan Huang 0008, Dongbo Zhang 0003, Zerui Chen, Liang Wang 0001
ICPR2
2020 Textual Dependency Embedding for Person Search by Language
abstract
Person search by language aims to associate the pedestrian images with free-form natural language descriptions. Although great efforts have been made to align images with sentences, most researchers neglect the difficulty of long-distance dependency modeling in textual encoding, which is very important for solving this problem because the description sentences are always long and have complex structures for distinguishing different pedestrians. In this work, we focus on the long-distance dependencies in a sentence for better textual encoding, and accordingly propose the Textual Dependency Embedding (TDE) method. We first employ the sentence analysis tools to figure out the long-distance syntactic dependencies from a dependent to its governor in a sentence. Then we embed the dependent representations to their governor adaptively in our Governor-guided Dependent Attention Module (GDAM) to model these long-distance relations. After that, we further consider the dependency types, which also tell the importance of different dependents semantically, and embed them together with the dependents' features to clarify their inequivalent contributions to their governor. Extensive experiments and analysis on person search by language and image-text matching have validated the effectiveness of our method, and we have obtained the state-of-the-art performance on the CUHK-PEDES and Flickr30K datasets.
Kai Niu 0002, Yan Huang 0008, Liang Wang 0001
ACM Multimedia2
2020 Unfolding the Alternating Optimization for Blind Super Resolution
abstract
Previous methods decompose blind super resolution (SR) problem into two sequential steps: \textit{i}) estimating blur kernel from given low-resolution (LR) image and \textit{ii}) restoring SR image based on estimated kernel. This two-step solution involves two independently trained models, which may not well compatible with each other. Small estimation error of the first step could cause severe performance drop of the second one. While on the other hand, the first step can only utilize limited information from LR image, which makes it difficult to predict highly accurate blur kernel. Towards these issues, instead of considering these two steps separately, we adopt an alternating optimization algorithm, which can estimate blur kernel and restore SR image in a single model. Specifically, we design two convolutional neural modules, namely \textit{Restorer} and \textit{Estimator}. \textit{Restorer} restores SR image based on predicted kernel, and \textit{Estimator} estimates blur kernel with the help of restored SR image. We alternate these two modules repeatedly and unfold this process to form an end-to-end trainable network. In this way, \textit{Estimator} utilizes information from both LR and SR images, which makes the estimation of blur kernel easier. More importantly, \textit{Restorer} is trained with the kernel estimated by \textit{Estimator}, instead of ground-truth kernel, thus \textit{Restorer} could be more tolerant to the estimation error of \textit{Estimator}. Extensive experiments on synthetic datasets and real-world images show that our model can largely outperform state-of-the-art methods and produce more visually favorable results at much higher speed. The source code will be publicly available.
Zhengxiong Luo 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
NeurIPS2
2020 Global Context Enhanced Multi-modal Fusion for Referring Image Segmentation
Yan Huang 0008, Linjiang Huang, Yunbo Wang, Zhanyu Ma, Liang Wang 0001
PRCV (1)2
2020 Frame-GAN: Increasing the frame rate of gait videos with generative adversarial networks
Hong Ai, Chunfeng Song, Yan Huang 0008, Liang Wang 0001
Neurocomputing5
2020 Image and Sentence Matching via Semantic Concepts and Order Learning
abstract
Image and sentence matching has made great progress recently, but it remains challenging due to the existing large visual-semantic discrepancy. This mainly arises from two aspects: 1) images consist of unstructured content which is not semantically abstract as the words in the sentences, so they are not directly comparable, and 2) arranging semantic concepts in different semantic order could lead to quite diverse meanings. The words in the sentences are sequentially arranged in a grammatical manner, while the semantic concepts in the images are usually unorganized. In this work, we propose a semantic concepts and order learning framework for image and sentence matching, which can improve the image representation by first predicting semantic concepts and then organizing them in a correct semantic order. Given an image, we first use a multi-regional multi-label CNN to predict its included semantic concepts in terms of object, property and action. These word-level semantic concepts are directly comparable with the words of noun, adjective and verb in the matched sentence. Then, to organize these concepts and make them express similar meanings as the matched sentence, we use a context-modulated attentional LSTM to learn the semantic order. It regards the predicted semantic concepts and image global scene as context at each timestep, and selectively attends to concept-related image regions by referring to the context in a sequential order. To further enhance the semantic order, we perform additional sentence generation on the image representation, by using the groundtruth order in the matched sentence as supervision. After obtaining the improved image representation, we learn the sentence representation with a conventional LSTM, and then jointly perform image and sentence matching and sentence generation for model learning. Extensive experiments demonstrate the effectiveness of our learned semantic concepts and order, by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Qi Wu 0001, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Re-ranking image-text matching by adaptive metric fusion
Kai Niu 0002, Yan Huang 0008, Liang Wang 0001
Pattern Recognit.2
2020 Long video question answering: A Matching-guided Attention Model
Weining Wang 0001, Yan Huang 0008, Liang Wang 0001
Pattern Recognit.2
2020 Improving Description-Based Person Re-Identification by Multi-Granularity Image-Text Alignments
abstract
Description-based person re-identification (Re-id) is an important task in video surveillance that requires discriminative cross-modal representations to distinguish different people. It is difficult to directly measure the similarity between images and descriptions due to the modality heterogeneity (the crossmodal problem). And all samples belonging to a single category (the fine-grained problem) makes this task even harder than the conventional image-description matching task. In this paper, we propose a Multi-granularity Image-text Alignments (MIA) model to alleviate the cross-modal fine-grained problem for better similarity evaluation in description-based person Re-id. Specifically, three different granularities, i.e., global-global, global-local and local-local alignments are carried out hierarchically. Firstly, the global-global alignment in the Global Contrast (GC) module is for matching the global contexts of images and descriptions. Secondly, the global-local alignment employs the potential relations between local components and global contexts to highlight the distinguishable components while eliminating the uninvolved ones adaptively in the Relation-guided Global-local Alignment (RGA) module. Thirdly, as for the local-local alignment, we match visual human parts with noun phrases in the Bi-directional Fine-grained Matching (BFM) module. The whole network combining multiple granularities can be end-to-end trained without complex preprocessing. To address the difficulties in training the combination of multiple granularities, an effective step training strategy is proposed to train these granularities step-by-step. Extensive experiments and analysis have shown that our method obtains the state-of-the-art performance on the CUHK-PEDES dataset and outperforms the previous methods by a significant margin.
Kai Niu 0002, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
IEEE Trans. Image Process.2
2019 Few-Shot Image and Sentence Matching via Gated Visual-Semantic Embedding
abstract
Although image and sentence matching has been widely studied, its intrinsic few-shot problem is commonly ignored, which has become a bottleneck for further performance improvement. In this work, we focus on this challenging problem of few-shot image and sentence matching, and propose a Gated Visual-Semantic Embedding (GVSE) model to deal with it. The model consists of three corporative modules in terms of uncommon VSE, common VSE, and gated metric fusion. The uncommon VSE exploits external auxiliary resources to extract generic features for representing uncommon instances and words in images and sentences, and then integrates them by modeling their semantic relation to obtain global representations for association analysis. To better model other common instances and words in rest content of images and sentences, the common VSE learns their discriminative representations directly from scratch. After obtaining two similarity metrics from the two VSE modules with different advantages, the gated metric fusion module adaptively fuses them by automatically balancing their relative importance. Based on the fused metric, we perform extensive experiments in terms of few-shot and conventional image and sentence matching, and demonstrate the effectiveness of the proposed model by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Yang Long 0001, Liang Wang 0001
AAAI1
2019 Learning Depth-aware Heatmaps for 3D Human Pose Estimation in the Wild
Zerui Chen, Yiru Guo, Yan Huang 0008, Liang Wang 0001
BMVC3
2019 Local Relationship Learning With Person-Specific Shape Regularization for Facial Action Unit Detection
abstract
Encoding individual facial expressions via action units (AUs) coded by the Facial Action Coding System (FACS) has been found to be an effective approach in resolving the ambiguity issue among different expressions. While a number of methods have been proposed for AU detection, robust AU detection in the wild remains a challenging problem because of the diverse baseline AU intensities across individual subjects, and the weakness of appearance signal of AUs. To resolve these issues, in this work, we propose a novel AU detection method by utilizing local information and the relationship of individual local face regions. Through such a local relationship learning, we expect to utilize rich local information to improve the AU detection robustness against the potential perceptual inconsistency of individual local regions. In addition, considering the diversity in the baseline AU intensities of individual subjects, we further regularize local relationship learning via person-specific face shape information, i.e., reducing the influence of person-specific shape information, and obtaining more AU discriminative features. The proposed approach outperforms the state-of-the-art methods on two widely used AU detection datasets in the public domain (BP4D and DISFA).
Xuesong Niu, Hu Han 0001, Songfan Yang, Yan Huang 0008, Shiguang Shan
CVPR4
2019 Box-Driven Class-Wise Region Masking and Filling Rate Guided Loss for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation has achieved huge progress via adopting deep Fully Convolutional Networks (FCN). However, the performance of FCN based models severely rely on the amounts of pixel-level annotations which are expensive and time-consuming. To address this problem, it is a good choice to learn to segment with weak supervision from bounding boxes. How to make full use of the class-level and region-level supervisions from bounding boxes is the critical challenge for the weakly supervised learning task. In this paper, we first introduce a box-driven class-wise masking model (BCM) to remove irrelevant regions of each class. Moreover, based on the pixel-level segment proposal generated from the bounding box supervision, we could calculate the mean filling rates of each class to serve as an important prior cue, then we propose a filling rate guided adaptive loss (FR-Loss) to help the model ignore the wrongly labeled pixels in proposals. Unlike previous methods directly training models with the fixed individual segment proposals, our method can adjust the model learning with global statistical information. Thus it can help reduce the negative impacts from wrongly labeled proposals. We evaluate the proposed method on the challenging PASCAL VOC 2012 benchmark and compare with other methods. Extensive experimental results show that the proposed method is effective and achieves the state-of-the-art results.
Chunfeng Song, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
CVPR2
2019 Language-Driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model
abstract
Current studies on action detection in untrimmed videos are mostly designed for action classes, where an action is described at word level such as jumping, tumbling, swing, etc. This paper focuses on a rarely investigated problem of localizing an activity via a sentence query which would be more challenging and practical. Considering that current methods are generally time-consuming due to the dense frame-processing manner, we propose a recurrent neural network based reinforcement learning model which selectively observes a sequence of frames and associates the given sentence with video content in a matching-based manner. However, directly matching sentences with video content performs poorly due to the large visual-semantic discrepancy. Thus, we extend the method to a semantic matching reinforcement learning (SM-RL) model by extracting semantic concepts of videos and then fusing them with global context features. Extensive experiments on three benchmark datasets, TACoS, Charades-STA and DiDeMo, show that our method achieves the state-of-the-art performance with a high detection speed, demonstrating both effectiveness and efficiency of our method.
Weining Wang 0001, Yan Huang 0008, Liang Wang 0001
CVPR2
2019 ACMM: Aligned Cross-Modal Memory for Few-Shot Image and Sentence Matching
abstract
Image and sentence matching has drawn much attention recently, but due to the lack of sufficient pairwise data for training, most previous methods still cannot well associate those challenging pairs of images and sentences containing rarely appeared regions and words, i.e., few-shot content. In this work, we study this challenging scenario as few-shot image and sentence matching, and accordingly propose an Aligned Cross-Modal Memory (ACMM) model to memorize the rarely appeared content. Given a pair of image and sentence, the model first includes an aligned memory controller network to produce two sets of semantically-comparable interface vectors through cross-modal alignment. Then the interface vectors are used by modality-specific read and update operations to alternatively interact with shared memory items. The memory items persistently memorize cross-modal shared semantic representations, which can be addressed out to better enhance the representation of few-shot content. We apply the proposed model to both conventional and few-shot image and sentence matching tasks, and demonstrate its effectiveness by achieving the state-of-the-art performance on two benchmark datasets.
Yan Huang 0008, Liang Wang 0001
ICCV1
2019 Hierarchical Graph Convolutional Network for Skeleton-Based Action Recognition
Linjiang Huang, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
ICIG (1)2
2019 Augmented Visual-Semantic Embeddings for Image and Sentence Matching
abstract
The task of image and sentence matching has witnessed significant progress recently, but it is still challenging arising from the tremendous semantic gap between a pixel-level image and its matched sentences. Due to limited training data, it is rather challenging to optimize the visual-semantic embeddings. In this work, we propose to augment visual-semantic embeddings via enlarging the training dataset. With more data, models can learn discriminative features with high-quality semantic concepts. More specifically, we augment data by generating sentences for given images. Our method consists of two steps. At first, to enlarge the training dataset, given an image, we perform image captioning. Instead of introducing redundancy to our augmented dataset, we hope that our generated sentences are in diverse style and maintain its fidelity at the same time. Therefore, we consult to generative adversarial networks (GANs) which can produce more flexible expressions compared to methods based on the maximum likelihood principle. Then, we augment visual-semantic embeddings with the augmented training dataset and obtain the model for the task of image and sentence matching. Experiments on the popular benchmark demonstrate the effectiveness of our method by achieving superior results compared to our baseline.
Zerui Chen, Yan Huang 0008, Liang Wang 0001
ICIP2
2019 Relational Network for Skeleton-Based Action Recognition
abstract
With the fast development of effective and low-cost human skeleton capture systems, skeleton-based action recognition has attracted much attention recently. Most existing methods use Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) to extract spatio-temporal information embedded in the skeleton sequences for action recognition. However, these approaches are limited in the ability of relational modeling in a single skeleton, due to the loss of important structural information when converting the raw skeleton data to adapt to the input format of CNN or RNN. In this paper, we propose an Attentional Recurrent Relational Network-LSTM (ARRN-LSTM) to simultaneously model spatial configurations and temporal dynamics in skeletons for action recognition. We introduce the Recurrent Relational Network to learn the spatial features in a single skeleton, followed by a multi-layer LSTM to learn the temporal features in the skeleton sequences. Between the two modules, we design an adaptive attentional module to focus attention on the most discriminative parts in the single skeleton. To exploit the complementarity from different geometries in the skeleton for sufficient relational modeling, we design a two-stream architecture to learn the structural features among joints and lines simultaneously. Extensive experiments are conducted on several popular skeleton datasets and the results show that the proposed approach achieves better results than most mainstream methods.
Wu Zheng, Zhaoxiang Zhang 0001, Yan Huang 0008, Liang Wang 0001
ICME4
2019 Recurrent Deconvolutional Generative Adversarial Networks with Application to Video Generation
Hongyuan Yu, Yan Huang 0008, Lihong Pi, Liang Wang 0001
PRCV (2)2
2019 A hierarchical contextual attention-based network for sequential recommendation
Qiang Cui 0002, Yan Huang 0008, Liang Wang 0001
Neurocomputing3
2019 Learning view invariant gait features with Two-Stream GAN
Yanyun Wang 0008, Chunfeng Song, Yan Huang 0008, Zhenyu Wang 0012, Liang Wang 0001
Neurocomputing3
2019 Part-aligned pose-guided recurrent network for action recognition
Linjiang Huang, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
Pattern Recognit.2
2019 GaitNet: An end-to-end network for gait based human identification
Chunfeng Song, Yongzhen Huang, Yan Huang 0008, Liang Wang 0001
Pattern Recognit.3
2018 Learning Semantic Concepts and Order for Image and Sentence Matching
abstract
Image and sentence matching has made great progress recently, but it remains challenging due to the large visual-semantic discrepancy. This mainly arises from that the representation of pixel-level image usually lacks of high-level semantic information as in its matched sentence. In this work, we propose a semantic-enhanced image and sentence matching model, which can improve the image representation by learning semantic concepts and then organizing them in a correct semantic order. Given an image, we first use a multi-regional multi-label CNN to predict its semantic concepts, including objects, properties, actions, etc. Then, considering that different orders of semantic concepts lead to diverse semantic meanings, we use a context-gated sentence generation scheme for semantic order learning. It simultaneously uses the image global context containing concept relations as reference and the groundtruth semantic order in the matched sentence as supervision. After obtaining the improved image representation, we learn the sentence representation with a conventional LSTM, and then jointly perform image and sentence matching and sentence generation for model learning. Extensive experiments demonstrate the effectiveness of our learned semantic concepts and order, by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Qi Wu 0001, Chunfeng Song, Liang Wang 0001
CVPR1
2018 Mask-Guided Contrastive Attention Model for Person Re-Identification
abstract
Person Re-identification (ReID) is an important yet challenging task in computer vision. Due to the diverse background clutters, variations on viewpoints and body poses, it is far from solved. How to extract discriminative and robust features invariant to background clutters is the core problem. In this paper, we first introduce the binary segmentation masks to construct synthetic RGB-Mask pairs as inputs, then we design a mask-guided contrastive attention model (MGCAM) to learn features separately from the body and background regions. Moreover, we propose a novel region-level triplet loss to restrain the features learnt from different regions, i.e., pulling the features from the full image and body region close, whereas pushing the features from backgrounds away. We may be the first one to successfully introduce the binary mask into person ReID task and the first one to propose region-level contrastive learning. We evaluate the proposed method on three public datasets, including MARS, Market-1501 and CUHK03. Extensive experimental results show that the proposed method is effective and achieves the state-of-the-art results. Mask and code will be released upon request.
Chunfeng Song, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
CVPR2
2018 M3: Multimodal Memory Modelling for Video Captioning
abstract
Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, video captioning has made great progress. However, learning an effective mapping from the visual sequence space to the language space is still a challenging problem due to the long-term multimodal dependency modelling and semantic misalignment. Inspired by the facts that memory modelling poses potential advantages to long-term sequential problems [35] and working memory is the key factor of visual attention [33], we propose a Multimodal Memory Model (M3) to describe videos, which builds a visual and textual shared memory to model the long-term visual-textual dependency and further guide visual attention on described visual targets to solve visual-textual alignments. Specifically, similar to [10], the proposed M3 attaches an external memory to store and retrieve both visual and textual contents by interacting with video and sentence with multiple read and write operations. To evaluate the proposed model, we perform experiments on two public datasets: MSVD and MSR-VTT. The experimental results demonstrate that our method outperforms most of the state-of-the-art methods in terms of BLEU and METEOR.
Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
CVPR3
2018 Cross-Modal Ranking with Soft Consistency and Noisy Labels for Robust RGB-T Tracking
Chenglong Li 0002, Chengli Zhu, Yan Huang 0008, Jin Tang 0001, Liang Wang 0001
ECCV (13)3
2018 Automatic Engagement Prediction with GAP Feature
abstract
In this paper, we propose an automatic engagement prediction method for the Engagement in the Wild sub-challenge of EmotiW 2018. We first design a novel Gaze-AU-Pose (GAP) feature taking into account the information of gaze, action units and head pose of a subject. The GAP feature is then used for the subsequent engagement level prediction. To efficiently predict the engagement level for a long-time video, we divide the long-time video into multiple overlapped video clips and extract GAP feature for each clip. A deep model consisting of a Gated Recurrent Unit (GRU) layer and a fully connected layer is used as the engagement predictor. Finally, a mean pooling layer is applied to the per-clip estimation to get the final engagement level of the whole video. Experimental results on the validation set and test set show the effectiveness of the proposed approach. In particular, our approach achieves a promising result with an MSE of 0.0724 on the test set of Engagement Prediction Challenge of EmotiW 2018.t with an MSE of 0.072391 on the test set of Engagement Prediction Challenge of EmotiW 2018.
Xuesong Niu, Hu Han 0001, Jiabei Zeng, Xuran Sun, Shiguang Shan, Yan Huang 0008, Songfan Yang, Xilin Chen 0001
ICMI6
2018 Towards Unconstrained Pointing Problem of Visual Question Answering: A Retrieval-based Method
abstract
The pointing problem of visual question answering (VQA) is that given an image and a question which asks for the location of the interested object, find a region that answers the question. It is an important research problem in VQA tasks and has many potential applications in our daily life. Most of the existing work on this task can only solve it in the form of multiple choices, i.e., given candidate answers in advance, and then selecting a correct one. In this paper, we propose a retrieval model, which can not only deal with the multiple-choices task, but also provide a feasible solution for the no-candidate-answer task. The principle of our method is to pull the question and correct answer close, and push the question and incorrect answer away in a common feature space. To our best knowledge, we are the first to use retrieval method to solve the unconstrained (no-candidate-answer) pointing problem of VQA. Furthermore, our proposed method outperforms the state-of-the-art methods on the Visual7W [1] dataset in terms of the pointing problem of VQA.
Yan Huang 0008, Liang Wang 0001
ICPR2
2018 Deep Temporal Feature Encoding for Action Recognition
abstract
Human action recognition is an important task in computer vision. Recently, deep learning methods for video action recognition have developed rapidly. A popular way to tackle this problem is known as two-stream methods which take both spatial and temporal modalities into consideration. These methods often treat sparsely-sampled frames as input and video labels as supervision. Because of such sampling strategy, they are typically limited to processing shorter sequences, which might cause the problems such as suffering from the confusion by partial observation. In this paper we propose a novel video feature representation method, called Deep Temporal Feature Encoding (DTE). It could aggregate frame-level features into a robust and global video-level representation. Firstly, we sample enough RGB frames and optical flow stacks across the whole video. Then we use a deep temporal feature encoding layer to construct a strong video feature. Lastly, end-to-end training is applied so that our video representation could be global and sequence-aware. Comprehensive experiments are conducted on two public datasets: HMDB51 and UCF101. Experimental results demonstrate that DTE achieves the competitive state-of-the-art performance on both datasets.
Zhaoxiang Zhang 0001, Yan Huang 0008, Liang Wang 0001
ICPR3
2018 Hierarchical Memory Modelling for Video Captioning
abstract
Translating videos into natural language sentences has drawn much attention recently. The framework of combining visual attention with Long Short-Term Memory (LSTM) based text decoder has achieved much progress. However, the vision-language translation still remains unsolved due to the semantic gap and misalignment between video content and described semantic concept. In this paper, we propose a Hierarchical Memory Model (HMM) - a novel deep video captioning architecture which unifies a textual memory, a visual memory and an attribute memory in a hierarchical way. These memories can guide attention for efficient video representation extraction and semantic attribute selection in addition to modelling the long-term dependency for video sequence and sentences, respectively. Compared with traditional vision-based text decoder, the proposed attribute-based text decoder can largely reduce the semantic discrepancy between video and sentence. To prove the effectiveness of the proposed model, we perform extensive experiments on two public benchmark datasets: MSVD and MSR-VTT. Experiments show that our model not only can discover appropriate video representation and semantic attributes but also can achieve comparable or superior performances than state-of-the-art methods on these datasets.
Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
ACM Multimedia3
2018 Video Super-Resolution via Bidirectional Recurrent Convolutional Networks
abstract
Super resolving a low-resolution video, namely video super-resolution (SR), is usually handled by either single-image SR or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video SR. Multi-Frame SR generally extracts motion information, e.g., optical flow, to model the temporal dependency, but often shows high computational cost. Considering that recurrent neural networks (RNNs) can model long-term temporal dependency of video sequences well, we propose a fully convolutional RNN named bidirectional recurrent convolutional network for efficient multi-frame SR. Different from vanilla RNNs, 1) the commonly-used full feedforward and recurrent connections are replaced with weight-sharing convolutional connections. So they can greatly reduce the large number of network parameters and well model the temporal dependency in a finer level, i.e., patch-based rather than frame-based, and 2) connections from input layers at previous timesteps to the current hidden layer are added by 3D feedforward convolutions, which aim to capture discriminate spatio-temporal patterns for short-term fast-varying motions in local adjacent frames. Due to the cheap convolutional operations, our model has a low computational complexity and runs orders of magnitude faster than other multi-frame SR methods. With the powerful temporal dependency modeling, our model can super resolve videos with complex motions and achieve well performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Instance-Aware Image and Sentence Matching with Selective Multimodal LSTM
abstract
Effective image and sentence matching depends on how to well measure their global visual-semantic similarity. Based on the observation that such a global similarity arises from a complex aggregation of multiple local similarities between pairwise instances of image (objects) and sentence (words), we propose a selective multimodal Long Short-Term Memory network (sm-LSTM) for instance-aware image and sentence matching. The sm-LSTM includes a multimodal context-modulated attention scheme at each timestep that can selectively attend to a pair of instances of image and sentence, by predicting pairwise instance-aware saliency maps for image and sentence. For selected pairwise instances, their representations are obtained based on the predicted saliency maps, and then compared to measure their local similarity. By similarly measuring multiple local similarities within a few timesteps, the sm-LSTM sequentially aggregates them with hidden states to obtain a final matching score as the desired global similarity. Extensive experiments show that our model can well match image and sentence with complex content, and achieve the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
CVPR1
2017 See the Forest for the Trees: Joint Spatial and Temporal Recurrent Neural Networks for Video-Based Person Re-identification
abstract
Surveillance cameras have been widely used in different scenes. Accordingly, a demanding need is to recognize a person under different cameras, which is called person re-identification. This topic has gained increasing interests in computer vision recently. However, less attention has been paid to video-based approaches, compared with image-based ones. Two steps are usually involved in previous approaches, namely feature learning and metric learning. But most of the existing approaches only focus on either feature learning or metric learning. Meanwhile, many of them do not take full use of the temporal and spatial information. In this paper, we concentrate on video-based person re-identification and build an end-to-end deep neural network architecture to jointly learn features and metrics. The proposed method can automatically pick out the most discriminative frames in a given video by a temporal attention model. Moreover, it integrates the surrounding information at each location by a spatial recurrent model when measuring the similarity with another pedestrian video. That is, our method handles spatial and temporal information simultaneously in a unified manner. The carefully designed experiments on three public datasets show the effectiveness of each component of the proposed deep network, performing better in comparison with the state-of-the-art methods.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
CVPR2
2017 Conditional High-Order Boltzmann Machines for Supervised Relation Learning
abstract
Relation learning is a fundamental problem in many vision tasks. Recently, high-order Boltzmann machine and its variants have shown their great potentials in learning various types of data relation in a range of tasks. But most of these models are learned in an unsupervised way, i.e., without using relation class labels, which are not very discriminative for some challenging tasks, e.g., face verification. In this paper, with the goal to perform supervised relation learning, we introduce relation class labels into conventional high-order multiplicative interactions with pairwise input samples, and propose a conditional high-order Boltzmann Machine (CHBM), which can learn to classify the data relation in a binary classification way. To be able to deal with more complex data relation, we develop two improved variants of CHBM: 1) latent CHBM, which jointly performs relation feature learning and classification, by using a set of latent variables to block the pathway from pairwise input samples to output relation labels and 2) gated CHBM, which untangles factors of variation in data relation, by exploiting a set of latent variables to multiplicatively gate the classification of CHBM. To reduce the large number of model parameters generated by the multiplicative interactions, we approximately factorize high-order parameter tensors into multiple matrices. Then, we develop efficient supervised learning algorithms, by first pretraining the models using joint likelihood to provide good parameter initialization, and then finetuning them using conditional likelihood to enhance the discriminant ability. We apply the proposed models to a series of tasks including invariant recognition, face verification, and action similarity labeling. Experimental results demonstrate that by exploiting supervised relation labels, our models can greatly improve the performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
IEEE Trans. Image Process.1
2015 Conditional High-Order Boltzmann Machine: A Supervised Learning Model for Relation Learning
abstract
Relation learning is a fundamental operation in many computer vision tasks. Recently, high-order Boltzmann machine and its variants have exhibited the great power of modelling various data relation. However, most of them are unsupervised learning models which are not very discriminative and thus cannot server as a standalone solution to relation learning tasks. In this paper, we explore supervised learning algorithms and propose a new model named Conditional High-order Boltzmann Machine (CHBM), which can be directly used as a bilinear classifier to assign similarity scores for pairwise images. Then, to better deal with complex data relation, we propose a gated version of CHBM which untangles factors of variation by exploiting a set of latent variables to gate classification. We perform four-order tensor factorization for parameter reduction, and present two efficient supervised learning algorithms from the perspectives of being generative and discriminative, respectively. The experimental results of image transformation visualization, binary-way classification and face verification demonstrate that, by performing supervised learning, our models can greatly improve the performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
ICCV1
2015 Bidirectional Recurrent Convolutional Networks for Multi-Frame Super-Resolution
abstract
Super resolving a low-resolution video is usually handled by either single-image super-resolution (SR) or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video super-resolution. Multi-Frame SR generally extracts motion information, e.g. optical flow, to model the temporal dependency, which often shows high computational cost. Considering that recurrent neural network (RNN) can model long-term contextual information of temporal sequences well, we propose a bidirectional recurrent convolutional network for efficient multi-frame SR.Different from vanilla RNN, 1) the commonly-used recurrent full connections are replaced with weight-sharing convolutional connections and 2) conditional convolutional connections from previous input layers to current hidden layer are added for enhancing visual-temporal dependency modelling. With the powerful temporal dependency modelling, our model can super resolve videos with complex motions and achieve state-of-the-art performance. Due to the cheap convolution operations, our model has a low computational complexity and runs orders of magnitude faster than other multi-frame methods.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
NIPS1
2015 Unconstrained Multimodal Multi-Label Learning
abstract
Multimodal learning has been mostly studied by assuming that multiple label assignments are independent of each other and all the modalities are available. In this paper, we consider a more general problem where the labels contain dependency relationships and some modalities are likely to be missing. To this end, we propose a multi-label conditional restricted Boltzmann machine (ML-CRBM), which handles modality completion , fusion, and multi-label prediction in a unified framework. The proposed model is able to generate missing modalities based on observed ones, by explicitly modelling and sampling their conditional distributions. After that, it can discriminatively fuse multiple modalities to obtain shared representations under the supervision of class labels. To consider the co-occurrence of the labels, the proposed model formulates the multi-label prediction as a max-margin-based multi-task learning problem. Model parameters can be jointly learned by seeking a balance between being generative for modality generation and being discriminative for label prediction. We perform a series of experiments in terms of classification, visualization, and retrieval, and the experimental results clearly demonstrate the effectiveness of our method.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Multim.1
2014 Deep Embedding Network for Clustering
abstract
Clustering is a fundamental technique widely used for exploring the inherent data structure in pattern recognition and machine learning. Most of the existing methods focus on modeling the similarity/dissimilarity relationship among instances, such as k-means and spectral clustering, and ignore to extract more effective representation for clustering. In this paper, we propose a deep embedding network for representation learning, which is more beneficial for clustering by considering two constraints on learned representations. We first utilize a deep auto encoder to learn the reduced representations from the raw data. To make the learned representations suitable for clustering, we first impose a locality-persevering constraint on the learned representations, which aims to embed original data into its underlying manifold space. Then, different from spectral clustering which extracts representations from the block diagonal similarity matrix, we apply a group sparsity constraint for the learned representations, and aim to learn block diagonal representations in which the nonzero groups correspond to its cluster. After obtaining the learned representations, we use k-means to cluster them. To evaluate the proposed deep embedding network, we compare its performance with k-means and spectral clustering on three commonly-used datasets. The experiments demonstrate that the proposed method achieves promising performance.
Peihao Huang, Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
ICPR2
2014 A General Nonlinear Embedding Framework Based on Deep Neural Network
abstract
Recently there has been increasing interest in deep neural network due to its powerful represent ability in several successful applications such as speech recognition and image classification. In this paper, we propose a general nonlinear embedding framework based on deep neural network which can be utilized to implement a family of dimensionality reduction algorithms. The objective function of our framework consists of two terms: 1) an embedding term which transforms the input to a low-dimensional representation with a multilayer network, and 2) a regularization term which computes the reconstruction error of the original input by unrolling the multilayer network to a deep auto encoder. We adopt a layer-by-layer pretraining procedure to obtain good initial weights for the network, and then minimize the objective function by back propagating derivatives of the two terms. To evaluate the proposed framework, we perform face recognition and digit classification experiments. The experiments demonstrate that the proposed framework achieves better results than the state-of-the-art algorithms. The success of our framework further verifies deep neural network's advantages in representation learning.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICPR1
2013 Multi-task deep neural network for multi-label learning
abstract
This paper proposes a multi-task deep neural network (MT-DNN) architecture to handle the multi-label learning problem, in which each label learning is defined as a binary classification task, i.e., a positive class for “an instance owns this label” and a negative class for “an instance does not own this label”. Multi-label learning is accordingly transformed to multiple binary-class classification tasks. Considering that a deep neural nets (DNN) architecture can learn good intermediate representations shared across tasks, we generalize one classification task of traditional DNN into multiple binary classification tasks through defining the output layer with a negative class node and a positive class node for each label. After a similar pretraining process to deep belief nets, we redefine the label assignment error of MT-DNN and perform the back-propagation algorithm to fine-tune the network. To evaluate the proposed model, we carry out image annotation experiments on two public image datasets, with 2000 images and 30,000 images respectively. The experiments demonstrate that the proposed model achieves the state-of-the-art performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICIP1
2012 An effective regional saliency model based on extended site entropy rate
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICPR1