Hantao Yao

dblp:167/3478 · DBLP profile ↗
← Back
55ranked-venue papers
13as first author
39since 2021 · last 2026
0000-0001-8125-2864ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 10 first-author · 29 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 16 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
abstract
Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the interface. We present GUI-Eyes, a reinforcement learning framework for active visual perception in GUI tasks. To acquire more informative observations, the agent learns to make strategic decisions on both whether and how to invoke visual tools, such as cropping or zooming, within a two-stage reasoning process. To support this behavior, we introduce a progressive perception strategy that decomposes the decision-making into coarse exploration and fine-grained grounding, coordinated by a two-level policy. In addition, we design a spatially continuous reward function tailored to tool usage, which integrates both location proximity and region overlap to provide dense supervision and alleviate the reward sparsity common in GUI environments. On the ScreenSpot-Pro benchmark, GUI-Eyes-3B achieves 44.8% grounding accuracy using only 3k labeled samples, significantly outperforming both supervised and RL-based baselines. These results highlight that tool-aware active perception, enabled by staged policy reasoning and fine-grained reward feedback, is critical for building robust and data-efficient GUI agents.
Jiawei Shao, Dakuan Lu, Haoyi Hu, Xiangcheng Liu, Hantao Yao, Wu Liu 0005
AAAI6
2026 GASim: A Graph-Accelerated Hybrid Framework for Social Simulation
abstract
Large-scale social simulators are essential for studying complex social patterns.Prior work explores hybrid methods to scale up simulations, combining large language models (LLM)-based agents with numerical agentbased models (ABM).However, this incurs high latency due to expensive memory retrieval and sequential ABM execution.To address this challenge, we propose GASim, a graph-accelerated hybrid multi-agent framework for large-scale social simulations.For core agents driven by LLM, GASim introduces Graph-Optimized Memory (GOM) to replace intensive LLM-based retrieval pipelines with lightweight propagation over a sparse memory graph.For the majority of ordinary agents, GASim employs Graph Message Passing (GMP), substituting sequential ABM execution with parallel updates by fine-grained feature aggregation and Graph Attention Network.We further introduce Entropy-Driven Grouping (EDG) that coordinates this hybrid partitioning, leveraging information entropy to dynamically identify emergent core agents situated in information-diverse neighborhoods.Extensive experiments show that GASim not only delivers a substantial 9.94× end-to-end speedup over the traditional hybrid framework but also consumes less than 20% of baseline tokens, significantly reducing costs while preserving strong alignment with real-world public opinion trends.Our code is available at https://github.com/Jasmine0201/GASim.
Yanhui Sun, Hantao Yao, Allen He, Yongdong Zhang 0001, Wu Liu 0005
ACL (1)3
2026 Generalizable large language model based human keypoint localization for emotion recognition
Chanho Eom, Hantao Yao, Jing Yuan 0004
Pattern Recognit.6
2025 SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. Traditional works focus on aligning visual features between domains to extract domain-invariant knowledge, and recent VLM-based DAOD methods leverage semantic information provided by the textual encoder to supplement domain-specific features for each domain. However, they overlook the role of semantic information in guiding the learning of visual features that are beneficial for adaptation. To solve the problem, we propose semantic entropy to quantify the semantic information contained in visual features, and design SEmantic ENtropy guided Domain-aware Attention (SEEN-DA) to adaptively refine visual features with the semantic information of two domains. Semantic entropy reflects the importance of features based on semantic information, which can serve as attention to select discriminative visual features and suppress semantically irrelevant redundant information. Guided by semantic entropy, we introduce domain-aware attention modules into the visual encoder in SEEN-DA. It utilizes an inter-domain attention branch to extract domain-invariant features and eliminate redundant information, and an intra-domain attention branch to supplement the domain-specific semantic information discriminative on each domain. Comprehensive experiments validate the effectiveness of SEEN-DA, demonstrating significant improvements in cross-domain object detection performance.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xin Zhang 0062, Yifan Hao 0001, Xinkai Song, Shaohui Peng, Yongwei Zhao 0001, Ling Li 0001
CVPR3
2025 Language Guided Concept Bottleneck Models for Interpretable Continual Learning
abstract
Continual learning (CL) aims to enable learning systems to acquire new knowledge constantly without forgetting previously learned information. CL faces the challenge of mitigating catastrophic forgetting while maintaining interpretability across tasks. Most existing CL methods focus primarily on preserving learned knowledge to improve model performance. However, as new information is introduced, the interpretability of the learning process becomes crucial for understanding the evolving decision-making process, yet it is rarely explored. In this paper, we introduce a novel framework that integrates language-guided Concept Bottleneck Models (CBMs) to address both challenges. Our approach leverages the Concept Bottleneck Layer, aligning semantic consistency with CLIP models to learn human-understandable concepts that can generalize across tasks. By focusing on interpretable concepts, our method not only enhances the model’s ability to retain knowledge over time but also provides transparent decision-making insights. We demonstrate the effectiveness of our approach by achieving superior performance on several datasets, outperforming state-of-the-art methods with an improvement of up to 3.06% in final average accuracy on ImageNet-subset. Additionally, we offer concept visualizations for model predictions, further advancing the understanding of interpretable continual learning. Code is available at https://github.com/FisherCats/CLG-CBM.
Lu Yu 0004, Zhe Tao, Hantao Yao, Changsheng Xu
CVPR4
2025 Leveraging Multiple Deep Experts for Online Class-incremental Learning
abstract
Online incremental learning aims to enable learning systems to continuously accumulate new knowledge from streaming data in a single-pass manner while preserving previously acquired information. This more realistic and challenging setting has gained increasing attention in recent years. The state-of-the-art methods treat each module of a model, from shallow to deep, as a separate sub-expert network and transfer all the shallow expert knowledge into the final deep expert network. Although this yields notable improvements, we argue that directly supervising shallow layers hampers their acquisition of task-invariant knowledge. Furthermore, explicitly designating the final expert as a student network to absorb knowledge from other experts lacks adaptability, considering that different experts may not excel uniformly across all tasks. To address the aforementioned limitations, we leverage the shallow layers of the model as a shared feature extractor, while the deeper layers form a set of experts capable of learning robust and diverse features. Moreover, to facilitate knowledge transfer between multiple experts, we introduce the LEEP score to assess the feature transferability of each expert on new tasks, thereby selecting the most suitable expert as the teacher network for the new task. Extensive experiments on two evaluation benchmarks verify the effectiveness of our method (e.g, up to 1.3% on Split CIFAR-100 and 2.5% on Split Tiny-ImageNet). Code is available at https://github.com/untitledunmastered1998/MDE-OIL.
Zhe Tao, Lu Yu 0004, Hantao Yao, Changsheng Xu
ICME3
2025 DATE: Dual Asymmetric Textual Embedding guided Person Re-Identification
abstract
Inspired by the development of the Visual-Language Models(VLM), the textual embedding generated from the learnable prompt or the textual-level description is explored to boost the visual embedding by synchronizing the fusion and alignment strategy for person re-identification(ReID). However, the synchronization strategy treats the learnable-based and description-based textual embedding equally, leading to the generated visual representation easily affected by the noise contained in each textual embedding, especially for the description-based textual embedding. To address the above shortcoming, we propose a novel Dual Asymmetric Textual Embedding(DATE) that uses learnable-based and description-based textual embedding to asymmetricly guide person representation learning. Since the description-based textual embeddings are controlled mainly by the automatically extracted textual-level descriptions, they contain a lot of noise and have less discriminative ability. Therefore, DATE treats the description-based textual embeddings as auxiliary clues to boost visual and textual representation learning. Moreover, the Textual-to-Visual Adapter and Textual-to-Textual Adapter are proposed to inject the description-based textual embedding into the learnable-based textual embedding and visual embedding. To reduce the effect of noise in textual description, the identity-aware description-based textual embeddings are generated by averaging the description-based textual embeddings belonging to the same identity, which is used to boost the discriminative to infer the learnable-based textual space used for aligning the visual representation learning. Extensive evaluations on Market-1501, DukeMTMC, and MSMT-17 validate the effectiveness of the proposed method.
Pengqi Yin, Hantao Yao, Changsheng Xu
ICME2
2025 Locality Preserving Markovian Transition for Instance Retrieval
abstract
Diffusion-based re-ranking methods are effective in modeling the data manifolds through similarity propagation in affinity graphs. However, positive signals tend to diminish over several steps away from the source, reducing discriminative power beyond local regions. To address this issue, we introduce the Locality Preserving Markovian Transition (LPMT) framework, which employs a long-term thermodynamic transition process with multiple states for accurate manifold distance measurement. The proposed LPMT first integrates diffusion processes across separate graphs using Bidirectional Collaborative Diffusion (BCD) to establish strong similarity relationships. Afterwards, Locality State Embedding (LSE) encodes each instance into a distribution for enhanced local consistency. These distributions are interconnected via the Thermodynamic Markovian Transition (TMT) process, enabling efficient global retrieval while maintaining local effectiveness. Experimental results across diverse tasks confirm the effectiveness of LPMT for instance retrieval.
Jifei Luo, Wenzheng Wu, Hantao Yao, Lu Yu 0004, Changsheng Xu
ICML3
2025 Bi-Modality Individual-Aware Prompt Tuning for Visual-Language Model
abstract
Prompt tuning is a valuable technique for adapting visual language models (VLMs) to different downstream tasks, such as domain generalization and learning from a few examples. Previous methods have utilized Context Optimization approaches to deduce domain-shared or cross-modality prompt tokens, which enhance generalization and discriminative ability in textual or visual contexts. However, these prompt tokens, inferred from training data, cannot adapt perfectly to the distribution of the test dataset. This work introduces a novel approach called Bi-modality Individual-aware Prompt Tuning (BIP) by explicitly incorporating the individual's essential prior knowledge into the learnable prompt to enhance their discriminability and generalization. The critical insight of BIP involves applying the Textual Knowledge Embedding (TKE) and Visual Knowledge Embedding (VKE) models to project the class-aware textual essential knowledge and the instance-aware essential knowledge into the class-aware prompt and instance-aware prompt, referred to as Textual-level Class-aware Prompt tuning (TCP) and Visual-level Instance-aware Prompt tuning (VIP). On the one hand, TCP integrates the generated class-aware prompts into the Text Encoder to produce a dynamic class-aware classifier to improve generalization on unseen domains. On the other hand, VIP uses the instance-aware prompt to generate the dynamic visual embedding of each instance, thereby enhancing the discriminative capability of visual embedding. Comprehensive evaluations demonstrate that BIP can be used as a plug-and-play module easily integrated with existing methods and achieves superior performance on 15 benchmarks across four tasks.
Hantao Yao, Rui Zhang 0040, Huaihai Lyu, Yongdong Zhang 0001, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Multiple Local Prompts Distillation for Domain Generalization
abstract
Prompt tuning has been proven effective for Domain Generalization (DG) by enhancing the generalization capability of visual-language models with fewer learnable tokens. Existing methods adopt mostly inferring global-level individual prompts for the whole dataset to capture domain-invariant knowledge across different domains. However, since domain shifts exist, a single global-level individual prompt is easily overfitted to source domain datasets, thus lacking generalizability to the whole dataset’s feature distribution. Moreover, fluctuations in the generalization performance during the training process in DG problems often pose significant challenges to model selection strategies. To address the aforementioned problems, inspired by the Mixture-of-Expert (MOE) and knowledge distillation, we propose a novel Multiple Local Prompts Distillation (MLPD) method to inject the knowledge of multiple local prompts into a unique global prompt, improving both the generalization and discriminative ability. To ensure the diversity of local prompts, we split the whole dataset into several subsets to infer the discriminative local prompts for each subset, which is further applied to generate the generability global prompt. Formally, for each subset, Meta Prompt Tuning (MPT) is proposed to constrain each local prompt to capture both the domain-specific and domain-shared generalization knowledge on the basis of the domain label and meta-learning mechanism. After that, Prompt Knowledge Distillation (PKD) is proposed to distill the knowledge captured in the local-level prompts into the global-level prompt with prompt-level and feature-level knowledge distillations. The final evaluation on multiple benchmarks underscores the effectiveness of the proposed MLPD, e.g, achieving mAPs of 97.3%, 84.8%, 85.2%, 57.3%, and 60.7% on PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet, respectively.
Huaihai Lyu, Hantao Yao, Changsheng Xu
IEEE Trans. Multim.2
2024 TCP: Textual-Based Class-Aware Prompt Tuning for Visual-Language Model
abstract
Prompt tuning represents a valuable technique for adapting pre-trained visual-language models (VLM) to various downstream tasks. Recent advancements in CoOp-based methods propose a set of learnable domain-shared or image-conditional textual tokens to facilitate the generation of task-specific textual classifiers. However, those textual tokens have a limited generalization ability regarding unseen domains, as they cannot dynamically adjust to the distribution of testing classes. To tackle this issue, we present a novel Textual-based Class-aware Prompt tuning(TCP) that explicitly incorporates prior knowledge about classes to en-hance their discriminability, The critical concept of TCP in-volves leveraging Textual Knowledge Embedding (TKE) to map the high generalizability of class-level textual knowledge into class-aware textual tokens. By seamlessly inte-grating these class-aware prompts into the Text Encoder, a dynamic class-aware classifier is generated to enhance dis-criminability for unseen domains. During inference, TKE dynamically generates class-aware prompts related to the unseen classes. Comprehensive evaluations demonstrate that TKE serves as a plug-and-play module effortlessly combinable with existing methods. Furthermore, TCP con-sistently achieves superior performance while demanding less training time11https://github.com/htyao89/Textual-based_Class-aware_prompt_tuning.
Hantao Yao, Rui Zhang 0040, Changsheng Xu
CVPR1
2024 Prompt-based Visual Alignment for Zero-shot Policy Transfer
abstract
Overfitting in RL has become one of the main obstacles to applications in reinforcement learning(RL). Existing methods do not provide explicit semantic constrain for the feature extractor, hindering the agent from learning a unified cross-domain representation and resulting in performance degradation on unseen domains. Besides, abundant data from multiple domains are needed. To address these issues, in this work, we propose prompt-based visual alignment (PVA), a robust framework to mitigate the detrimental domain bias in the image for zero-shot policy transfer. Inspired that Visual-Language Model (VLM) can serve as a bridge to connect both text space and image space, we leverage the semantic information contained in a text sequence as an explicit constraint to train a visual aligner. Thus, the visual aligner can map images from multiple domains to a unified domain and achieve good generalization performance. To better depict semantic information, prompt tuning is applied to learn a sequence of learnable tokens. With explicit constraints of semantic information, PVA can learn unified cross-domain representation under limited access to cross-domain data and achieves great zero-shot generalization ability in unseen domains. We verify PVA on a vision-based autonomous driving task with CARLA simulator. Experiments show that the agent generalizes well on unseen domains under limited access to multi-domain data.
Haihan Gao, Rui Zhang 0040, Qi Yi, Hantao Yao, Haochen Li 0002, Jiaming Guo, Shaohui Peng, Yunkai Gao 0001, QiCheng Wang, Xing Hu 0001, Yuanbo Wen 0001, Zidong Du, Ling Li 0001, Qi Guo 0001, Yunji Chen
ICML4
2024 Cluster-Aware Similarity Diffusion for Instance Retrieval
abstract
Diffusion-based re-ranking is a common method used for retrieving instances by performing similarity propagation in a nearest neighbor graph. However, existing techniques that construct the affinity graph based on pairwise instances can lead to the propagation of misinformation from outliers and other manifolds, resulting in inaccurate results. To overcome this issue, we propose a novel Cluster-Aware Similarity (CAS) diffusion for instance retrieval. The primary concept of CAS is to conduct similarity diffusion within local clusters, which can reduce the influence from other manifolds explicitly. To obtain a symmetrical and smooth similarity matrix, our Bidirectional Similarity Diffusion strategy introduces an inverse constraint term to the optimization objective of local cluster diffusion. Additionally, we have optimized a Neighbor-guided Similarity Smoothing approach to ensure similarity consistency among the local neighbors of each instance. Evaluations in instance retrieval and object re-identification validate the effectiveness of the proposed CAS, our code is publicly available.
Jifei Luo, Hantao Yao, Changsheng Xu
ICML2
2024 DA-Ada: Learning Domain-Aware Adapter for Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. As the visual-language models (VLMs) can provide essential general knowledge on unseen images, freezing the visual encoder and inserting a domain-agnostic adapter can learn domain-invariant knowledge for DAOD. However, the domain-agnostic adapter is inevitably biased to the source domain. It discards some beneficial knowledge discriminative on the unlabelled domain, \ie domain-specific knowledge of the target domain. To solve the issue, we propose a novel Domain-Aware Adapter (DA-Ada) tailored for the DAOD task. The key point is exploiting domain-specific knowledge between the essential general knowledge and domain-invariant knowledge. DA-Ada consists of the Domain-Invariant Adapter (DIA) for learning domain-invariant knowledge and the Domain-Specific Adapter (DSA) for injecting the domain-specific knowledge from the information discarded by the visual encoder. Comprehensive experiments over multiple DAOD tasks show that DA-Ada can efficiently infer a domain-aware visual encoder for boosting domain adaptive object detection. Our code is available at https://github.com/Therock90421/DA-Ada.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xin Zhang 0062, Yifan Hao 0001, Xinkai Song, Xiaqing Li, Yongwei Zhao 0001, Yunji Chen, Ling Li 0001
NeurIPS3
2024 Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition
abstract
Audio-visual video recognition (AVVR) integrates audio and visual cues to accurately categorize videos. While current methods using provided datasets achieve satisfactory results, they face challenges in retaining historical class knowledge when new classes appear in real-world situations. There are no dedicated methods to address this issue, prompting this paper to explore Class Incremental Audio-Visual Video Recognition (CIAVVR). CIAVVR aims to preserve historical knowledge contained in stored data and learned models to prevent catastrophic forgetting. Audio-visual data and models inherently have hierarchical structures, where the model contains both low-level and high-level semantic information, and data includes snippet-level, video-level, and distribution-level spatial information. It is crucial to fully exploit these hierarchical structures for data knowledge preservation and model knowledge preservation. However, existing image class incremental learning methods do not explicitly consider these hierarchical structures. Therefore, we introduce Hierarchical Augmentation and Distillation (HAD), which includes the Hierarchical Augmentation Module (HAM) and Hierarchical Distillation Module (HDM). These modules efficiently utilize the hierarchical structure of data and models. Specifically, HAM uses a novel augmentation strategy, segmental feature augmentation, to preserve hierarchical model knowledge. Simultaneously, HDM employs newly designed hierarchical logical distillation (video-distribution) and hierarchical correlative distillation (snippet-video) to maintain intra-sample and inter-sample hierarchical knowledge. Evaluations on four benchmarks (AVE, AVK-100, AVK-200, and AVK-400) show that HAD effectively captures hierarchical information, enhancing the preservation of historical class knowledge and performance. We also provide a theoretical analysis to support the segmental feature augmentation strategy.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Class Incremental Learning for Light-Weighted Networks
abstract
Despite deep neural networks (DNNs) show impressive performance across diverse tasks, they suffer from catastrophic forgetting when dealing with continuous data streams. Incremental learning aims to alleviate this phenomenon and enable DNNs to accumulate new knowledge to cope with the ever-changing world. Recently numerous advanced methods have been developed to enhance the incremental learning capabilities of neural networks. However, these methods mainly focus on the large networks, neglecting the unique needs of edged-device applications, which is surprisingly under-investigated in previous literature. In this paper, we propose two strategies for transferring knowledge from large teacher networks to light-weighted networks in class incremental learning. Specifically, in cases where the initial task contains a large number of categories, our static teacher strategy involves transferring knowledge from the teacher to the student network on the initial task to enhance the plasticity of the student network, and applying regularization constraints on the subsequent task to improve its stability. In a more challenging scenario where each task includes an equal number of categories, the dynamic teacher strategy continuously guides the student network on each task. We evaluate the proposed methods on CIFAR100, Tiny-ImageNet and ImageNet-subset datasets with different types of light-weighted networks (MobileNet, ShuffleNet). We observed that effective knowledge transfer resulting in the student network achieving performance comparable or even outperform the teacher network. Extensive and detailed experiments conducted on three datasets demonstrated the simplicity and effectiveness of our proposed method. Comprehensive analysis are also conducted including different factors and visualization.
Zhe Tao, Lu Yu 0004, Hantao Yao, Shucheng Huang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Source-Guided Target Feature Reconstruction for Cross-Domain Classification and Detection
abstract
Existing cross-domain classification and detection methods usually apply a consistency constraint between the target sample and its self-augmentation for unsupervised learning without considering the essential source knowledge. In this paper, we propose a Source-guided Target Feature Reconstruction (STFR) module for cross-domain visual tasks, which applies source visual words to reconstruct the target features. Since the reconstructed target features contain the source knowledge, they can be treated as a bridge to connect the source and target domains. Therefore, using them for consistency learning can enhance the target representation and reduce the domain bias. Technically, source visual words are selected and updated according to the source feature distribution, and applied to reconstruct the given target feature via a weighted combination strategy. After that, consistency constraints are built between the reconstructed and original target features for domain alignment. Furthermore, STFR is connected with the optimal transportation algorithm theoretically, which explains the rationality of the proposed module. Extensive experiments onnine benchmarksandtwo cross-domain visual tasksprove the effectiveness of the proposed STFR module,e.g., 1)cross-domain image classification: obtaining average accuracy of 91.0%, 73.9%, and 87.4% onOffice-31,Office-Home, andVisDA-2017, respectively; 2)cross-domain object detection: obtaining mAP of 44.50% onCityscapes→Foggy Cityscapes, AP on car of 78.10% onCityscapes→KITTI, MR-2of 8.63%, 12.27%, 22.10%, and 40.58% onCOCOPersons→Caltech,CityPersons→Caltech,COCOPersons→CityPersons, andCaltech→CityPersons, respectively.
Yifan Jiao, Hantao Yao, Bing-Kun Bao, Changsheng Xu
IEEE Trans. Image Process.2
2024 REACT: Remainder Adaptive Compensation for Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) aims to infer a robust detector on the target domain with the labelled source datasets. Recent studies utilize a feature extractor shared on the source and target domains to capture the domain-invariant features and the task-relevant information with both feature-alignment constraint and source annotations. However, the feature extractor shared across domains discards partial task-relevant information of the target domain due to the domain gap and lack of target annotations, leading to compromised discrimination capabilities within target domain. To this end, we propose a novel REmainder Adaptive CompensaTion network (REACT) to adaptively compensate the extracted features with the remainder features for generating task-relevant features. The key insight is that the remainder features contain the discarded task-relevant information, so they can be adapted to compensate for the inadequate target features. Especially, REACT introduces an additional remainder branch to regain the remainder features, and then adaptively utilizes them to compensate for the discarded task-relevant information, improving discrimination on the target domain. Extensive experiments over multiple cross-domain adaptation tasks with three baselines demonstrate that our approach gains significant improvements and achieves superior performance compared with highly-optimized state-of-the-art methods.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xin Zhang 0062, Yifan Hao 0001, Xinkai Song, Ling Li 0001
IEEE Trans. Image Process.3
2024 Camera-Incremental Object Re-Identification With Identity Knowledge Evolution
abstract
Object Re-identification (ReID) is a task focused on retrieving a probe object from a multitude of gallery images using a ReID model trained on a stationary, camera-free dataset. This training involves associating and aggregating identities across various camera views. However, when deploying ReID algorithms in real-world scenarios, several challenges, such as storage constraints, privacy considerations, and dynamic changes in camera setups, can hinder their generalizability and practicality. To address these challenges, we introduce a novel ReID task called Camera-Incremental Object Re-identification (CIOR). In CIOR, we treat each camera's data as a separate source and continually optimize the ReID model as new data streams come from various cameras. By associating and consolidating the knowledge of common identities, our aim is to enhance discrimination capabilities and mitigate the problem of catastrophic forgetting. Therefore, we propose a novel Identity Knowledge Evolution (IKE) framework for CIOR, consisting of Identity Knowledge Association (IKA), Identity Knowledge Distillation (IKD), and Identity Knowledge Update (IKU). IKA is proposed to discover common identities between the current identity and historical identities, facilitating the integration of previously acquired knowledge. IKD involves distilling historical identity knowledge from common identities, enabling rapid adaptation of the historical model to the current camera view. After each camera has been trained, IKU is applied to continually expand identity knowledge by combining historical and current identity memories. Market-CL and Veri-CL evaluations show the effectiveness of Identity Knowledge Evolution (IKE) for CIOR.Code:https://github.com/htyao89/Camera-Incremental-Object-ReID
Hantao Yao, Jifei Luo, Lu Yu 0004, Changsheng Xu
IEEE Trans. Multim.1
2024 Multi-object Tracking with Spatial-Temporal Tracklet Association
abstract
Recently, the tracking-by-detection methods have achieved excellent performance in Multi-Object Tracking (MOT), which focuses on obtaining a robust feature for each object and generating tracklets based on feature similarity. However, they are confronted with two issues: (1) unstable features in short-term occlusion and (2) insufficient matching in long-term occlusion. Specifically, the unstable feature is caused by the appearance variation under occlusion, and the association with the current unstable feature will lead to insufficient matching in long-term occlusion. To address the above issues, we propose a two-stage tracklet-level association method, Spatial-Temporal Tracklet Association (STTA), to effectively combine spatial-temporal context between feature extraction and data association. In the first stage, we propose the Tracklet-guided Spatial-Temporal Attention network (TSTA) to generate robust and stable features. Specifically, TSTA captures spatial-temporal context to obtain the most salient regions between the current and previous clips. In the second stage, we design the Bi-Tracklet Spatial-Temporal association (BTST) module to fully exploit the spatial-temporal context in data association. Specifically, we leverage BTST to merge different tracklets into long-term trajectories by jointly learning visual feature and spatial-temporal context and designing a bidirectional interpolation to recover the missed objects between matched tracklets. Extensive experiments of public and private detections on four benchmarks demonstrate the robustness of STTA. Furthermore, the proposed method is a model-agnostic method, which can be plugged and played with existing methods to boost their performance, e.g., obtain 11.0%, 10.1%, 2.9%, 3.2%, and 7.8% improvement on IDF1 in the MOT16 validation dataset for Tracktor, CenterTrack, Deepsort, JDE, and CTracker, respectively.
Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Incremental Audio-Visual Fusion for Person Recognition in Earthquake Scene
abstract
Earthquakes have a profound impact on social harmony and property, resulting in damage to buildings and infrastructure. Effective earthquake rescue efforts require rapid and accurate determination of whether any survivors are trapped in the rubble of collapsed buildings. While deep learning algorithms can enhance the speed of rescue operations using single-modal data (either visual or audio), they are confronted with two primary challenges: insufficient information provided by single-modal data and catastrophic forgetting. In particular, the complexity of earthquake scenes means that single-modal features may not provide adequate information. Additionally, catastrophic forgetting occurs when the model loses the information learned in a previous task after training on subsequent tasks, due to non-stationary data distributions in changing earthquake scenes. To address these challenges, we propose an innovative approach that utilizes an incremental audio-visual fusion model for person recognition in earthquake rescue scenarios. Firstly, we leverage a cross-modal hybrid attention network to capture discriminative temporal context embedding, which uses self-attention and cross-modal attention mechanisms to combine multi-modality information, enhancing the accuracy and reliability of person recognition. Secondly, an incremental learning model is proposed to overcome catastrophic forgetting, which includes elastic weight consolidation and feature replay modules. Specifically, the elastic weight consolidation module slows down learning on certain weights based on their importance to previously learned tasks. The feature replay module reviews the learned knowledge by reusing the features conserved from the previous task, thus preventing catastrophic forgetting in dynamic environments. To validate the proposed algorithm, we collected the Audio-Visual Earthquake Person Recognition (AVEPR) dataset from earthquake films and real scenes. Furthermore, the proposed method gets 85.41% accuracy while learning the 10th new task, which demonstrates the effectiveness of the proposed method and highlights its potential to significantly improve earthquake rescue efforts.
Sisi You, Yukun Zuo, Hantao Yao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Visual-Language Prompt Tuning with Knowledge-Guided Context Optimization
abstract
Prompt tuning is an effective way to adapt the pretrained visual-language model (VLM) to the downstream task using task-related textual tokens. Representative CoOp-based work combines the learnable textual tokens with the class tokens to obtain specific textual knowledge. However, the specific textual knowledge is worse generalization to the unseen classes because it forgets the essential general textual knowledge having a strong generalization ability. To tackle this issue, we introduce a novel Knowledge-guided Context Optimization (KgCoOp) to enhance the generalization ability of the learnable prompt for unseen classes. The key insight of KgCoOp is that the forgetting about essential knowledge can be alleviated by reducing the discrepancy between the learnable prompt and the hand-crafted prompt. Especially, KgCoOp minimizes the discrepancy between the textual embeddings generated by learned prompts and the hand-crafted prompts. Finally, adding the KgCoOp upon the contrastive loss can make a discriminative prompt for both seen and unseen tasks. Extensive evaluation of several benchmarks demonstrates that the proposed Knowledge-guided Context Optimization is an efficient method for prompt tuning, i.e., achieves better performance with less training time. code.
Hantao Yao, Rui Zhang 0040, Changsheng Xu
CVPR1
2023 UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement
abstract
Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. However, the independent identity association results in the identity-aware knowledge contained in the tracklet not be used to boost the detection and embedding modules. To overcome the limitations of existing methods, we introduce a novel Unified Tracking Model (UTM) to bridge those three components for generating a positive feedback loop with mutual benefits. The key insight of UTM is the Identity-Aware Feature Enhancement (IAFE), which is applied to bridge and benefit these three components by utilizing the identity-aware knowledge to boost detection and embedding. Formally, IAFE contains the Identity-Aware Boosting Attention (IABA) and the Identity-Aware Erasing Attention (IAEA), where IABA enhances the consistent regions between the current frame feature and identity-aware knowledge, and IAEA suppresses the distracted regions in the current frame feature. With better detections and embeddings, higher-quality tracklets can also be generated. Extensive experiments of public and private detections on three benchmarks demonstrate the robustness of UTM.
Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu
CVPR2
2023 Learning Domain-Aware Detection Head with Prompt Tuning
abstract
Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. However, existing methods focus on reducing the domain bias of the detection backbone by inferring a discriminative visual encoder, while ignoring the domain bias in the detection head. Inspired by the high generalization of vision-language models (VLMs), applying a VLM as the robust detection backbone following a domain-aware detection head is a reasonable way to learn the discriminative detector for each domain, rather than reducing the domain bias in traditional methods. To achieve the above issue, we thus propose a novel DAOD framework named Domain-Aware detection head with Prompt tuning (DA-Pro), which applies the learnable domain-adaptive prompt to generate the dynamic detection head for each domain. Formally, the domain-adaptive prompt consists of the domain-invariant tokens, domain-specific tokens, and the domain-related textual description along with the class label. Furthermore, two constraints between the source and target domains are applied to ensure that the domain-adaptive prompt can capture the domains-shared and domain-specific knowledge. A prompt ensemble strategy is also proposed to reduce the effect of prompt disturbance. Comprehensive experiments over multiple cross-domain adaptation tasks demonstrate that using the domain-adaptive prompt can produce an effectively domain-related detection head for boosting domain-adaptive object detection. Our code is available at https://github.com/Therock90421/DA-Pro.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xinkai Song, Yifan Hao 0001, Yongwei Zhao 0001, Ling Li 0001, Yunji Chen
NeurIPS3
2023 Dual Instance-Consistent Network for Cross-Domain Object Detection
abstract
Cross-domain object detection aims to transfer knowledge from a labeled dataset to an unlabeled dataset. Most existing methods apply a unified embedding model to generate the tightly coupled source and target descriptions for domain alignment, leading to the destroyed feature distribution of the target domain because the embedding model is mainly controlled by the source domain. To reduce the representation bias of the target domain, we apply two independent networks to extract two types of discriminative descriptions with mutual consistency, i.e., a novel Dual Instance-Consistent Network (DICN) is proposed for cross-domain object detection. Especially, Dual Instance-Consistent Module containing the instance mutual consistency between Primary Network and Auxiliary Network is applied to align two domains, where Primary and Auxiliary Networks are used to obtain the source-specific and target-specific information, respectively. The instance mutual consistency consists of two terms: feature consistency and detection consistency, which is applied to align the instance feature and the output of detection head, respectively. With the instance mutual consistency, optimizing the Primary (Auxiliary) Network only with source (target) images by fixing the Auxiliary (Primary) Network can generate the source(target)-specific description. Extensive experiments on several benchmarks demonstrate the effectiveness of the proposed DICN, e.g., obtaining mAP of 44.10% for Cityscapes$\rightarrow$Foggy Cityscapes, AP on car of 76.50% for Cityscapes$\rightarrow$KITTI, MR$^{-2}$of 8.87%, 12.66%, 22.27%, and 42.06% for COCOPersons$\rightarrow$Caltech, CityPersons$\rightarrow$Caltech, COCOPersons$\rightarrow$CityPersons, and Caltech$\rightarrow$CityPersons, respectively.
Yifan Jiao, Hantao Yao, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Dual Structural Knowledge Interaction for Domain Adaptation
abstract
Domain adaptation aims to transfer knowledge from a label-rich source domain to an unlabeled target domain. A common strategy is to assign pseudo-labels to unlabeled target samples for performing representation learning. However, most existing methods only apply the source-guided classifier to generate the source-biased pseudo-labels for self-training, leading to biased target representations. Moreover, the generated pseudo-labels ignore the manifold assumption that neighboring samples are likely to have the same labels. To address the above problem, we formulate a novel structural knowledge to assign target-oriented and manifold-guided pseudo-labels for unlabeled target samples. The structural knowledge consists of cluster-based knowledge and locality-based knowledge. The cluster-based knowledge denotes the label consistency between the target samples and the non-parametric target cluster centers, making the pseudo-labels target-oriented. The locality-based knowledge constrains the target sample and its neighbors to satisfy the manifold assumption. As the neighbors contain the source and target samples, the source and target locality-based knowledge are utilized to boost the descriptions. With the structural knowledge, we propose a novel Dual Structural Knowledge Interaction (DSKI) framework for domain adaptation. For generating aligned and discriminative features, knowledge consistency constraint and instance mutual constraint are proposed in DSKI. Evaluations on three benchmarks demonstrate the effectiveness of the Dual Structural Knowledge Interaction,e.g.,74.9%, 87.7%, and 90.8% for Office-Home, VisDa-2017, and Office-31, respectively.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Multim.2
2022 Multi-Object Tracking With Spatial-Temporal Topology-Based Detector
abstract
Multi-object tracking is a challenging task due to the occlusion of different targets. Existing methods focus on inferring a robust and discriminative feature for data association based on the targets generated by the existing detector. Unlike existing methods that consider each target independently during generating the trajectories, we propose a novel Spatial-Temporal Topology-based Detector (STTD) algorithm that treats the target and its nearest neighbors as a cluster and introduces a topology structure to describe the dynamics of moving targets belonging to the same cluster. With the public detections and the tracked objects in the previous frame, STTD firstly refines them by regression of detector to obtain the candidate proposals in the current frame. After that, the temporal topology constraint is proposed to recover the missed objects by considering the continuity and consistency of the topological structure. Based on the assumption that the targets belonging to the same topology should have a consistent characteristic, the spatial topology constraint is proposed to remove the inaccurate targets. Then we can obtain new candidate objects and construct the cost matrix used for data association. The evaluations on three MOTChallenge benchmarks verify the effectiveness of the proposed method.
Sisi You, Hantao Yao, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Margin-Based Adversarial Joint Alignment Domain Adaptation
abstract
Domain adaptation aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain, which has different data distribution with the source domain. Most of the existing methods focus on aligning the data distribution between the source and target domains but ignore the discrimination of the feature space among categories, leading the samples close to the decision boundary to be misclassified easily. To address the above issue, we propose a Margin-based Adversarial Joint Alignment (MAJA) to constrain the feature spaces of source and target domains to be aligned and discriminative. The proposed MAJA consists of two components: joint alignment module and margin-based generative module. The joint alignment module is proposed to align the source and target feature spaces by considering the joint distribution of features and labels. Therefore, the embedding features and the corresponding labels treated as pair data are applied for domain alignment. Furthermore, the margin-based generative module is proposed to boost the discrimination of the feature space,i.e.,make all samples as far away from the decision boundary as possible. The margin-based generative module first employs the Generative Adversarial Networks (GAN) to generate a lot of fake images for each category, then applies the adversarial learning to enlarge and reduce the category margin for the true images and generated fake images, respectively. The evaluations on three benchmarks,e.g.,small image datasets, VisDA-2017, and Office-31, verify the effectiveness of the proposed method.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Intra-Domain Consistency Enhancement for Unsupervised Person Re-Identification
abstract
Recently, unsupervised domain adaptation in person re-identification (ReID) has been widely studied to improve the generalization ability of the ReID model. Some existing methods focus on handling the intra-domain image variations caused by different camera configurations, pose, illumination, and background in target domain. However, they fail to fully mine the underlying consistency constraints contained in unlabeled target dataset. To comprehensively investigate the underlying constraints for unsupervised representation learning, we introduce two consistency constraints to deal with the intra-domain variations, namely instance-ensembling consistency and cross-granularity consistency. Specifically, the instance-ensembling consistency constraint aims to encourage similar features for a given instance and its positive samples. The cross-granularity consistency constraint is designed to enhance the collaboration of global clues and local clues in multi-granularity feature learning, which can overcome the negative effects caused by the noisy pseudo labels. By combining the advantages of the two constraints, we propose an iterative Intra-domain Consistency Enhancement (ICE) approach based on the Mean Teacher framework to fully mine the two underlying consistency constraints on multi-granularity features. The proposed ICE approach achieves significant improvement compared with the state-of-the-art, which demonstrates the superiority of the two consistency constraints.
Hantao Yao, Changsheng Xu
IEEE Trans. Multim.2
2022 Attribute-Induced Bias Eliminating for Transductive Zero-Shot Learning
abstract
Transductive zero-shot learning is designed to recognize unseen categories by aligning both visual and semantic information in a joint embedding space. Four types of domain biases exist in Transductive ZSL,i.e.,visual biasandsemantic biasin two domains, and twovisual-semantic biasesexist in the seen and unseen domains. However, the existing work has only focused on specific components of these topics, leading to severe semantic ambiguity during knowledge transfer. To solve this problem, we propose a novel attribute-induced bias eliminating (AIBE) module for Transductive ZSL. Specifically, for thevisual biasbetween the two domains, the mean-teacher module is first used to bridge the visual representation discrepancy between the two domains using unsupervised learning and unlabeled images. Then, an attentional graph attribute embedding process is proposed to reduce thesemantic biasbetween seen and unseen categories using a graph operation to describe the semantic relationship between categories. To reduce semantic-visual bias in the seen domain, we align the visual center of each category with the corresponding semantic attributes instead of with the individual visual data point, which preserves the semantic relationship in the embedding space. Finally, for the semantic-visual bias in the unseen domain, an unseen semantic alignment constraint is designed to align visual and semantic space using an unsupervised process. The evaluations on several benchmarks demonstrate the effectiveness of the proposed method,e.g.,82.8%/75.5%, 97.1%/82.5%, and 73.2%/52.1% for Conventional/Generalized ZSL settings for CUB, AwA2, and SUN datasets, respectively.
Hantao Yao, Shaobo Min, Yongdong Zhang 0001, Changsheng Xu
IEEE Trans. Multim.1
2022 Seek Common Ground While Reserving Differences: A Model-Agnostic Module for Noisy Domain Adaptation
abstract
Noisy domain adaptation aims to solve the problem that the source dataset contains noisy labels in domain adaptation. Previous methods handle noisy labels by selecting the small-loss samples with inconsistent predictions between two models and discarding the consistent samples, resulting in many noises contained in the selected samples. By jointly considering the consistent and inconsistent samples, we propose a model-agnostic module, named Seek Common Ground While Reserving Differences (SCGWRD), to reduce the impact of noisy samples. The proposed SCGWRD module consists of Seek Common Ground (SCG) component and Reserve Differences (RD) component by utilizing the outputs of two symmetrical domain adaptation models. As the common samples with consistent predictions between two models are more likely to be clean samples, the SCG component applies the small-loss strategy to select the reliable samples with consistent predictions. Unlike SCG, the RD component maintains the divergences between two models with mutual learning and reduces the effect of noisy data using the samples with different predictions and small losses. Evaluations on three benchmarks demonstrate the effectiveness and robustness of the proposed SCGWRD module for noisy domain adaptation.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Multim.2
2021 PEN: Pose-Embedding Network for Pedestrian Detection
abstract
In the past years, pedestrian detection has achieved significant progress via improving the visual description. However, the visual description is not robust to discover the occluded pedestrian, which is the bottleneck of the existing pedestrian methods. Targeting to overcome the shortcoming of visual description, we employ the human pose information, which is complementary to the visual description, to address the occlusion and false positive failure problems in pedestrian detection. The advantage of using human pose information is that the pose estimation model can localize the local part of the pedestrian once the pedestrian is occluded. By embedding the human pose information with the visual description, we propose a novel Pose-Embedding Network for pedestrian detection, which consists of two components: a Region Proposal Network, and a Pedestrian Recognization Network. The Region Proposal Network targets to generate lots of candidate proposals and corresponding confidence scores. Once obtaining the candidate proposals, the Pedestrian Recognization Network is proposed to distinguish pedestrian proposals by taking the visual information and pose information into consideration to refine the confidence scores and eliminate the false positives. Given the proposal image, the visual information is extracted with the Visual Feature Module. The Human Pose Module, which is proposed based on the pose estimation model, is used to predict the pose information. Further, the Classification Module is employed to fuse the visual and pose information and generates a pose-embedding pedestrian description. Extensive experiments on three challenging datasets, i.e., Caltech, CityPersons, and COCOPersons, show that the proposed approach achieves a significant improvement upon the state-of-the-art methods.
Yifan Jiao, Hantao Yao, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2021 Multi-Target Multi-Camera Tracking With Optical-Based Pose Association
abstract
Multi-target multi-camera tracking (MTMCT) targets to generate trajectories of the object that appeared under multiple cameras automatically. MTMCT can be treated as a combination of intra-camera tracking and cross-camera tracking. The existing work only employs the global description to perform the tracklet generating. However, the global description cannot model the local similarity between targets, leading to existing methods not to be robust to occlusion and fast motion. To handle the mentioned problem, we propose an online Optical-based Pose Association (OPA) for multi-target multi-camera tracking. The proposed method utilizes local pose matching to solve the occlusion problem, and applies optical flow to reduce the distance caused by fast motion. For optical-based pose association, we firstly employ OpenPose to generate human pose for each proposal. Then, we utilize the optical flow generated by PWC-Net to adjust the estimated pose for the previous frame. Finally, the modified Object Keypoint Similarity is used to compute the similarity between the pose of the current frame and adjusted pose in the prior frame. Once obtaining the optical-based pose similarity, we combine it with the visual and bounding box spatial similarities to generate the final similarity matrix, and apply the Kuhn-Munkras algorithm for data association. The experiments on the MTMCT and MOT datasets verify the rationality of using human pose information and prove the superiority of the proposed method.
Sisi You, Hantao Yao, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2021 SAN: Selective Alignment Network for Cross-Domain Pedestrian Detection
abstract
Cross-domain pedestrian detection, which has been attracting much attention, assumes that the training and test images are drawn from different data distributions. Existing methods focus on aligning the descriptions of whole candidate instances between source and target domains. Since there exists a giant visual difference among the candidate instances, aligning whole candidate instances between two domains cannot overcome the inter-instance difference. Compared with aligning the whole candidate instances, we consider that aligning each type of instances separately is a more reasonable manner. Therefore, we propose a novel Selective Alignment Network for cross-domain pedestrian detection, which consists of three components: a Base Detector, an Image-Level Adaptation Network, and an Instance-Level Adaptation Network. The Image-Level Adaptation Network and Instance-Level Adaptation Network can be regarded as the global-level and local-level alignments, respectively. Similar to the Faster R-CNN, the Base Detector, which is composed of a Feature module, an RPN module and a Detection module, is used to infer a robust pedestrian detector with the annotated source data. Once obtaining the image description extracted by the Feature module, the Image-Level Adaptation Network is proposed to align the image description with an adversarial domain classifier. Given the candidate proposals generated by the RPN module, the Instance-Level Adaptation Network firstly clusters the source candidate proposals into several groups according to their visual features, and thus generates the pseudo label for each candidate proposal. After generating the pseudo labels, we align the source and target domains by maximizing and minimizing the discrepancy between the prediction of two classifiers iteratively. Extensive evaluations on several benchmarks demonstrate the effectiveness of the proposed approach for cross-domain pedestrian detection.
Yifan Jiao, Hantao Yao, Changsheng Xu
IEEE Trans. Image Process.2
2021 TEST: Triplet Ensemble Student-Teacher Model for Unsupervised Person Re-Identification
abstract
The self-ensembling methods have achieved amazing performance for semi-supervised representation learning and domain adaptation. However, the disadvantage of these methods is that the teacher network is tightly coupled with the student network, which limits the descriptive ability of the self-ensembling model. To overcome the coupling effect between the teacher network and the student network, we propose a novel Triplet Ensemble Student-Teacher (TEST) model for unsupervised person re-identification, which consists of one teacher network T and two student networks S1 and S2 . Similar to the traditional self-ensembling model, the student network S1 is applied to update the teacher network T . Furthermore, a closed-loop learning mechanism is built in the TEST model by imposing an ensemble consistent constraint between T and S2 , and performing a heterogeneous co-teaching procedure between S1 and S2 . With the closed-loop learning mechanism, the TEST model can loosen the constraint between the teacher T and the student S1 , and enhance the descriptive ability of S1 . Besides, the knowledge exchange between S1 and S2 can ensure that the two student networks can elegantly deal with the noisy labels and avoid coupling. By training the TEST model with the clustering-generated pseudo labels, we can achieve effective and robust representation learning for unsupervised person re-identification. The evaluations on three widely-used benchmarks show that our approach can achieve significant performance compared with state-of-the-art methods.
Hantao Yao, Changsheng Xu
IEEE Trans. Image Process.2
2021 Joint Person Objectness and Repulsion for Person Search
abstract
Person search targets to search the probe person from the unconstrainted scene images, which can be treated as the combination of person detection and person matching. However, the existing methods based on the Detection-Matching framework ignore the person objectness and repulsion (OR) which are both beneficial to reduce the effect of distractor images. In this paper, we propose an OR similarity by jointly considering the objectness and repulsion information. Besides the traditional visual similarity term, the OR similarity also contains an objectness term and a repulsion term. The objectness term can reduce the similarity of distractor images that not contain a person and boost the performance of person search by improving the ranking of positive samples. Because the probe person has a different person ID with its neighbors, the gallery images having a higher similarity with the neighbors of probe should have a lower similarity with the probe person. Based on this repulsion constraint, the repulsion term is proposed to reduce the similarity of distractor images that are not most similar to the probe person. Treating the Faster R-CNN as the person detector, the OR similarity is evaluated on PRW and CUHK-SYSU datasets by the Detection-Matching framework with six description models. The extensive experiments demonstrate that the proposed OR similarity can effectively reduce the similarity of distractor samples and further boost the performance of person search, e.g., improve the mAP from 92.32% to 93.23% for CUHK-SYSY dataset, and from 50.91% to 52.30% for PRW datasets.
Hantao Yao, Changsheng Xu
IEEE Trans. Image Process.1
2021 Attention-Based Multi-Source Domain Adaptation
abstract
Multi-source domain adaptation (MSDA) aims to transfer knowledge from multi-source domains to one target domain. Inspired by single-source domain adaptation, existing methods solve MSDA by aligning the data distributions between the target domain and each source domain. However, aligning the target domain with the dissimilar source domain would harm the representation learning. To address the above issue, an intuitive motivation of MSDA is using the attention mechanism to enhance the positive effects of the similar domains, and suppress the negative effects of the dissimilar domains. Therefore, we propose Attention-Based Multi-Source Domain Adaptation (ABMSDA) by considering the domain correlations to alleviate the effects caused by dissimilar domains. To obtain the domain correlations between source and target domains, ABMSDA firstly trains a domain recognition model to calculate the probability that the target images belong to each source domain. Based on the domain correlations, Weighted Moment Distance (WMD) is proposed to pay more attention on the source domains with higher similarities. Furthermore, Attentive Classification Loss (ACL) is developed to constrain that the feature extractor can generate the alignment and discriminative visual representations. The evaluations on two benchmarks demonstrate the effectiveness of the proposed model, e.g., an average of 6.1% improvement on the challenging DomainNet dataset.
Yukun Zuo, Hantao Yao, Changsheng Xu
IEEE Trans. Image Process.2
2021 Domain-Oriented Semantic Embedding for Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) targets to recognize images from new classes. Existing methods focus on learning a projection function to associate the visual features and category descriptions in the seen domain, which is directly transferred to the unseen domain. However, due to the inherent domain shift, a single shared projection cannot fully capture the domain difference and similarity, thereby making the unseen samples tend to be recognized as seen categories. In this paper, we propose a novel Domain-Oriented Semantic Embedding (DOSE) network that learns specific projections for different domains to better capture the domain characteristics for unbiased ZSL. Besides a domain-shared projection, DOSE learns two auxiliary domain-specific sub-projections to model the semantic-visual association in respective seen and unseen domains. Specifically, the domain-specific projections are learned in a cycle consistency way to capture domain characteristics, and a domain division constraint is developed to penalize the margin between two domain embeddings. Furthermore, to boost semantic-visual association, a semantic-visual dual attention module is designed to automatically remove trivial information in both visual and semantic embeddings under a co-guidance learning manner. Experiments on four public benchmarks prove that the proposed DOSE is robust to the domain shift problem in ZSL and obtains an averaged 5.6% improvement in terms of harmonic mean.
Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001
IEEE Trans. Multim.2
2021 Part-based Structured Representation Learning for Person Re-identification
abstract
Person re-identification aims to match person of interest under non-overlapping camera views. Therefore, how to generate a robust and discriminative representation is crucial for person re-identification. Mining local clues from human body parts to describe pedestrians has been extensively studied in existing methods. However, existing methods locate human body parts coarsely and do not consider the relations among different local parts. To address the above problem, we propose a Part-based Structured Representation Learning (PSRL) for better exploiting local clues to improve the person representation. There are two important modules in our architecture: Local Semantic Feature Extraction and Structured Person Representation Learning. The Local Semantic Feature Extraction module is designed to extract local features from human body semantic regions. After obtaining the local features, the Structured Person Representation Learning is proposed to fuse the local features by considering the person structure. To model the underlying person structure, a graph convolutional network is employed to capture the relations of different semantic regions. The generated structured feature encodes underlying person structure information, and local semantic feature can solve the misalignment problem caused by pose variations in feature matching. By combining them together, we can improve the descriptive ability of the generated representation. Extensive evaluations on four standard benchmarks show that our proposed method achieves competitive performance against state-of-the-art methods.
Hantao Yao, Tianzhu Zhang 0001, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning aims to recognize images from seen and unseen domains. Recent methods focus on learning a unified semantic-aligned visual representation to transfer knowledge between two domains, while ignoring the effect of semantic-free visual representation in alleviating the biased recognition problem. In this paper, we propose a novel Domain-aware Visual Bias Eliminating (DVBE) network that constructs two complementary visual representations, i.e., semantic-free and semantic-aligned, to treat seen and unseen domains separately. Specifically, we explore cross-attentive second-order visual statistics to compact the semantic-free representation, and design an adaptive margin Softmax to maximize inter-class divergences. Thus, the semantic-free representation becomes discriminative enough to not only predict seen class accurately but also filter out unseen images, i.e., domain detection, based on the predicted class entropy. For unseen images, we automatically search an optimal semantic-visual alignment architecture, rather than manual designs, to predict unseen classes. With accurate domain detection, the biased recognition problem towards the seen domain is significantly reduced. Experiments on five benchmarks for classification and segmentation show that DVBE outperforms existing methods by averaged 5.7% improvement.
Shaobo Min, Hantao Yao, Hongtao Xie 0001, Chaoqun Wang 0011, Zhengjun Zha, Yongdong Zhang 0001
CVPR2
2020 Category-Level Adversarial Self-Ensembling for Domain Adaptation
abstract
Domain adaptation aims at learning from a source data distribution a well-performing model on a different target data distribution. Recently, the self-ensembling-based methods have been proved to be effective for unsupervised domain adaptation. However, they still have two shortcomings: 1) no explicitly constraint about the distributions between the source and target domains; 2) the Euclidean distance fails to measure the similarity between two distributions with no overlap. To solve those shortcomings, we propose a novel Category-level Adversarial Self-ensembling (CAS) model for domain adaptation, which contains two types of consistency constraints. The first one is how to constrain the descriptions for source and target domains to be aligned. Therefore, we adopt a minimax game with a discrepancy loss between the category information generated by two classifiers. As the self-ensembling consists of two sub-networks: student and teacher networks, the second one is the consistency between those two networks for the target samples. Aiming to overcome the disadvantage of the Euclidean metric, we employ the Wasserstein distance to measure the difference between two probabilistic distributions. Experiments on several benchmarks demonstrate that our proposed CAS is superior to existing methods.
Yukun Zuo, Hantao Yao, Changsheng Xu
ICME2
2020 Hierarchical Granularity Transfer Learning
abstract
In the real world, object categories usually have a hierarchical granularity tree. Nowadays, most researchers focus on recognizing categories in a specific granularity, \emph{e.g.,} basic-level or sub(ordinate)-level. Compared with basic-level categories, the sub-level categories provide more valuable information, but its training annotations are harder to acquire. Therefore, an attractive problem is how to transfer the knowledge learned from basic-level annotations to sub-level recognition. In this paper, we introduce a new task, named Hierarchical Granularity Transfer Learning (HGTL), to recognize sub-level categories with basic-level annotations and semantic descriptions for hierarchical categories. Different from other recognition tasks, HGTL has a serious granularity gap,~\emph{i.e.,} the two granularities share an image space but have different category domains, which impede the knowledge transfer. To this end, we propose a novel Bi-granularity Semantic Preserving Network (BigSPN) to bridge the granularity gap for robust knowledge transfer. Explicitly, BigSPN constructs specific visual encoders for different granularities, which are aligned with a shared semantic interpreter via a novel subordinate entropy loss. Experiments on three benchmarks with hierarchical granularities show that BigSPN is an effective framework for Hierarchical Granularity Transfer Learning.
Shaobo Min, Hongtao Xie 0001, Hantao Yao, Xuran Deng, Zhengjun Zha, Yongdong Zhang 0001
NeurIPS3
2020 Multi-Objective Matrix Normalization for Fine-Grained Visual Recognition
abstract
Bilinear pooling achieves great success in fine-grained visual recognition (FGVC). Recent methods have shown that the matrix power normalization can stabilize the second-order information in bilinear features, but some problems, e.g., redundant information and over-fitting, remain to be resolved. In this paper, we propose an efficient Multi-Objective Matrix Normalization (MOMN) method that can simultaneously normalize a bilinear representation in terms of square-root, low-rank, and sparsity. These three regularizers can not only stabilize the second-order information, but also compact the bilinear features and promote model generalization. In MOMN, a core challenge is how to jointly optimize three non-smooth regularizers of different convex properties. To this end, MOMN first formulates them into an augmented Lagrange formula with approximated regularizer constraints. Then, auxiliary variables are introduced to relax different constraints, which allow each regularizer to be solved alternately. Finally, several updating strategies based on gradient descent are designed to obtain consistent convergence and efficient implementation. Consequently, MOMN is implemented with only matrix multiplication, which is well-compatible with GPU acceleration, and the normalized bilinear features are stabilized and discriminative. Experiments on five public benchmarks for FGVC demonstrate that the proposed MOMN is superior to existing normalization-based methods in terms of both accuracy and efficiency. The code is available: https://github.com/mboboGO/MOMN.
Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001
IEEE Trans. Image Process.2
2019 Adaptive Feature Fusion via Graph Neural Network for Person Re-identification
abstract
Person Re-identification (ReID) targets to identify a probe person appeared under multiple camera views. Existing methods focus on proposing a robust model to capture the discriminative information. However, they all generate a representation by mining useful clues from a given single image, and ignore the intercommunication with other images. To address this issue, we propose a novel network named Feature-Fusing Graph Neural Network (FFGNN), which fully utilizes the relationships among the nearest neighbors of the given image, and allows message propagation to update the feature of the node during representation learning. Given an anchor image, the FFGNN firstly obtains its Top-K nearest images based on the feature generated by the trained Feature-Extracting Network(FEN). We then construct a graph G based on the obtained K+1 images, in which each node represents the feature of an image. The edge of the graph G is obtained by combing the visual similarity and Jaccard similarity between nodes. Within the constructed graph G, FFGNN conducts message propagation and adaptive feature fusion between nodes by iteratively performing graph convolutional operation on the input features. Finally, the FFGNN outputs a robust and discriminative representation which contains the information from its similar images. Extensive experiments on three public person ReID datasets including Market-1501, DukeMTMC-ReID, and CUHK03 demonstrate that the proposed model can achieve significant improvement against state-of-the-art methods.
Hantao Yao, Ling-Yu Duan, Hanxing Yao, Changsheng Xu
ACM Multimedia2
2019 Domain-Specific Embedding Network for Zero-Shot Recognition
abstract
Zero-Shot Learning (ZSL) seeks to recognize a sample from either seen or unseen domain by projecting the image data and semantic labels into a joint embedding space. However, most existing methods directly adapt a well-trained projection from one domain to another, thereby ignoring the serious bias problem caused by domain differences. To address this issue, we propose a novel Domain-Specific Embedding Network (DSEN) that can apply specific projections to different domains for unbiased embedding, as well as several domain constraints. In contrast to previous methods, the DSEN decomposes the domain-shared projection function into one domain-invariant and two domain-specific sub-functions to explore the similarities and differences between two domains. To prevent the two specific projections from breaking the semantic relationship, a semantic reconstruction constraint is proposed by applying the same decoder function to them in a cycle consistency way. Furthermore, a domain division constraint is developed to directly penalize the margin between real and pseudo image features in respective seen and unseen domains, which can enlarge the inter-domain difference of visual features. Extensive experiments on four public benchmarks demonstrate the effectiveness of DSEN with an average of $9.2%$ improvement in terms of harmonic mean. The code is available in \urlhttps://github.com/mboboGO/DSEN-for-GZSL.
Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001
ACM Multimedia2
2019 Adaptive Bilinear Pooling for Fine-grained Representation Learning
abstract
Fine-grained representation learning targets to generate discriminative description for fine-grained visual objects. Recently, the bilinear feature interaction has been proved effective in generating powerful high-order representation with spatially invariant information. However, the existing methods apply a fixed feature interaction strategy to all samples, which ignore the image and region heterogeneity in a dataset. To this end, we propose a generalized feature interaction method, named Adaptive Bilinear Pooling (ABP), which can adaptively infer a suitable pooling strategy for a given sample based on image content. Specifically, ABP consists of two learning strategies: p-order learning (P-net) and spatial attention learning (S-net). The p-order learning predicts an optimal exponential coefficient rather than a fixed order number to extract moderate visual information from an image. The spatial attention learning aims to infer a weighted score that measures the importance of each local region, which can compact the image representations. To make ABP compatible with kernelized bilinear feature interaction, a crossed two-branch structure is utilized to combine the P-net and S-net. This structure can facilitate complementary information exchange between two different visual branches. The experiments on three widely used benchmarks, including fine-grained object classification and action recognition, demonstrate the effectiveness of the proposed method.
Shaobo Min, Hongtao Xie 0001, Youliang Tian, Hantao Yao, Yongdong Zhang 0001
MMAsia4
2019 DR2-Net: Deep Residual Reconstruction Network for image compressive sensing
Hantao Yao, Shiliang Zhang, Yongdong Zhang 0001, Qi Tian 0001, Changsheng Xu
Neurocomputing1
2019 Deep Representation Learning With Part Loss for Person Re-Identification
abstract
Learning discriminative representations for unseen person images is critical for person Re-Identification (ReID). Most of current approaches learn deep representations in classification tasks, which essentially minimizes the empirical classification risk on the training set. As shown in our experiments, such representations easily get over-fitted on a discriminative human body part on the training set. To gain the discriminative power on unseen person images, we propose a deep representation learning procedure named Part Loss Network (PL-Net), to minimize both the empirical classification risk on training person images and the representation learning risk on unseen person images. The representation learning risk is evaluated by the proposed part loss, which automatically detects human body parts, and computes the person classification loss on each part separately. Compared with traditional global classification loss, simultaneously considering part loss enforces the deep network to learn representations for different body parts and gain the discriminative power on unseen persons. Experimental results on three person ReID datasets, i.e., Market1501, CUHK03, VIPeR, show that our representation outperforms existing deep representations.
Hantao Yao, Shiliang Zhang, Richang Hong, Yongdong Zhang 0001, Changsheng Xu, Qi Tian 0001
IEEE Trans. Image Process.1
2019 GLAD: Global-Local-Alignment Descriptor for Scalable Person Re-Identification
abstract
The huge variance of human pose and the misalign-ment of detected human images significantly increase the difficulty of pedestrian image matching in person Re-Identification (Re-ID). Moreover, the massive visual data being produced by surveillance video cameras requires highly efficient person Re-ID systems. Targeting to solve the first problem, this work proposes a robust and discriminative pedestrian image descriptor, namely, the Global-Local-Alignment Descriptor (GLAD). For the second problem, this work treats person Re-ID as image retrieval and proposes an efficient indexing and retrieval framework. GLAD explicitly leverages the local and global cues in the human body to generate a discriminative and robust representation. It consists of part extraction and descriptor learning modules, where several part regions are first detected and then deep neural networks are designed for representation learning on both the local and global regions. A hierarchical indexing and retrieval framework is designed to perform offline relevance mining to eliminate the huge person ID redundancy in the gallery set, and accelerate the online Re-ID procedure. Extensive experimental results on widely used public benchmark datasets show GLAD achieves competitive accuracy compared to the state-of-the-art methods. On a large-scale person, with the Re-ID dataset containing more than 520 K images, our retrieval framework significantly accelerates the online Re-ID procedure while also improving Re-ID accuracy. Therefore, this work has the potential to work better on person Re-ID tasks in real scenarios.
Longhui Wei, Shiliang Zhang, Hantao Yao, Wen Gao 0001, Qi Tian 0001
IEEE Trans. Multim.3
2018 AutoBD: Automated Bi-Level Description for Scalable Fine-Grained Visual Categorization
abstract
Compared with traditional image classification, fine-grained visual categorization is a more challenging task, because it targets to classify objects belonging to the same species, e.g., classify hundreds of birds or cars. In the past several years, researchers have made many achievements on this topic. However, most of them are heavily dependent on the artificial annotations, e.g., bounding boxes, part annotations, and so on. The requirement of artificial annotations largely hinders the scalability and application. Motivated to release such dependence, this paper proposes a robust and discriminative visual description named Automated Bi-level Description (AutoBD). “Bi-level” denotes two complementary part-level and object-level visual descriptions, respectively. AutoBD is “automated,” because it only requires the image-level labels of training images and does not need any annotations for testing images. Compared with the part annotations labeled by the human, the image-level labels can be easily acquired, which thus makes AutoBD suitable for large-scale visual categorization. Specifically, the part-level description is extracted by identifying the local region saliently representing the visual distinctiveness. The object-level description is extracted from object bounding boxes generated with a co-localization algorithm. Although only using the image-level labels, AutoBD outperforms the recent studies on two public benchmark, i.e., classification accuracy achieves 81.6% on CUB-200-2011 and 88.9% on Car-196, respectively. On the large-scale Birdsnap data set, AutoBD achieves the accuracy of 68%, which is currently the best performance to the best of our knowledge.
Hantao Yao, Shiliang Zhang, Chenggang Yan 0001, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001
IEEE Trans. Image Process.1
2017 Large-scale person re-identification as retrieval
abstract
This paper targets to bring together the research efforts on two fields that are growing actively in the past few years: multicamera person Re-Identification (ReID) and large-scale image retrieval. We demonstrate that the essentials of image retrieval and person ReID are the same, i.e., measuring the similarity between images. However, person ReID requires more discriminative and robust features to identify the subtle differences of different persons and overcome the large variance among images of the same person. Specifically, we propose a coarse-to-fine (C2F) framework and a Convolutional Neural Network structure named as Conv-Net to tackle the large-scale person ReID as an image retrieval task. Given a query person image, the C2F firstly employ Conv-Net to extract a compact descriptor and perform the coarse-level search. A robust descriptor conveying more spatial cues is hence extracted to perform the fine-level search. Extensive experimental results show that the proposed method outperforms existing methods on two public datasets. Further, the evaluation on a large-scale Person-520K dataset demonstrates that our work is significantly more efficient than existing works, e.g., only needs 180ms to identify a query person from 520K images.
Hantao Yao, Shiliang Zhang, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Yu Wang 0089, Qi Tian 0001
ICME1
2017 GLAD: Global-Local-Alignment Descriptor for Pedestrian Retrieval
abstract
The huge variance of human pose and the misalignment of detected human images significantly increase the difficulty of person Re-Identification (Re-ID). Moreover, efficient Re-ID systems are required to cope with the massive visual data being produced by video surveillance systems. Targeting to solve these problems, this work proposes a Global-Local-Alignment Descriptor (GLAD) and an efficient indexing and retrieval framework, respectively. GLAD explicitly leverages the local and global cues in human body to generate a discriminative and robust representation. It consists of part extraction and descriptor learning modules, where several part regions are first detected and then deep neural networks are designed for representation learning on both the local and global regions. A hierarchical indexing and retrieval framework is designed to eliminate the huge redundancy in the gallery set, and accelerate the online Re-ID procedure. Extensive experimental results show GLAD achieves competitive accuracy compared to the state-of-the-art methods. Our retrieval framework significantly accelerates the online Re-ID procedure without loss of accuracy. Therefore, this work has potential to work better on person Re-ID tasks in real scenarios.
Longhui Wei, Shiliang Zhang, Hantao Yao, Wen Gao 0001, Qi Tian 0001
ACM Multimedia3
2017 One-Shot Fine-Grained Instance Retrieval
abstract
Fine-Grained Visual Categorization (FGVC) has achieved significant progress recently. However, the number of fine-grained species could be huge and dynamically increasing in real scenarios, making it difficult to recognize unseen objects under the current FGVC framework. This raises an open issue to perform large-scale fine-grained identification without a complete training set. Aiming to conquer this issue, we propose a retrieval task named One-Shot Fine-Grained Instance Retrieval (OSFGIR). "One-Shot" denotes the ability of identifying unseen objects through a fine-grained retrieval task assisted with an incomplete auxiliary training set. This paper first presents the detailed description to OSFGIR task and our collected OSFGIR-378K dataset. Next, we propose the Convolutional and Normalization Networks (CN-Nets) learned on the auxiliary dataset to generate a concise and discriminative representation. Finally, we present a coarse-to-fine retrieval framework consisting of three components, i.e., coarse retrieval, fine-grained retrieval, and query expansion, respectively. The framework progressively retrieves images with similar semantics, and performs fine-grained identification. Experiments show our OSFGIR framework achieves significantly better accuracy and efficiency than existing FGVC and image retrieval methods, thus could be a better solution for large-scale fine-grained object identification.
Hantao Yao, Shiliang Zhang, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001
ACM Multimedia1
2017 DSP: Discriminative Spatial Part modeling for Fine-Grained Visual Categorization
Hantao Yao, Dongming Zhang 0004, Jintao Li 0001, Jianshe Zhou, Shiliang Zhang, Yongdong Zhang 0001
Image Vis. Comput.1
2016 Coarse-to-Fine Description for Fine-Grained Visual Categorization
abstract
Recent years have witnessed the significant advance in fine-grained visual categorization, which targets to classify the objects belonging to the same species. To capture enough subtle visual differences and build discriminative visual description, most of the existing methods heavily rely on the artificial part annotations, which are expensive to collect in real applications. Motivated to conquer this issue, this paper proposes a multi-level coarse-to-fine object description. This novel description only requires the original image as input, but could automatically generate visual descriptions discriminative enough for fine-grained visual categorization. This description is extracted from five sources representing coarse-to-fine visual clues: 1) original image is used as the source of global visual clue; 2) object bounding boxes are generated using convolutional neural network (CNN); 3) with the generated bounding box, foreground is segmented using the proposed k nearest neighbour-based co-segmentation algorithm; and 4) two types of part segmentations are generated by dividing the foreground with an unsupervised part learning strategy. The final description is generated by feeding these sources into CNN models and concatenating their outputs. Experiments on two public benchmark data sets show the impressive performance of this coarse-to-fine description, i.e., classification accuracy achieves 82.5% on CUB-200-2011, and 86.9% on fine-grained visual categorization-Aircraft, respectively, which outperform many recent works.
Hantao Yao, Shiliang Zhang, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001
IEEE Trans. Image Process.1