Nan Pu

dblp:210/5100 · DBLP profile ↗
← Back
45ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0002-2179-8301ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 26 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 21 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
abstract
The proliferation of synthetic facial imagery has intensified the need for robust Open-World DeepFake Attribution (OW-DFA), which aims to attribute both known and unknown forgeries using labeled data for known types and unlabeled data containing a mixture of known and novel types. However, existing OW-DFA methods face two critical limitations: 1) A confidence skew that leads to unreliable pseudo-labels for novel forgeries, resulting in biased training. 2) An unrealistic assumption that the number of unknown forgery types is known a priori. To address these challenges, we propose a Confidence-aware Asymmetric Learning (CAL) framework, which adaptively balances model confidence across known and novel forgery types. CAL mainly consists of two components: Confidence-aware Consistency Regularization (CCR) and Asymmetric Confidence Reinforcement (ACR). CCR mitigates pseudo-label bias by dynamically scaling sample losses based on normalized confidence, gradually shifting the training focus from high- to low-confidence samples. ACR complements this by separately calibrating confidence for known and novel classes through selective learning on high-confidence samples, guided by their confidence gap. Together, CCR and ACR form a mutually reinforcing loop that significantly improves the model's OW-DFA performance. Moreover, we introduce a Dynamic Prototype Pruning (DPP) strategy that automatically estimates the number of novel forgery types in a coarse-to-fine manner, removing the need for unrealistic prior assumptions and enhancing the scalability of our methods to real-world OW-DFA scenarios. Extensive experiments on the standard and OW-DFA benchmark and a newly extended benchmark incorporating advanced manipulations demonstrate that CAL consistently outperforms previous methods, achieving new state-of-the-art performance on both known and novel forgery attribution.
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
AAAI2
2026 Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model
Junhui Yin, Nan Pu, Lingfeng Yang, Zhun Zhong
Int. J. Comput. Vis.2
2026 Towards Stable Source-Free Domain Adaptive Semantic Segmentation
Dong Zhao 0007, Qi Zang, Nan Pu, Jinlong Li 0003, Shuang Wang 0001, Nicu Sebe, Zhun Zhong
Int. J. Comput. Vis.3
2026 ProtoSimNet: Exploiting prototype similarity for few-shot learning
abstract
Many few-shot learning (FSL) approaches exploit Euclidean Distance as the metric, as it provides a stable and well-understood geometric structure for prototype-based classification. Despite these advantages, we observe that Euclidean Distance metric may suffer from a metric-level limitation under few-shot supervision. Meanwhile, with limited supervision, cosine-based scoring often leads to a metric over-smoothing phenomenon, where similarity scores among competing prototypes become overly compressed, producing weak effective margins and less stable decision boundaries, resulting in more ambiguous predictions when distinguishing fine-grained novel categories. In this work, we identify this metric-level limitation and propose a problem-driven prototype similarity framework, termed ProtoSimNet, for FSL which explicitly performs score-level distribution shaping of prototype similarities under data-scarce episodic supervision. ProtoSimNet introduces a cosine sharpening transformation, combined with an elastic regularization, which together enhance discriminability by amplifying the margins between high- and low-confidence prototype similarity scores while regularizing the encoder toward compact and less redundant feature embeddings. Our method strengthens the reliability of cosine-based classification and improves the robustness of prototype matching when only a few labeled samples are available. Experimental results on FSL benchmark datasets show that ProtoSimNet achieves consistent improvements over standard Prototypical Networks and competitive or superior performance compared with representative state-of-the-art methods.
Nan Pu, Erwin M. Bakker, Michael S. Lew
Neurocomputing2
2026 Identity-Compensated Style Distillation for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-Identification (VI-ReID) that matches pedestrian images across visible and infrared modalities suffers from substantial modality discrepancies and intra-class variations. While existing methods typically address the modality gap via style alignment, they often lose identity-relevant semantics and overlook fine-grained inter-class nuances, such as body part contours and structural cues around the head, shoulders, or feet. To tackle these challenges, we propose an Identity-Compensated Style Distillation (ICSD) network that enforces cross-modality style consistency and enhances the discriminative power of modality-invariant features. Specifically, ICSD comprises two core components: (1) a Style Knowledge Distillation (SKD) module, which integrates Style Discrepancy Reduction (SDR) and Identity Knowledge Compensation (IKC) to align modality styles while preserving identity-relevant semantics; (2) an Identity Discrimination Amplification (IDA) module, which captures and enhances subtle inter-class differences by refining identity-specific cues, thereby facilitating more accurate discrimination between different pedestrians. Extensive experiments on three public benchmarks-SYSU-MM01, RegDB, and LLCM-demonstrate that ICSD consistently outperforms state-of-the-art methods, validating the effectiveness and complementarity of its components.
Yongguo Ling, Zihao Hu, Nan Pu, Zhun Zhong, Xudong Jiang 0001
IEEE Trans. Image Process.3
2025 Multi-Scale Global-Instance Prompt Tuning for Continual Test-Time Adaptation in Medical Image Segmentation
abstract
Distribution shift is a common challenge in medical images obtained from different clinical centers, significantly hindering the deployment of pre-trained semantic segmentation models in real-world applications across multiple domains. Continual Test-Time Adaptation (CTTA) has emerged as a promising approach to address cross-domain distribution shifts during continually evolving target domains. Most existing CTTA methods rely on incrementally updating model parameters, which inevitably suffer from error accumulation and catastrophic forgetting, especially in long-term adaptation. Recent prompt-tuning-based works have shown potential to mitigate the two issues above by updating only visual prompts. While these approaches have demonstrated promising performance, several limitations remain: 1) lacking multi-scale prompt diversity, 2) inadequate incorporation of instance-specific knowledge, and 3) risk of privacy leakage. To overcome these limitations, we propose Multi-scale Global-Instance Prompt Tuning (MGIPT), to enhance scale diversity of prompts as well as capture both globaland instance-level knowledge for robust CTTA. Specifically, MGIPT consists of an Adaptive-scale Instance Prompt (AIP) and a Multi-scale Global-level Prompt (MGP). AIP dynamically learns lightweight and instance-specific prompts to mitigate error accumulation with adaptive optimal-scale selection mechanism. MGP captures domain-level knowledge across different scales to ensure robust adaptation with anti-forgetting capabilities. These complementary components are combined through a weighted ensemble approach, enabling effective dual-level adaptation that integrates both global and local information. Extensive experiments on medical image segmentation benchmarks (five optic disc/cup datasets and four polyp datasets) demonstrate that our MGIPT outperforms state-of-the-art methods, achieving robust adaptation across continually changing target domains. Notably, our MGIPT exhibits particularly strong performance in longterm CTTA scenarios, showing great anti-forgetting ability.
Lingrui Li, Yanfeng Zhou, Nan Pu, Xin Chen 0003, Zhun Zhong
BIBM3
2025 Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery
abstract
In this paper, we investigate a practical yet challenging task: On-the-fly Category Discovery (OCD). This task focuses on the online identification of newly arriving stream data that may belong to both known and unknown categories, utilizing the category knowledge from only labeled data. Existing OCD methods are devoted to fully mining transferable knowledge from only labeled data. However, the transferability learned by these methods is limited because the knowledge contained in known categories is often insufficient, especially when few annotated data/categories are available in fine-grained recognition. To mitigate this limitation, we propose a diffusion-based OCD framework, dubbed DiffGRE, which integrates Generation, Refinement, and Encoding in a multi-stage fashion. Specifically, we first design an attribute-composition generation method based on cross-image interpolation in the diffusion latent space to synthesize novel samples. Then, we propose a diversity-driven refinement approach to select the synthesized images that differ from known categories for subsequent OCD model training. Finally, we leverage a semi-supervised leader encoding to inject additional category knowledge contained in synthesized data into the OCD models, which can benefit the discovery of both known and unknown categories during the on-the-fly inference process. Extensive experiments demonstrate the superiority of our DiffGRE over previous methods on six fine-grained datasets.
Nan Pu, Haiyang Zheng, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
ICCV2
2025 Robust Consensus Anchor Learning for Efficient Multi-view Subspace Clustering
abstract
As a leading unsupervised classification algorithm in artificial intelligence, multi-view subspace clustering segments unlabeled data from different subspaces. Recent works based on the anchor have been proposed to decrease the computation complexity for the datasets with large scales in multi-view clustering. The major differences among these methods lie on the objective functions they define. Despite considerable success, these works pay few attention to guaranting the robustness of learned consensus anchors via effective manner for efficient multi-view clustering and investigating the specific local distribution of cluster in the affine subspace. Besides, the robust consensus anchors as well as the common cluster structure shared by different views are not able to be simultaneously learned. In this paper, we propose Robust Consensus anchors learning for efficient multi-view Subspace Clustering (RCSC). We first show that if the data are sufficiently sampled from independent subspaces, and the objective function meets some conditions, the achieved anchor graph has the block-diagonal structure. As a special case, we provide a model based on Frobenius norm, non-negative and affine constraints in consensus anchors learning, which guarantees the robustness of learned consensus anchors for efficient multi-view clustering and investigates the specific local distribution of cluster in the affine subspace. While it is simple, we theoretically give the geometric analysis regarding the formulated RCSC. The union of these three constraints is able to restrict how each data point is described in the affine subspace with specific local distribution of cluster for guaranting the robustness of learned consensus anchors. RCSC takes full advantages of correlation among consensus anchors, which encourages the grouping effect and groups highly correlated consensus anchors together with the guidance of view-specific projection. The anchor graph construction, partition and robust anchor learning are jointly integrated into a unified framework. It ensures the mutual enhancement for these procedures and helps lead to more discriminative consensus anchors as well as the cluster indicator. We then adopt an alternative optimization strategy for solving the formulated problem. Experiments performed on eight multi-view datasets confirm the superiority of RCSC based on the effectiveness and efficiency.
Yalan Qin, Nan Pu, Guorui Feng, Nicu Sebe
ICML2
2025 Flexible Multi-view Clustering with Dynamic Views Generation
abstract
Multi-view clustering is one of the fundamental unsupervised multimedia analysis tasks. Recent studies have mainly focused on developing multi-view clustering approaches, which can achieve state-of-the-art clustering performance. However, most of the existing works just focus on multi-view clustering with fixed views, which lacks flexibility with guidance of the views dynamically generated. Besides, these works ignore to integrate generating views in a dynamic manner and learning the common representation shared by different views into a unified framework. To this end, we propose the Flexible Multi-view Clustering with Dynamic Views Generation (FMCDVG). Specifically, FMCDVG adopts the graph convolutional network and auto-encoder to dynamically generate the topological graph representation and node attribute representation as two different views, respectively. FMCDVG introduces the latent representation shared by different feature representations and integrates multiple feature representations based on node attributes and graph structure into the latent representation with reconstruction through reconstructed encoding networks (REN). FMCDVG jointly conducts generating views in a dynamic manner and learning the common representation shared by different views in a unified optimization framework. We demonstrate that FMCDVG is able to consistently achieve better clustering performance than the state-of-the-art methods through comprehensive experiments.
Yalan Qin, Nan Pu, Hanzhou Wu, Zhaoxin Fan
ACM Multimedia2
2025 Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
abstract
The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly reflect real-world manipulation patterns: practical attacks typically maintain semantic consistency across modalities, whereas current datasets artificially disrupt cross-modal alignment, creating easily detectable anomalies. To bridge this gap, we pioneer the detection of semantically-coordinated manipulations where visual edits are systematically paired with semantically consistent textual descriptions. Our approach begins with constructing the first Semantic-Aligned Multimodal Manipulation (SAMM) dataset, generated through a two-stage pipeline: 1) applying state-of-the-art image manipulations, followed by 2) generation of contextually-plausible textual narratives that reinforce the visual deception. Building on this foundation, we propose a Retrieval-Augmented Manipulation Detection and Grounding (RamDG) framework. RamDG commences by harnessing external knowledge repositories to retrieve contextual evidence, which serves as the auxiliary texts and encoded together with the inputs through our image forgery grounding and deep manipulation detection modules to trace all manipulations. Extensive experiments demonstrate our framework significantly outperforms existing methods, achieving 2.06% higher detection accuracy on SAMM compared to state-of-the-art approaches. The dataset and code are publicly available at https://github.com/shen8424/SAMM-RamDG-CAP.
Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu, Zhun Zhong
ACM Multimedia4
2025 Bilevel progressive homography estimation via correlative region-focused transformer
Qi Jia 0001, Xiaomei Feng, Wei Zhang 0339, Yu Liu 0012, Nan Pu, Nicu Sebe
Comput. Vis. Image Underst.5
2025 Few-shot event-based action recognition
abstract
Despite the evident superiority of event cameras in practical vision applications (e.g., action recognition), owing to their distinctive sensing mechanism, existing event-based action recognition methods rely heavily on large-scale training data. However, the expensive cost of camera deployment and the requirement of data privacy protection make it challenging to collect substantial data in real-world scenarios. To address this limitation, we explore a novel yet practical task, Few-Shot Event-Based Action Recognition (FSEAR), which aims at leveraging a minimal number of intractable event action data for model training and accurately classifying unlabeled data into a specific category. Accordingly, we design a new framework for FSEAR, including a Noise-Aware Event Encoder (NAE) and a Distilled Prototypical Distance Fusion (DPDF). The former efficiently filters noise within the spatiotemporal domain while retaining vital information related to action timing. The latter conducts multi-scale measurements across geometric, directional, and distributional dimensions. These two modules benefit mutually and thus effectively exploit the potential characteristics of event data. Extensive experiments on four distinct event action recognition datasets have demonstrated the significant advantages of our model over other few-shot learning methods. Our code and models will be publicly released.
Zanxi Ruan, Nan Pu, Jiangming Chen, Songqun Gao, Yanming Guo, Qiuyu Kong, Yuxiang Xie, Yingmei Wei
Neural Networks2
2025 SeCoV2: Semantic Connectivity-Driven Pseudo-Labeling for Robust Cross-Domain Semantic Segmentation
abstract
Pseudo-labeling is a dominant strategy for cross-domain semantic segmentation (CDSS), yet its effectiveness is limited by fragmented and noisy pixel-level predictions under severe domain shifts. To address this, we propose a semantic connectivity-driven pseudo-labeling framework, SeCo, which constructs and refines pseudo-labels at the connectivity level by aggregating high-confidence pixels into coherent semantic regions. The framework includes two key components: Pixel Semantic Aggregation (PSA), which leverages a dual prompting strategy to preserve category-specific granularity, and Semantic Connectivity Correction with Loss Distribution (SCC-LD), which filters noisy regions based on early-loss statistics. Building upon this foundation, we further present SeCoV2, which introduces SCC-Unc, a novel uncertainty-aware correction module that constructs a connectivity graph and enforces relational consistency for robust refinement in ambiguous regions. SeCoV2 also broadens the applicability of SeCo by extending evaluation to more challenging scenarios, including open-set and multimodal adaptation, semi-supervised domain generalization, and by validating compatibility with different interactive foundation segmentation models such as SAM Kirillov et al. 2023, SEEM Zou et al. 2023, and Fast-SAM Zhao et al. 2023. Extensive experiments across six CDSS tasks demonstrate that SeCoV2 achieves consistent improvements over previous methods, with an average performance gain of up to +4.6%, establishing new state-of-the-art results. These findings highlight the effectiveness and generalization ability for robust adaptation in diverse real-world environments.
Dong Zhao 0007, Qi Zang, Nan Pu, Shuang Wang 0001, Nicu Sebe, Zhun Zhong
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Latent Space Learning-Based Ensemble Clustering
abstract
Ensemble clustering fuses a set of base clusterings and shows promising capability in achieving more robust and better clustering results. The existing methods usually realize ensemble clustering by adopting a co-association matrix to measure how many times two data points are categorized into the same cluster based on the base clusterings. Though great progress has been achieved, the obtained co-association matrix is constructed based on the combination of different connective matrices or its variants. These methods ignore exploring the inherent latent space shared by multiple connective matrices and learning the corresponding co-association matrices according to this latent space. Moreover, these methods neglect to learn discriminative connective matrices, explore the high-order relation among these connective matrices and consider the latent space in a unified framework. In this paper, we propose a Latent spacE leArning baseD Ensemble Clustering (LEADEC), which introduces the latent space shared by different connective matrices and learns the corresponding connective matrices according to this latent space. Specifically, we factorize the original multiple connective matrices into a consensus latent space representation and the specific connective matrices. Meanwhile, the orthogonal constraint is imposed to make the latent space representation more discriminative. In addition, we collect the obtained connective matrices based on the latent space into a tensor with three orders to investigate the high-order relations among these connective matrices. The connective matrices learning, the high-order relation investigation among connective matrices and the latent space representation learning are integrated into a unified framework. Experiments on seven benchmark datasets confirm the superiority of LEADEC compared with the existing representive methods.
Yalan Qin, Nan Pu, Nicu Sebe, Guorui Feng
IEEE Trans. Image Process.2
2025 Margin-aware Noise-robust Contrastive Learning for Partially View-aligned Problem
abstract
In this article, we study a challenging problem in contrastive learning when just a portion of data is aligned in multi-view dataset due to temporal, spatial, or spatio-temporal asynchronism across views. It is important to study partially view-aligned data since this type of data is common in real-world application and easily leads to data inconsistency among different views. Such a Partially View-aligned Problem (PVP) in contrastive learning has been relatively less touched so far, especially in downstream tasks, i.e., classification and clustering. In order to solve this problem, we introduce a flexible margin and propose margin-aware noise-robust contrastive learning to simultaneously identify the within-category counterparts from the other view of one data point based on the established cross-view correspondence and learn a shared representation. To be specific, the proposed learning framework is built on a novel margin-aware noise-robust contrastive loss. Since data pairs are used as input for the proposed margin-aware noise-robust contrastive learning, we build positive pairs according to the known correspondences and negative pairs in the manner of random sampling. Our margin-aware noise-robust contrastive learning framework is able to effectively reduce or remove the impacts caused by the possible existing noise for the constructed pairs in a margin-aware manner, i.e., false negative pairs led by random sampling in PVP. We relax the proposed margin-aware noise-robust contrastive loss and then give a detailed mathematical analysis for the effectiveness of our loss. As an instantiation, we construct an example under the proposed margin-aware noise-robust contrastive learning framework for validation in this work. To the best of our knowledge, this is the first attempt of extending contrastive learning to a margin-aware noise-robust version for dealing with PVP. We also enrich the learning paradigm when there is noise in the data. Extensive experiments on different datasets demonstrate the promising performance of the proposed method in the classification and clustering tasks.
Yalan Qin, Nan Pu, Hanzhou Wu, Nicu Sebe
ACM Trans. Knowl. Discov. Data2
2025 Discriminative Anchor Learning for Efficient Multi-View Clustering
abstract
Multi-view clustering aims to study the complementary information across views and discover the underlying structure. For solving the relatively high computational cost for the existing approaches, works based on anchor have been presented recently. Even with acceptable clustering performance, these methods tend to map the original representation from multiple views into a fixed shared graph based on the original dataset. However, most studies ignore the discriminative property of the learned anchors, which ruin the representation capability of the built model. Moreover, the complementary information among anchors across views is neglected to be ensured by simply learning the shared anchor graph without considering the quality of view-specific anchors. In this paper, we propose discriminative anchor learning for multi-view clustering (DALMC) for handling the above issues. We learn discriminative view-specific feature representations according to the original dataset and build anchors from different views based on these representations, which increase the quality of the shared anchor graph. The discriminative feature learning and consensus anchor graph construction are integrated into a unified framework to improve each other for realizing the refinement. The optimal anchors from multiple views and the consensus anchor graph are learned with the orthogonal constraints. We give an iterative algorithm to deal with the formulated problem. Extensive experiments on different datasets show the effectiveness and efficiency of our method compared with other methods.
Yalan Qin, Nan Pu, Hanzhou Wu, Nicu Sebe
IEEE Trans. Multim.2
2024 Novel Class Discovery for Ultra-Fine-Grained Visual Categorization
abstract
Ultra-fine-grained visual categorization (Ultra-FGVC) aims at distinguishing highly similar sub-categories within fine-grained objects, such as different soybean cultivars. Compared to traditional fine-grained visual categorization, Ultra-FGVC encounters more hurdles due to the small inter-class and large intra-class variation. Given these challenges, relying on human annotation for Ultra-FGVC is impractical. To this end, our work introduces a novel task termed Ultra-Fine-Grained Novel Class Discovery (UFG-NCD), which leverages partially annotated data to identify new categories of unlabeled images for Ultra-FGVC. To tackle this problem, we devise a Region-Aligned Proxy Learning (RAPL) framework, which comprises a Channel-wise Region Alignment (CRA) module and a Semi-Supervised Proxy Learning (SemiPL) strategy. The CRA module is designed to extract and utilize discriminative features from local regions, facilitating knowledge transfer from labeled to unlabeled classes. Furthermore, SemiPL strengthens representation learning and knowledge transfer with proxy-guided supervised learning and proxy-guided contrastive learning. Such techniques leverage class distribution information in the embedding space, improving the mining of subtle differences between labeled and unlabeled ultra-fine-grained classes. Extensive experiments demonstrate that RAPL significantly outperforms baselines across various datasets, indicating its effectiveness in handling the challenges of UFG-NCD. Code is available at https://github.com/SSDUT-Caiyq/UFG-NCD.
Yu Liu 0012, Yaqi Cai, Qi Jia 0001, Binglin Qiu, Weimin Wang 0007, Nan Pu
CVPR6
2024 Federated Generalized Category Discovery
abstract
Generalized category discovery (GCD) aims at grouping unlabeled samples from known and unknown classes, given labeled data of known classes. To meet the recent decen-tralization trend in the community, we introduce a practical yet challenging task, Federated GCD (Fed-GCD), where the training data are distributed among local clients and cannot be shared among clients. Fed-GCD aims to train a generic GCD model by client collaboration under the privacy-protected constraint. The Fed-GCD leads to two challenges: 1) representation degradation caused by training each client model with fewer data than centralized GCD learning, and 2) highly heterogeneous label spaces across different clients. To this end, we propose a novel Asso-ciated Gaussian Contrastive Learning (AGCL) framework based on learnable GMMs, which consists of a Client Se-mantics Association (CSA) and a global-local GMM Contrastive Learning (GCL). On the server, CSA aggregates the heterogeneous categories of local-client GMMs to generate a global GMM containing more comprehensive category knowledge. On each client, GCL builds class-level contrastive learning with both local and global GMMs. The local GCL learns robust representation with limited local data. The global GCL encourages the model to produce more discriminative representation with the comprehensive category relationships that may not exist in local data. We build a benchmark based on six visual datasets to facilitate the study of Fed-GCD. Extensive experiments show that our AGCL outperforms multiple baselines on all datasets. Code is available at https://github.com/TPCD/FedGCD.
Nan Pu, Wenjing Li 0005, Xingyuan Ji, Yalan Qin, Nicu Sebe, Zhun Zhong
CVPR1
2024 Learning to Distinguish Samples for Generalized Category Discovery
Fengxiang Yang, Nan Pu, Wenjing Li 0005, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong
ECCV (65)2
2024 Textual Knowledge Matters: Cross-Modality Co-teaching for Generalized Visual Class Discovery
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
ECCV (52)2
2024 Fire and Smoke Detection with Burning Intensity Representation
Xiaoyi Han, Yanfei Wu, Nan Pu, Zunlei Feng, Qifei Zhang 0001, Yijun Bei, Lechao Cheng
MMAsia3
2024 Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery
abstract
In this paper, we study a practical yet challenging task, On-the-fly Category Discovery (OCD), aiming to online discover the newly-coming stream data that belong to both known and unknown classes, by leveraging only known category knowledge contained in labeled data. Previous OCD methods employ the hash-based technique to represent old/new categories by hash codes for instance-wise inference. However, directly mapping features into low-dimensional hash space not only inevitably damages the ability to distinguish classes and but also causes ``high sensitivity'' issue, especially for fine-grained classes, leading to inferior performance. To address these drawbacks, we propose a novel Prototypical Hash Encoding (PHE) framework consisting of Category-aware Prototype Generation (CPG) and Discriminative Category Encoding (DCE) to mitigate the sensitivity of hash code while preserving rich discriminative information contained in high-dimension feature space, in a two-stage projection fashion. CPG enables the model to fully capture the intra-category diversity by representing each category with multiple prototypes. DCE boosts the discrimination ability of hash code with the guidance of the generated category prototypes and the constraint of minimum separation distance. By jointly optimizing CPG and DCE, we demonstrate that these two components are mutually beneficial towards an effective OCD. Extensive experiments show the significant superiority of our PHE over previous methods, e.g. obtaining an improvement of +5.3% in ALL ACC averaged on all datasets. Moreover, due to the nature of the interpretable prototypes, we visually analyze the underlying mechanism of how PHE helps group certain samples into either known or unknown categories. Code is available at https://github.com/HaiyangZheng/PHE.
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
NeurIPS2
2024 Benchmarking Multi-Scene Fire and Smoke Detection
Xiaoyi Han, Nan Pu, Zunlei Feng, Yijun Bei, Qifei Zhang 0001, Lechao Cheng
PRCV (11)2
2024 PMGNet: Disentanglement and entanglement benefit mutually for compositional zero-shot learning
Yu Liu 0012, Yanyi Zhang, Qi Jia 0001, Weimin Wang 0007, Nan Pu, Nicu Sebe
Comput. Vis. Image Underst.6
2024 Elastic Multi-View Subspace Clustering With Pairwise and High-Order Correlations
abstract
Multi-view clustering has become an important research topic in machine learning and computer vision communities, which aims at achieving a consensus partition of data points across different views. However, the existing multi-view clustering methods fail to simultaneously consider the pairwise and high-order correlations among different views in the process of obtaining the final results. In this paper, we propose the Elastic multi-view Subspace Clustering with pairwise and high-order Correlations (ESCC) to solve this problem. ESCC simultaneously explores the pairwise and high-order correlations among different views, resulting in a more comprehensive shared representation. ESCC formulates these two kinds of correlations into a unified objective framework, which are able to be jointly optimized to refine each other. As an instantiation, we construct an example of ESCC (e-ESCC) in this work. To be specific, e-ESCC uses the multi-layer neural networks to study the pairwise correlation from multiple views with the guidance of the latent representation. It is also able to help obtain the nonlinear subspaces of the multi-view data. e-ESCC collects multi-view similarity matrices into a tensor and utilizes the low-rank tensor norm to exploit the high-order correlation among different views. The augmented Lagrangian multiplier is adopted to solve the formulated problem of e-ESCC. Experiments on seven data sets validate the superiority of our method over 13 state-of-the-art multi-view clustering methods under six metrics.
Yalan Qin, Nan Pu, Hanzhou Wu
IEEE Trans. Knowl. Data Eng.2
2024 EDMC: Efficient Multi-View Clustering via Cluster and Instance Space Learning
abstract
Multi-view subspace clustering aims to cluster the data lying in a union of subspaces with low dimensions. The commonly used spectral clustering performs the final clustering based on an n×n affinity graph, which suffers from relative high time and space complexity. Some existing works have chosen key anchors with uniform sampling strategy orK-means for dealing with large-scale datasets. However, few of them pay attention to the physical meaning of cluster representation in the column of the dataset for learning informative anchors, which is independent from the instance representation. In this paper, we propose efficient dual multi-view clustering (EDMC) with relative low complexity. To be specific, EDMC makes full use of cluster representation space in the column of the dataset to help produce informative anchors, which has a clear physical meaning and is independent of instance representation in the row. It simultaneously explores the cluster and instance subspace representations to learn anchors for large-scale datasets. We perform anchor learning and efficient multi-view clustering in a unified framework and then adopt an alternative optimization strategy for solving the formulated problem. Extensive experiments performed on different datasets in terms of several metrics validate the superiority of the proposed method.
Yalan Qin, Nan Pu, Hanzhou Wu
IEEE Trans. Multim.2
2023 COCA: COllaborative CAusal Regularization for Audio-Visual Question Answering
abstract
Audio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models are not only prone to over-exploit prevalent language bias, but also suffer from additional joint-modal biases caused by the shortcut relations between textual-auditory/visual co-occurrences and dominated answers. In this paper, we propose a COllabrative CAusal (COCA) Regularization to remedy this more challenging issue of data biases. Specifically, a novel Bias-centered Causal Regularization (BCR) is proposed to alleviate specific shortcut biases by intervening bias-irrelevant causal effects, and further introspect the predictions of AVQA models in counterfactual and factual scenarios. Based on the fact that the dominated bias impairing model robustness for different samples tends to be different, we introduce a Multi-shortcut Collaborative Debiasing (MCD) to measure how each sample suffers from different biases, and dynamically adjust their debiasing concentration to different shortcut correlations. Extensive experiments demonstrate the effectiveness as well as backbone-agnostic ability of our COCA strategy, and it achieves state-of-the-art performance on the large-scale MUSIC-AVQA dataset.
Mingrui Lao, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew
AAAI2
2023 Dynamic Conceptional Contrastive Learning for Generalized Category Discovery
abstract
Generalized category discovery (GCD) is a recently proposed open-world problem, which aims to automatically cluster partially labeled data. The main challenge is that the unlabeled data contain instances that are not only from known categories of the labeled data but also from novel categories. This leads traditional novel category discovery (NCD) methods to be incapacitated for GCD, due to their assumption of unlabeled data are only from novel categories. One effective way for GCD is applying self-supervised learning to learn discriminate representation for unlabeled data. However, this manner largely ignores underlying relationships between instances of the same concepts (e.g., class, super-class, and sub-class), which results in inferior representation learning. In this paper, we propose a Dynamic Conceptional Contrastive Learning (DCCL)framework, which can effectively improve clustering accuracy by alternately estimating underlying visual conceptions and learning conceptional representation. In addition, we design a dynamic conception generation and update mechanism, which is able to ensure consistent conception learning and thus further facilitate the optimization of DCCL. Extensive experiments show that DCCL achieves new state-of-the-art performances on six generic and fine-grained visual recognition datasets, especially on fine-grained ones. For example, our method significantly surpasses the best competitor by 16.2% on the new classes for the CUB-200 dataset. Code is available at https://github.com/TPCD/DCCL
Nan Pu, Zhun Zhong, Nicu Sebe
CVPR1
2023 Image Stitching Based on Multi-Scale Meshes
abstract
Generating high-quality stitching images with a natural structure is a challenging task in computer vision. Recent image stitching methods based on warps failed to suppress the distortion of the images. They often bend the salient lines in the image, which is inconsistent with human perception. In this paper, we succeed in proposing a novelty model called multi-perspective warps for natural image stitching which is related to the density of feature points in the images. With it we can get more precise matching results. Three new energy terms are developed to stitch quality to specify and balance the expected for aligning the vertices of the multi-scale mesh, which can constrain the transformation of the mesh. We also explore and introduce three feature point reconstruction algorithms to enrich the features in the images. Extensive experiments demonstrate that the proposed method outperforms most state-of-the-arts by effectively preserving the linear structure in the image and improving the robustness.
Qi Jia 0001, Nan Pu
ICIP4
2023 Multi-Domain Lifelong Visual Question Answering via Self-Critical Distillation
abstract
Visual Question Answering (VQA) has achieved significant success over the last few years, while most studies focus on training a VQA model on a stationary domain (e.g., a given dataset). In real-world application scenarios, however, these methods are often inefficient because VQA systems are always supposed to extend their knowledge and meet the ever-changing demands of users. In this paper, we introduce a new and challenging multi-domain lifelong VQA task, dubbed MDL-VQA, which encourages the VQA model to continuously learn across multiple domains while mitigating the forgetting on previously-learned domains. Furthermore, we propose a novel replay-free Self-Critical Distillation (SCD) framework tailor-made for MDL-VQA, which alleviates forgetting issue via transferring previous-domain knowledge from teacher to student models. First, we propose to introspect the teacher's understanding over original and counterfactual samples, thereby creating informative instance-relevant and domain-relevant knowledge for logits-based distillation. Second, on the side of feature-based distillation, we propose to introspect the reasoning behavior of student model to establish the harmful domain-specific knowledge acquired in current domain, and further leverage the metric learning strategy to encourage student to learn useful knowledge in new domain. Extensive experiments demonstrate that SCD framework outperforms state-of-the-art competitors with different training orders.
Mingrui Lao, Nan Pu, Yu Liu 0012, Zhun Zhong, Erwin M. Bakker, Nicu Sebe, Michael S. Lew
ACM Multimedia2
2023 FedVQA: Personalized Federated Visual Question Answering over Heterogeneous Scenes
abstract
This paper presents a new setting for visual question answering (VQA) called personalized federated VQA (FedVQA) that addresses the growing need for decentralization and data privacy protection. FedVQA is both practical and challenging, requiring clients to learn well-personalized models on scene-specific datasets with severe feature/label distribution skews. These models then collaborate to optimize a generic global model on a central server, which is desired to generalize well on both seen and unseen scenes without sharing raw data with the server and other clients. The primary challenge of FedVQA is that, client models tend to forget the global knowledge initialized from central server during the personalized training, which impairs their personalized capacity due to the potential overfitting issue on local data. This further leads to divergence issues when aggregating distinct personalized knowledge at the central server, resulting in an inferior generalization ability on unseen scenes. To address the challenge, we propose a novel federated pairwise preference preserving (FedP3) framework to improve personalized learning via preserving generic knowledge under FedVQA constraints. Specifically, we first design a differentiable pairwise preference (DPP) to improve knowledge preserving by formulating a flexible yet effective global knowledge. Then, we introduce a forgotten-knowledge filter (FKF) to encourage the client models to selectively consolidate easily-forgotten knowledge. By aggregating the DPP and the FKF, FedP3 coordinates the generic and the personalized knowledge to enhance the personalized ability of clients and generalizability of the server. Extensive experiments show that FedP3 consistently surpasses the state-of-the-art in FedVQA task.
Mingrui Lao, Nan Pu, Zhun Zhong, Nicu Sebe, Michael S. Lew
ACM Multimedia2
2023 Broaden Your Positives: A General Rectification Approach for Novel Class Discovery
Yaqi Cai, Nan Pu, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012
PRCV (4)2
2023 Dual selective knowledge transfer for few-shot classification
abstract
Abstract Few-shot learning aims at recognizing novel visual categories from very few labelled examples. Different from the existing few-shot classification methods that are mainly based on metric learning or meta-learning, in this work we focus on improving the representation capacity of feature extractors. For this purpose, we propose a new two-stage dual selective knowledge transfer (DSKT) framework, to guide models towards better optimization. Specifically, we first exploit an improved multi-task learning approach to train a feature extractor with robust representation capability as a teacher model. Then, we design an effective dual selective knowledge distillation method, which enables the student model to selectively learn knowledge from the teacher model and current samples, thereby improving the student model’s ability to generalize on unseen classes. Extensive experimental results show that our DSKT achieves competitive performances on four well-known few-shot classification benchmarks.
Nan Pu, Mingrui Lao, Erwin M. Bakker, Michael S. Lew
Appl. Intell.2
2023 A Memorizing and Generalizing Framework for Lifelong Person Re-Identification
abstract
In this paper, we introduce a challenging yet practical setting for person re-identification (ReID) task, named lifelong person re-identification (LReID), which aims to continuously train a ReID model across multiple domains and the trained model is required to generalize well on both seen and unseen domains. It is therefore critical to learn a ReID model that can learn a generalized representation without forgetting knowledge of seen domains. In this paper, we propose a new MEmorizing and GEneralizing framework (MEGE) for LReID, which can jointly prevent the model from forgetting and improve its generalization ability. Specifically, our MEGE is composed of two novel modules, i.e., Adaptive Knowledge Accumulation (AKA) and differentiable Ranking Consistency Distillation (RCD). Taking inspiration from the cognitive processes in the human brain, we endow AKA with two special capacities, knowledge representation and knowledge operation by graph convolution networks. AKA can effectively mitigate catastrophic forgetting on seen domains while improving the generalization ability to unseen domains. By considering the ranking factor that is specifically important in ReID, RCD is designed to distill the ranking knowledge in a differentiable manner, which can further prevent the catastrophic forgetting. To supporting the study of LReID, we build a new and large-scale benchmark with two practical evaluation protocols that consider the metrics of non-forgetting and generalization. Experiments demonstrate that 1) our MEGE framework can effectively improve the performance on seen and unseen domains under the domain-incremental learning constraint, and that 2) the proposed MEGE outperforms state-of-the-art competitors by large margins.
Nan Pu, Zhun Zhong, Nicu Sebe, Michael S. Lew
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Lifelong Fine-Grained Image Retrieval
abstract
Fine-grained image retrieval has been extensively explored in a zero-shot manner. A deep model is trained on the seen part and then evaluated the generalization performance on the unseen part. However, this setting is infeasible for many real-world applications since (1) the retrieval dataset can be non-fixed so that new data are added constantly, and (2) data samples of the seen categories are also common in practice and are important for evaluation. In this paper, we explore lifelong fine-grained image retrieval (LFGIR), which learns continuously on a sequence of new tasks with data from different datasets. We first use knowledge distillation to minimize catastrophic forgetting on old tasks. Training continuously on different datasets causes large domain shifts between the old and new tasks while image retrieval is sensitive to even small shifts in the features. This tends to weaken the effectiveness of knowledge distillation by the frozen teacher. To mitigate the impact of domain shifts, we use the network inversion method to generate images of the old tasks. In addition, we design an on-the-fly teacher which transfers knowledge captured on a new task to the student to improve better generalization performance, thereby achieving a better balance between old and new tasks in the end. We name the whole framework as Dual Knowledge Distillation (DKD), whose efficacy is demonstrated by extensive experimental results on sequential tasks including seven datasets.
Wei Chen 0072, Haoyang Xu, Nan Pu, Yu Liu 0012, Mingrui Lao, Weiping Wang 0002, Li Liu 0002, Michael S. Lew
IEEE Trans. Multim.3
2022 VQA-BC: Robust Visual Question Answering Via Bidirectional Chaining
abstract
Current VQA models are suffering from the problem of overdependence on language bias, which severely reduces their robustness in real-world scenarios. In this paper, we analyze VQA models from the view of forward/backward chaining in the inference engine, and propose to enhance their robustness via a novel Bidirectional Chaining (VQA-BC) framework. Specifically, we introduce a backward chaining with hardnegative contrastive learning to reason from the consequence (answers) to generate crucial known facts (question-related visual region features). Furthermore, to alleviate the overconfident problem in answer prediction (forward chaining), we present a novel introspective regularization to connect forward and backward chaining with label smoothing. Extensive experiments verify that VQA-BC not only effectively overcomes language bias on out-of-distribution dataset, but also alleviates the over-correct problem caused by ensemble-based method on in-distribution dataset. Compared with competitive debiasing strategies, our method achieves state-of-the-art performance to reduce language bias on VQA-CP v2 dataset.
Mingrui Lao, Yanming Guo, Wei Chen 0072, Nan Pu, Michael S. Lew
ICASSP4
2022 Meta Reconciliation Normalization for Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) is a challenging and emerging task, which concerns the ReID capability on both seen and unseen domains after learning across different domains continually. Existing works on LReID are devoted to introducing commonly-used lifelong learning approaches, while neglecting a serious side effect caused by using normalization layers in the context of domain-incremental learning. In this work, we aim to raise awareness of the importance of training proper batch normalization layers by proposing a new meta reconciliation normalization (MRN) method specifically designed for tackling LReID. Our MRN consists of grouped mixture standardization and additive rectified rescaling components, which are able to automatically maintain an optimal balance between domain-dependent and domain-independent statistics, and even adapt MRN for different testing instances. Furthermore, inspired by synaptic plasticity in human brain, we present a MRN-based meta-learning framework for mining the meta-knowledge shared across different domains, even without replaying any previous data, and further improve the model's LReID ability with theoretical analyses. Our method achieves new state-of-the-art performances on both balanced and imbalanced LReID benchmarks.
Nan Pu, Yu Liu 0012, Wei Chen 0072, Erwin M. Bakker, Michael S. Lew
ACM Multimedia1
2022 Feature Estimations Based Correlation Distillation for Incremental Image Retrieval
abstract
Deep learning for fine-grained image retrieval in an incremental context is less investigated. In this paper, we explore this task to realize the model’s continuous retrieval ability. That means, the model enables to perform well on new incoming data and reduce forgetting of the knowledge learned on preceding old tasks. For this purpose, we distill semantic correlations knowledge among the representations extracted from the new data only so as to regularize the parameters updates using the teacher-student framework. In particular, for the case of learning multiple tasks sequentially, aside from the correlations distilled from the penultimate model, we estimate the representations for all prior models and further their semantic correlations by using the representations extracted from the new data. To this end, the estimated correlations are used as an additional regularization and further prevent catastrophic forgetting over all previous tasks, and it is unnecessary to save the stream of models trained on these tasks. Extensive experiments demonstrate that the proposed method performs favorably for retaining performance on the already-trained old tasks and achieving good accuracy on the current task when new data are added at once or sequentially.
Wei Chen 0072, Yu Liu 0012, Nan Pu, Weiping Wang 0002, Li Liu 0002, Michael S. Lew
IEEE Trans. Multim.3
2021 Lifelong Person Re-Identification via Adaptive Knowledge Accumulation
abstract
Person re-identification (ReID) methods always learn through a stationary domain that is fixed by the choice of a given dataset. In many contexts (e.g., lifelong learning), those methods are ineffective because the domain is continually changing in which case incremental learning over multiple domains is required potentially. In this work we explore a new and challenging ReID task, namely lifelong person re-identification (LReID), which enables to learn continuously across multiple domains and even generalise on new and unseen domains. Following the cognitive processes in the human brain, we design an Adaptive Knowledge Accumulation (AKA) framework that is endowed with two crucial abilities: knowledge representation and knowledge operation. Our method alleviates catastrophic forgetting on seen domains and demonstrates the ability to generalize to unseen domains. Correspondingly, we also provide a new and large-scale benchmark for LReID. Extensive experiments demonstrate our method outperforms other competitors by a margin of 5.8% mAP in generalising evaluation. The codes will be available at https://github.com/TPCD/LifelongReID.
Nan Pu, Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew
CVPR1
2021 From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question Answering
abstract
Most Visual Question Answering (VQA) models are faced with language bias when learning to answer a given question, thereby failing to understand multimodal knowledge simultaneously. Based on the fact that VQA samples with different levels of language bias contribute differently for answer prediction, in this paper, we overcome the language prior problem by proposing a novel Language Bias driven Curriculum Learning (LBCL) approach, which employs an easy-to-hard learning strategy with a novel difficulty metric Visual Sensitive Coefficient (VSC). Specifically, in the initial training stage, the VQA model mainly learns the superficial textual correlations between questions and answers (easy concept) from more-biased examples, and then progressively focuses on learning the multimodal reasoning (hard concept) from less-biased examples in the following stages. The curriculum selection of examples on different stages is according to our proposed difficulty metric VSC, which is to evaluate the difficulty driven by the language bias of each VQA sample. Furthermore, to avoid the catastrophic forgetting of the learned concept during the multi-stage learning procedure, we propose to integrate knowledge distillation into the curriculum learning framework. Extensive experiments show that our LBCL can be generally applied to common VQA baseline models, and achieves remarkably better performance on the VQA-CP v1 and v2 datasets, with an overall 20% accuracy boost over baseline models.
Mingrui Lao, Yanming Guo, Yu Liu 0012, Wei Chen 0072, Nan Pu, Michael S. Lew
ACM Multimedia5
2021 Multi-stage hybrid embedding fusion network for visual question answering
Mingrui Lao, Yanming Guo, Nan Pu, Wei Chen 0072, Yu Liu 0012, Michael S. Lew
Neurocomputing3
2020 Comparison of deep learning and hand crafted features for mining simulation data
abstract
Computational Fluid Dynamics (CFD) simulations are a very important tool for many industrial applications, such as aerodynamic optimization of engineering designs like cars shapes, airplanes parts etc. The output of such simulations, in particular the calculated flow fields, are usually very complex and hard to interpret for realistic three-dimensional real-world applications, especially if time-dependent simulations are investigated. Automated data analysis methods are warranted but a non-trivial obstacle is given by the very large dimensionality of the data. A flow field typically consists of six measurement values for each point of the computational grid in 3D space and time (velocity vector values, turbulent kinetic energy, pressure and viscosity). In this paper we address the task of extracting meaningful results in an automated manner from such high dimensional data sets. We propose deep learning methods which are capable of processing such data and which can be trained to solve relevant tasks on simulation data, i.e. predicting drag and lift forces applied on an airfoil. We also propose an adaptation of the classical hand crafted features known from computer vision to address the same problem and compare a large variety of descriptors and detectors. Finally, we compile a large dataset of 2D simulations of the flow field around airfoils which contains 16000 flow fields with which we tested and compared approaches. Our results show that the deep learning-based methods, as well as hand crafted feature based approaches, are well-capable to accurately describe the content of the CFD simulation output on the proposed dataset.
Theodoros Georgiou 0001, Thomas Bäck, Nan Pu, Wei Chen 0072, Michael S. Lew
ICPR4
2020 Dual Gaussian-based Variational Subspace Disentanglement for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging and essential task in night-time intelligent surveillance systems. Except for the intra-modality variance that RGB-RGB person re-identification mainly overcomes, VI-ReID suffers from additional inter-modality variance caused by the inherent heterogeneous gap. To solve the problem, we present a carefully designed dual Gaussian-based variational auto-encoder (DG-VAE), which disentangles an identity-discriminable and an identity-ambiguous cross-modality feature subspace, following a mixture-of-Gaussians (MoG) prior and a standard Gaussian distribution prior, respectively. Disentangling cross-modality identity-discriminable features leads to more robust retrieval for VI-ReID. To achieve efficient optimization like conventional VAE, we theoretically derive two variational inference terms for the MoG prior under the supervised setting, which not only restricts the identity-discriminable subspace so that the model explicitly handles the cross-modality intra-identity variance, but also enables the MoG distribution to avoid posterior collapse. Furthermore, we propose a triplet swap reconstruction (TSR) strategy to promote the above disentangling process. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on two VI-ReID datasets. Codes will be available at https://github.com/TPCD/DG-VAE.
Nan Pu, Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew
ACM Multimedia1
2019 Domain Uncertainty Based On Information Theory for Cross-Modal Hash Retrieval
abstract
Cross-modal hash retrieval has received considerable interest in the area of deep learning. Here hash codes of data of different modalities are learned where pair-wise loss functions control feature similarity in a shared embedding space. In this paper we improve on feature similarity by using Shannon's information entropy with respect to the modality information that is left in learning superior hash codes. We introduce a novel network for predicting the domain from the learned features while the protagonist network uses a loss function based on Shannon's information entropy to learn to maximize the domain uncertainty and therefore the information content. Additionally, according to the number of common labels between each similar image-text pair, we define a multi-level similarity matrix as supervisory information, which constrains all similar pairs with different weights. We show with extensive experiments that our novel approach to domain uncertainty leads to a cross-modal hash retrieval that outperforms the state-of-the-art.
Wei Chen 0072, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew
ICME2
2019 Learning a Domain-Invariant Embedding for Unsupervised Person Re-identification
abstract
Person re-identification (Re-ID) aims at matching images of the same person where images are captured by non-overlapping camera views distributed at different locations. To solve this problem, most recent works require a large pre-labeled dataset for training a deep model. These methods are not always suitable for real-world applications, because the latter often lack labeled data. In order to tackle this drawback, we proposed a novel Domain-Invariant Embedding Network (DIEN) to learn a domain-invariant embedding (DIE) feature by introducing a multi-loss joint learning with Recurrent Top- Down Attention (RTDA) mechanism. Due to the improvement in traditional triplet loss, our proposed model can benefit from both source-domain (labeled) data and target-domain (unlabeled) data. Furthermore, the resulting DIE feature not only has improved class discrimination but also robustness to domain shift. We compared our method with recent competitive algorithms and also evaluated the effectiveness of the proposed modules.
Nan Pu, Theodoros Georgiou 0001, Erwin M. Bakker, Michael S. Lew
IJCNN1