VLDB 2026 Research / reviewers in the wild / expert
Zhen Fang 0001
dblp:62/3571-1
· DBLP profile ↗
62ranked-venue papers
8as first author
60since 2021 · last 2026
0000-0003-0602-6255ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 7 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language ModelsabstractAiming at identifying unexpected inputs from unknown classes, out-of-distribution (OOD) detection has emerged as a pivotal approach to enhancing the reliability of machine learning models. This paper focuses on the burgeoning paradigm of post-hoc OOD detection with pre-trained vision-language models (VLMs), where a popular pipeline is to detect OOD inputs by examining their affinities between ID labels and negative labels, i.e., those semantically different from ID labels. Due to the unavailability of target OOD labels, existing works predominantly rely on heuristic rules to mine negative labels from unlabeled wild corpus data. Despite the empirical success, we argue that the power of VLM-based OOD detection has yet to be fully unleashed since the notorious false negative problem is far from addressed in the literature. With this motivation, we are interested in addressing the challenge of mining true negative labels for OOD scoring. To this end, we develop a theoretical framework for correcting the sampling bias of negatives labels by indirectly approximating the distribution of negative labels. Perhaps surprisingly, we show that the debiased negative mining can be naturally converted into Monte-Carlo sampling based on ID labels and the unlabeled wild corpus data. Extensive experiments empirically manifest that our method establishes a new state-of-the-art in a variety of OOD detection setups. Jie Lu 0001, Guangquan Zhang 0001, Zhen Fang 0001 |
KDD (1) | 4 |
| 2026 | Toward Trustworthy Vision-language Models in the Wild: Theory, Algorithm and ApplicationabstractVision-Language Models (VLMs) have revolutionized multimedia applications by enabling open-vocabulary search and generation. While significant strides have been made in developing powerful VLMs, the field remains in the early stages of rigorously evaluating and understanding their real-world vulnerabilities, both empirically and theoretically. In response to this important but under-explored gap, the workshop aims to provide a forum for researchers to share advanced in theories, algorithms and applications of trustworthy multi-modal learning, fostering diverse viewpoints on core principles and emerging techniques for developing trustworthy VLMs in the wild. The workshop provides a platform for researchers to showcase their work, exchange ideas, and foster potential collaborations. Additionally, it serves as a valuable opportunity for practitioners to stay abreast of the latest developments in developing trustworthy vision-language models. Xuefeng Du, Zhen Fang 0001 |
ICMR | 3 |
| 2026 | Partial Label Learning-Inspired Denoising Implicit Feedback for Recommendation
Huilin Chen 0001, Jie Lu 0001, Kezhi Lu, Zhen Fang 0001, Guangquan Zhang 0001 |
SIGIR | 4 |
| 2026 | Cross-domain Few-shot Classification via Invariant-content Feature ReconstructionabstractAbstract In cross-domain few-shot classification (CFC), mainstream studies aim to train a simple module (e.g. a linear transformation head) to select or transform features (a.k.a., the high-level semantic features) for previously unseen domains with a few labeled training data available on top of a powerful pre-trained model. These studies usually assume that high-level semantic features are shared across these domains, and just simple feature selection or transformations are enough to adapt features to previously unseen domains. However, in this paper, we find that the simply transformed features are too general to fully cover the key content features regarding each class. Thus, we propose an effective method, invariant-content feature reconstruction (IFR), to train a simple module that simultaneously considers both high-level and fine-grained invariant-content features for the previously unseen domains. Specifically, the fine-grained invariant-content features are considered as a set of informative and discriminative features learned from a few labeled training data of tasks sampled from unseen domains and are extracted by retrieving features that are invariant to style modifications from a set of content-preserving augmented data in pixel level with an attention module. Extensive experiments on the Meta-Dataset benchmark show that IFR achieves good generalization performance on unseen domains, which demonstrates the effectiveness of the fusion of the high-level features and the fine-grained invariant-content features. Specifically, IFR improves the average accuracy on unseen domains by 1.6% and 6.5% respectively under two different cross-domain few-shot classification settings. Hongduan Tian, Feng Liu 0003, Ka Chun Cheung, Zhen Fang 0001, Simon See, Tongliang Liu, Bo Han 0003 |
Int. J. Comput. Vis. | 4 |
| 2026 | Pseudo-Label Refinement for Multimodal Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to improve a model’s performance on an unlabeled target domain by leveraging labeled data from a source domain. Traditional UDA methods often overlook the rich textual information inherent in class labels, limiting their effectiveness. Recent advances in visual language models (VLMs), especially the contrastive language–image pretraining (CLIP) model, provide a promising multimodal foundation by integrating both image and text representations. However, most CLIP-based UDA approaches rely on pseudo-labels generated in a zero-shot manner by the pretrained CLIP model, which may become suboptimal as training progresses. To address this limitation, we propose a pseudo-label refinement (PuRe) framework that begins training with zero-shot CLIP pseudo-labels and then transitions to self-training, where we iteratively refine the model’s own predictions as pseudo-labels. In this iterative process, we introduce label consistency learning to stabilize predictions under strong data augmentations, increasing robustness, and an information maximization (IM) loss to encourage high-confidence predictions while preserving prediction diversity. Together, these components progressively enhance pseudo-label reliability, leading to improved adaptation in the target domain. Furthermore, we propose a novel prompt learning (PL) method that fine-tunes only a minimal set of parameters, specifically a single context token and class embeddings. Extensive experiments demonstrate that PuRe surpasses existing UDA methods on multiple benchmarks. Kuo Shi, Jie Lu 0001, Zhen Fang 0001, Guangquan Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | NLPrompt: Noise-Label Prompt Learning for Vision-Language ModelsabstractThe emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite its promise, real-world datasets often contain noisy labels that can degrade prompt learning performance. In this paper, we demonstrate that using mean absolute error (MAE) loss in prompt learning, named PromptMAE, significantly enhances robustness against noisy labels while maintaining high accuracy. Though MAE is straightforward and recognized for its robustness, it is rarely used in noisy-label learning due to its slow convergence and poor performance outside prompt learning scenarios. To elucidate the robustness of PromptMAE, we leverage feature learning theory to show that MAE can suppress the influence of noisy samples, thereby improving the signal-to-noise ratio and enhancing overall robustness. Additionally, we introduce PromptOT, a prompt-based optimal transport data purification method to enhance the robustness further. PromptOT employs text features in vision-language models as prototypes to construct an optimal transportation matrix. This matrix effectively partitions datasets into clean and noisy subsets, allowing for the application of cross-entropy loss to the clean subset and MAE loss to the noisy subset. Our Noise-Label Prompt Learning method, named NLPrompt, offers a simple and efficient approach that leverages the expressive representations and precise alignment capabilities of vision-language models for robust prompt learning. We validate NLPrompt through extensive experiments across various noise settings, demonstrating significant performance improvements. Bikang Pan, Xiaoying Tang 0002, Wei Huang 0034, Zhen Fang 0001, Feng Liu 0003, Jingya Wang 0001, Jingyi Yu 0001, Ye Shi 0001 |
CVPR | 5 |
| 2025 | On the Provable Importance of Gradients for Autonomous Language-Assisted Image Clustering
Jie Lu 0001, Guangquan Zhang 0001, Zhen Fang 0001 |
ICCV | 4 |
| 2025 | Characterizing Submanifold Region for Out-of-Distribution Detection: (Extended Abstract)abstractDetecting out-of-distribution (OOD) samples poses a significant safety challenge when deploying models in open-world scenarios. Advanced works assume that OOD and in-distributional (ID) samples exhibit a distribution discrepancy, showing an encouraging direction in estimating the uncertainty with embedding features or predicting outputs. In this work, we propose a data structure-aware approach to mitigate the sensitivity of distances to the “curse of dimensionality”, where high-dimensional features are mapped to the manifold of ID samples, leveraging the well-known manifold assumption. Specifically, we present a novel distance termed as tangent distance, which tackles the issue of generalizing the meaningfulness of distances on testing samples to detect OOD inputs. Extensive experiments show that the tangent distance performs competitively with other post hoc OOD detection baselines on common and large-scale benchmarks. Zhen Fang 0001, Yonggang Zhang 0003, Jiajun Bu, Bo Han 0003, Haishuai Wang |
ICDE | 2 |
| 2025 | Deep Kernel Relative Test for Machine-generated Text DetectionabstractRecent studies demonstrate that two-sample test can effectively detect machine-generated texts (MGTs) with excellent adaptation ability to texts generated by newer LLMs. However, two-sample test-based detection relies on the assumption that human-written texts (HWTs) must follow the distribution of seen HWTs. As a result, it tends to make mistakes in identifying HWTs that deviate from the seen HWT distribution, limiting their use in sensitive areas like academic integrity verification. To address this issue, we propose to employ non-parametric kernel relative test to detect MGTs by testing whether it is statistically significant that the distribution of a text to be tested is closer to the distribution of HWTs than to the MGTs' distribution. We further develop a kernel optimisation algorithm in relative test to select the best kernel that can enhance the testing capability for MGT detection. As relative test does not assume that a text to be tested must belong exclusively to either MGTs or HWTs, relative test can largely reduce the false positive error compared to two-sample test, offering significant advantages in practice. Extensive experiments demonstrate the superior performance of our method, compared to state-of-the-art non-parametric and parametric detectors. The code and demo are available: https://github.com/xLearn-AU/R-Detect. Yiliao Song, Zhenqiao Yuan, Shuhai Zhang, Zhen Fang 0001, Feng Liu 0003 |
ICLR | 4 |
| 2025 | Release the Powers of Prompt Tuning: Cross-Modality Prompt TransferabstractPrompt Tuning adapts frozen models to new tasks by prepending a few learnable embeddings to the input.
However, it struggles with tasks that suffer from data scarcity.
To address this, we explore Cross-Modality Prompt Transfer, leveraging prompts pretrained on a data-rich modality to improve performance on data-scarce tasks in another modality.
As a pioneering study, we first verify the feasibility of cross-modality prompt transfer by directly applying frozen source prompts (trained on the source modality) to the target modality task.
To empirically study cross-modality prompt transferability, we train a linear layer to adapt source prompts to the target modality, thereby boosting performance and providing ground-truth transfer results.
Regarding estimating prompt transferability, existing methods show ineffectiveness in cross-modality scenarios where the gap between source and target tasks is larger.
We address this by decomposing the gap into the modality gap and the task gap, which we measure separately to autonomously select the best source prompt for a target task.
Additionally, we propose Attention Transfer to further reduce the gaps by injecting target knowledge into the prompt and reorganizing a top-transferable source prompt using an attention block.
We conduct extensive experiments involving prompt transfer from 13 source language tasks to 19 target vision tasks under three settings.
Our findings demonstrate that:
(i) cross-modality prompt transfer is feasible, supported by in-depth analysis;
(ii) measuring both the modality and task gaps is crucial for accurate prompt transferability estimation, a factor overlooked by previous studies;
(iii) cross-modality prompt transfer can significantly release the powers of prompt tuning on data-scarce tasks, as evidenced by comparisons with a newly released prompt-based benchmark. Ningyuan Zhang, Jie Lu 0001, Keqiuyin Li, Zhen Fang 0001, Guangquan Zhang 0001 |
ICLR | 4 |
| 2025 | Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic ApproachabstractMultimodal large language models (MLLMs) have shown promising capabilities but struggle under distribution shifts, where evaluation data differ from instruction tuning distributions. Although previous works have provided empirical evaluations, we argue that establishing a formal framework that can characterize and quantify the risk of MLLMs is necessary to ensure the safe and reliable application of MLLMs in the real world. By taking an information-theoretic perspective, we propose the first theoretical framework that enables the quantification of the maximum risk of MLLMs under distribution shifts. Central to our framework is the introduction of Effective Mutual Information (EMI), a principled metric that quantifies the relevance between input queries and model responses. We derive an upper bound for the EMI difference between in-distribution (ID) and out-of-distribution (OOD) data, connecting it to visual and textual distributional discrepancies. Extensive experiments on real benchmark datasets, spanning 61 shift scenarios, empirically validate our theoretical insights. Changdae Oh, Zhen Fang 0001, Shawn Im, Xuefeng Du, Yixuan Li 0001 |
ICML | 2 |
| 2025 | Distributional Prototype Learning for Out-of-distribution DetectionabstractOut-of-distribution (OOD) detection has emerged as a pivotal approach for enhancing the reliability of machine learning models, considering the potential for test data to be sampled from classes disparate from in-distribution (ID) data employed during model training. Detecting those OOD data is typically realized as a distance measurement problem, where those deviating far away from the training distribution in the learned feature space are considered OOD samples. Advanced works have shown great success in learning with prototypes for feature-based OOD detection methods, where each ID class is represented with single or multiple prototypes. However, modeling with a finite number of prototypes would fail to maximally capture intra-class variations. In view of this, this paper extends the existing prototype-based learning paradigm to an infinite setting. This motivates us to design two feasible formulations for the Distributional Prototype Learning (DPL) objective, where, to avoid intractable computation and exploding parameters caused by the infinity nature, our key idea is to model an infinite number of discrete prototypes of each ID class with a class-wise continuous distribution. We theoretically analyze both alternatives, identifying the more stable-converging version of the learning objective. We show that, by sampling prototypes from a mixture of class-conditioned Gaussian distributions, the objective can be efficiently computed in a closed form without resorting to the computationally expensive Monte-Carlo approximation of the involved expectation terms. Extensive evaluations across mainstream OOD detection benchmarks empirically manifest that our proposed DPL has established a new state-of-the-art in various OOD settings. Jie Lu 0001, Yonggang Zhang 0003, Guangquan Zhang 0001, Zhen Fang 0001 |
KDD (1) | 5 |
| 2025 | MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image DetectionabstractRecent advances in generative models have highlighted the need for robust detectors capable of distinguishing real images from AI-generated images. While existing methods perform well on known generators, their performance often declines when tested with newly emerging or unseen generative models due to overlapping feature embeddings that hinder accurate cross-generator classification. In this paper, we propose Multimodal Discriminative Representation Learning for Generalizable AI-generated Image Detection (MiraGe), a method designed to learn generator-invariant features. Motivated by theoretical insights on intra-class variation minimization and inter-class separation, MiraGe tightly aligns features within the same class while maximizing separation between classes, enhancing feature discriminability. Moreover, we apply multimodal prompt learning to further refine these principles into CLIP, leveraging text embeddings as semantic anchors for effective discriminative representation learning, thereby improving generalizability. Comprehensive experiments across multiple benchmarks show that MiraGe achieves state-of-the-art performance, maintaining robustness even against unseen generators like Sora. Kuo Shi, Jie Lu 0001, Shanshan Ye, Guangquan Zhang 0001, Zhen Fang 0001 |
ACM Multimedia | 5 |
| 2025 | An Information-theoretical Framework for Understanding Out-of-distribution Detection with Pretrained Vision-Language ModelsabstractOut-of-distribution (OOD) detection, recognized for its ability to identify samples of unknown classes, provides solid advantages in ensuring the reliability of machine learning models.
Among existing OOD detection methods, pre-trained vision-language models have emerged as powerful post-hoc OOD detectors by leveraging textual and visual information.
Despite the empirical success, there still remains a lack of research on a formal understanding of their effectiveness.
This paper bridges the gap by theoretically demonstrating that existing CLIP-based post-hoc methods effectively perform a stochastic estimation of the point-wise mutual information (PMI) between the input image and each in-distribution label. This estimation is then utilized to construct energy functions for modeling in-distribution distributions.
Different from prior methods that inherently consider PMI estimation as a whole task, we, motivated by the divide-and-conquer philosophy, decompose PMI estimation into multiple easier sub-tasks by applying the chain rule of PMI, which not only reduces the estimation complexity but also provably increases the estimation upper bound to reduce the underestimation bias.
Extensive evaluations across mainstream benchmarks empirically manifest that our method establishes a new state-of-the-art in a variety of OOD detection setups. Jie Lu 0001, Guangquan Zhang 0001, Zhen Fang 0001 |
NeurIPS | 4 |
| 2025 | Learning Robust Spectral Dynamics for Temporal Domain GeneralizationabstractModern machine learning models struggle to maintain performance in dynamic environments where temporal distribution shifts, \textit{i.e., concept drift}, are prevalent. Temporal Domain Generalization (TDG) seeks to enable model generalization across evolving domains, yet existing approaches typically assume smooth incremental changes, struggling with complex real-world drifts involving both long-term structure (incremental evolution/periodicity) and local uncertainties. To overcome these limitations, we introduce FreKoo, which tackles these challenges through a novel frequency-domain analysis of parameter trajectories. It leverages the Fourier transform to disentangle parameter evolution into distinct spectral bands. Specifically, the low-frequency components with dominant dynamics are learned and extrapolated using the Koopman operator, robustly capturing diverse drift patterns including both incremental and periodic drifts. Simultaneously, potentially disruptive high-frequency variations are smoothed via targeted temporal regularization, preventing overfitting to transient noise and domain uncertainties. In addition, this dual-spectral strategy is rigorously grounded through theoretical analysis, providing stability guarantees for the Koopman prediction, a principled Bayesian justification for the high-frequency regularization, and culminating in a multiscale generalization bound connecting spectral dynamics to improved generalization. Extensive experiments demonstrate FreKoo's significant superiority over state-of-the-art TDG methods, particularly excelling in real-world streaming scenarios with complex drifts and uncertainties. En Yu, Jie Lu 0001, Guangquan Zhang 0001, Zhen Fang 0001 |
NeurIPS | 5 |
| 2025 | Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied AgentsabstractPre-training vision-language representations on human action videos has emerged
as a promising approach to reduce reliance on large-scale expert demonstrations
for training embodied agents. However, prior methods often employ time con-
trastive learning based on goal-reaching heuristics, progressively aligning language
instructions from the initial to the final frame. This overemphasis on future frames
can result in erroneous vision-language associations, as actions may terminate
early or include irrelevant moments in the end. To address this issue, we propose
Action Temporal Coherence Learning (AcTOL) to learn ordered and continuous
vision-language representations without rigid goal-based constraint. AcTOL treats
a video as a continuous trajectory where it (1) contrasts semantic differences be-
tween frames to reflect their natural ordering, and (2) imposes a local Brownian
bridge constraint to ensure smooth transitions across intermediate frames. Exten-
sive imitation learning experiments on both simulated and real robots show that the
pretrained features significantly enhance downstream manipulation tasks with high
robustness to different linguistic styles of instructions, offering a viable pathway
toward generalized embodied agents. Our project page is at https://actol-pretrain.github.io/. Zhizhen Zhang, Lei Zhu 0002, Zhen Fang 0001, Zi Huang, Yadan Luo |
NeurIPS | 3 |
| 2025 | MetaGeno: a chromosome-wise multi-task genomic framework for ischaemic stroke risk predictionabstractCurrent genome-wide association studies provide valuable insights into the genetic basis of ischaemic stroke (IS) risk. However, polygenic risk scores, the most widely used method for genetic risk prediction, have notable limitations due to their linear nature and inability to capture complex, nonlinear interactions among genetic variants. While deep neural networks offer advantages in modeling these complex relationships, the multifactorial nature of IS and the influence of modifiable risk factors present additional challenges for genetic risk prediction. To address these challenges, we propose a Chromosome-wise Multi-task Genomic (MetaGeno) framework that utilizes genetic data from IS and five related diseases. The framework includes a chromosome-based embedding layer to model local and global interactions among adjacent variants, enabling a biologically informed approach. Incorporating multi-disease learning further enhances predictive accuracy by leveraging shared genetic information. Among various sequential models tested, the Transformer demonstrated superior performance, and outperformed other machine learning models and PRS baselines, achieving an AUROC of 0.809 on the UK Biobank dataset. Risk stratification identified a two-fold increased stroke risk (HR, 2.14; 95% CI: 1.81-2.46) in the top 1% risk group, with a nearly five-fold increase in those with modifiable risk factors such as atrial fibrillation and hypertension. Finally, the model was validated on the diverse All of Us dataset (AUROC = 0.764), highlighting ancestry and population differences while demonstrating effective generalization. This study introduces a predictive framework that identifies high-risk individuals and informs targeted prevention strategies, offering potential as a clinical decision-support tool. Yue Yang 0042, Kairui Guo, Yonggang Zhang 0003, Zhen Fang 0001, Mark Grosser, Deon Venter, Weihai Lu, Mengjia Wu, Dennis Cordato, Guangquan Zhang 0001, Jie Lu 0001 |
Briefings Bioinform. | 4 |
| 2025 | Out-of-Distribution Detection with Virtual Outlier SmoothingabstractAbstract Detecting out-of-distribution (OOD) inputs plays a crucial role in guaranteeing the reliability of deep neural networks (DNNs) when deployed in real-world scenarios. However, DNNs typically exhibit overconfidence in OOD samples, which is attributed to the similarity in patterns between OOD and in-distribution (ID) samples. To mitigate this overconfidence, advanced approaches suggest the incorporation of auxiliary OOD samples during model training, where the outliers are assigned with an equal likelihood of belonging to any category. However, identifying outliers that share patterns with ID samples poses a significant challenge. To address the challenge, we propose a novel method, V irtual O utlier S m o othing (VOSo), which constructs auxiliary outliers using ID samples, thereby eliminating the need to search for OOD samples. Specifically, VOSo creates these virtual outliers by perturbing the semantic regions of ID samples and infusing patterns from other ID samples. For instance, a virtual outlier might consist of a cat’s face with a dog’s nose, where the cat’s face serves as the semantic feature for model prediction. Meanwhile, VOSo adjusts the labels of virtual OOD samples based on the extent of semantic region perturbation, aligning with the notion that virtual outliers may contain ID patterns. Extensive experiments are conducted on diverse OOD detection benchmarks, demonstrating the effectiveness of the proposed VOSo. Our code will be available at https://github.com/junz-debug/VOSo . Jun Nie, Yadan Luo, Shanshan Ye, Yonggang Zhang 0003, Xinmei Tian 0001, Zhen Fang 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | Out-of-distribution detection with non-semantic explorationabstractOut-of-distribution (OOD) detection is crucial in modern deep learning applications, as it can identify OOD data drawn from distributions differing from those of the in-distribution (ID) data. Advanced OOD detection methods primarily rely on post-hoc strategies, which identify OOD data by analyzing the predictions of a model well-trained on ID data. However, deep models are known to be impacted by spurious features such as backgrounds, causing existing OOD detection methods to fail in identifying OOD data that share the same spurious features as ID data. Therefore, this paper studies how to mitigate spurious features to improve OOD detection. To address this challenge, we propose a novel method called N on- s emantic E xploration OOD D etection (NsED), which focuses on exploring and exploiting non-semantic features. In particular, NsED first explores non-semantic features in an OOD generalization manner. These non-semantic features are then used to train deep models to be more robust against spurious features. Through extensive experiments on representative benchmarks, we show that NsED significantly and consistently improves the detection performance of many representative post-hoc OOD detection methods. • NsED is proposed to address spurious correlation in OOD detection. • This paper first explores the link between OOD generalization and detection. • Experiments show NsED is robust and its components improve OOD detection. Zhen Fang 0001, Jie Lu 0001, Guangquan Zhang 0001 |
Inf. Sci. | 1 |
| 2025 | SENA: Leveraging set-level consistency adversarial learning for robust pre-trained language model adaptation
Jianqi Gao 0001, Jian Cao 0001, Hang Yu 0006, Yonggang Zhang 0003, Zhen Fang 0001 |
Knowl. Based Syst. | 5 |
| 2025 | Integrated Image-Text Augmentation for Few-Shot Learning in Vision-Language ModelsabstractVision-language models, such as the Contrastive Language-Image Pre-Training (CLIP) model, have achieved significant success in image classification tasks. CLIP demonstrates high expressive power in few-shot learning scenarios due to its pairing of text and image encoders. However, CLIP still faces over-fitting when trained with a limited number of samples. To mitigate this, image augmentation techniques have been proposed in few-shot learning tasks to prevent over-fitting by enriching the dataset. Existing image augmentation methods, primarily designed for single-modal image models, focus solely on transformations within the image itself. However, for CLIP, merely increasing visual variety without considering textual content can reduce generalization ability and may even mislead the model. To address this issue, we introduce a novel image augmentation approach—Integrated Image-Text Augmentation (ITA)— for CLIP model in few-shot learning tasks. This method generates new and diverse augmented images to increase the diversity of the training data and reduce over-fitting. Additionally, ITA establishes an alignment between the augmented images and their textual descriptions. Through this alignment, the model not only learns to recognize visual elements in the images but also understands the semantic connections between these elements and the text descriptions. This dual-modal approach enhances the model’s flexibility and accuracy in processing few-shot learning tasks. Extensive experiments in few-shot image classification scenarios have demonstrated that ITA shows significant improvements compared to various image augmentation techniques. Ran Wang 0016, Hua Zuo, Zhen Fang 0001, Jie Lu 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | Characterizing Submanifold Region for Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) samples poses a significant safety challenge when deploying models in open-world scenarios. Advanced works assume that OOD and in-distributional (ID) samples exhibit a distribution discrepancy, showing an encouraging direction in estimating the uncertainty with embedding features or predicting outputs. Besides incorporating auxiliary outlier as decision boundary, quantifying a “meaningful distance” in embedding space as uncertainty measurement is a promising strategy. However, these distances-based approaches overlook the data structure and heavily rely on the high-dimension features learned by deep neural networks, causing unreliable distances due to the “curse of dimensionality”. In this work, we propose a data structure-aware approach to mitigate the sensitivity of distances to the “curse of dimensionality”, where high-dimensional features are mapped to the manifold of ID samples, leveraging the well-known manifold assumption. Specifically, we present a novel distance termed astangent distance, which tackles the issue of generalizing the meaningfulness of distances on testing samples to detect OOD inputs. Inspired by manifold learning for adversarial examples, where adversarial region probability density is close to the orthogonal direction of the manifold, and both OOD and adversarial samples have common characteristic$-$imperceptible perturbations with shift distribution, we propose that OOD samples are relatively far away from the ID manifold, wheretangent distancedirectly computes the Euclidean distance between samples and the nearest submanifold space$-$instantiated as the linear approximation of local region on the manifold. We provide empirical and theoretical insights to demonstrate the effectiveness of OOD uncertainty measurements on the low-dimensional subspace. Extensive experiments show that thetangent distanceperforms competitively with other post hoc OOD detection baselines on common and large-scale benchmarks, and the theoretical analysis supports our claim that ID samples are likely to reside in high-density regions, explaining the effectiveness of internal connections among ID data. Zhen Fang 0001, Yonggang Zhang 0003, Jiajun Bu, Bo Han 0003, Haishuai Wang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Multiview Classification Through Learning From Interval-Valued DataabstractThe classification problem concerning crisp-valued data has been well resolved. However, interval-valued data, where all of the observations' features are described by intervals, are also a common data type in real-world scenarios. For example, the data extracted by many measuring devices are not exact numbers but intervals. In this article, we focus on a highly challenging problem called learning from interval-valued data (LIND), where we aim to learn a classifier with high performance on interval-valued observations. First, we obtain the estimation error bound of the LIND problem based on the Rademacher complexity. Then, we give the theoretical analysis to show the strengths of multiview learning on classification problems, which inspires us to construct a new algorithm called multiview interval information extraction (Mv-IIE) approach for improving classification accuracy on interval-valued data. The experiment comparisons with several baselines on both synthetic and real-world datasets illustrate the superiority of the proposed framework in handling interval-valued data. Moreover, we describe an application of Mv-IIE that we can prevent data privacy leakage by transforming crisp-valued (raw) data into interval-valued data. Guangzhi Ma, Jie Lu 0001, Zhen Fang 0001, Feng Liu 0003, Guangquan Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | How Does Unlabeled Data Provably Help Out-of-Distribution Detection?abstractUsing unlabeled data to regularize the machine learning models has demonstrated promise for improving safety and reliability in detecting out-of-distribution (OOD) data. Harnessing the power of unlabeled in-the-wild data is non-trivial due to the heterogeneity of both in-distribution (ID) and OOD data. This lack of a clean set of OOD samples poses significant challenges in learning an optimal OOD classifier. Currently, there is a lack of research on formally understanding how unlabeled data helps OOD detection. This paper bridges the gap by introducing a new learning framework SAL (Separate And Learn) that offers both strong theoretical guarantees and empirical effectiveness. The framework separates candidate outliers from the unlabeled data and then trains an OOD classifier using the candidate outliers and the labeled ID data. Theoretically, we provide rigorous error bounds from the lens of separability and learnability, formally justifying the two components in our algorithm. Our theory shows that SAL can separate the candidate outliers with small error rates, which leads to a generalization guarantee for the learned OOD classifier. Empirically, SAL achieves state-of-the-art performance on common benchmarks, reinforcing our theoretical insights. Code is publicly available at https://github.com/deeplearning-wisc/sal. Xuefeng Du, Zhen Fang 0001, Ilias Diakonikolas, Yixuan Li 0001 |
ICLR | 2 |
| 2024 | Negative Label Guided OOD Detection with Pretrained Vision-Language ModelsabstractOut-of-distribution (OOD) detection aims at identifying samples from unknown classes, playing a crucial role in trustworthy models against errors on unexpected inputs.
Extensive research has been dedicated to exploring OOD detection in the vision modality.
{Vision-language models (VLMs) can leverage both textual and visual information for various multi-modal applications, whereas few OOD detection methods take into account information from the text modality.
In this paper, we propose a novel post hoc OOD detection method, called NegLabel, which takes a vast number of negative labels from extensive corpus databases. We design a novel scheme for the OOD score collaborated with negative labels.
Theoretical analysis helps to understand the mechanism of negative labels. Extensive experiments demonstrate that our method NegLabel achieves state-of-the-art performance on various OOD detection benchmarks and generalizes well on multiple VLM architectures. Furthermore, our method NegLabel exhibits remarkable robustness against diverse domain shifts. The codes are available at https://github.com/tmlr-group/NegLabel. Feng Liu 0003, Zhen Fang 0001, Hong Chen 0004, Tongliang Liu, Feng Zheng 0001, Bo Han 0003 |
ICLR | 3 |
| 2024 | Out-of-Distribution Detection with Negative PromptsabstractOut-of-distribution (OOD) detection is indispensable for open-world machine learning models. Inspired by recent success in large pre-trained language-vision models, e.g., CLIP, advanced works have achieved impressive OOD detection results by matching the *similarity* between image features and features of learned prompts, i.e., positive prompts. However, existing works typically struggle with OOD samples having similar features with those of known classes. One straightforward approach is to introduce negative prompts to achieve a *dissimilarity* matching, which further assesses the anomaly level of image features by introducing the absence of specific features. Unfortunately, our experimental observations show that either employing a prompt like "not a photo of a" or learning a prompt to represent "not containing" fails to capture the dissimilarity for identifying OOD samples. The failure may be contributed to the diversity of negative features, i.e., tons of features could indicate features not belonging to a known class. To this end, we propose to learn a set of negative prompts for each class. The learned positive prompt (for all classes) and negative prompts (for each class) are leveraged to measure the similarity and dissimilarity in the feature space simultaneously, enabling more accurate detection of OOD samples. Extensive experiments are conducted on diverse OOD detection benchmarks, showing the effectiveness of our proposed method. Jun Nie, Yonggang Zhang 0003, Zhen Fang 0001, Tongliang Liu, Bo Han 0003, Xinmei Tian 0001 |
ICLR | 3 |
| 2024 | ConjNorm: Tractable Density Estimation for Out-of-Distribution DetectionabstractPost-hoc out-of-distribution (OOD) detection has garnered intensive attention in reliable machine learning. Many efforts have been dedicated to deriving score functions based on logits, distances, or rigorous data distribution assumptions to identify low-scoring OOD samples. Nevertheless, these estimate scores may fail to accurately reflect the true data density or impose impractical constraints. To provide a unified perspective on density-based score design, we propose a novel theoretical framework grounded in Bregman divergence, which extends distribution considerations to encompass an exponential family of distributions. Leveraging the conjugation constraint revealed in our theorem, we introduce a \textsc{ConjNorm} method, reframing density function design as a search for the optimal norm coefficient $p$ against the given dataset. In light of the computational challenges of normalization, we devise an unbiased and analytically tractable estimator of the partition function using the Monte Carlo-based importance sampling technique. Extensive experiments across OOD detection benchmarks empirically demonstrate that our proposed \textsc{ConjNorm} has established a new state-of-the-art in a variety of OOD detection setups, outperforming the current best method by up to 13.25\% and 28.19\% (FPR95) on CIFAR-100 and ImageNet-1K, respectively. Yadan Luo, Yonggang Zhang 0003, Yixuan Li 0001, Zhen Fang 0001 |
ICLR | 5 |
| 2024 | NoiseDiffusion: Correcting Noise for Image Interpolation with Diffusion Models beyond Spherical Linear InterpolationabstractImage interpolation based on diffusion models is promising in creating fresh and interesting images.
Advanced interpolation methods mainly focus on spherical linear interpolation, where images are encoded into the noise space and then interpolated for denoising to images.
However, existing methods face challenges in effectively interpolating natural images (not generated by diffusion models), thereby restricting their practical applicability.
Our experimental investigations reveal that these challenges stem from the invalidity of the encoding noise, which may no longer obey the expected noise distribution, e.g., a normal distribution.
To address these challenges, we propose a novel approach to correct noise for image interpolation, NoiseDiffusion. Specifically, NoiseDiffusion approaches the invalid noise to the expected distribution by introducing subtle Gaussian noise and introduces a constraint to suppress noise with extreme values. In this context, promoting noise validity contributes to mitigating image artifacts, but the constraint and introduced exogenous noise typically lead to a reduction in signal-to-noise ratio, i.e., loss of original image information. Hence, NoiseDiffusion performs interpolation within the noisy image space and injects raw images into these noisy counterparts to address the challenge of information loss. Consequently, NoiseDiffusion enables us to interpolate natural images without causing artifacts or information loss, thus achieving the best interpolation results. Yonggang Zhang 0003, Zhen Fang 0001, Tongliang Liu, Defu Lian, Bo Han 0003 |
ICLR | 3 |
| 2024 | Knowledge Distillation with Auxiliary VariableabstractKnowledge distillation (KD) provides an efficient framework for transferring knowledge from a teacher model to a student model by aligning their predictive distributions. The existing KD methods adopt the same strategy as the teacher to formulate the student’s predictive distribution. However, employing the same distribution-modeling strategy typically causes sub-optimal knowledge transfer due to the discrepancy in model capacity between teacher and student models. Designing student-friendly teachers contributes to alleviating the capacity discrepancy, while it requires either complicated or student-specific training schemes. To cast off this dilemma, we propose to introduce an auxiliary variable to promote the ability of the student to model predictive distribution. The auxiliary variable is defined to be related to target variables, which will boost the model prediction. Specifically, we reformulate the predictive distribution with the auxiliary variable, deriving a novel objective function of KD. Theoretically, we provide insights to explain why the proposed objective function can outperform the existing KD methods. Experimentally, we demonstrate that the proposed objective function can considerably and consistently outperform existing KD methods. Zhen Fang 0001, Guangquan Zhang 0001, Jie Lu 0001 |
ICML | 2 |
| 2024 | CLIP-Enhanced Unsupervised Domain Adaptation with Consistency RegularizationabstractUnsupervised domain adaptation (UDA) employs labeled data from a source domain to train classifiers for an unlabeled target domain. We utilize Contrastive Language-Image Pre-training (CLIP) models to exploit textual information in labels, enabling simultaneous matching of textual and image features. However, adapting CLIP models for UDA tasks poses a significant challenge and necessitates further investigation. To this end, we introduce CLIP-Enhanced Unsupervised Domain Adaptation with Consistency Regularization, which employs consistency regularization for concurrent training of CLIP’s prompts and image adapters. Our approach, particularly under consistency regularization, incorporates data augmentation to enhance the model’s generalization capability. During training, we maintain consistent pseudo-labels for target domain data, regardless of whether weak or strong augmentation techniques are applied. This strategy improves our model’s robustness in adapting to various domains. Additionally, the integration of domain-specific prompts and image adapters in our model optimizes the learning of domain-related textual and image features. Experiments on real-world datasets substantiate the effectiveness of our proposed method. The outcomes illustrate its superior performance compared to existing techniques across multiple benchmarks. Kuo Shi, Jie Lu 0001, Zhen Fang 0001, Guangquan Zhang 0001 |
IJCNN | 3 |
| 2024 | Prompt-Based Memory Bank for Continual Test-Time Domain Adaptation in Vision-Language ModelsabstractIn dynamic environments, the generalization capabilities of large-scale vision language models tend to decline. This is attributed to the evolving distribution of target domains over time, leading to misalignment between image and text pairings, affecting the model’s performance. Addressing this, Test-Time Adaptation (TTA) has been proposed to adapt pre-trained source models to these changing target domains during testing phases. However, traditional TTA approaches, which are designed for a single changing scenario and mainly depend on self-training and entropy minimization, are easily affected by extreme and novel samples in long-term environments, leading to error accumulation and catastrophic forgetting. Although previous Continual Test-Time Adaptation (Continual TTA) methods based on the teacher-student framework can effectively address long-term adaptation issues, they are not feasible for large-scale vision language models due to their high memory requirements. To overcome these challenges, we introduce a novel approach: Prompt-based memory bank for Continual Test-Time Adaptation (PCoTTA). PCoTTA uniquely freezes the CLIP image and text encoders, focusing on updating and storing trainable prompts, significantly reducing memory usage. By implementing a stable pseudo-label strategy and high gradient sensitivity updating, PCoTTA effectively learns new knowledge. In long-term dynamically changing environments, PCoTTA demonstrates high stability and accuracy and achieves a good balance between learning new information and retaining existing knowledge, significantly enhancing the adaptability and generalization capabilities of the CLIP model. Through extensive experimental comparisons, PCoTTA surpasses the current state-of-the-art methods, achieving an average 2% improvement in accuracy for both test-time adaptation and continual test-time adaptation tasks. Ran Wang 0016, Hua Zuo, Zhen Fang 0001, Jie Lu 0001 |
IJCNN | 3 |
| 2024 | Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot GeneralizationabstractIn the field of Vision-Language Models (VLM), the Contrastive Language-Image Pretraining (CLIP) model has yielded outstanding performance on many downstream tasks through prompt tuning. By integrating image and text representations, CLIP exhibits zero-shot generalization capabilities on unseen data. However, when new categories and distribution shifts occur, the pretrained text embeddings in CLIP may not align well with unseen images, potentially leading to a decrease in CLIP's zero-shot generalization performance. To address this issue, many existing methods use test samples to update the CLIP model during testing through a process known as Test-Time Adaptation (TTA). Previous TTA techniques, such as image augmentation, can lead to overfitting given outlying samples, while methods based on teacher-student distillation can increase memory use. Further, these methods significantly increase inference time, which is a crucial factor in the testing phase. To improve robustness, mitigate overfitting, and reduce bias toward outlying samples, we propose a novel method: Self-Text Distillation with Conjugate Pseudo-labels (SCP), designed to enhance CLIP's zero-shot generalization. SCP uses gradient information from conjugate pseudo-labels to enhance the model's robustness toward distribution shifts. It also innovates by using a fixed prompt list to distil learnable prompts from within the same model, acting as a self-regulation mechanism that minimizes overfitting. Additionally, SCP is a fully test-time adaptation method that does not require retraining. It directly improves CLIP's zero-shot generalization at test time without increasing either memory overheads or inference time. In evaluations across three zero-shot generalization scenarios, SCP surpasses existing state-of-the-art methods in performance and significantly reduces inference time. Ran Wang 0016, Hua Zuo, Zhen Fang 0001, Jie Lu 0001 |
ACM Multimedia | 3 |
| 2024 | Spatio-temporal Heterogeneous Federated Learning for Time Series Classification with Multi-view Orthogonal TrainingabstractFederated learning (FL) is undergoing significant traction due to its ability to perform privacy-preserving training on decentralized data. In this work, we focus on sensitive time series data collected by distributed sensors in real-world applications. However, time series data introduce the challenge of dual spatial-temporal feature skew due to their dynamic changes across domains and time, differing from computer vision. This key challenge includes inter-client spatial feature skew caused by heterogeneous sensor collection and intra-client temporal feature skew caused by dynamics in time series distribution. We follow the framework of Personalized Federated Learning (pFL) to handle dual feature drifts to enhance the capabilities of customized local models. Therefore, in this paper, we propose a method FedST to solve key challenges through orthogonal feature decoupling and regularization in both training and testing stages. During training, we collaborate time view and frequency view of time series data to enrich the mutual information and adopt orthogonal projection to disentangle and align the shared and personalized features between views, and between clients. During testing, we apply prototype-based predictions and model-based predictions to achieve model consistency based on shared features. Extensive experiments on multiple real-world classification datasets and multimodal time series datasets show our method consistently outperforms state-of-the-art baselines with clear advantages. Chenrui Wu 0002, Haishuai Wang, Xiang Zhang 0012, Zhen Fang 0001, Jiajun Bu |
ACM Multimedia | 4 |
| 2024 | Learning to Shape In-distribution Feature Space for Out-of-distribution DetectionabstractOut-of-distribution (OOD) detection is critical for deploying machine learning models in the open world. To design scoring functions that discern OOD data from the in-distribution (ID) cases from a pre-trained discriminative model, existing methods tend to make rigorous distributional assumptions either explicitly or implicitly due to the lack of knowledge about the learned feature space in advance.
The mismatch between the learned and assumed distributions motivates us to raise a fundamental yet under-explored question: \textit{Is it possible to deterministically model the feature distribution while pre-training a discriminative model?}
This paper gives an affirmative answer to this question by presenting a Distributional Representation Learning (\texttt{DRL}) framework for OOD detection. In particular, \texttt{DRL} explicitly enforces the underlying feature space to conform to a pre-defined mixture distribution, together with an online approximation of normalization constants to enable end-to-end training. Furthermore, we formulate \texttt{DRL} into a provably convergent Expectation-Maximization algorithm to avoid trivial solutions and rearrange the sequential sampling to guide the training consistency. Extensive evaluations across mainstream OOD detection benchmarks empirically manifest the superiority of the proposed \texttt{DRL} over its advanced counterparts. Yonggang Zhang 0003, Jie Lu 0001, Zhen Fang 0001, Yiu-Ming Cheung |
NeurIPS | 4 |
| 2024 | Source-Free Unsupervised Domain Adaptation: Current research and future directions
Ningyuan Zhang, Jie Lu 0001, Keqiuyin Li, Zhen Fang 0001, Guangquan Zhang 0001 |
Neurocomputing | 4 |
| 2024 | On the Learnability of Out-of-distribution DetectionabstractSupervised learning aims to train a classifier under the assumption that training and test data are from the same distribution. To ease the above assumption, researchers have studied a more realistic setting: out-of-distribution (OOD) detection, where test data may come from classes that are unknown during training (i.e., OOD data). Due to the unavailability and diversity of OOD data, good generalization ability is crucial for effective OOD detection algorithms, and corresponding learning theory is still an open problem. To study the generalization of OOD detection, this paper investigates the probably approximately correct (PAC) learning theory of OOD detection that fits the commonly used evaluation metrics in the literature. First, we find a necessary condition for the learnability of OOD detection. Then, using this condition, we prove several impossibility theorems for the learnability of OOD detection under some scenarios. Although the impossibility theorems are frustrating, we find that some conditions of these impossibility theorems may not hold in some practical scenarios. Based on this observation, we next give several necessary and sufficient conditions to characterize the learnability of OOD detection in some practical scenarios. Lastly, we offer theoretical support for representative OOD detection works based on our OOD theory. Zhen Fang 0001, Yixuan Li 0001, Feng Liu 0003, Bo Han 0003, Jie Lu 0001 |
J. Mach. Learn. Res. | 1 |
| 2024 | Where and How to Transfer: Knowledge Aggregation-Induced Transferability Perception for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation without accessing expensive annotation processes of target data has achieved remarkable successes in semantic segmentation. However, most existing state-of-the-art methods cannot explore whether semantic representations across domains are transferable or not, which may result in the negative transfer brought by irrelevant knowledge. To tackle this challenge, in this paper, we develop a novel Knowledge Aggregation-induced Transferability Perception (KATP) for unsupervised domain adaptation, which is a pioneering attempt to distinguish transferable or untransferable knowledge across domains. Specifically, the KATP module is designed to quantify which semantic knowledge across domains is transferable, by incorporating transferability information propagation from global category-wise prototypes. Based on KATP, we design a novel KATP Adaptation Network (KATPAN) to determine where and how to transfer. The KATPAN contains a transferable appearance translation module T_A() and a transferable representation augmentation module T_R(), where both modules construct a virtuous circle of performance promotion. T_A() develops a transferability-aware information bottleneck to highlight where to adapt transferable visual characterizations and modality information; T_R() explores how to augment transferable representations while abandoning untransferable information, and promotes the translation performance of T_A() in return. Experiments on several representative datasets and a medical dataset support the state-of-the-art performance of our model. Jiahua Dong 0001, Yang Cong, Gan Sun, Zhen Fang 0001, Zhengming Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | An Extremely Simple Algorithm for Source Domain ReconstructionabstractThe aim of unsupervised domain adaptation (UDA) is to utilize knowledge from a source domain to enhance the performance of a given target domain. Due to the lack of accessibility to the target domain's labels, UDA's efficacy is highly reliant on the source domain's quality. However, it is often impractical and expensive to obtain an appropriate transferable source domain. To address this issue, we propose a novel UDA setting, source domain reconstruction (SDR), which seeks to construct a new transferable source domain utilizing labeled source samples and unlabeled target samples. SDR has a significant advantage over the conventional method as it is much less expensive to construct a suitable pseudo-source domain rather than collecting an actual transferable source domain in real-world scenarios. To test the practice of SDR, we investigate SDR theoretically. We propose an easily implementable algorithm, the domain MixUp (DMU), which is motivated by the MixUp strategy, to solve the SDR problem. The algorithm can be used to design a UDA framework to significantly enhance the performance of several existing UDA algorithms. Results from extensive experiments conducted on seven benchmarks (66 UDA tasks) indicate that the reconstructed source domain has stronger transferability than the original source domain. Zhen Fang 0001, Jie Lu 0001, Guangquan Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | Multiclass Classification With Fuzzy-Feature Observations: Theory and AlgorithmsabstractThe theoretical analysis of multiclass classification has proved that the existing multiclass classification methods can train a classifier with high classification accuracy on the test set, when the instances are precise in the training and test sets with same distribution and enough instances can be collected in the training set. However, one limitation with multiclass classification has not been solved: how to improve the classification accuracy of multiclass classification problems when only imprecise observations are available. Hence, in this article, we propose a novel framework to address a new realistic problem called multiclass classification with imprecise observations (MCIMO), where we need to train a classifier with fuzzy-feature observations. First, we give the theoretical analysis of the MCIMO problem based on fuzzy Rademacher complexity. Then, two practical algorithms based on support vector machine and neural networks are constructed to solve the proposed new problem. The experiments on both synthetic and real-world datasets verify the rationality of our theoretical analysis and the efficacy of the proposed algorithms. Guangzhi Ma, Jie Lu 0001, Feng Liu 0003, Zhen Fang 0001, Guangquan Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2024 | Domain Adaptation With Interval-Valued Observations: Theory and AlgorithmsabstractUnsupervised Domain Adaptation (UDA) focuses on enhancing the model performance on an unlabeled target domain by leveraging knowledge from a source domain. The source and target domains usually share different distributions. Existing UDA research primarily concentrates on image data characterized by crisp-valued features. However, interval-valued data, where all the observations’ features are described by intervals, is also a common type of data in real-world scenarios. For instance, measurement instruments are unable to provide exact numerical outcomes, instead employing intervals to describe their results. Hence, this paper focuses on the highly challenging context known as domain adaptation with interval-valued observations. In this environment, the objective is to improve classification accuracy within an unlabeled target domain by capitalizing on knowledge gleaned from a labeled source domain, where both domains exclusively feature interval-valued observations. To address this, we first establish an upper bound on the risk in the interval-valued target domain, underpinning our analysis with rigorous theoretical insights. Subsequently, guided by our theoretical analysis, a new model based on Takagi-Sugeno Fuzzy rules and a Self-supervised Pseudo-labeling strategy (SP-TSF) is developed to address the proposed problem. Takagi-Sugeno fuzzy rules are harnessed to handle the inherent uncertainty intrinsic to interval-valued data, while a pseudo-labeling strategy is developed to augment distribution alignment between the source and target domains, each characterized by interval-valued observations. Extensive experiments on both synthetic and realworld datasets verify the rationality of our theoretical analysis and the efficacy of the proposed model. Guangzhi Ma, Jie Lu 0001, Feng Liu 0003, Zhen Fang 0001, Guangquan Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | Unsupervised Domain Adaptation Enhanced by Fuzzy Prompt LearningabstractUnsupervised Domain Adaptation (UDA) addresses the challenge of distribution shift between a labeled source domain and an unlabeled target domain by utilizing knowledge from the source. Traditional UDA methods mainly focus on single-modal scenarios, either vision or language, thus not fully exploring the advantages of multi-modal representations. Vision-Language models utilize multi-modal information, applying prompt learning techniques for addressing target domain tasks. Motivated by the recent advancements in pre-trained Vision-Language models, this paper expands the UDA framework to incorporate multi-modal approaches using fuzzy techniques. The adoption of fuzzy techniques, preferred over conventional domain adaptation methods, is based on two key aspects: the nature of prompt learning is intrinsically linked to fuzzy logic, and the superior capability of fuzzy techniques in processing soft information and effectively utilizing inherent relationships both within and across domains. To this end, we propose UDA enhanced byFuzzy PromptLearning (FUZZLE), a simple and effective method for aligning the source and target domains via domain-specific prompt learning. Specifically, we introduce a novel technique to enhance prompt learning in the target domain. This method integrates fuzzy C-means clustering and a novel instance-level fuzzy vector into the prompt learning loss function, minimizing the distance between prompt cluster centers and instance prompts, thereby enhancing the prompt learning process. In addition, we propose a Kullback-Leibler (KL) divergence-based loss function with a fuzzification factor. This function is designed to minimize the distribution discrepancy in the classification of similar cross-domain data, aligning domain-specific prompts during the training process. We contribute an in-depth analysis to understand the effectiveness of Extensive experiments demonstrate that our method achieves superior performance on standard UDA benchmarks. Kuo Shi, Jie Lu 0001, Zhen Fang 0001, Guangquan Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Continual Named Entity Recognition without Catastrophic ForgettingabstractContinual Named Entity Recognition (CNER) is a burgeoning area, which involves updating an existing model by incorporating new entity types sequentially.Nevertheless, continual learning approaches are often severely afflicted by catastrophic forgetting.This issue is intensified in CNER due to the consolidation of old entity types from previous steps into the non-entity type at each step, leading to what is known as the semantic shift problem of the non-entity type.In this paper, we introduce a pooled feature distillation loss that skillfully navigates the trade-off between retaining knowledge of old entity types and acquiring new ones, thereby more effectively mitigating the problem of catastrophic forgetting.Additionally, we develop a confidence-based pseudo-labeling for the non-entity type, i.e., predicting entity types using the old model to handle the semantic shift of the non-entity type.Following the pseudo-labeling process, we suggest an adaptive re-weighting type-balanced learning strategy to handle the issue of biased type distribution.We carried out comprehensive experiments on ten CNER settings using three different datasets.The results illustrate that our method significantly outperforms prior state-of-the-art approaches, registering an average improvement of 6.3% and 8.0% in Micro and Macro F1 scores, respectively.1 * Equal contributions.† The corresponding author is Dr. Duzhen Zhang, Wei Cong, Jiahua Dong 0001, Yahan Yu, Xiuyi Chen, Yonggang Zhang 0003, Zhen Fang 0001 |
EMNLP | 7 |
| 2023 | Kecor: Kernel Coding Rate Maximization for Active 3D Object DetectionabstractAchieving a reliable LiDAR-based object detector in autonomous driving is paramount, but its success hinges on obtaining large amounts of precise 3D annotations. Active learning (AL) seeks to mitigate the annotation burden through algorithms that use fewer labels and can attain performance comparable to fully supervised learning. Although AL has shown promise, current approaches prioritize the selection of unlabeled point clouds with high uncertainty and/or diversity, leading to the selection of more instances for labeling and reduced computational efficiency. In this paper, we resort to a novel kernel coding rate maximization (Kecor) strategy which aims to identify the most informative point clouds to acquire labels through the lens of information theory. Greedy search is applied to seek desired point clouds that can maximize the minimal number of bits required to encode the latent features. To determine the uniqueness and informativeness of the selected samples from the model perspective, we construct a proxy network of the 3D detector head and compute the outer product of Jacobians from all proxy layers to form the empirical neural tangent kernel (NTK) matrix. To accommodate both one-stage (i.e., Second) and two-stage detectors (i.e., Pv-rcnn), we further incorporate the classification entropy maximization and well trade-off between detection performance and the total number of bounding boxes selected for annotation. Extensive experiments conducted on two 3D benchmarks and a 2D detection dataset evidence the superiority and versatility of the proposed approach. Our results show that approximately 44% box-level annotation costs and 26% computational time are reduced compared to the state-of-the-art AL method, without compromising detection performance. Source code: https://github.com/Luoyadan/KECOR-active-3Ddet. Yadan Luo, Zhuoxiao Chen, Zhen Fang 0001, Zheng Zhang 0006, Mahsa Baktash, Zi Huang |
ICCV | 3 |
| 2023 | Meta OOD Learning For Continuously Adaptive OOD DetectionabstractOut-of-distribution (OOD) detection is crucial to modern deep learning applications by identifying and alerting about the OOD samples that should not be tested or used for making predictions. Current OOD detection methods have made significant progress when in-distribution (ID) and OOD samples are drawn from static distributions. However, this can be unrealistic when applied to real-world systems which often undergo continuous variations and shifts in ID and OOD distributions over time. Therefore, for an effective application in real-world systems, the development of OOD detection methods that can adapt to these dynamic and evolving distributions is essential. In this paper, we propose a novel and more realistic setting called continuously adaptive out-of-distribution (CAOOD) detection which targets on developing an OOD detection model that enables dynamic and quick adaptation to a new arriving distribution, with insufficient ID samples during deployment time. To address CAOOD, we develop meta OOD learning (MOL) by designing a learning-to-adapt diagram such that a good initialized OOD detection model is learned during the training process. In the testing process, MOL ensures OOD detection performance over shifting distributions by quickly adapting to new distributions with a few adaptations. Extensive experiments on several OOD benchmarks endorse the effectiveness of our method in preserving both ID classification accuracy and OOD detection performance on continuously shifting distributions. Xinheng Wu, Jie Lu 0001, Zhen Fang 0001, Guangquan Zhang 0001 |
ICCV | 3 |
| 2023 | Moderately Distributional Exploration for Domain GeneralizationabstractDomain generalization (DG) aims to tackle the distribution shift between training domains and unknown target domains. Generating new domains is one of the most effective approaches, yet its performance gain depends on the distribution discrepancy between the generated and target domains. Distributionally robust optimization is promising to tackle distribution discrepancy by exploring domains in an uncertainty set. However, the uncertainty set may be overwhelmingly large, leading to low-confidence prediction in DG. It is because a large uncertainty set could introduce domains containing semantically different factors from training domains. To address this issue, we propose to perform a $\textit{mo}$derately $\textit{d}$istributional $\textit{e}$xploration (MODE) for domain generalization. Specifically, MODE performs distribution exploration in an uncertainty $\textit{subset}$ that shares the same semantic factors with the training domains. We show that MODE can endow models with provable generalization performance on unknown target domains. The experimental results show that MODE achieves competitive performance compared to state-of-the-art baselines. Rui Dai 0005, Yonggang Zhang 0003, Zhen Fang 0001, Bo Han 0003, Xinmei Tian 0001 |
ICML | 3 |
| 2023 | Detecting Out-of-distribution Data through In-distribution Class PriorabstractGiven a pre-trained in-distribution (ID) model, the inference-time out-of-distribution (OOD) detection aims to recognize OOD data during the inference stage. However, some representative methods share an unproven assumption that the probability that OOD data belong to every ID class should be the same, i.e., these OOD-to-ID probabilities actually form a uniform distribution. In this paper, we show that this assumption makes the above methods incapable when the ID model is trained with class-imbalanced data.Fortunately, by analyzing the causal relations between ID/OOD classes and features, we identify several common scenarios where the OOD-to-ID probabilities should be the ID-class-prior distribution and propose two strategies to modify existing inference-time detection methods: 1) replace the uniform distribution with the ID-class-prior distribution if they explicitly use the uniform distribution; 2) otherwise, reweight their scores according to the similarity between the ID-class-prior distribution and the softmax outputs of the pre-trained model. Extensive experiments show that both strategies can improve the OOD detection performance when the ID model is pre-trained with imbalanced data, reflecting the importance of ID-class prior in OOD detection. Feng Liu 0003, Zhen Fang 0001, Hong Chen 0004, Tongliang Liu, Feng Zheng 0001, Bo Han 0003 |
ICML | 3 |
| 2023 | One-step Domain Adaptation Approach with Partial LabelabstractUnsupervised Domain adaptation (UDA) aims to train a target classifier by using massive accurate annotated source domain data and unlabeled target domain data. However, collecting massive accurate annotation can be labor-intensive, and even impractical especially when source domain training data shows label ambiguity. In this paper, we consider a novel domain adaptation setting where the model can be trained with partial-labeled source domain data, so that the cost of data labeling can be reduced. To alleviate the ambiguity induced by partial label, we propose a one-step domain adaptation approach trained from the partial-labeled source data and unlabeled target data. Our approach consists of two components: a feature extractor equipped with partial label loss that learns discriminative representation, and a domain classifier that learns domain-invariant representation across the source and target domains. Extensive experiments have shown that the proposed approach significantly outperforms a series of competitive baselines. Guohang Zeng, Zhen Fang 0001, Guangquan Zhang 0001, Jie Lu 0001 |
IJCNN | 2 |
| 2023 | Learning to Augment Distributions for Out-of-distribution DetectionabstractOpen-world classification systems should discern out-of-distribution (OOD) data whose labels deviate from those of in-distribution (ID) cases, motivating recent studies in OOD detection. Advanced works, despite their promising progress, may still fail in the open world, owing to the lacking knowledge about unseen OOD data in advance. Although one can access auxiliary OOD data (distinct from unseen ones) for model training, it remains to analyze how such auxiliary data will work in the open world. To this end, we delve into such a problem from a learning theory perspective, finding that the distribution discrepancy between the auxiliary and the unseen real OOD data is the key to affect the open-world detection performance. Accordingly, we propose Distributional-Augmented OOD Learning (DAOL), alleviating the OOD distribution discrepancy by crafting an OOD distribution set that contains all distributions in a Wasserstein ball centered on the auxiliary OOD distribution. We justify that the predictor trained over the worst OOD data in the ball can shrink the OOD distribution discrepancy, thus improving the open-world detection performance given only the auxiliary OOD data. We conduct extensive evaluations across representative OOD detection setups, demonstrating the superiority of our DAOL over its advanced counterparts. Zhen Fang 0001, Yonggang Zhang 0003, Feng Liu 0003, Yixuan Li 0001, Bo Han 0003 |
NeurIPS | 2 |
| 2023 | SODA: Robust Training of Test-Time Data AdaptorsabstractAdapting models deployed to test distributions can mitigate the performance degradation caused by distribution shifts. However, privacy concerns may render model parameters inaccessible. One promising approach involves utilizing zeroth-order optimization (ZOO) to train a data adaptor to adapt the test data to fit the deployed models. Nevertheless, the data adaptor trained with ZOO typically brings restricted improvements due to the potential corruption of data features caused by the data adaptor. To address this issue, we revisit ZOO in the context of test-time data adaptation. We find that the issue directly stems from the unreliable estimation of the gradients used to optimize the data adaptor, which is inherently due to the unreliable nature of the pseudo-labels assigned to the test data. Based on this observation, we propose pseudo-label-robust data adaptation (SODA) to improve the performance of data adaptation. Specifically, SODA leverages high-confidence predicted labels as reliable labels to optimize the data adaptor with ZOO for label prediction. For data with low-confidence predictions, SODA encourages the adaptor to preserve data information to mitigate data corruption. Empirical results indicate that SODA can significantly enhance the performance of deployed models in the presence of distribution shifts without requiring access to model parameters. Zige Wang, Yonggang Zhang 0003, Zhen Fang 0001, Long Lan, Wenjing Yang 0002, Bo Han 0003 |
NeurIPS | 3 |
| 2023 | Invariant Learning via Probability of Sufficient and Necessary CausesabstractOut-of-distribution (OOD) generalization is indispensable for learning models in the wild, where testing distribution typically unknown and different from the training. Recent methods derived from causality have shown great potential in achieving OOD generalization.
However, existing methods mainly focus on the invariance property of causes, while largely overlooking the property of sufficiency and necessity conditions. Namely, a necessary but insufficient cause (feature) is invariant to distribution shift, yet it may not have required accuracy. By contrast, a sufficient yet unnecessary cause (feature) tends to fit specific data well but may have a risk of adapting to a new domain.
To capture the information of sufficient and necessary causes, we employ a classical concept, the probability of sufficiency and necessary causes (PNS), which indicates the probability of whether one is the necessary and sufficient cause.
To associate PNS with OOD generalization, we propose PNS risk and formulate an algorithm to learn representation with a high PNS value. We theoretically analyze and prove the generalizability of the PNS risk. Experiments on both synthetic and real-world benchmarks demonstrate the effectiveness of the proposed method. The detailed implementation can be found at the GitHub repository: https://github.com/ymy4323460/CaSN. Mengyue Yang, Yonggang Zhang 0003, Zhen Fang 0001, Yali Du 0001, Furui Liu, Jean-Francois Ton, Jun Wang 0012 |
NeurIPS | 3 |
| 2023 | Out-of-distribution Detection Learning with Unreliable Out-of-distribution SourcesabstractOut-of-distribution (OOD) detection discerns OOD data where the predictor cannot make valid predictions as in-distribution (ID) data, thereby increasing the reliability of open-world classification. However, it is typically hard to collect real out-of-distribution (OOD) data for training a predictor capable of discerning ID and OOD patterns. This obstacle gives rise to *data generation-based learning methods*, synthesizing OOD data via data generators for predictor training without requiring any real OOD data.
Related methods typically pre-train a generator on ID data and adopt various selection procedures to find those data likely to be the OOD cases. However, generated data may still coincide with ID semantics, i.e., mistaken OOD generation remains, confusing the predictor between ID and OOD data. To this end, we suggest that generated data (with mistaken OOD generation) can be used to devise an *auxiliary OOD detection task* to facilitate real OOD detection. Specifically, we can ensure that learning from such an auxiliary task is beneficial if the ID and the OOD parts have disjoint supports, with the help of a well-designed training procedure for the predictor. Accordingly, we propose a powerful data generation-based learning method named *Auxiliary Task-based OOD Learning* (ATOL) that can relieve the mistaken OOD generation. We conduct extensive experiments under various OOD detection setups, demonstrating the effectiveness of our method against its advanced counterparts. Haotian Zheng 0001, Zhen Fang 0001, Xiaobo Xia, Feng Liu 0003, Tongliang Liu, Bo Han 0003 |
NeurIPS | 3 |
| 2023 | Semi-Supervised Heterogeneous Domain Adaptation: Theory and AlgorithmsabstractSemi-supervised heterogeneous domain adaptation (SsHeDA) aims to train a classifier for the target domain, in which only unlabeled and a small number of labeled data are available. This is done by leveraging knowledge acquired from a heterogeneous source domain. From algorithmic perspectives, several methods have been proposed to solve the SsHeDA problem; yet there is still no theoretical foundation to explain the nature of the SsHeDA problem or to guide new and better solutions. Motivated by compatibility condition in semi-supervised probably approximately correct (PAC) theory, we explain the SsHeDA problem by proving its generalization error - that is, why labeled heterogeneous source data and unlabeled target data help to reduce the target risk. Guided by our theory, we devise two algorithms as proof of concept. One, kernel heterogeneous domain alignment (KHDA), is a kernel-based algorithm; the other, joint mean embedding alignment (JMEA), is a neural network-based algorithm. When a dataset is small, KHDA's training time is less than JMEA's. When a dataset is large, JMEA is more accurate in the target domain. Comprehensive experiments with image/text classification tasks show KHDA to be the most accurate among all non-neural network baselines, and JMEA to be the most accurate among all baselines. Zhen Fang 0001, Jie Lu 0001, Feng Liu 0003, Guangquan Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Bridging the Theoretical Bound and Deep Algorithms for Open Set Domain AdaptationabstractIn the unsupervised open set domain adaptation (UOSDA), the target domain contains unknown classes that are not observed in the source domain. Researchers in this area aim to train a classifier to accurately: 1) recognize unknown target data (data with unknown classes) and 2) classify other target data. To achieve this aim, a previous study has proven an upper bound of the target-domain risk, and the open set difference, as an important term in the upper bound, is used to measure the risk on unknown target data. By minimizing the upper bound, a shallow classifier can be trained to achieve the aim. However, if the classifier is very flexible [e.g., deep neural networks (DNNs)], the open set difference will converge to a negative value when minimizing the upper bound, which causes an issue where most target data are recognized as unknown data. To address this issue, we propose a new upper bound of target-domain risk for UOSDA, which includes four terms: source-domain risk,$\epsilon $-open set difference ($\Delta _\epsilon $), distributional discrepancy between domains, and a constant. Compared with the open set difference,$\Delta _\epsilon $is more robust against the issue when it is being minimized, and thus we are able to use very flexible classifiers (i.e., DNNs). Then, we propose a new principle-guided deep UOSDA method that trains DNNs via minimizing the new upper bound. Specifically, source-domain risk and$\Delta _\epsilon $are minimized by gradient descent, and the distributional discrepancy is minimized via a novel open set conditional adversarial training strategy. Finally, compared with the existing shallow and deep UOSDA methods, our method shows the state-of-the-art performance on several benchmark datasets, including digit recognition [modified National Institute of Standards and Technology database (MNIST), the Street View House Number (SVHN), U.S. Postal Service (USPS)], object recognition (Office-31, Office-Home), and face recognition [pose, illumination, and expression (PIE)]. Zhong Li 0001, Zhen Fang 0001, Feng Liu 0003, Bo Yuan 0003, Guangquan Zhang 0001, Jie Lu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Federated Class-Incremental LearningabstractFederated learning (FL) has attracted growing attentions via data-private collaborative training on decentralized clients. However, most existing methods unrealistically assume object classes of the overall framework are fixed over time. It makes the global model suffer from significant catastrophic forgetting on old classes in real-world scenarios, where local clients often collect new classes continuously and have very limited storage memory to store old classes. Moreover, new clients with unseen new classes may participate in the FL training, further aggravating the catastrophic forgetting of global model. To address these challenges, we develop a novel Global-Local Forgetting Compensation (GLFC) model, to learn a global class-incremental model for alleviating the catastrophic forgetting from both local and global perspectives. Specifically, to address local forgetting caused by class imbalance at the local clients, we design a class-aware gradient compensation loss and a class-semantic relation distillation loss to balance the forgetting of old classes and distill consistent inter-class relations across tasks. To tackle the global forgetting brought by the non-i.i.d class imbalance across clients, we propose a proxy server that selects the best old global model to assist the local relation distillation. Moreover, a prototype gradient-based communication mechanism is developed to protect the privacy. Our model outperforms state-of-the-art methods by 4.4%~15.1% in terms of average accuracy on representative benchmark datasets. The code is available at https://github.com/conditionWang/FCIL. Jiahua Dong 0001, Lixu Wang, Zhen Fang 0001, Gan Sun, Shichao Xu, Xiao Wang 0012, Qi Zhu 0002 |
CVPR | 3 |
| 2022 | Is Out-of-Distribution Detection Learnable?abstractSupervised learning aims to train a classifier under the assumption that training and test data are from the same distribution. To ease the above assumption, researchers have studied a more realistic setting: out-of-distribution (OOD) detection, where test data may come from classes that are unknown during training (i.e., OOD data). Due to the unavailability and diversity of OOD data, good generalization ability is crucial for effective OOD detection algorithms. To study the generalization of OOD detection, in this paper, we investigate the probably approximately correct (PAC) learning theory of OOD detection, which is proposed by researchers as an open problem. First, we find a necessary condition for the learnability of OOD detection. Then, using this condition, we prove several impossibility theorems for the learnability of OOD detection under some scenarios. Although the impossibility theorems are frustrating, we find that some conditions of these impossibility theorems may not hold in some practical scenarios. Based on this observation, we next give several necessary and sufficient conditions to characterize the learnability of OOD detection in some practical scenarios. Lastly, we also offer theoretical supports for several representative OOD detection works based on our OOD theory. Zhen Fang 0001, Yixuan Li 0001, Jie Lu 0001, Jiahua Dong 0001, Bo Han 0003, Feng Liu 0003 |
NeurIPS | 1 |
| 2022 | Learning From a Complementary-Label Source Domain: Theory and AlgorithmsabstractIn unsupervised domain adaptation (UDA), a classifier for the target domain is trained with massive true-label data from the source domain and unlabeled data from the target domain. However, collecting true-label data in the source domain can be expensive and sometimes impractical. Compared to the true label (TL), a complementary label (CL) specifies a class that a pattern does not belong to, and hence, collecting CLs would be less laborious than collecting TLs. In this article, we propose a novel setting where the source domain is composed of complementary-label data, and a theoretical bound of this setting is provided. We consider two cases of this setting: one is that the source domain only contains complementary-label data [completely complementary UDA (CC-UDA)] and the other is that the source domain has plenty of complementary-label data and a small amount of true-label data [partly complementary UDA (PC-UDA)]. To this end, a complementary label adversarial network (CLARINET) is proposed to solve CC-UDA and PC-UDA problems. CLARINET maintains two deep networks simultaneously, with one focusing on classifying the complementary-label source data and the other taking care of the source-to-target distributional adaptation. Experiments show that CLARINET significantly outperforms a series of competent baselines on handwritten digit-recognition and object-recognition tasks. Yiyang Zhang 0005, Feng Liu 0003, Zhen Fang 0001, Bo Yuan 0003, Guangquan Zhang 0001, Jie Lu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | How Does the Combined Risk Affect the Performance of Unsupervised Domain Adaptation Approaches?abstractUnsupervised domain adaptation (UDA) aims to train a target classifier with labeled samples from the source domain and unlabeled samples from the target domain. Classical UDA learning bounds show that target risk is upper bounded by three terms: source risk, distribution discrepancy, and combined risk. Based on the assumption that the combined risk is a small fixed value, methods based on this bound train a target classifier by only minimizing estimators of the source risk and the distribution discrepancy. However, the combined risk may increase when minimizing both estimators, which makes the target risk uncontrollable. Hence the target classifier cannot achieve ideal performance if we fail to control the combined risk. To control the combined risk, the key challenge takes root in the unavailability of the labeled samples in the target domain. To address this key challenge, we propose a method named E-MixNet. E-MixNet employs enhanced mixup, a generic vicinal distribution, on the labeled source samples and pseudo-labeled target samples to calculate a proxy of the combined risk. Experiments show that the proxy can effectively curb the increase of the combined risk when minimizing the source risk and distribution discrepancy. Furthermore, we show that if the proxy of the combined risk is added into loss functions of four representative UDA methods, their performance is also improved. Zhong Li 0001, Zhen Fang 0001, Feng Liu 0003, Jie Lu 0001, Bo Yuan 0003, Guangquan Zhang 0001 |
AAAI | 2 |
| 2021 | Learning Bounds for Open-Set LearningabstractTraditional supervised learning aims to train a classifier in the closed-set world, where training and test samples share the same label space. In this paper, we target a more challenging and re_x0002_alistic setting: open-set learning (OSL), where there exist test samples from the classes that are unseen during training. Although researchers have designed many methods from the algorith_x0002_mic perspectives, there are few methods that pro_x0002_vide generalization guarantees on their ability to achieve consistent performance on different train_x0002_ing samples drawn from the same distribution. Motivated by the transfer learning and probably approximate correct (PAC) theory, we make a bold attempt to study OSL by proving its general_x0002_ization error-given training samples with size n, the estimation error will get close to order Op(1/$\sqrt{}$n). This is the first study to provide a generalization bound for OSL, which we do by theoretically investigating the risk of the tar_x0002_get classifier on unknown classes. According to our theory, a novel algorithm, called auxiliary open-set risk (AOSR) is proposed to address the OSL problem. Experiments verify the efficacy of AOSR. The code is available at github.com/AnjinLiu/Openset_Learning_AOSR. Zhen Fang 0001, Jie Lu 0001, Anjin Liu, Feng Liu 0003, Guangquan Zhang 0001 |
ICML | 1 |
| 2021 | Confident Anchor-Induced Multi-Source Free Domain AdaptationabstractUnsupervised domain adaptation has attracted appealing academic attentions by transferring knowledge from labeled source domain to unlabeled target domain. However, most existing methods assume the source data are drawn from a single domain, which cannot be successfully applied to explore complementarily transferable knowledge from multiple source domains with large distribution discrepancies. Moreover, they require access to source data during training, which are inefficient and unpractical due to privacy preservation and memory storage. To address these challenges, we develop a novel Confident-Anchor-induced multi-source-free Domain Adaptation (CAiDA) model, which is a pioneer exploration of knowledge adaptation from multiple source domains to the unlabeled target domain without any source data, but with only pre-trained source models. Specifically, a source-specific transferable perception module is proposed to automatically quantify the contributions of the complementary knowledge transferred from multi-source domains to the target domain. To generate pseudo labels for the target domain without access to the source data, we develop a confident-anchor-induced pseudo label generator by constructing a confident anchor group and assigning each unconfident target sample with a semantic-nearest confident anchor. Furthermore, a class-relationship-aware consistency loss is proposed to preserve consistent inter-class relationships by aligning soft confusion matrices across domains. Theoretical analysis answers why multi-source domains are better than a single source domain, and establishes a novel learning bound to show the effectiveness of exploiting multi-source domains. Experiments on several representative datasets illustrate the superiority of our proposed CAiDA model. The code is available at https://github.com/Learning-group123/CAiDA. Jiahua Dong 0001, Zhen Fang 0001, Anjin Liu, Gan Sun, Tongliang Liu |
NeurIPS | 2 |
| 2021 | Open Set Domain Adaptation: Theoretical Bound and AlgorithmabstractThe aim of unsupervised domain adaptation is to leverage the knowledge in a labeled (source) domain to improve a model's learning performance with an unlabeled (target) domain-the basic strategy being to mitigate the effects of discrepancies between the two distributions. Most existing algorithms can only handle unsupervised closed set domain adaptation (UCSDA), i.e., where the source and target domains are assumed to share the same label set. In this article, we target a more challenging but realistic setting: unsupervised open set domain adaptation (UOSDA), where the target domain has unknown classes that are not found in the source domain. This is the first study to provide learning bound for open set domain adaptation, which we do by theoretically investigating the risk of the target classifier on unknown classes. The proposed learning bound has a special term, namely, open set difference, which reflects the risk of the target classifier on unknown classes. Furthermore, we present a novel and theoretically guided unsupervised algorithm for open set domain adaptation, called distribution alignment with open difference (DAOD), which is based on regularizing this open set difference bound. The experiments on several benchmark data sets show the superior performance of the proposed UOSDA method compared with the state-of-the-art methods in the literature. Zhen Fang 0001, Jie Lu 0001, Feng Liu 0003, Junyu Xuan, Guangquan Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Clarinet: A One-step Approach Towards Budget-friendly Unsupervised Domain AdaptationabstractIn unsupervised domain adaptation (UDA), classifiers for the target domain are trained with massive true-label data from the source domain and unlabeled data from the target domain. However, it may be difficult to collect fully-true-label data in a source domain given limited budget. To mitigate this problem, we consider a novel problem setting where the classifier for the target domain has to be trained with complementary-label data from the source domain and unlabeled data from the target domain named budget-friendly UDA (BFUDA). The key benefit is that it is much less costly to collect complementary-label source data (required by BFUDA) than collecting the true-label source data (required by ordinary UDA). To this end, complementary label adversarial network (CLARINET) is proposed to solve the BFUDA problem. CLARINET maintains two deep networks simultaneously, where one focuses on classifying complementary-label source data and the other takes care of the source-to-target distributional adaptation. Experiments show that CLARINET significantly outperforms a series of competent baselines. Yiyang Zhang 0005, Feng Liu 0003, Zhen Fang 0001, Bo Yuan 0003, Guangquan Zhang 0001, Jie Lu 0001 |
IJCAI | 3 |
| 2019 | Unsupervised Domain Adaptation with Sphere Retracting TransformationabstractUnsupervised domain adaptation aims to leverage the knowledge in training data (source domain) to improve the performance of tasks in the remaining unlabeled data (target domain) by mitigating the effect of the distribution discrepancy. Existing approaches resolve this problem mainly by 1) mapping data into a latent space where the distribution discrepancy between two domains is reduced; or 2) reducing the domain shift by weighting the source domain. However, most of these approaches share a common issue that they neglect inter-class margins while matching distributions, which has a significant impact on classification performance. In this paper, we analyze the issue from the theoretical aspect and propose a novel unsupervised domain adaptation approach: Sphere Retracting Transformation (SRT), which reduces the distribution discrepancy and increases inter-class margins. We implement SRT, according to our theoretical analysis by (1) assigning class-specific weights for data in the source domain, and (2) minimizing the intra-class variations. Experiments confirm that the SRT approach outperforms several competitive approaches for standard domain adaptation benchmarks. Zhen Fang 0001, Jie Lu 0001, Feng Liu 0003, Guangquan Zhang 0001 |
IJCNN | 1 |