VLDB 2026 Research / reviewers in the wild / expert
Changjian Shui
dblp:215/5461
· DBLP profile ↗
29ranked-venue papers
10as first author
26since 2021 · last 2025
0000-0001-6447-6559ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reliably detecting model failures in deployment without labelsabstractThe distribution of data changes over time; models operating in dynamic environments need retraining. But knowing when to retrain, without access to labels, is an open challenge since some, but not all shifts degrade model performance. This paper formalizes and addresses the problem of post-deployment deterioration (PDD) monitoring. We propose D3M, a practical and efficient monitoring algorithm based on the disagreement of predictive models, achieving low false positive rates under non-deteriorating shifts and provides sample complexity bounds for high true positive rates under deteriorating shifts. Empirical results on both standard benchmark and a real-world large-scale internal medicine dataset demonstrate the effectiveness of the framework and highlight its viability as an alert mechanism for high-stakes machine learning pipelines. Viet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya, Amol A. Verma, Fahad Razak, Rahul G. Krishnan |
NeurIPS | 2 |
| 2025 | Unraveling the Mysteries of Label Noise in Source-Free Domain Adaptation: Theory and PracticeabstractRecent source-free domain adaptation (SFDA) methods have focused on learning meaningful cluster structures in feature space, successfully adapting the knowledge from the source domain to the unlabeled target domain without accessing the private source data. However, existing methods rely on pseudo-labels generated by source models that can be noisy due to domain shift, presenting a significant challenge to their efficacy. In this paper, we study SFDA from the perspective of learning with label noise (LLN) and prove that the label noise in SFDA, unlike in conventional LLN scenarios, follows a different distribution assumption. This discrepancy renders some existing LLN methods less effective in SFDA. To address this issue and comprehensively improve adaptation performance, we tackle label noise in SFDA from two perspectives. First, we demonstrate that the early-time training phenomenon (ETP), previously observed in LLN settings, still exists in SFDA. Hence, we introduce a simple yet effective approach to leveraging ETP to improve current SFDA algorithms. Second, we propose a noise and variance control module, mitigating the label noise discrepancy between SFDA and LLN and enhancing the effectiveness of LLN methods in SFDA. Extensive empirical evaluation and analysis of four benchmarks show that our methods substantially outperform existing baselines. Gezheng Xu, Pengcheng Xu 0008, Jiaqi Li 0005, Ruizhi Pu, Changjian Shui, A. Ian McLeod, Boyu Wang 0004, Charles Ling 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Generalizing across Temporal Domains with Koopman OperatorsabstractIn the field of domain generalization, the task of constructing a predictive model capable of generalizing to a target domain without access to target data remains challenging. This problem becomes further complicated when considering evolving dynamics between domains. While various approaches have been proposed to address this issue, a comprehensive understanding of the underlying generalization theory is still lacking. In this study, we contribute novel theoretic results that aligning conditional distribution leads to the reduction of generalization bounds. Our analysis serves as a key motivation for solving the Temporal Domain Generalization (TDG) problem through the application of Koopman Neural Operators, resulting in Temporal Koopman Networks (TKNets). By employing Koopman Neural Operators, we effectively address the time-evolving distributions encountered in TDG using the principles of Koopman theory, where measurement functions are sought to establish linear transition relations between evolving domains. Through empirical evaluations conducted on synthetic and real-world datasets, we validate the effectiveness of our proposed approach. Qiuhao Zeng, Wei Wang 0036, Fan Zhou 0006, Gezheng Xu, Ruizhi Pu, Changjian Shui, Christian Gagné 0001, Charles Ling 0001, Boyu Wang 0004 |
AAAI | 6 |
| 2024 | Towards Progressive Multi-Frequency Representation for Image WarpingabstractImage warping, a classic task in computer vision, aims to use geometric transformations to change the appearance of images. Recent methods learn the resampling kernels for warping through neural networks to estimate missing values in irregular grids, which, however, fail to capture local variations in deformed content and produce images with distortion and less high-frequency details. To address this issue, this paper proposes an effective method, namely MFR, to learn Multi-Frequency Representations from in-put images for image warping. Specifically, we propose a progressive filtering network to learn image representations from different frequency subbands and generate deformable images in a coarse-to-fine manner. Furthermore, we employ learnable Gabor wavelet filters to improve the model's capability to learn local spatial-frequency representations. Comprehensive experiments, including homography trans-formation, equirectangular to perspective projection, and asymmetric image super-resolution, demonstrate that the proposed MFR significantly outperforms state-of-the-art image warping methods. Our method also showcases superior generalization to out-of-distribution domains, where the generated images are equipped with rich details and less distortion, thereby high visual quality. The source code is available at https://github.com/junxiao01/MFR. Jun Xiao 0010, Zihang Lyu, Yakun Ju, Changjian Shui, Kin-Man Lam 0001 |
CVPR | 5 |
| 2024 | Learning Equilibrium Transformation for Gamut Expansion and Color Restoration
Jun Xiao 0010, Changjian Shui, Kin-Man Lam 0001 |
ECCV (71) | 2 |
| 2024 | Latent Trajectory Learning for Limited Timestamps under Distribution Shift over TimeabstractDistribution shifts over time are common in real-world machine-learning applications. This scenario is formulated as Evolving Domain Generalization (EDG), where models aim to generalize well to unseen target domains in a time-varying system by learning and leveraging the underlying evolving pattern of the distribution shifts across domains. However, existing methods encounter challenges due to the limited number of timestamps (every domain corresponds to a timestamp) in EDG datasets, leading to difficulties in capturing evolving dynamics and risking overfitting to the sparse timestamps, which hampers their generalization and adaptability to new tasks. To address this limitation, we propose a novel approach SDE-EDG that collects the Infinitely Fined-Grid Evolving Trajectory (IFGET) of the data distribution with continuous-interpolated samples to bridge temporal gaps (intervals between two successive timestamps). Furthermore, by leveraging the inherent capacity of Stochastic Differential Equations (SDEs) to capture continuous trajectories, we propose their use to align SDE-modeled trajectories with IFGET across domains, thus enabling the capture of evolving distribution trends. We evaluate our approach on several benchmark datasets and demonstrate that it can achieve superior performance compared to existing state-of-the-art methods. Qiuhao Zeng, Changjian Shui, Long-Kai Huang, Xi Chen 0009, Charles Ling 0001, Boyu Wang 0004 |
ICLR | 2 |
| 2024 | Intersectional Unfairness DiscoveryabstractAI systems have been shown to produce unfair results for certain subgroups of population, highlighting the need to understand bias on certain sensitive attributes. Current research often falls short, primarily focusing on the subgroups characterized by a single sensitive attribute, while neglecting the nature of intersectional fairness of multiple sensitive attributes. This paper focuses on its one fundamental aspect by discovering diverse high-bias intersectional sensitive attributes. Specifically, we propose a Bias-Guided Generative Network (BGGN). By treating each bias value as a reward, BGGN efficiently generates high-bias intersectional sensitive attributes. Experiments on real-world text and image datasets demonstrate a diverse and efficient discovery of BGGN. To further evaluate the generated unseen but possible unfair intersectional sensitive attributes, we formulate them as prompts and use modern generative AI to produce new text and images. The results of frequently generating biased data provides new insights of discovering potential unfairness in popular modern generative AI systems. Warning: This paper contains examples that are offensive in nature. Gezheng Xu, Qi Chen 0015, Charles Ling 0001, Boyu Wang 0004, Changjian Shui |
ICML | 5 |
| 2024 | Hessian Aware Low-Rank Perturbation for Order-Robust Continual LearningabstractContinual learning aims to learn a series of tasks sequentially without forgetting the knowledge acquired from the previous ones. In this work, we propose the Hessian Aware Low-Rank Perturbation algorithm for continual learning. By modeling the parameter transitions along the sequential tasks with the weight matrix transformation, we propose to apply the low-rank approximation on the task-adaptive parameters in each layer of the neural networks. Specifically, we theoretically demonstrate the quantitative relationship between the Hessian and the proposed low-rank approximation. The approximation ranks are then globally determined according to the marginal change of the empirical loss estimated by the layer-specific gradient and low-rank approximation error. Furthermore, we control the model capacity by pruning less important parameters to diminish the parameter growth. We conduct extensive experiments on various benchmarks, including a dataset with large-scale tasks, and compare our method against some recent state-of-the-art methods to demonstrate the effectiveness and scalability of our proposed method. Empirical results show that our method performs better on different benchmarks, especially in achieving task order robustness and handling the forgetting issue. Jiaqi Li 0005, Yuanhao Lai, Rui Wang 0121, Changjian Shui, Sabyasachi Sahoo, Charles Ling 0001, Boyu Wang 0004, Christian Gagné 0001, Fan Zhou 0006 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Mitigating Calibration Bias Without Fixed Attribute Grouping for Improved Fairness in Medical Imaging Analysis
Changjian Shui, Justin Szeto, Raghav Mehta, Douglas L. Arnold, Tal Arbel |
MICCAI (3) | 1 |
| 2023 | On the Stability-Plasticity Dilemma in Continual Meta-Learning: Theory and AlgorithmabstractWe focus on Continual Meta-Learning (CML), which targets accumulating and exploiting meta-knowledge on a sequence of non-i.i.d. tasks. The primary challenge is to strike a balance between stability and plasticity, where a model should be stable to avoid catastrophic forgetting in previous tasks and plastic to learn generalizable concepts from new tasks. To address this, we formulate the CML objective as controlling the average excess risk upper bound of the task sequence, which reflects the trade-off between forgetting and generalization. Based on the objective, we introduce a unified theoretical framework for CML in both static and shifting environments, providing guarantees for various task-specific learning algorithms. Moreover, we first present a rigorous analysis of a bi-level trade-off in shifting environments. To approach the optimal trade-off, we propose a novel algorithm that dynamically adjusts the meta-parameter and its learning rate w.r.t environment change. Empirical evaluations on synthetic and real datasets illustrate the effectiveness of the proposed theory and algorithm. Qi Chen 0015, Changjian Shui, Ligong Han, Mario Marchand |
NeurIPS | 2 |
| 2023 | Gap Minimization for Knowledge Sharing and TransferabstractLearning from multiple related tasks by knowledge sharing and transfer has become increasingly relevant over the last two decades. In order to successfully transfer information from one task to another, it is critical to understand the similarities and differences between the domains. In this paper, we introduce the notion of performance gap, an intuitive and novel measure of the distance between learning tasks. Unlike existing measures which are used as tools to bound the difference of expected risks between tasks (e.g., $\mathcal{H}$-divergence or discrepancy distance), we theoretically show that the performance gap can be viewed as a data- and algorithm-dependent regularizer, which controls the model complexity and leads to finer guarantees. More importantly, it also provides new insights and motivates a novel principle for designing strategies for knowledge sharing and transfer: gap minimization. We instantiate this principle with two algorithms: 1. gapBoost, a novel and principled boosting algorithm that explicitly minimizes the performance gap between source and target domains for transfer learning; and 2. gapMTNN, a representation learning algorithm that reformulates gap minimization as semantic conditional matching for multitask learning. Our extensive evaluation on both transfer learning and multitask learning benchmark data sets shows that our methods outperform existing baselines. Boyu Wang 0004, Jorge A. Mendez, Changjian Shui, Fan Zhou 0006, Di Wu 0044, Gezheng Xu, Christian Gagné 0001, Eric Eaton |
J. Mach. Learn. Res. | 3 |
| 2023 | Label shift conditioned hybrid querying for deep active learning
Jiaqi Li 0005, Haojia Kong, Gezheng Xu, Changjian Shui, Ruizhi Pu, Zhao Kang 0001, Charles Ling 0001, Boyu Wang 0004 |
Knowl. Based Syst. | 4 |
| 2023 | Episodic task agnostic contrastive training for multi-task learning
Fan Zhou 0006, Yuyi Chen, Jun Wen 0001, Qiuhao Zeng, Changjian Shui, Charles Ling 0001, Boyu Wang 0004 |
Neural Networks | 5 |
| 2023 | Lifelong Online Learning from Accumulated KnowledgeabstractIn this article, we formulate lifelong learning as an online transfer learning procedure over consecutive tasks, where learning a given task depends on the accumulated knowledge. We propose a novel theoretical principled framework, lifelong online learning, where the learning process for each task is in an incremental manner. Specifically, our framework is composed of two-level predictions: the prediction information that is solely from the current task; and the prediction from the knowledge base by previous tasks. Moreover, this article tackled several fundamental challenges: arbitrary or even non-stationary task generation process, an unknown number of instances in each task, and constructing an efficient accumulated knowledge base. Notably, we provide a provable bound of the proposed algorithm, which offers insights on the how the accumulated knowledge improves the predictions. Finally, empirical evaluations on both synthetic and real datasets validate the effectiveness of the proposed algorithm. Changjian Shui, William Wei Wang, Ihsen Hedhli, Chiman Wong, Feng Wan 0003, Boyu Wang 0004, Christian Gagné 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | Towards More General Loss and Setting in Unsupervised Domain AdaptationabstractIn this article, we present an analysis of unsupervised domain adaptation with a series of theoretical and algorithmic results. We derive a novel Rényi-$\alpha$divergence-based generalization bound, which is tailored to domain adaptation algorithms with arbitrary loss functions in a stochastic setting. Moreover, our theoretical results provide new insights into the assumptions for successful domain adaptation: the closeness between the conditional distributions of the domains and the Lipschitzness on the source domain. With these assumptions, we reveal the following: if their conditional generation distributions are close, the Lipschitzness property of the target domain can be transferred from the Lipschitzness on the source domain, without knowing the exact target distribution. Motivated by our analysis and assumptions, we further derive practical principles for deep domain adaptation: 1) Rényi-2 adversarial training for marginal distributions matching and 2) Lipschitz regularization for the classifier. Our experimental results on both synthetic and real-world datasets support our theoretical findings and the practical efficiency of the proposed principles. Changjian Shui, Ruizhi Pu, Gezheng Xu, Jun Wen 0001, Fan Zhou 0006, Christian Gagné 0001, Charles Ling 0001, Boyu Wang 0004 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | On the Benefits of Two Dimensional Metric LearningabstractIn this paper, we study two dimensional metric learning (2DML) for matrix data from both theoretical and algorithmic perspectives. We first investigate the generalization bounds of 2DML based on the notion of Rademacher complexity, which theoretically justifies the benefits of learning from matrices directly. Furthermore, we present a novel boosting-based algorithm that scales well with the feature dimension. Finally, we introduce an efficient rank-one correction algorithm, which is tailored to our boosting learning procedure to produce a low-rank solution to 2DML. As our algorithm works directly on the data in matrix representation, it scales well with the feature dimension, keeps the structure and dependence in the data, and has a more compact structure and much fewer parameters to optimize. Extensive evaluations on several benchmark data sets also empirically verify the effectiveness and efficiency of our algorithm. Di Wu 0044, Fan Zhou 0006, Boyu Wang 0004, Qicheng Lao, Chiman Wong, Changjian Shui, Yuan Zhou 0006, Feng Wan 0003 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Fair Representation Learning through Implicit Path AlignmentabstractWe consider a fair representation learning perspective, where optimal predictors, on top of the data representation, are ensured to be invariant with respect to different sub-groups. Specifically, we formulate this intuition as a bi-level optimization, where the representation is learned in the outer-loop, and invariant optimal group predictors are updated in the inner-loop. Moreover, the proposed bi-level objective is demonstrated to fulfill the sufficiency rule, which is desirable in various practical scenarios but was not commonly studied in the fair learning. Besides, to avoid the high computational and memory cost of differentiating in the inner-loop of bi-level objective, we propose an implicit path alignment algorithm, which only relies on the solution of inner optimization and the implicit differentiation rather than the exact optimization path. We further analyze the error gap of the implicit approach and empirically validate the proposed method in both classification and regression settings. Experimental results show the consistently better trade-off in prediction performance and fairness measurement. Changjian Shui, Qi Chen 0015, Jiaqi Li 0005, Boyu Wang 0004, Christian Gagné 0001 |
ICML | 1 |
| 2022 | On Learning Fairness and Accuracy on Multiple SubgroupsabstractWe propose an analysis in fair learning that preserves the utility of the data while reducing prediction disparities under the criteria of group sufficiency. We focus on the scenario where the data contains multiple or even many subgroups, each with limited number of samples. As a result, we present a principled method for learning a fair predictor for all subgroups via formulating it as a bilevel objective. Specifically, the subgroup specific predictors are learned in the lower-level through a small amount of data and the fair predictor. In the upper-level, the fair predictor is updated to be close to all subgroup specific predictors. We further prove that such a bilevel objective can effectively control the group sufficiency and generalization error. We evaluate the proposed framework on real-world datasets. Empirical evidence suggests the consistently improved fair predictions, as well as the comparable accuracy to the baselines. Changjian Shui, Gezheng Xu, Qi Chen 0015, Jiaqi Li 0005, Charles Ling 0001, Tal Arbel, Boyu Wang 0004, Christian Gagné 0001 |
NeurIPS | 1 |
| 2022 | A novel domain adaptation theory with Jensen-Shannon divergence
Changjian Shui, Qi Chen 0015, Jun Wen 0001, Fan Zhou 0006, Christian Gagné 0001, Boyu Wang 0004 |
Knowl. Based Syst. | 1 |
| 2022 | On the benefits of representation regularization in invariance based domain generalizationabstractA crucial aspect of reliable machine learning is to design a deployable system for generalizing new related but unobserved environments. Domain generalization aims to alleviate such a prediction gap between the observed and unseen environments. Previous approaches commonly incorporated learning the invariant representation for achieving good empirical performance. In this paper, we reveal that merely learning the invariant representation is vulnerable to the related unseen environment. To this end, we derive a novel theoretical analysis to control the unseen test environment error in the representation learning, which highlights the importance of controlling the smoothness of representation. In practice, our analysis further inspires an efficient regularization method to improve the robustness in domain generalization. The proposed regularization is orthogonal to and can be straightforwardly adopted in existing domain generalization algorithms that ensure invariant representation learning. Empirical results show that our algorithm outperforms the base versions in various datasets and invariance criteria. Changjian Shui, Boyu Wang 0004, Christian Gagné 0001 |
Mach. Learn. | 1 |
| 2021 | Aggregating From Multiple Target-Shifted SourcesabstractMulti-source domain adaptation aims at leveraging the knowledge from multiple tasks for predicting a related target domain. Hence, a crucial aspect is to properly combine different sources based on their relations. In this paper, we analyzed the problem for aggregating source domains with different label distributions, where most recent source selection approaches fail. Our proposed algorithm differs from previous approaches in two key ways: the model aggregates multiple sources mainly through the similarity of semantic conditional distribution rather than marginal distribution; the model proposes a unified framework to select relevant sources for three popular scenarios, i.e., domain adaptation with limited label on target domain, unsupervised domain adaptation and label partial unsupervised domain adaption. We evaluate the proposed method through extensive experiments. The empirical results significantly outperform the baselines. Changjian Shui, Zijian Li 0001, Jiaqi Li 0005, Christian Gagné 0001, Charles Ling 0001, Boyu Wang 0004 |
ICML | 1 |
| 2021 | Generalization Bounds For Meta-Learning: An Information-Theoretic AnalysisabstractWe derive a novel information-theoretic analysis of the generalization property of meta-learning algorithms. Concretely, our analysis proposes a generic understanding in both the conventional learning-to-learn framework \citep{amit2018meta} and the modern model-agnostic meta-learning (MAML) algorithms \citep{finn2017model}.Moreover, we provide a data-dependent generalization bound for the stochastic variant of MAML, which is \emph{non-vacuous} for deep few-shot learning. As compared to previous bounds that depend on the square norms of gradients, empirical validations on both simulated data and a well-known few-shot benchmark show that our bound is orders of magnitude tighter in most conditions. Qi Chen 0015, Changjian Shui, Mario Marchand |
NeurIPS | 2 |
| 2021 | Domain generalization via optimal transport with metric similarity learning
Fan Zhou 0006, Zhuqing Jiang, Changjian Shui, Boyu Wang 0004, Brahim Chaib-draa |
Neurocomputing | 3 |
| 2021 | Discriminative active learning for domain adaptation
Fan Zhou 0006, Changjian Shui, Bincheng Huang, Boyu Wang 0004, Brahim Chaib-draa |
Knowl. Based Syst. | 2 |
| 2021 | Common Spatial Pattern Reformulated for Regularizations in Brain-Computer InterfacesabstractCommon spatial pattern (CSP) is one of the most successful feature extraction algorithms for brain-computer interfaces (BCIs). It aims to find spatial filters that maximize the projected variance ratio between the covariance matrices of the multichannel electroencephalography (EEG) signals corresponding to two mental tasks, which can be formulated as a generalized eigenvalue problem (GEP). However, it is challenging in principle to impose additional regularization onto the CSP to obtain structural solutions (e.g., sparse CSP) due to the intrinsic nonconvexity and invariance property of GEPs. This article reformulates the CSP as a constrained minimization problem and establishes the equivalence of the reformulated and the original CSPs. An efficient algorithm is proposed to solve this optimization problem by alternately performing singular value decomposition (SVD) and least squares. Under this new formulation, various regularization techniques for linear regression can then be easily implemented to regularize the CSPs for different learning paradigms, such as the sparse CSP, the transfer CSP, and the multisubject CSP. Evaluations on three BCI competition datasets show that the regularized CSP algorithms outperform other baselines, especially for the high-dimensional small training set. The extensive results validate the efficiency and effectiveness of the proposed CSP formulation in different learning contexts. Boyu Wang 0004, Chiman Wong, Zhao Kang 0001, Feng Liu 0011, Changjian Shui, Feng Wan 0003, C. L. Philip Chen |
IEEE Trans. Cybern. | 5 |
| 2021 | Task Similarity Estimation Through Adversarial Multitask Neural NetworkabstractMultitask learning (MTL) aims at solving the related tasks simultaneously by exploiting shared knowledge to improve performance on individual tasks. Though numerous empirical results supported the notion that such shared knowledge among tasks plays an essential role in MTL, the theoretical understanding of the relationships between tasks and their impact on learning shared knowledge is still an open problem. In this work, we are developing a theoretical perspective of the benefits involved in using information similarity for MTL. To this end, we first propose an upper bound on the generalization error by implementing the Wasserstein distance as the similarity metric. This indicates the practical principles of applying the similarity information to control the generalization errors. Based on those theoretical results, we revisited the adversarial multitask neural network and proposed a new training algorithm to learn the task relation coefficients and neural network parameters automatically. The computer vision benchmarks reveal the abilities of the proposed algorithms to improve the empirical performance. Finally, we test the proposed approach on real medical data sets, showing its advantage for extracting task relations. Fan Zhou 0006, Changjian Shui, Mahdieh Abbasi, Louis-Émile Robitaille, Boyu Wang 0004, Christian Gagné 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Deep Active Learning: Unified and Principled Method for Query and TrainingabstractIn this paper, we are proposing a unified and principled method for both the querying and training processes in deep batch active learning. We are providing theoretical insights from the intuition of modeling the interactive procedure in active learning as distribution matching, by adopting the Wasserstein distance. As a consequence, we derived a new training loss from the theoretical analysis, which is decomposed into optimizing deep neural network parameters and batch query selection through alternative optimization. In addition, the loss for training a deep neural network is naturally formulated as a min-max optimization problem through leveraging the unlabeled data information. Moreover, the proposed principles also indicate an explicit uncertainty-diversity trade-off in the query batch selection. Finally, we evaluate our proposed method on different benchmarks, consistently showing better empirical performances and a better time-efficient query strategy compared to the baselines. Changjian Shui, Fan Zhou 0006, Christian Gagné 0001, Boyu Wang 0004 |
AISTATS | 1 |
| 2020 | Toward Metrics for Differentiating Out-of-Distribution SetsabstractVanilla CNNs, as uncalibrated classifiers, suffer from classifying out-of-distribution (OOD) samples nearly as confidently as in-distribution samples. To tackle this challenge, some recent works have demonstrated the gains of leveraging available OOD sets for training end-to-end calibrated CNNs. However, a critical question remains unanswered in these works: how to differentiate OOD sets for selecting the most effective one(s) that induce training such CNNs with high detection rates on unseen OOD sets? To address this pivotal question, we provide a criterion based on generalization errors of Augmented-CNN, a vanilla CNN with an added extra class employed for rejection, on in-distribution and unseen OOD sets. However, selecting the most effective OOD set by directly optimizing this criterion incurs a huge computational cost. Instead, we propose three novel computationally-efficient metrics for differentiating between OOD sets according to their level of in-distribution sub-manifolds. We empirically verify that the most protective OOD sets -- selected according to our metrics -- lead to A-CNNs with significantly lower generalization errors than the A-CNNs trained on the least protective ones. We also empirically show the effectiveness of a protective OOD set for training well-generalized confidence-calibrated vanilla CNNs. These results confirm that 1) all OOD sets are not equally effective for training well-performing end-to-end models (i.e., A-CNNs and calibrated CNNs) for OOD detection tasks and 2) the protection level of OOD sets is a viable factor for recognizing the most effective one. Finally, across the image classification tasks, we exhibit A-CNN trained on the most protective OOD set can also detect black-box FGS adversarial examples as their distance (measured by our metrics) is becoming larger from the protected sub-manifolds. Mahdieh Abbasi, Changjian Shui, Arezoo Rajabi, Christian Gagné 0001, Rakesh Bobba |
ECAI | 2 |
| 2019 | A Principled Approach for Learning Task Similarity in Multitask LearningabstractMultitask learning aims at solving a set of related tasks simultaneously, by exploiting the shared knowledge for improving the performance on individual tasks. Hence, an important aspect of multitask learning is to understand the similarities within a set of tasks. Previous works have incorporated this similarity information explicitly (e.g., weighted loss for each task) or implicitly (e.g., adversarial loss for feature adaptation), for achieving good empirical performances. However, the theoretical motivations for adding task similarity knowledge are often missing or incomplete. In this paper, we give a different perspective from a theoretical point of view to understand this practice. We first provide an upper bound on the generalization error of multitask learning, showing the benefit of explicit and implicit task similarity knowledge. We systematically derive the bounds based on two distinct task similarity metrics: H divergence and Wasserstein distance. From these theoretical results, we revisit the Adversarial Multi-task Neural Network, proposing a new training algorithm to learn the task relation coefficients and neural network parameters iteratively. We assess our new algorithm empirically on several benchmarks, showing not only that we find interesting and robust task relations, but that the proposed approach outperforms the baselines, reaffirming the benefits of theoretical insight in algorithm design. Changjian Shui, Mahdieh Abbasi, Louis-Émile Robitaille, Boyu Wang 0004, Christian Gagné 0001 |
IJCAI | 1 |