Christian Gagné 0001

dblp:80/5084-1 · DBLP profile ↗
← Back
57ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0003-3697-4184ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 A Layer Selection Approach to Test Time Adaptation
abstract
Test Time Adaptation (TTA) addresses the problem of distribution shift by adapting a pretrained model to a new domain during inference. When faced with challenging shifts, most methods collapse and perform worse than the original pretrained model. In this paper, we find that not all layers are equally receptive to the adaptation, and the layers with the most misaligned gradients often cause performance degradation. To address this, we propose GALA, a novel layer selection criterion to identify the most beneficial updates to perform during test time adaptation. This criterion can also filter out unreliable samples with noisy gradients. Its simplicity allows seamless integration with existing TTA loss functions, thereby preventing degradation and focusing adaptation on the most trainable layers. This approach also helps to regularize adaptation to preserve the pretrained features, which are crucial for handling unseen domains. Through extensive experiments, we demonstrate that the proposed layer selection framework improves the performance of existing TTA approaches across multiple datasets, domain shifts, model architectures, and TTA losses.
Sabyasachi Sahoo, Mostafa ElAraby, Jonas Ngnawé, Yann Pequignot, Frédéric Precioso, Christian Gagné 0001
AAAI6
2024 Generalizing across Temporal Domains with Koopman Operators
abstract
In the field of domain generalization, the task of constructing a predictive model capable of generalizing to a target domain without access to target data remains challenging. This problem becomes further complicated when considering evolving dynamics between domains. While various approaches have been proposed to address this issue, a comprehensive understanding of the underlying generalization theory is still lacking. In this study, we contribute novel theoretic results that aligning conditional distribution leads to the reduction of generalization bounds. Our analysis serves as a key motivation for solving the Temporal Domain Generalization (TDG) problem through the application of Koopman Neural Operators, resulting in Temporal Koopman Networks (TKNets). By employing Koopman Neural Operators, we effectively address the time-evolving distributions encountered in TDG using the principles of Koopman theory, where measurement functions are sought to establish linear transition relations between evolving domains. Through empirical evaluations conducted on synthetic and real-world datasets, we validate the effectiveness of our proposed approach.
Qiuhao Zeng, Wei Wang 0036, Fan Zhou 0006, Gezheng Xu, Ruizhi Pu, Changjian Shui, Christian Gagné 0001, Charles Ling 0001, Boyu Wang 0004
AAAI7
2024 Detecting Brittle Decisions for Free: Leveraging Margin Consistency in Deep Robust Classifiers
abstract
Despite extensive research on adversarial training strategies to improve robustness, the decisions of even the most robust deep learning models can still be quite sensitive to imperceptible perturbations, creating serious risks when deploying them for high-stakes real-world applications. While detecting such cases may be critical, evaluating a model's vulnerability at a per-instance level using adversarial attacks is computationally too intensive and unsuitable for real-time deployment scenarios. The input space margin is the exact score to detect non-robust samples and is intractable for deep neural networks. This paper introduces the concept of margin consistency -- a property that links the input space margins and the logit margins in robust models -- for efficient detection of vulnerable samples. First, we establish that margin consistency is a necessary and sufficient condition to use a model's logit margin as a score for identifying non-robust samples. Next, through comprehensive empirical analysis of various robustly trained models on CIFAR10 and CIFAR100 datasets, we show that they indicate high margin consistency with a strong correlation between their input space margins and the logit margins. Then, we show that we can effectively use the logit margin to confidently detect brittle decisions with such models. Finally, we address cases where the model is not sufficiently margin-consistent by learning a pseudo-margin from the feature representation. Our findings highlight the potential of leveraging deep representations to efficiently assess adversarial vulnerability in deployment scenarios.
Jonas Ngnawé, Sabyasachi Sahoo, Yann Pequignot, Frédéric Precioso, Christian Gagné 0001
NeurIPS5
2024 Hessian Aware Low-Rank Perturbation for Order-Robust Continual Learning
abstract
Continual learning aims to learn a series of tasks sequentially without forgetting the knowledge acquired from the previous ones. In this work, we propose the Hessian Aware Low-Rank Perturbation algorithm for continual learning. By modeling the parameter transitions along the sequential tasks with the weight matrix transformation, we propose to apply the low-rank approximation on the task-adaptive parameters in each layer of the neural networks. Specifically, we theoretically demonstrate the quantitative relationship between the Hessian and the proposed low-rank approximation. The approximation ranks are then globally determined according to the marginal change of the empirical loss estimated by the layer-specific gradient and low-rank approximation error. Furthermore, we control the model capacity by pruning less important parameters to diminish the parameter growth. We conduct extensive experiments on various benchmarks, including a dataset with large-scale tasks, and compare our method against some recent state-of-the-art methods to demonstrate the effectiveness and scalability of our proposed method. Empirical results show that our method performs better on different benchmarks, especially in achieving task order robustness and handling the forgetting issue.
Jiaqi Li 0005, Yuanhao Lai, Rui Wang 0121, Changjian Shui, Sabyasachi Sahoo, Charles Ling 0001, Boyu Wang 0004, Christian Gagné 0001, Fan Zhou 0006
IEEE Trans. Knowl. Data Eng.9
2023 Gap Minimization for Knowledge Sharing and Transfer
abstract
Learning from multiple related tasks by knowledge sharing and transfer has become increasingly relevant over the last two decades. In order to successfully transfer information from one task to another, it is critical to understand the similarities and differences between the domains. In this paper, we introduce the notion of performance gap, an intuitive and novel measure of the distance between learning tasks. Unlike existing measures which are used as tools to bound the difference of expected risks between tasks (e.g., $\mathcal{H}$-divergence or discrepancy distance), we theoretically show that the performance gap can be viewed as a data- and algorithm-dependent regularizer, which controls the model complexity and leads to finer guarantees. More importantly, it also provides new insights and motivates a novel principle for designing strategies for knowledge sharing and transfer: gap minimization. We instantiate this principle with two algorithms: 1. gapBoost, a novel and principled boosting algorithm that explicitly minimizes the performance gap between source and target domains for transfer learning; and 2. gapMTNN, a representation learning algorithm that reformulates gap minimization as semantic conditional matching for multitask learning. Our extensive evaluation on both transfer learning and multitask learning benchmark data sets shows that our methods outperform existing baselines.
Boyu Wang 0004, Jorge A. Mendez, Changjian Shui, Fan Zhou 0006, Di Wu 0044, Gezheng Xu, Christian Gagné 0001, Eric Eaton
J. Mach. Learn. Res.7
2023 Lifelong Online Learning from Accumulated Knowledge
abstract
In this article, we formulate lifelong learning as an online transfer learning procedure over consecutive tasks, where learning a given task depends on the accumulated knowledge. We propose a novel theoretical principled framework, lifelong online learning, where the learning process for each task is in an incremental manner. Specifically, our framework is composed of two-level predictions: the prediction information that is solely from the current task; and the prediction from the knowledge base by previous tasks. Moreover, this article tackled several fundamental challenges: arbitrary or even non-stationary task generation process, an unknown number of instances in each task, and constructing an efficient accumulated knowledge base. Notably, we provide a provable bound of the proposed algorithm, which offers insights on the how the accumulated knowledge improves the predictions. Finally, empirical evaluations on both synthetic and real datasets validate the effectiveness of the proposed algorithm.
Changjian Shui, William Wei Wang, Ihsen Hedhli, Chiman Wong, Feng Wan 0003, Boyu Wang 0004, Christian Gagné 0001
ACM Trans. Knowl. Discov. Data7
2023 Towards More General Loss and Setting in Unsupervised Domain Adaptation
abstract
In this article, we present an analysis of unsupervised domain adaptation with a series of theoretical and algorithmic results. We derive a novel Rényi-$\alpha$divergence-based generalization bound, which is tailored to domain adaptation algorithms with arbitrary loss functions in a stochastic setting. Moreover, our theoretical results provide new insights into the assumptions for successful domain adaptation: the closeness between the conditional distributions of the domains and the Lipschitzness on the source domain. With these assumptions, we reveal the following: if their conditional generation distributions are close, the Lipschitzness property of the target domain can be transferred from the Lipschitzness on the source domain, without knowing the exact target distribution. Motivated by our analysis and assumptions, we further derive practical principles for deep domain adaptation: 1) Rényi-2 adversarial training for marginal distributions matching and 2) Lipschitz regularization for the classifier. Our experimental results on both synthetic and real-world datasets support our theoretical findings and the practical efficiency of the proposed principles.
Changjian Shui, Ruizhi Pu, Gezheng Xu, Jun Wen 0001, Fan Zhou 0006, Christian Gagné 0001, Charles Ling 0001, Boyu Wang 0004
IEEE Trans. Knowl. Data Eng.6
2022 Matching Feature Sets for Few-Shot Image Classification
abstract
In image classification, it is common practice to train deep networks to extract a single feature vector per input image. Few-shot classification methods also mostly follow this trend. In this work, we depart from this established direction and instead propose to extract sets of feature vectors for each image. We argue that a set-based representation intrinsically builds a richer representation of images from the base classes, which can subsequently better transfer to the few-shot classes. To do so, we propose to adapt existing feature extractors to instead produce sets of feature vectors from images. Our approach, dubbed SetFeat, embeds shallow self-attention mechanisms inside existing encoder architectures. The attention modules are lightweight, and as such our method results in encoders that have approximately the same number of parameters as their original versions. During training and inference, a set-to-set matching metric is used to perform image classification. The effectiveness of our proposed architecture and metrics is demonstrated via thorough experiments on standard few-shot datasets-namely miniImageNet, tieredImageNet, and CUB-in both the 1- and 5-shot scenarios. In all cases but one, our method outperforms the state-of-the-art.
Arman Afrasiyabi, Hugo Larochelle, Jean-François Lalonde, Christian Gagné 0001
CVPR4
2022 Fair Representation Learning through Implicit Path Alignment
abstract
We consider a fair representation learning perspective, where optimal predictors, on top of the data representation, are ensured to be invariant with respect to different sub-groups. Specifically, we formulate this intuition as a bi-level optimization, where the representation is learned in the outer-loop, and invariant optimal group predictors are updated in the inner-loop. Moreover, the proposed bi-level objective is demonstrated to fulfill the sufficiency rule, which is desirable in various practical scenarios but was not commonly studied in the fair learning. Besides, to avoid the high computational and memory cost of differentiating in the inner-loop of bi-level objective, we propose an implicit path alignment algorithm, which only relies on the solution of inner optimization and the implicit differentiation rather than the exact optimization path. We further analyze the error gap of the implicit approach and empirically validate the proposed method in both classification and regression settings. Experimental results show the consistently better trade-off in prediction performance and fairness measurement.
Changjian Shui, Qi Chen 0015, Jiaqi Li 0005, Boyu Wang 0004, Christian Gagné 0001
ICML5
2022 On Learning Fairness and Accuracy on Multiple Subgroups
abstract
We propose an analysis in fair learning that preserves the utility of the data while reducing prediction disparities under the criteria of group sufficiency. We focus on the scenario where the data contains multiple or even many subgroups, each with limited number of samples. As a result, we present a principled method for learning a fair predictor for all subgroups via formulating it as a bilevel objective. Specifically, the subgroup specific predictors are learned in the lower-level through a small amount of data and the fair predictor. In the upper-level, the fair predictor is updated to be close to all subgroup specific predictors. We further prove that such a bilevel objective can effectively control the group sufficiency and generalization error. We evaluate the proposed framework on real-world datasets. Empirical evidence suggests the consistently improved fair predictions, as well as the comparable accuracy to the baselines.
Changjian Shui, Gezheng Xu, Qi Chen 0015, Jiaqi Li 0005, Charles Ling 0001, Tal Arbel, Boyu Wang 0004, Christian Gagné 0001
NeurIPS8
2022 A novel domain adaptation theory with Jensen-Shannon divergence
Changjian Shui, Qi Chen 0015, Jun Wen 0001, Fan Zhou 0006, Christian Gagné 0001, Boyu Wang 0004
Knowl. Based Syst.5
2022 On the benefits of representation regularization in invariance based domain generalization
abstract
A crucial aspect of reliable machine learning is to design a deployable system for generalizing new related but unobserved environments. Domain generalization aims to alleviate such a prediction gap between the observed and unseen environments. Previous approaches commonly incorporated learning the invariant representation for achieving good empirical performance. In this paper, we reveal that merely learning the invariant representation is vulnerable to the related unseen environment. To this end, we derive a novel theoretical analysis to control the unseen test environment error in the representation learning, which highlights the importance of controlling the smoothness of representation. In practice, our analysis further inspires an efficient regularization method to improve the robustness in domain generalization. The proposed regularization is orthogonal to and can be straightforwardly adopted in existing domain generalization algorithms that ensure invariant representation learning. Empirical results show that our algorithm outperforms the base versions in various datasets and invariance criteria.
Changjian Shui, Boyu Wang 0004, Christian Gagné 0001
Mach. Learn.3
2021 Mixture-based Feature Space Learning for Few-shot Image Classification
abstract
We introduce Mixture-based Feature Space Learning (MixtFSL) for obtaining a rich and robust feature representation in the context of few-shot image classification. Previous works have proposed to model each base class either with a single point or with a mixture model by relying on offline clustering algorithms. In contrast, we propose to model base classes with mixture models by simultaneously training the feature extractor and learning the mixture model parameters in an online manner. This results in a richer and more discriminative feature space which can be employed to classify novel examples from very few samples. Two main stages are proposed to train the MixtFSL model. First, the multimodal mixtures for each base class and the feature extractor parameters are learned using a combination of two loss functions. Second, the resulting network and mixture models are progressively refined through a leader-follower learning procedure, which uses the current estimate as a "target" network. This target network is used to make a consistent assignment of instances to mixture components, which increases performance and stabilizes training. The effectiveness of our end-to-end feature space learning approach is demonstrated with extensive experiments on four standard datasets and four backbones. Notably, we demon-strate that when we combine our robust representation with recent alignment-based approaches, we achieve new state-of-the-art results in the inductive setting, with an absolute accuracy for 5-shot classification of 82.45% on miniImageNet, 88.20% with tieredImageNet, and 60.70% in FC100 using the ResNet-12 backbone.
Arman Afrasiyabi, Jean-François Lalonde, Christian Gagné 0001
ICCV3
2021 Aggregating From Multiple Target-Shifted Sources
abstract
Multi-source domain adaptation aims at leveraging the knowledge from multiple tasks for predicting a related target domain. Hence, a crucial aspect is to properly combine different sources based on their relations. In this paper, we analyzed the problem for aggregating source domains with different label distributions, where most recent source selection approaches fail. Our proposed algorithm differs from previous approaches in two key ways: the model aggregates multiple sources mainly through the similarity of semantic conditional distribution rather than marginal distribution; the model proposes a unified framework to select relevant sources for three popular scenarios, i.e., domain adaptation with limited label on target domain, unsupervised domain adaptation and label partial unsupervised domain adaption. We evaluate the proposed method through extensive experiments. The empirical results significantly outperform the baselines.
Changjian Shui, Zijian Li 0001, Jiaqi Li 0005, Christian Gagné 0001, Charles Ling 0001, Boyu Wang 0004
ICML4
2021 Task Similarity Estimation Through Adversarial Multitask Neural Network
abstract
Multitask learning (MTL) aims at solving the related tasks simultaneously by exploiting shared knowledge to improve performance on individual tasks. Though numerous empirical results supported the notion that such shared knowledge among tasks plays an essential role in MTL, the theoretical understanding of the relationships between tasks and their impact on learning shared knowledge is still an open problem. In this work, we are developing a theoretical perspective of the benefits involved in using information similarity for MTL. To this end, we first propose an upper bound on the generalization error by implementing the Wasserstein distance as the similarity metric. This indicates the practical principles of applying the similarity information to control the generalization errors. Based on those theoretical results, we revisited the adversarial multitask neural network and proposed a new training algorithm to learn the task relation coefficients and neural network parameters automatically. The computer vision benchmarks reveal the abilities of the proposed algorithms to improve the empirical performance. Finally, we test the proposed approach on real medical data sets, showing its advantage for extracting task relations.
Fan Zhou 0006, Changjian Shui, Mahdieh Abbasi, Louis-Émile Robitaille, Boyu Wang 0004, Christian Gagné 0001
IEEE Trans. Neural Networks Learn. Syst.6
2020 Deep Active Learning: Unified and Principled Method for Query and Training
abstract
In this paper, we are proposing a unified and principled method for both the querying and training processes in deep batch active learning. We are providing theoretical insights from the intuition of modeling the interactive procedure in active learning as distribution matching, by adopting the Wasserstein distance. As a consequence, we derived a new training loss from the theoretical analysis, which is decomposed into optimizing deep neural network parameters and batch query selection through alternative optimization. In addition, the loss for training a deep neural network is naturally formulated as a min-max optimization problem through leveraging the unlabeled data information. Moreover, the proposed principles also indicate an explicit uncertainty-diversity trade-off in the query batch selection. Finally, we evaluate our proposed method on different benchmarks, consistently showing better empirical performances and a better time-efficient query strategy compared to the baselines.
Changjian Shui, Fan Zhou 0006, Christian Gagné 0001, Boyu Wang 0004
AISTATS3
2020 Toward Metrics for Differentiating Out-of-Distribution Sets
abstract
Vanilla CNNs, as uncalibrated classifiers, suffer from classifying out-of-distribution (OOD) samples nearly as confidently as in-distribution samples. To tackle this challenge, some recent works have demonstrated the gains of leveraging available OOD sets for training end-to-end calibrated CNNs. However, a critical question remains unanswered in these works: how to differentiate OOD sets for selecting the most effective one(s) that induce training such CNNs with high detection rates on unseen OOD sets? To address this pivotal question, we provide a criterion based on generalization errors of Augmented-CNN, a vanilla CNN with an added extra class employed for rejection, on in-distribution and unseen OOD sets. However, selecting the most effective OOD set by directly optimizing this criterion incurs a huge computational cost. Instead, we propose three novel computationally-efficient metrics for differentiating between OOD sets according to their level of in-distribution sub-manifolds. We empirically verify that the most protective OOD sets -- selected according to our metrics -- lead to A-CNNs with significantly lower generalization errors than the A-CNNs trained on the least protective ones. We also empirically show the effectiveness of a protective OOD set for training well-generalized confidence-calibrated vanilla CNNs. These results confirm that 1) all OOD sets are not equally effective for training well-performing end-to-end models (i.e., A-CNNs and calibrated CNNs) for OOD detection tasks and 2) the protection level of OOD sets is a viable factor for recognizing the most effective one. Finally, across the image classification tasks, we exhibit A-CNN trained on the most protective OOD set can also detect black-box FGS adversarial examples as their distance (measured by our metrics) is becoming larger from the protected sub-manifolds.
Mahdieh Abbasi, Changjian Shui, Arezoo Rajabi, Christian Gagné 0001, Rakesh Bobba
ECAI4
2020 Associative Alignment for Few-Shot Image Classification
Arman Afrasiyabi, Jean-François Lalonde, Christian Gagné 0001
ECCV (5)3
2020 Input Dropout for Spatially Aligned Modalities
abstract
Computer vision datasets containing multiple modalities such as color, depth, and thermal properties are now commonly accessible and useful for solving a wide array of challenging tasks. However, deploying multi-sensor heads is not possible in many scenarios. As such many practical solutions tend to be based on simpler sensors, mostly for cost, simplicity and robustness considerations. In this work, we propose a training methodology to take advantage of these additional modalities available in datasets, even if they are not available at test time. By assuming that the modalities have a strong spatial correlation, we propose Input Dropout, a simple technique that consists in stochastic hiding of one or many input modalities at training time, while using only the canonical (e.g. RGB) modalities at test time. We demonstrate that Input Dropout trivially combines with existing deep convolutional architectures, and improves their performance on a wide range of computer vision tasks such as dehazing, 6-DOF object tracking, pedestrian detection and object classification.
Sébastien de Blois, Mathieu Garon, Christian Gagné 0001, Jean-François Lalonde
ICIP3
2020 The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
abstract
Evolution provides a creative fount of complex and subtle adaptations that often surprise the scientists who discover them. However, the creativity of evolution is not limited to the natural world: Artificial organisms evolving in computational environments have also elicited surprise and wonder from the researchers studying them. The process of evolution is an algorithmic process that transcends the substrate in which it occurs. Indeed, many researchers in the field of digital evolution can provide examples of how their evolving algorithms and organisms have creatively subverted their expectations or intentions, exposed unrecognized bugs in their code, produced unexpectedly adaptations, or engaged in behaviors and outcomes, uncannily convergent with ones found in nature. Such stories routinely reveal surprise and creativity by evolution in these digital worlds, but they rarely fit into the standard scientific narrative. Instead they are often treated as mere obstacles to be overcome, rather than results that warrant study in their own right. Bugs are fixed, experiments are refocused, and one-off surprises are collapsed into a single data point. The stories themselves are traded among researchers through oral tradition, but that mode of information transmission is inefficient and prone to error and outright loss. Moreover, the fact that these stories tend to be shared only among practitioners means that many natural scientists do not realize how interesting and lifelike digital organisms are and how natural their evolution can be. To our knowledge, no collection of such anecdotes has been published before. This article is the crowd-sourced product of researchers in the fields of artificial life and evolutionary computation who have provided first-hand accounts of such cases. It thus serves as a written, fact-checked collection of scientifically important and even entertaining stories. In doing so we also present here substantial evidence that the existence and importance of evolutionary surprises extends beyond the natural world, and may indeed be a universal property of all complex evolving systems.
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J. Bentley, Samuel Bernard, Guillaume Beslon, David M. Bryson, Nicholas Cheney, Patryk Chrabaszcz, Antoine Cully, Stéphane Doncieux, Fred C. Dyer, Kai Olav Ellefsen, Robert Feldt, Stephan Fischer 0002, Stephanie Forrest, Antoine Frénoy, Christian Gagné 0001, Leni K. Le Goff, Laura M. Grabowski, Babak Hodjat, Frank Hutter, Laurent Keller, Carole Knibbe, Peter Krcah, Richard E. Lenski, Hod Lipson, Robert MacCurdy, Carlos Maestre, Risto Miikkulainen, Sara Mitri, David E. Moriarty, Jean-Baptiste Mouret, Anh Totti Nguyen, Charles Ofria, Marc Parizeau, David P. Parsons, Robert T. Pennock, William F. Punch, Thomas S. Ray, Marc Schoenauer, Eric Schulte, Karl Sims, Kenneth O. Stanley, François Taddei, Danesh Tarapore, Simon Thibault, Richard A. Watson, Westley Weimer, Jason Yosinski
Artif. Life21
2019 Deep Parametric Indoor Lighting Estimation
abstract
We present a method to estimate lighting from a single image of an indoor scene. Previous work has used an environment map representation that does not account for the localized nature of indoor lighting. Instead, we represent lighting as a set of discrete 3D lights with geometric and photometric parameters. We train a deep neural network to regress these parameters from a single image, on a dataset of environment maps annotated with depth. We propose a differentiable layer to convert these parameters to an environment map to compute our loss; this bypasses the challenge of establishing correspondences between estimated and ground truth lights. We demonstrate, via quantitative and qualitative evaluations, that our representation and training scheme lead to more accurate results compared to previous work, while allowing for more realistic 3D object compositing with spatially-varying lighting.
Marc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagné 0001, Jean-François Lalonde
ICCV4
2019 A Principled Approach for Learning Task Similarity in Multitask Learning
abstract
Multitask learning aims at solving a set of related tasks simultaneously, by exploiting the shared knowledge for improving the performance on individual tasks. Hence, an important aspect of multitask learning is to understand the similarities within a set of tasks. Previous works have incorporated this similarity information explicitly (e.g., weighted loss for each task) or implicitly (e.g., adversarial loss for feature adaptation), for achieving good empirical performances. However, the theoretical motivations for adding task similarity knowledge are often missing or incomplete. In this paper, we give a different perspective from a theoretical point of view to understand this practice. We first provide an upper bound on the generalization error of multitask learning, showing the benefit of explicit and implicit task similarity knowledge. We systematically derive the bounds based on two distinct task similarity metrics: H divergence and Wasserstein distance. From these theoretical results, we revisit the Adversarial Multi-task Neural Network, proposing a new training algorithm to learn the task relation coefficients and neural network parameters iteratively. We assess our new algorithm empirically on several benchmarks, showing not only that we find interesting and robust task relations, but that the proposed approach outperforms the baselines, reaffirming the benefits of theoretical insight in algorithm design.
Changjian Shui, Mahdieh Abbasi, Louis-Émile Robitaille, Boyu Wang 0004, Christian Gagné 0001
IJCAI5
2018 Learning to Become an Expert: Deep Networks Applied to Super-Resolution Microscopy
abstract
With super-resolution optical microscopy, it is now possible to observe molecular interactions in living cells. The obtained images have a very high spatial precision but their overall quality can vary a lot depending on the structure of interest and the imaging parameters. Moreover, evaluating this quality is often difficult for non-expert users. In this work, we tackle the problem of learning the quality function of super-resolution images from scores provided by experts. More specifically, we are proposing a system based on a deep neural network that can provide a quantitative quality measure of a STED image of neuronal structures given as input. We conduct a user study in order to evaluate the quality of the predictions of the neural network against those of a human expert. Results show the potential while highlighting some of the limits of the proposed approach.
Louis-Émile Robitaille, Audrey Durand, Marc-André Gardner, Christian Gagné 0001, Paul De Koninck, Flavie Lavoie-Cardinal
AAAI4
2018 Rating Super-Resolution Microscopy Images With Deep Learning
Louis-Émile Robitaille, Audrey Durand, Marc-André Gardner, Christian Gagné 0001, Paul De Koninck, Flavie Lavoie-Cardinal
AAAI4
2018 A two-step approach for mining patient treatment pathways in administrative healthcare databases
Ahmed Najjar, Daniel Reinharz, Catherine Girouard, Christian Gagné 0001
Artif. Intell. Medicine4
2017 Bayesian optimization for conditional hyperparameter spaces
abstract
Hyperparameter optimization is now widely applied to tune the hyperparameters of learning algorithms. The hyperparameters can have structure, resulting in hyperparameters depending on conditions, or on the values of other hyperparameters. We target the problem of combined algorithm selection and hyperparameter optimization, which includes at least one conditional hyperparameter: the choice of the learning algorithm. In this work, we show that Bayesian optimization with Gaussian processes can be used for the optimization of conditional spaces with the injection of knowledge concerning conditions in the kernel. We propose and examine the behavior of two kernels, a conditional kernel which forces the similarity of two samples from different condition branches to be zero, and the Laplace kernel, based on similarities with Mondrian processes and random forests. We show the benefit of using such kernels, as well as proper imputation of inactive hyperparameters, on a benchmark of scikit-learn models.
Julien-Charles Levesque, Audrey Durand, Christian Gagné 0001, Robert Sabourin
IJCNN3
2017 Learning to predict indoor illumination from a single image
abstract
We propose an automatic method to infer high dynamic range illumination from a single, limited field-of-view, low dynamic range photograph of an indoor scene. In contrast to previous work that relies on specialized image capture, user input, and/or simple scene models, we train an end-to-end deep neural network that directly regresses a limited field-of-view photo to HDR illumination, without strong assumptions on scene geometry, material properties, or lighting. We show that this can be accomplished in a three step process: 1) we train a robust lighting classifier to automatically annotate the location of light sources in a large dataset of LDR environment maps, 2) we use these annotations to train a deep neural network that predicts the location of lights in a scene from a single limited field-of-view photo, and 3) we fine-tune this network using a small dataset of HDR environment maps to predict light intensities. This allows us to automatically recover high-quality HDR illumination estimates that significantly outperform previous state-of-the-art methods. Consequently, using our illumination estimates for applications like 3D object insertion, produces photo-realistic results that we validate via a perceptual user study.
Marc-André Gardner, Kalyan Sunkavalli, Ersin Yumer, Xiaohui Shen, Emiliano Gambaretto, Christian Gagné 0001, Jean-François Lalonde
ACM Trans. Graph.6
2016 Stream clustering of tweets
abstract
This paper proposes an approach to cluster social media posts. It aims at taking full advantage of this recent source of newsworthy information and at facilitating the work of users who need to monitor public events in real-time. The emphasis is on developing a stream clustering algorithm able to process incoming tweets. A first implementation of the algorithm, focusing on the tweets' text, was tuned and tested on a dataset of manually annotated messages. Results show that the algorithm produces a partition of tweets similar to the manual partition obtained from humans. In future work, we plan to extend this algorithm with additional features and integrate the resulting analytical capabilities to a real-time social media monitoring platform called CrowdStack.
Sophie Baillargeon, Simon Hallé, Christian Gagné 0001
ASONAM3
2016 Sensor control for temporal coverage optimization
abstract
This paper proposes a new sensor control algorithm to adapt the operation parameters of the sensors in order to optimize temporal coverage. Traditional sensor control algorithms rely on non-probabilistic sensor coverage models and simple target trajectory prediction methods, while the new algorithm overcomes all these limitations. The proposed approach make an optimization of one sensor at the time, processed in a random order. The performance of the proposed algorithm is compared with a state-of-the-art general purpose optimization method (i.e., CMA-ES). Results show that our algorithm produces the results in a much shorter time in all the cases considered, while the final coverage of the network was superior in a smaller map and competitive in larger maps. Therefore, the new algorithm can be used in scenarios where response within some time limit is of great importance (e.g., real-time control of the sensors).
Vahab Akbarzadeh, Christian Gagné 0001, Marc Parizeau
CEC2
2016 Bayesian Hyperparameter Optimization for Ensemble Learning
Julien-Charles Levesque, Christian Gagné 0001, Robert Sabourin
UAI2
2016 A double-layer ELM with added feature selection ability using a sparse Bayesian approach
Farkhondeh Kiaee, Christian Gagné 0001, Hamid Sheikhzadeh
Neurocomputing2
2015 Monocular 3D Human Pose Estimation with a Semi-supervised Graph-Based Method
abstract
In this paper, a semi-supervised graph-based method for estimating 3D body pose from a sequence of silhouettes, is presented. The performance of graph-based methods is highly dependent on the quality of the constructed graph. In the case of the human pose estimation problem, the missing depth information from silhouettes intensifies the occurrence of shortcut edges within the graph. To identify and remove these shortcut edges, we measure the similarity of each pair of connected vertices through the use of sliding temporal windows. Furthermore, by exploiting the relationships between labeled and unlabeled data, the proposed method can estimate the 3D body poses, with a small set of labeled data. We evaluated the proposed method on several activities and compared the results with other recent methods. Our method significantly reduced the mean squared error, showing the positive effect of removing shortcut edges.
Mahdieh Abbasi, Hamid R. Rabiee 0001, Christian Gagné 0001
3DV3
2015 Two-Step Heterogeneous Finite Mixture Model Clustering for Mining Healthcare Databases
abstract
Dealing with real-life databases often implies handling sets of heterogeneous variables. We are proposing in this paper a methodology for exploring and analyzing such databases, with an application in the specific domain of healthcare data analytics. We are thus proposing a two-step heterogeneous finite mixture model, with a first step involving a joint mixture of Gaussian and multinomial distribution to handle numerical (i.e., real and integer numbers) and categorical variables (i.e., discrete values), and a second step featuring a mixture of hidden Markov models to handle sequences of categorical values (e.g., series of events). This approach is evaluated on a real-world application, the clustering of administrative healthcare databases from Québec, with results illustrating the good performances of the proposed method.
Ahmed Najjar, Christian Gagné 0001, Daniel Reinharz
ICDM2
2015 Multisensor placement in 3D environments via visibility estimation and derivative-free optimization
abstract
This paper proposes a complete system for robotic sensor placement in initially unknown arbitrary three-dimensional environments. The system uses a novel approach for computing the quality of acquisition of a mobile sensor group in such environments. The quality of acquisition is based on a geometric model of a camera which allows accurate sensor models and simple occlusion computation. The proposed system combines this new metric with a global derivative-free optimization algorithm to find simultaneously the number of sensors and their configuration to sense accordingly the environment. The presented framework compares favourably with current techniques working in two-dimensional environments. Furthermore, simulation and experimental results demonstrate the ability of the system to cope with full three-dimensional environments, a domain still unexplored by previous methods.
François-Michel De Rainville, Jean-Philippe Mercier, Christian Gagné 0001, Philippe Giguère, Denis Laurendeau
ICRA3
2015 Kernel density estimation for target trajectory prediction
abstract
This paper proposes the use of a kernel density estimation to measure similarities between trajectories. The similarities are then used to predict the future locations of a target. For a given environment with a history of previous target trajectories, the goal is to establish a probabilistic framework to predict the future trajectory of currently observed targets based on their recent moves. Instead of clustering trajectories into groups, we calculate the similarity between a given test trajectory and the set of all past trajectories in a dataset. Next, we use a weighted mechanism for prediction, that can be used in target tracking and collision avoidance applications. The proposed method is compared with two other commonly used similarity models (PCA and LCSS) over a dataset of simulated trajectories, and two datasets of real observations. Results show that the proposed method significantly outperforms the existing models for those datasets and experimental settings.
Vahab Akbarzadeh, Christian Gagné 0001, Marc Parizeau
IROS2
2015 Training subset selection in Hourly Ontario Energy Price forecasting using time series clustering-based stratification
Karol Lina López Bedoya, Christian Gagné 0001, Germán Castellanos-Domínguez, Mauricio Orozco-Alzate
Neurocomputing2
2013 Evolutionary multiobjective optimization for selecting members of an ensemble streamflow forecasting model
abstract
We are proposing to use the Nondominated Sorting Genetic Algorithm II (NSGA-II) for optimizing a hydrological forecasting model of 800 simultaneous streamflow predictors. The optimization is based on the selection of the best 48 predictors from the 800 that jointly define the "best" ensemble in terms of two probabilistic criteria. Results showed that the difficulties in simplifying the ensembles mainly originate from the preservation of the system reliability. We conclude that Pareto fronts generated with NSGA-II allow the development of a decision process based explicitly on the trade-off between different probabilistic properties. In other words, evolutionary multiobjective optimization offers more flexibility to the operational hydrologists than a priori methods that produce only one selection.
Darwin Brochero, Christian Gagné 0001, François Anctil
GECCO2
2013 Sustainable cooperative coevolution with a multi-armed bandit
abstract
This paper proposes a self-adaptation mechanism to manage the resources allocated to the different species comprising a cooperative coevolutionary algorithm. The proposed approach relies on a dynamic extension to the well-known multi-armed bandit framework. At each iteration, the dynamic multi-armed bandit makes a decision on which species to evolve for a generation, using the history of progress made by the different species to guide the decisions. We show experimentally, on a benchmark and a real-world problem, that evolving the different populations at different paces allows not only to identify solutions more rapidly, but also improves the capacity of cooperative coevolution to solve more complex problems.
François-Michel De Rainville, Michèle Sebag, Christian Gagné 0001, Marc Schoenauer, Denis Laurendeau
GECCO3
2012 Multi-objective evolutionary optimization for generating ensembles of classifiers in the ROC space
abstract
In this paper, we propose a novel approach for the multi-objective optimization of classifier ensembles in the ROC space. We first evolve a pool of simple classifiers with NSGA-II using values of the ROC curves as the optimization objectives. These simple classifiers are then combined at the decision level using the Iterative Boolean Combination method (IBC). This method produces multiple ensembles of classifiers optimized for various operating conditions. We perform a rigorous series of experiments to demonstrate the properties and behaviour of this approach. This allows us to propose interesting venues for future research on optimizing ensembles of classifiers using multi-objective evolutionary algorithms.
Julien-Charles Levesque, Audrey Durand, Christian Gagné 0001, Robert Sabourin
GECCO3
2012 DEAP: evolutionary algorithms made easy
Félix-Antoine Fortin, François-Michel De Rainville, Marc-André Gardner, Marc Parizeau, Christian Gagné 0001
J. Mach. Learn. Res.5
2012 Population-Based Simulation for Public Health: Generic Software Infrastructure and Its Application to Osteoporosis
abstract
Policy-making in public health has great socio-economical consequences and must be done using the best available knowledge on the possible options. These processes are often too complex to be evaluated through analytical methods, such that computer simulations are often the best way to produce quantitative evaluations of their performances. For that purpose, we are proposing a complete software infrastructure for the simulation of public health processes. This software stack includes a generic population-based simulator called SynCHroNous Agent- and Population-based Simulator, which has a modern object-oriented software architecture, and is completely configured through eXtensible Markup Language files. These configuration files can themselves be produced by a graphical user interface that allows modeling of public health simulation by nonprogrammers. This software infrastructure has been illustrated with the real-life case study of osteoporosis prevention in adult women populations. This example, which is of great interest for Quebec health decision makers, provides insightful results for comparing several prevention strategies on a realistic population.
Audrey Durand, Christian Gagné 0001, Léon Nshimyumukiza, Mathieu Gagnon, Yves Giguère, Daniel Reinharz
IEEE Trans. Syst. Man Cybern. Part A2
2010 Topography-Aware Sensor Deployment Optimization with CMA-ES
Vahab Akbarzadeh, Albert Hung-Ren Ko, Christian Gagné 0001, Marc Parizeau
PPSN (2)3
2009 A Statistical Learning Perspective of Genetic Programming
Nur Merve Amil, Nicolas Bredèche, Christian Gagné 0001, Sylvain Gelly, Marc Schoenauer, Olivier Teytaud
EuroGP3
2009 Co-evolutionary information gathering for a cooperative unmanned aerial vehicle team
Jean Berger, Jens Happe, Christian Gagné 0001, Martin Lau
FUSION3
2009 Improving genetic algorithms performance via deterministic population shrinkage
abstract
Despite the intuition that the same population size is not needed throughout the run of an Evolutionary Algorithm (EA), most EAs use a fixed population size. This paper presents an empirical study on the possible benefits of a Simple Variable Population Sizing (SVPS) scheme on the performance of Genetic Algorithms (GAs). It consists in decreasing the population for a GA run following a predetermined schedule, configured by a speed and a severity parameter. The method uses as initial population size an estimation of the minimum size needed to supply enough building blocks, using a fixed-size selectorecombinative GA converging within some confidence interval toward good solutions for a particular problem. Following this methodology, a scalability analysis is conducted on deceptive, quasi-deceptive, and non-deceptive trap functions in order to assess whether SVPS-GA improves performances compared to a fixed-size GA under different problem instances and difficulty levels. Results show several combinations of speed-severity where SVPS-GA preserves the solution quality while improving performances, by reducing the number of evaluations needed for success.
Juan Luis Jiménez Laredo, Carlos M. Fernandes 0001, Juan Julián Merelo Guervós, Christian Gagné 0001
GECCO4
2009 Optimizing low-discrepancy sequences with an evolutionary algorithm
abstract
Many fields rely on some stochastic sampling of a given complex space. Low-discrepancy sequences are methods aiming at producing samples with better space-filling properties than uniformly distributed random numbers, hence allowing a more efficient sampling of that space. State-of-the-art methods like nearly orthogonal Latin hypercubes and scrambled Halton sequences are configured by permutations of internal parameters, where permutations are commonly done randomly. This paper proposes the use of evolutionary algorithms to evolve these permutations, in order to optimize a discrepancy measure. Results show that an evolutionary method is able to generate low-discrepancy sequences of significantly better space-filling properties compared to sequences configured with purely random permutations.
François-Michel De Rainville, Christian Gagné 0001, Olivier Teytaud, Denis Laurendeau
GECCO2
2007 Ensemble learning for free with evolutionary algorithms?
abstract
Evolutionary Learning proceeds by evolving a population of classifiers, from which it generally returns (with some notable exceptions) the single best-of-run classifier as final result. In the meanwhile, Ensemble Learning, one of the most efficient approaches in supervised Machine Learning for the last decade, proceeds by building a population of diverse classifiers. Ensemble Learning with Evolutionary Computation thus receives increasing attention. The Evolutionary Ensemble Learning (EEL) approach presented in this paper features two contributions. First, a new fitness function, inspired by co-evolution and enforcing the classifier diversity, is presented. Further, a new selection criterion based on the classification margin is proposed. This criterion is used to extract the classifier ensemble from the final population only (Off-EEL) or incrementally along evolution (On-EEL). Experiments on a set of benchmark problems show that Off-EEL outperforms single-hypothesis evolutionary learning and state-of-art Boosting and generates smaller classifier ensembles.
Christian Gagné 0001, Michèle Sebag, Marc Schoenauer, Marco Tomassini
GECCO1
2007 Coevolution of Nearest Neighbor Classifiers
abstract
This paper presents experiments of Nearest Neighbor (NN) classifier design using different evolutionary computation methods. Through multiobjective and coevolution techniques, it combines genetic algorithms and genetic programming to both select NN prototypes and design a neighborhood proximity measure, in order to produce a more efficient and robust classifier. The proposed approach is compared with the standard NN classifier, with and without the use of classic prototype selection methods, and classic data normalization. Results on both synthetic and real data sets show that the proposed methodology performs as well or better than other methods on all tested data sets.
Christian Gagné 0001, Marc Parizeau
Int. J. Pattern Recognit. Artif. Intell.1
2006 Resource-Aware Parameterizations of EDA
abstract
This paper presents a framework for the theoretical analysis of Estimation of Distribution Algorithms (EDA). Using this framework, derived from the VC-theory, we propose non-asymptotic bounds which depend on: 1) the population size, 2) the selection rate, 3) the families of distributions used for the modelling, 4) the dimension, and 5) the number of iterations. To validate these results, optimization algorithms are applied to a context where bounds on resources are crucial, namely Design of Experiments, that is a black-box optimization with very few fitness-values evaluations.
Sylvain Gelly, Olivier Teytaud, Christian Gagné 0001
IEEE Congress on Evolutionary Computation3
2006 Genetic Programming, Validation Sets, and Parsimony Pressure
Christian Gagné 0001, Marc Schoenauer, Marc Parizeau, Marco Tomassini
EuroGP1
2006 Genetic Programming for Kernel-Based Learning with Co-evolving Subsets Selection
Christian Gagné 0001, Marc Schoenauer, Michèle Sebag, Marco Tomassini
PPSN1
2006 Genetic engineering of hierarchical fuzzy regional representations for handwritten character recognition
Christian Gagné 0001, Marc Parizeau
Int. J. Document Anal. Recognit.1
2006 Analysis of a master-slave architecture for distributed evolutionary computations
abstract
This paper introduces a new mathematical model of the master-slave architecture for distributed evolutionary computations (EC). This model is validated using a concrete implementation based on the Distributed BEAGLE C++ framework. Results show that contrary to (current) popular belief, master-slave architectures are able to scale well over local area networks of workstations using off-the-shelf networking equipment. The main properties of the master-slave are also compared with those of the more mainstream island-model.
Marc Dubreuil, Christian Gagné 0001, Marc Parizeau
IEEE Trans. Syst. Man Cybern. Part B2
2003 The Master-Slave Architecture for Evolutionary Computations Revisited
Christian Gagné 0001, Marc Parizeau, Marc Dubreuil
GECCO1
2002 Lens System Design And Re-engineering With Evolutionary Algorithms
Julie Beaulieu, Christian Gagné 0001, Marc Parizeau
GECCO2
2002 Open BEAGLE: A New C++ Evolutionary Computation Framework
Christian Gagné 0001, Marc Parizeau
GECCO1
2001 Character Recognition Experiments Using Unipen Data
abstract
This paper presents experiments that compare the performances of several versions of a regional-fuzzy representation (RFR) developed for cursive handwriting recognition (CHR). These experiments are conducted using a common neural network classifier namely a multilayer perceptron (MLP) trained with backpropagation. Results are given for isolated digits, isolated lower-case letters and lower-case letters extracted from phrases, from the Unipen database. Data set Train-R01/V07 is used for training while DevTest-R01/V02 is used for testing. The best overall representation yields recognition rates of respectively 97.0% and 85.6% for isolated digits and lower case, and 84.4% for lower-case extracted from phrases.
Marc Parizeau, Alexandre Lemieux, Christian Gagné 0001
ICDAR3