Sébastien Lachapelle

dblp:224/0080 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-8290-5951ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Representation and self-supervised learning · 51% Probabilistic and Bayesian machine learning · 20% Transfer learning and domain adaptation · 13%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
1.522025
Interaction Asymmetry: A General Principle for Learning Composable Abstractions · ICLR 2025
Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023
Machine learning › Representation and self-supervised learning
causal representation learning
1.522024
A Sparsity Principle for Partially Observable Causal Representation Learning · ICML 2024
Multi-View Causal Representation Learning with Partial Observability · ICLR 2024
Machine learning › Representation and self-supervised learning › blind source separation › independent component analysis
nonlinear ICA
1.422024
Multi-View Causal Representation Learning with Partial Observability · ICLR 2024
Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.922020
Differentiable Causal Discovery from Interventional Data · NeurIPS 2020
Gradient-Based Neural DAG Learning · ICLR 2020
Natural language and speech › Language models and text generation
compositional generalization
0.912025
Interaction Asymmetry: A General Principle for Learning Composable Abstractions · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
disentanglement
0.812024
Multi-View Causal Representation Learning with Partial Observability · ICLR 2024
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability
0.812024
A Sparsity Principle for Partially Observable Causal Representation Learning · ICML 2024
Machine learning › Transfer learning and domain adaptation
few-shot classification
0.712023
Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent variable identification
0.712023
Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.712023
Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.412020
Differentiable Causal Discovery from Interventional Data · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
directed acyclic graph learning
0.412020
Gradient-Based Neural DAG Learning · ICLR 2020
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-transfer learning
0.412020
A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms · ICLR 2020
Machine learning › Generative modeling
normalizing flow
0.412020
Differentiable Causal Discovery from Interventional Data · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.412020
Gradient-Based Neural DAG Learning · ICLR 2020
Machine learning › Optimization for machine learning
bilevel optimization
0.212023
Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023
Machine learning › Optimization for machine learning
constrained optimization
0.112020
Differentiable Causal Discovery from Interventional Data · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

regularizer · 0.9neural network · 0.9autoencoder · 0.9Transformer-based VAE · 0.9sparsity regularization · 0.8piecewise linear mixing · 0.8contrastive learning · 0.8multi-class SVM · 0.7group lasso · 0.7dual formulation · 0.7
YearPublicationVenuePosition
2025 All or None: Identifiable Linear Properties of Next-Token Predictors in Language Modeling
abstract
We analyze identifiability as a possible explanation for the ubiquity of linear properties across language models, such as the vector difference between the representations of “easy” and “easiest” being parallel to that between “lucky” and “luckiest”. For this, we ask whether finding a linear property in one model implies that any model that induces the same distribution has that property, too. To answer that, we first prove an identifiability result to characterize distribution-equivalent next-token predictors, lifting a diversity requirement of previous results. Second, based on a refinement of relational linearity [Paccanaro and Hinton, 2001; Hernandez et al., 2024], we show how many notions of linearity are amenable to our analysis. Finally, we show that under suitable conditions, these linear properties either hold in all or none distribution equivalent next-token predictors.
Emanuele Marconato, Sébastien Lachapelle, Sebastian Weichwald, Luigi Gresele
AISTATS2
2025 Interaction Asymmetry: A General Principle for Learning Composable Abstractions
abstract
Learning disentangled representations of concepts and re-composing them in unseen ways is crucial for generalizing to out-of-domain situations. However, the underlying properties of concepts that enable such disentanglement and compositional generalization remain poorly understood. In this work, we propose the principle of interaction asymmetry which states: "Parts of the same concept have more complex interactions than parts of different concepts". We formalize this via block diagonality conditions on the $(n+1)$th order derivatives of the generator mapping concepts to observed data, where different orders of "complexity" correspond to different $n$. Using this formalism, we prove that interaction asymmetry enables both disentanglement and compositional generalization. Our results unify recent theoretical results for learning concepts of objects, which we show are recovered as special cases with $n=0$ or $1$. We provide results for up to $n=2$, thus extending these prior works to more flexible generator functions, and conjecture that the same proof strategies generalize to larger $n$. Practically, our theory suggests that, to disentangle concepts, an autoencoder should penalize its latent capacity and the interactions between concepts during decoding. We propose an implementation of these criteria using a flexible Transformer-based VAE, with a novel regularizer on the attention weights of the decoder. On synthetic image datasets consisting of objects, we provide evidence that this model can achieve comparable object disentanglement to existing models that use more explicit object-centric priors.
Jack Brady, Julius von Kügelgen, Sébastien Lachapelle, Simon Buchholz, Thomas Kipf, Wieland Brendel
ICLR3
2024 Multi-View Causal Representation Learning with Partial Observability
abstract
We present a unified framework for studying the identifiability of representations learned from simultaneously observed views, such as different data modalities. We allow a partially observed setting in which each view constitutes a nonlinear mixture of a subset of underlying latent variables, which can be causally related. We prove that the information shared across all subsets of any number of views can be learned up to a smooth bijection using contrastive learning and a single encoder per view. We also provide graphical criteria indicating which latent variables can be identified through a simple set of rules, which we refer to as identifiability algebra. Our general framework and theoretical results unify and extend several previous work on multi-view nonlinear ICA, disentanglement, and causal representation learning. We experimentally validate our claims on numerical, image, and multi-modal data sets. Further, we demonstrate that the performance of prior methods is recovered in different special cases of our setup. Overall, we find that access to multiple partial views offers unique opportunities for identifiable representation learning, enabling the discovery of latent structures from purely observational data.
Dingling Yao, Danru Xu, Sébastien Lachapelle, Sara Magliacane, Perouz Taslakian, Georg Martius, Julius von Kügelgen, Francesco Locatello
ICLR3
2024 A Sparsity Principle for Partially Observable Causal Representation Learning
abstract
Causal representation learning aims at identifying high-level causal variables from perceptual data. Most methods assume that all latent causal variables are captured in the high-dimensional observations. We instead consider a partially observed setting, in which each measurement only provides information about a subset of the underlying causal state. Prior work has studied this setting with multiple domains or views, each depending on a fixed subset of latents. Here, we focus on learning from unpaired observations from a dataset with an instance-dependent partial observability pattern. Our main contribution is to establish two identifiability results for this setting: one for linear mixing functions without parametric assumptions on the underlying causal model, and one for piecewise linear mixing functions with Gaussian latent causal variables. Based on these insights, we propose two methods for estimating the underlying causal variables by enforcing sparsity in the inferred representation. Experiments on different simulated datasets and established benchmarks highlight the effectiveness of our approach in recovering the ground-truth latents.
Danru Xu, Dingling Yao, Sébastien Lachapelle, Perouz Taslakian, Julius von Kügelgen, Francesco Locatello, Sara Magliacane
ICML3
2023 Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning
abstract
Although disentangled representations are often said to be beneficial for downstream tasks, current empirical and theoretical understanding is limited. In this work, we provide evidence that disentangled representations coupled with sparse task-specific predictors improve generalization. In the context of multi-task learning, we prove a new identifiability result that provides conditions under which maximally sparse predictors yield disentangled representations. Motivated by this theoretical result, we propose a practical approach to learn disentangled representations based on a sparsity-promoting bi-level optimization problem. Finally, we explore a meta-learning version of this algorithm based on group Lasso multiclass SVM predictors, for which we derive a tractable dual formulation. It obtains competitive results on standard few-shot classification benchmarks, while each task is using only a fraction of the learned representations.
Sébastien Lachapelle, Tristan Deleu, Divyat Mahajan, Ioannis Mitliagkas, Yoshua Bengio, Simon Lacoste-Julien, Quentin Bertrand
ICML1
2023 Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation
abstract
We tackle the problems of latent variables identification and "out-of-support'' image generation in representation learning. We show that both are possible for a class of decoders that we call additive, which are reminiscent of decoders used for object-centric representation learning (OCRL) and well suited for images that can be decomposed as a sum of object-specific images. We provide conditions under which exactly solving the reconstruction problem using an additive decoder is guaranteed to identify the blocks of latent variables up to permutation and block-wise invertible transformations. This guarantee relies only on very weak assumptions about the distribution of the latent factors, which might present statistical dependencies and have an almost arbitrarily shaped support. Our result provides a new setting where nonlinear independent component analysis (ICA) is possible and adds to our theoretical understanding of OCRL methods. We also show theoretically that additive decoders can generate novel images by recombining observed factors of variations in novel ways, an ability we refer to as Cartesian-product extrapolation. We show empirically that additivity is crucial for both identifiability and extrapolation on simulated data.
Sébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, Simon Lacoste-Julien
NeurIPS1
2022 On the Convergence of Continuous Constrained Optimization for Structure Learning
abstract
Recently, structure learning of directed acyclic graphs (DAGs) has been formulated as a continuous optimization problem by leveraging an algebraic characterization of acyclicity. The constrained problem is solved using the augmented Lagrangian method (ALM) which is often preferred to the quadratic penalty method (QPM) by virtue of its standard convergence result that does not require the penalty coefficient to go to infinity, hence avoiding ill-conditioning. However, the convergence properties of these methods for structure learning, including whether they are guaranteed to return a DAG solution, remain unclear, which might limit their practical applications. In this work, we examine the convergence of ALM and QPM for structure learning in the linear, nonlinear, and confounded cases. We show that the standard convergence result of ALM does not hold in these settings, and demonstrate empirically that its behavior is akin to that of the QPM which is prone to ill-conditioning. We further establish the convergence guarantee of QPM to a DAG solution, under mild conditions. Lastly, we connect our theoretical results with existing approaches to help resolve the convergence issue, and verify our findings in light of an empirical comparison of them.
Ignavier Ng, Sébastien Lachapelle, Nan Rosemary Ke, Simon Lacoste-Julien, Kun Zhang 0001
AISTATS2
2022 Predicting Tactical Solutions to Operational Planning Problems Under Imperfect Information
abstract
This paper offers a methodological contribution at the intersection of machine learning and operations research. Namely, we propose a methodology to quickly predict expected tactical descriptions of operational solutions (TDOSs). The problem we address occurs in the context of two-stage stochastic programming, where the second stage is demanding computationally. We aim to predict at a high speed the expected TDOS associated with the second-stage problem, conditionally on the first-stage variables. This may be used in support of the solution to the overall two-stage problem by avoiding the online generation of multiple second-stage scenarios and solutions. We formulate the tactical prediction problem as a stochastic optimal prediction program, whose solution we approximate with supervised machine learning. The training data set consists of a large number of deterministic operational problems generated by controlled probabilistic sampling. The labels are computed based on solutions to these problems (solved independently and offline), employing appropriate aggregation and subselection methods to address uncertainty. Results on our motivating application on load planning for rail transportation show that deep learning models produce accurate predictions in very short computing time (milliseconds or less). The predictive accuracy is close to the lower bounds calculated based on sample average approximation of the stochastic prediction programs.
Eric Larsen, Sébastien Lachapelle, Yoshua Bengio, Emma Frejinger, Simon Lacoste-Julien, Andrea Lodi 0001
INFORMS J. Comput.2
2020 A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, Christopher Joseph Pal
ICLR5
2020 Gradient-Based Neural DAG Learning
Sébastien Lachapelle, Philippe Brouillard, Tristan Deleu, Simon Lacoste-Julien
ICLR1
2020 Differentiable Causal Discovery from Interventional Data
abstract
Learning a causal directed acyclic graph from data is a challenging task that involves solving a combinatorial problem for which the solution is not always identifiable. A new line of work reformulates this problem as a continuous constrained optimization one, which is solved via the augmented Lagrangian method. However, most methods based on this idea do not make use of interventional data, which can significantly alleviate identifiability issues. This work constitutes a new step in this direction by proposing a theoretically-grounded method based on neural networks that can leverage interventional data. We illustrate the flexibility of the continuous-constrained framework by taking advantage of expressive neural architectures such as normalizing flows. We show that our approach compares favorably to the state of the art in a variety of settings, including perfect and imperfect interventions for which the targeted nodes may even be unknown.
Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien, Alexandre Drouin
NeurIPS2