EDBT 2026 Demo / reviewers in the wild / expert
Sébastien Lachapelle
dblp:224/0080
· DBLP profile ↗
11ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-8290-5951ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Representation and self-supervised learning · 51% Probabilistic and Bayesian machine learning · 20% Transfer learning and domain adaptation · 13% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
1.5 | 2 | 2025 | Interaction Asymmetry: A General Principle for Learning Composable Abstractions · ICLR 2025 Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023 |
Machine learning › Representation and self-supervised learning
causal representation learning |
1.5 | 2 | 2024 | A Sparsity Principle for Partially Observable Causal Representation Learning · ICML 2024 Multi-View Causal Representation Learning with Partial Observability · ICLR 2024 |
Machine learning › Representation and self-supervised learning › blind source separation › independent component analysis
nonlinear ICA |
1.4 | 2 | 2024 | Multi-View Causal Representation Learning with Partial Observability · ICLR 2024 Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.9 | 2 | 2020 | Differentiable Causal Discovery from Interventional Data · NeurIPS 2020 Gradient-Based Neural DAG Learning · ICLR 2020 |
Natural language and speech › Language models and text generation
compositional generalization |
0.9 | 1 | 2025 | Interaction Asymmetry: A General Principle for Learning Composable Abstractions · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
disentanglement |
0.8 | 1 | 2024 | Multi-View Causal Representation Learning with Partial Observability · ICLR 2024 |
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability |
0.8 | 1 | 2024 | A Sparsity Principle for Partially Observable Causal Representation Learning · ICML 2024 |
Machine learning › Transfer learning and domain adaptation
few-shot classification |
0.7 | 1 | 2023 | Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent variable identification |
0.7 | 1 | 2023 | Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.7 | 1 | 2023 | Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
0.7 | 1 | 2023 | Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.4 | 1 | 2020 | Differentiable Causal Discovery from Interventional Data · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
directed acyclic graph learning |
0.4 | 1 | 2020 | Gradient-Based Neural DAG Learning · ICLR 2020 |
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-transfer learning |
0.4 | 1 | 2020 | A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms · ICLR 2020 |
Machine learning › Generative modeling
normalizing flow |
0.4 | 1 | 2020 | Differentiable Causal Discovery from Interventional Data · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning |
0.4 | 1 | 2020 | Gradient-Based Neural DAG Learning · ICLR 2020 |
Machine learning › Optimization for machine learning
bilevel optimization |
0.2 | 1 | 2023 | Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning · ICML 2023 |
Machine learning › Optimization for machine learning
constrained optimization |
0.1 | 1 | 2020 | Differentiable Causal Discovery from Interventional Data · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
regularizer · 0.9neural network · 0.9autoencoder · 0.9Transformer-based VAE · 0.9sparsity regularization · 0.8piecewise linear mixing · 0.8contrastive learning · 0.8multi-class SVM · 0.7group lasso · 0.7dual formulation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | All or None: Identifiable Linear Properties of Next-Token Predictors in Language ModelingabstractWe analyze identifiability as a possible explanation for the ubiquity of linear properties across language models, such as the vector difference between the representations of “easy” and “easiest” being parallel to that between “lucky” and “luckiest”. For this, we ask whether finding a linear property in one model implies that any model that induces the same distribution has that property, too. To answer that, we first prove an identifiability result to characterize distribution-equivalent next-token predictors, lifting a diversity requirement of previous results. Second, based on a refinement of relational linearity [Paccanaro and Hinton, 2001; Hernandez et al., 2024], we show how many notions of linearity are amenable to our analysis. Finally, we show that under suitable conditions, these linear properties either hold in all or none distribution equivalent next-token predictors. Emanuele Marconato, Sébastien Lachapelle, Sebastian Weichwald, Luigi Gresele |
AISTATS | 2 |
| 2025 | Interaction Asymmetry: A General Principle for Learning Composable AbstractionsabstractLearning disentangled representations of concepts and re-composing them in unseen ways is crucial for generalizing to out-of-domain situations. However, the underlying properties of concepts that enable such disentanglement and compositional generalization remain poorly understood. In this work, we propose the principle of interaction asymmetry which states: "Parts of the same concept have more complex interactions than parts of different concepts". We formalize this via block diagonality conditions on the $(n+1)$th order derivatives of the generator mapping concepts to observed data, where different orders of "complexity" correspond to different $n$. Using this formalism, we prove that interaction asymmetry enables both disentanglement and compositional generalization. Our results unify recent theoretical results for learning concepts of objects, which we show are recovered as special cases with $n=0$ or $1$. We provide results for up to $n=2$, thus extending these prior works to more flexible generator functions, and conjecture that the same proof strategies generalize to larger $n$. Practically, our theory suggests that, to disentangle concepts, an autoencoder should penalize its latent capacity and the interactions between concepts during decoding. We propose an implementation of these criteria using a flexible Transformer-based VAE, with a novel regularizer on the attention weights of the decoder. On synthetic image datasets consisting of objects, we provide evidence that this model can achieve comparable object disentanglement to existing models that use more explicit object-centric priors. Jack Brady, Julius von Kügelgen, Sébastien Lachapelle, Simon Buchholz, Thomas Kipf, Wieland Brendel |
ICLR | 3 |
| 2024 | Multi-View Causal Representation Learning with Partial ObservabilityabstractWe present a unified framework for studying the identifiability of representations learned from simultaneously observed views, such as different data modalities. We allow a partially observed setting in which each view constitutes a nonlinear mixture of a subset of underlying latent variables, which can be causally related.
We prove that the information shared across all subsets of any number of views can be learned up to a smooth bijection using contrastive learning and a single encoder per view.
We also provide graphical criteria indicating which latent variables can be identified through a simple set of rules, which we refer to as identifiability algebra. Our general framework and theoretical results unify and extend several previous work on multi-view nonlinear ICA, disentanglement, and causal representation learning. We experimentally validate our claims on numerical, image, and multi-modal data sets. Further, we demonstrate that the performance of prior methods is recovered in different special cases of our setup.
Overall, we find that access to multiple partial views offers unique opportunities for identifiable representation learning, enabling the discovery of latent structures from purely observational data. Dingling Yao, Danru Xu, Sébastien Lachapelle, Sara Magliacane, Perouz Taslakian, Georg Martius, Julius von Kügelgen, Francesco Locatello |
ICLR | 3 |
| 2024 | A Sparsity Principle for Partially Observable Causal Representation LearningabstractCausal representation learning aims at identifying high-level causal variables from perceptual data. Most methods assume that all latent causal variables are captured in the high-dimensional observations. We instead consider a partially observed setting, in which each measurement only provides information about a subset of the underlying causal state. Prior work has studied this setting with multiple domains or views, each depending on a fixed subset of latents. Here, we focus on learning from unpaired observations from a dataset with an instance-dependent partial observability pattern. Our main contribution is to establish two identifiability results for this setting: one for linear mixing functions without parametric assumptions on the underlying causal model, and one for piecewise linear mixing functions with Gaussian latent causal variables. Based on these insights, we propose two methods for estimating the underlying causal variables by enforcing sparsity in the inferred representation. Experiments on different simulated datasets and established benchmarks highlight the effectiveness of our approach in recovering the ground-truth latents. Danru Xu, Dingling Yao, Sébastien Lachapelle, Perouz Taslakian, Julius von Kügelgen, Francesco Locatello, Sara Magliacane |
ICML | 3 |
| 2023 | Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task LearningabstractAlthough disentangled representations are often said to be beneficial for downstream tasks, current empirical and theoretical understanding is limited. In this work, we provide evidence that disentangled representations coupled with sparse task-specific predictors improve generalization. In the context of multi-task learning, we prove a new identifiability result that provides conditions under which maximally sparse predictors yield disentangled representations. Motivated by this theoretical result, we propose a practical approach to learn disentangled representations based on a sparsity-promoting bi-level optimization problem. Finally, we explore a meta-learning version of this algorithm based on group Lasso multiclass SVM predictors, for which we derive a tractable dual formulation. It obtains competitive results on standard few-shot classification benchmarks, while each task is using only a fraction of the learned representations. Sébastien Lachapelle, Tristan Deleu, Divyat Mahajan, Ioannis Mitliagkas, Yoshua Bengio, Simon Lacoste-Julien, Quentin Bertrand |
ICML | 1 |
| 2023 | Additive Decoders for Latent Variables Identification and Cartesian-Product ExtrapolationabstractWe tackle the problems of latent variables identification and "out-of-support'' image generation in representation learning. We show that both are possible for a class of decoders that we call additive, which are reminiscent of decoders used for object-centric representation learning (OCRL) and well suited for images that can be decomposed as a sum of object-specific images. We provide conditions under which exactly solving the reconstruction problem using an additive decoder is guaranteed to identify the blocks of latent variables up to permutation and block-wise invertible transformations. This guarantee relies only on very weak assumptions about the distribution of the latent factors, which might present statistical dependencies and have an almost arbitrarily shaped support. Our result provides a new setting where nonlinear independent component analysis (ICA) is possible and adds to our theoretical understanding of OCRL methods. We also show theoretically that additive decoders can generate novel images by recombining observed factors of variations in novel ways, an ability we refer to as Cartesian-product extrapolation. We show empirically that additivity is crucial for both identifiability and extrapolation on simulated data. Sébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, Simon Lacoste-Julien |
NeurIPS | 1 |
| 2022 | On the Convergence of Continuous Constrained Optimization for Structure LearningabstractRecently, structure learning of directed acyclic graphs (DAGs) has been formulated as a continuous optimization problem by leveraging an algebraic characterization of acyclicity. The constrained problem is solved using the augmented Lagrangian method (ALM) which is often preferred to the quadratic penalty method (QPM) by virtue of its standard convergence result that does not require the penalty coefficient to go to infinity, hence avoiding ill-conditioning. However, the convergence properties of these methods for structure learning, including whether they are guaranteed to return a DAG solution, remain unclear, which might limit their practical applications. In this work, we examine the convergence of ALM and QPM for structure learning in the linear, nonlinear, and confounded cases. We show that the standard convergence result of ALM does not hold in these settings, and demonstrate empirically that its behavior is akin to that of the QPM which is prone to ill-conditioning. We further establish the convergence guarantee of QPM to a DAG solution, under mild conditions. Lastly, we connect our theoretical results with existing approaches to help resolve the convergence issue, and verify our findings in light of an empirical comparison of them. Ignavier Ng, Sébastien Lachapelle, Nan Rosemary Ke, Simon Lacoste-Julien, Kun Zhang 0001 |
AISTATS | 2 |
| 2022 | Predicting Tactical Solutions to Operational Planning Problems Under Imperfect InformationabstractThis paper offers a methodological contribution at the intersection of machine learning and operations research. Namely, we propose a methodology to quickly predict expected tactical descriptions of operational solutions (TDOSs). The problem we address occurs in the context of two-stage stochastic programming, where the second stage is demanding computationally. We aim to predict at a high speed the expected TDOS associated with the second-stage problem, conditionally on the first-stage variables. This may be used in support of the solution to the overall two-stage problem by avoiding the online generation of multiple second-stage scenarios and solutions. We formulate the tactical prediction problem as a stochastic optimal prediction program, whose solution we approximate with supervised machine learning. The training data set consists of a large number of deterministic operational problems generated by controlled probabilistic sampling. The labels are computed based on solutions to these problems (solved independently and offline), employing appropriate aggregation and subselection methods to address uncertainty. Results on our motivating application on load planning for rail transportation show that deep learning models produce accurate predictions in very short computing time (milliseconds or less). The predictive accuracy is close to the lower bounds calculated based on sample average approximation of the stochastic prediction programs. Eric Larsen, Sébastien Lachapelle, Yoshua Bengio, Emma Frejinger, Simon Lacoste-Julien, Andrea Lodi 0001 |
INFORMS J. Comput. | 2 |
| 2020 | A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, Christopher Joseph Pal |
ICLR | 5 |
| 2020 | Gradient-Based Neural DAG Learning
Sébastien Lachapelle, Philippe Brouillard, Tristan Deleu, Simon Lacoste-Julien |
ICLR | 1 |
| 2020 | Differentiable Causal Discovery from Interventional DataabstractLearning a causal directed acyclic graph from data is a challenging task that involves solving a combinatorial problem for which the solution is not always identifiable. A new line of work reformulates this problem as a continuous constrained optimization one, which is solved via the augmented Lagrangian method. However, most methods based on this idea do not make use of interventional data, which can significantly alleviate identifiability issues. This work constitutes a new step in this direction by proposing a theoretically-grounded method based on neural networks that can leverage interventional data. We illustrate the flexibility of the continuous-constrained framework by taking advantage of expressive neural architectures such as normalizing flows. We show that our approach compares favorably to the state of the art in a variety of settings, including perfect and imperfect interventions for which the targeted nodes may even be unknown. Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien, Alexandre Drouin |
NeurIPS | 2 |