Sebastian Weichwald

dblp:158/0010 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-0169-7244ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 73% Knowledge representation and reasoning · 20% Representation and self-supervised learning · 7%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Theoretical computer science
3 papers
Mathematical optimization · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
additive noise model
1.222023
A Scale-Invariant Sorting Criterion to Find a Causal Order in Additive Noise Models · NeurIPS 2023
Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to Game · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
1.222023
A Scale-Invariant Sorting Criterion to Find a Causal Order in Additive Noise Models · NeurIPS 2023
Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to Game · NeurIPS 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
1.222023
A Scale-Invariant Sorting Criterion to Find a Causal Order in Additive Noise Models · NeurIPS 2023
Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to Game · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.922024
Identifying Causal Effects using Instrumental Time Series: Nuisance IV and Correcting for the Past · J. Mach. Learn. Res. 2024
Robustifying Independent Component Analysis by Adjusting for Group-Wise Stationary Noise · J. Mach. Learn. Res. 2019
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable regression
0.812024
Identifying Causal Effects using Instrumental Time Series: Nuisance IV and Correcting for the Past · J. Mach. Learn. Res. 2024
Bioinformatics and computational biology › single-cell analysis
cytometry data analysis
0.812024
<tt>spillR</tt> : spillover compensation in mass cytometry data · Bioinform. 2024
Machine learning › Representation and self-supervised learning › blind source separation
independent component analysis
0.412019
Robustifying Independent Component Analysis by Adjusting for Group-Wise Stationary Noise · J. Mach. Learn. Res. 2019
Mathematical optimization
riemannian optimization
0.212016
Pymanopt: A Python Toolbox for Optimization on Manifolds using Automatic Differentiation · J. Mach. Learn. Res. 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212024
Identifying Causal Effects using Instrumental Time Series: Nuisance IV and Correcting for the Past · J. Mach. Learn. Res. 2024
Mathematical optimization › statistical estimation
regression
0.212023
A Scale-Invariant Sorting Criterion to Find a Causal Order in Additive Noise Models · NeurIPS 2023
Mathematical optimization
continuous optimization
0.112021
Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to Game · NeurIPS 2021
Mathematical optimization
automatic differentiation
0.112016
Pymanopt: A Python Toolbox for Optimization on Manifolds using Automatic Differentiation · J. Mach. Learn. Res. 2016

Methods — techniques the papers use, named apart from their topics

r²-sortnregress · 1.3coefficient of determination · 1.3continuous structure learning · 1.0baseline variance sorting · 1.0vector autoregressive processes · 0.8nuisance covariates · 0.8nonparametric finite mixture model · 0.8graph marginalization · 0.8expectation-maximization · 0.8identifiability analysis · 0.4group-wise stationary noise model · 0.4manifold geometry · 0.2automatic differentiation · 0.2
YearPublicationVenuePosition
2025 All or None: Identifiable Linear Properties of Next-Token Predictors in Language Modeling
abstract
We analyze identifiability as a possible explanation for the ubiquity of linear properties across language models, such as the vector difference between the representations of “easy” and “easiest” being parallel to that between “lucky” and “luckiest”. For this, we ask whether finding a linear property in one model implies that any model that induces the same distribution has that property, too. To answer that, we first prove an identifiability result to characterize distribution-equivalent next-token predictors, lifting a diversity requirement of previous results. Second, based on a refinement of relational linearity [Paccanaro and Hinton, 2001; Hernandez et al., 2024], we show how many notions of linearity are amenable to our analysis. Finally, we show that under suitable conditions, these linear properties either hold in all or none distribution equivalent next-token predictors.
Emanuele Marconato, Sébastien Lachapelle, Sebastian Weichwald, Luigi Gresele
AISTATS3
2024 Adjustment Identification Distance: A gadjid for Causal Structure Learning
abstract
Evaluating graphs learned by causal discovery algorithms is difficult: The number of edges that differ between two graphs does not reflect how the graphs differ with respect to the identifying formulas they suggest for causal effects. We introduce a framework for developing causal distances between graphs which includes the structural intervention distance for directed acyclic graphs as a special case. We use this framework to develop improved adjustment-based distances as well as extensions to completed partially directed acyclic graphs and causal orders. We develop new reachability algorithms to compute the distances efficiently and to prove their low polynomial time complexity. In our package gadjid (open source at https://github.com/CausalDisco/gadjid), we provide implementations of our distances; they are orders of magnitude faster with proven lower time complexity than the structural intervention distance and thereby provide a success metric for causal discovery that scales to graph sizes that were previously prohibitive.
Leonard Henckel, Theo Würtzen, Sebastian Weichwald
UAI3
2024 <tt>spillR</tt> : spillover compensation in mass cytometry data
abstract
MOTIVATION: Channel interference in mass cytometry can cause spillover and may result in miscounting of protein markers. Chevrier et al. introduce an experimental and computational procedure to estimate and compensate for spillover implemented in their R package CATALYST. They assume spillover can be described by a spillover matrix that encodes the ratio between the signal in the unstained spillover receiving and stained spillover emitting channel. They estimate the spillover matrix from experiments with beads. We propose to skip the matrix estimation step and work directly with the full bead distributions. We develop a nonparametric finite mixture model and use the mixture components to estimate the probability of spillover. Spillover correction is often a pre-processing step followed by downstream analyses, and choosing a flexible model reduces the chance of introducing biases that can propagate downstream. RESULTS: We implement our method in an R package spillR using expectation-maximization to fit the mixture model. We test our method on simulated, semi-simulated, and real data from CATALYST. We find that our method compensates low counts accurately, does not introduce negative counts, avoids overcompensating high counts, and preserves correlations between markers that may be biologically meaningful. AVAILABILITY AND IMPLEMENTATION: Our new R package spillR is on bioconductor at bioconductor.org/packages/spillR. All experiments and plots can be reproduced by compiling the R markdown file spillR_paper.Rmd at github.com/ChristofSeiler/spillR_paper.
Marco Guazzini, Alexander G. Reisach, Sebastian Weichwald, Christof Seiler
Bioinform.3
2024 Identifying Causal Effects using Instrumental Time Series: Nuisance IV and Correcting for the Past
abstract
Instrumental variable (IV) regression relies on instruments to infer causal effects from observational data with unobserved confounding. We consider IV regression in time series models, such as vector auto-regressive (VAR) processes. Direct applications of i.i.d. techniques are generally inconsistent as they do not correctly adjust for dependencies in the past. In this paper, we outline the difficulties that arise due to time structure and propose methodology for constructing identifying equations that can be used for consistent parametric estimation of causal effects in time series data. One method uses extra nuisance covariates to obtain identifiability (an idea that can be of interest even in the i.i.d. case). We further propose a graph marginalization framework that allows us to apply nuisance IV and other IV methods in a principled way to time series. Our methods make use of a version of the global Markov property, which we prove holds for VAR(p) processes. For VAR(1) processes, we prove identifiability conditions that relate to Jordan forms and are different from the well-known rank conditions in the i.i.d. case (they do not require as many instruments as covariates, for example). We provide methods, prove their consistency, and show how the inferred causal effect can be used for distribution generalization. Simulation experiments corroborate our theoretical results. We provide ready-to-use Python code.
Nikolaj Thams, Rikke Søndergaard, Sebastian Weichwald, Jonas Peters
J. Mach. Learn. Res.3
2023 A Scale-Invariant Sorting Criterion to Find a Causal Order in Additive Noise Models
abstract
Additive Noise Models (ANMs) are a common model class for causal discovery from observational data. Due to a lack of real-world data for which an underlying ANM is known, ANMs with randomly sampled parameters are commonly used to simulate data for the evaluation of causal discovery algorithms. While some parameters may be fixed by explicit assumptions, fully specifying an ANM requires choosing all parameters. Reisach et al. (2021) show that, for many ANM parameter choices, sorting the variables by increasing variance yields an ordering close to a causal order and introduce ‘var-sortability’ to quantify this alignment. Since increasing variances may be unrealistic and cannot be exploited when data scales are arbitrary, ANM data are often rescaled to unit variance in causal discovery benchmarking. We show that synthetic ANM data are characterized by another pattern that is scale-invariant and thus persists even after standardization: the explainable fraction of a variable’s variance, as captured by the coefficient of determination $R^2$, tends to increase along the causal order. The result is high ‘$R^2$-sortability’, meaning that sorting the variables by increasing $R^2$ yields an ordering close to a causal order. We propose a computationally efficient baseline algorithm termed ‘$R^2$-SortnRegress’ that exploits high $R^2$-sortability and that can match and exceed the performance of established causal discovery algorithms. We show analytically that sufficiently high edge weights lead to a relative decrease of the noise contributions along causal chains, resulting in increasingly deterministic relationships and high $R^2$. We characterize $R^2$-sortability on synthetic data with different simulation parameters and find high values in common settings. Our findings reveal high $R^2$-sortability as an assumption about the data generating process relevant to causal discovery and implicit in many ANM sampling schemes. It should be made explicit, as its prevalence in real-world data is an open question. For causal discovery benchmarking, we provide implementations of $R^2$-sortability, the $R^2$-SortnRegress algorithm, and ANM simulation procedures in our library CausalDisco at https://causaldisco.github.io/CausalDisco/.
Alexander G. Reisach, Myriam Tami, Christof Seiler, Antoine Chambaz, Sebastian Weichwald
NeurIPS5
2021 Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to Game
abstract
Simulated DAG models may exhibit properties that, perhaps inadvertently, render their structure identifiable and unexpectedly affect structure learning algorithms. Here, we show that marginal variance tends to increase along the causal order for generically sampled additive noise models. We introduce varsortability as a measure of the agreement between the order of increasing marginal variance and the causal order. For commonly sampled graphs and model parameters, we show that the remarkable performance of some continuous structure learning algorithms can be explained by high varsortability and matched by a simple baseline method. Yet, this performance may not transfer to real-world data where varsortability may be moderate or dependent on the choice of measurement scales. On standardized data, the same algorithms fail to identify the ground-truth DAG or its Markov equivalence class. While standardization removes the pattern in marginal variance, we show that data generating processes that incur high varsortability also leave a distinct covariance pattern that may be exploited even after standardization. Our findings challenge the significance of generic benchmarks with independently drawn parameters. The code is available at https://github.com/Scriddie/Varsortability.
Alexander G. Reisach, Christof Seiler, Sebastian Weichwald
NeurIPS3
2021 Compositional abstraction error and a category of causal models
abstract
Interventional causal models describe several joint distributions over some variables used to describe a system, one for each intervention setting. They provide a formal recipe for how to move between the different joint distributions and make predictions about the variables upon intervening on the system. Yet, it is difficult to formalise how we may change the underlying variables used to describe the system, say moving from fine-grained to coarse-grained variables. Here, we argue that compositionality is a desideratum for such model transformations and the associated errors: When abstracting a reference model M iteratively, first obtaining M’ and then further simplifying that to obtain M”, we expect the composite transformation from M to M” to exist and its error to be bounded by the errors incurred by each individual transformation step. Category theory, the study of mathematical objects via compositional transformations between them, offers a natural language to develop our framework for model transformations and abstractions. We introduce a category of finite interventional causal models and, leveraging theory of enriched categories, prove the desired compositionality properties for our framework.
Eigil Fjeldgren Rischel, Sebastian Weichwald
UAI2
2019 Robustifying Independent Component Analysis by Adjusting for Group-Wise Stationary Noise
abstract
We introduce coroICA, confounding-robust independent component analysis, a novel ICA algorithm which decomposes linearly mixed multivariate observations into independent components that are corrupted (and rendered dependent) by hidden group-wise stationary confounding. It extends the ordinary ICA model in a theoretically sound and explicit way to incorporate group-wise (or environment-wise) confounding. We show that our proposed general noise model allows to perform ICA in settings where other noisy ICA procedures fail. Additionally, it can be used for applications with grouped data by adjusting for different stationary noise within each group. Our proposed noise model has a natural relation to causality and we explain how it can be applied in the context of causal inference. In addition to our theoretical framework, we provide an efficient estimation procedure and prove identifiability of the unmixing matrix under mild assumptions. Finally, we illustrate the performance and robustness of our method on simulated data, provide audible and visual examples, and demonstrate the applicability to real-world scenarios by experiments on publicly available Antarctic ice core data as well as two EEG data sets. We provide a scikit-learn compatible pip-installable Python package coroICA as well as R and Matlab implementations accompanied by a documentation at https://sweichwald.de/coroICA/
Niklas Pfister, Sebastian Weichwald, Peter Bühlmann, Bernhard Schölkopf
J. Mach. Learn. Res.2
2017 Personalized brain-computer interface models for motor rehabilitation
abstract
We propose to fuse two currently separate research lines on novel therapies for stroke rehabilitation: brain-computer interface (BCI) training and transcranial electrical stimulation (TES). Specifically, we show that BCI technology can be used to learn personalized decoding models that relate the global configuration of brain rhythms in individual subjects (as measured by EEG) to their motor performance during 3D reaching movements. We demonstrate that our models capture substantial across-subject heterogeneity, and argue that this heterogeneity is a likely cause of limited effect sizes observed in TES for enhancing motor performance. We conclude by discussing how our personalized models can be used to derive optimal TES parameters, e.g., stimulation site and frequency, for individual patients.
Atalanti-Anastasia Mastakouri, Sebastian Weichwald, Ozan Özdenizci, Timm Meyer, Bernhard Schölkopf, Moritz Grosse-Wentrup
SMC2
2017 Causal Consistency of Structural Equation Models
Paul K. Rubenstein, Sebastian Weichwald, Stephan Bongers, Joris M. Mooij, Dominik Janzing, Moritz Grosse-Wentrup, Bernhard Schölkopf
UAI2
2016 Pymanopt: A Python Toolbox for Optimization on Manifolds using Automatic Differentiation
abstract
Optimization on manifolds is a class of methods for optimization of an objective function, subject to constraints which are smooth, in the sense that the set of points which satisfy the constraints admits the structure of a differentiable manifold. While many optimization problems are of the described form, technicalities of differential geometry and the laborious calculation of derivatives pose a significant barrier for experimenting with these methods. We introduce Pymanopt (available at pymanopt.github.io), a toolbox for optimization on manifolds, implemented in Python, that---similarly to the Manopt Matlab toolbox---implements several manifold geometries and optimization algorithms. Moreover, we lower the barriers to users further by using automated differentiation for calculating derivative information, saving users time and saving them from potential calculation and implementation errors.
James Townsend, Niklas Koep, Sebastian Weichwald
J. Mach. Learn. Res.3