EDBT 2026 Demo / reviewers in the wild / expert
Dominik Janzing
dblp:17/6280
· DBLP profile ↗
68ranked-venue papers
11as first author
21since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Theory of computation · 2 · 2 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks
Fabien Chraim, Dominik Janzing, John Evans |
ICC | 2 |
| 2025 | Toward Falsifying Causal Graphs Using a Permutation-Based TestabstractUnderstanding causal relationships among the variables of a system is paramount to explain and control its behavior. For many real-world systems, however, the true causal graph is not readily available and one must resort to predictions made by algorithms or domain experts. Therefore, metrics that quantitatively assess the goodness of a causal graph provide helpful checks before using it in downstream tasks. Existing metrics provide an absolute number of inconsistencies between the graph and the observed data, and without a baseline, practitioners are left to answer the hard question of how many such inconsistencies are acceptable or expected. Here, we propose a novel consistency metric by constructing a baseline through node permutations. By comparing the number of inconsistencies with those on the baseline, we derive an interpretable metric that captures whether the graph is significantly better than random. Evaluating on both simulated and real data sets from various domains, including biology and cloud monitoring, we demonstrate that the true graph is not falsified by our metric, whereas the wrong graphs given by a hypothetical user are likely to be falsified. Elias Eulig, Atalanti-Anastasia Mastakouri, Patrick Blöbaum, Michaela Hardt, Dominik Janzing |
AAAI | 5 |
| 2025 | Root Cause Analysis of Outliers with Missing Structural KnowledgeabstractThe goal of Root Cause Analysis (RCA) is to explain why an anomaly occurred by identifying where the fault originated. Several recent works model the anomalous event as resulting from a change in the causal mechanism at the root cause, i.e., as a soft intervention. RCA is then the task of identifying which causal mechanism changed. In real-world applications, one often has either few or only a single sample from the post-intervention distribution: a severe limitation for most methods, which assume one knows or can estimate the distribution. However, even those that do not are statistically ill-posed due to the need to probe regression models in regions of low probability density. In this paper, we propose simple, efficient methods to overcome both difficulties in the case where there is a single root cause and the causal graph is a polytree. When one knows the causal graph, we give guarantees for a traversal algorithm that requires only marginal anomaly scores and does not depend on specifying an arbitrary anomaly score cut-off. When one does not know the causal graph, we show that the heuristic of identifying root causes as the variables with the highest marginal anomaly scores is causally justified. To this end, we prove that anomalies with small scores are unlikely to cause those with larger scores in polytrees and give upper bounds for the likelihood of causal pathways with non-monotonic anomaly scores. William Roy Orchard, Nastaran Okati, Sergio Hernan Garrido Mejia, Patrick Blöbaum, Dominik Janzing |
NeurIPS | 5 |
| 2025 | Toward Universal Laws of Outlier PropagationabstractWhen a variety of anomalous features motivate flagging different samples as *outliers*, Algorithmic Information Theory (AIT) offers a principled way to unify them in terms of a sample’s *randomness deficiency*. Subject to the Independence of Mechanisms Principle, we show that for a joint sample on the nodes of a causal Bayesian network, the randomness deficiency decomposes into a sum of randomness deficiencies at each causal mechanism. Consequently, anomalous observations can be attributed to their root causes, i.e., the mechanisms that behaved anomalously. As an extension of Levin’s law of randomness conservation, we show that weak outliers cannot cause strong ones. We show how these information theoretic laws clarify our understanding of outlier detection and attribution, in the context of more specialized outlier scores from prior literature. Aram Ebtekar, Dominik Janzing |
UAI | 3 |
| 2024 | Self-Compatibility: Evaluating Causal Discovery without Ground Truth
Philipp Michael Faller, Leena C. Vankadara, Atalanti-Anastasia Mastakouri, Francesco Locatello, Dominik Janzing |
AISTATS | 5 |
| 2024 | Quantifying intrinsic causal contributions via structure preserving interventionsabstractWe propose a notion of causal influence that describes the ‘intrinsic’ part of the contribution of a node on a target node in a DAG. By recursively writing each node as a function of the upstream noise terms, we separate the intrinsic information added by each node from the one obtained from its ancestors. To interpret the intrinsic information as a causal contribution, we consider ‘structure-preserving interventions’ that randomize each node in a way that mimics the usual dependence on the parents and does not perturb the observed joint distribution. To get a measure that is invariant across arbitrary orderings of nodes we use Shapley based symmetrization and show that it reduces in the linear case to simple ANOVA after resolving the target node into noise variables. We describe our contribution analysis for variance and entropy, but contributions for other target metrics can be defined analogously. Dominik Janzing, Patrick Blöbaum, Atalanti-Anastasia Mastakouri, Philipp Michael Faller, Lenon Minorics, Kailash Budhathoki |
AISTATS | 1 |
| 2024 | Causal vs. Anticausal merging of predictorsabstractWe study the differences arising from merging predictors in the causal and anticausal directions using the same data.
In particular we study the asymmetries that arise in a simple model where we merge the predictors using one binary variable as target and two continuous variables as predictors.
We use Causal Maximum Entropy (CMAXENT) as inductive bias to merge the predictors, however, we expect similar differences to hold also when we use other merging methods that take into account asymmetries between cause and effect.
We show that if we observe all bivariate distributions, the CMAXENT solution reduces to a logistic regression in the causal direction and Linear Discriminant Analysis (LDA) in the anticausal direction.
Furthermore, we study how the decision boundaries of these two solutions differ whenever we observe only some of the bivariate distributions implications for Out-Of-Variable (OOV) generalisation. Sergio Hernan Garrido Mejia, Patrick Blöbaum, Bernhard Schölkopf, Dominik Janzing |
NeurIPS | 4 |
| 2024 | DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal ModelsabstractWe present DoWhy-GCM, an extension of the DoWhy Python library, which leverages graphical causal models. Unlike existing causality libraries, which mainly focus on effect estimation, DoWhy-GCM addresses diverse causal queries, such as identifying the root causes of outliers and distributional changes, attributing causal influences to the data generating process of each node, or diagnosis of causal structures. With DoWhy-GCM, users typically specify cause-effect relations via a causal graph, fit causal mechanisms, and pose causal queries---all with just a few lines of code. The general documentation is available at https://www.pywhy.org/dowhy and the DoWhy-GCM specific code at https://github.com/py-why/dowhy/tree/main/dowhy/gcm. Patrick Blöbaum, Peter Götz, Kailash Budhathoki, Atalanti-Anastasia Mastakouri, Dominik Janzing |
J. Mach. Learn. Res. | 5 |
| 2023 | Assumption violations in causal discovery and the robustness of score matchingabstractWhen domain knowledge is limited and experimentation is restricted by ethical, financial, or time constraints, practitioners turn to observational causal discovery methods to recover the causal structure, exploiting the statistical properties of their data. Because causal discovery without further assumptions is an ill-posed problem, each algorithm comes with its own set of usually untestable assumptions, some of which are hard to meet in real datasets. Motivated by these considerations, this paper extensively benchmarks the empirical performance of recent causal discovery methods on observational _iid_ data generated under different background conditions, allowing for violations of the critical assumptions required by each selected approach.
Our experimental findings show that score matching-based methods demonstrate surprising performance in the false positive and false negative rate of the inferred graph in these challenging scenarios, and we provide theoretical insights into their performance. This work is also the first effort to benchmark the stability of causal discovery algorithms with respect to the values of their hyperparameters. Finally, we hope this paper will set a new standard for the evaluation of causal discovery methods and can serve as an accessible entry point for practitioners interested in the field, highlighting the empirical implications of different algorithm choices. Francesco Montagna, Atalanti-Anastasia Mastakouri, Elias Eulig, Nicoletta Noceti, Lorenzo Rosasco, Dominik Janzing, Bryon Aragam, Francesco Locatello |
NeurIPS | 6 |
| 2023 | Causal information splitting: Engineering proxy features for robustness to distribution shiftsabstractStatistical prediction models are often trained on data that is drawn from different probability distributions than their eventual use cases. One approach to proactively prepare for these shifts harnesses the intuition that causal mechanisms should remain invariant between environments. Here we focus on a challenging setting in which the causal and anticausal variables of the target are unobserved. Leaning on information theory, we develop feature selection and engineering techniques for the observed downstream variables that act as proxies. We identify proxies that help to build stable models and moreover utilize auxiliary training tasks to extract stability-enhancing information from proxies. We demonstrate the effectiveness of our techniques on synthetic and real data. Bijan Mazaheri, Atalanti-Anastasia Mastakouri, Dominik Janzing, Michaela Hardt |
UAI | 3 |
| 2022 | Obtaining Causal Information by Merging Datasets with MAXENTabstractThe investigation of the question "which treatment has a causal effect on a target variable?" is of particular relevance in a large number of scientific disciplines. This challenging task becomes even more difficult if not all treatment variables were or even can not be observed jointly with the target variable. In this paper, we discuss how causal knowledge can be obtained without having observed all variables jointly, but by merging the statistical information from different datasets. We show how the maximum entropy principle can be used to identify edges among random variables when assuming causal sufficiency and an extended version of faithfulness, and when only subsets of the variables have been observed jointly. Sergio Hernan Garrido Mejia, Elke Kirschbaum, Dominik Janzing |
AISTATS | 3 |
| 2022 | Testing Granger Non-Causality in Panels with Cross-Sectional DependenciesabstractThis paper proposes a new approach for testing Granger non-causality on panel data. Instead of aggregating panel member statistics, we aggregate their corresponding p-values and show that the resulting p-value approximately bounds the type I error by the chosen significance level even if the panel members are dependent. We compare our approach against the most widely used Granger causality algorithm on panel data and show that our approach yields lower FDR at the same power for large sample sizes and panels with cross sectional dependencies. Finally, we examine COVID-19 data about confirmed cases and deaths measured in countries/regions worldwide and show that our approach is able to discover the true causal relation between confirmed cases and deaths while state-of-the-art approaches fail. Lenon Minorics, Ali Caner Türkmen, David Kernert, Patrick Blöbaum, Laurent Callot, Dominik Janzing |
AISTATS | 6 |
| 2022 | You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory Prediction
Osama Makansi, Julius von Kügelgen, Francesco Locatello, Peter V. Gehler, Dominik Janzing, Thomas Brox, Bernhard Schölkopf |
ICLR | 5 |
| 2022 | Causal structure-based root cause analysis of outliersabstractCurrent techniques for explaining outliers cannot tell what caused the outliers. We present a formal method to identify "root causes" of outliers, amongst variables. The method requires a causal graph of the variables along with the functional causal model. It quantifies the contribution of each variable to the target outlier score, which explains to what extent each variable is a "root cause" of the target outlier. We study the empirical performance of the method through simulations and present a real-world case study identifying "root causes" of extreme river flows. Kailash Budhathoki, Lenon Minorics, Patrick Blöbaum, Dominik Janzing |
ICML | 4 |
| 2022 | Causal Inference Through the Structural Causal Marginal ProblemabstractWe introduce an approach to counterfactual inference based on merging information from multiple datasets. We consider a causal reformulation of the statistical marginal problem: given a collection of marginal structural causal models (SCMs) over distinct but overlapping sets of variables, determine the set of joint SCMs that are counterfactually consistent with the marginal ones. We formalise this approach for categorical SCMs using the response function formulation and show that it reduces the space of allowed marginal and joint SCMs. Our work thus highlights a new mode of falsifiability through additional variables, in contrast to the statistical one via additional data. Luigi Gresele, Julius von Kügelgen, Jonas M. Kübler, Elke Kirschbaum, Bernhard Schölkopf, Dominik Janzing |
ICML | 6 |
| 2022 | On Measuring Causal Contributions via do-interventionsabstractCausal contributions measure the strengths of different causes to a target quantity. Understanding causal contributions is important in empirical sciences and data-driven disciplines since it allows to answer practical queries like “what are the contributions of each cause to the effect?” In this paper, we develop a principled method for quantifying causal contributions. First, we provide desiderata of properties axioms that causal contribution measures should satisfy and propose the do-Shapley values (inspired by do-interventions [Pearl, 2000]) as a unique method satisfying these properties. Next, we develop a criterion under which the do-Shapley values can be efficiently inferred from non-experimental data. Finally, we provide do-Shapley estimators exhibiting consistency, computational feasibility, and statistical robustness. Simulation results corroborate with the theory. Yonghan Jung, Shiva Prasad Kasiviswanathan, Jin Tian 0001, Dominik Janzing, Patrick Blöbaum, Elias Bareinboim |
ICML | 4 |
| 2022 | Score Matching Enables Causal Discovery of Nonlinear Additive Noise ModelsabstractThis paper demonstrates how to recover causal graphs from the score of the data distribution in non-linear additive (Gaussian) noise models. Using score matching algorithms as a building block, we show how to design a new generation of scalable causal discovery methods. To showcase our approach, we also propose a new efficient method for approximating the score’s Jacobian, enabling to recover the causal graph. Empirically, we find that the new algorithm, called SCORE, is competitive with state-of-the-art causal discovery methods while being significantly faster. Paul Rolland, Volkan Cevher, Matthäus Kleindessner, Chris Russell 0001, Dominik Janzing, Bernhard Schölkopf, Francesco Locatello |
ICML | 5 |
| 2022 | Causal forecasting: generalization bounds for autoregressive modelsabstractDespite the increasing relevance of forecasting methods, causal implications of these algorithms remain largely unexplored. This is concerning considering that, even under simplifying assumptions such as causal sufficiency, the statistical risk of a model can differ significantly from its causal risk. Here, we study the problem of causal generalization—generalizing from the observational to interventional distributions—in forecasting. Our goal is to find answers to the question: How does the efficacy of an autoregressive (VAR) model in predicting statistical associations compare with its ability to predict under interventions? To this end, we introduce the framework of causal learning theory for forecasting. Using this framework, we obtain a characterization of the difference between statistical and causal risks, which helps identify sources of divergence between them. Under causal sufficiency, the problem of causal generalization amounts to learning under covariate shifts albeit with additional structure (restriction to interventional distributions under the VAR model). This structure allows us to obtain uniform convergence bounds on causal generalizability for the class of VAR models. To the best of our knowledge, this is the first work that provides theoretical guarantees for causal generalization in the time-series setting. Leena C. Vankadara, Philipp Michael Faller, Michaela Hardt, Lenon Minorics, Debarghya Ghoshdastidar, Dominik Janzing |
UAI | 6 |
| 2021 | A Theory of Independent Mechanisms for Extrapolation in Generative ModelsabstractGenerative models can be trained to emulate complex empirical data, but are they useful to make predictions in the context of previously unobserved environments? An intuitive idea to promote such extrapolation capabilities is to have the architecture of such model reflect a causal graph of the true data generating process, such that one can intervene on each node independently of the others. However, the nodes of this graph are usually unobserved, leading to overparameterization and lack of identifiability of the causal structure. We develop a theoretical framework to address this challenging situation by defining a weaker form of identifiability, based on the principle of independence of mechanisms. We demonstrate on toy examples that classical stochastic gradient descent can hinder the model's extrapolation capabilities, suggesting independence of mechanisms should be enforced explicitly during training. Experiments on deep generative models trained on real world data support these insights and illustrate how the extrapolation capabilities of such models can be leveraged. Michel Besserve, Rémy Sun, Dominik Janzing, Bernhard Schölkopf |
AAAI | 3 |
| 2021 | Why did the distribution change?abstractWe describe a formal approach based on graphical causal models to identify the "root causes" of the change in the probability distribution of variables. After factorizing the joint distribution into conditional distributions of each variable, given its parents (the "causal mechanisms"), we attribute the change to changes of these causal mechanisms. This attribution analysis accounts for the fact that mechanisms often change independently and sometimes only some of them change. Through simulations, we study the performance of our distribution change attribution proposal. We then present a real-world case study identifying the drivers of the difference in the income distribution between men and women. Kailash Budhathoki, Dominik Janzing, Patrick Blöbaum, Hoiyi Ng |
AISTATS | 2 |
| 2021 | Necessary and sufficient conditions for causal feature selection in time series with latent common causesabstractWe study the identification of direct and indirect causes on time series with latent variables, and provide a constrained-based causal feature selection method, which we prove that is both sound and complete under some graph constraints. Our theory and estimation algorithm require only two conditional independence tests for each observed candidate time series to determine whether or not it is a cause of an observed target time series. Furthermore, our selection of the conditioning set is such that it improves signal to noise ratio. We apply our method on real data, and on a wide range of simulated experiments, which yield very low false positive and relatively low false negative rates. Atalanti-Anastasia Mastakouri, Bernhard Schölkopf, Dominik Janzing |
ICML | 3 |
| 2020 | Feature relevance quantification in explainable AI: A causal problemabstractWe discuss promising recent contributions on quantifying feature relevance using Shapley values, where we observed some confusion on which probability distribution is the right one for dropped features. We argue that the confusion is based on not carefully distinguishing between observational and interventional conditional probabilities and try a clarification based on Pearl’s seminal work on causality. We conclude that unconditional rather than conditional expectations provide the right notion of dropping features. This contradicts the view of the authors of the software package SHAP. In that work, unconditional expectations (which we argue to be conceptually right) are only used as approximation for the conditional ones, which encouraged others to ’improve’ SHAP in a way that we believe to be flawed. Dominik Janzing, Lenon Minorics, Patrick Blöbaum |
AISTATS | 1 |
| 2019 | Causal RegularizationabstractWe argue that regularizing terms in standard regression methods not only help against overfitting finite data, but sometimes also help in getting better causal models. We first consider a multi-dimensional variable linearly influencing a target variable with some multi-dimensional unobserved common cause, where the confounding effect can be decreased by keeping the penalizing term in Ridge and Lasso regression even in the population limit. The reason is a close analogy between overfitting and confounding observed for our toy model. In the case of overfitting, we can choose regularization constants via cross validation, but here we choose the regularization constant by first estimating the strength of confounding, which yielded reasonable results for simulated and real data. Further, we show a ‘causal generalization bound’ which states (subject to our particular model of confounding) that the error made by interpreting any non-linear regression as causal model can be bounded from above whenever functions are taken from a not too rich class. Dominik Janzing |
NeurIPS | 1 |
| 2019 | Selecting causal brain features with a single conditional independence test per featureabstractWe propose a constraint-based causal feature selection method for identifying causes of a given target variable, selecting from a set of candidate variables, while there can also be hidden variables acting as common causes with the target. We prove that if we observe a cause for each candidate cause, then a single conditional independence test with one conditioning variable is sufficient to decide whether a candidate associated with the target is indeed causing it. We thus improve upon existing methods by significantly simplifying statistical testing and requiring a weaker version of causal faithfulness. Our main assumption is inspired by neuroscience paradigms where the activity of a single neuron is considered to be also caused by its own previous state. We demonstrate successful application of our method to simulated, as well as encephalographic data of twenty-one participants, recorded in Max Planck Institute for intelligent Systems. The detected causes of motor performance are in accordance with the latest consensus about the neurophysiological pathways, and can provide new insights into personalised brain stimulation. Atalanti-Anastasia Mastakouri, Bernhard Schölkopf, Dominik Janzing |
NeurIPS | 3 |
| 2019 | Perceiving the arrow of time in autoregressive motionabstractUnderstanding the principles of causal inference in the visual system has a long history at least since the seminal studies by Albert Michotte. Many cognitive and machine learning scientists believe that intelligent behavior requires agents to possess causal models of the world. Recent ML algorithms exploit the dependence structure of additive noise terms for inferring causal structures from observational data, e.g. to detect the direction of time series; the arrow of time. This raises the question whether the subtle asymmetries between the time directions can also be perceived by humans. Here we show that human observers can indeed discriminate forward and backward autoregressive motion with non-Gaussian additive independent noise, i.e. they appear sensitive to subtle asymmetries between the time directions. We employ a so-called frozen noise paradigm enabling us to compare human performance with four different algorithms on a trial-by-trial basis: A causal inference algorithm exploiting the dependence structure of additive noise terms, a neurally inspired network, a Bayesian ideal observer model as well as a simple heuristic. Our results suggest that all human observers use similar cues or strategies to solve the arrow of time motion discrimination task, but the human algorithm is significantly different from the three machine algorithms we compared it to. In fact, our simple heuristic appears most similar to our human observers. Kristof Meding, Dominik Janzing, Bernhard Schölkopf, Felix A. Wichmann |
NeurIPS | 2 |
| 2018 | Group invariance principles for causal generative modelsabstractThe postulate of independence of cause and mechanism (ICM) has recently led to several new causal discovery algorithms. The interpretation of independence and the way it is utilized, however, varies across these methods. Our aim in this paper is to propose a group theoretic framework for ICM to unify and generalize these approaches. In our setting, the cause-mechanism relationship is assessed by perturbing it with random group transformations. We show that the group theoretic view encompasses previous ICM approaches and provides a very general tool to study the structure of data generating mechanisms with direct applications to machine learning. Michel Besserve, Naji Shajarisales, Bernhard Schölkopf, Dominik Janzing |
AISTATS | 4 |
| 2018 | Cause-Effect Inference by Comparing Regression ErrorsabstractWe address the problem of inferring the causal relation between two variables by comparing the least-squares errors of the predictions in both possible causal directions. Under the assumption of an independence between the function relating cause and effect, the conditional noise distribution, and the distribution of the cause, we show that the errors are smaller in causal direction if both variables are equally scaled and the causal relation is close to deterministic. Based on this, we provide an easily applicable method that only requires a regression in both possible causal directions. The performance of this method is compared with different related causal inference methods in various artificial and real-world data sets. Patrick Blöbaum, Dominik Janzing, Takashi Washio, Shohei Shimizu, Bernhard Schölkopf |
AISTATS | 2 |
| 2018 | Detecting non-causal artifacts in multivariate linear regression modelsabstractWe consider linear models where d potential causes X_1,...,X_d are correlated with one target quantity Y and propose a method to infer whether the association is causal or whether it is an artifact caused by overfitting or hidden common causes. We employ the idea that in the former case the vector of regression coefficients has ‘generic’ orientation relative to the covariance matrix Sigma_{XX} of X. Using an ICA based model for confounding, we show that both confounding and overfitting yield regression vectors that concentrate mainly in the space of low eigenvalues of Sigma_{XX}. Dominik Janzing, Bernhard Schölkopf |
ICML | 1 |
| 2017 | Avoiding Discrimination through Causal ReasoningabstractRecent work on fairness in machine learning has focused on various statistical discrimination criteria and how they trade off. Most of these criteria are observational: They depend only on the joint distribution of predictor, protected attribute, features, and outcome. While convenient to work with, observational criteria have severe inherent limitations that prevent them from resolving matters of fairness conclusively. Going beyond observational criteria, we frame the problem of discrimination based on protected attributes in the language of causal reasoning. This viewpoint shifts attention from "What is the right fairness criterion?" to "What do we want to assume about our model of the causal data generating process?" Through the lens of causality, we make several contributions. First, we crisply articulate why and when observational criteria fail, thus formalizing what was before a matter of opinion. Second, our approach exposes previously ignored subtleties and why they are fundamental to the problem. Finally, we put forward natural causal non-discrimination criteria and develop algorithms that satisfy them. Niki Kilbertus, Mateo Rojas-Carulla, Giambattista Parascandolo, Moritz Hardt, Dominik Janzing, Bernhard Schölkopf |
NIPS | 5 |
| 2017 | Causal Consistency of Structural Equation Models
Paul K. Rubenstein, Sebastian Weichwald, Stephan Bongers, Joris M. Mooij, Dominik Janzing, Moritz Grosse-Wentrup, Bernhard Schölkopf |
UAI | 5 |
| 2016 | Distinguishing Cause from Effect Using Observational Data: Methods and BenchmarksabstractThe discovery of causal relationships from purely observational data is a fundamental problem in science. The most elementary form of such a causal discovery problem is to decide whether $X$ causes $Y$ or, alternatively, $Y$ causes $X$, given joint observations of two variables $X,Y$. An example is to decide whether altitude causes temperature, or vice versa, given only joint measurements of both variables. Even under the simplifying assumptions of no confounding, no feedback loops, and no selection bias, such bivariate causal discovery problems are challenging. Nevertheless, several approaches for addressing those problems have been proposed in recent years. We review two families of such methods: methods based on Additive Noise Models (ANMs) and Information Geometric Causal Inference (IGCI). We present the benchmark CauseEffectPairs that consists of data for 100 different cause-effect pairs selected from 37 data sets from various domains (e.g., meteorology, biology, medicine, engineering, economy, etc.) and motivate our decisions regarding the ground truth causal directions of all pairs. We evaluate the performance of several bivariate causal discovery methods on these real-world benchmark data and in addition on artificially simulated data. Our empirical results on real-world data indicate that certain methods are indeed able to distinguish cause from effect using only purely observational data, although more benchmark data would be needed to obtain statistically significant conclusions. One of the best performing methods overall is the method based on Additive Noise Models that has originally been proposed by Hoyer et al. (2009), which obtains an accuracy of 63 $\pm$ 10 % and an AUC of 0.74 $\pm$ 0.05 on the real-world benchmark. As the main theoretical contribution of this work we prove the consistency of that method. Joris M. Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, Bernhard Schölkopf |
J. Mach. Learn. Res. | 3 |
| 2015 | Inference of Cause and Effect with Unsupervised Inverse RegressionabstractWe address the problem of causal discovery in the two-variable case given a sample from their joint distribution. The proposed method is based on a known assumption that, if X -> Y (X causes Y), the marginal distribution of the cause, P(X), contains no information about the conditional distribution P(Y|X). Consequently, estimating P(Y|X) from P(X) should not be possible. However, estimating P(X|Y) based on P(Y) may be possible. This paper employs this asymmetry to propose CURE, a causal discovery method which decides upon the causal direction by comparing the accuracy of the estimations of P(Y|X) and P(X|Y). To this end, we propose a method for estimating a conditional from samples of the corresponding marginal, which we call unsupervised inverse GP regression. We evaluate CURE on synthetic and real data. On the latter, our method outperforms existing causal inference methods. Eleni Sgouritsa, Dominik Janzing, Philipp Hennig, Bernhard Schölkopf |
AISTATS | 2 |
| 2015 | Causal Inference by Identification of Vector Autoregressive Processes with Hidden ComponentsabstractA widely applied approach to causal inference from a time series X, often referred to as “(linear) Granger causal analysis”, is to simply regress present on past and interpret the regression matrix \hatB causally. However, if there is an unmeasured time series Z that influences X, then this approach can lead to wrong causal conclusions, i.e., distinct from those one would draw if one had additional information such as Z. In this paper we take a different approach: We assume that X together with some hidden Z forms a first order vector autoregressive (VAR) process with transition matrix A, and argue why it is more valid to interpret A causally instead of \hatB. Then we examine under which conditions the most important parts of A are identifiable or almost identifiable from only X. Essentially, sufficient conditions are (1) non-Gaussian, independent noise or (2) no influence from X to Z. We present two estimation algorithms that are tailored towards conditions (1) and (2), respectively, and evaluate them on synthetic and real-world data. We discuss how to check the model using X. Philipp Geiger, Kun Zhang 0001, Bernhard Schölkopf, Mingming Gong, Dominik Janzing |
ICML | 5 |
| 2015 | Removing systematic errors for exoplanet search via latent causesabstractWe describe a method for removing the effect of confounders in order to reconstruct a latent quantity of interest. The method, referred to as half-sibling regression, is inspired by recent work in causal inference using additive noise models. We provide a theoretical justification and illustrate the potential of the method in a challenging astronomy application. Bernhard Schölkopf, David W. Hogg, Dun Wang, Daniel Foreman-Mackey, Dominik Janzing, Carl-Johann Simon-Gabriel, Jonas Peters |
ICML | 5 |
| 2015 | Telling cause from effect in deterministic linear dynamical systemsabstractTelling a cause from its effect using observed time series data is a major challenge in natural and social sciences. Assuming the effect is generated by the cause through a linear system, we propose a new approach based on the hypothesis that nature chooses the “cause” and the “mechanism generating the effect from the cause” independently of each other. Specifically we postulate that the power spectrum of the “cause” time series is uncorrelated with the square of the frequency response of the linear filter (system) generating the effect. While most causal discovery methods for time series mainly rely on the noise, our method relies on asymmetries of the power spectral density properties that exist even in deterministic systems. We describe mathematical assumptions in a deterministic model under which the causal direction is identifiable. In particular, we show a scenario where the method works but Granger causality fails. Experiments show encouraging results on synthetic as well as real-world data. Overall, this suggests that the postulate of Independence of Cause and Mechanism is a promising principle for causal inference on observed time series. Naji Shajarisales, Dominik Janzing, Bernhard Schölkopf, Michel Besserve |
ICML | 2 |
| 2015 | Semi-supervised interpolation in an anticausal learning scenario
Dominik Janzing, Bernhard Schölkopf |
J. Mach. Learn. Res. | 1 |
| 2014 | Consistency of Causal Inference under the Additive Noise ModelabstractWe analyze a family of methods for statistical causal inference from sample under the so-called Additive Noise Model. While most work on the subject has concentrated on establishing the soundness of the Additive Noise Model, the statistical consistency of the resulting inference methods has received little attention. We derive general conditions under which the given family of inference methods consistently infers the causal direction in a nonparametric setting. Samory Kpotufe, Eleni Sgouritsa, Dominik Janzing, Bernhard Schölkopf |
ICML | 3 |
| 2014 | Inferring latent structures via information inequalities
Rafael Chaves, Lukas Luft, Thiago O. Maciel, David Gross 0003, Dominik Janzing, Bernhard Schölkopf |
UAI | 5 |
| 2014 | Estimating Causal Effects by Bounding Confounding
Philipp Geiger, Dominik Janzing, Bernhard Schölkopf |
UAI | 2 |
| 2014 | Causal discovery with continuous additive noise models
Jonas Peters, Joris M. Mooij, Dominik Janzing, Bernhard Schölkopf |
J. Mach. Learn. Res. | 3 |
| 2013 | Causal Inference on Time Series using Restricted Structural Equation ModelsabstractCausal inference uses observational data to infer the causal structure of the data generating system. We study a class of restricted Structural Equation Models for time series that we call Time Series Models with Independent Noise (TiMINo). These models require independent residual time series, whereas traditional methods like Granger causality exploit the variance of residuals. This work contains two main contributions: (1) Theoretical: By restricting the model class (e.g. to additive noise) we provide more general identifiability results than existing ones. The results cover lagged and instantaneous effects that can be nonlinear and unfaithful, and non-instantaneous feedbacks between the time series. (2) Practical: If there are no feedback loops between time series, we propose an algorithm based on non-linear independence tests of time series. When the data are causally insufficient, or the data generating process does not satisfy the model assumptions, this algorithm may still give partial results, but mostly avoids incorrect answers. The Structural Equation Model point of view allows us to extend both the theoretical and the algorithmic part to situations in which the time series have been measured with different time delays (as may happen for fMRI data, for example). TiMINo outperforms existing methods on artificial and real data. Code is provided. Jonas Peters, Dominik Janzing, Bernhard Schölkopf |
NIPS | 2 |
| 2013 | From Ordinary Differential Equations to Structural Causal Models: the deterministic case
Joris M. Mooij, Dominik Janzing, Bernhard Schölkopf |
UAI | 2 |
| 2013 | Identifying Finite Mixtures of Nonparametric Product Distributions and Causal Inference of Confounders
Eleni Sgouritsa, Dominik Janzing, Jonas Peters, Bernhard Schölkopf |
UAI | 2 |
| 2012 | On causal and anticausal learning
Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang 0001, Joris M. Mooij |
ICML | 2 |
| 2012 | Information-geometric approach to inferring causal directions
Dominik Janzing, Joris M. Mooij, Kun Zhang 0001, Jan Lemeire, Jakob Zscheischler, Povilas Daniusis, Bastian Steudel, Bernhard Schölkopf |
Artif. Intell. | 1 |
| 2011 | Finding dependencies between frequencies with the kernel cross-spectral densityabstractCross-spectral density (CSD), is widely used to find linear dependency between two real or complex valued time series. We define a non-linear extension of this measure by mapping the time series into two Reproducing Kernel Hilbert Spaces. The dependency is quantified by the Hilbert Schmidt norm of a cross-spectral density operator between these two spaces. We prove that, by choosing a characteristic kernel for the mapping, this quantity detects any pairwise dependency between the time series. Then we provide a fast estimator for the Hilbert-Schmidt norm based on the Fast Fourier Trans form. We demonstrate the interest of this approach to quantify non-linear dependencies between frequency bands of simulated signals and intra-cortical neural recordings. Michel Besserve, Dominik Janzing, Nikos K. Logothetis, Bernhard Schölkopf |
ICASSP | 2 |
| 2011 | On Causal Discovery with Cyclic Additive Noise ModelsabstractWe study a particular class of cyclic causal models, where each variable is a (possibly nonlinear) function of its parents and additive noise. We prove that the causal graph of such models is generically identifiable in the bivariate, Gaussian-noise case. We also propose a method to learn such models from observational data. In the acyclic case, the method reduces to ordinary regression, but in the more challenging cyclic case, an additional term arises in the loss function, which makes it a special case of nonlinear independent component analysis. We illustrate the proposed method on synthetic data. Joris M. Mooij, Dominik Janzing, Tom Heskes, Bernhard Schölkopf |
NIPS | 2 |
| 2011 | Detecting low-complexity unobserved causes
Dominik Janzing, Eleni Sgouritsa, Oliver Stegle, Jonas Peters, Bernhard Schölkopf |
UAI | 1 |
| 2011 | Identifiability of Causal Graphs using Functional Models
Jonas Peters, Joris M. Mooij, Dominik Janzing, Bernhard Schölkopf |
UAI | 3 |
| 2011 | Kernel-based Conditional Independence Test and Application in Causal Discovery
Kun Zhang 0001, Jonas Peters, Dominik Janzing, Bernhard Schölkopf |
UAI | 3 |
| 2011 | Testing whether linear equations are causal: A free probability theory approach
Jakob Zscheischler, Dominik Janzing, Kun Zhang 0001 |
UAI | 2 |
| 2011 | Causal Inference on Discrete Data Using Additive Noise ModelsabstractInferring the causal structure of a set of random variables from a finite sample of the joint distribution is an important problem in science. The case of two random variables is particularly challenging since no (conditional) independences can be exploited. Recent methods that are based on additive noise models suggest the following principle: Whenever the joint distribution P((X,Y)) admits such a model in one direction, e.g., Y = f(X)+N, N ⊥ X, but does not admit the reversed model X=g(Y)+Ñ, Ñ ⊥ Y, one infers the former direction to be causal (i.e., X → Y). Up to now, these approaches only dealt with continuous variables. In many situations, however, the variables of interest are discrete or even have only finitely many states. In this work, we extend the notion of additive noise models to these cases. We prove that it almost never occurs that additive noise models can be fit in both directions. We further propose an efficient algorithm that is able to perform this way of causal inference on finite samples of discrete variables. We show that the algorithm works on both synthetic and real data sets. Jonas Peters, Dominik Janzing, Bernhard Schölkopf |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Causal Markov Condition for Submodular Information Measures
Bastian Steudel, Dominik Janzing, Bernhard Schölkopf |
COLT | 2 |
| 2010 | Telling cause from effect based on high-dimensional observations
Dominik Janzing, Patrik O. Hoyer, Bernhard Schölkopf |
ICML | 1 |
| 2010 | Probabilistic latent variable models for distinguishing between cause and effectabstractWe propose a novel method for inferring whether X causes Y or vice versa from joint observations of X and Y. The basic idea is to model the observed data using probabilistic latent variable models, which incorporate the effects of unobserved noise. To this end, we consider the hypothetical effect variable to be a function of the hypothetical cause variable and an independent noise term (not necessarily additive). An important novel aspect of our work is that we do not restrict the model class, but instead put general non-parametric priors on this function and on the distribution of the cause. The causal direction can then be inferred by using standard Bayesian model selection. We evaluate our approach on synthetic data and real-world data and report encouraging results. Joris M. Mooij, Oliver Stegle, Dominik Janzing, Kun Zhang 0001, Bernhard Schölkopf |
NIPS | 3 |
| 2010 | Inferring deterministic causal relations
Povilas Daniusis, Dominik Janzing, Joris M. Mooij, Jakob Zscheischler, Bastian Steudel, Kun Zhang 0001, Bernhard Schölkopf |
UAI | 2 |
| 2010 | Invariant Gaussian Process Latent Variable Models and Application in Causal Discovery
Kun Zhang 0001, Bernhard Schölkopf, Dominik Janzing |
UAI | 3 |
| 2010 | Causal inference using the algorithmic Markov conditionabstractInferring the causal structure that links n observables is usually based upon detecting statistical dependences and choosing simple graphs that make the joint measure Markovian. Here we argue why causal inference is also possible when the sample size is one. We develop a theory how to generate causal graphs explaining similarities between single objects. To this end, we replace the notion of conditional stochastic independence in the causal Markov condition with the vanishing of conditional algorithmic mutual information and describe the corresponding causal inference rules. We explain why a consistent reformulation of causal inference in terms of algorithmic complexity implies a new inference principle that takes into account also the complexity of conditional probability densities, making it possible to select among Markov equivalent causal graphs. This insight provides a theoretical foundation of a heuristic principle proposed in earlier work. We also sketch some ideas on how to replace Kolmogorov complexity with decidable complexity criteria. This can be seen as an algorithmic analog of replacing the empirically undecidable question of statistical independence with practical independence tests that are based on implicit or explicit assumptions on the underlying distribution. Dominik Janzing, Bernhard Schölkopf |
IEEE Trans. Inf. Theory | 1 |
| 2009 | Regression by dependence minimization and its application to causal inference in additive noise modelsabstractMotivated by causal inference problems, we propose a novel method for regression that minimizes the statistical dependence between regressors and residuals. The key advantage of this approach to regression is that it does not assume a particular distribution of the noise, i.e., it is non-parametric with respect to the noise distribution. We argue that the proposed regression method is well suited to the task of causal inference in additive noise models. A practical disadvantage is that the resulting optimization problem is generally non-convex and can be difficult to solve. Nevertheless, we report good results on one of the tasks of the NIPS 2008 Causality Challenge, where the goal is to distinguish causes from effects in pairs of statistically dependent variables. In addition, we propose an algorithm for efficiently inferring causal models from observational data for more than two variables. The required number of regressions and independence tests is quadratic in the number of variables, which is a significant improvement over the simple method that tests all possible DAGs. Joris M. Mooij, Dominik Janzing, Jonas Peters, Bernhard Schölkopf |
ICML | 2 |
| 2009 | Detecting the direction of causal time seriesabstractWe propose a method that detects the true direction of time series, by fitting an autoregressive moving average model to the data. Whenever the noise is independent of the previous samples for one ordering of the observations, but dependent for the opposite ordering, we infer the former direction to be the true one. We prove that our method works in the population case as long as the noise of the process is not normally distributed (for the latter case, the direction is not identifiable). A new and important implication of our result is that it confirms a fundamental conjecture in causal reasoning --- if after regression the noise is independent of signal for one direction and dependent for the other, then the former represents the true causal direction --- in the case of time series. We test our approach on two types of data: simulated data sets conforming to our modeling assumptions, and real world EEG time series. Our method makes a decision for a significant fraction of both data sets, and these decisions are mostly correct. For real world data, our approach outperforms alternative solutions to the problem of time direction recovery. Jonas Peters, Dominik Janzing, Arthur Gretton, Bernhard Schölkopf |
ICML | 2 |
| 2009 | Identifying confounders using additive noise models
Dominik Janzing, Jonas Peters, Joris M. Mooij, Bernhard Schölkopf |
UAI | 1 |
| 2008 | Nonlinear causal discovery with additive noise modelsabstractThe discovery of causal relationships between a set of observed variables is a fundamental problem in science. For continuous-valued data linear acyclic causal models are often used because these models are well understood and there are well-known methods to fit them to data. In reality, of course, many causal relationships are more or less nonlinear, raising some doubts as to the applicability and usefulness of purely linear methods. In this contribution we show that in fact the basic linear framework can be generalized to nonlinear models with additive noise. In this extended framework, nonlinearities in the data-generating process are in fact a blessing rather than a curse, as they typically provide information on the underlying causal system and allow more aspects of the true data-generating mechanisms to be identified. In addition to theoretical results we show simulations and some simple real data experiments illustrating the identification power provided by nonlinearities. Patrik O. Hoyer, Dominik Janzing, Joris M. Mooij, Jonas Peters, Bernhard Schölkopf |
NIPS | 2 |
| 2008 | Causal reasoning by evaluating the complexity of conditional densities with kernel methods
Xiaohai Sun, Dominik Janzing, Bernhard Schölkopf |
Neurocomputing | 2 |
| 2007 | Learning causality by identifying common effects with kernel-based dependence measures
Xiaohai Sun, Dominik Janzing |
ESANN | 2 |
| 2007 | Exploring the causal order of binary variables via exponential hierarchies of Markov kernels
Xiaohai Sun, Dominik Janzing |
ESANN | 2 |
| 2007 | Distinguishing between cause and effect via kernel-based complexity measures for conditional distributions
Xiaohai Sun, Dominik Janzing, Bernhard Schölkopf |
ESANN | 2 |
| 2007 | A kernel-based causal learning algorithmabstractWe describe a causal learning method, which employs measuring the strength of statistical dependences in terms of the Hilbert-Schmidt norm of kernel-based cross-covariance operators. Following the line of the common faithfulness assumption of constraint-based causal learning, our approach assumes that a variable Z is likely to be a common effect of X and Y, if conditioning on Z increases the dependence between X and Y. Based on this assumption, we collect "votes" for hypothetical causal directions and orient the edges by the majority principle. In most experiments with known causal structures, our method provided plausible results and outperformed the conventional constraint-based PC algorithm. Xiaohai Sun, Dominik Janzing, Bernhard Schölkopf, Kenji Fukumizu |
ICML | 2 |
| 2003 | Quasi-order of clocks and their synchronism and quantum bounds for copying timing informationabstractThe statistical state of any (classical or quantum) system with nontrivial time evolution can be interpreted as the pointer of a clock. The quality of such a clock is given by the statistical distinguishability of its states at different times. If a clock is used as a resource for producing another one the latter can at most have the quality of the resource. We show that this principle, formalized by a quasi-order, implies constraints on many physical processes. Similarly, the degree to which two (quantum or classical) clocks are synchronized can be formalized by a quasi-order of synchronism. Copying timing information is restricted by quantum no-cloning and no-broadcasting theorems since classical clocks can only exist in the limit of infinite energy. We show this quantitatively by comparing the Fisher timing information of two output systems to the input's timing information. For classical signal processing in the quantum regime our results imply that a signal looses its localization in time if it is amplified and distributed to many devices. Dominik Janzing, Thomas Beth |
IEEE Trans. Inf. Theory | 1 |