VLDB 2026 Research / reviewers in the wild / expert
Mathias Drton
dblp:78/3067
· DBLP profile ↗
26ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-5614-3025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Score MatchingabstractProposed in Hyv{ä}rinen (2005), score matching is a statistical estimation procedure that does not require computation of distributional normalizing constants. In this work we utilize the geometric median of means to develop a robust score matching procedure that yields consistent parameter estimates in settings where the observed data has been contaminated. A special appeal of the proposed method is that it retains convexity in exponential family models. The new method is therefore particularly attractive for non-Gaussian, exponential family graphical models where evaluation of normalizing constants is intractable. Support recovery guarantees for such models when contamination is present are provided. Additionally, support recovery is studied in numerical experiments and on a precipitation dataset. We demonstrate that the proposed robust score matching estimator performs comparably to the standard score matching estimator when no contamination is present but greatly outperforms this estimator in a setting with contamination. Richard Schwank, Andrew McCormack, Mathias Drton |
AISTATS | 3 |
| 2025 | Causal Effect Identification in lvLiNGAM from Higher-Order CumulantsabstractThis paper investigates causal effect identification in latent variable Linear Non-Gaussian Acyclic Models (lvLiNGAM) using higher-order cumulants, addressing two prominent setups that are challenging in the presence of latent confounding: (1) a single proxy variable that may causally influence the treatment and (2) underspecified instrumental variable cases where fewer instruments exist than treatments. We prove that causal effects are identifiable with a single proxy or instrument and provide corresponding estimation methods. Experimental results demonstrate the accuracy and robustness of our approaches compared to existing methods, advancing the theoretical and practical understanding of causal inference in linear systems with latent confounders. Daniele Tramontano, Yaroslav Kivva, Saber Salehkaleybar, Negar Kiyavash, Mathias Drton |
ICML | 5 |
| 2025 | Causal Discovery for Linear Non-Gaussian Models with Disjoint CyclesabstractThe paradigm of linear structural equation modeling readily allows one to incorporate causal feedback loops in the model specification. These appear as directed cycles in the common graphical representation of the models. However, the presence of cycles entails difficulties such as the fact that models need no longer be characterized by conditional independence relations. As a result, learning cyclic causal structures remains a challenging problem. In this paper, we offer new insights on this problem in the context of linear non-Gaussian models. First, we precisely characterize when two directed graphs determine the same linear non-Gaussian model. Next, we take up a setting of cycle-disjoint graphs, for which we are able to show that simple quadratic and cubic polynomial relations among low-order moments of a non-Gaussian distribution allow one to locate source cycles. Complementing this with a strategy of decorrelating cycles and multivariate regression allows one to infer a block-topological order among the directed cycles, which leads to a consistent and computationally efficient algorithm for learning causal structures with disjoint cycles. Mathias Drton, Marina Garrote-López, Niko Nikov, Elina Robeva, Y. Samuel Wang |
UAI | 1 |
| 2025 | Nonlinear Causal Discovery for Grouped DataabstractInferring cause-effect relationships from observational data has gained significant attention in recent years, but most methods are limited to scalar random variables. In many important domains, including neuroscience, psychology, social science, and industrial manufacturing, the causal units of interest are groups of variables rather than individual scalar measurements. Motivated by these applications, we extend nonlinear additive noise models to handle random vectors, establishing a two-step approach for causal graph learning: First, infer the causal order among random vectors. Second, perform model selection to identify the best graph consistent with this order. We introduce effective and novel solutions for both steps in the vector case, demonstrating strong performance in simulations. Finally, we apply our method to real-world assembly line data with partial knowledge of causal ordering among variable groups. Konstantin Göbler, Tobias Windisch, Mathias Drton |
UAI | 3 |
| 2025 | Single-cell differential expression analysis between conditions within nested settingsabstractDifferential expression analysis provides insights into fundamental biological processes and with the advent of single-cell transcriptomics, gene expression can now be studied at the level of individual cells. Many analyses treat cells as samples and assume statistical independence. As cells are pseudoreplicates, this assumption does not hold, leading to reduced robustness, reproducibility, and an inflated type 1 error rate. In this study, we investigate various methods for differential expression analysis on single-cell data, conduct extensive benchmarking, and give recommendations for method choice. The tested methods include DESeq2, MAST, DREAM, scVI, the permutation test, distinct, and the t-test. We additionally adapt hierarchical bootstrapping to differential expression analysis on single-cell data and include it in our benchmark. We found that differential expression analysis methods designed specifically for single-cell data do not offer performance advantages over conventional pseudobulk methods such as DESeq2 when applied to individual datasets. In addition, they mostly require significantly longer run times. For atlas-level analysis, permutation-based methods excel in performance but show poor runtime, suggesting to use DREAM as a compromise between quality and runtime. Overall, our study offers the community a valuable benchmark of methods across diverse scenarios and offers guidelines on method selection. Leon Hafner, Gregor Sturm, Sarah Lumpp, Mathias Drton, Markus List |
Briefings Bioinform. | 4 |
| 2025 | Identifying total causal effects in linear models under partial homoscedasticityabstractA fundamental challenge of scientific research is inferring causal relations based on observed data. One commonly used approach involves utilizing structural causal models that postulate noisy functional relations among interacting variables. A directed graph naturally represents these models and reflects the underlying causal structure. However, classical identifiability results suggest that, without conducting additional experiments, this causal graph can only be identified up to a Markov equivalence class of indistinguishable models. Recent research has shown that focusing on linear relations with equal error variances can enable the identification of the causal structure from mere observational data. Nonetheless, practitioners are often primarily interested in the effects of specific interventions, rendering the complete identification of the causal structure unnecessary. In this work, we investigate the extent to which less restrictive assumptions of partial homoscedasticity are sufficient for identifying the causal effects of interest. Furthermore, we construct mathematically rigorous confidence regions for total causal effects under structure uncertainty and explore the performance gain of relying on stricter error assumptions in a simulation study. • We demonstrate that partial homoscedasticity is sufficient to generically identify total causal effects. • We construct rigorous confidence regions for total causal effects under structure uncertainty. • We evaluate the performance gained by relying on stricter error assumptions in a simulation study. David Strieder, Mathias Drton |
Int. J. Approx. Reason. | 2 |
| 2024 | Causal Effect Identification in LiNGAM Models with Latent ConfoundersabstractWe study the generic identifiability of causal effects in linear non-Gaussian acyclic models (LiNGAM) with latent variables. We consider the problem in two main settings: When the causal graph is known a priori, and when it is unknown. In both settings, we provide a complete graphical characterization of the identifiable direct or total causal effects among observed variables. Moreover, we propose efficient algorithms to certify the graphical conditions. Finally, we propose an adaptation of the reconstruction independent component analysis (RICA) algorithm that estimates the causal effects from the observational data given the causal graph. Experimental results show the effectiveness of the proposed method in estimating the causal effects. Daniele Tramontano, Yaroslav Kivva, Saber Salehkaleybar, Mathias Drton, Negar Kiyavash |
ICML | 4 |
| 2023 | Rank-Based Causal Discovery for Post-Nonlinear ModelsabstractLearning causal relationships from empirical observations is a central task in scientific research. A common method is to employ structural causal models that postulate noisy functional relations among a set of interacting variables. To ensure unique identifiability of causal directions, researchers consider restricted subclasses of structural causal models. Post-nonlinear (PNL) causal models constitute one of the most flexible options for such restricted subclasses, containing in particular the popular additive noise models as a further subclass. However, learning PNL models is not well studied beyond the bivariate case. The existing methods learn non-linear functional relations by minimizing residual dependencies and subsequently test independence from residuals to determine causal orientations. However, these methods can be prone to overfitting and, thus, difficult to tune appropriately in practice. As an alternative, we propose a new approach for PNL causal discovery that uses rank-based methods to estimate the functional parameters. This new approach exploits natural invariances of PNL models and disentangles the estimation of the non-linear functions from the independence tests used to find causal orientations. We prove consistency of our method and validate our results in numerical experiments. Grigor Keropyan, David Strieder, Mathias Drton |
AISTATS | 3 |
| 2023 | Unpaired Multi-Domain Causal Representation LearningabstractThe goal of causal representation learning is to find a representation of data that consists of causally related latent variables. We consider a setup where one has access to data from multiple domains that potentially share a causal representation. Crucially, observations in different domains are assumed to be unpaired, that is, we only observe the marginal distribution in each domain but not their joint distribution. In this paper, we give sufficient conditions for identifiability of the joint distribution and the shared causal graph in a linear setup. Identifiability holds if we can uniquely recover the joint distribution and the shared causal representation from the marginal distributions in each domain. We transform our results into a practical method to recover the shared latent causal graph. Nils Sturma, Chandler Squires, Mathias Drton, Caroline Uhler |
NeurIPS | 3 |
| 2023 | Causal Discovery with Unobserved Confounding and Non-Gaussian DataabstractWe consider recovering causal structure from multivariate observational data. We assume the data arise from a linear structural equation model (SEM) in which the idiosyncratic errors are allowed to be dependent in order to capture possible latent confounding. Each SEM can be represented by a graph where vertices represent observed variables, directed edges represent direct causal effects, and bidirected edges represent dependence among error terms. Specifically, we assume that the true model corresponds to a bow-free acyclic path diagram; i.e., a graph that has at most one edge between any pair of nodes and is acyclic in the directed part. We show that when the errors are non-Gaussian, the exact causal structure encoded by such a graph, and not merely an equivalence class, can be recovered from observational data. The method we propose for this purpose uses estimates of suitable moments, but, in contrast to previous results, does not require specifying the number of latent variables a priori. We also characterize the output of our procedure when the assumptions are violated and the true graph is acyclic, but not bow-free. We illustrate the effectiveness of our procedure in simulations and an application to an ecology data set. Y. Samuel Wang, Mathias Drton |
J. Mach. Learn. Res. | 2 |
| 2022 | Learning linear non-Gaussian polytree modelsabstractIn the context of graphical causal discovery, we adapt the versatile framework of linear non-Gaussian acyclic models (LiNGAMs) to propose new algorithms to efficiently learn graphs that are polytrees. Our approach combines the Chow–Liu algorithm, which first learns the undirected tree structure, with novel schemes to orient the edges. The orientation schemes assess algebraic relations among moments of the data-generating distribution and are computationally inexpensive. We establish high-dimensional consistency results for our approach and compare different algorithmic versions in numerical experiments. Daniele Tramontano, Anthea Monod, Mathias Drton |
UAI | 3 |
| 2021 | Confidence in causal discovery with linear causal modelsabstractStructural causal models postulate noisy functional relations among a set of interacting variables. The causal structure underlying each such model is naturally represented by a directed graph whose edges indicate for each variable which other variables it causally depends upon. Under a number of different model assumptions, it has been shown that this causal graph and, thus also, causal effects are identifiable from mere observational data. For these models, practical algorithms have been devised to learn the graph. Moreover, when the graph is known, standard techniques may be used to give estimates and confidence intervals for causal effects. We argue, however, that a two-step method that first learns a graph and then treats the graph as known yields confidence intervals that are overly optimistic and can drastically fail to account for the uncertain causal structure. To address this issue we lay out a framework based on test inversion that allows us to give confidence regions for total causal effects that capture both sources of uncertainty: causal structure and numerical size of nonzero effects. Our ideas are developed in the context of bivariate linear causal models with homoscedastic errors, but as we exemplify they are generalizable to larger systems as well as other settings such as, in particular, linear non-Gaussian models. David Strieder, Tobias Freidling, Stefan Haffner, Mathias Drton |
UAI | 4 |
| 2021 | CorDiffViz: an R package for visualizing multi-omics differential correlation networksabstractBACKGROUND: Differential correlation networks are increasingly used to delineate changes in interactions among biomolecules. They characterize differences between omics networks under two different conditions, and can be used to delineate mechanisms of disease initiation and progression. RESULTS: We present a new R package, CorDiffViz, that facilitates the estimation and visualization of differential correlation networks using multiple correlation measures and inference methods. The software is implemented in R, HTML and Javascript, and is available at https://github.com/sqyu/CorDiffViz . Visualization has been tested for the Chrome and Firefox web browsers. A demo is available at https://diffcornet.github.io/CorDiffViz/demo.html . CONCLUSIONS: Our software offers considerable flexibility by allowing the user to interact with the visualization and choose from different estimation methods and visualizations. It also allows the user to easily toggle between correlation networks for samples under one condition and differential correlations between samples under two conditions. Moreover, the software facilitates integrative analysis of cross-correlation networks between two omics data sets. Shiqing Yu, Mathias Drton, Daniel E. L. Promislow, Ali Shojaie |
BMC Bioinform. | 2 |
| 2020 | Structure Learning for Cyclic Linear Causal ModelsabstractWe consider the problem of structure learning for linear causal models based on observational data. We treat models given by possibly cyclic mixed graphs, which allow for feedback loops and effects of latent confounders. Generalizing related work on bow-free acyclic graphs, we assume that the underlying graph is simple. This entails that any two observed variables can be related through at most one direct causal effect and that (confounding-induced) correlation between error terms in structural equations occurs only in absence of direct causal effects.We show that, despite new subtleties in the cyclic case, the considered simple cyclic models are of expected dimension and that a previously considered criterion for distributional equivalence of bow-free acyclic graphs has an analogue in the cyclic case. Our result on model dimension justifies in particular score-based methods for structure learning of linear Gaussian mixed graph models, which we implement via greedy search. Carlos Améndola, Philipp Dettling, Mathias Drton, Federica Onori |
UAI | 3 |
| 2019 | Generalized Score Matching for Non-Negative DataabstractA common challenge in estimating parameters of probability density functions is the intractability of the normalizing constant. While in such cases maximum likelihood estimation may be implemented using numerical integration, the approach becomes computationally intensive. The score matching method of Hyvärinen (2005) avoids direct calculation of the normalizing constant and yields closed-form estimates for exponential families of continuous distributions over $\mathbb{R}^m$. Hyvärinen (2007) extended the approach to distributions supported on the non-negative orthant, $\mathbb{R}_+^m$. In this paper, we give a generalized form of score matching for non-negative data that improves estimation efficiency. As an example, we consider a general class of pairwise interaction models. Addressing an overlooked inexistence problem, we generalize the regularized score matching method of Lin et al. (2016) and improve its theoretical guarantees for non-negative Gaussian graphical models. Shiqing Yu, Mathias Drton, Ali Shojaie |
J. Mach. Learn. Res. | 2 |
| 2018 | Graphical Models for Non-Negative Data Using Generalized Score MatchingabstractA common challenge in estimating parameters of probability density functions is the intractability of the normalizing constant. While in such cases maximum likelihood estimation may be implemented using numerical integration, the approach becomes computationally intensive. In contrast, the score matching method of Hyvärinen (2005) avoids direct calculation of the normalizing constant and yields closed-form estimates for exponential families of continuous distributions over $\mathbb{R}^m$. Hyvärinen (2007) extended the approach to distributions supported on the non-negative orthant $\mathbb{R}_+^m$. In this paper, we give a generalized form of score matching for non-negative data that improves estimation efficiency. We also generalize the regularized score matching method of Lin et al. (2016) for non-negative Gaussian graphical models, with improved theoretical guarantees. Shiqing Yu, Mathias Drton, Ali Shojaie |
AISTATS | 2 |
| 2018 | Algebraic tests of general Gaussian latent tree modelsabstractWe consider general Gaussian latent tree models in which the observed variables are not restricted to be leaves of the tree. Extending related recent work, we give a full semi-algebraic description of the set of covariance matrices of any such model. In other words, we find polynomial constraints that characterize when a matrix is the covariance matrix of a distribution in a given latent tree model. However, leveraging these constraints to test a given such model is often complicated by the number of constraints being large and by singularities of individual polynomials, which may invalidate standard approximations to relevant probability distributions. Illustrating with the star tree, we propose a new testing methodology that circumvents singularity issues by trading off some statistical estimation efficiency and handles cases with many constraints through recent advances on Gaussian approximation for maxima of sums of high-dimensional random vectors. Our test avoids the need to maximize the possibly multimodal likelihood function of such models and is applicable to models with larger number of variables. These points are illustrated in numerical experiments. Dennis Leung, Mathias Drton |
NeurIPS | 2 |
| 2013 | PC algorithm for nonparanormal graphical models
Naftali Harris, Mathias Drton |
J. Mach. Learn. Res. | 2 |
| 2012 | Nonparametric Reduced Rank RegressionabstractWe propose an approach to multivariate nonparametric regression that generalizes reduced rank regression for linear models. An additive model is estimated for each dimension of a $q$-dimensional response, with a shared $p$-dimensional predictor variable. To control the complexity of the model, we employ a functional form of the Ky-Fan or nuclear norm, resulting in a set of function estimates that have low rank. Backfitting algorithms are derived and justified using a nonparametric form of the nuclear norm subdifferential. Oracle inequalities on excess risk are derived that exhibit the scaling behavior of the procedure in the high dimensional setting. The methods are illustrated on gene expression data. Rina Foygel Barber, Michael Horrell, Mathias Drton, John D. Lafferty |
NIPS | 3 |
| 2010 | Extended Bayesian Information Criteria for Gaussian Graphical ModelsabstractGaussian graphical models with sparsity in the inverse covariance matrix are of significant interest in many modern applications. For the problem of recovering the graphical structure, information criteria provide useful optimization objectives for algorithms searching through sets of graphs or for selection of tuning parameters of other methods such as the graphical lasso, which is a likelihood penalization technique. In this paper we establish the asymptotic consistency of an extended Bayesian information criterion for Gaussian graphical models in a scenario where both the number of variables p and the sample size n grow. Compared to earlier work on the regression case, our treatment allows for growth in the number of non-zero parameters in the true model, which is necessary in order to cover connected graphs. We demonstrate the performance of this criterion on simulated data when used in conjuction with the graphical lasso, and verify that the criterion indeed performs better than either cross-validation or the ordinary Bayesian information criterion when p and the number of non-zero parameters q both scale with n. Rina Foygel Barber, Mathias Drton |
NIPS | 2 |
| 2009 | Robust Graphical Modeling with t-Distributions
Michael Finegold, Mathias Drton |
UAI | 2 |
| 2009 | Computing Maximum Likelihood Estimates in Recursive Linear Models with Correlated Errors
Mathias Drton, Michael Eichler, Thomas Richardson 0001 |
J. Mach. Learn. Res. | 1 |
| 2008 | Graphical Methods for Efficient Likelihood Inference in Gaussian Covariance Models
Mathias Drton, Thomas Richardson 0001 |
J. Mach. Learn. Res. | 1 |
| 2006 | Computing all roots of the likelihood equations of seemingly unrelated regressions
Mathias Drton |
J. Symb. Comput. | 1 |
| 2004 | Iterative Conditional Fitting for Gaussian Ancestral Graph Models
Mathias Drton, Thomas Richardson 0001 |
UAI | 1 |
| 2003 | A New Algorithm for Maximum Likelihood Estimation in Gaussian Graphical Models for Marginal Independence
Mathias Drton, Thomas Richardson 0001 |
UAI | 1 |