VLDB 2026 Research / reviewers in the wild / expert
Clark Glymour
dblp:88/16
· DBLP profile ↗
33ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
Probabilistic and Bayesian machine learning · 68% Transfer learning and domain adaptation · 9% Knowledge representation and reasoning · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 55% Computational social science and digital humanities · 45% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
4.9 | 15 | 2024 | Generalized Independent Noise Condition for Estimating Causal Structure with Latent Variables · J. Mach. Learn. Res. 2024 Latent Hierarchical Causal Structure Discovery with Rank Constraints · NeurIPS 2022 Causal Discovery from Heterogeneous/Nonstationary Data · J. Mach. Learn. Res. 2020 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
1.1 | 6 | 2020 | Causal Discovery from Heterogeneous/Nonstationary Data · J. Mach. Learn. Res. 2020 Triad Constraints for Learning Causal Structure of Latent Variables · NeurIPS 2019 Search for Additive Nonlinear Time Series Causal Models · J. Mach. Learn. Res. 2008 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
latent variable causal models |
0.8 | 1 | 2024 | Generalized Independent Noise Condition for Estimating Causal Structure with Latent Variables · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
latent variable causal discovery |
0.6 | 1 | 2022 | Latent Hierarchical Causal Structure Discovery with Rank Constraints · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning |
0.6 | 1 | 2022 | Action-Sufficient State Representation Learning for Control with Structural Constraints · ICML 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint satisfaction
structural constraints |
0.6 | 1 | 2022 | Action-Sufficient State Representation Learning for Control with Structural Constraints · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.4 | 1 | 2020 | Domain Adaptation as a Problem of Inference on Graphical Models · NeurIPS 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graph discovery |
0.4 | 1 | 2020 | Causal Discovery from Multiple Data Sets with Non-Identical Variable Sets · AAAI 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.4 | 1 | 2020 | Generalized Independent Noise Condition for Estimating Latent Variable Causal Graphs · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference |
0.4 | 1 | 2020 | Domain Adaptation as a Problem of Inference on Graphical Models · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders |
0.4 | 1 | 2020 | Generalized Independent Noise Condition for Estimating Latent Variable Causal Graphs · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.4 | 1 | 2020 | Generalized Independent Noise Condition for Estimating Latent Variable Causal Graphs · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation |
0.4 | 1 | 2020 | Domain Adaptation as a Problem of Inference on Graphical Models · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.4 | 1 | 2020 | Domain Adaptation as a Problem of Inference on Graphical Models · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning
clustering |
0.4 | 1 | 2019 | Specific and Shared Causal Relation Modeling and Mechanism-Based Clustering · NeurIPS 2019 |
Machine learning › Time series and sequential data › time series analysis
non-stationary time series |
0.4 | 1 | 2019 | Causal Discovery and Forecasting in Nonstationary Environments with State-Space Models · ICML 2019 |
Machine learning › Deep learning architectures and training
state space model |
0.4 | 1 | 2019 | Causal Discovery and Forecasting in Nonstationary Environments with State-Space Models · ICML 2019 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.4 | 1 | 2019 | Causal Discovery and Forecasting in Nonstationary Environments with State-Space Models · ICML 2019 |
Computational social science and digital humanities
causal inference |
0.4 | 1 | 2019 | Mixed graphical models for integrative causal analysis with application to chronic lung disease diagnosis and prognosis · Bioinform. 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › conditional independence
conditional independence testing |
0.3 | 1 | 2018 | Generalized Score Functions for Causal Discovery · KDD 2018 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
score-based causal discovery |
0.3 | 1 | 2018 | Generalized Score Functions for Causal Discovery · KDD 2018 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
causal direction detection |
0.3 | 1 | 2017 | Behind Distribution Shift: Mining Driving Forces of Changes and Causal Arrows · ICDM 2017 |
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift |
0.2 | 1 | 2016 | Domain Adaptation with Conditional Transferable Components · ICML 2016 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2016 | Domain Adaptation with Conditional Transferable Components · ICML 2016 |
Machine learning › Generative modeling
variational autoencoder |
0.2 | 1 | 2022 | Action-Sufficient State Representation Learning for Control with Structural Constraints · ICML 2022 |
Machine learning › Generative modeling
generative adversarial network |
0.1 | 1 | 2020 | Causal Discovery from Multiple Data Sets with Non-Identical Variable Sets · AAAI 2020 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference |
0.1 | 1 | 2020 | Causal Discovery from Multiple Data Sets with Non-Identical Variable Sets · AAAI 2020 |
Machine learning › Time series and sequential data
time series |
0.1 | 1 | 2008 | Search for Additive Nonlinear Time Series Causal Models · J. Mach. Learn. Res. 2008 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
latent confounder identification |
0.1 | 1 | 2006 | Learning the Structure of Linear Latent Variable Models · J. Mach. Learn. Res. 2006 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.0 | 1 | 2003 | A Statistical Problem for Inference to Regulatory Structure from Associations of Gene Expression Measurements with Microarrays · Bioinform. 2003 |
Methods — techniques the papers use, named apart from their topics
bayesian inference · 0.8non-gaussian acyclic models · 0.8generalized independent noise condition · 0.8constraint-based causal discovery · 0.7variational autoencoder · 0.6structural constraints · 0.6rank deficiency constraints · 0.6markov equivalence class · 0.6generative environment model · 0.6distribution shift handling · 0.4directed graphical models · 0.4CausalMGM · 0.4correlation analysis · 0.0conditional independence testing · 0.0bayesian updating · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Generalized Independent Noise Condition for Estimating Causal Structure with Latent VariablesabstractWe investigate the challenging task of learning causal structure in the presence of latent variables, including locating latent variables, determining their quantity, and identifying causal relationships among both latent and observed variables. To address this, we propose a Generalized Independent Noise (GIN) condition for linear non-Gaussian acyclic causal models that incorporate latent variables, which establishes the independence between a linear combination of certain measured variables and some other measured variables. Specifically, for two observed random vectors $\bf{Y}$ and $\bf{Z}$, GIN holds if and only if $\omega^{\intercal}\mathbf{Y}$ and $\mathbf{Z}$ are statistically independent, where $\omega$ is a non-zero parameter vector determined by the cross-covariance between $\mathbf{Y}$ and $\mathbf{Z}$. We then give necessary and sufficient graphical criteria of the GIN condition in linear non-Gaussian acyclic causal models. From a graphical perspective, roughly speaking, GIN implies the existence of a set $\mathcal{S}$ such that $\mathcal{S}$ is causally earlier (w.r.t. the causal ordering) than $\mathbf{Y}$, and that every active (collider-free) path between $\mathbf{Y}$ and $\mathbf{Z}$ must contain a node from $\mathcal{S}$. Interestingly, we find that the independent noise condition (i.e., if there is no confounder, causes are independent of the residual derived from regressing the effect on the causes) can be seen as a special case of GIN. With such a connection between GIN and latent causal structures, we further leverage the proposed GIN condition, together with a well-designed search procedure, to efficiently estimate Linear, Non-Gaussian Latent Hierarchical Models (LiNGLaHs), where latent confounders may also be causally related and may even follow a hierarchical structure. We show that the underlying causal structure of a LiNGLaH is identifiable in light of GIN conditions under mild assumptions. Experimental results on both synthetic and three real-world data sets show the effectiveness of the proposed approach. Feng Xie 0002, Biwei Huang, Zhengming Chen 0002, Ruichu Cai, Clark Glymour, Zhi Geng, Kun Zhang 0001 |
J. Mach. Learn. Res. | 5 |
| 2022 | Action-Sufficient State Representation Learning for Control with Structural ConstraintsabstractPerceived signals in real-world scenarios are usually high-dimensional and noisy, and finding and using their representation that contains essential and sufficient information required by downstream decision-making tasks will help improve computational efficiency and generalization ability in the tasks. In this paper, we focus on partially observable environments and propose to learn a minimal set of state representations that capture sufficient information for decision-making, termed Action-Sufficient state Representations (ASRs). We build a generative environment model for the structural relationships among variables in the system and present a principled way to characterize ASRs based on structural constraints and the goal of maximizing cumulative reward in policy learning. We then develop a structured sequential Variational Auto-Encoder to estimate the environment model and extract ASRs. Our empirical results on CarRacing and VizDoom demonstrate a clear advantage of learning and using ASRs for policy learning. Moreover, the estimated environment model and ASRs allow learning behaviors from imagined outcomes in the compact latent space to improve sample efficiency. Biwei Huang, Chaochao Lu, Liu Leqi, José Miguel Hernández-Lobato, Clark Glymour, Bernhard Schölkopf, Kun Zhang 0001 |
ICML | 5 |
| 2022 | Latent Hierarchical Causal Structure Discovery with Rank ConstraintsabstractMost causal discovery procedures assume that there are no latent confounders in the system, which is often violated in real-world problems. In this paper, we consider a challenging scenario for causal structure identification, where some variables are latent and they may form a hierarchical graph structure to generate the measured variables; the children of latent variables may still be latent and only leaf nodes are measured, and moreover, there can be multiple paths between every pair of variables (i.e., it is beyond tree structure). We propose an estimation procedure that can efficiently locate latent variables, determine their cardinalities, and identify the latent hierarchical structure, by leveraging rank deficiency constraints over the measured variables. We show that the proposed algorithm can find the correct Markov equivalence class of the whole graph asymptotically under proper restrictions on the graph structure and with linear causal relations. Biwei Huang, Charles Jia Han Low, Feng Xie 0002, Clark Glymour, Kun Zhang 0001 |
NeurIPS | 4 |
| 2020 | Causal Discovery from Multiple Data Sets with Non-Identical Variable SetsabstractA number of approaches to causal discovery assume that there are no hidden confounders and are designed to learn a fixed causal model from a single data set. Over the last decade, with closer cooperation across laboratories, we are able to accumulate more variables and data for analysis, while each lab may only measure a subset of them, due to technical constraints or to save time and cost. This raises a question of how to handle causal discovery from multiple data sets with non-identical variable sets, and at the same time, it would be interesting to see how more recorded variables can help to mitigate the confounding problem. In this paper, we propose a principled method to uniquely identify causal relationships over the integrated set of variables from multiple data sets, in linear, non-Gaussian cases. The proposed method also allows distribution shifts across data sets. Theoretically, we show that the causal structure over the integrated set of variables is identifiable under testable conditions. Furthermore, we present two types of approaches to parameter estimation: one is based on maximum likelihood, and the other is likelihood free and leverages generative adversarial nets to improve scalability of the estimation procedure. Experimental results on various synthetic and real-world data sets are presented to demonstrate the efficacy of our methods. Biwei Huang, Kun Zhang 0001, Mingming Gong, Clark Glymour |
AAAI | 4 |
| 2020 | Domain Adaptation as a Problem of Inference on Graphical ModelsabstractThis paper is concerned with data-driven unsupervised domain adaptation, where it is unknown in advance how the joint distribution changes across domains, i.e., what factors or modules of the data distribution remain invariant or change across domains. To develop an automated way of domain adaptation with multiple source domains, we propose to use a graphical model as a compact way to encode the change property of the joint distribution, which can be learned from data, and then view domain adaptation as a problem of Bayesian inference on the graphical models. Such a graphical model distinguishes between constant and varied modules of the distribution and specifies the properties of the changes across domains, which serves as prior knowledge of the changing modules for the purpose of deriving the posterior of the target variable $Y$ in the target domain. This provides an end-to-end framework of domain adaptation, in which additional knowledge about how the joint distribution changes, if available, can be directly incorporated to improve the graphical representation. We discuss how causality-based domain adaptation can be put under this umbrella. Experimental results on both synthetic and real data demonstrate the efficacy of the proposed framework for domain adaptation. Kun Zhang 0001, Mingming Gong, Petar Stojanov, Biwei Huang, Clark Glymour |
NeurIPS | 6 |
| 2020 | Generalized Independent Noise Condition for Estimating Latent Variable Causal GraphsabstractCausal discovery aims to recover causal structures or models underlying the observed data. Despite its success in certain domains, most existing methods focus on causal relations between observed variables, while in many scenarios the observed ones may not be the underlying causal variables (e.g., image pixels), but are generated by latent causal variables or confounders that are causally related. To this end, in this paper, we consider Linear, Non-Gaussian Latent variable Models (LiNGLaMs), in which latent confounders are also causally related, and propose a Generalized Independent Noise (GIN) condition to estimate such latent variable graphs. Specifically, for two observed random vectors $\mathbf{Y}$ and $\mathbf{Z}$, GIN holds if and only if $\omega^{\intercal}\mathbf{Y}$ and $\mathbf{Z}$ are statistically independent, where $\omega$ is a parameter vector characterized from the cross-covariance between $\mathbf{Y}$ and $\mathbf{Z}$. From the graphical view, roughly speaking, GIN implies that causally earlier latent common causes of variables in $\mathbf{Y}$ d-separate $\mathbf{Y}$ from $\mathbf{Z}$. Interestingly, we find that the independent noise condition, i.e., if there is no confounder, causes are independent from the error of regressing the effect on the causes, can be seen as a special case of GIN. Moreover, we show that GIN helps locate latent variables and identify their causal structure, including causal directions. We further develop a recursive learning algorithm to achieve these goals. Experimental results on synthetic and real-world data demonstrate the effectiveness of our method. Feng Xie 0002, Ruichu Cai, Biwei Huang, Clark Glymour, Zhifeng Hao 0004, Kun Zhang 0001 |
NeurIPS | 4 |
| 2020 | Causal Discovery from Heterogeneous/Nonstationary DataabstractIt is commonplace to encounter heterogeneous or nonstationary data, of which the underlying generating process changes across domains or over time. Such a distribution shift feature presents both challenges and opportunities for causal discovery. In this paper, we develop a framework for causal discovery from such data, called Constraint-based causal Discovery from heterogeneous/NOnstationary Data (CD-NOD), to find causal skeleton and directions and estimate the properties of mechanism changes. First, we propose an enhanced constraint-based procedure to detect variables whose local mechanisms change and recover the skeleton of the causal structure over observed variables. Second, we present a method to determine causal orientations by making use of independent changes in the data distribution implied by the underlying causal model, benefiting from information carried by changing distributions. After learning the causal structure, next, we investigate how to efficiently estimate the “driving force” of the nonstationarity of a causal mechanism. That is, we aim to extract from data a low-dimensional representation of changes. The proposed methods are nonparametric, with no hard restrictions on data distributions and causal mechanisms, and do not rely on window segmentation. Furthermore, we find that data heterogeneity benefits causal structure identification even with particular types of confounders. Finally, we show the connection between heterogeneity/nonstationarity and soft intervention in causal discovery. Experimental results on various synthetic and real-world data sets (task-fMRI and stock market data) are presented to demonstrate the efficacy of the proposed methods. Biwei Huang, Kun Zhang 0001, Jiji Zhang, Joseph D. Ramsey, Ruben Sanchez-Romero, Clark Glymour, Bernhard Schölkopf |
J. Mach. Learn. Res. | 6 |
| 2019 | Causal Discovery and Forecasting in Nonstationary Environments with State-Space ModelsabstractIn many scientific fields, such as economics and neuroscience, we are often faced with nonstationary time series, and concerned with both finding causal relations and forecasting the values of variables of interest, both of which are particularly challenging in such nonstationary environments. In this paper, we study causal discovery and forecasting for nonstationary time series. By exploiting a particular type of state-space model to represent the processes, we show that nonstationarity helps to identify the causal structure, and that forecasting naturally benefits from learned causal knowledge. Specifically, we allow changes in both causal strengths and noise variances in the nonlinear state-space models, which, interestingly, renders both the causal structure and model parameters identifiable. Given the causal model, we treat forecasting as a problem in Bayesian inference in the causal model, which exploits the time-varying property of the data and adapts to new observations in a principled manner. Experimental results on synthetic and real-world data sets demonstrate the efficacy of the proposed methods. Biwei Huang, Kun Zhang 0001, Mingming Gong, Clark Glymour |
ICML | 4 |
| 2019 | Triad Constraints for Learning Causal Structure of Latent VariablesabstractLearning causal structure from observational data has attracted much attention, and it is notoriously challenging to find the underlying structure in the presence of confounders (hidden direct common causes of two variables). In this paper, by properly leveraging the non-Gaussianity of the data, we propose to estimate the structure over latent variables with the so-called Triad constraints: we design a form of "pseudo-residual" from three variables, and show that when causal relations are linear and noise terms are non-Gaussian, the causal direction between the latent variables for the three observed variables is identifiable by checking a certain kind of independence relationship. In other words, the Triad constraints help us to locate latent confounders and determine the causal direction between them. This goes far beyond the Tetrad constraints and reveals more information about the underlying structure from non-Gaussian data. Finally, based on the Triad constraints, we develop a two-step algorithm to learn the causal structure corresponding to measurement models. Experimental results on both synthetic and real data demonstrate the effectiveness and reliability of our method. Ruichu Cai, Feng Xie 0002, Clark Glymour, Zhifeng Hao 0004, Kun Zhang 0001 |
NeurIPS | 3 |
| 2019 | Specific and Shared Causal Relation Modeling and Mechanism-Based ClusteringabstractState-of-the-art approaches to causal discovery usually assume a fixed underlying causal model. However, it is often the case that causal models vary across domains or subjects, due to possibly omitted factors that affect the quantitative causal effects. As a typical example, causal connectivity in the brain network has been reported to vary across individuals, with significant differences across groups of people, such as autistics and typical controls. In this paper, we develop a unified framework for causal discovery and mechanism-based group identification. In particular, we propose a specific and shared causal model (SSCM), which takes into account the variabilities of causal relations across individuals/groups and leverages their commonalities to achieve statistically reliable estimation. The learned SSCM gives the specific causal knowledge for each individual as well as the general trend over the population. In addition, the estimated model directly provides the group information of each individual. Experimental results on synthetic and real-world data demonstrate the efficacy of the proposed method. Biwei Huang, Kun Zhang 0001, Pengtao Xie, Mingming Gong, Eric P. Xing, Clark Glymour |
NeurIPS | 6 |
| 2019 | Mixed graphical models for integrative causal analysis with application to chronic lung disease diagnosis and prognosisabstractMOTIVATION: Integration of data from different modalities is a necessary step for multi-scale data analysis in many fields, including biomedical research and systems biology. Directed graphical models offer an attractive tool for this problem because they can represent both the complex, multivariate probability distributions and the causal pathways influencing the system. Graphical models learned from biomedical data can be used for classification, biomarker selection and functional analysis, while revealing the underlying network structure and thus allowing for arbitrary likelihood queries over the data. RESULTS: In this paper, we present and test new methods for finding directed graphs over mixed data types (continuous and discrete variables). We used this new algorithm, CausalMGM, to identify variables directly linked to disease diagnosis and progression in various multi-modal datasets, including clinical datasets from chronic obstructive pulmonary disease (COPD). COPD is the third leading cause of death and a major cause of disability and thus determining the factors that cause longitudinal lung function decline is very important. Applied on a COPD dataset, mixed graphical models were able to confirm and extend previously described causal effects and provide new insights on the factors that potentially affect the longitudinal lung function decline of COPD patients. AVAILABILITY AND IMPLEMENTATION: The CausalMGM package is available on http://www.causalmgm.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andrew J. Sedgewick, Kristina Buschur, Ivy Shi, Joseph D. Ramsey, Vineet K. Raghu, Dimitris V. Manatakis, Yingze Zhang, Jessica Bon, Divay Chandra, Chad Karoleski, Frank C. Sciurba, Peter Spirtes, Clark Glymour, Panayiotis V. Benos |
Bioinform. | 13 |
| 2018 | Generalized Score Functions for Causal DiscoveryabstractDiscovery of causal relationships from observational data is a fundamental problem. Roughly speaking, there are two types of methods for causal discovery, constraint-based ones and score-based ones. Score-based methods avoid the multiple testing problem and enjoy certain advantages compared to constraint-based ones. However, most of them need strong assumptions on the functional forms of causal mechanisms, as well as on data distributions, which limit their applicability. In practice the precise information of the underlying model class is usually unknown. If the above assumptions are violated, both spurious and missing edges may result. In this paper, we introduce generalized score functions for causal discovery based on the characterization of general (conditional) independence relationships between random variables, without assuming particular model classes. In particular, we exploit regression in RKHS to capture the dependence in a non-parametric way. The resulting causal discovery approach produces asymptotically correct results in rather general cases, which may have nonlinear causal mechanisms, a wide class of data distributions, mixed continuous and discrete data, and multidimensional variables. Experimental results on both synthetic and real-world data demonstrate the efficacy of our proposed approach. Biwei Huang, Kun Zhang 0001, Yizhu Lin, Bernhard Schölkopf, Clark Glymour |
KDD | 5 |
| 2018 | Causal Discovery with Linear Non-Gaussian Models under Measurement Error: Structural Identifiability Results
Kun Zhang 0001, Mingming Gong, Joseph D. Ramsey, Kayhan Batmanghelich, Peter Spirtes, Clark Glymour |
UAI | 6 |
| 2017 | Behind Distribution Shift: Mining Driving Forces of Changes and Causal ArrowsabstractWe address two important issues in causal discovery from nonstationary or heterogeneous data, where parameters associated with a causal structure may change over time or across data sets. First, we investigate how to efficiently estimate the "driving force" of the nonstationarity of a causal mechanism. That is, given a causal mechanism that varies over time or across data sets and whose qualitative structure is known, we aim to extract from data a low-dimensional and interpretable representation of the main components of the changes. For this purpose we develop a novel kernel embedding of nonstationary conditional distributions that does not rely on sliding windows. Second, the embedding also leads to a measure of dependence between the changes of causal modules that can be used to determine the directions of many causal arrows. We demonstrate the power of our methods with experiments on both synthetic and real data. Biwei Huang, Kun Zhang 0001, Jiji Zhang, Ruben Sanchez-Romero, Clark Glymour, Bernhard Schölkopf |
ICDM | 5 |
| 2017 | Causal Discovery from Nonstationary/Heterogeneous Data: Skeleton Estimation and Orientation DeterminationabstractIt is commonplace to encounter nonstationary or heterogeneous data, of which the underlying generating process changes over time or across data sets (the data sets may have different experimental conditions or data collection conditions). Such a distribution shift feature presents both challenges and opportunities for causal discovery. In this paper we develop a principled framework for causal discovery from such data, called Constraint-based causal Discovery from Nonstationary/heterogeneous Data (CD-NOD), which addresses two important questions. First, we propose an enhanced constraint-based procedure to detect variables whose local mechanisms change and recover the skeleton of the causal structure over observed variables. Second, we present a way to determine causal orientations by making use of independence changes in the data distribution implied by the underlying causal model, benefiting from information carried by changing distributions. Experimental results on various synthetic and real-world data sets are presented to demonstrate the efficacy of our methods. Kun Zhang 0001, Biwei Huang, Jiji Zhang, Clark Glymour, Bernhard Schölkopf |
IJCAI | 4 |
| 2017 | Causal Discovery from Temporally Aggregated Time Series
Mingming Gong, Kun Zhang 0001, Bernhard Schölkopf, Clark Glymour, Dacheng Tao |
UAI | 4 |
| 2016 | Domain Adaptation with Conditional Transferable ComponentsabstractDomain adaptation arises in supervised learning when the training (source domain) and test (target domain) data have different distributions. Let X and Y denote the features and target, respectively, previous work on domain adaptation considers the covariate shift situation where the distribution of the features P(X) changes across domains while the conditional distribution P(Y|X) stays the same. To reduce domain discrepancy, recent methods try to find invariant components \mathcalT(X) that have similar P(\mathcalT(X)) by explicitly minimizing a distribution discrepancy measure. However, it is not clear if P(Y|\mathcalT(X)) in different domains is also similar when P(Y|X) changes. Furthermore, transferable components do not necessarily have to be invariant. If the change in some components is identifiable, we can make use of such components for prediction in the target domain. In this paper, we focus on the case where P(X|Y) and P(Y) both change in a causal system in which Y is the cause for X. Under appropriate assumptions, we aim to extract conditional transferable components whose conditional distribution P(\mathcalT(X)|Y) is invariant after proper location-scale (LS) transformations, and identify how P(Y) changes between domains simultaneously. We provide theoretical analysis and empirical evaluation on both synthetic and real-world data to show the effectiveness of our method. Mingming Gong, Kun Zhang 0001, Tongliang Liu, Dacheng Tao, Clark Glymour, Bernhard Schölkopf |
ICML | 5 |
| 2016 | On the Identifiability and Estimation of Functional Causal Models in the Presence of Outcome-Dependent Selection
Kun Zhang 0001, Jiji Zhang, Biwei Huang, Bernhard Schölkopf, Clark Glymour |
UAI | 5 |
| 2015 | The center for causal discovery of biomedical knowledge from big dataabstractThe Big Data to Knowledge (BD2K) Center for Causal Discovery is developing and disseminating an integrated set of open source tools that support causal modeling and discovery of biomedical knowledge from large and complex biomedical datasets. The Center integrates teams of biomedical and data scientists focused on the refinement of existing and the development of new constraint-based and Bayesian algorithms based on causal Bayesian networks, the optimization of software for efficient operation in a supercomputing environment, and the testing of algorithms and software developed using real data from 3 representative driving biomedical projects: cancer driver mutations, lung disease, and the functional connectome of the human brain. Associated training activities provide both biomedical and data scientists with the knowledge and skills needed to apply and extend these tools. Collaborative activities with the BD2K Consortium further advance causal discovery tools and integrate tools and resources developed by other centers. Gregory F. Cooper, Ivet Bahar, Michael J. Becich, Panayiotis V. Benos, Jeremy M. Berg, Jeremy U. Espino, Clark Glymour, Rebecca S. Jacobson, Michelle Kienholz, Adrian V. Lee, Xinghua Lu 0001, Richard Scheines |
J. Am. Medical Informatics Assoc. | 7 |
| 2008 | Integrating Locally Learned Causal Structures with Overlapping VariablesabstractIn many domains, data are distributed among datasets that share only some variables; other recorded variables may occur in only one dataset. There are several asymptotically correct, informative algorithms that search for causal information given a single dataset, even with missing values and hidden variables. There are, however, no such reliable procedures for distributed data with overlapping variables, and only a single heuristic procedure (Structural EM). This paper describes an asymptotically correct procedure, ION, that provides all the information about structure obtainable from the marginal independence relations. Using simulated and real data, the accuracy of ION is compared with that of Structural EM, and with inference on complete, unified data. Robert E. Tillman, David Danks, Clark Glymour |
NIPS | 3 |
| 2008 | Search for Additive Nonlinear Time Series Causal Models
Tianjiao Chu, Clark Glymour |
J. Mach. Learn. Res. | 2 |
| 2006 | Learning the Structure of Linear Latent Variable ModelsabstractWe describe anytime search procedures that (1) find disjoint subsets of recorded variables for which the members of each subset are d-separated by a single common unrecorded cause, if such exists; (2) return information about the causal relations among the latent factors so identified. We prove the procedure is point-wise consistent assuming (a) the causal relations can be represented by a directed acyclic graph (DAG) satisfying the Markov Assumption and the Faithfulness Assumption; (b) unrecorded variables are not caused by recorded variables; and (c) dependencies are linear. We compare the procedure with standard approaches over a variety of simulated structures and sample sizes, and illustrate its practical value with brief studies of social science data sets. Finally, we consider generalizations for non-linear systems. Ricardo Bezerra de Andrade e Silva, Richard Scheines, Clark Glymour, Peter Spirtes |
J. Mach. Learn. Res. | 3 |
| 2005 | On the Number of Experiments Sufficient and in the Worst Case Necessary to Identify All Causal Relations Among N Variables
Frederick Eberhardt, Clark Glymour, Richard Scheines |
UAI | 2 |
| 2003 | Learning Measurement Models for Unobserved Variables
Ricardo Bezerra de Andrade e Silva, Richard Scheines, Clark Glymour, Peter Spirtes |
UAI | 3 |
| 2003 | A Statistical Problem for Inference to Regulatory Structure from Associations of Gene Expression Measurements with MicroarraysabstractMOTIVATION: One approach to inferring genetic regulatory structure from microarray measurements of mRNA transcript hybridization is to estimate the associations of gene expression levels measured in repeated samples. The associations may be estimated by correlation coefficients or by conditional frequencies (for discretized measurements) or by some other statistic. Although these procedures have been successfully applied to other areas, their validity when applied to microarray measurements has yet to be tested. RESULTS: This paper describes an elementary statistical difficulty for all such procedures, no matter whether based on Bayesian updating, conditional independence testing, or other machine learning procedures such as simulated annealing or neural net pruning. The difficulty obtains if a number of cells from a common population are aggregated in a measurement of expression levels. Although there are special cases where the conditional associations are preserved under aggregation, in general inference of genetic regulatory structure based on conditional association is unwarranted Tianjiao Chu, Clark Glymour, Richard Scheines, Peter Spirtes |
Bioinform. | 2 |
| 2002 | Automated Remote Sensing with Near Infrared Reflectance Spectra: Carbonate Recognition
Joseph D. Ramsey, Paul Gazis, Ted Roush, Peter Spirtes, Clark Glymour |
Data Min. Knowl. Discov. | 5 |
| 2002 | Classification and filtering of spectra: A case study in mineralogy
Jonathan Moody, Ricardo Bezerra de Andrade e Silva, Joseph Vanderwaart, Joseph D. Ramsey, Clark Glymour |
Intell. Data Anal. | 5 |
| 2001 | Linearity Properties of Bayes Nets with Binary Variables
David Danks, Clark Glymour |
UAI | 2 |
| 1998 | Psychological and Normative Theories of Causal Power and the Probabilities of Causes
Clark Glymour |
UAI | 1 |
| 1997 | An evaluation of machine-learning methods for predicting pneumonia mortality
Gregory F. Cooper, Constantin F. Aliferis, Richard Ambrosino, John M. Aronis, Bruce G. Buchanan, Rich Caruana, Michael J. Fine, Clark Glymour, Geoffrey J. Gordon, Barbara H. Hanusa, Janine E. Janosky, Christopher Meek, Tom M. Mitchell, Thomas Richardson 0001, Peter Spirtes |
Artif. Intell. Medicine | 8 |
| 1997 | Statistical Themes and Lessons for Data Mining
Clark Glymour, David Madigan, Daryl Pregibon, Padhraic Smyth |
Data Min. Knowl. Discov. | 1 |
| 1995 | Available Technology for Discovering Causal Models, Building Bayes Nets, and Selecting Predictors: The TETRAD II Program
Clark Glymour |
KDD | 1 |
| 1985 | Independence Assumptions and Bayesian Updating
Clark Glymour |
Artif. Intell. | 1 |