VLDB 2026 Research / reviewers in the wild / expert
Gregory F. Cooper
dblp:96/2706
· DBLP profile ↗
128ranked-venue papers
21as first author
9since 2021 · last 2025
0000-0002-9276-773XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 66 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 52 · 11 first-author · 5 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IBI-DT: a novel approach combining individualized Bayesian inference and decision tree for identifying cancer drivers and their interactionsabstractCancer is mainly caused by a relatively small portion of somatic genome alterations (SGAs), called cancer drivers. Despite success in identifying a good number of cancer drivers, many more remain to be discovered to explain various cancers. Moreover, limited tools are available to identify potential interactions among cancer drivers for a better understanding of oncogenesis. To tackle these challenges, we have developed a novel approach called individualized Bayesian inference using a decision tree (IBI-DT). IBI-DT recognizes the genetic heterogeneity among cancer patients, where different individuals or patient subgroups of distinct genomic makeup may have different drivers. IBI-DT works by constructing smaller subgroups with similar genetic makeup (i.e. patient-like-me subgroups) using a decision tree structure and analyzing multiple trees to identify the SGAs that play a significant role in regulating downstream gene expression patterns at the subgroup and individual levels. This is distinct from population-based approaches, which tend to evaluate the influence of an SGA for the entire population, thereby likely missing low-frequency SGAs that may well explain a small subgroup of cancer patients. Also importantly, IBI-DT can efficiently identify cancer drivers that may have functional interactions. We applied IBI-DT to identify cancer drivers regulating the downstream differential gene expression in cancer patients and compared it to the standard, population-based method of expression quantitative trait loci analysis. Our results show that IBI-DT performs well in identifying both important cancer drivers, especially the low-frequency drivers, and their interactions, allowing for a better understanding of the cancer signaling pathways. Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu 0001, Jinling Liu |
Briefings Bioinform. | 2 |
| 2023 | Learning Treatment Effects from Observational and Experimental DataabstractDecision making often depends on causal effect estimation. For example, clinical decisions are often based on estimates of the probability of post-treatment outcomes. Experimental data from randomized controlled trials allow for unbiased estimation of these probabilities. However, such data are usually limited in the number of samples and the set of measured covariates. Observational data, such as electronic medical records, contain many more samples and a richer set of measured covariates, which can be used to estimate more personalized treatment effects; however, these estimates may be biased due to latent confounding. In this work, we propose a Bayesian method for combining observational and experimental data for unbiased conditional treatment effect estimation. Our method addresses the following question: Given observational data $D_o$ measuring a set of covariates $\mathbf V$, and experimental data $D_e$ measuring a possibly smaller set of covariates $\mathbf{V_b}\subseteq \mathbf{V}$, which set of covariates $\mathbf{Z}$ leads to the optimal, unbiased prediction of the post-intervention outcome $P(Y |do(X), \mathbf{Z})$, and when can we use observational data for this estimation? In simulated data, we show that our method improves the prediction of post-intervention outcomes. Sofia Triantafyllou, Fattaneh Jabbari, Gregory F. Cooper |
AISTATS | 3 |
| 2023 | A new method for estimating the probability of causal relationships from observational data: Application to the study of the short-term effects of air pollution on cardiovascular and respiratory disease
Bryan Andrews, Chirayu Wongchokprasitti, Shyam Visweswaran, Chirag M. Lakhani, Chirag J. Patel, Gregory F. Cooper |
Artif. Intell. Medicine | 6 |
| 2023 | A voice-based digital assistant for intelligent prompting of evidence-based practices during ICU rounds
Andrew J. King 0002, Derek C. Angus, Gregory F. Cooper, Danielle L. Mowery, Jennifer B. Seaman, Kelly M. Potter, Leigh A. Bukowski, Ali Al-Khafaji, Scott R. Gunn, Jeremy M. Kahn |
J. Biomed. Informatics | 3 |
| 2022 | An individualized causal framework for learning intercellular communication networks that define microenvironments of individual tumorsabstractCells within a tumor microenvironment (TME) dynamically communicate and influence each other's cellular states through an intercellular communication network (ICN). In cancers, intercellular communications underlie immune evasion mechanisms of individual tumors. We developed an individualized causal analysis framework for discovering tumor specific ICNs. Using head and neck squamous cell carcinoma (HNSCC) tumors as a testbed, we first mined single-cell RNA-sequencing data to discover gene expression modules (GEMs) that reflect the states of transcriptomic processes within tumor and stromal single cells. By deconvoluting bulk transcriptomes of HNSCC tumors profiled by The Cancer Genome Atlas (TCGA), we estimated the activation states of these transcriptomic processes in individual tumors. Finally, we applied individualized causal network learning to discover an ICN within each tumor. Our results show that cellular states of cells in TMEs are coordinated through ICNs that enable multi-way communications among epithelial, fibroblast, endothelial, and immune cells. Further analyses of individual ICNs revealed structural patterns that were shared across subsets of tumors, leading to the discovery of 4 different subtypes of networks that underlie disparate TMEs of HNSCC. Patients with distinct TMEs exhibited significantly different clinical outcomes. Our results show that the capability of estimating individual ICNs reveals heterogeneity of ICNs and sheds light on the importance of intercellular communication in impacting disease development and progression. Xueer Chen, Lujia Chen 0001, Cornelius H. L. Kürten, Fattaneh Jabbari, Lazar Vujanovic, Ying Ding 0003, Binfeng Lu, Aditi Kulkarni, Tracy Tabib, Robert Lafyatis, Gregory F. Cooper, Robert Ferris, Xinghua Lu 0001 |
PLoS Comput. Biol. | 12 |
| 2021 | Learning Adjustment Sets from Observational and Limited Experimental Data
Sofia Triantafyllou, Gregory F. Cooper |
AAAI | 2 |
| 2021 | Patient-Specific Modeling with Lazy Random Forest (LazyRF)
Adriana Johnson, Gregory F. Cooper, Shyam Visweswaran |
AMIA | 2 |
| 2021 | The KDD 2021 Workshop on Causal Discovery (CD2021)abstractAs a basic and effective tool for explanation, prediction and decision making, causal relationships have been utilized in almost all disciplines. Traditionally, causal relationships are identified by making use of interventions or randomized controlled experiments. However, conducting such experiments is often expensive or even impossible due to cost or ethical concerns. Therefore, there has been an increasing interest in discovering causal relationships based on observational data, and in the past few decades, significant contributions have been made to this field by computer scientists. Thuc Duy Le, Jiuyong Li, Gregory F. Cooper, Sofia Triantafyllou, Elias Bareinboim, Huan Liu 0001, Negar Kiyavash |
KDD | 3 |
| 2021 | Causal and interventional Markov boundariesabstractFeature selection is an important problem in machine learning, which aims to select variables that lead to an optimal predictive model. In this paper, we focus on feature selection for post-intervention outcome prediction from pre-intervention variables. We are motivated by healthcare settings, where the goal is often to select the treatment that will maximize a specific patient’s outcome; however, we often do not have sufficient randomized control trial data to identify well the conditional treatment effect. We show how we can use observational data to improve feature selection and effect estimation in two cases: (a) using observational data when we know the causal graph, and (b) when we do not know the causal graph but have observational and limited experimental data. Our paper extends the notion of Markov boundary to treatment-outcome pairs. We provide theoretical guarantees for the methods we introduce. In simulated data, we show that combining observational and experimental data improves feature selection and effect estimation. Sofia Triantafyllou, Fattaneh Jabbari, Gregory F. Cooper |
UAI | 3 |
| 2020 | Lung Cancer Survival Prediction Using Instance-Specific Bayesian Networks
Fattaneh Jabbari, Liza C. Villaruz, Gregory F. Cooper |
AIME | 4 |
| 2020 | Patient-Specific Modeling with Personalized Decision Paths
Adriana Johnson, Gregory F. Cooper, Shyam Visweswaran |
AMIA | 2 |
| 2020 | An Instance-Specific Algorithm for Learning the Structure of Causal Bayesian Networks Containing Latent VariablesabstractAlmost all of the existing algorithms for learning a causal Bayesian network structure (CBN) from observational data recover a structure that models the causal relationships that are shared by the instances in a population. Although learning such population-wide CBNs accurately is useful, it is important to learn CBNs that are specific to each instance in domains in which different instances may have varying causal structures, such as in human biology. For example, a breast cancer tumor in a patient (instance) is often a composite of causal mechanisms, where each of these individual causal mechanisms may appear relatively frequently in breast-cancer tumors of other patients, but the particular combination of mechanisms is unique to the current tumor. Therefore, it is critical to discover the specific set of causal mechanisms that are operating in each patient to understand and treat that particular patient effectively. We previously introduced an instance-specific CBN structure learning method that builds a causal model for a given instance T from the features we know about T and from a training set of data on many other instances [12]. However, that method assumes that there are no latent (hidden) confounders, that is, there are no latent variables that cause two or more of the measured variables. Unfortunately, this assumption rarely holds in practice. In the current paper, we introduce a novel instance-specific causal structure learning algorithm that uses partial ancestral graphs (PAGs) to model latent confounders. Simulations support that the proposed instance-specific method improves structure-discovery performance compared to an existing PAG-learning method called GFCI, which is not instance-specific. We also report results that provide support for instance-specific causal relationships existing in real-world datasets. Fattaneh Jabbari, Gregory F. Cooper |
SDM | 2 |
| 2020 | Explicit representation of protein activity states significantly improves causal discovery of protein phosphorylation networksabstractBACKGROUND: Protein phosphorylation networks play an important role in cell signaling. In these networks, phosphorylation of a protein kinase usually leads to its activation, which in turn will phosphorylate its downstream target proteins. A phosphorylation network is essentially a causal network, which can be learned by causal inference algorithms. Prior efforts have applied such algorithms to data measuring protein phosphorylation levels, assuming that the phosphorylation levels represent protein activity states. However, the phosphorylation status of a kinase does not always reflect its activity state, because interventions such as inhibitors or mutations can directly affect its activity state without changing its phosphorylation status. Thus, when cellular systems are subjected to extensive perturbations, the statistical relationships between phosphorylation states of proteins may be disrupted, making it difficult to reconstruct the true protein phosphorylation network. Here, we describe a novel framework to address this challenge. RESULTS: We have developed a causal discovery framework that explicitly represents the activity state of each protein kinase as an unmeasured variable and developed a novel algorithm called "InferA" to infer the protein activity states, which allows us to incorporate the protein phosphorylation level, pharmacological interventions and prior knowledge. We applied our framework to simulated datasets and to a real-world dataset. The simulation experiments demonstrated that explicit representation of activity states of protein kinases allows one to effectively represent the impact of interventions and thus enabled our framework to accurately recover the ground-truth causal network. Results from the real-world dataset showed that the explicit representation of protein activity states allowed an effective and data-driven integration of the prior knowledge by InferA, which further leads to the recovery of a phosphorylation network that is more consistent with experiment results. CONCLUSIONS: Explicit representation of the protein activity states by our novel framework significantly enhances causal discovery of protein phosphorylation networks. Jinling Liu, Gregory F. Cooper, Xinghua Lu 0001 |
BMC Bioinform. | 3 |
| 2019 | Insights from a dissertation on the development of a Learning Electronic Medical Record System: data-driven, context-aware learning
Andrew J. King 0002, Shyam Visweswaran, Harry Hochheiser, Gilles Clermont, Gregory F. Cooper |
AMIA | 5 |
| 2019 | An Empirical Investigation of Instance-Specific Causal Bayesian Network LearningabstractSignificant progress has been made in developing algorithms for learning graphical causal models from data. Most of these algorithms learn a causal structure that is shared by all the instances (e.g., patients) in the training dataset. However, different instances may not all share the same causal structure. We introduced an instance-specific method called IGES [15] that learns a causal model for each instance T by using the features of T and the instances in the training dataset. In the current paper, we study the empirical performance of the IGES method on several biomedical datasets. The results provide support that instance-specific structure exists and is important to model in these real domains. Fattaneh Jabbari, Shyam Visweswaran, Gregory F. Cooper |
BIBM | 3 |
| 2019 | Using machine learning to selectively highlight patient information
Andrew J. King 0002, Gregory F. Cooper, Gilles Clermont, Harry Hochheiser, Milos Hauskrecht, Dean F. Sittig, Shyam Visweswaran |
J. Biomed. Informatics | 2 |
| 2019 | Systematic discovery of the functional impact of somatic genome alterations in individual tumors through tumor-specific causal inferenceabstractCancer is mainly caused by somatic genome alterations (SGAs). Precision oncology involves identifying and targeting tumor-specific aberrations resulting from causative SGAs. We developed a novel tumor-specific computational framework that finds the likely causative SGAs in an individual tumor and estimates their impact on oncogenic processes, which suggests the disease mechanisms that are acting in that tumor. This information can be used to guide precision oncology. We report a tumor-specific causal inference (TCI) framework, which estimates causative SGAs by modeling causal relationships between SGAs and molecular phenotypes (e.g., transcriptomic, proteomic, or metabolomic changes) within an individual tumor. We applied the TCI algorithm to tumors from The Cancer Genome Atlas (TCGA) and estimated for each tumor the SGAs that causally regulate the differentially expressed genes (DEGs) in that tumor. Overall, TCI identified 634 SGAs that are predicted to cause cancer-related DEGs in a significant number of tumors, including most of the previously known drivers and many novel candidate cancer drivers. The inferred causal relationships are statistically robust and biologically sensible, and multiple lines of experimental evidence support the predicted functional impact of both the well-known and the novel candidate drivers that are predicted by TCI. TCI provides a unified framework that integrates multiple types of SGAs and molecular phenotypes to estimate which genome perturbations are causally influencing one or more molecular/cellular phenotypes in an individual tumor. By identifying major candidate drivers and revealing their functional impact in an individual tumor, TCI sheds light on the disease mechanisms of that tumor, which can serve to advance our basic knowledge of cancer biology and to support precision oncology that provides tailored treatment of individual tumors. Chunhui Cai, Gregory F. Cooper, Kevin N. Lu, Shuping Xu, Zhenlong Zhao, Xueer Chen, Adrian V. Lee, Nathan Clark, Vicky Chen, Songjian Lu, Lujia Chen 0001, Liyue Yu, Harry Hochheiser, Xia Jiang, Q. Jane Wang, Xinghua Lu 0001 |
PLoS Comput. Biol. | 2 |
| 2018 | Design of a Learning Electronic Medical Record: A Qualitative Study of ICU Clinicians' Information Needs and Practices
Luca Calzoni, Gilles Clermont, Gregory F. Cooper, Shyam Visweswaran, Harry Hochheiser |
AMIA | 3 |
| 2018 | Using Machine Learning to Predict the Information Seeking Behavior of Clinicians Using an Electronic Medical Record System
Andrew J. King 0002, Gregory F. Cooper, Harry Hochheiser, Gilles Clermont, Milos Hauskrecht, Shyam Visweswaran |
AMIA | 2 |
| 2018 | Binary classifier calibration using an ensemble of piecewise linear regression models
Mahdi Pakdaman Naeini, Gregory F. Cooper |
Knowl. Inf. Syst. | 2 |
| 2017 | Exploring Novel Graphical Representations of Clinical Data in a Learning EMR
Luca Calzoni, Gilles Clermont, Gregory F. Cooper, Shyam Visweswaran, Harry Hochheiser |
AMIA | 3 |
| 2017 | Discovery of Causal Models that Contain Latent Variables Through Bayesian Scoring of Independence Constraints
Fattaneh Jabbari, Joseph D. Ramsey, Peter Spirtes, Gregory F. Cooper |
ECML/PKDD (2) | 4 |
| 2017 | A Bayesian system to detect and characterize overlapping outbreaks
John M. Aronis, Nicholas Millett, Michael M. Wagner 0001, Fu-Chiang Tsui, Ye Ye 0002, Jeffrey P. Ferraro, Peter J. Haug, Per H. Gesteland, Gregory F. Cooper |
J. Biomed. Informatics | 9 |
| 2016 | Binary Classifier Calibration Using an Ensemble of Near Isotonic Regression ModelsabstractLearning accurate probabilistic models from data is crucial in many practical tasks in data mining. In this paper we present a new non-parametric calibration method called Ensemble of Near Isotonic Regression (ENIR). The method can be considered as an extension of BBQ, a recently proposed calibration method, as well as the commonly used calibration method based on isotonic regression (IsoRegC). ENIR is designed to address the key limitation of IsoRegC which is the monotonicity assumption of the predictions. Similar to BBQ, the method post-processes the output of a binary classifier to obtain calibrated probabilities. Thus it can be used with many existing classification models to generate accurate probabilistic predictions. We demonstrate the performance of ENIR on synthetic and real datasets for commonly applied binary classification models. Experimental results show that the method outperforms several common binary classifier calibration methods. In particular on the real data, ENIR commonly performs statistically significantly better than the other methods, and never worse. It is able to improve the calibration power of classifiers, while retaining their discrimination power. The method is also computationally tractable for large scale datasets, as it is O(N log N) time, where N is the number of samples. Mahdi Pakdaman Naeini, Gregory F. Cooper |
ICDM | 2 |
| 2016 | Binary Classifier Calibration Using an Ensemble of Linear Trend EstimationabstractLearning accurate probabilistic models from data is crucial in many practical tasks in data mining. In this paper we present a new non-parametric calibration method called ensemble of linear trend estimation (ELiTE). ELiTE utilizes the recently proposed ℓ1 trend filtering signal approximation method [22] to find the mapping from uncalibrated classification scores to the calibrated probability estimates. ELiTE is designed to address the key limitations of the histogram binning-based calibration methods which are (1) the use of a piecewise constant form of the calibration mapping using bins, and (2) the assumption of independence of predicted probabilities for the instances that are located in different bins. The method post-processes the output of a binary classifier to obtain calibrated probabilities. Thus, it can be applied with many existing classification models. We demonstrate the performance of ELiTE on real datasets for commonly used binary classification models. Experimental results show that the method outperforms several common binary-classifier calibration methods. In particular, ELiTE commonly performs statistically significantly better than the other methods, and never worse. Moreover, it is able to improve the calibration power of classifiers, while retaining their discrimination power. The method is also computationally tractable for large scale datasets, as it is practically O(N log N) time, where N is the number of samples. Mahdi Pakdaman Naeini, Gregory F. Cooper |
SDM | 2 |
| 2016 | Outlier-based detection of unusual patient-management actions: An ICU study
Milos Hauskrecht, Iyad Batal, Charmgil Hong, Gregory F. Cooper, Shyam Visweswaran, Gilles Clermont |
J. Biomed. Informatics | 5 |
| 2016 | An efficient pattern mining approach for event detection in multivariate temporal data
Iyad Batal, Gregory F. Cooper, Dmitriy Fradkin, James H. Harrison Jr., Fabian Mörchen, Milos Hauskrecht |
Knowl. Inf. Syst. | 2 |
| 2015 | Obtaining Well Calibrated Probabilities Using Bayesian BinningabstractLearning probabilistic predictive models that are well calibrated is critical for many prediction and decision-making tasks in artificial intelligence. In this paper we present a new non-parametric calibration method called Bayesian Binning into Quantiles (BBQ) which addresses key limitations of existing calibration methods. The method post processes the output of a binary classification algorithm; thus, it can be readily combined with many existing classification algorithms. The method is computationally tractable, and empirically accurate, as evidenced by the set of experiments reported here on both real and simulated datasets. Mahdi Pakdaman Naeini, Gregory F. Cooper, Milos Hauskrecht |
AAAI | 2 |
| 2015 | Development and Preliminary Evaluation of a Prototype of a Learning Electronic Medical Record System
Andrew J. King 0002, Gregory F. Cooper, Harry Hochheiser, Gilles Clermont, Shyam Visweswaran |
AMIA | 2 |
| 2015 | A Bayesian Approach for Identifying Multivariate Differences Between Groups
Yuriy Sverchkov, Gregory F. Cooper |
IDA | 2 |
| 2015 | Binary Classifier Calibration Using a Bayesian Non-Parametric ApproachabstractLearning probabilistic predictive models that are well calibrated is critical for many prediction and decision-making tasks in Data mining. This paper presents two new non-parametric methods for calibrating outputs of binary classification models: a method based on the Bayes optimal selection and a method based on the Bayesian model averaging. The advantage of these methods is that they are independent of the algorithm used to learn a predictive model, and they can be applied in a post-processing step, after the model is learned. This makes them applicable to a wide variety of machine learning models and methods. These calibration methods, as well as other methods, are tested on a variety of datasets in terms of both discrimination and calibration performance. The results show the methods either outperform or are comparable in performance to the state-of-the-art calibration methods. Mahdi Pakdaman Naeini, Gregory F. Cooper, Milos Hauskrecht |
SDM | 2 |
| 2015 | The center for causal discovery of biomedical knowledge from big dataabstractThe Big Data to Knowledge (BD2K) Center for Causal Discovery is developing and disseminating an integrated set of open source tools that support causal modeling and discovery of biomedical knowledge from large and complex biomedical datasets. The Center integrates teams of biomedical and data scientists focused on the refinement of existing and the development of new constraint-based and Bayesian algorithms based on causal Bayesian networks, the optimization of software for efficient operation in a supercomputing environment, and the testing of algorithms and software developed using real data from 3 representative driving biomedical projects: cancer driver mutations, lung disease, and the functional connectome of the human brain. Associated training activities provide both biomedical and data scientists with the knowledge and skills needed to apply and extend these tools. Collaborative activities with the BD2K Consortium further advance causal discovery tools and integrate tools and resources developed by other centers. Gregory F. Cooper, Ivet Bahar, Michael J. Becich, Panayiotis V. Benos, Jeremy M. Berg, Jeremy U. Espino, Clark Glymour, Rebecca S. Jacobson, Michelle Kienholz, Adrian V. Lee, Xinghua Lu 0001, Richard Scheines |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | A method for detecting and characterizing outbreaks of infectious disease from clinical reports
Gregory F. Cooper, Ricardo Villamarín-Salomón, Fu-Chiang Tsui, Nicholas Millett, Jeremy U. Espino, Michael M. Wagner 0001 |
J. Biomed. Informatics | 1 |
| 2015 | Comparison of machine learning classifiers for influenza detection from emergency department free-text reports
Arturo L. Pineda, Ye Ye 0002, Shyam Visweswaran, Gregory F. Cooper, Michael M. Wagner 0001, Fu-Chiang Tsui |
J. Biomed. Informatics | 4 |
| 2014 | Application of Bayesian Logistic Regression to Mining Biomedical Data
Viji R. Avali, Gregory F. Cooper, Vanathi Gopalakrishnan |
AMIA | 2 |
| 2013 | Decision Path Models for Patient-Specific Modeling of Patient Outcomes
Antonio Luiz S. Ferreira, Gregory F. Cooper, Shyam Visweswaran |
AMIA | 2 |
| 2013 | Data-driven identification of unusual clinical actions in the ICU
Milos Hauskrecht, Shyam Visweswaran, Gregory F. Cooper, Gilles Clermont |
AMIA | 3 |
| 2013 | Outlier detection for patient monitoring and alerting
Milos Hauskrecht, Iyad Batal, Michal Valko, Shyam Visweswaran, Gregory F. Cooper, Gilles Clermont |
J. Biomed. Informatics | 5 |
| 2013 | A method for estimating from thermometer sales the incidence of diseases that are symptomatically similar to influenza
Ricardo Villamarín-Salomón, Gregory F. Cooper, Michael M. Wagner 0001, Fu-Chiang Tsui, Jeremy U. Espino |
J. Biomed. Informatics | 2 |
| 2013 | A temporal pattern mining approach for classifying electronic health record dataabstractWe study the problem of learning classification models from complex multivariate temporal data encountered in electronic health record systems. The challenge is to define a good set of features that are able to represent well the temporal aspect of the data. Our method relies on temporal abstractions and temporal pattern mining to extract the classification features. Temporal pattern mining usually returns a large number of temporal patterns, most of which may be irrelevant to the classification task. To address this problem, we present the Minimal Predictive Temporal Patterns framework to generate a small set of predictive and non-spurious patterns. We apply our approach to the real-world clinical task of predicting patients who are at risk of developing heparin induced thrombocytopenia. The results demonstrate the benefit of our approach in efficiently learning accurate classifiers, which is a key step for developing intelligent clinical monitoring systems. Iyad Batal, Hamed Valizadegan, Gregory F. Cooper, Milos Hauskrecht |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2012 | A Bayesian Scoring Technique for Mining Predictive and Non-Spurious Rules
Iyad Batal, Gregory F. Cooper, Milos Hauskrecht |
ECML/PKDD (2) | 2 |
| 2012 | Improving the Prediction of Clinical Outcomes from Genomic Data Using Multiresolution AnalysisabstractThe prediction of patient's future clinical outcome, such as Alzheimer's and cardiac disease, using only genomic information is an open problem. In cases when genome-wide association studies (GWASs) are able to find strong associations between genomic predictors (e.g., SNPs) and disease, pattern recognition methods may be able to predict the disease well. Furthermore, by using signal processing methods, we can capitalize on latent multivariate interactions of genomic predictors. Such an approach to genomic pattern recognition for prediction of clinical outcomes is investigated in this work. In particular, we show how multiresolution transforms can be applied to genomic data to extract cues of multivariate interactions and, in some cases, improve on the predictive performance of clinical outcomes of standard classification methods. Our results show, for example, that an improvement of about 6 percent increase of the area under the ROC curve can be achieved using multiresolution spaces to train logistic regression to predict late-onset Alzheimer's disease (LOAD) compared to logistic regression applied directly on SNP data. Pablo H. Hennings-Yeomans, Gregory F. Cooper |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | A Pattern Mining Approach for Classifying Multivariate Temporal DataabstractWe study the problem of learning classification models from complex multivariate temporal data encountered in electronic health record systems. The challenge is to define a good set of features that are able to represent well the temporal aspect of the data. Our method relies on temporal abstractions and temporal pattern mining to extract the classification features. Temporal pattern mining usually returns a large number of temporal patterns, most of which may be irrelevant to the classification task. To address this problem, we present the minimal predictive temporal patterns framework to generate a small set of predictive and non-spurious patterns. We apply our approach to the real-world clinical task of predicting patients who are at risk of developing heparin induced thrombocytopenia. The results demonstrate the benefit of our approach in learning accurate classifiers, which is a key step for developing intelligent clinical monitoring systems. Iyad Batal, Hamed Valizadegan, Gregory F. Cooper, Milos Hauskrecht |
BIBM | 3 |
| 2011 | Conditional Anomaly Detection with Soft Harmonic FunctionsabstractIn this paper, we consider the problem of conditional anomaly detection that aims to identify data instances with an unusual response or a class label. We develop a new non-parametric approach for conditional anomaly detection based on the soft harmonic solution, with which we estimate the confidence of the label to detect anomalous mislabeling. We further regularize the solution to avoid the detection of isolated examples and examples on the boundary of the distribution support. We demonstrate the efficacy of the proposed method on several synthetic and UCI ML datasets in detecting unusual labels when compared to several baseline approaches. We also evaluate the performance of our method on a real-world electronic health record dataset where we seek to identify unusual patient-management decisions. Michal Valko, Branislav Kveton, Hamed Valizadegan, Gregory F. Cooper, Milos Hauskrecht |
ICDM | 4 |
| 2011 | Application of an efficient Bayesian discretization method to biomedical dataabstractBACKGROUND: Several data mining methods require data that are discrete, and other methods often perform better with discrete data. We introduce an efficient Bayesian discretization (EBD) method for optimal discretization of variables that runs efficiently on high-dimensional biomedical datasets. The EBD method consists of two components, namely, a Bayesian score to evaluate discretizations and a dynamic programming search procedure to efficiently search the space of possible discretizations. We compared the performance of EBD to Fayyad and Irani's (FI) discretization method, which is commonly used for discretization. RESULTS: On 24 biomedical datasets obtained from high-throughput transcriptomic and proteomic studies, the classification performances of the C4.5 classifier and the naïve Bayes classifier were statistically significantly better when the predictor variables were discretized using EBD over FI. EBD was statistically significantly more stable to the variability of the datasets than FI. However, EBD was less robust, though not statistically significantly so, than FI and produced slightly more complex discretizations than FI. CONCLUSIONS: On a range of biomedical datasets, a Bayesian discretization method (EBD) yielded better classification performance and stability but was less robust than the widely used FI discretization method. The EBD discretization method is easy to implement, permits the incorporation of prior knowledge and belief, and is sufficiently fast for application to high-dimensional data. Jonathan L. Lustgarten, Shyam Visweswaran, Vanathi Gopalakrishnan, Gregory F. Cooper |
BMC Bioinform. | 4 |
| 2011 | The application of naive Bayes model averaging to predict Alzheimer's disease from genome-wide dataabstractOBJECTIVE: Predicting patient outcomes from genome-wide measurements holds significant promise for improving clinical care. The large number of measurements (eg, single nucleotide polymorphisms (SNPs)), however, makes this task computationally challenging. This paper evaluates the performance of an algorithm that predicts patient outcomes from genome-wide data by efficiently model averaging over an exponential number of naive Bayes (NB) models. DESIGN: This model-averaged naive Bayes (MANB) method was applied to predict late onset Alzheimer's disease in 1411 individuals who each had 312,318 SNP measurements available as genome-wide predictive features. Its performance was compared to that of a naive Bayes algorithm without feature selection (NB) and with feature selection (FSNB). MEASUREMENT: Performance of each algorithm was measured in terms of area under the ROC curve (AUC), calibration, and run time. RESULTS: The training time of MANB (16.1 s) was fast like NB (15.6 s), while FSNB (1684.2 s) was considerably slower. Each of the three algorithms required less than 0.1 s to predict the outcome of a test case. MANB had an AUC of 0.72, which is significantly better than the AUC of 0.59 by NB (p<0.00001), but not significantly different from the AUC of 0.71 by FSNB. MANB was better calibrated than NB, and FSNB was even better in calibration. A limitation was that only one dataset and two comparison algorithms were included in this study. CONCLUSION: MANB performed comparatively well in predicting a clinical outcome from a high-dimensional genome-wide dataset. These results provide support for including MANB in the methods used to predict outcomes from large, genome-wide datasets. Shyam Visweswaran, Gregory F. Cooper |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Bayesian rule learning for biomedical data miningabstractMOTIVATION: Disease state prediction from biomarker profiling studies is an important problem because more accurate classification models will potentially lead to the discovery of better, more discriminative markers. Data mining methods are routinely applied to such analyses of biomedical datasets generated from high-throughput 'omic' technologies applied to clinical samples from tissues or bodily fluids. Past work has demonstrated that rule models can be successfully applied to this problem, since they can produce understandable models that facilitate review of discriminative biomarkers by biomedical scientists. While many rule-based methods produce rules that make predictions under uncertainty, they typically do not quantify the uncertainty in the validity of the rule itself. This article describes an approach that uses a Bayesian score to evaluate rule models. RESULTS: We have combined the expressiveness of rules with the mathematical rigor of Bayesian networks (BNs) to develop and evaluate a Bayesian rule learning (BRL) system. This system utilizes a novel variant of the K2 algorithm for building BNs from the training data to provide probabilistic scores for IF-antecedent-THEN-consequent rules using heuristic best-first search. We then apply rule-based inference to evaluate the learned models during 10-fold cross-validation performed two times. The BRL system is evaluated on 24 published 'omic' datasets, and on average it performs on par or better than other readily available rule learning methods. Moreover, BRL produces models that contain on average 70% fewer variables, which means that the biomarker panels for disease prediction contain fewer markers for further verification and validation by bench scientists. Vanathi Gopalakrishnan, Jonathan L. Lustgarten, Shyam Visweswaran, Gregory F. Cooper |
Bioinform. | 4 |
| 2010 | A real-time temporal Bayesian architecture for event surveillance and its application to patient-specific multiple disease outbreak detection
Xia Jiang, Gregory F. Cooper |
Data Min. Knowl. Discov. | 2 |
| 2010 | A Bayesian network model for spatial event surveillance
Xia Jiang, Daniel B. Neill, Gregory F. Cooper |
Int. J. Approx. Reason. | 3 |
| 2010 | A Bayesian spatio-temporal method for disease outbreak detectionabstractA system that monitors a region for a disease outbreak is called a disease outbreak surveillance system. A spatial surveillance system searches for patterns of disease outbreak in spatial subregions of the monitored region. A temporal surveillance system looks for emerging patterns of outbreak disease by analyzing how patterns have changed during recent periods of time. If a non-spatial, non-temporal system could be converted to a spatio-temporal one, the performance of the system might be improved in terms of early detection, accuracy, and reliability. A Bayesian network framework is proposed for a class of space-time surveillance systems called BNST. The framework is applied to a non-spatial, non-temporal disease outbreak detection system called PC in order to create the spatio-temporal system called PCTS. Differences in the detection performance of PC and PCTS are examined. The results show that the spatio-temporal Bayesian approach performs well, relative to the non-spatial, non-temporal approach. Xia Jiang, Gregory F. Cooper |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Learning patient-specific predictive models from clinical data
Shyam Visweswaran, Derek C. Angus, Margaret Hsieh, Lisa A. Weissfeld, Donald Yealy, Gregory F. Cooper |
J. Biomed. Informatics | 6 |
| 2010 | Learning Instance-Specific Predictive Models
Shyam Visweswaran, Gregory F. Cooper |
J. Mach. Learn. Res. | 2 |
| 2010 | A multivariate Bayesian scan statistic for early event detection and characterization
Daniel B. Neill, Gregory F. Cooper |
Mach. Learn. | 2 |
| 2009 | Generalized AMOC Curves For Evaluation and Improvement of Event Surveillance
Xia Jiang, Gregory F. Cooper, Daniel B. Neill |
AMIA | 2 |
| 2009 | Bayesian Modeling of Unknown Diseases for Biosurveillance
Yanna Shen, Gregory F. Cooper |
AMIA | 2 |
| 2009 | Bayesian prediction of an epidemic curve
Xia Jiang, Garrick L. Wallstrom, Gregory F. Cooper, Michael M. Wagner 0001 |
J. Biomed. Informatics | 3 |
| 2008 | Analysis of a Failed Clinical Decision Support System for Management of Congestive Heart Failure
Rajiv Wadhwa, Douglas B. Fridsma, Melissa I. Saul, Louis E. Penrod, Shyam Visweswaran, Gregory F. Cooper, Wendy W. Chapman |
AMIA | 6 |
| 2008 | Hierarchical explanation of inference in Bayesian networks that represent a population of independent agentsabstractThis paper describes a novel method for explaining Bayesian network (BN) inference when the network is modeling a population of conditionally independent agents, each of which is modeled as a subnetwork. For example, consider disease-outbreak detection, in which the agents are patients who are modeled as independent, conditioned on the factors that cause disease spread. Given evidence about these patients, such as their symptoms, suppose that the BN system infers that a respiratory anthrax outbreak is highly likely. A public-health official who received such a report would generally want to know why anthrax is being given a high posterior probability. This paper describes the design of a system that explains such inferences. The explanation approach is applicable in general to inference in BNs that model conditionally independent agents; it complements previous approaches for explaining inference on BNs that model a single agent (e.g., explaining the diagnostic inference for a single patient using a BN that models just that patient). Peter Sutovskú, Gregory F. Cooper |
ECAI | 2 |
| 2008 | Evaluation of preprocessing techniques for chief complaint classification
Jagan Dara, John N. Dowling, Debbie A. Travers, Gregory F. Cooper, Wendy W. Chapman |
J. Biomed. Informatics | 4 |
| 2008 | Estimating the joint disease outbreak-detection time when an automated biosurveillance system is augmenting traditional clinical case finding
Yanna Shen, Christina Adamou, John N. Dowling, Gregory F. Cooper |
J. Biomed. Informatics | 4 |
| 2007 | Evidence-based Anomaly Detection in Clinical Domains
Milos Hauskrecht, Michal Valko, Branislav Kveton, Shyam Visweswaran, Gregory F. Cooper |
AMIA | 5 |
| 2007 | A Recursive Algorithm for Spatial Cluster Detection
Xia Jiang, Gregory F. Cooper |
AMIA | 2 |
| 2006 | A Theoretical Study of Y Structures for Causal Discovery
Subramani Mani, Gregory F. Cooper, Peter Spirtes |
UAI | 2 |
| 2006 | A control study to evaluate a computer-based microarray experiment design recommendation system for gene-regulation pathways discovery
Changwon Yoo, Gregory F. Cooper |
J. Biomed. Informatics | 2 |
| 2005 | Deriving the Expected Utility of a Predictive Model When the Utilities Are Uncertain
Gregory F. Cooper, Shyam Visweswaran |
AMIA | 1 |
| 2005 | A Software Tool to Assist Researchers in Coding Free Text Clinical Reports
Manoj Ramachandran, Gregory F. Cooper, Wendy W. Chapman, John N. Dowling |
AMIA | 2 |
| 2005 | Patient-Specific Models for Predicting the Outcomes of Patients with Community Acquired Pneumonia
Shyam Visweswaran, Gregory F. Cooper |
AMIA | 2 |
| 2005 | A Bayesian Spatial Scan StatisticabstractWe propose a new Bayesian method for spatial cluster detection, the “Bayesian spatial scan statistic,” and compare this method to the standard (frequentist) scan statistic approach. We demonstrate that the Bayesian statistic has several advantages over the frequentist approach, including increased power to detect clusters and (since randomization testing is unnecessary) much faster runtime. We evaluate the Bayesian and fre- quentist methods on the task of prospective disease surveillance: detect- ing spatial clusters of disease cases resulting from emerging disease out- breaks. We demonstrate that our Bayesian methods are successful in rapidly detecting outbreaks while keeping number of false positives low. Daniel B. Neill, Andrew W. Moore 0001, Gregory F. Cooper |
NIPS | 3 |
| 2005 | Editorial Comments: Defining a Workable Strategy to Stimulate Widespread Adoption of Electronic Health Records in the United StatesabstractThis issue of JAMIA contains three articles that are based on discussions held at the 2004 Annual Symposium of the American Medical Informatics Association's College of Informatics. The College's agenda for the 2004 Symposium was to discuss and define a workable strategy to stimulate widespread adoption of electronic health records (EHRs) in the United States. The remainder of this introduction provides background information that sets the context for reading the three associated papers in this issue. In September 2003, a general call was issued to the College membership for topics for the 2004 Symposium. On an e-mail discussion list, Dr. Ed Hammond suggested that the Symposium discuss and develop concrete ideas for how to stimulate the adoption of EHRs in the U.S. His suggestion generated dozens of e-mail messages that explored many facets of the issue. It was clear that the topic was of great interest. Dr. Hammond's suggestion for a symposium theme was accepted by the College's Scientific Program Committee, which consisted of Drs. Suzanne Bakken, James Brinkley, Gregory Cooper (Chair), Lawrence Fagan, Antoine Geissbuhler, Isaac Kohane, Peter Haug, Alexa McCray, Blackford Middleton, and Frank Sonnenberg. The Committee proceeded to organize the content and format of the Symposium. Particularly helpful in shaping the agenda for the meeting were Drs. Blackford Middleton and Randolph Miller and Mr. Jeff Williamson, who arranged logistical support. The Symposium sessions took place on February 13–15, 2004, at the Paradise Point Resort in San Diego, CA. Approximately 50 College of Informatics members participated in the sessions. Three topics about the EHR were discussed on each of three days: (1) Where are we and how did we get here? (Session Chairs: Drs. Don Detmer and Don Simborg), (2) What are the factors and forces affecting EHR system adoption? (Session Chairs: Drs. Joan Ash and David Bates), and (3) How can the College of Informatics best contribute to the widespread adoption of EHRs in the United States? (Session Chairs: Drs. Patricia Brennan, Ed Hammond, and Blackford Middleton). The three papers in this issue parallel the three session topics at the Symposium. Although these papers are based on discussions that occurred at the Symposium, they do not represent an official (or unofficial) position of the College, in that no formal vote of the membership approved their content. In addition, the papers could not present all the ideas that were discussed at the Symposium; conversely, they mention some ideas that were inspired by the Symposium but not specifically discussed there. The first paper, by Berner, Detmer, and Simborg, provides a succinct overview of the history of adoption of EHRs in the United States from the 1960s to the present. The paper notes that prospects seem good for widespread use of EHRs in the United States and explores reasons why a basis exists for optimism today more than in the past. Nonetheless, it cautions that considerable effort will be required to overcome resistance to making such a fundamental change in the health care delivery process. The second paper by Ash and Bates describes some key factors influencing EHR system adoption in the United States, including environmental, organizational, personal, and technical factors. It summarizes relevant survey results, particularly regarding attitudes toward and adoption of computerized physician order entry. The paper contains suggestions for greater education, communication, alignment of incentives, and standardization. Some of these suggestions are discussed in additional detail in the third paper. The third paper, by Middleton, Hammond, Brennan, and Cooper, discusses the Symposium theme of defining a workable strategy to stimulate widespread adoption of EHRs in the United States. It posits that EHR adoption is stymied by a fundamental market failure of health care information technology in the United States. Building on issues presented in the first and second papers, this paper elaborates on reasons for failure, including misaligned incentives among key market players, the lack of both broad standards adoption and definitions of basic product features, and the rapid-cycle turnover of health information technology companies. The paper proposes four areas in which advances are likely to facilitate United States EHR adoption: (1) providing financial incentives to the EHR marketplace, (2) setting and adopting EHR functional and related informatics standards, (3) creating policies for EHR adoption, and (4) engaging in educational, marketing, and supporting activities. Each of these areas is discussed in detail. We are at an exciting point in the evolution of the EHR in the United States. Circumstances are aligning in a way that makes the widespread adoption of EHRs more likely than ever before. The three articles in this issue of JAMIA provide helpful background and suggestions for achieving that long-awaited goal. Members of the American Medical Informatics Association, its College of Informatics, and many others are working to make the goal a reality. Although much work remains to be done, it seems clear that the benefits of success are likely to be immense. Gregory F. Cooper |
J. Am. Medical Informatics Assoc. | 1 |
| 2005 | Viewpoint Paper: Accelerating U.S. EHR Adoption: How to Get There From Here. Recommendations Based on the 2004 ACMI RetreatabstractDespite growing support for the adoption of electronic health records (EHR) to improve U.S. healthcare delivery, EHR adoption in the United States is slow to date due to a fundamental failure of the healthcare information technology marketplace. Reasons for the slow adoption of healthcare information technology include a misalignment of incentives, limited purchasing power among providers, variability in the viability of EHR products and companies, and limited demonstrated value of EHRs in practice. At the 2004 American College of Medical Informatics (ACMI) Retreat, attendees discussed the current state of EHR adoption in this country and identified steps that could be taken to stimulate adoption. In this paper, based upon the ACMI retreat, and building upon the experiences of the authors developing EHR in academic and commercial settings we identify a set of recommendations to stimulate adoption of EHR, including financial incentives, promotion of EHR standards, enabling policy, and educational, marketing, and supporting activities for both the provider community and healthcare consumers. Blackford Middleton, William Edward Hammond, Patricia Flatley Brennan, Gregory F. Cooper |
J. Am. Medical Informatics Assoc. | 4 |
| 2005 | Predicting dire outcomes of patients with community acquired pneumonia
Gregory F. Cooper, Vijoy Abraham, Constantin F. Aliferis, John M. Aronis, Bruce G. Buchanan, Rich Caruana, Michael J. Fine, Janine E. Janosky, Gary Livingston, Tom M. Mitchell |
J. Biomed. Informatics | 1 |
| 2005 | What's Strange About Recent Events (WSARE): An Algorithm for the Early Detection of Disease OutbreaksabstractTraditional biosurveillance algorithms detect disease outbreaks by looking for peaks in a univariate time series of health-care data. Current health-care surveillance data, however, are no longer simply univariate data streams. Instead, a wealth of spatial, temporal, demographic and symptomatic information is available. We present an early disease outbreak detection algorithm called What's Strange About Recent Events (WSARE), which uses a multivariate approach to improve its timeliness of detection. WSARE employs a rule-based technique that compares recent health-care data against data from a baseline distribution and finds subgroups of the recent data whose proportions have changed the most from the baseline data. In addition, health-care data also pose difficulties for surveillance algorithms because of inherent temporal trends such as seasonal effects and day of week variations. WSARE approaches this problem using a Bayesian network to produce a baseline distribution that accounts for these temporal trends. The algorithm itself incorporates a wide range of ideas, including association rules, Bayesian networks, hypothesis testing and permutation tests to produce a detection algorithm that is careful to evaluate the significance of the alarms that it raises. Weng-Keen Wong, Andrew W. Moore 0001, Gregory F. Cooper, Michael M. Wagner 0001 |
J. Mach. Learn. Res. | 3 |
| 2004 | Instance-Specific Bayesian Model Averaging for ClassificationabstractClassification algorithms typically induce population-wide models that are trained to perform well on average on expected future instances. We introduce a Bayesian framework for learning instance-specific models from data that are optimized to predict well for a particular instance. Based on this framework, we present a that performs selective model averaging over a restricted class of Bayesian networks. On experimental evaluation, this algorithm shows superior performance over model selection. We intend to apply such instance-specific algorithms to improve the performance of patient-specific predictive models induced from medical data. instance-specific algorithm called ISA Shyam Visweswaran, Gregory F. Cooper |
NIPS | 2 |
| 2004 | Bayesian Biosurveillance of Disease Outbreaks
Gregory F. Cooper, Denver Dash, John D. Levander, Weng-Keen Wong, William R. Hogan, Michael M. Wagner 0001 |
UAI | 1 |
| 2004 | An evaluation of a system that recommends microarray experiments to perform to discover gene-regulation pathways
Changwon Yoo, Gregory F. Cooper |
Artif. Intell. Medicine | 2 |
| 2004 | Model Averaging for Prediction with Discrete Bayesian Networks
Denver Dash, Gregory F. Cooper |
J. Mach. Learn. Res. | 2 |
| 2003 | Detecting Adverse Drug Events in Discharge Summaries Using Variations on the Simple Bayes Model
Shyam Visweswaran, Paul Hanbury, Melissa I. Saul, Gregory F. Cooper |
AMIA | 4 |
| 2003 | A Computer-Based Microarray Experiment Design-System for Gene-Regulation Pathway Discovery
Changwon Yoo, Gregory F. Cooper |
AMIA | 2 |
| 2003 | Bayesian Network Anomaly Pattern Detection for Disease Outbreaks
Weng-Keen Wong, Andrew W. Moore 0001, Gregory F. Cooper, Michael M. Wagner 0001 |
ICML | 3 |
| 2003 | Research Paper: Creating a Text Classifier to Detect Radiology Reports Describing Mediastinal Findings Associated with Inhalational Anthrax and Other DisordersabstractOBJECTIVE: The aim of this study was to create a classifier for automatic detection of chest radiograph reports consistent with the mediastinal findings of inhalational anthrax. DESIGN: The authors used the Identify Patient Sets (IPS) system to create a key word classifier for detecting reports describing mediastinal findings consistent with anthrax and compared their performances on a test set of 79,032 chest radiograph reports. MEASUREMENTS: Area under the ROC curve was the main outcome measure of the IPS classifier. Sensitivity and specificity of an initial IPS model were calculated based on an existing key word search and were compared against a Boolean version of the IPS classifier. RESULTS: The IPS classifier received an area under the ROC curve of 0.677 (90% CI = 0.628 to 0.772) with a specificity of 0.99 and maximum sensitivity of 0.35. The initial IPS model attained a specificity of 1.0 and a sensitivity of 0.04. CONCLUSION: The IPS system is a useful tool for helping domain experts create a statistical key word classifier for textual reports that is a potentially useful component in surveillance of radiographic findings suspicious for anthrax. Wendy W. Chapman, Gregory F. Cooper, Paul Hanbury, Brian E. Chapman, Lee H. Harrison, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2002 | Creating a Software Tool for the Clinical Researcher - the IPS System
Bruce G. Buchanan, Wendy W. Chapman, Gregory F. Cooper, Paul Hanbury, Mehmet Kayaalp 0002, Manoj Ramachandran, Melissa I. Saul |
AMIA | 3 |
| 2002 | Discovery of gene-regulation pathways using local causal search
Changwon Yoo, Gregory F. Cooper |
AMIA | 2 |
| 2002 | Exact model averaging with naive Bayesian classifiers
Denver Dash, Gregory F. Cooper |
ICML | 2 |
| 2002 | A Bayesian Network Scoring Metic that Is Based on Globally Uniform Parameter Priors
Mehmet Kayaalp 0002, Gregory F. Cooper |
UAI | 2 |
| 2001 | Evaluation of negation phrases in narrative clinical reports
Wendy W. Chapman, Will Bridewell, Paul Hanbury, Gregory F. Cooper, Bruce G. Buchanan |
AMIA | 4 |
| 2001 | IPS: A System That Uses Machine Learning to Help Locate Patient Records for Clinical Research
Gregory F. Cooper, Bruce G. Buchanan, Wendy W. Chapman, Paul Hanbury, Mehmet Kayaalp 0002, Melissa I. Saul |
AMIA | 1 |
| 2001 | A Simple Algorithm for Identifying Negated Findings and Diseases in Discharge Summaries
Wendy W. Chapman, Will Bridewell, Paul Hanbury, Gregory F. Cooper, Bruce G. Buchanan |
J. Biomed. Informatics | 4 |
| 2000 | A Retrospective Study of the Effect of a Prognostic Model on the Hospital Admission Decision for Patients with Low Risk Community-acquired Pneumonia
John M. Aronis, Gregory F. Cooper, Michael J. Fine |
AMIA | 2 |
| 2000 | Predicting ICU mortality: a comparison of stationary and nonstationary temporal models
Mehmet Kayaalp 0002, Gregory F. Cooper, Gilles Clermont |
AMIA | 2 |
| 2000 | Causal discovery from medical textual data
Subramani Mani, Gregory F. Cooper |
AMIA | 2 |
| 2000 | A Bayesian Method for Causal Modeling and Discovery Under Selection
Gregory F. Cooper |
UAI | 1 |
| 1999 | Identifying patient subgroups with simple Bayes'
John M. Aronis, Gregory F. Cooper, Mehmet Kayaalp 0002, Bruce G. Buchanan |
AMIA | 2 |
| 1999 | A study in causal discovery from population-based infant birth and death records
Subramani Mani, Gregory F. Cooper |
AMIA | 2 |
| 1999 | Causal Discovery from a Mixture of Experimental and Observational Data
Gregory F. Cooper, Changwon Yoo |
UAI | 1 |
| 1999 | A Bayesian Network Classifier that Combines a Finite Mixture Model and a NaIve Bayes Model
Stefano Monti, Gregory F. Cooper |
UAI | 2 |
| 1998 | Temporal representation design principles: an assessment in the domain of liver transplantation
Constantin F. Aliferis, Gregory F. Cooper |
AMIA | 2 |
| 1998 | Using computer modeling to help identify patient subgroups in clinical data repositories
Gregory F. Cooper, Bruce G. Buchanan, Mehmet Kayaalp 0002, Melissa I. Saul, John K. Vries |
AMIA | 1 |
| 1998 | The impact of modeling the dependencies among patient findings on classification accuracy and calibration
Stefano Monti, Gregory F. Cooper |
AMIA | 2 |
| 1998 | A Multivariate Discretization Method for Learning Bayesian Networks from Mixed Data
Stefano Monti, Gregory F. Cooper |
UAI | 2 |
| 1998 | Research Paper: An Experiment Comparing Lexical and Statistical Methods for Extracting MeSH Terms from Clinical Free TextabstractOBJECTIVE: A primary goal of the University of Pittsburgh's 1990-94 UMLS-sponsored effort was to develop and evaluate PostDoc (a lexical indexing system) and Pindex (a statistical indexing system) comparatively, and then in combination as a hybrid system. Each system takes as input a portion of the free text from a narrative part of a patient's electronic medical record and returns a list of suggested MeSH terms to use in formulating a Medline search that includes concepts in the text. This paper describes the systems and reports an evaluation. The intent is for this evaluation to serve as a step toward the eventual realization of systems that assist healthcare personnel in using the electronic medical record to construct patient-specific searches of Medline. DESIGN: The authors tested the performances of PostDoc, Pindex, and a hybrid system, using text taken from randomly selected clinical records, which were stratified to include six radiology reports, six pathology reports, and six discharge summaries. They identified concepts in the clinical records that might conceivably be used in performing a patient-specific Medline search. Each system was given the free text of each record as an input. The extent to which a system-derived list of MeSH terms captured the relevant concepts in these documents was determined based on blinded assessments by the authors. RESULTS: PostDoc output a mean of approximately 19 MeSH terms per report, which included about 40% of the relevant report concepts. Pindex output a mean of approximately 57 terms per report and captured about 45% of the relevant report concepts. A hybrid system captured approximately 66% of the relevant concepts and output about 71 terms per report. CONCLUSION: The outputs of PostDoc and Pindex are complementary in capturing MeSH terms from clinical free text. The results suggest possible approaches to reduce the number of terms output while maintaining the percentage of terms captured, including the use of UMLS semantic types to constrain the output list to contain only clinically relevant MeSH terms. Gregory F. Cooper, Randolph A. Miller |
J. Am. Medical Informatics Assoc. | 1 |
| 1997 | INKBLOT: A neurological diagnostic decision support system integrating causal and anatomical knowledge
Gil Citro, Gordon Banks, Gregory F. Cooper |
Artif. Intell. Medicine | 3 |
| 1997 | An evaluation of machine-learning methods for predicting pneumonia mortality
Gregory F. Cooper, Constantin F. Aliferis, Richard Ambrosino, John M. Aronis, Bruce G. Buchanan, Rich Caruana, Michael J. Fine, Clark Glymour, Geoffrey J. Gordon, Barbara H. Hanusa, Janine E. Janosky, Christopher Meek, Tom M. Mitchell, Thomas Richardson 0001, Peter Spirtes |
Artif. Intell. Medicine | 1 |
| 1997 | A Simple Constraint-Based Algorithm for Efficiently Mining Observational Databases for Causal Relationships
Gregory F. Cooper |
Data Min. Knowl. Discov. | 1 |
| 1996 | Learning Bayesian Belief Networks with Neural Network Estimators
Stefano Monti, Gregory F. Cooper |
NIPS | 2 |
| 1996 | A Structurally and Temporally Extended Bayesian Belief Network Model: Definitions, Properties, and Modeling Techniques
Constantin F. Aliferis, Gregory F. Cooper |
UAI | 2 |
| 1996 | Bounded recursive decomposition: a search-based method for belief-network inference under limited resources
Stefano Monti, Gregory F. Cooper |
Int. J. Approx. Reason. | 2 |
| 1996 | Research Paper: A Temporal Analysis of QMRabstractOBJECTIVE: To understand better the trade-offs of not incorporating explicit time in Quick Medical Reference (QMR), a diagnostic system in the domain of general internal medicine, along the dimensions of expressive power and diagnostic accuracy. DESIGN: The study was conducted in two phases. Phase I was a descriptive analysis of the temporal abstractions incorporated in QMR's terms. Phase II was a pseudo-prospective controlled experiment, measuring the effect of history and physical examination temporal content on the diagnostic accuracy of QMR. MEASUREMENTS: For each QMR finding that would fit our operational definition of temporal finding, several parameters describing the temporal nature of the finding were assessed, the most important ones being: temporal primitives, time units, temporal uncertainty, processes, and patterns. The history, physical examination, and initial laboratory results of 105 consecutive patients admitted to the Pittsburgh University Presbyterian Hospital were analyzed for temporal content and factors that could potentially influence diagnostic accuracy (these included: rareness of primary diagnosis, case length, uncertainty, spatial/causal information, and multiple diseases). RESULTS: 776 findings were identified as temporal. The authors developed an ontology describing the terms utilized by QMR developers to express temporal knowledge. The authors classified the temporal abstractions found in QMR in 116 temporal types, 11 temporal templates, and a temporal hierarchy. The odds of QMR's making a correct diagnosis in high temporal complexity cases is 0.7 the odds when the temporal complexity is lower, but this result is not statistically significant (95% confidence interval = 0.27-1.83). CONCLUSIONS: QMR contains extensive implicit time modeling. These results support the conclusion that the abstracted encoding of time in the medical knowledge of QMR does not induce a diagnostic performance penalty. Constantin F. Aliferis, Gregory F. Cooper, Randolph A. Miller, Bruce G. Buchanan, Richard Bankowitz, Nunzia Bettinsoli Giuse |
J. Am. Medical Informatics Assoc. | 2 |
| 1995 | A Bayesian Method for Learning Belief Networks that Contain Hidden Variables
Gregory F. Cooper |
J. Intell. Inf. Syst. | 1 |
| 1994 | An Evaluation of an Algorithm for Inductive Learning of Bayesian Belief Networks Using Simulated Data Sets
Constantin F. Aliferis, Gregory F. Cooper |
UAI | 2 |
| 1993 | Probabilistic and decision-theoretic systems in medicine
Gregory F. Cooper |
Artif. Intell. Medicine | 1 |
| 1992 | A Bayesian Method for the Induction of Probabilistic Networks from Data
Gregory F. Cooper, Edward Herskovits |
Mach. Learn. | 1 |
| 1991 | A Bayesian Method for Constructing Bayesian Belief Networks from Databases
Gregory F. Cooper, Edward Herskovits |
UAI | 1 |
| 1991 | Initialization for the Method of Conditioning in Bayesian Belief Networks
Henri Jacques Suermondt, Gregory F. Cooper |
Artif. Intell. | 2 |
| 1991 | A combination of exact algorithms for inference on Bayesian belief networks
Henri Jacques Suermondt, Gregory F. Cooper |
Int. J. Approx. Reason. | 2 |
| 1990 | an entropy-driven system for construction of probabilistic expert systems from databases
Edward Herskovits, Gregory F. Cooper |
UAI | 2 |
| 1990 | A combination of cutset conditioning with clique-tree propagation in the Pathfinder system
Henri Jacques Suermondt, Gregory F. Cooper, David Heckerman |
UAI | 2 |
| 1990 | The Computational Complexity of Probabilistic Inference Using Bayesian Belief NetworksabstractBayesian belief networks provide a natural, efficient method for representing probabilistic dependencies among a set of variables. For these reasons, numerous researchers are exploring the use of belief networks as a knowledge representation in artificial intelligence. Algorithms have been developed previously for efficient probabilistic inference using special classes of belief networks. More general classes of belief networks, however, have eluded efforts to develop efficient inference algorithms. We show that probabilistic inference using belief networks is NP-hard. Therefore, it seems unlikely that an exact algorithm can be developed to perform probabilistic inference efficiently over all classes of belief networks. This result suggests that research should be directed away from the search for a general, efficient probabilistic inference algorithm, and toward the design of efficient special-case, average-case, and approximation algorithms. Gregory F. Cooper |
Artif. Intell. | 1 |
| 1990 | The 1990 AAAI Spring Symposium on Artificial Intelligence in Medicine
Gregory F. Cooper, Mark A. Musen |
Artif. Intell. Medicine | 1 |
| 1990 | Probabilistic inference in multiply connected belief networks using loop cutsets
Henri Jacques Suermondt, Gregory F. Cooper |
Int. J. Approx. Reason. | 2 |
| 1990 | A randomized approximation algorithm for probabilistic inference on bayesian belief networksabstractAbstract Researchers in decision analysis and artificial intelligence (AI) have used Bayesian belief networks to build probabilistic expert systems. Using standard methods drawn from the theory of computational complexity, workers in the field have shown that the problem of probabilistic inference in belief networks is difficult and almost certainly intractable. We have developed a randomized approximation scheme, BN‐RAS, for doing probabilistic inference in belief networks. The algorithm can, in many circumstances, perform efficient approximate inference in large and richly interconnected models. Unlike previously described stochastic algorithms for probabilistic inference, the randomized approximation scheme (ras) computes a priori bounds on running time by analyzing the structure and contents of the belief network. In this article, we describe BN‐RAS precisely and analyze its performance mathematically. R. Martin Chavez, Gregory F. Cooper |
Networks | 2 |
| 1989 | The ALARM Monitoring System: A Case Study with two Probabilistic Inference Techniques for Belief Networks
Ingo A. Beinlich, Henri Jacques Suermondt, R. Martin Chavez, Gregory F. Cooper |
AIME | 4 |
| 1989 | Reflection and Action Under Scarce Resources: Theoretical Principles and Empirical Study
Eric Horvitz, Gregory F. Cooper, David Heckerman |
IJCAI | 2 |
| 1989 | An Empirical Evaluation of a Randomized Algorithm for Probabilistic Inference
R. Martin Chavez, Gregory F. Cooper |
UAI | 2 |
| 1988 | KNET: integrating hypermedia and normative bayesian modeling
R. Martin Chavez, Gregory F. Cooper |
UAI | 2 |
| 1988 | Stochastic simulation of Bayesian belief networks
Homer L. Chin, Gregory F. Cooper |
Int. J. Approx. Reason. | 2 |
| 1988 | An algorithm for computing probabilistic propositions
Gregory F. Cooper |
Int. J. Approx. Reason. | 1 |
| 1987 | Bayesian Belief Network Inference Using Simulation
Homer L. Chin, Gregory F. Cooper |
UAI | 2 |
| 1987 | An Algorithm for Computing Probabilistic Propositions
Gregory F. Cooper |
UAI | 1 |