David Page

dblp:p/DavidPage · also C. David Page, C. David Page Jr. · DBLP profile ↗
← Back
100ranked-venue papers
7as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 4 since 2021Databases, data management, data science and information retrieval · 17Theory of computation · 12 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-authorSecurity and privacy · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 MPAC: a computational framework for inferring pathway activities from multi-omic data
abstract
MOTIVATION: Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. RESULTS: We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g. associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell composition. Our MPAC R package enables similar multi-omic analyses on new datasets. AVAILABILITY AND IMPLEMENTATION: The MPAC package is available at Bioconductor https://bioconductor.org/packages/MPAC.
David Page, Paul Ahlquist, Irene M. Ong, Anthony Gitter
Bioinform.2
2024 On Neural Networks as Infinite Tree-Structured Probabilistic Graphical Models
abstract
Deep neural networks (DNNs) lack the precise semantics and definitive probabilistic interpretation of probabilistic graphical models (PGMs). In this paper, we propose an innovative solution by constructing infinite tree-structured PGMs that correspond exactly to neural networks. Our research reveals that DNNs, during forward propagation, indeed perform approximations of PGM inference that are precise in this alternative PGM structure. Not only does our research complement existing studies that describe neural networks as kernel machines or infinite-sized Gaussian processes, it also elucidates a more direct approximation that DNNs make to exact inference in PGMs. Potential benefits include improved pedagogy and interpretation of DNNs, and algorithms that can merge the strengths of PGMs and DNNs.
Boyao Li, Alexander Thomson, Houssam Nassif, Matthew Engelhard, David Page
NeurIPS5
2024 Soft phenotyping for sepsis via EHR time-aware soft clustering
Shiyi Jiang, Xin Gai, Miriam M. Treggiari, William W. Stead, Yuankang Zhao, David Page, Anru Zhang
J. Biomed. Informatics6
2023 Variable importance matching for causal inference
abstract
Our goal is to produce methods for observational causal inference that are auditable, easy to troubleshoot, yield accurate treatment effect estimates, and scalable to high-dimensional data. We describe a general framework called Model-to-Match that achieves these goals by (i) learning a distance metric via outcome modeling, (ii) creating matched groups using the distance metric, and (iii) using the matched groups to estimate treatment effects. Model-to-Match uses variable importance measurements to construct a distance metric, making it a flexible framework that can be adapted to various applications. Concentrating on the scalability of the problem in the number of potential confounders, we operationalize the Model-to-Match framework with LASSO. We derive performance guarantees for settings where LASSO outcome modeling consistently identifies all confounders (importantly without requiring the linear model to be correctly specified). We also provide experimental results demonstrating the auditability of matches, as well as extensions to more general nonparametric outcome modeling.
Quinn Lanners, Harsh Parikh, Alexander Volfovsky, Cynthia Rudin, David Page
UAI5
2021 Predicting Drug-Drug Interactions from Heterogeneous Data: An Embedding Approach
Devendra Singh Dhami, Siwen Yan, Gautam Kunapuli, David Page, Sriraam Natarajan
AIME4
2021 E-Pedigrees: a large-scale automatic family pedigree prediction application
abstract
MOTIVATION: The use and functionality of Electronic Health Records (EHR) have increased rapidly in the past few decades. EHRs are becoming an important depository of patient health information and can capture family data. Pedigree analysis is a longstanding and powerful approach that can gain insight into the underlying genetic and environmental factors in human health, but traditional approaches to identifying and recruiting families are low-throughput and labor-intensive. Therefore, high-throughput methods to automatically construct family pedigrees are needed. RESULTS: We developed a stand-alone application: Electronic Pedigrees, or E-Pedigrees, which combines two validated family prediction algorithms into a single software package for high throughput pedigrees construction. The convenient platform considers patients' basic demographic information and/or emergency contact data to infer high-accuracy parent-child relationship. Importantly, E-Pedigrees allows users to layer in additional pedigree data when available and provides options for applying different logical rules to improve accuracy of inferred family relationships. This software is fast and easy to use, is compatible with different EHR data sources, and its output is a standard PED file appropriate for multiple downstream analyses. AVAILABILITY AND IMPLEMENTATION: The Python 3.3+ version E-Pedigrees application is freely available on: https://github.com/xiayuan-huang/E-pedigrees.
Xiayuan Huang, Nicholas P. Tatonetti, Katie Larow, Brooke Delgoffe, John Mayer, David Page, Scott J. Hebbring
Bioinform.6
2020 CAUSE: Learning Granger Causality from Event Sequences using Attribution Methods
abstract
We study the problem of learning Granger causality between event types from asynchronous, interdependent, multi-type event sequences. Existing work suffers from either limited model flexibility or poor model explainability and thus fails to uncover Granger causality across a wide variety of event sequences with diverse event interdependency. To address these weaknesses, we propose CAUSE (Causality from AttribUtions on Sequence of Events), a novel framework for the studied task. The key idea of CAUSE is to first implicitly capture the underlying event interdependency by fitting a neural point process, and then extract from the process a Granger causality statistic using an axiomatic attribution method. Across multiple datasets riddled with diverse event interdependency, we demonstrate that CAUSE achieves superior performance on correctly inferring the inter-type Granger causality over a range of state-of-the-art methods.
Wei Zhang 0058, Thomas Kobber Panum, Somesh Jha, Prasad Chalasani, David Page
ICML5
2020 AutoBlock: A Hands-off Blocking Framework for Entity Matching
abstract
Entity matching seeks to identify data records over one or multiple data sources that refer to the same real-world entity. Virtually every entity matching task on large datasets requires blocking, a step that reduces the number of record pairs to be matched. However, most of the traditional blocking methods are learning-free and key-based, and their successes are largely built on laborious human effort in cleaning data and designing blocking keys.
Wei Zhang 0058, Bunyamin Sisman, Xin Dong 0001, Christos Faloutsos, David Page
WSDM6
2019 AUCμ: A Performance Metric for Multi-Class Machine Learning Models
abstract
The area under the receiver operating characteristic curve (AUC) is arguably the most common metric in machine learning for assessing the quality of a two-class classification model. As the number and complexity of machine learning applications grows, so too does the need for measures that can gracefully extend to classification models trained for more than two classes. Prior work in this area has proven computationally intractable and/or inconsistent with known properties of AUC, and thus there is still a need for an improved multi-class efficacy metric. We provide in this work a multi-class extension of AUC that we call AUC{\textmu} that is derived from first principles of the binary class AUC. AUC{\textmu} has similar computational complexity to AUC and maintains the properties of AUC critical to its interpretation and use.
Ross Kleiman, David Page
ICML2
2019 Machine Learning to Predict Developmental Neurotoxicity with High-Throughput Data from 2D Bio-Engineered Tissues
abstract
animal studies, and assays of animal and human primary cell cultures, suffer from challenges related to time, cost, and applicability to human physiology. Prior work has demonstrated success employing machine learning to predict developmental neurotoxicity using gene expression data collected from human 3D tissue models exposed to various compounds. The 3D model is biologically similar to developing neural structures, but its complexity necessitates extensive expertise and effort to employ. By instead focusing solely on constructing an assay of developmental neurotoxicity, we propose that a simpler 2D tissue model may prove sufficient. We thus compare the accuracy of predictive models trained on data from a 2D tissue model with those trained on data from a 3D tissue model, and find the 2D model to be substantially more accurate. Furthermore, we find the 2D model to be more robust under stringent gene set selection, whereas the 3D model suffers substantial accuracy degradation. While both approaches have advantages and disadvantages, we propose that our described 2D approach could be a valuable tool for decision makers when prioritizing neurotoxicity screening.
Finn Kuusisto, Vítor Santos Costa, Zhonggang Hou, James A. Thomson, David Page, Ron M. Stewart
ICMLA5
2019 Machine learning for phenotyping opioid overdose events
Jonathan C. Badger, Eric LaRose, John Mayer, Fereshteh S. Bashiri, David Page, Peggy L. Peissig
J. Biomed. Informatics5
2018 Privacy-Preserving Ridge Regression with only Linearly-Homomorphic Encryption
Irene Giacomelli, Somesh Jha, Marc Joye, David Page, Kyonghwan Yoon
ACNS4
2018 Improving breast cancer risk prediction by using demographic risk factors, abnormality features on mammograms and genetic variants
Shara Feld, Kaitlin M. Woo, Roxana Alexandridis, Yirong Wu, Jie Liu 0006, Peggy L. Peissig, Adedayo A. Onitilo, Jennifer Cox, David Page, Elizabeth S. Burnside
AMIA9
2018 Use of Electronic Health Record to Predict Family Relationships for Phenome-wide Research
Xiayuan Huang, Robert C. Elston, John Mayer, Zhan Ye, David Page, Scott J. Hebbring
AMIA6
2018 Temporal Poisson Square Root Graphical Models
abstract
We propose temporal Poisson square root graphical models (TPSQRs), a generalization of Poisson square root graphical models (PSQRs) specifically designed for modeling longitudinal event data. By estimating the temporal relationships for all possible pairs of event types, TPSQRs can offer a holistic perspective about whether the occurrences of any given event type could excite or inhibit any other type. A TPSQR is learned by estimating a collection of interrelated PSQRs that share the same template parameterization. These PSQRs are estimated jointly in a pseudo-likelihood fashion, where Poisson pseudo-likelihood is used to approximate the original more computationally intensive pseudo-likelihood problem stemming from PSQRs. Theoretically, we demonstrate that under mild assumptions, the Poisson pseudolikelihood approximation is sparsistent for recovering the underlying PSQR. Empirically, we learn TPSQRs from a real-world large-scale electronic health record (EHR) with millions of drug prescription and condition diagnosis events, for adverse drug reaction (ADR) detection. Experimental results demonstrate that the learned TPSQRs can recover ADR signals from the EHR effectively and efficiently.
Sinong Geng, Zhaobin Kuang, Peggy L. Peissig, David Page
ICML4
2018 Recursive Feature Elimination by Sensitivity Testing
abstract
There is great interest in methods to improve human insight into trained non-linear models. Leading approaches include producing a ranking of the most relevant features, a non-trivial task for non-linear models. We show theoretically and empirically the benefit of a novel version of recursive feature elimination (RFE) as often used with SVMs; the key idea is a simple twist on the kinds of sensitivity testing employed in computational learning theory with membership queries (e.g., [1]). With membership queries, one can check whether changing the value of a feature in an example changes the label. In the real-world, we usually cannot get answers to such queries, so our approach instead makes these queries to a trained (imperfect) non-linear model. Because SVMs are widely used in bioinformatics, our empirical results use a real-world cancer genomics problem; because ground truth is not known for this task, we discuss the potential insights provided. We also evaluate on synthetic data where ground truth is known.
Nicholas Sean Escanilla, Lisa Hellerstein, Ross Kleiman, Zhaobin Kuang, James Shull, David Page
ICMLA6
2018 Stochastic Learning for Sparse Discrete Markov Random Fields with Controlled Gradient Approximation Error
Sinong Geng, Zhaobin Kuang, Jie Liu 0006, Stephen J. Wright 0001, David Page
UAI5
2018 Applying family analyses to electronic health records to facilitate genetic research
abstract
Motivation: Pedigree analysis is a longstanding and powerful approach to gain insight into the underlying genetic factors in human health, but identifying, recruiting and genotyping families can be difficult, time consuming and costly. Development of high throughput methods to identify families and foster downstream analyses are necessary. Results: This paper describes simple methods that allowed us to identify 173 368 family pedigrees with high probability using basic demographic data available in most electronic health records (EHRs). We further developed and validate a novel statistical method that uses EHR data to identify families more likely to have a major genetic component to their diseases risk. Lastly, we showed that incorporating EHR-linked family data into genetic association testing may provide added power for genetic mapping without additional recruitment or genotyping. The totality of these results suggests that EHR-linked families can enable classical genetic analyses in a high-throughput manner. Availability and implementation: Pseudocode is provided as supplementary information. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Xiayuan Huang, Robert C. Elston, Guilherme J. M. Rosa, John Mayer, Zhan Ye, Terrie E. Kitchner, Murray H. Brilliant, David Page, Scott J. Hebbring
Bioinform.8
2017 Identifying Parkinson's Patients: A Functional Gradient Boosting Approach
Devendra Singh Dhami, Ameet Soni, David Page, Sriraam Natarajan
AIME3
2017 bigNN: An open-source big data toolkit focused on biomedical sentence classification
abstract
Every single day, a massive amount of text data is generated by different medical data sources, such as scientific literature, medical web pages, health-related social media, clinical notes, and drug reviews. Processing this wealth of data is indeed a daunting task, and it forces us to adopt smart and scalable computational strategies, including machine intelligence, big data analytics, and distributed architecture. In this contribution, we designed and developed an open-source big data neural network toolkit, namely bigNN which tackles the problem of large-scale biomedical text classification in an efficient fashion, facilitating fast prototyping and reproducible text analytics researches. bigNN scales up a word2vec-based neural network model over Apache Spark 2.10 and Hadoop Distributed File System (HDFS) 2.7.3, allowing for more efficient big data sentence classification. The toolkit supports big data computing, and simplifies rapid application development in sentence analysis by allowing users to configure and examine different internal parameters of both Apache Spark and the neural network model. bigNN is fully documented, and it is publicly and freely available at https://github.com/bircatmcri/bigNN.
Ahmad P. Tafti, Ehsun Behravesh, Mehdi Assefi, Eric LaRose, Jonathan C. Badger, John Mayer, AnHai Doan, David Page, Peggy L. Peissig
IEEE BigData8
2017 Employing spaceborne multispectral stereo pairs and pedestrian flow modeling to support disaster response activities in urban environments
abstract
Pedestrian flow modeling applies geocomputational techniques to understand the patterns of human movement within an environment. A key principle is that the “cost” (time, distance, energy, etc.) of travel along different routes is affected by environmental factors, such as terrain variation or land cover. High-probability travel routes can therefore be estimated by performing a least-cost analysis on base geographic data layers. The primary data input to these methods are 3D structural information related to the Earth topography and the objects upon its surface. However, generating and analyzing such data at high resolution (1 m) has only recently become viable at scale. We present a data fusion approach that combines recent advances in dense stereo reconstruction, in conjunction with multispectral-derived land cover, towards exploring the travel routes of pedestrians in urban environments. Most previous studies focus on rural areas; this analysis provides a novel entry into understanding the flow of pedestrians in more densely populated areas, with applications to humanitarian assistance and disaster response (HADR).
David Kelbe, Devin White, David Page, Kristin Safi, Andrew Hardin, Amy N. Rose
IGARSS3
2017 Pharmacovigilance via Baseline Regularization with Large-Scale Longitudinal Observational Data
abstract
Several prominent public health incidents that occurred at the beginning of this century due to adverse drug events (ADEs) have raised international awareness of governments and industries about pharmacovigilance (PhV), the science and activities to monitor and prevent adverse events caused by pharmaceutical products after they are introduced to the market. A major data source for PhV is large-scale longitudinal observational databases (LODs) such as electronic health records (EHRs) and medical insurance claim databases. Inspired by the Multiple Self-Controlled Case Series (MSCCS) model, arguably the leading method for ADE discovery from LODs, we propose baseline regularization, a regularized generalized linear model that leverages the diverse health profiles available in LODs across different individuals at different times. We apply the proposed method as well as MSCCS to the Marshfield Clinic EHR. Experimental results suggest that incorporating the heterogeneity among different patients and different times help to improve the performance in identifying benchmark ADEs from the Observational Medical Outcomes Partnership ground truth
Zhaobin Kuang, Peggy L. Peissig, Vítor Santos Costa, Richard Maclin, David Page
KDD5
2017 A Screening Rule for l1-Regularized Ising Model Estimation
abstract
We discover a screening rule for l1-regularized Ising model estimation. The simple closed-form screening rule is a necessary and sufficient condition for exactly recovering the blockwise structure of a solution under any given regularization parameters. With enough sparsity, the screening rule can be combined with various optimization procedures to deliver solutions efficiently in practice. The screening rule is especially suitable for large-scale exploratory data analysis, where the number of variables in the dataset can be thousands while we are only interested in the relationship among a handful of variables within moderate-size clusters for interpretability. Experimental results on various datasets demonstrate the efficiency and insights gained from the introduction of the screening rule.
Zhaobin Kuang, Sinong Geng, David Page
NIPS3
2017 Markov logic networks for adverse drug event extraction from text
Sriraam Natarajan, Vishal Bangera, Tushar Khot, Jose Picado, Anurag Wazalwar, Vítor Santos Costa, David Page, Michael Caldwell
Knowl. Inf. Syst.7
2016 Baseline Regularization for Computational Drug Repositioning with Longitudinal Observational Data
Zhaobin Kuang, James A. Thomson, Michael Caldwell, Peggy L. Peissig, Ron M. Stewart, David Page
IJCAI6
2016 Computational Drug Repositioning Using Continuous Self-Controlled Case Series
abstract
Computational Drug Repositioning (CDR) is the task of discovering potential new indications for existing drugs by mining large-scale heterogeneous drug-related data sources. Leveraging the patient-level temporal ordering information between numeric physiological measurements and various drug prescriptions provided in Electronic Health Records (EHRs), we propose a Continuous Self-controlled Case Series (CSCCS) model for CDR. As an initial evaluation, we look for drugs that can control Fasting Blood Glucose (FBG) level in our experiments. Applying CSCCS to the Marshfield Clinic EHR, well-known drugs that are indicated for controlling blood glucose level are rediscovered. Furthermore, some drugs with recent literature support for the potential effect of blood glucose level control are also identified.
Zhaobin Kuang, James A. Thomson, Michael Caldwell, Peggy L. Peissig, Ron M. Stewart, David Page
KDD6
2016 Structure-Leveraged Methods in Breast Cancer Risk Prediction
abstract
Predicting breast cancer risk has long been a goal of medical research in the pursuit of precision medicine. The goal of this study is to develop novel penalized methods to improve breast cancer risk prediction by leveraging structure information in electronic health records. We conducted a retrospective case- control study, garnering 49 mammography descriptors and 77 high- frequency/low-penetrance single-nucleotide polymorphisms (SNPs) from an existing personalized medicine data repository. Structured mammography reports and breast imaging features have long been part of a standard electronic health record (EHR), and genetic markers likely will be in the near future. Lasso and its variants are widely used approaches to integrated learning and feature selection, and our methodological contribution is to incorporate the dependence structure among the features into these approaches. More specifically, we propose a new methodology by combining group penalty and $\ell^p$ ($1\leq p\leq2$) fusion penalty to improve breast cancer risk prediction, taking into account structure information in mammography descriptors and SNPs. We demonstrate that our method provides benefits that are both statistically significant and potentially significant to people's lives.
Yirong Wu, Ming Yuan 0001, David Page, Jie Liu 0006, Irene M. Ong, Peggy L. Peissig, Elizabeth S. Burnside
J. Mach. Learn. Res.4
2015 Learning to Reject Sequential Importance Steps for Continuous-Time Bayesian Networks
abstract
Applications of graphical models often require the use of approximate inference, such as sequential importance sampling (SIS), for estimation of the model distribution given partial evidence, i.e., the target distribution. However, when SIS proposal and target distributions are dissimilar, such procedures lead to biased estimates or require a prohibitive number of samples. We introduce ReBaSIS, a method that better approximates the target distribution by sampling variable by variable from existing importance samplers and accepting or rejecting each proposed assignment in the sequence: a choice made based on anticipating upcoming evidence. We relate the per-variable proposal and model distributions by expected weight ratios of sequence completions and show that we can learn accurate models of optimal acceptance probabilities from local samples. In a continuous-time domain, our method improves upon previous importance samplers by transforming an SIS problem into a machine learning one.
Jeremy C. Weiss, Sriraam Natarajan, David Page
AAAI3
2015 Extracting Adverse Drug Events from Text Using Human Advice
Phillip Odom, Vishal Bangera, Tushar Khot, David Page, Sriraam Natarajan
AIME4
2015 Machine Learning for Treatment Assignment: Improving Individualized Risk Attribution
Jeremy C. Weiss, Finn Kuusisto, Kendrick Boyd, Jie Liu 0006, David Page
AMIA5
2015 Sparse modeling of spatial environmental variables associated with asthma
Timothy S. Chang, Ronald E. Gangnon, David Page, William R. Buckingham, Aman Tandias, Kelly J. Cowan, Carrie D. Tomasallo, Brian G. Arndt, Lawrence P. Hanrahan, Theresa W. Guilbert
J. Biomed. Informatics3
2014 Learning Heterogeneous Hidden Markov Random Fields
abstract
Hidden Markov random fields (HMRFs) are conventionally assumed to be homogeneous in the sense that the potential functions are invariant across different sites. However in some biological applications, it is desirable to make HMRFs heterogeneous, especially when there exists some background knowledge about how the potential functions vary. We formally define heterogeneous HMRFs and propose an EM algorithm whose M-step combines a contrastive divergence learner with a kernel smoothing step to incorporate the background knowledge. Simulations show that our algorithm is effective for learning heterogeneous HMRFs and outperforms alternative binning methods. We learn a heterogeneous HMRF in a real-world study.
Jie Liu 0006, Elizabeth S. Burnside, David Page
AISTATS4
2014 Comparing the Value of Mammographic Features and Genetic Variants in Breast Cancer Risk Prediction
Yirong Wu, Jie Liu 0006, David Page, Peggy L. Peissig, Catherine A. McCarty, Adedayo A. Onitilo, Elizabeth S. Burnside
AMIA3
2014 Multiple Testing under Dependence via Semiparametric Graphical Models
abstract
It has been shown that graphical models can be used to leverage the dependence in large-scale multiple testing problems with significantly improved performance (Sun & Cai, 2009; Liu et al., 2012). These graphical models are fully parametric and require that we know the parameterization of f1, the density function of the test statistic under the alternative hypothesis. However in practice, f1 is often heterogeneous, and cannot be estimated with a simple parametric distribution. We propose a novel semiparametric approach for multiple testing under dependence, which estimates f1 adaptively. This semiparametric approach exactly generalizes the local FDR procedure (Efron et al., 2001) and connects with the BH procedure (Benjamini & Hochberg, 1995). A variety of simulations show that our semiparametric approach outperforms classical procedures which assume independence and the parametric approaches which capture dependence.
Jie Liu 0006, Elizabeth S. Burnside, David Page
ICML4
2014 Support Vector Machines for Differential Prediction
Finn Kuusisto, Vítor Santos Costa, Houssam Nassif, Elizabeth S. Burnside, David Page, Jude W. Shavlik
ECML/PKDD (2)5
2014 Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing
Matt Fredrikson, Eric Lantz, Somesh Jha, Simon M. Lin, David Page, Thomas Ristenpart
USENIX Security Symposium5
2014 Relational machine learning for electronic health record-driven phenotyping
Peggy L. Peissig, Vítor Santos Costa, Michael Caldwell, Carla Rottscheit, Richard L. Berg, Eneida A. Mendonça, David Page
J. Biomed. Informatics7
2014 QuickFOIL: Scalable Inductive Logic Programming
abstract
Inductive Logic Programming (ILP) is a classic machine learning technique that learns first-order rules from relational-structured data. However, to-date most ILP systems can only be applied to small datasets (tens of thousands of examples). A long-standing challenge in the field is to scale ILP methods to larger data sets. This paper presents a method called QuickFOIL that addresses this limitation. QuickFOIL employs a new scoring function and a novel pruning strategy that enables the algorithm to find high-quality rules. QuickFOIL can also be implemented as an in-RDBMS algorithm. Such an implementation presents a host of query processing and optimization challenges that we address in this paper. Our empirical evaluation shows that QuickFOIL can scale to large datasets consisting of hundreds of millions tuples, and is often more than order of magnitude more efficient than other existing approaches.
Qiang Zeng 0002, Jignesh M. Patel, David Page
Proc. VLDB Endow.3
2013 Genetic Variants Improve Breast Cancer Risk Prediction on Mammograms
Jie Liu 0006, David Page, Houssam Nassif, Jude W. Shavlik, Peggy L. Peissig, Catherine A. McCarty, Adedayo A. Onitilo, Elizabeth S. Burnside
AMIA2
2013 On Differentially Private Inductive Logic Programming
Eric Lantz, Jeffrey F. Naughton, David Page
ILP4
2013 Bayesian Estimation of Latently-grouped Parameters in Undirected Graphical Models
abstract
In large-scale applications of undirected graphical models, such as social networks and biological networks, similar patterns occur frequently and give rise to similar parameters. In this situation, it is beneficial to group the parameters for more efficient learning. We show that even when the grouping is unknown, we can infer these parameter groups during learning via a Bayesian approach. We impose a Dirichlet process prior on the parameters. Posterior inference usually involves calculating intractable terms, and we propose two approximation algorithms, namely a Metropolis-Hastings algorithm with auxiliary variables and a Gibbs sampling algorithm with stripped Beta approximation (GibbsSBA). Simulations show that both algorithms outperform conventional maximum likelihood estimation (MLE). GibbsSBA's performance is close to Gibbs sampling with exact likelihood calculation. Models learned with Gibbs_SBA also generalize better than the models learned by MLE on real-world Senate voting data.
Jie Liu 0006, David Page
NIPS2
2013 Area under the Precision-Recall Curve: Point Estimates and Confidence Intervals
Kendrick Boyd, Kevin H. Eng, David Page
ECML/PKDD (3)3
2013 Erratum: Area under the Precision-Recall Curve: Point Estimates and Confidence Intervals
Kendrick Boyd, Kevin H. Eng, David Page
ECML/PKDD (3)3
2013 Score As You Lift (SAYL): A Statistical Relational Learning Approach to Uplift Modeling
Houssam Nassif, Finn Kuusisto, Elizabeth S. Burnside, David Page, Jude W. Shavlik, Vítor Santos Costa
ECML/PKDD (3)4
2013 Forest-Based Point Process for Event Prediction from Electronic Health Records
Jeremy C. Weiss, David Page
ECML/PKDD (3)2
2012 Identifying Adverse Drug Events by Relational Learning
abstract
The pharmaceutical industry, consumer protection groups, users of medications and government oversight agencies are all strongly interested in identifying adverse reactions to drugs. While a clinical trial of a drug may use only a thousand patients, once a drug is released on the market it may be taken by millions of patients. As a result, in many cases adverse drug events (ADEs) are observed in the broader population that were not identified during clinical trials. Therefore, there is a need for continued, postmarketing surveillance of drugs to identify previously-unanticipated ADEs. This paper casts this problem as a reverse machine learning task, related to relational subgroup discovery and provides an initial evaluation of this approach based on experiments with an actual EMR/EHR and known adverse drug events.
David Page, Vítor Santos Costa, Sriraam Natarajan, Aubrey Barnard, Peggy L. Peissig, Michael Caldwell
AAAI1
2012 Logical Differential Prediction Bayes Net, improving breast cancer diagnosis for older women
Houssam Nassif, Yirong Wu, David Page, Elizabeth S. Burnside
AMIA3
2012 Extracting BI-RADS features from Portuguese clinical texts
abstract
In this work we build the first BI-RADS parser for Portuguese free texts, modeled after existing approaches to extract BI-RADS features from English medical records. Our concept finder uses a semantic grammar based on the BIRADS lexicon and on iterative transferred expert knowledge. We compare the performance of our algorithm to manual annotation by a specialist in mammography. Our results show that our parser's performance is comparable to the manual method.
Houssam Nassif, Filipe Cunha, Inês C. Moreira, Ricardo João Cruz Correia, Eliana Sousa, David Page, Elizabeth S. Burnside, Inês de Castro Dutra
BIBM6
2012 Statistical Relational Learning to Predict Primary Myocardial Infarction from Electronic Health Records
abstract
Electronic health records (EHRs) are an emerging relational domain with large potential to improve clinical outcomes. We apply two statistical relational learning (SRL) algorithms to the task of predicting primary myocardial infarction. We show that one SRL algorithm, relational functional gradient boosting, outperforms propositional learners particularly in the medically-relevant high recall region. We observe that both SRL algorithms predict outcomes better than their propositional analogs and suggest how our methods can augment current epidemiological practices.
Jeremy C. Weiss, Sriraam Natarajan, Peggy L. Peissig, Catherine A. McCarty, David Page
IAAI5
2012 Unachievable Region in Precision-Recall Space and Its Effect on Empirical Evaluation
Kendrick Boyd, Jesse Davis, David Page, Vítor Santos Costa
ICML3
2012 Demand-Driven Clustering in Relational Domains for Predicting Adverse Drug Events
Jesse Davis, Vítor Santos Costa, Elizabeth Berg, David Page, Peggy L. Peissig, Michael Caldwell
ICML4
2012 Multiplicative Forests for Continuous-Time Processes
abstract
Learning temporal dependencies between variables over continuous time is an important and challenging task. Continuous-time Bayesian networks effectively model such processes but are limited by the number of conditional intensity matrices, which grows exponentially in the number of parents per variable. We develop a partition-based representation using regression trees and forests whose parameter spaces grow linearly in the number of node splits. Using a multiplicative assumption we show how to update the forest likelihood in closed form, producing efficient model updates. Our results show multiplicative forests can be learned from few temporal trajectories with large gains in performance and scalability.
Jeremy C. Weiss, Sriraam Natarajan, David Page
NIPS3
2012 Relational Differential Prediction
Houssam Nassif, Vítor Santos Costa, Elizabeth S. Burnside, David Page
ECML/PKDD (1)4
2012 Graphical-model Based Multiple Testing under Dependence, with Applications to Genome-wide Association Studies
Jie Liu 0006, Catherine A. McCarty, Peggy L. Peissig, Elizabeth S. Burnside, David Page
UAI6
2012 Automated identification of protein-ligand interaction features using Inductive Logic Programming: a hexose binding case study
abstract
BACKGROUND: There is a need for automated methods to learn general features of the interactions of a ligand class with its diverse set of protein receptors. An appropriate machine learning approach is Inductive Logic Programming (ILP), which automatically generates comprehensible rules in addition to prediction. The development of ILP systems which can learn rules of the complexity required for studies on protein structure remains a challenge. In this work we use a new ILP system, ProGolem, and demonstrate its performance on learning features of hexose-protein interactions. RESULTS: The rules induced by ProGolem detect interactions mediated by aromatics and by planar-polar residues, in addition to less common features such as the aromatic sandwich. The rules also reveal a previously unreported dependency for residues cys and leu. They also specify interactions involving aromatic and hydrogen bonding residues. This paper shows that Inductive Logic Programming implemented in ProGolem can derive rules giving structural features of protein/ligand interactions. Several of these rules are consistent with descriptions in the literature. CONCLUSIONS: In addition to confirming literature results, ProGolem's model has a 10-fold cross-validated predictive accuracy that is superior, at the 95% confidence level, to another ILP system previously used to study protein/hexose interactions and is comparable with state-of-the-art statistical learners.
Jose Santos 0001, Houssam Nassif, David Page, Stephen H. Muggleton, Michael J. E. Sternberg
BMC Bioinform.3
2011 Integrating knowledge capture and supervised learning through a human-computer interface
abstract
Some supervised-learning algorithms can make effective use of domain knowledge in addition to the input-output pairs commonly used in machine learning. However, formulating this additional information often requires an in-depth understanding of the specific knowledge representation used by a given learning algorithm. The requirement to use a formal knowledge-representation language means that most domain experts will not be able to articulate their expertise, even when a learning algorithm is capable of exploiting such valuable information. We investigate a method to ease this knowledge acquisition through the use of a graphical, human-computer interface. Our interface allows users to easily provide advice about specific examples, rather than requiring them to provide general rules; we leave the task of properly generalizing such advice to the learning algorithms. We demonstrate the effectiveness of our approach using the Wargus real-time strategy game, comparing learning with no advice to learning with concrete advice provided through our interface, as well as comparing to using generalized advice written by an AI expert. Our results show that our approach of combining a GUI-based advice language with an advice-taking learning algorithm is an effective way to capture domain knowledge.
Trevor Walker, Gautam Kunapuli, Noah Larsen, David Page, Jude W. Shavlik
K-CAP4
2010 Automating the ILP Setup Task: Converting User Advice about Specific Examples into General Background Knowledge
Trevor Walker, Ciaran O'Reilly, Gautam Kunapuli, Sriraam Natarajan, Richard Maclin, David Page, Jude W. Shavlik
ILP6
2009 An Inductive Logic Programming Approach to Validate Hexose Binding Biochemical Knowledge
Houssam Nassif, Hassan Al-Ali, Sawsan Khuri, Walid Keirouz, David Page
ILP5
2009 Exploiting Product Distributions to Identify Relevant Variables of Correlation Immune Functions
Lisa Hellerstein, Bernard Rosell, Eric Bach 0001, Soumya Ray, David Page
J. Mach. Learn. Res.5
2008 Matching isotopic distributions from metabolically labeled samples
abstract
MOTIVATION: In recent years stable isotopic labeling has become a standard approach for quantitative proteomic analyses. Among the many available isotopic labeling strategies, metabolic labeling is attractive for the excellent internal control it provides. However, analysis of data from metabolic labeling experiments can be complicated because the spacing between labeled and unlabeled forms of each peptide depends on its sequence, and is thus variable from analyte to analyte. As a result, one generally needs to know the sequence of a peptide to identify its matching isotopic distributions in an automated fashion. In some experimental situations it would be necessary or desirable to match pairs of labeled and unlabeled peaks from peptides of unknown sequence. This article addresses this largely overlooked problem in the analysis of quantitative mass spectrometry data by presenting an algorithm that not only identifies isotopic distributions within a mass spectrum, but also annotates matches between natural abundance light isotopic distributions and their metabolically labeled counterparts. This algorithm is designed in two stages: first we annotate the isotopic peaks using a modified version of the IDM algorithm described last year; then we use a probabilistic classifier that is supplemented by dynamic programming to find the metabolically labeled matched isotopic pairs. Such a method is needed for high-throughput quantitative proteomic metabolomic experiments measured via mass spectrometry. RESULTS: The primary result of this article is that the dynamic programming approach performs well given perfect isotopic distribution annotations. Our algorithm achieves a true positive rate of 99% and a false positive rate of 1% using perfect isotopic distribution annotations. When the isotopic distributions are annotated given 'expert' selected peaks, the same algorithm gets a true positive rate of 77% and a false positive rate of 1%. Finally, when annotating using machine selected peaks, which may contain noise, the dynamic programming algorithm gives a true positive rate of 36% and a false positive rate of 1%. It is important to mention that these rates arise from the requirement of exact annotations of both the light and heavy isotopic distributions. In our evaluations, a match is considered 'entirely incorrect' if it is missing even one peak or containing an extraneous peak. If we only require that the 'monoisotopic' peaks exist within the two matched distributions, our algorithm obtains a positive rate of 45% and a false positive rate of 1% on the 'machine' selected data. Changes to the algorithm's scoring function and training example generation improves our 'monoisotopic' peak score true positive rate to 65% while obtaining a false positive rate of 2%. All results were obtained within 10-fold cross-validation of 41 mass spectra with a mass-to-charge range of 800-4000 m/z. There are a total of 713 isotopic distributions and 255 matched isotopic pairs that are hand-annotated for this study. AVAILABILITY: Programs are available via http://www.cs.wisc.edu/~mcilwain/IDM/.
Sean McIlwain, David Page, Edward L. Huttlin, Michael R. Sussman
ISMB2
2007 An integrated approach to feature invention and model construction for drug activity prediction
abstract
We present a new machine learning approach for 3D-QSAR, the task of predicting binding affinities of molecules to target proteins based on 3D structure. Our approach predicts binding affinity by using regression on substructures discovered by relational learning. We make two contributions to the state-of-the-art. First, we use multiple-instance (MI) regression, which represents a molecule as a set of 3D conformations, to model activity. Second, the relational learning component employs the "Score As You Use" (SAYU) method to select substructures for their ability to improve the regression model. This is the first application of SAYU to multiple-instance, real-valued prediction. We evaluate our approach on three tasks and demonstrate that (i) SAYU outperforms standard coverage measures when selecting features for regression, (ii) the MI representation improves accuracy over standard single feature-vector encodings and (iii) combining SAYU with MI regression is more accurate for 3D-QSAR than either approach by itself.
Jesse Davis, Vítor Santos Costa, Soumya Ray, David Page
ICML4
2007 Change of Representation for Statistical Relational Learning
Jesse Davis, Irene M. Ong, Jan Struyf, Elizabeth S. Burnside, David Page, Vítor Santos Costa
IJCAI5
2007 Learning Bayesian Network Structure from Correlation-Immune Data
Eric Lantz, Soumya Ray, David Page
UAI3
2006 An Efficient Approximation to Lookahead in Relational Learners
Jan Struyf, Jesse Davis, David Page
ECML3
2006 Inferring Regulatory Networks from Time Series Expression Data and Relational Data Via Inductive Logic Programming
Irene M. Ong, Scott E. Topper, David Page, Vítor Santos Costa
ILP3
2006 ILP Through Propositionalization and Stochastic k-Term DNF Learning
Aline Paes, Filip Zelezný, Gerson Zaverucha, David Page, Ashwin Srinivasan 0001
ILP4
2006 Quantitative pharmacophore models with inductive logic programming
Ashwin Srinivasan 0001, David Page, Rui Camacho, Ross D. King
Mach. Learn.2
2006 Randomised restarted search in ILP
Filip Zelezný, Ashwin Srinivasan 0001, David Page
Mach. Learn.3
2005 Knowledge Discovery from Structured Mammography Reports Using Inductive Logic Programming
Elizabeth S. Burnside, Jesse Davis, Vítor Santos Costa, Inês de Castro Dutra, Charles E. Kahn Jr., Jason Fine, David Page
AMIA7
2005 An Integrated Approach to Learning Bayesian Networks of Rules
Jesse Davis, Elizabeth S. Burnside, Inês de Castro Dutra, David Page, Vítor Santos Costa
ECML4
2005 Mode Directed Path Finding
Irene M. Ong, Inês de Castro Dutra, David Page, Vítor Santos Costa
ECML3
2005 Multi-instance tree learning
abstract
We introduce a novel algorithm for decision tree learning in the multi-instance setting as originally defined by Dietterich et al. It differs from existing multi-instance tree learners in a few crucial, well-motivated details. Experiments on synthetic and real-life datasets confirm the beneficial effect of these differences and show that the resulting system outperforms the existing multi-instance decision tree learners.
Hendrik Blockeel, David Page, Ashwin Srinivasan 0001
ICML2
2005 Generalized skewing for functions with continuous and nominal attributes
abstract
This paper extends previous work on skewing, an approach to problematic functions in decision tree induction. The previous algorithms were applicable only to functions of binary variables. In this paper, we extend skewing to directly handle functions of continuous and nominal variables. We present experiments with randomly generated functions and a number of real world datasets to evaluate the algorithm's accuracy. Our results indicate that our algorithm almost always outperforms an Information Gain-based decision tree learner.
Soumya Ray, David Page
ICML2
2005 Why skewing works: learning difficult Boolean functions with greedy tree learners
abstract
We analyze skewing, an approach that has been empirically observed to enable greedy decision tree learners to learn "difficult" Boolean functions, such as parity, in the presence of irrelevant variables. We prove tha, in an idealized setting, for any function and choice of skew parameters, skewing finds relevant variables with probability 1. We present experiments exploring how different parameter choices affect the success of skewing in empirical settings. Finally, we analyze a variant of skewing called Sequential Skewing.
Bernard Rosell, Lisa Hellerstein, Soumya Ray, David Page
ICML4
2005 View Learning for Statistical Relational Learning: With an Application to Mammography
Jesse Davis, Elizabeth S. Burnside, Inês de Castro Dutra, David Page, Raghu Ramakrishnan 0001, Vítor Santos Costa, Jude W. Shavlik
IJCAI4
2005 A Framework for Set-Oriented Computation in Inductive Logic Programming and Its Application in Generalizing Inverse Entailment
Héctor Corrada Bravo, David Page, Raghu Ramakrishnan 0001, Jude W. Shavlik, Vítor Santos Costa
ILP2
2004 Sequential skewing: an improved skewing algorithm
abstract
This paper extends previous work on the Skewing algorithm, a promising approach that allows greedy decision tree induction algorithms to handle problematic functions such as parity functions with a lower run-time penalty than Lookahead. A deficiency of the previously proposed algorithm is its inability to scale up to high dimensional problems. In this paper, we describe a modified algorithm that scales better with increasing numbers of variables. We present experiments with randomly generated Boolean functions that evaluate the algorithm's response to increasing dimensions. We also evaluate the algorithm on a challenging real world biomedical problem, that of SH3 domain binding. Our results indicate that our algorithm almost always outperforms an information gain-based decision tree learner.
Soumya Ray, David Page
ICML2
2004 A Monte Carlo Study of Randomised Restarted Search in ILP
Filip Zelezný, Ashwin Srinivasan 0001, David Page
ILP3
2003 Toward Automatic Management of Embarrassingly Parallel Applications
Inês de Castro Dutra, David Page, Vítor Santos Costa, Jude W. Shavlik, Michael Waddell
Euro-Par2
2003 Skewing: An Efficient Alternative to Lookahead for Decision Tree Induction
David Page, Soumya Ray
IJCAI1
2003 The Role of Declarative Languages in Mining Biological Databases
David Page
PADL1
2003 CLP(BN): Constraint Logic Programming for Probabilistic Knowledge
Vítor Santos Costa, David Page, Maleeha Qazi, James Cussens
UAI2
2003 A Bayesian Network Approach to Operon Prediction
abstract
MOTIVATION: In order to understand transcription regulation in a given prokaryotic genome, it is critical to identify operons, the fundamental units of transcription, in such species. While there are a growing number of organisms whose sequence and gene coordinates are known, by and large their operons are not known. RESULTS: We present a probabilistic approach to predicting operons using Bayesian networks. Our approach exploits diverse evidence sources such as sequence and expression data. We evaluate our approach on the Escherichia coli K-12 genome where our results indicate we are able to identify over 78% of its operons at a 10% false positive rate. Also, empirical evaluation using a reduced set of data sources suggests that our approach may have significant value for organisms that do not have as rich of evidence sources as E.coli. AVAILABILITY: Our E.coli K-12 operon predictions are available at http://www.biostat.wisc.edu/gene-regulation.
Joseph Bockhorst, Mark W. Craven, David Page, Jude W. Shavlik, Jeremy D. Glasner
Bioinform.3
2003 ILP: A Short Look Back and a Longer Look Forward
David Page, Ashwin Srinivasan 0001
J. Mach. Learn. Res.1
2002 An Empirical Evaluation of Bagging in Inductive Logic Programming
Inês de Castro Dutra, David Page, Vítor Santos Costa, Jude W. Shavlik
ILP2
2002 Lattice-Search Runtime Distributions May Be Heavy-Tailed
Filip Zelezný, Ashwin Srinivasan 0001, David Page
ILP3
2002 Modelling regulatory pathways in E. coli from time series expression profiles
abstract
MOTIVATION: Cells continuously reprogram their gene expression network as they move through the cell cycle or sense changes in their environment. In order to understand the regulation of cells, time series expression profiles provide a more complete picture than single time point expression profiles. Few analysis techniques, however, are well suited to modelling such time series data. RESULTS: We describe an approach that naturally handles time series data with the capabilities of modelling causality, feedback loops, and environmental or hidden variables using a Dynamic Bayesian network. We also present a novel way of combining prior biological knowledge and current observations to improve the quality of analysis and to model interactions between sets of genes rather than individual genes. Our approach is evaluated on time series expression data measured in response to physiological changes that affect tryptophan metabolism in E. coli. Results indicate that this approach is capable of finding correlations between sets of related genes.
Irene M. Ong, Jeremy D. Glasner, David Page
ISMB3
2001 Multiple Instance Regression
Soumya Ray, David Page
ICML2
2000 Using Multiple Levels of Learning and Diverse Evidence to Uncover Coordinately Controlled Genes
Mark W. Craven, David Page, Jude W. Shavlik, Joseph Bockhorst, Jeremy D. Glasner
ICML2
2000 ILP: Just Do It
David Page
ILP1
2000 A Probabilistic Learning Approach to Whole-Genome Operon Prediction
Mark W. Craven, David Page, Jude W. Shavlik, Joseph Bockhorst, Jeremy D. Glasner
ISMB2
2000 Parallel data mining for pharmacophore discovery
abstract
Rapid and effective design of new drugs to combat new strains of antibiotic resistant organisms, more effectively treat chronic conditions, and provide other life sustaining treatment is a key challenge for the medical industry. Current drug design methodologies can take several years just in the initial chemical evaluation stages before compounds can be created for animal and human testing. This paper presents some recent research results in a new parallel machine learning approach that can expedite the drug design cycle. An inductive logic programming search has been reformulated and parallelized to run on an eight node Beowulf cluster. Initial testing with several data sets indicate almost linear speedup using the cluster.
James Graham, David Page, Alan Wild
SMC2
1998 Pharmacophore Discovery Using the Inductive Logic Programming System PROGOL
Paul W. Finn, Stephen H. Muggleton, David Page, Ashwin Srinivasan 0001
Mach. Learn.3
1997 Guest Editors' Introduction
Stephen H. Muggleton, David Page
Mach. Learn.2
1995 Building Theories into Instantiation
Alan M. Frisch, David Page
IJCAI2
1994 Prefix Grammars: An Alternative Characterization of the Regular Languages
Michael Frazier, David Page
Inf. Process. Lett.2
1993 Learnability in Inductive Logic Programrning: Some Basic Results and Techniques
Michael Frazier, David Page
AAAI2
1991 Learning Constrained Atoms
David Page, Alan M. Frisch
ML1
1991 Generalizing Atoms in Constraint Logic
David Page, Alan M. Frisch
KR1
1990 Generalization with Taxonomic Information
Alan M. Frisch, David Page
AAAI2