Barbara E. Engelhardt

dblp:53/328 · also Barbara Engelhardt · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-6139-7334ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 52% Reinforcement learning · 24% Efficient and distributed learning · 18%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 86% Computational science and engineering · 14%
Theoretical computer science
2 papers
Mathematical optimization · 50% Algorithms and data structures · 50%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
1.622025
Bayesian Multi-Group Gaussian Process Models for Heterogeneous Group-Structured Data · J. Mach. Learn. Res. 2025
Active Learning for Derivative-Based Global Sensitivity Analysis with Gaussian Processes · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
kernel design
0.912025
Bayesian Multi-Group Gaussian Process Models for Heterogeneous Group-Structured Data · J. Mach. Learn. Res. 2025
Machine learning › Efficient and distributed learning
active learning
0.812024
Active Learning for Derivative-Based Global Sensitivity Analysis with Gaussian Processes · NeurIPS 2024
Computational science and engineering › statistical computing
data imputation
0.412019
netNMF-sc: A Network Regularization Algorithm for Dimensionality Reduction and Imputation of Single-Cell Expression Data · RECOMB 2019
Bioinformatics and computational biology › statistical genetics
quantitative genetics
0.412019
Statistical tests for detecting variance effects in quantitative trait studies · Bioinform. 2019
Bioinformatics and computational biology › single-cell analysis
single-cell transcriptomics
0.412019
netNMF-sc: A Network Regularization Algorithm for Dimensionality Reduction and Imputation of Single-Cell Expression Data · RECOMB 2019
Machine learning › Reinforcement learning › bandit
contextual bandit
0.312018
PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits · NeurIPS 2018
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.312018
PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits · NeurIPS 2018
Machine learning › Reinforcement learning
regret minimization
0.312018
PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits · NeurIPS 2018
Machine learning › Reinforcement learning
thompson sampling
0.312018
PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits · NeurIPS 2018
Bioinformatics and computational biology › statistical genetics › genetic association study
association mapping
0.312017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Bioinformatics and computational biology › statistical genetics
linear mixed model
0.312017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.312017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Algorithms and data structures › numerical linear algebra
dimensionality reduction
0.312017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Algorithms and data structures › numerical linear algebra › dimensionality reduction
principal component analysis
0.312017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Bioinformatics and computational biology
gene expression analysis
0.312025
Bayesian Multi-Group Gaussian Process Models for Heterogeneous Group-Structured Data · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor model
0.212016
Bayesian group factor analysis with structured sparsity · J. Mach. Learn. Res. 2016
Machine learning › Representation and self-supervised learning › matrix factorization
nonnegative matrix factorization
0.212016
Hierarchical Compound Poisson Factorization · ICML 2016
Machine learning › Efficient and distributed learning › model compression › sparsity
structured sparsity
0.212016
Bayesian group factor analysis with structured sparsity · J. Mach. Learn. Res. 2016
Recommender systems
collaborative filtering
0.212016
Hierarchical Compound Poisson Factorization · ICML 2016
Recommender systems › collaborative filtering
matrix factorization
0.212016
Hierarchical Compound Poisson Factorization · ICML 2016
Recommender systems › collaborative filtering › factor models
poisson factorization
0.212016
Hierarchical Compound Poisson Factorization · ICML 2016
Bioinformatics and computational biology › gene regulation › transcription factor binding
transcription factor binding specificity
0.212013
Stability selection for regression-based models of transcription factor-DNA binding specificity · Bioinform. 2013
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.112018
PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits · NeurIPS 2018
Algorithms and data structures › numerical linear algebra
matrix compression
0.112017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Algorithms and data structures › numerical linear algebra
randomized numerical linear algebra
0.112017
Adaptive Randomized Dimension Reduction on Massive Data · J. Mach. Learn. Res. 2017
Recommender systems › collaborative filtering
implicit feedback
0.112016
Hierarchical Compound Poisson Factorization · ICML 2016
Bioinformatics and computational biology
protein function prediction
0.112006
A graphical model for predicting protein molecular function · ICML 2006
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › domain-independent planning
factored planning
0.012003
Factored Planning · IJCAI 2003
Bioinformatics and computational biology › protein function prediction › protein classification
protein family classification
0.012006
A graphical model for predicting protein molecular function · ICML 2006

Methods — techniques the papers use, named apart from their topics

bayesian inference · 2.4gaussian process regression · 1.7information gain acquisition · 1.5gaussian process surrogate models · 1.5regularization · 0.6randomized low-rank approximation · 0.6variational inference · 0.5gamma-poisson model · 0.5non-negative matrix factorization · 0.4network regularization · 0.4double generalized linear model · 0.4polya-gamma augmentation · 0.3gibbs sampling · 0.3parameter expansion · 0.2expectation-maximization · 0.2regression · 0.2position weight matrix · 0.2feature selection · 0.2
YearPublicationVenuePosition
2025 Preference-Guided Diffusion for Multi-Objective Offline Optimization
abstract
Offline multi-objective optimization aims to identify Pareto-optimal solutions given a dataset of designs and their objective values. In this work, we propose a preference-guided diffusion model that generates Pareto-optimal designs by leveraging a classifier-based guidance mechanism. Our guidance classifier is a preference model trained to predict the probability that one design dominates another, directing the diffusion model toward optimal regions of the design space. Crucially, this preference model generalizes beyond the training distribution, enabling the discovery of Pareto-optimal solutions outside the observed dataset. We introduce a novel diversity-aware preference guidance, augmenting Pareto dominance preference with diversity criteria. This ensures that generated solutions are optimal and well-distributed across the objective space, a capability absent in prior generative methods for offline multi-objective optimization. We evaluate our approach on various continuous offline multi-objective optimization tasks and find that it consistently outperforms other inverse/generative approaches while remaining competitive with forward/ surrogate-based optimization methods. Our results highlight the effectiveness of classifier-guided diffusion models in generating diverse and high-quality solutions that approximate the Pareto front well.
Yashas Annadani, Syrine Belakaria, Stefano Ermon, Stefan Bauer, Barbara E. Engelhardt
NeurIPS5
2025 Bayesian Multi-Group Gaussian Process Models for Heterogeneous Group-Structured Data
abstract
Gaussian processes are pervasive in functional data analysis, machine learning, and spatial statistics for modeling complex dependencies. Scientific data are often heterogeneous in their inputs and contain multiple known discrete groups of samples; thus, it is desirable to leverage the similarity among groups while accounting for heterogeneity across groups. We propose multi-group Gaussian processes (MGGPs) defined over $\mathbb{R}^p\times \mathscr{C}$, where $\mathscr{C}$ is a finite set representing the group label, by developing general classes of valid (positive definite) covariance functions on such domains. MGGPs are able to accurately recover relationships between the groups and efficiently share strength across samples from all groups during inference, while capturing distinct group-specific behaviors in the conditional posterior distributions. We demonstrate inference in MGGPs through simulation experiments, and we apply our proposed MGGP regression framework to gene expression data to illustrate the behavior and enhanced inferential capabilities of multi-group Gaussian processes by jointly modeling continuous and categorical variables.
Didong Li, Sudipto Banerjee, Barbara E. Engelhardt
J. Mach. Learn. Res.4
2024 Automating Transparency Mechanisms in the Judicial System Using LLMs: Opportunities and Challenges
abstract
Bringing more transparency to the judicial system for the purposes of increasing accountability often demands extensive effort from auditors who must meticulously sift through numerous disorganized legal case files to detect patterns of bias and errors. For example, the high-profile investigation into the Curtis Flowers case took seven reporters a full year to assemble evidence about the prosecutor's history of selecting racially biased juries. LLMs have the potential to automate and scale these transparency pipelines, especially given their demonstrated capabilities to extract information from unstructured documents. We discuss the opportunities and challenges of using LLMs to provide transparency in two important court processes: jury selection in criminal trials and housing eviction cases.
Ishana Shastri, Shomik Jain, Barbara E. Engelhardt, Ashia Wilson
AIES (1)3
2024 Active Learning for Derivative-Based Global Sensitivity Analysis with Gaussian Processes
abstract
We consider the problem of active learning for global sensitivity analysis of expensive black-box functions. Our aim is to efficiently learn the importance of different input variables, e.g., in vehicle safety experimentation, we study the impact of the thickness of various components on safety objectives. Since function evaluations are expensive, we use active learning to prioritize experimental resources where they yield the most value. We propose novel active learning acquisition functions that directly target key quantities of derivative-based global sensitivity measures (DGSMs) under Gaussian process surrogate models. We showcase the first application of active learning directly to DGSMs, and develop tractable uncertainty reduction and information gain acquisition functions for these measures. Through comprehensive evaluation on synthetic and real-world problems, our study demonstrates how these active learning acquisition strategies substantially enhance the sample efficiency of DGSM estimation, particularly with limited evaluation budgets. Our work paves the way for more efficient and accurate sensitivity analysis in various scientific and engineering applications.
Syrine Belakaria, Benjamin Letham, Janardhan Rao Doppa, Barbara E. Engelhardt, Stefano Ermon, Eytan Bakshy
NeurIPS4
2024 Answering open questions in biology using spatial genomics and structured methods
abstract
Genomics methods have uncovered patterns in a range of biological systems, but obscure important aspects of cell behavior: the shapes, relative locations, movement, and interactions of cells in space. Spatial technologies that collect genomic or epigenomic data while preserving spatial information have begun to overcome these limitations. These new data promise a deeper understanding of the factors that affect cellular behavior, and in particular the ability to directly test existing theories about cell state and variation in the context of morphology, location, motility, and signaling that could not be tested before. Rapid advancements in resolution, ease-of-use, and scale of spatial genomics technologies to address these questions also require an updated toolkit of statistical methods with which to interrogate these data. We present a framework to respond to this new avenue of research: four open biological questions that can now be answered using spatial genomics data paired with methods for analysis. We outline spatial data modalities for each open question that may yield specific insights, discuss how conflicting theories may be tested by comparing the data to conceptual models of biological behavior, and highlight statistical and machine learning-based tools that may prove particularly helpful to recover biological understanding.
Siddhartha G. Jena, Archit Verma, Barbara E. Engelhardt
BMC Bioinform.3
2022 Variance Minimization in the Wasserstein Space for Invariant Causal Prediction
abstract
Selecting powerful predictors for an outcome is a cornerstone task for machine learning. However, some types of questions can only be answered by identifying the predictors that causally affect the outcome. A recent approach to this causal inference problem leverages the invariance property of a causal mechanism across differing experimental environments (Peters et al., 2016; Heinze-Deml et al., 2018). This method, invariant causal prediction (ICP), has a substantial computational defect – the runtime scales exponentially with the number of possible causal variables. In this work, we show that the approach taken in ICP may be reformulated as a series of nonparametric tests that scales linearly in the number of predictors. Each of these tests relies on the minimization of a novel loss function – the Wasserstein variance – that is derived from tools in optimal transport theory and is used to quantify distributional variability across environments. We prove under mild assumptions that our method is able to recover the set of identifiable direct causes, and we demonstrate in our experiments that it is competitive with other benchmark causal discovery algorithms.
Guillaume Martinet, Alexander Strzalkowski, Barbara E. Engelhardt
AISTATS3
2022 A Poisson reduced-rank regression model for association mapping in sequencing data
abstract
BACKGROUND: Single-cell RNA-sequencing (scRNA-seq) technologies allow for the study of gene expression in individual cells. Often, it is of interest to understand how transcriptional activity is associated with cell-specific covariates, such as cell type, genotype, or measures of cell health. Traditional approaches for this type of association mapping assume independence between the outcome variables (or genes), and perform a separate regression for each. However, these methods are computationally costly and ignore the substantial correlation structure of gene expression. Furthermore, count-based scRNA-seq data pose challenges for traditional models based on Gaussian assumptions. RESULTS: We aim to resolve these issues by developing a reduced-rank regression model that identifies low-dimensional linear associations between a large number of cell-specific covariates and high-dimensional gene expression readouts. Our probabilistic model uses a Poisson likelihood in order to account for the unique structure of scRNA-seq counts. We demonstrate the performance of our model using simulations, and we apply our model to a scRNA-seq dataset, a spatial gene expression dataset, and a bulk RNA-seq dataset to show its behavior in three distinct analyses. CONCLUSION: We show that our statistical modeling approach, which is based on reduced-rank regression, captures associations between gene expression and cell- and sample-specific covariates by leveraging low-dimensional representations of transcriptional states.
Tiana Fitzgerald, Barbara E. Engelhardt
BMC Bioinform.3
2021 Latent variable modeling with random features
abstract
Gaussian process-based latent variable models are flexible and theoretically grounded tools for nonlinear dimension reduction, but generalizing to non-Gaussian data likelihoods within this nonlinear framework is statistically challenging. Here, we use random features to develop a family of nonlinear dimension reduction models that are easily extensible to non-Gaussian data likelihoods; we call these random feature latent variable models (RFLVMs). By approximating a nonlinear relationship between the latent space and the observations with a function that is linear with respect to random features, we induce closed-form gradients of the posterior distribution with respect to the latent variable. This allows the RFLVM framework to support computationally tractable nonlinear latent variable models for a variety of data likelihoods in the exponential family without specialized derivations. Our generalized RFLVMs produce results comparable with other state-of-the-art dimension reduction methods on diverse types of data, including neural spike train recordings, images, and text data.
Gregory W. Gundersen, Michael Minyi Zhang, Barbara E. Engelhardt
AISTATS3
2021 Active multi-fidelity Bayesian online changepoint detection
abstract
Online algorithms for detecting changepoints, or abrupt shifts in the behavior of a time series, are often deployed with limited resources, e.g., to edge computing settings such as mobile phones or industrial sensors. In these scenarios it may be beneficial to trade the cost of collecting an environmental measurement against the quality or “fidelity” of this measurement and how the measurement affects changepoint estimation. For instance, one might decide between inertial measurements or GPS to determine changepoints for motion. A Bayesian approach to changepoint detection is particularly appealing because we can represent our posterior uncertainty about changepoints and make active, cost-sensitive decisions about data fidelity to reduce this posterior uncertainty. Moreover, the total cost could be dramatically lowered through active fidelity switching, while remaining robust to changes in data distribution. We propose a multi-fidelity approach that makes cost-sensitive decisions about which data fidelity to collect based on maximizing information gain with respect to changepoints. We evaluate this framework on synthetic, video, and audio data and show that this information-based approach results in accurate predictions while reducing total cost.
Gregory W. Gundersen, Diana Cai, Chuteng Zhou, Barbara E. Engelhardt, Ryan P. Adams
UAI4
2021 Causal network inference from gene transcriptional time-series response to glucocorticoids
abstract
Gene regulatory network inference is essential to uncover complex relationships among gene pathways and inform downstream experiments, ultimately enabling regulatory network re-engineering. Network inference from transcriptional time-series data requires accurate, interpretable, and efficient determination of causal relationships among thousands of genes. Here, we develop Bootstrap Elastic net regression from Time Series (BETS), a statistical framework based on Granger causality for the recovery of a directed gene network from transcriptional time-series data. BETS uses elastic net regression and stability selection from bootstrapped samples to infer causal relationships among genes. BETS is highly parallelized, enabling efficient analysis of large transcriptional data sets. We show competitive accuracy on a community benchmark, the DREAM4 100-gene network inference challenge, where BETS is one of the fastest among methods of similar performance and additionally infers whether causal effects are activating or inhibitory. We apply BETS to transcriptional time-series data of differentially-expressed genes from A549 cells exposed to glucocorticoids over a period of 12 hours. We identify a network of 2768 genes and 31,945 directed edges (FDR ≤ 0.2). We validate inferred causal network edges using two external data sources: Overexpression experiments on the same glucocorticoid system, and genetic variants associated with inferred edges in primary lung tissue in the Genotype-Tissue Expression (GTEx) v6 project. BETS is available as an open source software package at https://github.com/lujonathanh/BETS.
Jonathan Lu, Bianca Dumitrascu, Ian C. McDowell, Brian Jo, Alejandro Barrera, Linda K. Hong, Sarah M. Leichter, Timothy E. Reddy, Barbara E. Engelhardt
PLoS Comput. Biol.9
2020 Patient-Specific Effects of Medication Using Latent Force Models with Gaussian Processes
abstract
A multi-output Gaussian process (GP) is a flexible Bayesian nonparametric framework that has proven useful in jointly modeling the physiological states of patients in medical time series data. However, capturing the short-term effects of drugs and therapeutic interventions on patient physiological state remains challenging. We propose a novel approach that models the effect of interventions as a hybrid Gaussian process composed of a GP capturing patient baseline physiology convolved with a latent force model capturing effects of treatments on specific physiological features. The combination of a multi-output GP with a time-marked kernel GP leads to a well-characterized model of patients’ physiological state across a hospital stay, including response to interventions. Our model leads to analytically tractable cross-covariance functions that allow for scalable inference. Our hierarchical model includes estimates of patient-specific effects but allows sharing of support across patients. Our approach achieves competitive predictive performance on challenging hospital data, where we recover patient-specific response to the administration of three common drugs: one antihypertensive drug and two anticoagulants.
Li-Fang Cheng, Bianca Dumitrascu, Michael Minyi Zhang, Corey Chivers, Michael Draugelis, Kai Li 0001, Barbara E. Engelhardt
AISTATS7
2020 A robust nonlinear low-dimensional manifold for single cell RNA-seq data
abstract
BACKGROUND: Modern developments in single-cell sequencing technologies enable broad insights into cellular state. Single-cell RNA sequencing (scRNA-seq) can be used to explore cell types, states, and developmental trajectories to broaden our understanding of cellular heterogeneity in tissues and organs. Analysis of these sparse, high-dimensional experimental results requires dimension reduction. Several methods have been developed to estimate low-dimensional embeddings for filtered and normalized single-cell data. However, methods have yet to be developed for unfiltered and unnormalized count data that estimate uncertainty in the low-dimensional space. We present a nonlinear latent variable model with robust, heavy-tailed error and adaptive kernel learning to estimate low-dimensional nonlinear structure in scRNA-seq data. RESULTS: Gene expression in a single cell is modeled as a noisy draw from a Gaussian process in high dimensions from low-dimensional latent positions. This model is called the Gaussian process latent variable model (GPLVM). We model residual errors with a heavy-tailed Student's t-distribution to estimate a manifold that is robust to technical and biological noise found in normalized scRNA-seq data. We compare our approach to common dimension reduction tools across a diverse set of scRNA-seq data sets to highlight our model's ability to enable important downstream tasks such as clustering, inferring cell developmental trajectories, and visualizing high throughput experiments on available experimental data. CONCLUSION: We show that our adaptive robust statistical approach to estimate a nonlinear manifold is well suited for raw, unfiltered gene counts from high-throughput sequencing technologies for visualization, exploration, and uncertainty estimation of cell states.
Archit Verma, Barbara E. Engelhardt
BMC Bioinform.2
2019 netNMF-sc: A Network Regularization Algorithm for Dimensionality Reduction and Imputation of Single-Cell Expression Data
Rebecca Elyanow, Bianca Dumitrascu, Barbara E. Engelhardt, Benjamin J. Raphael
RECOMB3
2019 End-to-end Training of Deep Probabilistic CCA on Paired Biomedical Observations
Gregory W. Gundersen, Bianca Dumitrascu, Jordan T. Ash, Barbara E. Engelhardt
UAI4
2019 Statistical tests for detecting variance effects in quantitative trait studies
abstract
Motivation: Identifying variants, both discrete and continuous, that are associated with quantitative traits, or QTs, is the primary focus of quantitative genetics. Most current methods are limited to identifying mean effects, or associations between genotype or covariates and the mean value of a quantitative trait. It is possible, however, that a variant may affect the variance of the quantitative trait in lieu of, or in addition to, affecting the trait mean. Here, we develop a general methodology to identify covariates with variance effects on a quantitative trait using a Bayesian heteroskedastic linear regression model (BTH). We compare BTH with existing methods to detect variance effects across a large range of simulations drawn from scenarios common to the analysis of quantitative traits. Results: We find that BTH and a double generalized linear model (dglm) outperform classical tests used for detecting variance effects in recent genomic studies. We show BTH and dglm are less likely to generate spurious discoveries through simulations and application to identifying methylation variance QTs and expression variance QTs. We identify four variance effects of sex in the Cardiovascular and Pharmacogenetics study. Our work is the first to offer a comprehensive view of variance identifying methodology. We identify shortcomings in previously used methodology and provide a more conservative and robust alternative. We extend variance effect analysis to a wide array of covariates that enables a new statistical dimension in the study of sex and age specific quantitative trait effects. Availability and implementation: https://github.com/b2du/bth. Supplementary information: Supplementary data are available at Bioinformatics online.
Bianca Dumitrascu, Gregory Darnell, Julien Ayroles, Barbara E. Engelhardt
Bioinform.4
2018 PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits
abstract
We address the problem of regret minimization in logistic contextual bandits, where a learner decides among sequential actions or arms given their respective contexts to maximize binary rewards. Using a fast inference procedure with Polya-Gamma distributed augmentation variables, we propose an improved version of Thompson Sampling, a Bayesian formulation of contextual bandits with near-optimal performance. Our approach, Polya-Gamma augmented Thompson Sampling (PG-TS), achieves state-of-the-art performance on simulated and real data. PG-TS explores the action space efficiently and exploits high-reward arms, quickly converging to solutions of low regret. Its explicit estimation of the posterior distribution of the context feature covariance leads to substantial empirical gains over approximate approaches. PG-TS is the first approach to demonstrate the benefits of Polya-Gamma augmentation in bandits and to propose an efficient Gibbs sampler for approximating the analytically unsolvable integral of logistic contextual bandits.
Bianca Dumitrascu, Karen Feng, Barbara E. Engelhardt
NeurIPS3
2018 How algorithmic confounding in recommendation systems increases homogeneity and decreases utility
abstract
Recommendation systems are ubiquitous and impact many domains; they have the potential to influence product consumption, individuals' perceptions of the world, and life-altering decisions. These systems are often evaluated or trained with data from users already exposed to algorithmic recommendations; this creates a pernicious feedback loop. Using simulations, we demonstrate how using data confounded in this way homogenizes user behavior without increasing utility.
Allison June-Barlow Chaney, Brandon M. Stewart, Barbara E. Engelhardt
RecSys3
2018 Clustering gene expression time series data using an infinite Gaussian process mixture model
abstract
Transcriptome-wide time series expression profiling is used to characterize the cellular response to environmental perturbations. The first step to analyzing transcriptional response data is often to cluster genes with similar responses. Here, we present a nonparametric model-based method, Dirichlet process Gaussian process mixture model (DPGP), which jointly models data clusters with a Dirichlet process and temporal dependencies with Gaussian processes. We demonstrate the accuracy of DPGP in comparison to state-of-the-art approaches using hundreds of simulated data sets. To further test our method, we apply DPGP to published microarray data from a microbial model organism exposed to stress and to novel RNA-seq data from a human cell line exposed to the glucocorticoid dexamethasone. We validate our clusters by examining local transcription factor binding and histone modifications. Our results demonstrate that jointly modeling cluster number and temporal dependencies can reveal shared regulatory mechanisms. DPGP software is freely available online at https://github.com/PrincetonUniversity/DP_GP_cluster.
Ian C. McDowell, Dinesh Manandhar, Christopher M. Vockley, Amy K. Schmid, Timothy E. Reddy, Barbara E. Engelhardt
PLoS Comput. Biol.6
2017 Dynamic Collaborative Filtering With Compound Poisson Factorization
abstract
Model-based collaborative filtering (CF) analyzes user–item interactions to infer latent factors that represent user preferences and item characteristics in order to predict future interactions. Most CF approaches assume that these latent factors are static; however, in most CF data, user preferences and item perceptions drift over time. Here, we propose a new conjugate and numerically stable dynamic matrix factorization (DCPF) based on hierarchical Poisson factorization that models the smoothly drifting latent factors using gamma-Markov chains. We propose a conjugate gamma chain construction that is numerically stable within our compound-Poisson framework. We then derive a scalable stochastic variational inference approach to estimate the parameters of our model. We apply our model to time-stamped ratings data sets from Netflix, Yelp, and Last.fm. We empirically demonstrate that DCPF achieves a higher predictive accuracy than state-of-the-art static and dynamic factorization algorithms.
Ghassen Jerfel, Mehmet Emin Basbug, Barbara E. Engelhardt
AISTATS3
2017 A Reinforcement Learning Approach to Weaning of Mechanical Ventilation in Intensive Care Units
Niranjani Prasad, Li-Fang Cheng, Corey Chivers, Michael Draugelis, Barbara E. Engelhardt
UAI5
2017 Adaptive Randomized Dimension Reduction on Massive Data
abstract
The scalability of statistical estimators is of increasing importance in modern applications. One approach to implementing scalable algorithms is to compress data into a low dimensional latent space using dimension reduction methods. In this paper, we develop an approach for dimension reduction that exploits the assumption of low rank structure in high dimensional data to gain both computational and statistical advantages. We adapt recent randomized low-rank approximation algorithms to provide an efficient solution to principal component analysis (PCA), and we use this efficient solver to improve estimation in large- scale linear mixed models (LMM) for association mapping in statistical genomics. A key observation in this paper is that randomization serves a dual role, improving both computational and statistical performance by implicitly regularizing the covariance matrix estimate of the random effect in an LMM. These statistical and computational advantages are highlighted in our experiments on simulated data and large-scale genomic studies.
Gregory Darnell, Stoyan Georgiev, Sayan Mukherjee 0001, Barbara E. Engelhardt
J. Mach. Learn. Res.4
2016 Hierarchical Compound Poisson Factorization
abstract
Non-negative matrix factorization models based on a hierarchical Gamma-Poisson structure capture user and item behavior effectively in extremely sparse data sets, making them the ideal choice for collaborative filtering applications. Hierarchical Poisson factorization (HPF) in particular has proved successful for scalable recommendation systems with extreme sparsity. HPF, however, suffers from a tight coupling of sparsity model (absence of a rating) and response model (the value of the rating), which limits the expressiveness of the latter. Here, we introduce hierarchical compound Poisson factorization (HCPF) that has the favorable Gamma-Poisson structure and scalability of HPF to high-dimensional extremely sparse matrices. More importantly, HCPF decouples the sparsity model from the response model, allowing us to choose the most suitable distribution for the response. HCPF can capture binary, non-negative discrete, non-negative continuous, and zero-inflated continuous responses. We compare HCPF with HPF on nine discrete and three continuous data sets and conclude that HCPF captures the relationship between sparsity and response better than HPF.
Mehmet Emin Basbug, Barbara E. Engelhardt
ICML2
2016 Bayesian group factor analysis with structured sparsity
abstract
Latent factor models are the canonical statistical tool for exploratory analyses of low-dimensional linear structure for a matrix of $p$ features across $n$ samples. We develop a structured Bayesian group factor analysis model that extends the factor model to multiple coupled observation matrices; in the case of two observations, this reduces to a Bayesian model of canonical correlation analysis. Here, we carefully define a structured Bayesian prior that encourages both element-wise and column-wise shrinkage and leads to desirable behavior on high- dimensional data. In particular, our model puts a structured prior on the joint factor loading matrix, regularizing at three levels, which enables element-wise sparsity and unsupervised recovery of latent factors corresponding to structured variance across arbitrary subsets of the observations. In addition, our structured prior allows for both dense and sparse latent factors so that covariation among either all features or only a subset of features can be recovered. We use fast parameter-expanded expectation-maximization for parameter estimation in this model. We validate our method on simulated data with substantial structure. We show results of our method applied to three high- dimensional data sets, comparing results against a number of state-of-the-art approaches. These results illustrate useful properties of our model, including i) recovering sparse signal in the presence of dense effects; ii) the ability to scale naturally to large numbers of observations; iii) flexible observation- and factor-specific regularization to recover factors with a wide variety of sparsity levels and percentage of variance explained; and iv) tractable inference that scales to modern genomic and text data sizes.
Shiwen Zhao, Chuan Gao, Sayan Mukherjee 0001, Barbara E. Engelhardt
J. Mach. Learn. Res.4
2016 Context Specific and Differential Gene Co-expression Networks via Bayesian Biclustering
abstract
Identifying latent structure in high-dimensional genomic data is essential for exploring biological processes. Here, we consider recovering gene co-expression networks from gene expression data, where each network encodes relationships between genes that are co-regulated by shared biological mechanisms. To do this, we develop a Bayesian statistical model for biclustering to infer subsets of co-regulated genes that covary in all of the samples or in only a subset of the samples. Our biclustering method, BicMix, allows overcomplete representations of the data, computational tractability, and joint modeling of unknown confounders and biological signals. Compared with related biclustering methods, BicMix recovers latent structure with higher precision across diverse simulation scenarios as compared to state-of-the-art biclustering methods. Further, we develop a principled method to recover context specific gene co-expression networks from the estimated sparse biclustering matrices. We apply BicMix to breast cancer gene expression data and to gene expression data from a cardiovascular study cohort, and we recover gene co-expression networks that are differential across ER+ and ER- samples and across male and female samples. We apply BicMix to the Genotype-Tissue Expression (GTEx) pilot data, and we find tissue specific gene networks. We validate these findings by using our tissue specific networks to identify trans-eQTLs specific to one of four primary tissues.
Chuan Gao, Ian C. McDowell, Shiwen Zhao, Christopher D. Brown, Barbara E. Engelhardt
PLoS Comput. Biol.5
2013 Stability selection for regression-based models of transcription factor-DNA binding specificity
abstract
MOTIVATION: The DNA binding specificity of a transcription factor (TF) is typically represented using a position weight matrix model, which implicitly assumes that individual bases in a TF binding site contribute independently to the binding affinity, an assumption that does not always hold. For this reason, more complex models of binding specificity have been developed. However, these models have their own caveats: they typically have a large number of parameters, which makes them hard to learn and interpret. RESULTS: We propose novel regression-based models of TF-DNA binding specificity, trained using high resolution in vitro data from custom protein-binding microarray (PBM) experiments. Our PBMs are specifically designed to cover a large number of putative DNA binding sites for the TFs of interest (yeast TFs Cbf1 and Tye7, and human TFs c-Myc, Max and Mad2) in their native genomic context. These high-throughput quantitative data are well suited for training complex models that take into account not only independent contributions from individual bases, but also contributions from di- and trinucleotides at various positions within or near the binding sites. To ensure that our models remain interpretable, we use feature selection to identify a small number of sequence features that accurately predict TF-DNA binding specificity. To further illustrate the accuracy of our regression models, we show that even in the case of paralogous TF with highly similar position weight matrices, our new models can distinguish the specificities of individual factors. Thus, our work represents an important step toward better sequence-based models of individual TF-DNA binding specificity. AVAILABILITY: Our code is available at http://genome.duke.edu/labs/gordan/ISMB2013. The PBM data used in this article are available in the Gene Expression Omnibus under accession number GSE47026.
Fantine Mordelet, John Horton, Alexander J. Hartemink, Barbara E. Engelhardt, Raluca Gordân
Bioinform.4
2006 A graphical model for predicting protein molecular function
abstract
We present a simple statistical model of molecular function evolution to predict protein function. The model description encodes general knowledge of how molecular function evolves within a phylogenetic tree based on the proteins' sequence. Inputs are a phylogeny for a set of evolutionarily related protein sequences and any available function characterizations for those proteins. Posterior probabilities for each protein are used to predict the molecular function of that protein. We present results from applying our model to three protein families, and compare our prediction results on the extant proteins to other available protein function prediction methods. For the deaminase family, our method achieves 93.9% where related methods BLAST achieves 72.7%, GOtcha achieves 87.9%, and Orthostrapper achieves 72.7% in prediction accuracy.
Barbara E. Engelhardt, Michael I. Jordan, Steven E. Brenner
ICML1
2005 Protein Molecular Function Prediction by Bayesian Phylogenomics
abstract
We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.
Barbara E. Engelhardt, Michael I. Jordan, Kathryn E. Muratore, Steven E. Brenner
PLoS Comput. Biol.1
2003 Factored Planning
Eyal Amir, Barbara E. Engelhardt
IJCAI2
2001 The RadarSAT-MAMM Automated Mission Planner
Benjamin D. Smith, Barbara E. Engelhardt, Darren H. Mutz
IAAI2
2001 Balancing deliberation and reaction, planning and execution for space robotic applications
abstract
Intelligent behavior for robotic agents requires a careful balance of fast reactions and deliberate consideration of long-term ramifications. The need for this balance is particularly acute in space applications, where hostile environments demand fast reactions, and remote locations dictate careful management of consumables that cannot be replenished. However, fast reactions typically require procedural representations with limited scope and handling long-term considerations in a general fashion is often computationally expensive. We describe three major areas for autonomous systems for space exploration: free-flying spacecraft, planetary rovers, and ground communications stations. In each of these broad applications areas, we identify operational considerations requiring rapid response and considerations of long-term ramifications. We describe these issues in the context of ongoing efforts to deploy autonomous systems using planning and task execution systems.
Russell Knight, Forest Fisher, Tara A. Estlin, Barbara E. Engelhardt, Steve A. Chien
IROS4