Brent A. Coull

dblp:74/266 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-1808-4156ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 42% Probabilistic and Bayesian machine learning · 31% Kernel, tree and ensemble methods · 21%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.612022
Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical Guarantees · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability
feature importance
0.612022
Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical Guarantees · NeurIPS 2022
Machine learning › Trustworthy machine learning
interpretability
0.612022
Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical Guarantees · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior consistency
0.612022
Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical Guarantees · NeurIPS 2022
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.412019
Accurate Uncertainty Estimation and Decomposition in Ensemble Learning · NeurIPS 2019
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty decomposition
0.412019
Accurate Uncertainty Estimation and Decomposition in Ensemble Learning · NeurIPS 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412019
Accurate Uncertainty Estimation and Decomposition in Ensemble Learning · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.312017
Robust Hypothesis Test for Nonlinear Effect with Gaussian Processes · NIPS 2017
Machine learning › Learning theory
hypothesis testing
0.312017
Robust Hypothesis Test for Nonlinear Effect with Gaussian Processes · NIPS 2017
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.312017
Robust Hypothesis Test for Nonlinear Effect with Gaussian Processes · NIPS 2017
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space
0.312017
Robust Hypothesis Test for Nonlinear Effect with Gaussian Processes · NIPS 2017

Methods — techniques the papers use, named apart from their topics

integrated partial derivative · 0.6bayesian nonparametric theory · 0.6ensemble learning · 0.4bayesian nonparametric modeling · 0.4score test · 0.3ensemble estimator · 0.3
YearPublicationVenuePosition
2025 Group-wise normalization in differential abundance analysis of microbiome samples
abstract
BACKGROUND: A key challenge in differential abundance analysis (DAA) of microbial sequencing data is that the counts for each sample are compositional, resulting in potentially biased comparisons of the absolute abundance across study groups. Normalization-based DAA methods rely on external normalization factors that account for compositionality by standardizing the counts onto a common numerical scale. However, existing normalization methods have struggled to maintain the false discovery rate in settings where the variance or compositional bias is large. This article proposes a novel framework for normalization that can reduce bias in DAA by re-conceptualizing normalization as a group-level task. We present two new normalization methods within the group-wise framework: group-wise relative log expression (G-RLE) and fold-truncated sum scaling (FTSS). RESULTS: G-RLE and FTSS achieve higher statistical power for identifying differentially abundant taxa than existing methods in model-based and synthetic data simulation settings. The two novel methods also maintain the false discovery rate in challenging scenarios where existing methods suffer. The best results are obtained from using FTSS normalization with the DAA method MetagenomeSeq. CONCLUSION: Compared with other methods for normalizing compositional sequence count data prior to DAA, the proposed group-level normalization frameworks offer more robust statistical inference. With a solid mathematical foundation, validated performance in numerical studies, and publicly available software, these new methods can help improve rigor and reproducibility in microbiome research.
Dylan Clark-Boucher, Brent A. Coull, Harrison T. Reeder, Fenglei Wang, Jacqueline R. Starr, Kyu Ha Lee
BMC Bioinform.2
2022 Bayesian Nonparametric Model Averaging Using Scalable Gaussian Process Representations
abstract
Bayesian nonparametric methods provide a convenient and well-founded framework for constructing spatio-temporally evolving ensemble models. They not only provide a flexible way to weight different underlying models in an ensemble according to the strengths of each model on different regions of the input space, but also allow for quantification of the ensemble’s uncertainty. However, computational costs can pose a challenge to kernel-based ensemble methods when spatio-temporal resolution is high, such as environmental models. We propose a Bayesian Nonparametric Ensemble (BNE) method for spatio-temporal ensemble learning based on the Gaussian process. Our ensemble relies on theoretically well-founded linearized approximations to the Gaussian process to adaptively weight the underlying models in a way that is scalable and amenable to stochastic learning methods. We investigate both the random Fourier feature and Nystrom approaches in this setting. We demonstrate the practicality and usefulness of the approximate model on the problem of air pollution prediction across the contiguous USA over a 6 year period.
John W. Paisley, Sebastian Rowland, Jeremiah Z. Liu, Brent A. Coull, Marianthi-Anna Kioumourtzoglou
IEEE Big Data4
2022 Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical Guarantees
abstract
We develop a simple and unified framework for nonlinear variable importance estimation that incorporates uncertainty in the prediction function and is compatible with a wide range of machine learning models (e.g., tree ensembles, kernel methods, neural networks, etc). In particular, for a learned nonlinear model $f(\mathbf{x})$, we consider quantifying the importance of an input variable $\mathbf{x}^j$ using the integrated partial derivative $\Psi_j = \Vert \frac{\partial}{\partial \mathbf{x}^j} f(\mathbf{x})\Vert^2_{P_\mathcal{X}}$. We then (1) provide a principled approach for quantifying uncertainty in variable importance by deriving its posterior distribution, and (2) show that the approach is generalizable even to non-differentiable models such as tree ensembles. Rigorous Bayesian nonparametric theorems are derived to guarantee the posterior consistency and asymptotic uncertainty of the proposed approach. Extensive simulations and experiments on healthcare benchmark datasets confirm that the proposed algorithm outperforms existing classical and recent variable selection methods.
Wenying Deng, Beau Coker, Rajarshi Mukherjee, Jeremiah Z. Liu, Brent A. Coull
NeurIPS5
2019 Accurate Uncertainty Estimation and Decomposition in Ensemble Learning
abstract
Ensemble learning is a standard approach to building machine learning systems that capture complex phenomena in real-world data. An important aspect of these systems is the complete and valid quantification of model uncertainty. We introduce a Bayesian nonparametric ensemble (BNE) approach that augments an existing ensemble model to account for different sources of model uncertainty. BNE augments a model’s prediction and distribution functions using Bayesian nonparametric machinery. It has a theoretical guarantee in that it robustly estimates the uncertainty patterns in the data distribution, and can decompose its overall predictive uncertainty into distinct components that are due to different sources of noise and error. We show that our method achieves accurate uncertainty estimates under complex observational noise, and illustrate its real-world utility in terms of uncertainty decomposition and model bias detection for an ensemble in predict air pollution exposures in Eastern Massachusetts, USA.
Jeremiah Z. Liu, John W. Paisley, Marianthi-Anna Kioumourtzoglou, Brent A. Coull
NeurIPS4
2017 Robust Hypothesis Test for Nonlinear Effect with Gaussian Processes
abstract
This work constructs a hypothesis test for detecting whether an data-generating function $h: \real^p \rightarrow \real$ belongs to a specific reproducing kernel Hilbert space $\mathcal{H}_0$, where the structure of $\mathcal{H}_0$ is only partially known. Utilizing the theory of reproducing kernels, we reduce this hypothesis to a simple one-sided score test for a scalar parameter, develop a testing procedure that is robust against the mis-specification of kernel functions, and also propose an ensemble-based estimator for the null model to guarantee test performance in small samples. To demonstrate the utility of the proposed method, we apply our test to the problem of detecting nonlinear interaction between groups of continuous features. We evaluate the finite-sample performance of our test under different data-generating functions and estimation strategies for the null model. Our results revealed interesting connection between notions in machine learning (model underfit/overfit) and those in statistical inference (i.e. Type I error/power of hypothesis test), and also highlighted unexpected consequences of common model estimating strategies (e.g. estimating kernel hyperparameters using maximum likelihood estimation) on model inference.
Jeremiah Z. Liu, Brent A. Coull
NIPS2