Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jeffrey W. Miller

dblp:28/9509 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-7718-1581ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 84% Efficient and distributed learning · 9% Learning theory · 7%
Databases, data mining, and information retrieval
2 papers
Data mining · 89% Data integration and cleaning · 11%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
1.032023
Consistent Model-based Clustering using the Quasi-Bernoulli Stick-breaking Process · J. Mach. Learn. Res. 2023
Inconsistency of Pitman-Yor process mixtures for the number of components · J. Mach. Learn. Res. 2014
A simple example of Dirichlet process mixture inconsistency for the number of components · NIPS 2013
Bioinformatics and computational biology
feature selection
0.912025
Nonparametric IPSS: fast, flexible feature selection with false discovery control · Bioinform. 2025
Data mining › statistical analysis
false discovery rate control
0.912025
Nonparametric IPSS: fast, flexible feature selection with false discovery control · Bioinform. 2025
Data mining › dimensionality reduction
feature selection
0.912025
Nonparametric IPSS: fast, flexible feature selection with false discovery control · Bioinform. 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection
0.712023
Bayesian Data Selection · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning
clustering
0.712023
Consistent Model-based Clustering using the Quasi-Bernoulli Stick-breaking Process · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning
data selection
0.712023
Bayesian Data Selection · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › clustering
model-based clustering
0.712023
Consistent Model-based Clustering using the Quasi-Bernoulli Stick-breaking Process · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › divergence measure
stein discrepancy
0.712023
Bayesian Data Selection · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
stick-breaking process
0.712023
Consistent Model-based Clustering using the Quasi-Bernoulli Stick-breaking Process · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.632016
Flexible Models for Microclustering with Application to Entity Resolution · NIPS 2016
Inconsistency of Pitman-Yor process mixtures for the number of components · J. Mach. Learn. Res. 2014
A simple example of Dirichlet process mixture inconsistency for the number of components · NIPS 2013
Machine learning › Learning theory › statistical estimation › asymptotic estimation theory
asymptotic normality
0.512021
Asymptotic Normality, Concentration, and Coverage of Generalized Posteriors · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.512021
Asymptotic Normality, Concentration, and Coverage of Generalized Posteriors · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior concentration
0.512021
Asymptotic Normality, Concentration, and Coverage of Generalized Posteriors · J. Mach. Learn. Res. 2021
Data mining
clustering
0.212016
Flexible Models for Microclustering with Application to Entity Resolution · NIPS 2016
Data integration and cleaning
entity resolution
0.212016
Flexible Models for Microclustering with Application to Entity Resolution · NIPS 2016
Bioinformatics and computational biology › single-cell analysis
single-cell RNA sequencing
0.212023
Bayesian Data Selection · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
pitman-yor process
0.212014
Inconsistency of Pitman-Yor process mixtures for the number of components · J. Mach. Learn. Res. 2014
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.212013
A simple example of Dirichlet process mixture inconsistency for the number of components · NIPS 2013
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior consistency
0.012013
A simple example of Dirichlet process mixture inconsistency for the number of components · NIPS 2013

Methods — techniques the papers use, named apart from their topics

random forest · 1.7integrated path stability selection · 1.7gradient boosting · 1.7probabilistic principal components analysis · 1.3kernelized stein discrepancy · 1.3gibbs sampling · 0.7dirichlet process · 0.7pitman-yor process mixture model · 0.5laplace approximation · 0.5dirichlet process mixture model · 0.5bernstein-von mises · 0.5asymptotic analysis · 0.2
YearPublicationVenuePosition
2025 Nonparametric IPSS: fast, flexible feature selection with false discovery control
abstract
MOTIVATION: Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery control, or (iii) identify few true positives. RESULTS: We introduce a general feature selection method with finite-sample false discovery control based on applying integrated path stability selection (IPSS) to arbitrary feature importance scores. The method is nonparametric whenever the importance scores are nonparametric, and it estimates q-values, which are better suited to high-dimensional data than P-values. We focus on two special cases using importance scores from gradient boosting (IPSSGB) and random forests (IPSSRF). Extensive nonlinear simulations with RNA sequencing data show that both methods accurately control the false discovery rate and detect more true positives than existing methods. Both methods are also efficient, running in under 20 s when there are 500 samples and 5000 features. We apply IPSSGB and IPSSRF to detect microRNAs and genes related to cancer, finding that they yield better predictions with fewer features than existing approaches. AVAILABILITY AND IMPLEMENTATION: All code and data used in this work are available on GitHub (https://github.com/omelikechi/ipss_bioinformatics) and permanently archived on Zenodo (https://doi.org/10.5281/zenodo.15335289). A Python package for implementing IPSS is available on GitHub (https://github.com/omelikechi/ipss) and PyPI (https://pypi.org/project/ipss/). An R implementation of IPSS is also available on GitHub (https://github.com/omelikechi/ipssR).
Omar Melikechi, David B. Dunson, Jeffrey W. Miller
Bioinform.3
2023 Bayesian Data Selection
abstract
Insights into complex, high-dimensional data can be obtained by discovering features of the data that match or do not match a model of interest. To formalize this task, we introduce the "data selection" problem: finding a lower-dimensional statistic - such as a subset of variables - that is well fit by a given parametric model of interest. A fully Bayesian approach to data selection would be to parametrically model the value of the statistic, nonparametrically model the remaining "background" components of the data, and perform standard Bayesian model selection for the choice of statistic. However, fitting a nonparametric model to high-dimensional data tends to be highly inefficient, statistically and computationally. We propose a novel score for performing data selection, the "Stein volume criterion (SVC)", that does not require fitting a nonparametric model. The SVC takes the form of a generalized marginal likelihood with a kernelized Stein discrepancy in place of the Kullback-Leibler divergence. We prove that the SVC is consistent for data selection, and establish consistency and asymptotic normality of the corresponding generalized posterior on parameters. We apply the SVC to the analysis of single-cell RNA sequencing data sets using probabilistic principal components analysis and a spin glass model of gene regulation.
Eli N. Weinstein, Jeffrey W. Miller
J. Mach. Learn. Res.2
2023 Consistent Model-based Clustering using the Quasi-Bernoulli Stick-breaking Process
abstract
In mixture modeling and clustering applications, the number of components and clusters is often not known. A stick-breaking mixture model, such as the Dirichlet process mixture model, is an appealing construction that assumes infinitely many components, while shrinking the weights of most of the unused components to near zero. However, it is well-known that this shrinkage is inadequate: even when the component distribution is correctly specified, spurious weights appear and give an inconsistent estimate of the number of clusters. In this article, we propose a simple solution: when breaking each mixture weight stick into two pieces, the length of the second piece is multiplied by a quasi-Bernoulli random variable, taking value one or a small constant close to zero. This effectively creates a soft truncation and further shrinks the unused weights. Asymptotically, we show that as long as this small constant diminishes to zero at a rate faster than $o(1/n^2)$, with $n$ the sample size and given data from a finite mixture model, the posterior distribution will converge to the true number of clusters. In comparison, we rigorously explore Dirichlet process mixture models using a concentration parameter that is either constant or rapidly diminishes to zero---both of which lead to inconsistency for the number of clusters. Our proposed model is easy to implement, requiring only a small modification of a standard Gibbs sampler for mixture models. In simulations and a data application of clustering brain networks, our proposed method recovers the ground-truth number of clusters, and leads to a small number of clusters.
Jeffrey W. Miller, Leo L. Duan
J. Mach. Learn. Res.2
2021 Asymptotic Normality, Concentration, and Coverage of Generalized Posteriors
abstract
Generalized likelihoods are commonly used to obtain consistent estimators with attractive computational and robustness properties. Formally, any generalized likelihood can be used to define a generalized posterior distribution, but an arbitrarily defined "posterior" cannot be expected to appropriately quantify uncertainty in any meaningful sense. In this article, we provide sufficient conditions under which generalized posteriors exhibit concentration, asymptotic normality (Bernstein-von Mises), an asymptotically correct Laplace approximation, and asymptotically correct frequentist coverage. We apply our results in detail to generalized posteriors for a wide array of generalized likelihoods, including pseudolikelihoods in general, the Gaussian Markov random field pseudolikelihood, the fully observed Boltzmann machine pseudolikelihood, the Ising model pseudolikelihood, the Cox proportional hazards partial likelihood, and a median-based likelihood for robust inference of location. Further, we show how our results can be used to easily establish the asymptotics of standard posteriors for exponential families and generalized linear models. We make no assumption of model correctness so that our results apply with or without misspecification.
Jeffrey W. Miller
J. Mach. Learn. Res.1
2020 Identifying longevity associated genes by integrating gene expression and curated annotations
abstract
Aging is a complex process with poorly understood genetic mechanisms. Recent studies have sought to classify genes as pro-longevity or anti-longevity using a variety of machine learning algorithms. However, it is not clear which types of features are best for optimizing classification performance and which algorithms are best suited to this task. Further, performance assessments based on held-out test data are lacking. We systematically compare five popular classification algorithms using gene ontology and gene expression datasets as features to predict the pro-longevity versus anti-longevity status of genes for two model organisms (C. elegans and S. cerevisiae) using the GenAge database as ground truth. We find that elastic net penalized logistic regression performs particularly well at this task. Using elastic net, we make novel predictions of pro- and anti-longevity genes that are not currently in the GenAge database.
F. William Townes, Kareem Carr, Jeffrey W. Miller
PLoS Comput. Biol.3
2016 Flexible Models for Microclustering with Application to Entity Resolution
abstract
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable clustering models. However, for some applications, this assumption is inappropriate. For example, when performing entity resolution, the size of each cluster should be unrelated to the size of the data set, and each cluster should contain a negligible fraction of the total number of data points. These applications require models that yield clusters whose sizes grow sublinearly with the size of the data set. We address this requirement by defining the microclustering property and introducing a new class of models that can exhibit this property. We compare models within this class to two commonly used clustering models using four entity-resolution data sets.
Brenda Betancourt, Giacomo Zanella, Jeffrey W. Miller, Hanna M. Wallach, Abbas Zaidi, Rebecca C. Steorts
NIPS3
2014 Inconsistency of Pitman-Yor process mixtures for the number of components
Jeffrey W. Miller, Matthew T. Harrison
J. Mach. Learn. Res.1
2013 A simple example of Dirichlet process mixture inconsistency for the number of components
abstract
For data assumed to come from a finite mixture with an unknown number of components, it has become common to use Dirichlet process mixtures (DPMs) not only for density estimation, but also for inferences about the number of components. The typical approach is to use the posterior distribution on the number of components occurring so far --- that is, the posterior on the number of clusters in the observed data. However, it turns out that this posterior is not consistent --- it does not converge to the true number of components. In this note, we give an elementary demonstration of this inconsistency in what is perhaps the simplest possible setting: a DPM with normal components of unit variance, applied to data from a mixture" with one standard normal component. Further, we find that this example exhibits severe inconsistency: instead of going to 1, the posterior probability that there is one cluster goes to 0."
Jeffrey W. Miller, Matthew T. Harrison
NIPS1