Andrew C. Miller

dblp:190/7517 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Probabilistic and Bayesian machine learning · 50% Trustworthy machine learning · 16% Language models and text generation · 15%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Medical and health informatics · 52% Computational science and engineering · 43% Computational social science and digital humanities · 5%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.142018
Semi-Amortized Variational Autoencoders · ICML 2018
Reducing Reparameterization Gradient Variance · NIPS 2017
Variational Boosting: Iteratively Refining Posterior Approximations · ICML 2017
Machine learning › Trustworthy machine learning
interpretability
1.022025
Do LLMs "know" internally when they follow instructions? · ICLR 2025
Discriminative Regularization for Latent Variable Models with Applications to Electrocardiography · ICML 2019
Natural language and speech › Language models and text generation
instruction following
0.912025
Do LLMs "know" internally when they follow instructions? · ICLR 2025
Machine learning › Deep learning architectures and training
foundation model
0.812024
Large-scale Training of Foundation Models for Wearable Biosignals · ICLR 2024
Computational science and engineering
astronomy
0.422015
A Gaussian Process Model of Quasar Spectral Energy Distributions · NIPS 2015
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412019
Discriminative Regularization for Latent Variable Models with Applications to Electrocardiography · ICML 2019
Medical and health informatics
electrocardiography
0.412019
Discriminative Regularization for Latent Variable Models with Applications to Electrocardiography · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
amortized variational inference
0.312018
Semi-Amortized Variational Autoencoders · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
stochastic variational inference
0.312018
Semi-Amortized Variational Autoencoders · ICML 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Semi-Amortized Variational Autoencoders · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.312017
Variational Boosting: Iteratively Refining Posterior Approximations · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
boosting variational inference
0.312017
Variational Boosting: Iteratively Refining Posterior Approximations · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
control variates
0.312017
Reducing Reparameterization Gradient Variance · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
reparameterization gradient
0.312017
Reducing Reparameterization Gradient Variance · NIPS 2017
Machine learning › Optimization for machine learning
variance reduction
0.312017
Reducing Reparameterization Gradient Variance · NIPS 2017
Medical and health informatics › digital health
digital biomarkers
0.212024
Large-scale Training of Foundation Models for Wearable Biosignals · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.212015
A Gaussian Process Model of Quasar Spectral Energy Distributions · NIPS 2015
Natural language and speech › Language models and text generation
generative inference
0.212015
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Data mining
dimensionality reduction
0.212014
Factorized Point Process Intensities: A Spatial Analysis of Professional Basketball · ICML 2014
Data mining › dimensionality reduction
nonnegative matrix factorization
0.212014
Factorized Point Process Intensities: A Spatial Analysis of Professional Basketball · ICML 2014
Data mining › probabilistic model
point process
0.212014
Factorized Point Process Intensities: A Spatial Analysis of Professional Basketball · ICML 2014
Data mining › structured data mining
spatial data mining
0.212014
Factorized Point Process Intensities: A Spatial Analysis of Professional Basketball · ICML 2014
Machine learning › Trustworthy machine learning › interpretability
model explanation
0.112019
Discriminative Regularization for Latent Variable Models with Applications to Electrocardiography · ICML 2019
Computational social science and digital humanities
sports analytics
0.112014
Factorized Point Process Intensities: A Spatial Analysis of Professional Basketball · ICML 2014

Methods — techniques the papers use, named apart from their topics

stochastic augmentation · 2.3self-supervised learning · 2.3momentum training · 2.3contrastive learning · 2.3representation analysis · 0.9prompt engineering · 0.9generative modeling · 0.8discriminative regularization · 0.8variational inference · 0.7gradient-based optimization · 0.3poisson model · 0.2gaussian process · 0.2bayesian inference · 0.2point process · 0.2non-negative matrix factorization · 0.2
YearPublicationVenuePosition
2025 Do LLMs "know" internally when they follow instructions?
abstract
Instruction-following is crucial for building AI agents with large language models (LLMs), as these models must adhere strictly to user-provided constraints and guidelines. However, LLMs often fail to follow even simple and clear instructions. To improve instruction-following behavior and prevent undesirable outputs, a deeper understanding of how LLMs' internal states relate to these outcomes is required. In this work, we investigate whether LLMs encode information in their representations that correlates with instruction-following success—a property we term ``knowing internally''. Our analysis identifies a direction in the input embedding space, termed the instruction-following dimension, that predicts whether a response will comply with a given instruction. We find that this dimension generalizes well across unseen tasks but not across unseen instruction types. We demonstrate that modifying representations along this dimension improves instruction-following success rates compared to random changes, without compromising response quality. Further investigation reveals that this dimension is more closely related to the phrasing of prompts rather than the inherent difficulty of the task or instructions. This discovery also suggests explanations for why LLMs sometimes fail to follow clear instructions and why prompt engineering is often effective, even when the content remains largely unchanged. This work provides insight into the internal workings of LLMs' instruction-following, paving the way for reliable LLM agents.
Juyeon Heo, Christina Heinze-Deml, Oussama Elachqar, Kwan Ho Ryan Chan, Shirley You Ren, Andrew C. Miller, Udhyakumar Nallasamy, Jaya Narain
ICLR6
2024 Large-scale Training of Foundation Models for Wearable Biosignals
abstract
Tracking biosignals is crucial for monitoring wellness and preempting the development of severe medical conditions. Today, wearable devices can conveniently record various biosignals, creating the opportunity to monitor health status without disruption to one's daily routine. Despite widespread use of wearable devices and existing digital biomarkers, the absence of curated data with annotated medical labels hinders the development of new biomarkers to measure common health conditions. In fact, medical datasets are usually small in comparison to other domains, which is an obstacle for developing neural network models for biosignals. To address this challenge, we have employed self-supervised learning using the unlabeled sensor data collected under informed consent from the large longitudinal Apple Heart and Movement Study (AHMS) to train foundation models for two common biosignals: photoplethysmography (PPG) and electrocardiogram (ECG) recorded on Apple Watch. We curated PPG and ECG datasets from AHMS that include data from ${\sim} 141$K participants spanning ${\sim} 3$ years. Our self-supervised learning framework includes participant level positive pair selection, stochastic augmentation module and a regularized contrastive loss optimized with momentum training, and generalizes well to both PPG and ECG modalities. We show that the pre-trained foundation models readily encode information regarding participants' demographics and health conditions. To the best of our knowledge, this is the first study that builds foundation models using large-scale PPG and ECG data collected via wearable consumer devices $\textendash$ prior works have commonly used smaller-size datasets collected in clinical and experimental settings. We believe PPG and ECG foundation models can enhance future wearable devices by reducing the reliance on labeled data and hold the potential to help the users improve their health.
Salar Abbaspourazad, Oussama Elachqar, Andrew C. Miller, Saba Emrani, Udhyakumar Nallasamy, Ian Shapiro
ICLR3
2019 Discriminative Regularization for Latent Variable Models with Applications to Electrocardiography
abstract
Generative models often use latent variables to represent structured variation in high-dimensional data, such as images and medical waveforms. However, these latent variables may ignore subtle, yet meaningful features in the data. Some features may predict an outcome of interest (e.g. heart attack) but account for only a small fraction of variation in the data. We propose a generative model training objective that uses a black-box discriminative model as a regularizer to learn representations that preserve this predictive variation. With these discriminatively regularized latent variable models, we visualize and measure variation in the data that influence a black-box predictive model, enabling an expert to better understand each prediction. With this technique, we study models that use electrocardiograms to predict outcomes of clinical interest. We measure our approach on synthetic and real data with statistical summaries and an experiment carried out by a physician.
Andrew C. Miller, Ziad Obermeyer, John P. Cunningham, Sendhil Mullainathan
ICML1
2018 Semi-Amortized Variational Autoencoders
abstract
Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal variational parameters. We propose a hybrid approach, to use AVI to initialize the variational parameters and run stochastic variational inference (SVI) to refine them. Crucially, the local SVI procedure is itself differentiable, so the inference network and generative model can be trained end-to-end with gradient-based optimization. This semi-amortized approach enables the use of rich generative models without experiencing the posterior-collapse phenomenon common in training VAEs for problems like text generation. Experiments show this approach outperforms strong autoregressive and variational baselines on standard text and image datasets.
Sam Wiseman, Andrew C. Miller, David A. Sontag, Alexander M. Rush
ICML3
2017 Bayesian Learning and Inference in Recurrent Switching Linear Dynamical Systems
abstract
Many natural systems, such as neurons firing in the brain or basketball teams traversing a court, give rise to time series data with complex, nonlinear dynamics. We can gain insight into these systems by decomposing the data into segments that are each explained by simpler dynamic units. Building on switching linear dynamical systems (SLDS), we develop a model class and Bayesian inference algorithms that not only discover these dynamical units but also, by learning how transition probabilities depend on observations or continuous latent states, explain their switching behavior. Our key innovation is to design these recurrent SLDS models to enable recent Pólya-gamma auxiliary variable techniques and thus make approximate Bayesian learning and inference in these models easy, fast, and scalable.
Scott W. Linderman, Matthew J. Johnson 0002, Andrew C. Miller, Ryan P. Adams, David M. Blei, Liam Paninski
AISTATS3
2017 Variational Boosting: Iteratively Refining Posterior Approximations
abstract
We propose a black-box variational inference method to approximate intractable distributions with an increasingly rich approximating class. Our method, variational boosting, iteratively refines an existing variational approximation by solving a sequence of optimization problems, allowing a trade-off between computation time and accuracy. We expand the variational approximating class by incorporating additional covariance structure and by introducing new components to form a mixture. We apply variational boosting to synthetic and real statistical models, and show that the resulting posterior inferences compare favorably to existing variational algorithms.
Andrew C. Miller, Nicholas J. Foti, Ryan P. Adams
ICML1
2017 Reducing Reparameterization Gradient Variance
abstract
Optimization with noisy gradients has become ubiquitous in statistics and machine learning. Reparameterization gradients, or gradient estimates computed via the ``reparameterization trick,'' represent a class of noisy gradients often used in Monte Carlo variational inference (MCVI). However, when these gradient estimators are too noisy, the optimization procedure can be slow or fail to converge. One way to reduce noise is to generate more samples for the gradient estimate, but this can be computationally expensive. Instead, we view the noisy gradient as a random variable, and form an inexpensive approximation of the generating procedure for the gradient sample. This approximation has high correlation with the noisy gradient by construction, making it a useful control variate for variance reduction. We demonstrate our approach on a non-conjugate hierarchical model and a Bayesian neural net where our method attained orders of magnitude (20-2{,}000$\times$) reduction in gradient variance resulting in faster and more stable optimization.
Andrew C. Miller, Nick Foti, Alexander D'Amour, Ryan P. Adams
NIPS1
2015 Celeste: Variational inference for a generative model of astronomical images
abstract
We present a new, fully generative model of optical telescope image sets, along with a variational procedure for inference. Each pixel intensity is treated as a Poisson random variable, with a rate parameter dependent on latent properties of stars and galaxies. Key latent properties are themselves random, with scientific prior distributions constructed from large ancillary data sets. We check our approach on synthetic images. We also run it on images from a major sky survey, where it exceeds the performance of the current state-of-the-art method for locating celestial bodies and measuring their colors.
Jeffrey Regier, Andrew C. Miller, Jon D. McAuliffe, Ryan P. Adams, Matthew Hoffman 0001, Dustin Lang, David Schlegel, Prabhat
ICML2
2015 A Gaussian Process Model of Quasar Spectral Energy Distributions
abstract
We propose a method for combining two sources of astronomical data, spectroscopy and photometry, that carry information about sources of light (e.g., stars, galaxies, and quasars) at extremely different spectral resolutions. Our model treats the spectral energy distribution (SED) of the radiation from a source as a latent variable that jointly explains both photometric and spectroscopic observations. We place a flexible, nonparametric prior over the SED of a light source that admits a physically interpretable decomposition, and allows us to tractably perform inference. We use our model to predict the distribution of the redshift of a quasar from five-band (low spectral resolution) photometric data, the so called ``photo-z'' problem. Our method shows that tools from machine learning and Bayesian statistics allow us to leverage multiple resolutions of information to make accurate predictions with well-characterized uncertainties.
Andrew C. Miller, Albert Wu, Jeffrey Regier, Jon D. McAuliffe, Dustin Lang, Prabhat, David Schlegel, Ryan P. Adams
NIPS1
2014 Factorized Point Process Intensities: A Spatial Analysis of Professional Basketball
abstract
We develop a machine learning approach to represent and analyze the underlying spatial structure that governs shot selection among professional basketball players in the NBA. Typically, NBA players are discussed and compared in an heuristic, imprecise manner that relies on unmeasured intuitions about player behavior. This makes it difficult to draw comparisons between players and make accurate player specific predictions. Modeling shot attempt data as a point process, we create a low dimensional representation of offensive player types in the NBA. Using non-negative matrix factorization (NMF), an unsupervised dimensionality reduction technique, we show that a low-rank spatial decomposition summarizes the shooting habits of NBA players. The spatial representations discovered by the algorithm correspond to intuitive descriptions of NBA player types, and can be used to model other spatial effects, such as shooting accuracy.
Andrew C. Miller, Luke Bornn, Ryan P. Adams, Kirk Goldsberry
ICML1