Nicholas J. Foti

dblp:122/3118 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Probabilistic and Bayesian machine learning · 90% Generative modeling · 5% Deep learning architectures and training · 4%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.612022
Neural Granger Causality · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality
0.612022
Neural Granger Causality · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.522017
Stochastic Gradient MCMC Methods for Hidden Markov Models · ICML 2017
Stochastic variational inference for hidden Markov models · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.522017
Variational Boosting: Iteratively Refining Posterior Approximations · ICML 2017
Stochastic variational inference for hidden Markov models · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.422017
Stochastic Gradient MCMC Methods for Hidden Markov Models · ICML 2017
Slice sampling normalized kernel-weighted completely random measure mixture models · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.422015
A Survey of Non-Exchangeable Priors for Bayesian Nonparametric Models · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Slice sampling normalized kernel-weighted completely random measure mixture models · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent factor model
interpretable latent factor model
0.312018
oi-VAE: Output Interpretable VAEs for Nonlinear Group Factor Analysis · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor model
0.312018
oi-VAE: Output Interpretable VAEs for Nonlinear Group Factor Analysis · ICML 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
oi-VAE: Output Interpretable VAEs for Nonlinear Group Factor Analysis · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.312017
Variational Boosting: Iteratively Refining Posterior Approximations · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
boosting variational inference
0.312017
Variational Boosting: Iteratively Refining Posterior Approximations · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.312017
Stochastic Gradient MCMC Methods for Hidden Markov Models · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
stochastic gradient MCMC
0.312017
Stochastic Gradient MCMC Methods for Hidden Markov Models · ICML 2017
Machine learning › Deep learning architectures and training
foundation model
0.312025
Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212014
Stochastic variational inference for hidden Markov models · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212014
Stochastic variational inference for hidden Markov models · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
stochastic variational inference
0.212014
Stochastic variational inference for hidden Markov models · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
slice sampling
0.112012
Slice sampling normalized kernel-weighted completely random measure mixture models · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › prior modeling
sparsity-inducing prior
0.112018
oi-VAE: Output Interpretable VAEs for Nonlinear Group Factor Analysis · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process
0.112015
A Survey of Non-Exchangeable Priors for Bayesian Nonparametric Models · IEEE Trans. Pattern Anal. Mach. Intell. 2015

Methods — techniques the papers use, named apart from their topics

tokenization · 1.7foundation models · 0.9foundation model · 0.9variational inference · 0.6recurrent neural network · 0.6multi-layer perceptron · 0.6group-lasso penalty · 0.6sparsity regularization · 0.3mini-batch sampling · 0.3marginal likelihood · 0.3boosting · 0.3
YearPublicationVenuePosition
2025 Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
abstract
Wearable devices record physiological and behavioral signals that can improve health predictions. While foundation models are increasingly used for such predictions, they have been primarily applied to low-level sensor data, despite behavioral data often being more informative due to their alignment with physiologically relevant timescales and quantities. We develop foundation models of such behavioral signals using over 2.5B hours of wearable data from 162K individuals, systematically optimizing architectures and tokenization strategies for this unique dataset. Evaluated on 57 health-related tasks, our model shows strong performance across diverse real-world applications including individual-level classification and time-varying health state prediction. The model excels in behavior-driven tasks like sleep prediction, and improves further when combined with representations of raw sensor data. These results underscore the importance of tailoring foundation model design to wearables and demonstrate the potential to enable new health applications.
Eray Erturk, Fahad Kamran, Salar Abbaspourazad, Sean Jewell, Sinead Williamson, Nicholas J. Foti, Joseph Futoma
ICML8
2022 Neural Granger Causality
abstract
While most classical approaches to Granger causality detection assume linear dynamics, many interactions in real-world applications, like neuroscience and genomics, are inherently nonlinear. In these cases, using linear models may lead to inconsistent estimation of Granger causal interactions. We propose a class of nonlinear methods by applying structured multilayer perceptrons (MLPs) or recurrent neural networks (RNNs) combined with sparsity-inducing penalties on the weights. By encouraging specific sets of weights to be zero-in particular, through the use of convex group-lasso penalties-we can extract the Granger causal structure. To further contrast with traditional approaches, our framework naturally enables us to efficiently capture long-range dependencies between series either via our RNNs or through an automatic lag selection in the MLP. We show that our neural Granger causality methods outperform state-of-the-art nonlinear Granger causality methods on the DREAM3 challenge data. This data consists of nonlinear gene expression and regulation time courses with only a limited number of time points. The successes we show in this challenging dataset provide a powerful example of how deep learning can be useful in cases that go beyond prediction on large datasets. We likewise illustrate our methods in detecting nonlinear interactions in a human motion capture dataset.
Alex Tank, Ian Covert, Nicholas J. Foti, Ali Shojaie, Emily B. Fox
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Adaptively Truncating Backpropagation Through Time to Control Gradient Bias
Christopher Aicher, Nicholas J. Foti, Emily B. Fox
UAI2
2018 oi-VAE: Output Interpretable VAEs for Nonlinear Group Factor Analysis
abstract
Deep generative models have recently yielded encouraging results in producing subjectively realistic samples of complex data. Far less attention has been paid to making these generative models interpretable. In many scenarios, ranging from scientific applications to finance, the observed variables have a natural grouping. It is often of interest to understand systems of interaction amongst these groups, and latent factor models (LFMs) are an attractive approach. However, traditional LFMs are limited by assuming a linear correlation structure. We present an output interpretable VAE (oi-VAE) for grouped data that models complex, nonlinear latent-to-observed relationships. We combine a structured VAE comprised of group-specific generators with a sparsity-inducing prior. We demonstrate that oi-VAE yields meaningful notions of interpretability in the analysis of motion capture and MEG data. We further show that in these situations, the regularization inherent to oi-VAE can actually lead to improved generalization and learned generative processes.
Samuel K. Ainsworth, Nicholas J. Foti, Adrian K. C. Lee, Emily B. Fox
ICML2
2018 The cultural evolution of national constitutions
abstract
We explore how ideas from infectious disease and genetics can be used to uncover patterns of cultural inheritance and innovation in a corpus of 591 national constitutions spanning 1789–2008. Legal “ideas” are encoded as “topics”—words statistically linked in documents—derived from topic modeling the corpus of constitutions. Using these topics we derive a diffusion network for borrowing from ancestral constitutions back to the US Constitution of 1789 and reveal that constitutions are complex cultural recombinants. We find systematic variation in patterns of borrowing from ancestral texts and “biological”‐like behavior in patterns of inheritance, with the distribution of “offspring” arising through a bounded preferential‐attachment process. This process leads to a small number of highly innovative (influential) constitutions some of which have yet to have been identified as so in the current literature. Our findings thus shed new light on the critical nodes of the constitution‐making network. The constitutional network structure reflects periods of intense constitution creation, and systematic patterns of variation in constitutional lifespan and temporal influence.
Daniel N. Rockmore, Nicholas J. Foti, Tom Ginsburg, David C. Krakauer
J. Assoc. Inf. Sci. Technol.3
2017 Stochastic Gradient MCMC Methods for Hidden Markov Models
abstract
Stochastic gradient MCMC (SG-MCMC) algorithms have proven useful in scaling Bayesian inference to large datasets under an assumption of i.i.d data. We instead develop an SG-MCMC algorithm to learn the parameters of hidden Markov models (HMMs) for time-dependent data. There are two challenges to applying SG-MCMC in this setting: The latent discrete states, and needing to break dependencies when considering minibatches. We consider a marginal likelihood representation of the HMM and propose an algorithm that harnesses the inherent memory decay of the process. We demonstrate the effectiveness of our algorithm on synthetic experiments and an ion channel recording data, with runtimes significantly outperforming batch MCMC.
Yi-An Ma, Nicholas J. Foti, Emily B. Fox
ICML2
2017 Variational Boosting: Iteratively Refining Posterior Approximations
abstract
We propose a black-box variational inference method to approximate intractable distributions with an increasingly rich approximating class. Our method, variational boosting, iteratively refines an existing variational approximation by solving a sequence of optimization problems, allowing a trade-off between computation time and accuracy. We expand the variational approximating class by incorporating additional covariance structure and by introducing new components to form a mixture. We apply variational boosting to synthetic and real statistical models, and show that the resulting posterior inferences compare favorably to existing variational algorithms.
Andrew C. Miller, Nicholas J. Foti, Ryan P. Adams
ICML2
2015 Streaming Variational Inference for Bayesian Nonparametric Mixture Models
abstract
In theory, Bayesian nonparametric (BNP) models are well suited to streaming data scenarios due to their ability to adapt model complexity based on the amount of data observed. Unfortunately, such benefits have not been fully realized in practice; existing inference algorithms either are not applicable to streaming applications or are not extensible to nonparametric models. For the special case of Dirichlet processes, streaming inference has been considered. However, there is growing interest in more flexible BNP models, in particular building on the class of normalized random measures (NRMs). We work within this general framework and present a streaming variational inference algorithm for NRM mixture models based on assumed density filtering. Extensions to expectation propagation algorithms are possible in the batch data setting. We demonstrate the efficacy of the algorithm on clustering documents in large, streaming text corpora.
Alex Tank, Nicholas J. Foti, Emily B. Fox
AISTATS2
2015 Bayesian Structure Learning for Stationary Time Series
Alex Tank, Nicholas J. Foti, Emily B. Fox
UAI2
2015 A Survey of Non-Exchangeable Priors for Bayesian Nonparametric Models
abstract
Dependent nonparametric processes extend distributions over measures, such as the Dirichlet process and the beta process, to give distributions over collections of measures, typically indexed by values in some covariate space. Such models are appropriate priors when exchangeability assumptions do not hold, and instead we want our model to vary fluidly with some set of covariates. Since the concept of dependent nonparametric processes was formalized by MacEachern, there have been a number of models proposed and used in the statistics and machine learning literatures. Many of these models exhibit underlying similarities, an understanding of which, we hope, will help in selecting an appropriate prior, developing new models, and leveraging inference techniques.
Nicholas J. Foti, Sinead Williamson
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Stochastic variational inference for hidden Markov models
Nicholas J. Foti, Dillon Laird, Emily B. Fox
NIPS1
2013 A unifying representation for a class of dependent random measures
abstract
We present a general construction for dependent random measures based on thinning Poisson processes on an augmented space. The framework is not restricted to dependent versions of a specific nonparametric model, but can be applied to all models that can be represented using completely random measures. Several existing dependent random measures can be seen as specific cases of this framework. Interesting properties of the resulting measures are derived and the efficacy of the framework is demonstrated by constructing a covariate-dependent latent feature model and topic model that obtain superior predictive performance.
Nicholas J. Foti, Joseph Futoma, Daniel N. Rockmore, Sinead Williamson
AISTATS1
2012 Slice sampling normalized kernel-weighted completely random measure mixture models
abstract
A number of dependent nonparametric processes have been proposed to model non-stationary data with unknown latent dimensionality. However, the inference algorithms are often slow and unwieldy, and are in general highly specific to a given model formulation. In this paper, we describe a wide class of nonparametric processes, including several existing models, and present a slice sampler that allows efficient inference across this class of models.
Nicholas J. Foti, Sinead Williamson
NIPS1