VLDB 2026 Research / reviewers in the wild / expert
Diana Cai
dblp:191/6693
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Probabilistic and Bayesian machine learning · 91% Learning theory · 5% Generative modeling · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › gradient-based variational inference
black-box variational inference |
2.4 | 3 | 2025 | Fisher meets Feynman: score-based variational inference with a product of experts · NeurIPS 2025 EigenVI: score-based variational inference with orthogonal function expansions · NeurIPS 2024 Batch and match: black-box variational inference with a score-based divergence · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
2.4 | 3 | 2025 | Fisher meets Feynman: score-based variational inference with a product of experts · NeurIPS 2025 EigenVI: score-based variational inference with orthogonal function expansions · NeurIPS 2024 Batch and match: black-box variational inference with a score-based divergence · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
1.1 | 2 | 2022 | Multi-fidelity Monte Carlo: a pseudo-marginal approach · NeurIPS 2022 Slice Sampling Reparameterization Gradients · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection |
0.5 | 1 | 2021 | Finite mixture models do not reliably learn the number of components · ICML 2021 |
Machine learning › Learning theory › statistical estimation
misspecified models |
0.5 | 1 | 2021 | Finite mixture models do not reliably learn the number of components · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior consistency |
0.5 | 1 | 2021 | Finite mixture models do not reliably learn the number of components · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
reparameterization gradient |
0.5 | 1 | 2021 | Slice Sampling Reparameterization Gradients · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
slice sampling |
0.5 | 1 | 2021 | Slice Sampling Reparameterization Gradients · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.3 | 1 | 2018 | A Bayesian Nonparametric View on Count-Min Sketch · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process |
0.3 | 1 | 2018 | A Bayesian Nonparametric View on Count-Min Sketch · NeurIPS 2018 |
Data stream processing › frequency estimation
count-min sketch |
0.3 | 1 | 2018 | A Bayesian Nonparametric View on Count-Min Sketch · NeurIPS 2018 |
Data stream processing
frequency estimation |
0.3 | 1 | 2018 | A Bayesian Nonparametric View on Count-Min Sketch · NeurIPS 2018 |
Machine learning › Generative modeling › generative model › probabilistic generative model
product of experts |
0.3 | 1 | 2025 | Fisher meets Feynman: score-based variational inference with a product of experts · NeurIPS 2025 |
Machine learning › Graph learning
graph generation |
0.2 | 1 | 2016 | Edge-exchangeable graphs and sparsity · NIPS 2016 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference |
0.2 | 1 | 2024 | EigenVI: score-based variational inference with orthogonal function expansions · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference |
0.2 | 1 | 2024 | Batch and match: black-box variational inference with a score-based divergence · ICML 2024 |
Computational science and engineering
uncertainty quantification |
0.2 | 1 | 2022 | Multi-fidelity Monte Carlo: a pseudo-marginal approach · NeurIPS 2022 |
Graph algorithms and graph theory
random graph models |
0.1 | 1 | 2016 | Edge-exchangeable graphs and sparsity · NIPS 2016 |
Mathematical optimization
sparsity |
0.1 | 1 | 2016 | Edge-exchangeable graphs and sparsity · NIPS 2016 |
Methods — techniques the papers use, named apart from their topics
fisher divergence minimization · 1.6pseudo-marginal MCMC · 1.1feynman identity · 0.9dirichlet latent variables · 0.9convex quadratic programming · 0.9score-based divergence · 0.8proximal update · 0.8orthogonal function expansions · 0.8eigenvalue problem · 0.8ELBO · 0.8telescoping series · 0.6random truncation · 0.6stable beta process · 0.3dirichlet process · 0.3bayesian inference · 0.3latent random measures · 0.2exchangeability · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Batch, match, and patch: low-rank approximations for score-based variational inferenceabstractBlack-box variational inference (BBVI) scales poorly to high-dimensional problems when it is used to estimate a multivariate Gaussian approximation with a full covariance matrix. In this paper, we extend the \emph{batch-and-match} (BaM) framework for score-based BBVI to problems where it is prohibitively expensive to store such covariance matrices, let alone to estimate them. Unlike classical algorithms for BBVI, which use stochastic gradient descent to minimize the reverse Kullback-Leibler divergence, BaM uses more specialized updates to match the scores of the target density and its Gaussian approximation. We extend the updates for BaM by integrating them with a more compact parameterization of full covariance matrices. In particular, borrowing ideas from factor analysis, we add an extra step to each iteration of BaM—a \emph{patch}—that projects each newly updated covariance matrix into a more efficiently parameterized family of diagonal plus low rank matrices. We evaluate this approach on a variety of synthetic target distributions and real-world problems in high-dimensional inference. Chirag Modi 0002, Diana Cai, Lawrence K. Saul |
AISTATS | 2 |
| 2025 | Fisher meets Feynman: score-based variational inference with a product of expertsabstractWe introduce a highly expressive yet distinctly tractable family for black-box variational inference (BBVI). Each member of this family is a weighted product of experts (PoE), and each weighted expert in the product is proportional to a multivariate $t$-distribution. These products of experts can model distributions with skew, heavy tails, and multiple modes, but to use them for BBVI, we must be able to sample from their densities. We show how to do this by reformulating these products of experts as latent variable models with auxiliary Dirichlet random variables. These Dirichlet variables emerge from a Feynman identity, originally developed for loop integrals in quantum field theory, that expresses the product of multiple fractions (or in our case, $t$-distributions) as an integral over the simplex. We leverage this simplicial latent space to draw weighted samples from these products of experts---samples which BBVI then uses to find the PoE that best approximates a target density. Given a collection of experts, we derive an iterative procedure to optimize the exponents that determine their geometric weighting in the PoE. At each iteration, this procedure minimizes a regularized Fisher divergence to match the scores of the variational and target densities at a batch of samples drawn from the current approximation. This minimization reduces to a convex quadratic program, and we prove under general conditions that these updates converge exponentially fast to a near-optimal weighting of experts. We conclude by evaluating this approach on a variety of synthetic and real-world target distributions. Diana Cai, Robert M. Gower, David M. Blei, Lawrence K. Saul |
NeurIPS | 1 |
| 2024 | Batch and match: black-box variational inference with a score-based divergenceabstractMost leading implementations of black-box variational inference (BBVI) are based on optimizing a stochastic evidence lower bound (ELBO). But such approaches to BBVI often converge slowly due to the high variance of their gradient estimates and their sensitivity to hyperparameters. In this work, we propose _batch and match_ (BaM), an alternative approach to BBVI based on a score-based divergence. Notably, this score-based divergence can be optimized by a closed-form proximal update for Gaussian variational families with full covariance matrices. We analyze the convergence of BaM when the target distribution is Gaussian, and we prove that in the limit of infinite batch size the variational parameter updates converge exponentially quickly to the target mean and covariance. We also evaluate the performance of BaM on Gaussian and non-Gaussian target distributions that arise from posterior inference in hierarchical and deep generative models. In these experiments, we find that BaM typically converges in fewer (and sometimes significantly fewer) gradient evaluations than leading implementations of BBVI based on ELBO maximization. Diana Cai, Chirag Modi 0002, Loucas Pillaud-Vivien, Charles C. Margossian, Robert M. Gower, David M. Blei, Lawrence K. Saul |
ICML | 1 |
| 2024 | EigenVI: score-based variational inference with orthogonal function expansionsabstractWe develop EigenVI, an eigenvalue-based approach for black-box variational inference (BBVI). EigenVI constructs its variational approximations from orthogonal function expansions. For distributions over $\mathbb{R}^D$, the lowest order term in these expansions provides a Gaussian variational approximation, while higher-order terms provide a systematic way to model non-Gaussianity. These approximations are flexible enough to model complex distributions (multimodal, asymmetric), but they are simple enough that one can calculate their low-order moments and draw samples from them. EigenVI can also model other types of random variables (e.g., nonnegative, bounded) by constructing variational approximations from different families of orthogonal functions. Within these families, EigenVI computes the variational approximation that best matches the score function of the target distribution by minimizing a stochastic estimate of the Fisher divergence. Notably, this optimization reduces to solving a minimum eigenvalue problem, so that EigenVI effectively sidesteps the iterative gradient-based optimizations that are required for many other BBVI algorithms. (Gradient-based methods can be sensitive to learning rates, termination criteria, and other tunable hyperparameters.) We use EigenVI to approximate a variety of target distributions, including a benchmark suite of Bayesian models from posteriordb. On these distributions, we find that EigenVI is more accurate than existing methods for Gaussian BBVI. Diana Cai, Chirag Modi 0002, Charles C. Margossian, Robert M. Gower, David M. Blei, Lawrence K. Saul |
NeurIPS | 1 |
| 2022 | Multi-fidelity Monte Carlo: a pseudo-marginal approachabstractMarkov chain Monte Carlo (MCMC) is an established approach for uncertainty quantification and propagation in scientific applications. A key challenge in applying MCMC to scientific domains is computation: the target density of interest is often a function of expensive computations, such as a high-fidelity physical simulation, an intractable integral, or a slowly-converging iterative algorithm. Thus, using an MCMC algorithms with an expensive target density becomes impractical, as these expensive computations need to be evaluated at each iteration of the algorithm. In practice, these computations often approximated via a cheaper, low-fidelity computation, leading to bias in the resulting target density. Multi-fidelity MCMC algorithms combine models of varying fidelities in order to obtain an approximate target density with lower computational cost. In this paper, we describe a class of asymptotically exact multi-fidelity MCMC algorithms for the setting where a sequence of models of increasing fidelity can be computed that approximates the expensive target density of interest. We take a pseudo-marginal MCMC approach for multi-fidelity inference that utilizes a cheaper, randomized-fidelity unbiased estimator of the target fidelity constructed via random truncation of a telescoping series of the low-fidelity sequence of models. Finally, we discuss and evaluate the proposed multi-fidelity MCMC approach on several applications, including log-Gaussian Cox process modeling, Bayesian ODE system identification, PDE-constrained optimization, and Gaussian process parameter inference. Diana Cai, Ryan P. Adams |
NeurIPS | 1 |
| 2021 | Finite mixture models do not reliably learn the number of componentsabstractScientists and engineers are often interested in learning the number of subpopulations (or components) present in a data set. A common suggestion is to use a finite mixture model (FMM) with a prior on the number of components. Past work has shown the resulting FMM component-count posterior is consistent; that is, the posterior concentrates on the true, generating number of components. But consistency requires the assumption that the component likelihoods are perfectly specified, which is unrealistic in practice. In this paper, we add rigor to data-analysis folk wisdom by proving that under even the slightest model misspecification, the FMM component-count posterior diverges: the posterior probability of any particular finite number of components converges to 0 in the limit of infinite data. Contrary to intuition, posterior-density consistency is not sufficient to establish this result. We develop novel sufficient conditions that are more realistic and easily checkable than those common in the asymptotics literature. We illustrate practical consequences of our theory on simulated and real data. Diana Cai, Trevor Campbell, Tamara Broderick |
ICML | 1 |
| 2021 | Slice Sampling Reparameterization GradientsabstractMany probabilistic modeling problems in machine learning use gradient-based optimization in which the objective takes the form of an expectation. These problems can be challenging when the parameters to be optimized determine the probability distribution under which the expectation is being taken, as the na\"ive Monte Carlo procedure is not differentiable. Reparameterization gradients make it possible to efficiently perform optimization of these Monte Carlo objectives by transforming the expectation to be differentiable, but the approach is typically limited to distributions with simple forms and tractable normalization constants. Here we describe how to differentiate samples from slice sampling to compute \textit{slice sampling reparameterization gradients}, enabling a richer class of Monte Carlo objective functions to be optimized. Slice sampling is a Markov chain Monte Carlo algorithm for simulating samples from probability distributions; it only requires a density function that can be evaluated point-wise up to a normalization constant, making it applicable to a variety of inference problems and unnormalized models. Our approach is based on the observation that when the slice endpoints are known, the sampling path is a deterministic and differentiable function of the pseudo-random variables, since the algorithm is rejection-free. We evaluate the method on synthetic examples and apply it to a variety of applications with reparameterization of unnormalized probability distributions. David M. Zoltowski, Diana Cai, Ryan P. Adams |
NeurIPS | 2 |
| 2021 | Active multi-fidelity Bayesian online changepoint detectionabstractOnline algorithms for detecting changepoints, or abrupt shifts in the behavior of a time series, are often deployed with limited resources, e.g., to edge computing settings such as mobile phones or industrial sensors. In these scenarios it may be beneficial to trade the cost of collecting an environmental measurement against the quality or “fidelity” of this measurement and how the measurement affects changepoint estimation. For instance, one might decide between inertial measurements or GPS to determine changepoints for motion. A Bayesian approach to changepoint detection is particularly appealing because we can represent our posterior uncertainty about changepoints and make active, cost-sensitive decisions about data fidelity to reduce this posterior uncertainty. Moreover, the total cost could be dramatically lowered through active fidelity switching, while remaining robust to changes in data distribution. We propose a multi-fidelity approach that makes cost-sensitive decisions about which data fidelity to collect based on maximizing information gain with respect to changepoints. We evaluate this framework on synthetic, video, and audio data and show that this information-based approach results in accurate predictions while reducing total cost. Gregory W. Gundersen, Diana Cai, Chuteng Zhou, Barbara E. Engelhardt, Ryan P. Adams |
UAI | 2 |
| 2018 | A Bayesian Nonparametric View on Count-Min SketchabstractThe count-min sketch is a time- and memory-efficient randomized data structure that provides a point estimate of the number of times an item has appeared in a data stream. The count-min sketch and related hash-based data structures are ubiquitous in systems that must track frequencies of data such as URLs, IP addresses, and language n-grams. We present a Bayesian view on the count-min sketch, using the same data structure, but providing a posterior distribution over the frequencies that characterizes the uncertainty arising from the hash-based approximation. In particular, we take a nonparametric approach and consider tokens generated from a Dirichlet process (DP) random measure, which allows for an unbounded number of unique tokens. Using properties of the DP, we show that it is possible to straightforwardly compute posterior marginals of the unknown true counts and that the modes of these marginals recover the count-min sketch estimator, inheriting the associated probabilistic guarantees. Using simulated data with known ground truth, we investigate the properties of these estimators. Lastly, we also study a modified problem in which the observation stream consists of collections of tokens (i.e., documents) arising from a random measure drawn from a stable beta process, which allows for power law scaling behavior in the number of unique tokens. Diana Cai, Michael Mitzenmacher, Ryan P. Adams |
NeurIPS | 1 |
| 2016 | Edge-exchangeable graphs and sparsityabstractMany popular network models rely on the assumption of (vertex) exchangeability, in which the distribution of the graph is invariant to relabelings of the vertices. However, the Aldous-Hoover theorem guarantees that these graphs are dense or empty with probability one, whereas many real-world graphs are sparse. We present an alternative notion of exchangeability for random graphs, which we call edge exchangeability, in which the distribution of a graph sequence is invariant to the order of the edges. We demonstrate that edge-exchangeable models, unlike models that are traditionally vertex exchangeable, can exhibit sparsity. To do so, we outline a general framework for graph generative models; by contrast to the pioneering work of Caron and Fox (2015), models within our framework are stationary across steps of the graph sequence. In particular, our model grows the graph by instantiating more latent atoms of a single random measure as the dataset size increases, rather than adding new atoms to the measure. Diana Cai, Trevor Campbell, Tamara Broderick |
NIPS | 1 |