VLDB 2026 Research / reviewers in the wild / expert
Pavel Sountsov
dblp:155/8031
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 30% Probabilistic and Bayesian machine learning · 29% 3D vision · 25% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d scene reconstruction |
0.8 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Machine learning › Generative modeling › diffusion model
diffusion model conditioning |
0.8 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Computer vision › 3D vision
inverse rendering |
0.8 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference |
0.8 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting |
0.7 | 1 | 2023 | Training Chain-of-Thought via Latent-Variable Inference · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization |
0.7 | 1 | 2023 | Training Chain-of-Thought via Latent-Variable Inference · NeurIPS 2023 |
Machine learning › Generative modeling
energy-based model |
0.6 | 1 | 2022 | MCMC Should Mix: Learning Energy-Based Model with Neural Transport Latent Space MCMC · ICLR 2022 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.6 | 1 | 2022 | MCMC Should Mix: Learning Energy-Based Model with Neural Transport Latent Space MCMC · ICLR 2022 |
Computer vision › 3D vision
neural radiance field |
0.2 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Computer vision › Segmentation and scene understanding
scene priors |
0.2 | 1 | 2024 | Robust Inverse Graphics via Probabilistic Inference · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
markov chain monte carlo · 1.2normalizing flow · 0.8neural radiance field · 0.8diffusion model · 0.8bayesian inference · 0.8wake-sleep · 0.7control variates · 0.7neural transport · 0.6global conditioning · 0.2beam search · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Robust Inverse Graphics via Probabilistic InferenceabstractHow do we infer a 3D scene from a single image in the presence of corruptions like rain, snow or fog? Straightforward domain randomization relies on knowing the family of corruptions ahead of time. Here, we propose a Bayesian approach—dubbed robust inverse graphics (RIG)—that relies on a strong scene prior and an uninformative uniform corruption prior, making it applicable to a wide range of corruptions. Given a single image, RIG performs posterior inference jointly over the scene and the corruption. We demonstrate this idea by training a neural radiance field (NeRF) scene prior and using a secondary NeRF to represent the corruptions over which we place an uninformative prior. RIG, trained only on clean data, outperforms depth estimators and alternative NeRF approaches that perform point estimation instead of full inference. The results hold for a number of scene prior architectures based on normalizing flows and diffusion models. For the latter, we develop reconstruction-guidance with auxiliary latents (ReGAL)—a diffusion conditioning algorithm that is applicable in the presence of auxiliary latent variables such as the corruption. RIG demonstrates how scene priors can be used beyond generation tasks. Pavel Sountsov, Matthew Hoffman 0001, Ben Lee, Brian Patton, Rif A. Saurous |
ICML | 2 |
| 2023 | ProbNeRF: Uncertainty-Aware Inference of 3D Shapes from 2D ImagesabstractThe problem of inferring object shape from a single 2D image is underconstrained. Prior knowledge about what objects are plausible can help, but even given such prior knowledge there may still be uncertainty about the shapes of occluded parts of objects. Recently, conditional neural radiance field (NeRF) models have been developed that can learn to infer good point estimates of 3D models from single 2D images. The problem of inferring uncertainty estimates for these models has received less attention. In this work, we propose probabilistic NeRF (ProbNeRF), a model and inference strategy for learning probabilistic generative models of 3D objects’ shapes and appearances, and for doing posterior inference to recover those properties from 2D images. ProbNeRF is trained as a variational autoencoder, but at test time we use Hamiltonian Monte Carlo (HMC) for inference. Given one or a few 2D images of an object (which may be partially occluded), ProbNeRF is able not only to accurately model the parts it sees, but also to propose realistic and diverse hypotheses about the parts it does not see. We show that key to the success of ProbNeRF are (i) a deterministic rendering scheme, (ii) an annealed-HMC strategy, (iii) a hypernetwork-based decoder architecture, and (iv) doing inference over a full set of NeRF weights, rather than just a low-dimensional code. Videos and code are available at https://probnerf.github.io. Matthew Hoffman 0001, Pavel Sountsov, Christopher Suter, Ben Lee, Vikash Mansinghka 0001, Rif A. Saurous |
AISTATS | 3 |
| 2023 | Adaptive Tuning for Metropolis Adjusted Langevin TrajectoriesabstractHamiltonian Monte Carlo (HMC) is a widely used sampler for continuous probability distributions. In many cases, the underlying Hamiltonian dynamics exhibit a phenomenon of resonance which decreases the efficiency of the algorithm and makes it very sensitive to hyperparameter values. This issue can be tackled efficiently, either via the use of trajectory length randomization (RHMC) or via partial momentum refreshment. The second approach is connected to the kinetic Langevin diffusion, and has been mostly investigated through the use of Generalized HMC (GHMC). However, GHMC induces momentum flips upon rejections causing the sampler to backtrack and waste computational resources. In this work we focus on a recent algorithm bypassing this issue, named Metropolis Adjusted Langevin Trajectories (MALT). We build upon recent strategies for tuning the hyperparameters of RHMC which target a bound on the Effective Sample Size (ESS) and adapt it to MALT, thereby enabling the first user-friendly deployment of this algorithm. We construct a method to optimize a sharper bound on the ESS and reduce the estimator variance. Easily compatible with parallel implementation, the resultant Adaptive MALT algorithm is competitive in terms of ESS rate and hits useful tradeoffs in memory usage when compared to GHMC, RHMC and NUTS. Lionel Riou-Durand, Pavel Sountsov, Jure Vogrinc, Charles C. Margossian, Sam Power |
AISTATS | 2 |
| 2023 | Training Chain-of-Thought via Latent-Variable InferenceabstractLarge language models (LLMs) solve problems more accurately and interpretably when instructed to work out the answer step by step using a "chain-of-thought" (CoT) prompt. One can also improve LLMs' performance on a specific task by supervised fine-tuning, i.e., by using gradient ascent on some tunable parameters to maximize the average log-likelihood of correct answers from a labeled training set.
Naively combining CoT with supervised tuning requires supervision not just of the correct answers, but also of detailed rationales that lead to those answers; these rationales are expensive to produce by hand. Instead, we propose a fine-tuning strategy that tries to maximize the \emph{marginal} log-likelihood of generating a correct answer using CoT prompting, approximately averaging over all possible rationales. The core challenge is sampling from the posterior over rationales conditioned on the correct answer; we address it using a simple Markov-chain Monte Carlo (MCMC) expectation-maximization (EM) algorithm inspired by the self-taught reasoner (STaR), memoized wake-sleep, Markovian score climbing, and persistent contrastive divergence. This algorithm also admits a novel control-variate technique that drives the variance of our gradient estimates to zero as the model improves. Applying our technique to GSM8K and the tasks in BIG-Bench Hard, we find that this MCMC-EM fine-tuning technique typically improves the model's accuracy on held-out examples more than STaR or prompt-tuning with or without CoT. Matthew Hoffman 0001, Du Phan, David Dohan, Sholto Douglas, Aaron Parisi, Pavel Sountsov, Charles Sutton, Sharad Vikram, Rif A. Saurous |
NeurIPS | 7 |
| 2022 | Tuning-Free Generalized Hamiltonian Monte CarloabstractHamiltonian Monte Carlo (HMC) has become a go-to family of Markov chain Monte Carlo (MCMC) algorithms for Bayesian inference problems, in part because we have good procedures for automatically tuning its parameters. Much less attention has been paid to automatic tuning of generalized HMC (GHMC), in which the auxiliary momentum vector is partially updated frequently instead of being completely resampled infrequently. Since GHMC spreads progress over many iterations, it is not straightforward to tune GHMC based on quantities typically used to tune HMC such as average acceptance rate and squared jumped distance. In this work, we propose an ensemble-chain adaptation (ECA) algorithm for GHMC that automatically selects values for all of GHMC’s tunable parameters each iteration based on statistics collected from a population of many chains. This algorithm is designed to make good use of SIMD hardware accelerators such as GPUs, allowing most chains to be updated in parallel each iteration. Unlike typical adaptive-MCMC algorithms, our ECA algorithm does not perturb the chain’s stationary distribution, and therefore does not need to be “frozen” after warmup. Empirically, we find that the proposed algorithm quickly converges to its stationary distribution, producing accurate estimates of posterior expectations with relatively few gradient evaluations per chain. Matthew Hoffman 0001, Pavel Sountsov |
AISTATS | 2 |
| 2022 | MCMC Should Mix: Learning Energy-Based Model with Neural Transport Latent Space MCMC
Erik Nijkamp, Ruiqi Gao, Pavel Sountsov, Srinivas Vasudevan, Bo Pang 0004, Song-Chun Zhu, Ying Nian Wu |
ICLR | 3 |
| 2021 | An Adaptive-MCMC Scheme for Setting Trajectory Lengths in Hamiltonian Monte CarloabstractHamiltonian Monte Carlo (HMC) is a powerful MCMC algorithm based on simulating Hamiltonian dynamics. Its performance depends strongly on choosing appropriate values for two parameters: the step size used in the simulation, and how long the simulation runs for. The step-size parameter can be tuned using standard adaptive-MCMC strategies, but it is less obvious how to tune the simulation-length parameter. The no-U-turn sampler (NUTS) eliminates this problematic simulation-length parameter, but NUTS’s relatively complex control flow makes it difficult to efficiently run many parallel chains on accelerators such as GPUs. NUTS also spends some extra gradient evaluations relative to HMC in order to decide how long to run each iteration without violating detailed balance. We propose ChEES-HMC, a simple adaptive-MCMC scheme for automatically tuning HMC’s simulation-length parameter, which minimizes a proxy for the autocorrelation of the state’s second moments. We evaluate ChEES-HMC and NUTS on many tasks, and find that ChEES-HMC typically yields larger effective sample sizes per gradient evaluation than NUTS does. When running many chains on a GPU, ChEES-HMC can also run significantly more gradient evaluations per second than NUTS, allowing it to quickly provide accurate estimates of posterior expectations. Matthew Hoffman 0001, Alexey Radul, Pavel Sountsov |
AISTATS | 3 |
| 2020 | Hamiltonian Monte Carlo SwindlesabstractHamiltonian Monte Carlo (HMC) is a powerful Markov chain Monte Carlo (MCMC) algorithm for estimating expectations with respect to continuous un-normalized probability distributions. MCMC estimators typically have higher variance than classical Monte Carlo with i.i.d. samples due to autocorrelations; most MCMC research tries to reduce these autocorrelations. In this work, we explore a complementary approach to variance reduction based on two classical Monte Carlo ’swindles’: first, running an auxiliary coupled chain targeting a tractable approximation to the target distribution, and using the auxiliary samples as control variates; and second, generating anti-correlated ("antithetic") samples by running two chains with flipped randomness. Both ideas have been explored previously in the context of Gibbs samplers and random-walk Metropolis algorithms, but we argue that they are ripe for adaptation to HMC in light of recent coupling results from the HMC theory literature. For many posterior distributions, we find that these swindles generate effective sample sizes orders of magnitude larger than plain HMC, as well as being more efficient than analogous swindles for Metropolis-adjusted Langevin algorithm and random-walk Metropolis. Dan Piponi, Matthew Hoffman 0001, Pavel Sountsov |
AISTATS | 3 |
| 2016 | Length bias in Encoder Decoder Models and a Case for Global ConditioningabstractEncoder-decoder networks are popular for modeling sequences probabilistically in many applications.These models use the power of the Long Short-Term Memory (LSTM) architecture to capture the full dependence among variables, unlike earlier models like CRFs that typically assumed conditional independence among non-adjacent variables.However in practice encoder-decoder models exhibit a bias towards short sequences that surprisingly gets worse with increasing beam size.In this paper we show that such phenomenon is due to a discrepancy between the full sequence margin and the per-element margin enforced by the locally conditioned training objective of a encoder-decoder model.The discrepancy more adversely impacts long sequences, explaining the bias towards predicting short sequences.For the case where the predicted sequences come from a closed set, we show that a globally conditioned model alleviates the above problems of encoder-decoder models.From a practical point of view, our proposed model also eliminates the need for a beam-search during inference, which reduces to an efficient dot-product based search in a vector-space. Pavel Sountsov, Sunita Sarawagi |
EMNLP | 1 |