Shreyas Padhy

dblp:267/9851 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Probabilistic and Bayesian machine learning · 38% Generative modeling · 24% Trustworthy machine learning · 13%
Theoretical computer science
2 papers
Mathematical optimization · 50% Algorithms and data structures · 50%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
2.642024
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes · NeurIPS 2024
Stochastic Gradient Descent for Gaussian Processes Done Right · ICLR 2024
Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent · NeurIPS 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
1.122023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness · NeurIPS 2020
Machine learning › Generative modeling
diffusion model
1.022024
DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024
Transport meets Variational Inference: Controlled Monte Carlo Diffusions · ICLR 2024
Machine learning › Generative modeling › diffusion model
conditional sampling
0.812024
DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
0.812024
Stochastic Gradient Descent for Gaussian Processes Done Right · ICLR 2024
Machine learning › Optimization for machine learning
hyperparameter optimization
0.812024
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
inverse problem solving
0.812024
DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
monte carlo diffusion
0.812024
Transport meets Variational Inference: Controlled Monte Carlo Diffusions · ICLR 2024
Machine learning › Representation and self-supervised learning
symmetry learning
0.812024
A Generative Model of Symmetry Transformations · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.812024
Transport meets Variational Inference: Controlled Monte Carlo Diffusions · ICLR 2024
Mathematical optimization › iterative methods
conjugate gradient method
0.812024
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes · NeurIPS 2024
Algorithms and data structures › numerical linear algebra
linear system solving
0.812024
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes · NeurIPS 2024
Algorithms and data structures
numerical linear algebra
0.812024
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes · NeurIPS 2024
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent
0.812024
Stochastic Gradient Descent for Gaussian Processes Done Right · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.712023
Sampling-based inference for large linear models, with application to linearised Laplace · ICLR 2023
Machine learning › Trustworthy machine learning › uncertainty estimation › predictive uncertainty
distance-aware uncertainty
0.712023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Learning theory
implicit bias
0.712023
Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling
0.712023
Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › sampling
sampling-based inference
0.712023
Sampling-based inference for large linear models, with application to linearised Laplace · ICLR 2023
Machine learning › Optimization for machine learning
stochastic gradient descent
0.712023
Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent · NeurIPS 2023
Computer vision › 3D vision › depth perception
distance perception
0.412020
Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness · NeurIPS 2020
Machine learning › Deep learning architectures and training › normalization
spectral normalization
0.412020
Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
linearized laplace approximation
0.212023
Sampling-based inference for large linear models, with application to linearised Laplace · ICLR 2023
Machine learning › Trustworthy machine learning › calibration
neural network calibration
0.212023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.212023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023

Methods — techniques the papers use, named apart from their topics

variational approximation · 1.5stochastic dual descent · 1.5preconditioned conjugate gradient · 1.5minimax learning · 1.1warm-starting · 0.8score-based annealing · 0.8schroedinger bridges · 0.8pathwise gradient estimator · 0.8optimal transport · 0.8group theory · 0.8early stopping · 0.8doob's h-transform · 0.8data augmentation · 0.8
YearPublicationVenuePosition
2024 Transport meets Variational Inference: Controlled Monte Carlo Diffusions
abstract
Connecting optimal transport and variational inference, we present a principled and systematic framework for sampling and generative modelling centred around divergences on path space. Our work culminates in the development of the Controlled Monte Carlo Diffusion sampler (CMCD) for Bayesian computation, a score-based annealing technique that crucially adapts both forward and backward dynamics in a diffusion model. On the way, we clarify the relationship between the EM-algorithm and iterative proportional fitting (IPF) for Schroedinger bridges, deriving as well a regularised objective that bypasses the iterative bottleneck of standard IPF-updates. Finally, we show that CMCD has a strong foundation in the Jarzinsky and Crooks identities from statistical physics, and that it convincingly outperforms competing approaches across a wide array of experiments.
Francisco Vargas 0001, Shreyas Padhy, Denis Blessing, Nikolas Nüsken
ICLR2
2024 Stochastic Gradient Descent for Gaussian Processes Done Right
abstract
As is well known, both sampling from the posterior and computing the mean of the posterior in Gaussian process regression reduces to solving a large linear system of equations. We study the use of stochastic gradient descent for solving this linear system, and show that when done right---by which we mean using specific insights from the optimisation and kernel communities---stochastic gradient descent is highly effective. To that end, we introduce a particularly simple stochastic dual descent algorithm, explain its design in an intuitive manner and illustrate the design choices through a series of ablation studies. Further experiments demonstrate that our new method is highly competitive. In particular, our evaluations on the UCI regression tasks and on Bayesian optimisation set our approach apart from preconditioned conjugate gradients and variational Gaussian process approximations. Moreover, our method places Gaussian process regression on par with state-of-the-art graph neural networks for molecular binding affinity prediction.
Jihao Andreas Lin, Shreyas Padhy, Javier Antorán, Austin Tripp, Alexander Terenin, Csaba Szepesvári, José Miguel Hernández-Lobato, David Janz
ICLR2
2024 A Generative Model of Symmetry Transformations
abstract
Correctly capturing the symmetry transformations of data can lead to efficient models with strong generalization capabilities, though methods incorporating symmetries often require prior knowledge. While recent advancements have been made in learning those symmetries directly from the dataset, most of this work has focused on the discriminative setting. In this paper, we take inspiration from group theoretic ideas to construct a generative model that explicitly aims to capture the data's approximate symmetries. This results in a model that, given a prespecified broad set of possible symmetries, learns to what extent, if at all, those symmetries are actually present. Our model can be seen as a generative process for data augmentation. We provide a simple algorithm for learning our generative model and empirically demonstrate its ability to capture symmetries under affine and color transformations, in an interpretable way. Combining our symmetry model with standard generative models results in higher marginal test-log-likelihoods and improved data efficiency.
James Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán, David Krueger 0001, Richard E. Turner, Eric T. Nalisnick, José Miguel Hernández-Lobato
NeurIPS3
2024 DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform
abstract
Generative modelling paradigms based on denoising diffusion processes have emerged as a leading candidate for conditional sampling in inverse problems. In many real-world applications, we often have access to large, expensively trained unconditional diffusion models, which we aim to exploit for improving conditional sampling. Most recent approaches are motivated heuristically and lack a unifying framework, obscuring connections between them. Further, they often suffer from issues such as being very sensitive to hyperparameters, being expensive to train or needing access to weights hidden behind a closed API. In this work, we unify conditional training and sampling using the mathematically well-understood Doob's h-transform. This new perspective allows us to unify many existing methods under a common umbrella. Under this framework, we propose DEFT (Doob's h-transform Efficient FineTuning), a new approach for conditional generation that simply fine-tunes a very small network to quickly learn the conditional $h$-transform, while keeping the larger unconditional network unchanged. DEFT is much faster than existing baselines while achieving state-of-the-art performance across a variety of linear and non-linear benchmarks. On image reconstruction tasks, we achieve speedups of up to 1.6$\times$, while having the best perceptual quality on natural images and reconstruction performance on medical images. Further, we also provide initial experiments on protein motif scaffolding and outperform reconstruction guidance methods.
Alexander Denker, Francisco Vargas 0001, Shreyas Padhy, Kieran Didi, Simon V. Mathis, Riccardo Barbano, Vincent Dutordoir, Emile Mathieu, Urszula Julia Komorowska, Pietro Liò
NeurIPS3
2024 Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes
abstract
Scaling hyperparameter optimisation to very large datasets remains an open problem in the Gaussian process community. This paper focuses on iterative methods, which use linear system solvers, like conjugate gradients, alternating projections or stochastic gradient descent, to construct an estimate of the marginal likelihood gradient. We discuss three key improvements which are applicable across solvers: (i) a pathwise gradient estimator, which reduces the required number of solver iterations and amortises the computational cost of making predictions, (ii) warm starting linear system solvers with the solution from the previous step, which leads to faster solver convergence at the cost of negligible bias, (iii) early stopping linear system solvers after a limited computational budget, which synergises with warm starting, allowing solver progress to accumulate over multiple marginal likelihood steps. These techniques provide speed-ups of up to $72\times$ when solving to tolerance, and decrease the average residual norm by up to $7\times$ when stopping early.
Jihao Andreas Lin, Shreyas Padhy, Bruno Mlodozeniec, Javier Antorán, José Miguel Hernández-Lobato
NeurIPS2
2023 Sampling-based inference for large linear models, with application to linearised Laplace
Javier Antorán, Shreyas Padhy, Riccardo Barbano, Eric T. Nalisnick, David Janz, José Miguel Hernández-Lobato
ICLR2
2023 Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent
abstract
Gaussian processes are a powerful framework for quantifying uncertainty and for sequential decision-making but are limited by the requirement of solving linear systems. In general, this has a cubic cost in dataset size and is sensitive to conditioning. We explore stochastic gradient algorithms as a computationally efficient method of approximately solving these linear systems: we develop low-variance optimization objectives for sampling from the posterior and extend these to inducing points. Counterintuitively, stochastic gradient descent often produces accurate predictions, even in cases where it does not converge quickly to the optimum. We explain this through a spectral characterization of the implicit bias from non-convergence. We show that stochastic gradient descent produces predictive distributions close to the true posterior both in regions with sufficient data coverage, and in regions sufficiently far away from the data. Experimentally, stochastic gradient descent achieves state-of-the-art performance on sufficiently large-scale or ill-conditioned regression tasks. Its uncertainty estimates match the performance of significantly more expensive baselines on a large-scale Bayesian~optimization~task.
Jihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz, José Miguel Hernández-Lobato, Alexander Terenin
NeurIPS3
2023 A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness
abstract
Accurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distribution (OOD) inputs. The most popular approaches to estimate predictive uncertainty in deep learning are methods that combine predictions from multiple neural networks, such as Bayesian neural networks (BNNs) and deep ensembles. However their practicality in real-time, industrial-scale applications are limited due to the high memory and computational cost. Furthermore, ensembles and BNNs do not necessarily fix all the issues with the underlying member networks. In this work, we study principled approaches to improve the uncertainty property of a single network, based on a single, deterministic representation. By formalizing the uncertainty quantification as a minimax learning problem, we first identify distance awareness, i.e., the model's ability to quantify the distance of a testing example from the training data, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs with two simple changes: (1) applying spectral normalization to hidden weights to enforce bi-Lipschitz smoothness in representations and (2) replacing the last output layer with a Gaussian process layer. On a suite of vision and language understanding benchmarks and on modern architectures (Wide-ResNet and BERT), SNGP consistently outperforms other single-model approaches in prediction, calibration and out-of-domain detection. Furthermore, SNGP provides complementary benefits to popular techniques such as deep ensembles and data augmentation, making it a simple and scalable building block for probabilistic deep learning.
Jeremiah Z. Liu, Shreyas Padhy, Jie Ren 0006, Zi Lin, Yeming Wen, Ghassen Jerfel, Zachary Nado, Jasper Snoek, Dustin Tran, Balaji Lakshminarayanan
J. Mach. Learn. Res.2
2020 Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
abstract
Bayesian neural networks (BNN) and deep ensembles are principled approaches to estimate the predictive uncertainty of a deep learning model. However their practicality in real-time, industrial-scale applications are limited due to their heavy memory and inference cost. This motivates us to study principled approaches to high-quality uncertainty estimation that require only a single deep neural network (DNN). By formalizing the uncertainty quantification as a minimax learning problem, we first identify input distance awareness, i.e., the model’s ability to quantify the distance of a testing example from the training data in the input space, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs, by adding a weight normalization step during training and replacing the output layer. On a suite of vision and language understanding tasks and on modern architectures (Wide-ResNet and BERT), SNGP is competitive with deep ensembles in prediction, calibration and out-of-domain detection, and outperforms the other single-model approaches.
Jeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax-Weiss, Balaji Lakshminarayanan
NeurIPS3