Aniket Das

dblp:248/8281 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Optimization for machine learning · 29% Learning theory · 28% Probabilistic and Bayesian machine learning · 18%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 61% Hardware accelerators and domain-specific architectures · 30% Memory systems · 9%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 26 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
statistical estimation
1.422024
Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD · NeurIPS 2024
Near Optimal Heteroscedastic Regression with Symbiotic Learning · COLT 2023
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion probabilistic model
0.912025
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation · ICML 2025
Machine learning › Generative modeling
diffusion model
0.912025
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation · ICML 2025
Machine learning › Efficient and distributed learning
parallel inference
0.912025
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation · ICML 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.912025
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation · ICML 2025
Storage systems › computational storage
in-storage computing
0.912025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Storage systems
key-value storage
0.912025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Hardware accelerators and domain-specific architectures › accelerator integration
near-storage accelerator
0.912025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Machine learning › Optimization for machine learning
stochastic approximation
0.922023
Provably Fast Finite Particle Variants of SVGD via Virtual Particle Stochastic Approximation · NeurIPS 2023
Utilising the CLT Structure in Stochastic Gradient based Sampling : Improved Analysis and Faster Algorithms · COLT 2023
Machine learning › Learning theory › distribution learning
heavy-tailed distribution learning
0.812024
Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD · NeurIPS 2024
Machine learning › Optimization for machine learning › convex optimization
stochastic convex optimization
0.812024
Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD · NeurIPS 2024
Algorithms and data structures › data streams
streaming algorithms
0.812024
Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD · NeurIPS 2024
Machine learning › Optimization for machine learning
convergence analysis
0.712023
Utilising the CLT Structure in Stochastic Gradient based Sampling : Improved Analysis and Faster Algorithms · COLT 2023
Machine learning › Learning theory › probability metric
KL divergence
0.712023
Utilising the CLT Structure in Stochastic Gradient based Sampling : Improved Analysis and Faster Algorithms · COLT 2023
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates
0.712023
Near Optimal Heteroscedastic Regression with Symbiotic Learning · COLT 2023
Machine learning › Probabilistic and Bayesian machine learning
sampling
0.712023
Utilising the CLT Structure in Stochastic Gradient based Sampling : Improved Analysis and Faster Algorithms · COLT 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › particle-based variational inference
stein variational gradient descent
0.712023
Provably Fast Finite Particle Variants of SVGD via Virtual Particle Stochastic Approximation · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo › langevin dynamics
stochastic gradient langevin dynamics
0.712023
Utilising the CLT Structure in Stochastic Gradient based Sampling : Improved Analysis and Faster Algorithms · COLT 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.712023
Provably Fast Finite Particle Variants of SVGD via Virtual Particle Stochastic Approximation · NeurIPS 2023
Machine learning › Optimization for machine learning
minimax optimization
0.612022
Sampling without Replacement Leads to Faster Rates in Finite-Sum Minimax Optimization · NeurIPS 2022
Machine learning › Optimization for machine learning › mini-batch sampling
sampling without replacement
0.612022
Sampling without Replacement Leads to Faster Rates in Finite-Sum Minimax Optimization · NeurIPS 2022
Machine learning › Optimization for machine learning
stochastic gradient descent
0.612022
Sampling without Replacement Leads to Faster Rates in Finite-Sum Minimax Optimization · NeurIPS 2022
Machine learning › Generative modeling
autoregressive model
0.312025
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation · ICML 2025
Memory systems › processing-in-memory
near-data processing
0.312025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Machine learning › Optimization for machine learning
alternating minimization
0.212023
Near Optimal Heteroscedastic Regression with Symbiotic Learning · COLT 2023
Machine learning › Optimization for machine learning
non-convex optimization
0.212023
Near Optimal Heteroscedastic Regression with Symbiotic Learning · COLT 2023

Methods — techniques the papers use, named apart from their topics

martingale concentration · 1.5iterative refinement · 1.5clipped SGD · 1.5stochastic localization · 0.9speculative decoding · 0.9exchangeability · 0.9weighted least squares · 0.7pseudogradient descent · 0.7lecam's two-point method · 0.7central limit theorem · 0.7
YearPublicationVenuePosition
2025 Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation
abstract
Denoising Diffusion Probabilistic Models (DDPMs) have emerged as powerful tools for generative modeling. However, their sequential computation requirements lead to significant inference-time bottlenecks. In this work, we utilize the connection between DDPMs and Stochastic Localization to prove that, under an appropriate reparametrization, the increments of DDPM satisfy an exchangeability property. This general insight enables near-black-box adaptation of various performance optimization techniques from autoregressive models to the diffusion setting. To demonstrate this, we introduce Autospeculative Decoding (ASD), an extension of the widely used speculative decoding algorithm to DDPMs that does not require any auxiliary draft models. Our theoretical analysis shows that ASD achieves a $\tilde{O}(K^{\frac{1}{3}})$ parallel runtime speedup over the $K$ step sequential DDPM. We also demonstrate that a practical implementation of autospeculative decoding accelerates DDPM inference significantly in various domains.
Hengyuan Hu, Aniket Das, Dorsa Sadigh, Nima Anari
ICML2
2025 ANVIL: An In-Storage Accelerator for Name-Value Data Stores
abstract
Name-value pairs (NVPs) are a widely-used abstraction to organize data in millions of applications.At a high level, an NVP associates a name (e.g., array index, key, hash) with each value in a collection of data.Specific NVP data store formats can vary widely, ranging from simple arrays/dictionaries and lookup tables to key-value stores and data mining workloads.Despite their importance, existing optimizations for NVPs are limited to only a single data store format, as the broad definition of NVPs allows for significant heterogeneity in encoding and implementation.We propose ANVIL, the first end-to-end system that allows programmers to broadly accelerate most formats of NVPs.With a conventional solid-state drive (SSD), large-scale NVP lookups can saturate both external and internal SSD bandwidth, as every NVP in the data store needs to be sent back to the host CPU to check for a matching name.ANVIL makes use of in-storage processing to avoid reading out any data for names that do not match, by performing name match checks directly inside the SSD's NAND flash chips.We demonstrate that ANVIL can substantially reduce disk I/O, reduce metadata overheads, and provide speedups of 4.0×, 25×, and 14.6% over a conventional SSD, for three different NVP workloads (database transactions, analytics, and graph processing).
Ryan Wong 0001, Nikita Kim, Aniket Das, Kevin Higgs, Engin Ipek, Sapan Agarwal, Saugata Ghose, Ben Feinberg
ISCA3
2024 Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD
abstract
$\newcommand{\Tr}{\mathsf{Tr}}$ We consider the problem of high-dimensional heavy-tailed statistical estimation in the streaming setting, which is much harder than the traditional batch setting due to memory constraints. We cast this problem as stochastic convex optimization with heavy tailed stochastic gradients, and prove that the widely used Clipped-SGD algorithm attains near-optimal sub-Gaussian statistical rates whenever the second moment of the stochastic gradient noise is finite. More precisely, with $T$ samples, we show that Clipped-SGD, for smooth and strongly convex objectives, achieves an error of $\sqrt{\frac{\Tr(\Sigma)+\sqrt{\Tr(\Sigma)\\|\Sigma\\|_2}\ln(\tfrac{\ln(T)}{\delta})}{T}}$ with probability $1-\delta$, where $\Sigma$ is the covariance of the clipped gradient. Note that the fluctuations (depending on $\tfrac{1}{\delta}$) are of lower order than the term $\Tr(\Sigma)$. This improves upon the current best rate of $\sqrt{\frac{\Tr(\Sigma)\ln(\tfrac{1}{\delta})}{T}}$ for Clipped-SGD, known \emph{only} for smooth and strongly convex objectives. Our results also extend to smooth convex and lipschitz convex objectives. Key to our result is a novel iterative refinement strategy for martingale concentration, improving upon the PAC-Bayes approach of \citet{catoni2018dimension}.
Aniket Das, Dheeraj Nagaraj, Soumyabrata Pal, Arun Suggala, Prateek Varshney
NeurIPS1
2023 Near Optimal Heteroscedastic Regression with Symbiotic Learning
abstract
We consider the classical problem of heteroscedastic linear regression where, given $n$ i.i.d. samples $(\vx_i, y_i)$ drawn from the model $y_i = \dotp{\wstar}{\vx_i} + \epsilon_i \cdot \dotp{\fstar}{\vx_i}, \vx_i \sim \cN(0, \vI), \epsilon_i \sim \cN(0, 1)$, our aim is to \emph{estimate the regressor} $\wstar$ \emph{without prior knowledge of the noise parameter} $\fstar$. In addition to classical applications of such models in statistics \citep{jobson1980least}, econometrics \citep{harvey1976estimating}, time series analysis \citep{engle1982autoregressive} etc., it is also particularly relevant in machine learning problems where data is collected from multiple sources of varying (but apriori unknown) quality, e.g., in the training of large models \citep{devlin2019bert} on web-scale data. In this work, we develop an algorithm called \emph{\ouralg} (short for \emph{Symb}iotic \emph{Learn}ing) which estimates $\wstar$ in squared norm upto an error of $\Otilde(\| \fstar \|^2 \cdot (\nicefrac{1}{n} + (\nicefrac{d}{n})^2))$, and prove that this rate is minimax optimal modulo logarithmic factors. This represents a substantial improvement upon the previous best known upper bound of $\Otilde(\| \fstar \|^2 \cdot \nicefrac{d}{n})$. Our algorithm is essentially an alternating minimization procedure which comprises of two key subroutines 1. An adaptation of the classical weighted least squares heuristic to estimate $\wstar$ (dating back to at least \citet{davidian1987variance}), for which our work presents the first non-asymptotic guarantee; 2. A novel non-convex pseudogradient descent procedure for estimating $\fstar$, which draws inspiration from the phase retrieval literature. As corollaries of our analysis, we obtain fast non-asymptotic rates for two important problems, linear regression with multiplicative noise, and phase retrieval with multiplicative noise, both of which could be of independent interest. Beyond this, the proof of our lower bound, which involves a novel adaptation of LeCam’s two point method for handling infinite mutual information quantities (thereby preventing a direct application of standard techniques such as Fano’s method), could also be of broader interest for establishing lower bounds for other heteroscedastic or heavy tailed statistical problems.
Aniket Das, Dheeraj Nagaraj, Praneeth Netrapalli, Dheeraj Baby
COLT1
2023 Utilising the CLT Structure in Stochastic Gradient based Sampling : Improved Analysis and Faster Algorithms
abstract
We consider stochastic approximations of sampling algorithms, such as Stochastic Gradient Langevin Dynamics (SGLD) and the Random Batch Method (RBM) for Interacting Particle Dynamcs (IPD). We observe that the noise introduced by the stochastic approximation is nearly Gaussian due to the Central Limit Theorem (CLT) while the driving Brownian motion is exactly Gaussian. We harness this structure to absorb the stochastic approximation error inside the diffusion process, and obtain improved convergence guarantees for these algorithms. For SGLD, we prove the first stable convergence rate in KL divergence without requiring uniform warm start, assuming the target density satisfies a Log-Sobolev Inequality. Our result implies superior first-order oracle complexity compared to prior works, under significantly milder assumptions. We also prove the first guarantees for SGLD under even weaker conditions such as Hölder smoothness and Poincare Inequality, thus bridging the gap between the state-of-the-art guarantees for LMC and SGLD. Our analysis motivates a new algorithm called covariance correction, which corrects for the additional noise introduced by the stochastic approximation by rescaling the strength of the diffusion. Finally, we apply our techniques to analyze RBM, and significantly improve upon the guarantees in prior works (such as removing exponential dependence on horizon), under minimal assumptions.
Aniket Das, Dheeraj Nagaraj, Anant Raj
COLT1
2023 Provably Fast Finite Particle Variants of SVGD via Virtual Particle Stochastic Approximation
abstract
Stein Variational Gradient Descent (SVGD) is a popular particle-based variational inference algorithm with impressive empirical performance across various domains. Although the population (i.e, infinite-particle) limit dynamics of SVGD is well characterized, its behavior in the finite-particle regime is far less understood. To this end, our work introduces the notion of *virtual particles* to develop novel stochastic approximations of population-limit SVGD dynamics in the space of probability measures, that are exactly realizable using finite particles. As a result, we design two computationally efficient variants of SVGD, namely VP-SVGD and GB-SVGD, with provably fast finite-particle convergence rates. Our algorithms can be viewed as specific random-batch approximations of SVGD, which are computationally more efficient than ordinary SVGD. We show that the $n$ particles output by VP-SVGD and GB-SVGD, run for $T$ steps with batch-size $K$, are at-least as good as i.i.d samples from a distribution whose Kernel Stein Discrepancy to the target is at most $O(\tfrac{d^{1/3}}{(KT)^{1/6}})$ under standard assumptions. Our results also hold under a mild growth condition on the potential function, which is much weaker than the isoperimetric (e.g. Poincare Inequality) or information-transport conditions (e.g. Talagrand's Inequality $\mathsf{T}_1$) generally considered in prior works. As a corollary, we analyze the convergence of the empirical measure (of the particles output by VP-SVGD and GB-SVGD) to the target distribution and demonstrate a **double exponential improvement** over the best known finite-particle analysis of SVGD. Beyond this, our results present the **first known oracle complexities for this setting with polynomial dimension dependence**, thereby completely eliminating the curse of dimensionality exhibited by previously known finite-particle rates.
Aniket Das, Dheeraj Nagaraj
NeurIPS1
2022 Sampling without Replacement Leads to Faster Rates in Finite-Sum Minimax Optimization
abstract
We analyze the convergence rates of stochastic gradient algorithms for smooth finite-sum minimax optimization and show that, for many such algorithms, sampling the data points \emph{without replacement} leads to faster convergence compared to sampling with replacement. For the smooth and strongly convex-strongly concave setting, we consider gradient descent ascent and the proximal point method, and present a unified analysis of two popular without-replacement sampling strategies, namely \emph{Random Reshuffling} (RR), which shuffles the data every epoch, and \emph{Single Shuffling} or \emph{Shuffle Once} (SO), which shuffles only at the beginning. We obtain tight convergence rates for RR and SO and demonstrate that these strategies lead to faster convergence than uniform sampling. Moving beyond convexity, we obtain similar results for smooth nonconvex-nonconcave objectives satisfying a two-sided Polyak-\L{}ojasiewicz inequality. Finally, we demonstrate that our techniques are general enough to analyze the effect of \emph{data-ordering attacks}, where an adversary manipulates the order in which data points are supplied to the optimizer. Our analysis also recovers tight rates for the \emph{incremental gradient} method, where the data points are not shuffled at all.
Aniket Das, Bernhard Schölkopf, Michael Muehlebach
NeurIPS1
2020 Jointly Trained Image and Video Generation using Residual Vectors
abstract
In this work, we propose a modeling technique for jointly training image and video generation models by simultaneously learning to map latent variables with a fixed prior onto real images and interpolate over images to generate videos. The proposed approach models the variations in representations using residual vectors encoding the change at each time step over a summary vector for the entire video. We utilize the technique to jointly train an image generation model with a fixed prior along with a video generation model lacking constraints such as disentanglement. The joint training enables the image generator to exploit temporal information while the video generation model learns to flexibly share information across frames. Moreover, experimental results verify our approach's compatibility with pre-training on videos or images and training on datasets containing a mixture of both. A comprehensive set of quantitative and qualitative evaluations reveal the improvements in sample quality and diversity over both video generation and image generation baselines. We further demonstrate the technique's capabilities of exploiting similarity in features across frames by applying it to a model based on decomposing the video into motion and content. The proposed model allows minor variations in content across frames while maintaining the temporal dependence through latent vectors encoding the pose or motion features.
Yatin Dandi, Aniket Das, Soumye Singhal, Vinay P. Namboodiri, Piyush Rai
WACV2