Promit Ghosal

dblp:319/5056 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 65% Learning theory · 23% Optimization for machine learning · 12%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › particle-based variational inference
stein variational gradient descent
1.522025
Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent · ICLR 2025
Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.522025
Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent · ICLR 2025
Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent · NeurIPS 2023
Machine learning › Learning theory › statistical learning theory
asymptotic analysis
1.012026
Statistical Inference for Linear Functionals of Online Least-Squares SGD When t ≳ d1+δ · IEEE Trans. Inf. Theory 2026
Machine learning › Optimization for machine learning
stochastic gradient descent
1.012026
Statistical Inference for Linear Functionals of Online Least-Squares SGD When t ≳ d1+δ · IEEE Trans. Inf. Theory 2026
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.812024
Hybrid Top-Down Global Causal Discovery with Local Search for Linear and Nonlinear Additive Noise Models · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › sampling
particle-based sampling
0.712023
Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning
sampling
0.712023
Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent · NeurIPS 2023
Mathematical optimization
optimal transport
0.512021
Rates of Estimation of Optimal Transport Maps using Plug-in Estimators via Barycentric Projections · NeurIPS 2021
Mathematical optimization › optimal transport
wasserstein distance estimation
0.512021
Rates of Estimation of Optimal Transport Maps using Plug-in Estimators via Barycentric Projections · NeurIPS 2021
Mathematical optimization › optimal transport
wasserstein barycenter
0.112021
Rates of Estimation of Optimal Transport Maps using Plug-in Estimators via Barycentric Projections · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

stein variational gradient descent · 2.4propagation of chaos · 1.7kernel methods · 1.7online variance estimation · 1.0berry-esseen bound · 1.0topological sorting · 0.8constraint-based algorithm · 0.8additive noise model · 0.8mean-field approximation · 0.7bilinear kernel · 0.7wavelet smoothing · 0.5kernel smoothing · 0.5barycentric projection · 0.5
YearPublicationVenuePosition
2026 Statistical Inference for Linear Functionals of Online Least-Squares SGD When t ≳ d1+δ
abstract
Stochastic Gradient Descent (SGD) has become a cornerstone method in modern data science. However, deploying SGD in high-stakes applications necessitates rigorous quantification of its inherent uncertainty. In this work, we establishnon-asymptotic Berry–Esseen boundsfor linear functionals of online least-squares SGD, thereby providing a Gaussian Central Limit Theorem (CLT) in agrowing-dimensional regime. Existing approaches to high-dimensional inference for projection parameters, such as [1], rely on inverting empirical covariance matrices and require at leastt≳d3/2iterations to achieve finite-sample Berry–Esseen guarantees, rendering them computationally expensive and restrictive in the allowable dimensional scaling. In contrast, we show that a CLT holds for SGD iterates when the number of iterations grows ast≳d1+δfor any δ > 0, significantly extending the dimensional regime permitted by prior works while improving computational efficiency. The proposed online SGD-based procedure operates inO(td) time and requires onlyO(d) memory, in contrast to theO(td2+d3) runtime of covariance-inversion methods. To render the theory practically applicable, we further develop anonline variance estimatorfor the asymptotic variance appearing in the CLT and establishhigh-probability deviation boundsfor this estimator. Collectively, these results yield the first fully online and data-driven framework for constructing confidence intervals for SGD iterates in the near-optimal scaling regimet≳d1+δ.
Bhavya Agrawalla, Krishnakumar Balasubramanian 0002, Promit Ghosal
IEEE Trans. Inf. Theory3
2025 Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent
abstract
We provide finite-particle convergence rates for the Stein Variational Gradient Descent (SVGD) algorithm in the Kernelized Stein Discrepancy ($\KSD$) and Wasserstein-2 metrics. Our key insight is that the time derivative of the relative entropy between the joint density of $N$ particle locations and the $N$-fold product target measure, starting from a regular initial distribution, splits into a dominant 'negative part' proportional to $N$ times the expected $\KSD^2$ and a smaller 'positive part'. This observation leads to $\KSD$ rates of order $1/\sqrt{N}$, in both continuous and discrete time, providing a near optimal (in the sense of matching the corresponding i.i.d. rates) double exponential improvement over the recent result by~\cite{shi2024finite}. Under mild assumptions on the kernel and potential, these bounds also grow polynomially in the dimension $d$. By adding a bilinear component to the kernel, the above approach is used to further obtain Wasserstein-2 convergence in continuous time. For the case of `bilinear + Mat\'ern' kernels, we derive Wasserstein-2 rates that exhibit a curse-of-dimensionality similar to the i.i.d. setting. We also obtain marginal convergence and long-time propagation of chaos results for the time-averaged particle laws.
Sayan Banerjee, Krishna Balasubramanian, Promit Ghosal
ICLR3
2025 When Additive Noise Meets Unobserved Mediators: Bivariate Denoising Diffusion for Causal Discovery
abstract
Distinguishing cause and effect from bivariate observational data is a foundational problem in many disciplines, but challenging without additional assumptions. Additive noise models (ANMs) are widely used to enable sample-efficient bivariate causal discovery. However, conventional ANM-based methods fail when unobserved mediators corrupt the causal relationship between variables. This paper makes three key contributions: first, we rigorously characterize why standard ANM approaches break down in the presence of unmeasured mediators. Second, we demonstrate that prior solutions for hidden mediation are brittle in finite sample settings, limiting their practical utility. To address these gaps, we propose Bivariate Denoising Diffusion (BiDD) for causal discovery, a method designed to handle latent noise introduced by unmeasured mediators. Unlike prior methods that infer directionality through mean squared error loss comparisons, our approach introduces a novel independence test statistic: during the noising and denoising processes for each variable, we condition on the other variable as input and evaluate the independence of the predicted noise relative to this input. We prove asymptotic consistency of BiDD under the ANM, and conjecture that it performs well under hidden mediation. Experiments on synthetic and real-world data demonstrate consistent performance, outperforming existing methods in mediator-corrupted settings while maintaining strong performance in mediator-free settings.
Dominik Meier, Sujai Hiremath, Promit Ghosal, Kyra Gan
NeurIPS3
2025 LoSAM: Local Search in Additive Noise Models with Mixed Mechanisms and General Noise for Global Causal Discovery
abstract
Inferring causal relationships from observational data is crucial when experiments are costly or infeasible. Additive noise models (ANMs) enable unique directed acyclic graph (DAG) identification, but existing sample-efficient ANM methods often rely on restrictive assumptions on the data generating process, limiting their applicability to real-world settings. We propose local search in additive noise models, LoSAM, a topological ordering method for learning a unique DAG in ANMs with mixed causal mechanisms and general noise distributions. We introduce new causal substructures and criteria for identifying roots and leaves, enabling efficient top-down learning. We prove asymptotic consistency and polynomial runtime, ensuring scalability and sample efficiency. We test LoSAM on synthetic and real-world data, demonstrating state-of-the-art performance across all mixed mechanism settings.
Sujai Hiremath, Promit Ghosal, Kyra Gan
UAI2
2024 Hybrid Top-Down Global Causal Discovery with Local Search for Linear and Nonlinear Additive Noise Models
abstract
Learning the unique directed acyclic graph corresponding to an unknown causal model is a challenging task. Methods based on functional causal models can identify a unique graph, but either suffer from the curse of dimensionality or impose strong parametric assumptions. To address these challenges, we propose a novel hybrid approach for global causal discovery in observational data that leverages local causal substructures. We first present a topological sorting algorithm that leverages ancestral relationships in linear structural causal models to establish a compact top-down hierarchical ordering, encoding more causal information than linear orderings produced by existing methods. We demonstrate that this approach generalizes to nonlinear settings with arbitrary noise. We then introduce a nonparametric constraint-based algorithm that prunes spurious edges by searching for local conditioning sets, achieving greater accuracy than current methods. We provide theoretical guarantees for correctness and worst-case polynomial time complexities, with empirical validation on synthetic data.
Sujai Hiremath, Jacqueline R. M. A. Maasch, Mengxiao Gao, Promit Ghosal, Kyra Gan
NeurIPS4
2023 Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent
abstract
Stein Variational Gradient Descent (SVGD) is a nonparametric particle-based deterministic sampling algorithm. Despite its wide usage, understanding the theoretical properties of SVGD has remained a challenging problem. For sampling from a Gaussian target, the SVGD dynamics with a bilinear kernel will remain Gaussian as long as the initializer is Gaussian. Inspired by this fact, we undertake a detailed theoretical study of the Gaussian-SVGD, i.e., SVGD projected to the family of Gaussian distributions via the bilinear kernel, or equivalently Gaussian variational inference (GVI) with SVGD. We present a complete picture by considering both the mean-field PDE and discrete particle systems. When the target is strongly log-concave, the mean-field Gaussian-SVGD dynamics is proven to converge linearly to the Gaussian distribution closest to the target in KL divergence. In the finite-particle setting, there is both uniform in time convergence to the mean-field limit and linear convergence in time to the equilibrium if the target is Gaussian. In the general case, we propose a density-based and a particle-based implementation of the Gaussian-SVGD, and show that several recent algorithms for GVI, proposed from different perspectives, emerge as special cases of our unified framework. Interestingly, one of the new particle-based instance from this framework empirically outperforms existing approaches. Our results make concrete contributions towards obtaining a deeper understanding of both SVGD and GVI.
Promit Ghosal, Krishnakumar Balasubramanian 0002, Natesh S. Pillai
NeurIPS2
2021 Rates of Estimation of Optimal Transport Maps using Plug-in Estimators via Barycentric Projections
abstract
Optimal transport maps between two probability distributions $\mu$ and $\nu$ on $\R^d$ have found extensive applications in both machine learning and statistics. In practice, these maps need to be estimated from data sampled according to $\mu$ and $\nu$. Plug-in estimators are perhaps most popular in estimating transport maps in the field of computational optimal transport. In this paper, we provide a comprehensive analysis of the rates of convergences for general plug-in estimators defined via barycentric projections. Our main contribution is a new stability estimate for barycentric projections which proceeds under minimal smoothness assumptions and can be used to analyze general plug-in estimators. We illustrate the usefulness of this stability estimate by first providing rates of convergence for the natural discrete-discrete and semi-discrete estimators of optimal transport maps. We then use the same stability estimate to show that, under additional smoothness assumptions of Besov type or Sobolev type, wavelet based or kernel smoothed plug-in estimators respectively speed up the rates of convergence and significantly mitigate the curse of dimensionality suffered by the natural discrete-discrete/semi-discrete estimators. As a by-product of our analysis, we also obtain faster rates of convergence for plug-in estimators of $W_2(\mu,\nu)$, the Wasserstein distance between $\mu$ and $\nu$, under the aforementioned smoothness assumptions, thereby complementing recent results in Chizat et al. (2020). Finally, we illustrate the applicability of our results in obtaining rates of convergence for Wasserstein barycenters between two probability distributions and obtaining asymptotic detection thresholds for some recent optimal-transport based tests of independence.
Nabarun Deb, Promit Ghosal, Bodhisattva Sen
NeurIPS2