EDBT 2026 Demo / reviewers in the wild / expert
Grant M. Rotskoff
dblp:220/5367
· DBLP profile ↗
10ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-7772-5179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Generative modeling · 43% Learning theory · 19% Optimization for machine learning · 12% | |
| Theoretical computer science
3 papers |
Algorithmic game theory and mechanism design · 69% Information theory · 21% Mathematical optimization · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
discrete diffusion model |
1.7 | 2 | 2025 | Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms · NeurIPS 2025 How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
1.6 | 2 | 2025 | How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework · ICLR 2025 Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024 |
Machine learning › Learning theory › statistical learning theory › statistical physics of learning
mean-field analysis |
1.1 | 3 | 2020 | A Dynamical Central Limit Theorem for Shallow Neural Networks · NeurIPS 2020 Neuron birth-death dynamics accelerates gradient descent and converges asymptotically · ICML 2019 Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks · NeurIPS 2018 |
Machine learning › Optimization for machine learning
convergence analysis |
1.1 | 2 | 2024 | Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024 Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks · NeurIPS 2018 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | Aligning Transformers with Continuous Feedback via Energy Rank Alignment · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks |
0.9 | 1 | 2025 | Features are fate: a theory of transfer learning in high-dimensional regression · ICML 2025 |
Machine learning › Generative modeling › diffusion model
diffusion model inference |
0.9 | 1 | 2025 | Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Features are fate: a theory of transfer learning in high-dimensional regression · ICML 2025 |
Machine learning › Generative modeling
molecular generation |
0.9 | 1 | 2025 | Aligning Transformers with Continuous Feedback via Energy Rank Alignment · NeurIPS 2025 |
Machine learning › Reinforcement learning
policy alignment |
0.9 | 1 | 2025 | Aligning Transformers with Continuous Feedback via Energy Rank Alignment · NeurIPS 2025 |
Machine learning › Generative modeling › protein design
protein sequence generation |
0.9 | 1 | 2025 | Aligning Transformers with Continuous Feedback via Energy Rank Alignment · NeurIPS 2025 |
Machine learning › Learning theory
transfer learning theory |
0.9 | 1 | 2025 | Features are fate: a theory of transfer learning in high-dimensional regression · ICML 2025 |
Machine learning › Deep learning architectures and training
neural network estimator |
0.8 | 1 | 2024 | Statistical Spatially Inhomogeneous Diffusion Inference · AAAI 2024 |
Machine learning › Learning theory › statistical estimation
nonparametric estimation |
0.8 | 1 | 2024 | Statistical Spatially Inhomogeneous Diffusion Inference · AAAI 2024 |
Machine learning › Generative modeling › diffusion model
parallel sampling |
0.8 | 1 | 2024 | Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024 |
Machine learning › Reinforcement learning
sample efficiency |
0.8 | 1 | 2024 | Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024 |
Machine learning › Learning theory
generalization bounds |
0.4 | 1 | 2020 | A Dynamical Central Limit Theorem for Shallow Neural Networks · NeurIPS 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | A mean-field analysis of two-player zero-sum games · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › dynamical system › neural dynamics
neural network dynamics |
0.4 | 1 | 2020 | A Dynamical Central Limit Theorem for Shallow Neural Networks · NeurIPS 2020 |
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts › nash equilibrium
mixed nash equilibrium |
0.4 | 1 | 2020 | A mean-field analysis of two-player zero-sum games · NeurIPS 2020 |
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium |
0.4 | 1 | 2020 | A mean-field analysis of two-player zero-sum games · NeurIPS 2020 |
Machine learning › Optimization for machine learning
convergence acceleration |
0.4 | 1 | 2019 | Neuron birth-death dynamics accelerates gradient descent and converges asymptotically · ICML 2019 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.4 | 1 | 2019 | Neuron birth-death dynamics accelerates gradient descent and converges asymptotically · ICML 2019 |
Machine learning › Deep learning architectures and training
training dynamics |
0.4 | 1 | 2019 | Neuron birth-death dynamics accelerates gradient descent and converges asymptotically · ICML 2019 |
Machine learning › Learning theory › approximation theory
neural network approximation |
0.3 | 1 | 2018 | Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks · NeurIPS 2018 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.3 | 1 | 2018 | Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks · NeurIPS 2018 |
Mathematical optimization › optimal transport
wasserstein gradient flow |
0.1 | 1 | 2020 | A Dynamical Central Limit Theorem for Shallow Neural Networks · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
tau-leaping · 2.6levy-type stochastic integrals · 1.7girsanov theorem · 1.7proximal policy optimization · 0.9phase diagram analysis · 0.9high-order numerical scheme · 0.9high-dimensional regression · 0.9gibbs-boltzmann distribution · 0.9energy rank alignment · 0.9direct preference optimization · 0.9neural network · 0.8minimax optimal rate analysis · 0.8wasserstein gradient flow · 0.4mirror descent · 0.4mean-field analysis · 0.4langevin dynamics · 0.4gradient descent-ascent · 0.4central limit theorem · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral FrameworkabstractDiscrete diffusion models have gained increasing attention for their ability to model complex distributions with tractable sampling and inference. However, the error analysis for discrete diffusion models remains less well-understood. In this work, we propose a comprehensive framework for the error analysis of discrete diffusion models based on Lévy-type stochastic integrals. By generalizing the Poisson random measure to that with a time-independent and state-dependent intensity, we rigorously establish a stochastic integral formulation of discrete diffusion models and provide the corresponding change of measure theorems that are intriguingly analogous to Itô integrals and Girsanov's theorem for their continuous counterparts. Our framework unifies and strengthens the current theoretical results on discrete diffusion models and obtains the first error bound for the $\tau$-leaping scheme in KL divergence. With error sources clearly identified, our analysis gives new insight into the mathematical properties of discrete diffusion models and offers guidance for the design of efficient and accurate algorithms for real-world discrete diffusion model applications. Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Lexing Ying |
ICLR | 3 |
| 2025 | Features are fate: a theory of transfer learning in high-dimensional regressionabstractWith the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity.
Fine-tuning, preference optimization, and transfer learning have all been successfully employed for these purposes when the target task closely resembles the source task, but a precise theoretical understanding of ``task similarity'' is still lacking.
We adopt a \emph{feature-centric} viewpoint on transfer learning and establish a number of theoretical results that demonstrate that when the target task is well represented by the feature space of the pre-trained model, transfer learning outperforms training from scratch.
We study deep linear networks as a minimal model of transfer learning in which we can analytically characterize the transferability phase diagram as a function of the target dataset size and the feature space overlap.
For this model, we establish rigorously that when the feature space overlap between the source and target tasks is sufficiently strong, both linear transfer and fine-tuning improve performance, especially in the low data limit.
These results build on an emerging understanding of feature learning dynamics in deep linear networks, and we demonstrate numerically that the rigorous results we derive for the linear case also apply to nonlinear networks. Javan Tahir, Surya Ganguli, Grant M. Rotskoff |
ICML | 3 |
| 2025 | Aligning Transformers with Continuous Feedback via Energy Rank AlignmentabstractSearching through chemical space is an exceptionally challenging problem because the number of possible molecules grows combinatorially with the number of atoms. Large, autoregressive models trained on databases of chemical compounds have yielded powerful generators, but we still lack robust strategies for generating molecules with desired properties. This molecular search problem closely resembles the "alignment" problem for large language models, though for many chemical tasks we have a specific and easily evaluable reward function. Here, we introduce an algorithm called energy rank alignment (ERA) that leverages an explicit reward function to produce a gradient-based objective that we use to optimize autoregressive policies.
We show theoretically that this algorithm is closely related to proximal policy optimization (PPO) and direct preference optimization (DPO), but has a minimizer that converges to an ideal Gibbs-Boltzmann distribution with the reward playing the role of an energy function.
Furthermore, this algorithm is highly scalable, does not require reinforcement learning, and performs well relative to DPO when the number of preference observations per pairing is small. We deploy this approach to align molecular transformers and protein language models to generate molecules and protein sequences, respectively, with externally specified
properties and find that it does so robustly, searching through diverse parts of chemical space. Shriram Chennakesavalu, Frank Hu, Sebastian Ibarraran, Grant M. Rotskoff |
NeurIPS | 4 |
| 2025 | Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsabstractDiscrete diffusion models have emerged as a powerful generative modeling framework for discrete data with successful applications spanning from text generation to image synthesis. However, their deployment faces challenges due to the high dimensionality of the state space, necessitating the development of efficient inference algorithms. Current inference approaches mainly fall into two categories: exact simulation and approximate methods such as $\tau$-leaping. While exact methods suffer from unpredictable inference time and redundant function evaluations, $\tau$-leaping is limited by its first-order accuracy. In this work, we advance the latter category by tailoring the first extension of high-order numerical inference schemes to discrete diffusion models, enabling larger step sizes while reducing error. We rigorously analyze the proposed schemes and establish the second-order accuracy of the $\theta$-Trapezoidal method in KL divergence. Empirical evaluations on GSM8K-level math-reasoning, GPT-2-level text, and ImageNet-level image generation tasks demonstrate that our method achieves superior sample quality compared to existing approaches under equivalent computational constraints, with consistent performance gains across models ranging from 200M to 8B. Our code is available at https://github.com/yuchen-zhu-zyc/DiscreteFastSolver Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Molei Tao, Lexing Ying |
NeurIPS | 6 |
| 2024 | Statistical Spatially Inhomogeneous Diffusion InferenceabstractInferring a diffusion equation from discretely observed measurements is a statistical challenge of significant importance in a variety of fields, from single-molecule tracking in biophysical systems to modeling financial instruments. Assuming that the underlying dynamical process obeys a d-dimensional stochastic differential equation of the form dx_t = b(x_t)dt + \Sigma(x_t)dw_t, we propose neural network-based estimators of both the drift b and the spatially-inhomogeneous diffusion tensor D = \Sigma\Sigma^T/2 and provide statistical convergence guarantees when b and D are s-Hölder continuous. Notably, our bound aligns with the minimax optimal rate N^{-\frac{2s}{2s+d}} for nonparametric function estimation even in the presence of correlation within observational data, which necessitates careful handling when establishing fast-rate generalization bounds. Our theoretical results are bolstered by numerical experiments demonstrating accurate inference of spatially-inhomogeneous diffusion tensors. Yinuo Ren, Yiping Lu 0001, Lexing Ying, Grant M. Rotskoff |
AAAI | 4 |
| 2024 | Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time ComplexityabstractDiffusion models have become a leading method for generative modeling of both image and scientific data.
As these models are costly to train and \emph{evaluate}, reducing the inference cost for diffusion models remains a major goal.
Inspired by the recent empirical success in accelerating diffusion models via the parallel sampling technique~\cite{shih2024parallel}, we propose to divide the sampling process into $\mathcal{O}(1)$ blocks with parallelizable Picard iterations within each block. Rigorous theoretical analysis reveals that our algorithm achieves $\widetilde{\mathcal{O}}(\mathrm{poly} \log d)$ overall time complexity, marking \emph{the first implementation with provable sub-linear complexity w.r.t. the data dimension $d$}. Our analysis is based on a generalized version of Girsanov's theorem and is compatible with both the SDE and probability flow ODE implementations. Our results shed light on the potential of fast and efficient sampling of high-dimensional data on fast-evolving modern large-memory GPU clusters. Haoxuan Chen, Yinuo Ren, Lexing Ying, Grant M. Rotskoff |
NeurIPS | 4 |
| 2020 | A Dynamical Central Limit Theorem for Shallow Neural NetworksabstractRecent theoretical work has characterized the dynamics and convergence properties for wide shallow neural networks trained via gradient descent; the asymptotic regime in which the number of parameters tends towards infinity has been dubbed the "mean-field" limit. At initialization, the randomly sampled parameters lead to a deviation from the mean-field limit that is dictated by the classical central limit theorem (CLT). However, the dynamics of training introduces correlations among the parameters raising the question of how the fluctuations evolve during training. Here, we analyze the mean-field dynamics as a Wasserstein gradient flow and prove that the deviations from the mean-field evolution scaled by the width, in the width-asymptotic limit, remain bounded throughout training. This observation has implications for both the approximation rate and the generalization: the upper bound we obtain is controlled by a Monte-Carlo type resampling error, which importantly does not depend on dimension. We also relate the bound on the fluctuations to the total variation norm of the measure to which the dynamics converges, which in turn controls the generalization error. Zhengdao Chen, Grant M. Rotskoff, Joan Bruna, Eric Vanden-Eijnden |
NeurIPS | 2 |
| 2020 | A mean-field analysis of two-player zero-sum gamesabstractFinding Nash equilibria in two-player zero-sum continuous games is a central problem in machine learning, e.g. for training both GANs and robust models. The existence of pure Nash equilibria requires strong conditions which are not typically met in practice. Mixed Nash equilibria exist in greater generality and may be found using mirror descent. Yet this approach does not scale to high dimensions. To address this limitation, we parametrize mixed strategies as mixtures of particles, whose positions and weights are updated using gradient descent-ascent. We study this dynamics as an interacting gradient flow over measure spaces endowed with the Wasserstein-Fisher-Rao metric. We establish global convergence to an approximate equilibrium for the related Langevin gradient-ascent dynamic. We prove a law of large numbers that relates particle dynamics to mean-field dynamics. Our method identifies mixed equilibria in high dimensions and is demonstrably effective for training mixtures of GANs. Carles Domingo-Enrich, Samy Jelassi, Arthur Mensch, Grant M. Rotskoff, Joan Bruna |
NeurIPS | 4 |
| 2019 | Neuron birth-death dynamics accelerates gradient descent and converges asymptoticallyabstractNeural networks with a large number of parameters admit a mean-field description, which has recently served as a theoretical explanation for the favorable training properties of models with a large number of parameters. In this regime, gradient descent obeys a deterministic partial differential equation (PDE) that converges to a globally optimal solution for networks with a single hidden layer under appropriate assumptions. In this work, we propose a non-local mass transport dynamics that leads to a modified PDE with the same minimizer. We implement this non-local dynamics as a stochastic neuronal birth/death process and we prove that it accelerates the rate of convergence in the mean-field limit. We subsequently realize this PDE with two classes of numerical schemes that converge to the mean-field equation, each of which can easily be implemented for neural networks with finite numbers of parameters. We illustrate our algorithms with two models to provide intuition for the mechanism through which convergence is accelerated. Grant M. Rotskoff, Samy Jelassi, Joan Bruna, Eric Vanden-Eijnden |
ICML | 1 |
| 2018 | Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networksabstractThe performance of neural networks on high-dimensional data distributions suggests that it may be possible to parameterize a representation of a given high-dimensional function with controllably small errors, potentially outperforming standard interpolation methods. We demonstrate, both theoretically and numerically, that this is indeed the case. We map the parameters of a neural network to a system of particles relaxing with an interaction potential determined by the loss function. We show that in the limit that the number of parameters $n$ is large, the landscape of the mean-squared error becomes convex and the representation error in the function scales as $O(n^{-1})$. In this limit, we prove a dynamical variant of the universal approximation theorem showing that the optimal representation can be attained by stochastic gradient descent, the algorithm ubiquitously used for parameter optimization in machine learning. In the asymptotic regime, we study the fluctuations around the optimal representation and show that they arise at a scale $O(n^{-1})$. These fluctuations in the landscape identify the natural scale for the noise in stochastic gradient descent. Our results apply to both single and multi-layer neural networks, as well as standard kernel methods like radial basis functions. Grant M. Rotskoff, Eric Vanden-Eijnden |
NeurIPS | 1 |