VLDB 2026 Research / reviewers in the wild / expert
Elliot Paquette
dblp:126/6986
· DBLP profile ↗
10ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Optimization for machine learning · 66% Deep learning architectures and training · 21% Learning theory · 14% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
stochastic gradient descent |
4.0 | 6 | 2025 | To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions · ICLR 2025 4+3 Phases of Compute-Optimal Neural Scaling Laws · NeurIPS 2024 The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms · NeurIPS 2024 |
Machine learning › Optimization for machine learning
stochastic optimization |
2.2 | 3 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effects · ICML 2025 Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models · NeurIPS 2021 |
Machine learning › Optimization for machine learning
convergence analysis |
1.8 | 3 | 2024 | The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms · NeurIPS 2024 Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High Dimensions · NeurIPS 2022 Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › training dynamics
optimization dynamics |
1.7 | 2 | 2025 | Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effects · ICML 2025 To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions · ICLR 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
1.6 | 2 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 4+3 Phases of Compute-Optimal Neural Scaling Laws · NeurIPS 2024 |
Machine learning › Optimization for machine learning › stochastic gradient descent
stochastic gradient descent with momentum |
1.4 | 2 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High Dimensions · NeurIPS 2022 |
Machine learning › Optimization for machine learning
gradient clipping |
0.9 | 1 | 2025 | To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions · ICLR 2025 |
Machine learning › Learning theory › high-dimensional statistics
high-dimensional analysis |
0.9 | 1 | 2025 | To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions · ICLR 2025 |
Machine learning › Optimization for machine learning › stochastic gradient descent
signSGD |
0.9 | 1 | 2025 | Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effects · ICML 2025 |
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate |
0.8 | 1 | 2024 | The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › scaling laws
compute-optimal scaling |
0.8 | 1 | 2024 | 4+3 Phases of Compute-Optimal Neural Scaling Laws · NeurIPS 2024 |
Machine learning › Learning theory
learning dynamics |
0.8 | 1 | 2024 | 4+3 Phases of Compute-Optimal Neural Scaling Laws · NeurIPS 2024 |
Machine learning › Learning theory
generalization bounds |
0.6 | 1 | 2022 | Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions · NeurIPS 2022 |
Machine learning › Optimization for machine learning
implicit regularization |
0.6 | 1 | 2022 | Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions · NeurIPS 2022 |
Machine learning › Learning theory › computational complexity
average-case analysis |
0.5 | 1 | 2021 | SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality · COLT 2021 |
Machine learning › Optimization for machine learning › stochastic optimization
stochastic momentum methods |
0.5 | 1 | 2021 | Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
stochastic gradient descent · 3.0random matrix theory · 1.6stochastic differential equation · 1.4nesterov acceleration · 1.4volterra integral equation · 1.1ordinary differential equation · 0.9gradient clipping · 0.9mean squared loss · 0.8exact line search · 0.8adagrad-norm · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-DimensionsabstractThe success of modern machine learning is due in part to the adaptive optimization methods that have been developed to deal with the difficulties of training large models over complex datasets. One such method is gradient clipping: a practical procedure with limited theoretical underpinnings. In this work, we study clipping in a least squares problem under streaming SGD. We develop a theoretical analysis of the learning dynamics in the limit of large intrinsic dimension—a model and dataset dependent notion of dimensionality. In this limit we find a deterministic equation that describes the evolution of the loss and demonstrate that this equation predicts the path of clipped SGD on synthetic, CIFAR10, and Wikitext2 data. We show that with Gaussian noise clipping cannot improve SGD performance. Yet, in other noisy settings, clipping can provide benefits with tuning of the clipping threshold. We propose a simple heuristic for near optimal scheduling of the clipping threshold which requires the tuning of only one hyperparameter. We conclude with a discussion about the links between high-dimensional clipping and neural network training. Noah Marshall, Ke Liang Xiao, Atish Agarwala, Elliot Paquette |
ICLR | 4 |
| 2025 | Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effectsabstractIn recent years, SignSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consensus that SignSGD acts to precondition optimization and reshapes noise, quantitatively understanding these effects in theoretically solvable settings remains difficult. We present an analysis of SignSGD in a high dimensional limit, and derive a limiting SDE and ODE to describe the risk. Using this framework we quantify four effects of SignSGD: effective learning rate, noise compression, diagonal preconditioning, and gradient noise reshaping. Our analysis is consistent with experimental observations but moves beyond that by quantifying the dependence of these effects on the data and noise distributions. We conclude with a conjecture on how these results might be extended to Adam. Ke Liang Xiao, Noah Marshall, Atish Agarwala, Elliot Paquette |
ICML | 4 |
| 2025 | Dimension-adapted Momentum Outscales SGDabstractWe investigate scaling laws for stochastic momentum algorithms on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying data-target complexities. While traditional stochastic gradient descent with momentum (SGD-M) yields identical scaling law exponents to SGD, dimension-adapted Nesterov acceleration (DANA) improves these exponents by scaling momentum hyperparameters based on model size and data complexity. This outscaling phenomenon, which also improves compute-optimal scaling behavior, is achieved by DANA across a broad range of data and target complexities, while traditional methods fall short. Extensive experiments on high-dimensional synthetic quadratics validate our theoretical predictions and large-scale text experiments with LSTMs show DANA's improved loss exponents over SGD hold in a practical setting. Damien Ferbach, Katie Everett, Gauthier Gidel, Elliot Paquette, Courtney Paquette |
NeurIPS | 4 |
| 2024 | The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate AlgorithmsabstractWe develop a framework for analyzing the training and learning rate dynamics on a large class of high-dimensional optimization problems, which we call the high line, trained using one-pass stochastic gradient descent (SGD) with adaptive learning rates. We give exact expressions for the risk and learning rate curves in terms of a deterministic solution to a system of ODEs. We then investigate in detail two adaptive learning rates -- an idealized exact line search and AdaGrad-Norm -- on the least squares problem. When the data covariance matrix has strictly positive eigenvalues, this idealized exact line search strategy can exhibit arbitrarily slower convergence when compared to the optimal fixed learning rate with SGD. Moreover we exactly characterize the limiting learning rate (as time goes to infinity) for line search in the setting where the data covariance has only two distinct eigenvalues. For noiseless targets, we further demonstrate that the AdaGrad-Norm learning rate converges to a deterministic constant inversely proportional to the average eigenvalue of the data covariance matrix, and identify a phase transition when the covariance density of eigenvalues follows a power law distribution. We provide
our code for evaluation at https://github.com/amackenzie1/highline2024. Elizabeth Collins-Woodfin, Inbar Seroussi, Begoña García Malaxechebarría, Andrew W. Mackenzie, Elliot Paquette, Courtney Paquette |
NeurIPS | 5 |
| 2024 | 4+3 Phases of Compute-Optimal Neural Scaling LawsabstractWe consider the solvable neural scaling model with three parameters: data complexity, target complexity, and model-parameter-count. We use this neural scaling model to derive new predictions about the compute-limited, infinite-data scaling law regime. To train the neural scaling model, we run one-pass stochastic gradient descent on a mean-squared loss. We derive a representation of the loss curves which holds over all iteration counts and improves in accuracy as the model parameter count grows. We then analyze the compute-optimal model-parameter-count, and identify 4 phases (+3 subphases) in the data-complexity/target-complexity phase-plane. The phase boundaries are determined by the relative importance of model capacity, optimizer noise, and embedding of the features. We furthermore derive, with mathematical proof and extensive numerical evidence, the scaling-law exponents in all of these phases, in particular computing the optimal model-parameter-count as a function of floating point operation budget. We include a colab notebook https://tinyurl.com/2saj6bkj, nanoChinchilla, that reproduces some key results of the paper. Elliot Paquette, Courtney Paquette, Lechao Xiao, Jeffrey Pennington |
NeurIPS | 1 |
| 2022 | Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High DimensionsabstractWe analyze the dynamics of large batch stochastic gradient descent with momentum (SGD+M) on the least squares problem when both the number of samples and dimensions are large. In this setting, we show that the dynamics of SGD+M converge to a deterministic discrete Volterra equation as dimension increases, which we analyze. We identify a stability measurement, the implicit conditioning ratio (ICR), which regulates the ability of SGD+M to accelerate the algorithm. When the batch size exceeds this ICR, SGD+M converges linearly at a rate of $\mathcal{O}(1/\sqrt{\kappa})$, matching optimal full-batch momentum (in particular performing as well as a full-batch but with a fraction of the size). For batch sizes smaller than the ICR, in contrast, SGD+M has rates that scale like a multiple of the single batch SGD rate. We give explicit choices for the learning rate and momentum parameter in terms of the Hessian spectra that achieve this performance. Kiwon Lee, Andrew N. Cheng, Elliot Paquette, Courtney Paquette |
NeurIPS | 3 |
| 2022 | Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High DimensionsabstractStochastic gradient descent (SGD) is a pillar of modern machine learning, serving as the go-to optimization algorithm for a diverse array of problems. While the empirical success of SGD is often attributed to its computational efficiency and favorable generalization behavior, neither effect is well understood and disentangling them remains an open problem. Even in the simple setting of convex quadratic problems, worst-case analyses give an asymptotic convergence rate for SGD that is no better than full-batch gradient descent (GD), and the purported implicit regularization effects of SGD lack a precise explanation. In this work, we study the dynamics of multi-pass SGD on high-dimensional convex quadratics and establish an asymptotic equivalence to a stochastic differential equation, which we call homogenized stochastic gradient descent (HSGD), whose solutions we characterize explicitly in terms of a Volterra integral equation. These results yield precise formulas for the learning and risk trajectories, which reveal a mechanism of implicit conditioning that explains the efficiency of SGD relative to GD. We also prove that the noise from SGD negatively impacts generalization performance, ruling out the possibility of any type of implicit regularization in this context. Finally, we show how to adapt the HSGD formalism to include streaming SGD, which allows us to produce an exact prediction for the excess risk of multi-pass SGD relative to that of streaming SGD (bootstrap risk). Courtney Paquette, Elliot Paquette, Ben Adlam, Jeffrey Pennington |
NeurIPS | 2 |
| 2021 | SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize CriticalityabstractWe propose a new framework, inspired by random matrix theory, for analyzing the dynamics of stochastic gradient descent (SGD) when both number of samples and dimensions are large. This framework applies to any fixed stepsize and on the least squares problem with random data (finite-sum). Using this new framework, we show that the dynamics of SGD become deterministic in the large sample and dimensional limit. Furthermore, the limiting dynamics are governed by a Volterra integral equation. This model predicts that SGD undergoes a phase transition at an explicitly given critical stepsize that ultimately affects its convergence rate, which is also verified experimentally. Finally, when input data is isotropic, we provide explicit expressions for the dynamics and average-case convergence rates (i.e., the complexity of an algorithm averaged over all possible inputs). These rates show significant improvement over classical worst-case complexities. Courtney Paquette, Kiwon Lee, Fabian Pedregosa, Elliot Paquette |
COLT | 4 |
| 2021 | Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic ModelsabstractWe analyze a class of stochastic gradient algorithms with momentum on a high-dimensional random least squares problem. Our framework, inspired by random matrix theory, provides an exact (deterministic) characterization for the sequence of function values produced by these algorithms which is expressed only in terms of the eigenvalues of the Hessian. This leads to simple expressions for nearly-optimal hyperparameters, a description of the limiting neighborhood, and average-case complexity. As a consequence, we show that (small-batch) stochastic heavy-ball momentum with a fixed momentum parameter provides no actual performance improvement over SGD when step sizes are adjusted correctly. For contrast, in the non-strongly convex setting, it is possible to get a large improvement over SGD using momentum. By introducing hyperparameters that depend on the number of samples, we propose a new algorithm sDANA (stochastic dimension adjusted Nesterov acceleration) which obtains an asymptotically optimal average-case complexity while remaining linearly convergent in the strongly convex setting without adjusting parameters. Courtney Paquette, Elliot Paquette |
NeurIPS | 2 |
| 2017 | The Threshold for Integer Homology in Random d-Complexes
Christopher Hoffman, Matthew Kahle, Elliot Paquette |
Discret. Comput. Geom. | 3 |