VLDB 2026 Research / reviewers in the wild / expert
Scott Pesme
dblp:268/7836
· DBLP profile ↗
9ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Learning theory · 36% Optimization for machine learning · 34% Deep learning architectures and training · 17% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
implicit bias |
1.9 | 3 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 Saddle-to-Saddle Dynamics in Diagonal Linear Networks · NeurIPS 2023 Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity · NeurIPS 2021 |
Machine learning › Optimization for machine learning
gradient flow |
1.5 | 2 | 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025 Saddle-to-Saddle Dynamics in Diagonal Linear Networks · NeurIPS 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
1.5 | 3 | 2023 | (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of Stability · NeurIPS 2023 Online Robust Regression via SGD on the l1 loss · NeurIPS 2020 On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent · ICML 2020 |
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks |
1.2 | 2 | 2023 | Saddle-to-Saddle Dynamics in Diagonal Linear Networks · NeurIPS 2023 Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity · NeurIPS 2021 |
Machine learning › Optimization for machine learning
implicit regularization |
0.9 | 2 | 2025 | (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of Stability · NeurIPS 2023 A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › training dynamics
grokking |
0.9 | 1 | 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference |
0.9 | 1 | 2025 | MAP Estimation with Denoisers: Convergence Rates and Guarantees · NeurIPS 2025 |
Machine learning › Learning theory
margin maximization |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › large margin methods
maximum margin classifiers |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Learning theory › implicit bias
mirror flow |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › training dynamics
edge of stability |
0.7 | 1 | 2023 | (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of Stability · NeurIPS 2023 |
Machine learning › Learning theory › implicit bias
implicit bias of gradient descent |
0.7 | 1 | 2023 | (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of Stability · NeurIPS 2023 |
Machine learning › Learning theory
sparse recovery |
0.7 | 1 | 2023 | Saddle-to-Saddle Dynamics in Diagonal Linear Networks · NeurIPS 2023 |
Machine learning › Learning theory
generalization |
0.5 | 1 | 2021 | Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity · NeurIPS 2021 |
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate |
0.4 | 1 | 2020 | On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent · ICML 2020 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
convergence diagnostics |
0.4 | 1 | 2020 | On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent · ICML 2020 |
Machine learning › Learning theory › statistical estimation › robust statistics
robust regression |
0.4 | 1 | 2020 | Online Robust Regression via SGD on the l1 loss · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
continuous-time analysis · 1.3convergence analysis · 1.1weight decay · 0.9riemannian gradient flow · 0.9log-concavity analysis · 0.9gradient flow · 0.9gradient descent on smoothed proximal objectives · 0.9mirror descent · 0.8overparametrized regression · 0.7arc-length time-reparametrization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationabstractWe study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase behaviour as $\lambda \to 0$. During the initial fast phase, the trajectory follows the unregularised gradient flow and converges to a manifold of critical points of $F$. Then, at time of order $1/\lambda$, the trajectory enters a slow drift phase and follows a Riemannian gradient flow minimising the $\ell_2$-norm of the parameters. This purely optimisation-based phenomenon offers a natural explanation for the \textit{grokking} effect observed in deep learning, where the training loss rapidly reaches zero while the test loss plateaus for an extended period before suddenly improving. We argue that this generalisation jump can be attributed to the slow norm reduction induced by weight decay, as explained by our analysis. We validate this mechanism empirically on several synthetic regression tasks. Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir |
NeurIPS | 2 |
| 2025 | MAP Estimation with Denoisers: Convergence Rates and GuaranteesabstractDenoiser models have become powerful tools for inverse problems, enabling the use of pretrained networks to approximate the score of a smoothed prior distribution. These models are often used in heuristic iterative schemes aimed at solving Maximum a Posteriori (MAP) optimisation problems, where the proximal operator of the negative log-prior plays a central role. In practice, this operator is intractable, and practitioners plug in a pretrained denoiser as a surrogate—despite the lack of general theoretical justification for this substitution. In this work, we show that a simple algorithm, closely related to several used in practice, provably converges to the proximal operator under a log-concavity assumption on the prior $p$. We show that this algorithm can be interpreted as a gradient descent on smoothed proximal objectives. Our analysis thus provides a theoretical foundation for a class of empirically successful but previously heuristic methods Scott Pesme, Giacomo Meanti, Michael Arbel, Julien Mairal |
NeurIPS | 1 |
| 2024 | Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear NetworksabstractIn this work, we investigate the effect of momentum on the optimisation trajectory of gradient descent. We leverage a continuous-time approach in the analysis of momentum gradient descent with step size $\gamma$ and momentum parameter $\beta$ that allows us to identify an intrinsic quantity $\lambda = \frac{ \gamma }{ (1 - \beta)^2 }$ which uniquely defines the optimisation path and provides a simple acceleration rule. When training a $2$-layer diagonal linear network in an overparametrised regression setting, we characterise the recovered solution through an implicit regularisation problem. We then prove that small values of $\lambda$ help to recover sparse solutions. Finally, we give similar but weaker results for stochastic momentum gradient descent. We provide numerical experiments which support our claims. Hristo Papazov, Scott Pesme, Nicolas Flammarion |
AISTATS | 2 |
| 2024 | Implicit Bias of Mirror Flow on Separable DataabstractWe examine the continuous-time counterpart of mirror descent, namely mirror flow, on classification problems which are linearly separable. Such problems are minimised ‘at infinity’ and have many possible solutions; we study which solution is preferred by the algorithm depending on the mirror potential. For exponential tailed losses and under mild assumptions on the potential, we show that the iterates converge in direction towards a $\phi_\infty$-maximum margin classifier. The function $\phi_\infty$ is the horizon function of the mirror potential and characterises its shape ‘at infinity’. When the potential is separable, a simple formula allows to compute this function. We analyse several examples of potentials and provide numerical experiments highlighting our results. Scott Pesme, Radu-Alexandru Dragomir, Nicolas Flammarion |
NeurIPS | 1 |
| 2023 | (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityabstractIn this paper, we investigate the impact of stochasticity and large stepsizes on the implicit regularisation of gradient descent (GD) and stochastic gradient descent (SGD) over $2$-layer diagonal linear networks. We prove the convergence of GD and SGD with macroscopic stepsizes in an overparametrised regression setting and characterise their solutions through an implicit regularisation problem. Our crisp characterisation leads to qualitative insights about the impact of stochasticity and stepsizes on the recovered solution. Specifically, we show that large stepsizes consistently benefit SGD for sparse regression problems, while they can hinder the recovery of sparse solutions for GD. These effects are magnified for stepsizes in a tight window just below the divergence threshold, in the ``edge of stability'' regime. Our findings are supported by experimental results. Mathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas Flammarion |
NeurIPS | 2 |
| 2023 | Saddle-to-Saddle Dynamics in Diagonal Linear NetworksabstractIn this paper we fully describe the trajectory of gradient flow over $2$-layer diagonal linear networks for the regression setting in the limit of vanishing initialisation. We show that the limiting flow successively jumps from a saddle of the training loss to another until reaching the minimum $\ell_1$-norm solution. We explicitly characterise the visited saddles as well as the jump times through a recursive algorithm reminiscent of the LARS algorithm used for computing the Lasso path. Starting from the zero vector, coordinates are successively activated until the minimum $\ell_1$-norm solution is recovered, revealing an incremental learning. Our proof leverages a convenient arc-length time-reparametrisation which enables to keep track of the transitions between the jumps. Our analysis requires negligible assumptions on the data, applies to both under and overparametrised settings and covers complex cases where there is no monotonicity of the number of active coordinates. We provide numerical experiments to support our findings. Scott Pesme, Nicolas Flammarion |
NeurIPS | 1 |
| 2021 | Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityabstractUnderstanding the implicit bias of training algorithms is of crucial importance in order to explain the success of overparametrised neural networks. In this paper, we study the dynamics of stochastic gradient descent over diagonal linear networks through its continuous time version, namely stochastic gradient flow. We explicitly characterise the solution chosen by the stochastic flow and prove that it always enjoys better generalisation properties than that of gradient flow.Quite surprisingly, we show that the convergence speed of the training loss controls the magnitude of the biasing effect: the slower the convergence, the better the bias. To fully complete our analysis, we provide convergence guarantees for the dynamics. We also give experimental results which support our theoretical claims. Our findings highlight the fact that structured noise can induce better generalisation and they help explain the greater performances of stochastic gradient descent over gradient descent observed in practice. Scott Pesme, Loucas Pillaud-Vivien, Nicolas Flammarion |
NeurIPS | 1 |
| 2020 | On Convergence-Diagnostic based Step Sizes for Stochastic Gradient DescentabstractConstant step-size Stochastic Gradient Descent exhibits two phases: a transient phase during which iterates make fast progress towards the optimum, followed by a stationary phase during which iterates oscillate around the optimal point. In this paper, we show that efficiently detecting this transition and appropriately decreasing the step size can lead to fast convergence rates. We analyse the classical statistical test proposed by Pflug (1983), based on the inner product between consecutive stochastic gradients. Even in the simple case where the objective function is quadratic we show that this test cannot lead to an adequate convergence diagnostic. We then propose a novel and simple statistical procedure that accurately detects stationarity and we provide experimental results showing state-of-the-art performance on synthetic and real-word datasets. Scott Pesme, Aymeric Dieuleveut, Nicolas Flammarion |
ICML | 1 |
| 2020 | Online Robust Regression via SGD on the l1 lossabstractWe consider the robust linear regression problem in the online setting where we have access to the data in a streaming manner, one data point after the other. More specifically, for a true parameter $ \theta^* $, we consider the corrupted Gaussian linear model $y = + \varepsilon + b$ where the adversarial noise $b$ can take any value with probability $\eta$ and equals zero otherwise. We consider this adversary to be oblivious (i.e., $b$ independent of the data) since this is the only contamination model under which consistency is possible. Current algorithms rely on having the whole data at hand in order to identify and remove the outliers. In contrast, we show in this work that stochastic gradient descent on the l1 loss converges to the true parameter vector at a $\tilde{O}( 1 / (1 - \eta)^2 n )$ rate which is independent of the values of the contaminated measurements. Our proof relies on the elegant smoothing of the l1 loss by the Gaussian data and a classical non-asymptotic analysis of Polyak-Ruppert averaged SGD. In addition, we provide experimental evidence of the efficiency of this simple and highly scalable algorithm. Scott Pesme, Nicolas Flammarion |
NeurIPS | 1 |