EDBT 2026 Demo / reviewers in the wild / expert
Aditya Varre
dblp:224/6338 · also Aditya Vardhan Varre
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Optimization for machine learning · 45% Deep learning architectures and training · 29% Learning theory · 13% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
implicit regularization |
2.8 | 4 | 2024 | SGD vs GD: Rank Deficiency in Linear Networks · NeurIPS 2024 Why Do We Need Weight Decay in Modern Deep Learning? · NeurIPS 2024 On the spectral bias of two-layer linear networks · NeurIPS 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025 |
Machine learning › Deep learning architectures and training
loss landscape |
0.9 | 1 | 2025 | Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025 |
Machine learning › Optimization for machine learning › optimization landscape
stationary points |
0.9 | 1 | 2025 | Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025 |
Machine learning › Deep learning architectures and training
training dynamics |
0.8 | 1 | 2024 | Why Do We Need Weight Decay in Modern Deep Learning? · NeurIPS 2024 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | SGD with Large Step Sizes Learns Sparse Features · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
sparse feature learning |
0.7 | 1 | 2023 | SGD with Large Step Sizes Learns Sparse Features · ICML 2023 |
Machine learning › Learning theory › inductive bias
spectral bias |
0.7 | 1 | 2023 | On the spectral bias of two-layer linear networks · NeurIPS 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.7 | 1 | 2023 | SGD with Large Step Sizes Learns Sparse Features · ICML 2023 |
Machine learning › Optimization for machine learning › stochastic gradient descent
accelerated stochastic gradient descent |
0.6 | 1 | 2022 | Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.6 | 1 | 2022 | Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022 |
Mathematical optimization › continuous optimization
convex optimization |
0.6 | 1 | 2022 | Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022 |
Mathematical optimization
least squares |
0.6 | 1 | 2022 | Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022 |
Machine learning › Learning theory
statistical learning theory |
0.5 | 1 | 2021 | Last iterate convergence of SGD for Least-Squares in the Interpolation regime · NeurIPS 2021 |
Machine learning › Optimization for machine learning › convergence guarantees
stochastic gradient descent convergence |
0.5 | 1 | 2021 | Last iterate convergence of SGD for Least-Squares in the Interpolation regime · NeurIPS 2021 |
Natural language and speech › Language models and text generation
large language model training |
0.2 | 1 | 2024 | Why Do We Need Weight Decay in Modern Deep Learning? · NeurIPS 2024 |
Machine learning › Optimization for machine learning › stochastic gradient descent
step size schedule |
0.2 | 1 | 2023 | SGD with Large Step Sizes Learns Sparse Features · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
stochastic gradient descent · 1.9gradient flow · 1.4accelerated gradient descent · 1.1gradient analysis · 0.9cross-entropy loss · 0.9weight decay · 0.8stochastic differential equation · 0.8variational characterization · 0.7mirror flow · 0.7loss stabilization analysis · 0.7lower bound analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary PointsabstractMotivated by empirical observations of prolonged plateaus and stage-wise progression during training, we investigate the loss landscape of transformer models trained on in-context next-token prediction tasks. In particular, we focus on learning in-context $n$-gram language models under cross-entropy loss, and establish a sufficient condition for parameter configurations to be stationary points. We then construct a set of parameter configurations for a simplified transformer model that represent $k$-gram estimators (for $k \leq n$), and show that the gradient of the population loss at these solutions vanishes in the limit of infinite sequence length and parameter norm. This reveals a key property of the loss landscape: sub-$n$-grams are near-stationary points of the population cross-entropy loss, offering theoretical insight into widely observed phenomena such as stage-wise learning dynamics and emergent phase transitions. These insights are further supported by numerical experiments that illustrate the learning dynamics of $n$-grams, characterized by discrete transitions between near-stationary solutions. Aditya Varre, Gizem Yüce, Nicolas Flammarion |
ICML | 1 |
| 2024 | Why Do We Need Weight Decay in Modern Deep Learning?abstractWeight decay is a broadly used technique for training state-of-the-art deep networks from image classification to large language models. Despite its widespread usage and being extensively studied in the classical literature, its role remains poorly understood for deep learning. In this work, we highlight that the role of weight decay in modern deep learning is different from its regularization effect studied in classical learning theory. For deep networks on vision tasks trained with multipass SGD, we show how weight decay modifies the optimization dynamics enhancing the ever-present implicit regularization of SGD via the *loss stabilization mechanism*. In contrast, for large language models trained with nearly one-epoch training, we describe how weight decay balances the *bias-variance tradeoff* in stochastic optimization leading to lower training loss and improved training stability.
Overall, we present a unifying perspective from ResNets on vision tasks to LLMs: weight decay is never useful as an explicit regularizer but instead changes the training dynamics in a desirable way. Francesco D'Angelo, Maksym Andriushchenko, Aditya Varre, Nicolas Flammarion |
NeurIPS | 3 |
| 2024 | SGD vs GD: Rank Deficiency in Linear NetworksabstractIn this article, we study the behaviour of continuous-time gradient methods on a two-layer linear network with square loss. A dichotomy between SGD and GD is revealed: GD preserves the rank at initialization while (label noise) SGD diminishes the rank regardless of the initialization. We demonstrate this rank deficiency by studying the time evolution of the *determinant* of a matrix of parameters. To further understand this phenomenon, we derive the stochastic differential equation (SDE) governing the eigenvalues of the parameter matrix. This SDE unveils a *replusive force* between the eigenvalues: a key regularization mechanism which induces rank deficiency. Our results are well supported by experiments illustrating the phenomenon beyond linear networks and regression tasks. Aditya Varre, Margarita Sagitova, Nicolas Flammarion |
NeurIPS | 1 |
| 2023 | SGD with Large Step Sizes Learns Sparse FeaturesabstractWe showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks. We present empirical observations that commonly used large step sizes (i) may lead the iterates to jump from one side of a valley to the other causing *loss stabilization*, and (ii) this stabilization induces a hidden stochastic dynamics that *biases it implicitly* toward simple predictors. Furthermore, we show empirically that the longer large step sizes keep SGD high in the loss landscape valleys, the better the implicit regularization can operate and find sparse representations. Notably, no explicit regularization is used: the regularization effect comes solely from the SGD dynamics influenced by the large step sizes schedule. Therefore, these observations unveil how, through the step size schedules, both gradient and noise drive together the SGD dynamics through the loss landscape of neural networks. We justify these findings theoretically through the study of simple neural network models as well as qualitative arguments inspired from stochastic processes. This analysis allows us to shed new light on some common practices and observed phenomena when training deep networks. Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, Nicolas Flammarion |
ICML | 2 |
| 2023 | On the spectral bias of two-layer linear networksabstractThis paper studies the behaviour of two-layer fully connected networks with linear activations trained with gradient flow on the square loss. We show how the optimization process carries an implicit bias on the parameters that depends on the scale of its initialization. The main result of the paper is a variational characterization of the loss minimizers retrieved by the gradient flow for a specific initialization shape. This characterization reveals that, in the small scale initialization regime, the linear neural network's hidden layer is biased toward having a low-rank structure. To complement our results, we showcase a hidden mirror flow that tracks the dynamics of the singular values of the weights matrices and describe their time evolution. We support our findings with numerical experiments illustrating the phenomena. Aditya Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, Nicolas Flammarion |
NeurIPS | 1 |
| 2022 | Accelerated SGD for Non-Strongly-Convex Least SquaresabstractWe consider stochastic approximation for the least squares regression problem in the non-strongly convex setting. We present the first practical algorithm that achieves the optimal prediction error rates in terms of dependence on the noise of the problem, as $O(d/t)$ while accelerating the forgetting of the initial conditions to $O(d/t^2)$. Our new algorithm is based on a simple modification of the accelerated gradient descent. We provide convergence results for both the averaged and the last iterate of the algorithm. In order to describe the tightness of these new bounds, we present a matching lower bound in the noiseless setting and thus show the optimality of our algorithm. Aditya Varre, Nicolas Flammarion |
COLT | 1 |
| 2021 | Last iterate convergence of SGD for Least-Squares in the Interpolation regimeabstractMotivated by the recent successes of neural networks that have the ability to fit the data perfectly \emph{and} generalize well, we study the noiseless model in the fundamental least-squares setup. We assume that an optimum predictor perfectly fits the inputs and outputs $\langle \theta_* , \phi(X) \rangle = Y$, where $\phi(X)$ stands for a possibly infinite dimensional non-linear feature map. To solve this problem, we consider the estimator given by the last iterate of stochastic gradient descent (SGD) with constant step-size. In this context, our contribution is two fold: (i) \emph{from a (stochastic) optimization perspective}, we exhibit an archetypal problem where we can show explicitly the convergence of SGD final iterate for a non-strongly convex problem with constant step-size whereas usual results use some form of average and (ii) \emph{from a statistical perspective}, we give explicit non-asymptotic convergence rates in the over-parameterized setting and leverage a \emph{fine-grained} parameterization of the problem to exhibit polynomial rates that can be faster than $O(1/T)$. The link with reproducing kernel Hilbert spaces is established. Aditya Varre, Loucas Pillaud-Vivien, Nicolas Flammarion |
NeurIPS | 1 |
| 2019 | Variants of Homomorphism Polynomials Complete for Algebraic Complexity Classes
Prasad Chaugule, Nutan Limaye, Aditya Varre |
COCOON | 3 |