Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Aditya Varre

dblp:224/6338 · also Aditya Vardhan Varre · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Optimization for machine learning · 45% Deep learning architectures and training · 29% Learning theory · 13%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
implicit regularization
2.842024
SGD vs GD: Rank Deficiency in Linear Networks · NeurIPS 2024
Why Do We Need Weight Decay in Modern Deep Learning? · NeurIPS 2024
On the spectral bias of two-layer linear networks · NeurIPS 2023
Natural language and speech › Language models and text generation
in-context learning
0.912025
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025
Machine learning › Deep learning architectures and training
loss landscape
0.912025
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025
Machine learning › Optimization for machine learning › optimization landscape
stationary points
0.912025
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points · ICML 2025
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Why Do We Need Weight Decay in Modern Deep Learning? · NeurIPS 2024
Machine learning › Learning theory
generalization
0.712023
SGD with Large Step Sizes Learns Sparse Features · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
sparse feature learning
0.712023
SGD with Large Step Sizes Learns Sparse Features · ICML 2023
Machine learning › Learning theory › inductive bias
spectral bias
0.712023
On the spectral bias of two-layer linear networks · NeurIPS 2023
Machine learning › Optimization for machine learning
stochastic gradient descent
0.712023
SGD with Large Step Sizes Learns Sparse Features · ICML 2023
Machine learning › Optimization for machine learning › stochastic gradient descent
accelerated stochastic gradient descent
0.612022
Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022
Machine learning › Optimization for machine learning
stochastic optimization
0.612022
Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022
Mathematical optimization › continuous optimization
convex optimization
0.612022
Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022
Mathematical optimization
least squares
0.612022
Accelerated SGD for Non-Strongly-Convex Least Squares · COLT 2022
Machine learning › Learning theory
statistical learning theory
0.512021
Last iterate convergence of SGD for Least-Squares in the Interpolation regime · NeurIPS 2021
Machine learning › Optimization for machine learning › convergence guarantees
stochastic gradient descent convergence
0.512021
Last iterate convergence of SGD for Least-Squares in the Interpolation regime · NeurIPS 2021
Natural language and speech › Language models and text generation
large language model training
0.212024
Why Do We Need Weight Decay in Modern Deep Learning? · NeurIPS 2024
Machine learning › Optimization for machine learning › stochastic gradient descent
step size schedule
0.212023
SGD with Large Step Sizes Learns Sparse Features · ICML 2023

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 1.9gradient flow · 1.4accelerated gradient descent · 1.1gradient analysis · 0.9cross-entropy loss · 0.9weight decay · 0.8stochastic differential equation · 0.8variational characterization · 0.7mirror flow · 0.7loss stabilization analysis · 0.7lower bound analysis · 0.6
YearPublicationVenuePosition
2025 Learning In-context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points
abstract
Motivated by empirical observations of prolonged plateaus and stage-wise progression during training, we investigate the loss landscape of transformer models trained on in-context next-token prediction tasks. In particular, we focus on learning in-context $n$-gram language models under cross-entropy loss, and establish a sufficient condition for parameter configurations to be stationary points. We then construct a set of parameter configurations for a simplified transformer model that represent $k$-gram estimators (for $k \leq n$), and show that the gradient of the population loss at these solutions vanishes in the limit of infinite sequence length and parameter norm. This reveals a key property of the loss landscape: sub-$n$-grams are near-stationary points of the population cross-entropy loss, offering theoretical insight into widely observed phenomena such as stage-wise learning dynamics and emergent phase transitions. These insights are further supported by numerical experiments that illustrate the learning dynamics of $n$-grams, characterized by discrete transitions between near-stationary solutions.
Aditya Varre, Gizem Yüce, Nicolas Flammarion
ICML1
2024 Why Do We Need Weight Decay in Modern Deep Learning?
abstract
Weight decay is a broadly used technique for training state-of-the-art deep networks from image classification to large language models. Despite its widespread usage and being extensively studied in the classical literature, its role remains poorly understood for deep learning. In this work, we highlight that the role of weight decay in modern deep learning is different from its regularization effect studied in classical learning theory. For deep networks on vision tasks trained with multipass SGD, we show how weight decay modifies the optimization dynamics enhancing the ever-present implicit regularization of SGD via the *loss stabilization mechanism*. In contrast, for large language models trained with nearly one-epoch training, we describe how weight decay balances the *bias-variance tradeoff* in stochastic optimization leading to lower training loss and improved training stability. Overall, we present a unifying perspective from ResNets on vision tasks to LLMs: weight decay is never useful as an explicit regularizer but instead changes the training dynamics in a desirable way.
Francesco D'Angelo, Maksym Andriushchenko, Aditya Varre, Nicolas Flammarion
NeurIPS3
2024 SGD vs GD: Rank Deficiency in Linear Networks
abstract
In this article, we study the behaviour of continuous-time gradient methods on a two-layer linear network with square loss. A dichotomy between SGD and GD is revealed: GD preserves the rank at initialization while (label noise) SGD diminishes the rank regardless of the initialization. We demonstrate this rank deficiency by studying the time evolution of the *determinant* of a matrix of parameters. To further understand this phenomenon, we derive the stochastic differential equation (SDE) governing the eigenvalues of the parameter matrix. This SDE unveils a *replusive force* between the eigenvalues: a key regularization mechanism which induces rank deficiency. Our results are well supported by experiments illustrating the phenomenon beyond linear networks and regression tasks.
Aditya Varre, Margarita Sagitova, Nicolas Flammarion
NeurIPS1
2023 SGD with Large Step Sizes Learns Sparse Features
abstract
We showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks. We present empirical observations that commonly used large step sizes (i) may lead the iterates to jump from one side of a valley to the other causing *loss stabilization*, and (ii) this stabilization induces a hidden stochastic dynamics that *biases it implicitly* toward simple predictors. Furthermore, we show empirically that the longer large step sizes keep SGD high in the loss landscape valleys, the better the implicit regularization can operate and find sparse representations. Notably, no explicit regularization is used: the regularization effect comes solely from the SGD dynamics influenced by the large step sizes schedule. Therefore, these observations unveil how, through the step size schedules, both gradient and noise drive together the SGD dynamics through the loss landscape of neural networks. We justify these findings theoretically through the study of simple neural network models as well as qualitative arguments inspired from stochastic processes. This analysis allows us to shed new light on some common practices and observed phenomena when training deep networks.
Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, Nicolas Flammarion
ICML2
2023 On the spectral bias of two-layer linear networks
abstract
This paper studies the behaviour of two-layer fully connected networks with linear activations trained with gradient flow on the square loss. We show how the optimization process carries an implicit bias on the parameters that depends on the scale of its initialization. The main result of the paper is a variational characterization of the loss minimizers retrieved by the gradient flow for a specific initialization shape. This characterization reveals that, in the small scale initialization regime, the linear neural network's hidden layer is biased toward having a low-rank structure. To complement our results, we showcase a hidden mirror flow that tracks the dynamics of the singular values of the weights matrices and describe their time evolution. We support our findings with numerical experiments illustrating the phenomena.
Aditya Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, Nicolas Flammarion
NeurIPS1
2022 Accelerated SGD for Non-Strongly-Convex Least Squares
abstract
We consider stochastic approximation for the least squares regression problem in the non-strongly convex setting. We present the first practical algorithm that achieves the optimal prediction error rates in terms of dependence on the noise of the problem, as $O(d/t)$ while accelerating the forgetting of the initial conditions to $O(d/t^2)$. Our new algorithm is based on a simple modification of the accelerated gradient descent. We provide convergence results for both the averaged and the last iterate of the algorithm. In order to describe the tightness of these new bounds, we present a matching lower bound in the noiseless setting and thus show the optimality of our algorithm.
Aditya Varre, Nicolas Flammarion
COLT1
2021 Last iterate convergence of SGD for Least-Squares in the Interpolation regime
abstract
Motivated by the recent successes of neural networks that have the ability to fit the data perfectly \emph{and} generalize well, we study the noiseless model in the fundamental least-squares setup. We assume that an optimum predictor perfectly fits the inputs and outputs $\langle \theta_* , \phi(X) \rangle = Y$, where $\phi(X)$ stands for a possibly infinite dimensional non-linear feature map. To solve this problem, we consider the estimator given by the last iterate of stochastic gradient descent (SGD) with constant step-size. In this context, our contribution is two fold: (i) \emph{from a (stochastic) optimization perspective}, we exhibit an archetypal problem where we can show explicitly the convergence of SGD final iterate for a non-strongly convex problem with constant step-size whereas usual results use some form of average and (ii) \emph{from a statistical perspective}, we give explicit non-asymptotic convergence rates in the over-parameterized setting and leverage a \emph{fine-grained} parameterization of the problem to exhibit polynomial rates that can be faster than $O(1/T)$. The link with reproducing kernel Hilbert spaces is established.
Aditya Varre, Loucas Pillaud-Vivien, Nicolas Flammarion
NeurIPS1
2019 Variants of Homomorphism Polynomials Complete for Algebraic Complexity Classes
Prasad Chaugule, Nutan Limaye, Aditya Varre
COCOON3