Salma Tarmoun

dblp:292/8008 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Optimization for machine learning · 63% Deep learning architectures and training · 23% Learning theory · 13%
Theoretical computer science
2 papers
Mathematical optimization · 73% Algorithms and data structures · 27%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
gradient flow
1.022021
Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021
On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks · ICML 2021
Machine learning › Deep learning architectures and training › training dynamics
edge of stability
0.912025
Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.912025
Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025
Machine learning › Optimization for machine learning
convergence analysis
0.512021
On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks · ICML 2021
Machine learning › Learning theory › over-parameterization
overparameterized models
0.512021
Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021
Mathematical optimization
nonconvex optimization
0.312025
Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025
Algorithms and data structures › numerical linear algebra
matrix factorization
0.112021
Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021
Mathematical optimization › control theory
riccati equation
0.112021
Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021

Methods — techniques the papers use, named apart from their topics

riemannian manifold analysis · 1.7overparametrised least squares · 1.7conservation laws analysis · 1.0overparametrization analysis · 0.5min-norm solution · 0.5gradient flow · 0.5
YearPublicationVenuePosition
2025 Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization
abstract
Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theoretically analyzing the learning dynamics of LoRA for matrix factorization (MF) under gradient flow (GF), emphasizing the crucial role of initialization. For small initialization, we theoretically show that GF converges to a neighborhood of the optimal solution, with smaller initialization leading to lower final error. Our analysis shows that the final error is affected by the misalignment between the singular spaces of the pre-trained model and the target matrix, and reducing the initialization scale improves alignment. To address this misalignment, we propose a spectral initialization for LoRA in MF and theoretically prove that GF with small spectral initialization converges to the fine-tuning task with arbitrary precision. Numerical experiments from MF and image classification validate our findings.
Ziqing Xu, Hancheng Min, Lachlan E. MacDonald, Jinqi Luo, Salma Tarmoun, Enrique Mallada, René Vidal
AISTATS5
2025 Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares
abstract
Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the objective decreases non-monotonically with an observed implicit bias towards flat minima. In this paper, we take a step toward quantifying this phenomenon by providing convergence rates for gradient descent with large learning rates in an overparametrised least squares setting. The key insight behind our analysis is that, as a consequence of overparametrisation, the set of global minimisers forms a Riemannian manifold $M$, which enables the decomposition of the GD dynamics into components parallel and orthogonal to $M$. The parallel component corresponds to Riemannian gradient descent on the objective sharpness, while the orthogonal component corresponds to a quadratic dynamical system. This insight allows us to derive convergence rates in three regimes characterised by the learning rate size: the subcritical regime, in which transient instability is overcome in finite time before linear convergence to a suboptimally flat global minimum; the critical regime, in which instability persists for all time with a power-law convergence toward the optimally flat global minimum; the supercritical regime, in which instability persists for all time with linear convergence to an oscillation of period two centred on the optimally flat global minimum.
Lachlan E. MacDonald, Hancheng Min, Leandro Palma, Salma Tarmoun, Ziqing Xu, René Vidal
NeurIPS4
2023 Linear Convergence of Gradient Descent For Finite Width Over-parametrized Linear Networks With General Initialization
abstract
Recent theoretical analyses of the convergence of gradient descent (GD) to a global minimum for over-parametrized neural networks make strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (spectral, balanced). In this work, we relax these assumptions and derive a linear convergence rate for two-layer linear networks trained using GD on the squared loss in the case of finite step size, finite width and general initialization. Despite the generality of our analysis, our rate estimates are significantly tighter than those of prior work. Moreover, we provide a time-varying step size rule that monotonically improves the convergence rate as the loss function decreases to zero. Numerical experiments validate our findings.
Ziqing Xu, Hancheng Min, Salma Tarmoun, Enrique Mallada, René Vidal
AISTATS3
2021 On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks
abstract
Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon is to study how initialization and overparametrization affect convergence and implicit bias of training algorithms. In this paper, we present a novel analysis of single-hidden-layer linear networks trained under gradient flow, which connects initialization, optimization, and overparametrization. Firstly, we show that the squared loss converges exponentially to its optimum at a rate that depends on the level of imbalance of the initialization. Secondly, we show that proper initialization constrains the dynamics of the network parameters to lie within an invariant set. In turn, minimizing the loss over this set leads to the min-norm solution. Finally, we show that large hidden layer width, together with (properly scaled) random initialization, ensures proximity to such an invariant set during training, allowing us to derive a novel non-asymptotic upper-bound on the distance between the trained network and the min-norm solution.
Hancheng Min, Salma Tarmoun, René Vidal, Enrique Mallada
ICML2
2021 Understanding the Dynamics of Gradient Flow in Overparameterized Linear models
abstract
We provide a detailed analysis of the dynamics ofthe gradient flow in overparameterized two-layerlinear models. A particularly interesting featureof this model is that its nonlinear dynamics can beexactly solved as a consequence of a large num-ber of conservation laws that constrain the systemto follow particular trajectories. More precisely,the gradient flow preserves the difference of theGramian matrices of the input and output weights,and its convergence to equilibrium depends onboth the magnitude of that difference (which isfixed at initialization) and the spectrum of the data.In addition, and generalizing prior work, we proveour results without assuming small, balanced orspectral initialization for the weights. Moreover,we establish interesting mathematical connectionsbetween matrix factorization problems and differ-ential equations of the Riccati type.
Salma Tarmoun, Guilherme França, Benjamin D. Haeffele, René Vidal
ICML1