VLDB 2026 Research / reviewers in the wild / expert
Salma Tarmoun
dblp:292/8008
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Optimization for machine learning · 63% Deep learning architectures and training · 23% Learning theory · 13% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 73% Algorithms and data structures · 27% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
gradient flow |
1.0 | 2 | 2021 | Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021 On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks · ICML 2021 |
Machine learning › Deep learning architectures and training › training dynamics
edge of stability |
0.9 | 1 | 2025 | Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.9 | 1 | 2025 | Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025 |
Machine learning › Optimization for machine learning
convergence analysis |
0.5 | 1 | 2021 | On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks · ICML 2021 |
Machine learning › Learning theory › over-parameterization
overparameterized models |
0.5 | 1 | 2021 | Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021 |
Mathematical optimization
nonconvex optimization |
0.3 | 1 | 2025 | Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025 |
Algorithms and data structures › numerical linear algebra
matrix factorization |
0.1 | 1 | 2021 | Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021 |
Mathematical optimization › control theory
riccati equation |
0.1 | 1 | 2021 | Understanding the Dynamics of Gradient Flow in Overparameterized Linear models · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
riemannian manifold analysis · 1.7overparametrised least squares · 1.7conservation laws analysis · 1.0overparametrization analysis · 0.5min-norm solution · 0.5gradient flow · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix FactorizationabstractDespite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theoretically analyzing the learning dynamics of LoRA for matrix factorization (MF) under gradient flow (GF), emphasizing the crucial role of initialization. For small initialization, we theoretically show that GF converges to a neighborhood of the optimal solution, with smaller initialization leading to lower final error. Our analysis shows that the final error is affected by the misalignment between the singular spaces of the pre-trained model and the target matrix, and reducing the initialization scale improves alignment. To address this misalignment, we propose a spectral initialization for LoRA in MF and theoretically prove that GF with small spectral initialization converges to the fine-tuning task with arbitrary precision. Numerical experiments from MF and image classification validate our findings. Ziqing Xu, Hancheng Min, Lachlan E. MacDonald, Jinqi Luo, Salma Tarmoun, Enrique Mallada, René Vidal |
AISTATS | 5 |
| 2025 | Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least SquaresabstractClassical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the objective decreases non-monotonically with an observed implicit bias towards flat minima. In this paper, we take a step toward quantifying this phenomenon by providing convergence rates for gradient descent with large learning rates in an overparametrised least squares setting. The key insight behind our analysis is that, as a consequence of overparametrisation, the set of global minimisers forms a Riemannian manifold $M$, which enables the decomposition of the GD dynamics into components parallel and orthogonal to $M$. The parallel component corresponds to Riemannian gradient descent on the objective sharpness, while the orthogonal component corresponds to a quadratic dynamical system. This insight allows us to derive convergence rates in three regimes characterised by the learning rate size: the subcritical regime, in which transient instability is overcome in finite time before linear convergence to a suboptimally flat global minimum; the critical regime, in which instability persists for all time with a power-law convergence toward the optimally flat global minimum; the supercritical regime, in which instability persists for all time with linear convergence to an oscillation of period two centred on the optimally flat global minimum. Lachlan E. MacDonald, Hancheng Min, Leandro Palma, Salma Tarmoun, Ziqing Xu, René Vidal |
NeurIPS | 4 |
| 2023 | Linear Convergence of Gradient Descent For Finite Width Over-parametrized Linear Networks With General InitializationabstractRecent theoretical analyses of the convergence of gradient descent (GD) to a global minimum for over-parametrized neural networks make strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (spectral, balanced). In this work, we relax these assumptions and derive a linear convergence rate for two-layer linear networks trained using GD on the squared loss in the case of finite step size, finite width and general initialization. Despite the generality of our analysis, our rate estimates are significantly tighter than those of prior work. Moreover, we provide a time-varying step size rule that monotonically improves the convergence rate as the loss function decreases to zero. Numerical experiments validate our findings. Ziqing Xu, Hancheng Min, Salma Tarmoun, Enrique Mallada, René Vidal |
AISTATS | 3 |
| 2021 | On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksabstractNeural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon is to study how initialization and overparametrization affect convergence and implicit bias of training algorithms. In this paper, we present a novel analysis of single-hidden-layer linear networks trained under gradient flow, which connects initialization, optimization, and overparametrization. Firstly, we show that the squared loss converges exponentially to its optimum at a rate that depends on the level of imbalance of the initialization. Secondly, we show that proper initialization constrains the dynamics of the network parameters to lie within an invariant set. In turn, minimizing the loss over this set leads to the min-norm solution. Finally, we show that large hidden layer width, together with (properly scaled) random initialization, ensures proximity to such an invariant set during training, allowing us to derive a novel non-asymptotic upper-bound on the distance between the trained network and the min-norm solution. Hancheng Min, Salma Tarmoun, René Vidal, Enrique Mallada |
ICML | 2 |
| 2021 | Understanding the Dynamics of Gradient Flow in Overparameterized Linear modelsabstractWe provide a detailed analysis of the dynamics ofthe gradient flow in overparameterized two-layerlinear models. A particularly interesting featureof this model is that its nonlinear dynamics can beexactly solved as a consequence of a large num-ber of conservation laws that constrain the systemto follow particular trajectories. More precisely,the gradient flow preserves the difference of theGramian matrices of the input and output weights,and its convergence to equilibrium depends onboth the magnitude of that difference (which isfixed at initialization) and the spectrum of the data.In addition, and generalizing prior work, we proveour results without assuming small, balanced orspectral initialization for the weights. Moreover,we establish interesting mathematical connectionsbetween matrix factorization problems and differ-ential equations of the Riccati type. Salma Tarmoun, Guilherme França, Benjamin D. Haeffele, René Vidal |
ICML | 1 |