Radu-Alexandru Dragomir

dblp:254/2756 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-4600-7748ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Optimization for machine learning · 40% Learning theory · 35% Deep learning architectures and training · 13%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
gradient flow
0.912025
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025
Machine learning › Deep learning architectures and training › training dynamics
grokking
0.912025
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025
Machine learning › Learning theory
implicit bias
0.812024
Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024
Machine learning › Learning theory
margin maximization
0.812024
Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods › large margin methods
maximum margin classifiers
0.812024
Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024
Machine learning › Learning theory › implicit bias
mirror flow
0.812024
Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024
Machine learning › Optimization for machine learning › gradient-based optimization › proximal gradient method
bregman proximal gradient
0.512021
Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021
Machine learning › Optimization for machine learning
stochastic gradient methods
0.512021
Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021
Machine learning › Optimization for machine learning
variance reduction
0.512021
Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021
Mathematical optimization › continuous optimization
convex optimization
0.512021
Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021
Machine learning › Optimization for machine learning
implicit regularization
0.312025
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

variance reduction · 1.0bregman divergence · 1.0weight decay · 0.9riemannian gradient flow · 0.9gradient flow · 0.9mirror descent · 0.8continuous-time analysis · 0.8
YearPublicationVenuePosition
2025 A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
abstract
We study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase behaviour as $\lambda \to 0$. During the initial fast phase, the trajectory follows the unregularised gradient flow and converges to a manifold of critical points of $F$. Then, at time of order $1/\lambda$, the trajectory enters a slow drift phase and follows a Riemannian gradient flow minimising the $\ell_2$-norm of the parameters. This purely optimisation-based phenomenon offers a natural explanation for the \textit{grokking} effect observed in deep learning, where the training loss rapidly reaches zero while the test loss plateaus for an extended period before suddenly improving. We argue that this generalisation jump can be attributed to the slow norm reduction induced by weight decay, as explained by our analysis. We validate this mechanism empirically on several synthetic regression tasks.
Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir
NeurIPS3
2024 Implicit Bias of Mirror Flow on Separable Data
abstract
We examine the continuous-time counterpart of mirror descent, namely mirror flow, on classification problems which are linearly separable. Such problems are minimised ‘at infinity’ and have many possible solutions; we study which solution is preferred by the algorithm depending on the mirror potential. For exponential tailed losses and under mild assumptions on the potential, we show that the iterates converge in direction towards a $\phi_\infty$-maximum margin classifier. The function $\phi_\infty$ is the horizon function of the mirror potential and characterises its shape ‘at infinity’. When the potential is separable, a simple formula allows to compute this function. We analyse several examples of potentials and provide numerical experiments highlighting our results.
Scott Pesme, Radu-Alexandru Dragomir, Nicolas Flammarion
NeurIPS2
2021 Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction
abstract
We study the problem of minimizing a relatively-smooth convex function using stochastic Bregman gradient methods. We first prove the convergence of Bregman Stochastic Gradient Descent (BSGD) to a region that depends on the noise (magnitude of the gradients) at the optimum. In particular, BSGD quickly converges to the exact minimizer when this noise is zero (interpolation setting, in which the data is fit perfectly). Otherwise, when the objective has a finite sum structure, we show that variance reduction can be used to counter the effect of noise. In particular, fast convergence to the exact minimizer can be obtained under additional regularity assumptions on the Bregman reference function. We illustrate the effectiveness of our approach on two key applications of relative smoothness: tomographic reconstruction with Poisson noise and statistical preconditioning for distributed optimization.
Radu-Alexandru Dragomir, Mathieu Even, Hadrien Hendrikx
ICML1