EDBT 2026 Demo / reviewers in the wild / expert
Radu-Alexandru Dragomir
dblp:254/2756
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-4600-7748ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Optimization for machine learning · 40% Learning theory · 35% Deep learning architectures and training · 13% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
gradient flow |
0.9 | 1 | 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › training dynamics
grokking |
0.9 | 1 | 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025 |
Machine learning › Learning theory
implicit bias |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Learning theory
margin maximization |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › large margin methods
maximum margin classifiers |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Learning theory › implicit bias
mirror flow |
0.8 | 1 | 2024 | Implicit Bias of Mirror Flow on Separable Data · NeurIPS 2024 |
Machine learning › Optimization for machine learning › gradient-based optimization › proximal gradient method
bregman proximal gradient |
0.5 | 1 | 2021 | Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021 |
Machine learning › Optimization for machine learning
stochastic gradient methods |
0.5 | 1 | 2021 | Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021 |
Machine learning › Optimization for machine learning
variance reduction |
0.5 | 1 | 2021 | Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021 |
Mathematical optimization › continuous optimization
convex optimization |
0.5 | 1 | 2021 | Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction · ICML 2021 |
Machine learning › Optimization for machine learning
implicit regularization |
0.3 | 1 | 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
variance reduction · 1.0bregman divergence · 1.0weight decay · 0.9riemannian gradient flow · 0.9gradient flow · 0.9mirror descent · 0.8continuous-time analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationabstractWe study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase behaviour as $\lambda \to 0$. During the initial fast phase, the trajectory follows the unregularised gradient flow and converges to a manifold of critical points of $F$. Then, at time of order $1/\lambda$, the trajectory enters a slow drift phase and follows a Riemannian gradient flow minimising the $\ell_2$-norm of the parameters. This purely optimisation-based phenomenon offers a natural explanation for the \textit{grokking} effect observed in deep learning, where the training loss rapidly reaches zero while the test loss plateaus for an extended period before suddenly improving. We argue that this generalisation jump can be attributed to the slow norm reduction induced by weight decay, as explained by our analysis. We validate this mechanism empirically on several synthetic regression tasks. Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir |
NeurIPS | 3 |
| 2024 | Implicit Bias of Mirror Flow on Separable DataabstractWe examine the continuous-time counterpart of mirror descent, namely mirror flow, on classification problems which are linearly separable. Such problems are minimised ‘at infinity’ and have many possible solutions; we study which solution is preferred by the algorithm depending on the mirror potential. For exponential tailed losses and under mild assumptions on the potential, we show that the iterates converge in direction towards a $\phi_\infty$-maximum margin classifier. The function $\phi_\infty$ is the horizon function of the mirror potential and characterises its shape ‘at infinity’. When the potential is separable, a simple formula allows to compute this function. We analyse several examples of potentials and provide numerical experiments highlighting our results. Scott Pesme, Radu-Alexandru Dragomir, Nicolas Flammarion |
NeurIPS | 2 |
| 2021 | Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance ReductionabstractWe study the problem of minimizing a relatively-smooth convex function using stochastic Bregman gradient methods. We first prove the convergence of Bregman Stochastic Gradient Descent (BSGD) to a region that depends on the noise (magnitude of the gradients) at the optimum. In particular, BSGD quickly converges to the exact minimizer when this noise is zero (interpolation setting, in which the data is fit perfectly). Otherwise, when the objective has a finite sum structure, we show that variance reduction can be used to counter the effect of noise. In particular, fast convergence to the exact minimizer can be obtained under additional regularity assumptions on the Bregman reference function. We illustrate the effectiveness of our approach on two key applications of relative smoothness: tomographic reconstruction with Poisson noise and statistical preconditioning for distributed optimization. Radu-Alexandru Dragomir, Mathieu Even, Hadrien Hendrikx |
ICML | 1 |