VLDB 2026 Research / reviewers in the wild / expert
Amit Attia
dblp:284/8167
· DBLP profile ↗
7ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Optimization for machine learning · 72% Learning theory · 17% Learning paradigms · 5% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
stochastic optimization |
2.3 | 3 | 2025 | Faster Stochastic Optimization with Arbitrary Delays via Adaptive Asynchronous Mini-Batching · ICML 2025 How Free is Parameter-Free Stochastic Optimization? · ICML 2024 SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine Variance · ICML 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
1.5 | 2 | 2025 | Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime · NeurIPS 2025 SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine Variance · ICML 2023 |
Machine learning › Optimization for machine learning
convergence analysis |
1.4 | 2 | 2024 | How Free is Parameter-Free Stochastic Optimization? · ICML 2024 SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine Variance · ICML 2023 |
Machine learning › Optimization for machine learning › stochastic optimization
asynchronous stochastic optimization |
0.9 | 1 | 2025 | Faster Stochastic Optimization with Arbitrary Delays via Adaptive Asynchronous Mini-Batching · ICML 2025 |
Machine learning › Learning paradigms
continual learning |
0.9 | 1 | 2025 | Optimal Rates in Continual Linear Regression via Increasing Regularization · NeurIPS 2025 |
Machine learning › Optimization for machine learning
delayed gradient |
0.9 | 1 | 2025 | Faster Stochastic Optimization with Arbitrary Delays via Adaptive Asynchronous Mini-Batching · ICML 2025 |
Machine learning › Optimization for machine learning
implicit regularization |
0.9 | 1 | 2025 | Optimal Rates in Continual Linear Regression via Increasing Regularization · NeurIPS 2025 |
Machine learning › Learning theory › over-parameterization › interpolation
interpolation regime |
0.9 | 1 | 2025 | Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime · NeurIPS 2025 |
Machine learning › Optimization for machine learning › convergence guarantees
last-iterate convergence |
0.9 | 1 | 2025 | Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
regularization |
0.9 | 1 | 2025 | Optimal Rates in Continual Linear Regression via Increasing Regularization · NeurIPS 2025 |
Machine learning › Optimization for machine learning › adaptive optimization
parameter-free optimization |
0.8 | 1 | 2024 | How Free is Parameter-Free Stochastic Optimization? · ICML 2024 |
Machine learning › Optimization for machine learning › stochastic gradient descent
adaptive step size |
0.7 | 1 | 2023 | SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine Variance · ICML 2023 |
Machine learning › Learning theory
empirical risk minimization |
0.6 | 1 | 2022 | Uniform Stability for First-Order Empirical Risk Minimization · COLT 2022 |
Machine learning › Optimization for machine learning
gradient-based optimization |
0.6 | 1 | 2022 | Uniform Stability for First-Order Empirical Risk Minimization · COLT 2022 |
Machine learning › Learning theory › generalization › stability and generalization
uniform stability |
0.6 | 1 | 2022 | Uniform Stability for First-Order Empirical Risk Minimization · COLT 2022 |
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
accelerated gradient descent |
0.5 | 1 | 2021 | Algorithmic Instabilities of Accelerated Gradient Descent · NeurIPS 2021 |
Machine learning › Learning theory › generalization bounds
algorithmic stability |
0.5 | 1 | 2021 | Algorithmic Instabilities of Accelerated Gradient Descent · NeurIPS 2021 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.5 | 1 | 2021 | Algorithmic Instabilities of Accelerated Gradient Descent · NeurIPS 2021 |
Machine learning › Learning theory
generalization bounds |
0.2 | 1 | 2022 | Uniform Stability for First-Order Empirical Risk Minimization · COLT 2022 |
Methods — techniques the papers use, named apart from their topics
stochastic gradient descent · 1.7stale gradient filtering · 0.9regularization · 0.9excess risk analysis · 0.9adaptive mini-batching · 0.9hyperparameter search · 0.8affine variance noise model · 0.7adagrad · 0.7mirror descent · 0.6black-box conversion · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Faster Stochastic Optimization with Arbitrary Delays via Adaptive Asynchronous Mini-BatchingabstractWe consider the problem of asynchronous stochastic optimization, where an optimization algorithm makes updates based on stale stochastic gradients of the objective that are subject to an arbitrary (possibly adversarial) sequence of delays. We present a procedure which, for any given $q \in (0,1]$, transforms any standard stochastic first-order method to an asynchronous method with convergence guarantee depending on the $q$-quantile delay of the sequence. This approach leads to convergence rates of the form $O(\tau_q/qT+\sigma/\sqrt{qT})$ for non-convex and $O(\tau_q^2/(q T)^2+\sigma/\sqrt{qT})$ for convex smooth problems, where $\tau_q$ is the $q$-quantile delay, generalizing and improving on existing results that depend on the average delay. We further show a method that automatically adapts to all quantiles simultaneously, without any prior knowledge of the delays, achieving convergence rates of the form $O(\inf_{q} \tau_q/qT+\sigma/\sqrt{qT})$ for non-convex and $O(\inf_{q} \tau_q^2/(q T)^2+\sigma/\sqrt{qT})$ for convex smooth problems. Our technique is based on asynchronous mini-batching with a careful batch-size selection and filtering of stale gradients. Amit Attia, Ofir Gaash, Tomer Koren |
ICML | 1 |
| 2025 | Fast Last-Iterate Convergence of SGD in the Smooth Interpolation RegimeabstractWe study population convergence guarantees of stochastic gradient descent (SGD) for smooth convex objectives in the interpolation regime, where the noise at optimum is zero or near zero. The behavior of the last iterate of SGD in this setting---particularly with large (constant) stepsizes---has received growing attention in recent years due to implications for the training of over-parameterized models, as well as to analyzing forgetting in continual learning and to understanding the convergence of the randomized Kaczmarz method for solving linear systems. We establish that after $T$ steps of SGD on $\beta$-smooth convex loss functions with stepsize $0 < \eta < 2/\beta$, the last iterate exhibits expected excess risk $\widetilde{O}(\tfrac{1}{\eta (2-\beta \eta) T^{1-\beta\eta/2}} + \tfrac{\eta}{(2-\beta\eta)^2} T^{\beta\eta/2} \sigma_\star^2)$, where $\sigma_\star^2$ denotes the variance of the stochastic gradients at the optimum. In particular, for a well-tuned stepsize we obtain a near optimal $\widetilde{O}(1/T + \sigma_\star/\sqrt T)$ rate for the last iterate, extending the results of Varre et al. (2021) beyond least squares regression; and when $\sigma_\star=0$ we obtain a rate of $O(1/\sqrt T)$ with $\eta=1/\beta$, improving upon the best-known $O(T^{-1/4})$ rate recently established by Evron et al. (2025) in the special case of realizable linear regression. Amit Attia, Matan Schliserman, Uri Sherman, Tomer Koren |
NeurIPS | 1 |
| 2025 | Optimal Rates in Continual Linear Regression via Increasing RegularizationabstractWe study realizable continual linear regression under random task orderings, a common setting for developing continual learning theory.
In this setup, the worst-case expected loss after $k$ learning iterations admits a lower bound of $\Omega(1/k)$.
However, prior work using an unregularized scheme has only established an upper bound of $O(1/k^{1/4})$, leaving a significant gap.
Our paper proves that this gap can be narrowed, or even closed, using two frequently used regularization schemes:
(1) explicit isotropic $\ell_2$ regularization, and (2) implicit regularization via finite step budgets.
We show that these approaches, which are used in practice to mitigate forgetting, reduce to stochastic gradient descent (SGD) on carefully defined surrogate losses.
Through this lens, we identify a fixed regularization strength that yields a near-optimal rate of $O(\log k / k)$.
Formalizing and analyzing a generalized variant of SGD for time-varying functions, we derive an increasing regularization strength schedule that provably achieves an optimal rate of $O(1/k)$.
This suggests that schedules that increase the regularization coefficient or decrease the number of steps per task are beneficial, at least in the worst case. Ran Levinstein, Amit Attia, Matan Schliserman, Uri Sherman, Daniel Soudry, Tomer Koren, Itay Evron |
NeurIPS | 2 |
| 2024 | How Free is Parameter-Free Stochastic Optimization?abstractWe study the problem of parameter-free stochastic optimization, inquiring whether, and under what conditions, do fully parameter-free methods exist: these are methods that achieve convergence rates competitive with optimally tuned methods, without requiring significant knowledge of the true problem parameters. Existing parameter-free methods can only be considered ``partially'' parameter-free, as they require some non-trivial knowledge of the true problem parameters, such as a bound on the stochastic gradient norms, a bound on the distance to a minimizer, etc. In the non-convex setting, we demonstrate that a simple hyperparameter search technique results in a fully parameter-free method that outperforms more sophisticated state-of-the-art algorithms. We also provide a similar result in the convex setting with access to noisy function values under mild noise assumptions. Finally, assuming only access to stochastic gradients, we establish a lower bound that renders fully parameter-free stochastic convex optimization infeasible, and provide a method which is (partially) parameter-free up to the limit indicated by our lower bound. Amit Attia, Tomer Koren |
ICML | 1 |
| 2023 | SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine VarianceabstractWe study Stochastic Gradient Descent with AdaGrad stepsizes: a popular adaptive (self-tuning) method for first-order stochastic optimization. Despite being well studied, existing analyses of this method suffer from various shortcomings: they either assume some knowledge of the problem parameters, impose strong global Lipschitz conditions, or fail to give bounds that hold with high probability. We provide a comprehensive analysis of this basic method without any of these limitations, in both the convex and non-convex (smooth) cases, that additionally supports a general ``affine variance'' noise model and provides sharp rates of convergence in both the low-noise and high-noise regimes. Amit Attia, Tomer Koren |
ICML | 1 |
| 2022 | Uniform Stability for First-Order Empirical Risk MinimizationabstractWe consider the problem of designing uniformly stable first-order optimization algorithms for empirical risk minimization. Uniform stability is often used to obtain generalization error bounds for optimization algorithms, and we are interested in a general approach to achieve it. For Euclidean geometry, we suggest a black-box conversion which given a smooth optimization algorithm, produces a uniformly stable version of the algorithm while maintaining its convergence rate up to logarithmic factors. Using this reduction we obtain a (nearly) optimal algorithm for smooth optimization with convergence rate $\tilde{O}(1/T^2)$ and uniform stability $O(T^2/n)$, resolving an open problem of Chen et al. (2018); Attia and Koren (2021). For more general geometries, we develop a variant of Mirror Descent for smooth optimization with convergence rate $\tilde{O}(1/T)$ and uniform stability $O(T/n)$, leaving open the question of devising a general conversion method as in the Euclidean case. Amit Attia, Tomer Koren |
COLT | 1 |
| 2021 | Algorithmic Instabilities of Accelerated Gradient DescentabstractWe study the algorithmic stability of Nesterov's accelerated gradient method. For convex quadratic objectives, Chen et al. (2018) proved that the uniform stability of the method grows quadratically with the number of optimization steps, and conjectured that the same is true for the general convex and smooth case. We disprove this conjecture and show, for two notions of algorithmic stability (including uniform stability), that the stability of Nesterov's accelerated method in fact deteriorates exponentially fast with the number of gradient steps. This stands in sharp contrast to the bounds in the quadratic case, but also to known results for non-accelerated gradient methods where stability typically grows linearly with the number of steps. Amit Attia, Tomer Koren |
NeurIPS | 1 |