Abdurakhmon Sadiev

dblp:264/9455 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Optimization for machine learning · 79% Efficient and distributed learning · 16% Learning theory · 5%
Theoretical computer science
1 paper
Mathematical optimization · 50% Computational complexity · 50%

Topics — the 22 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
stochastic optimization
3.242025
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits · NeurIPS 2025
Error Feedback under (L0, L1)-Smoothness: Normalization and Momentum · NeurIPS 2025
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise · ICML 2024
Machine learning › Optimization for machine learning
distributed optimization
3.042025
Error Feedback under (L0, L1)-Smoothness: Normalization and Momentum · NeurIPS 2025
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise · ICML 2024
Machine learning › Optimization for machine learning › stochastic optimization
high-probability convergence
1.422024
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise · ICML 2024
High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance · ICML 2023
Machine learning › Optimization for machine learning
variational inequality
1.422024
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise · ICML 2024
High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance · ICML 2023
Machine learning › Optimization for machine learning › stochastic optimization
heavy-tailed noise
1.122025
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits · NeurIPS 2025
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise · ICML 2024
Machine learning › Efficient and distributed learning
federated learning
0.922024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact Prox · NeurIPS 2022
Machine learning › Optimization for machine learning › distributed optimization
error feedback
0.912025
Error Feedback under (L0, L1)-Smoothness: Normalization and Momentum · NeurIPS 2025
Machine learning › Learning theory › sample complexity
sample complexity lower bounds
0.912025
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits · NeurIPS 2025
Machine learning › Optimization for machine learning
second-order optimization
0.912025
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.812024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Optimization for machine learning › stochastic gradient descent
random reshuffling
0.812024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Efficient and distributed learning › federated learning
communication-efficient federated learning
0.612022
Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact Prox · NeurIPS 2022
Machine learning › Optimization for machine learning › distributed optimization
local gradient method
0.612022
Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact Prox · NeurIPS 2022
Machine learning › Optimization for machine learning
primal-dual methods
0.612022
Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact Prox · NeurIPS 2022
Computational complexity › complexity classes › approximation classes › optimization complexity
lower complexity bounds
0.612022
Optimal Algorithms for Decentralized Stochastic Variational Inequalities · NeurIPS 2022
Mathematical optimization › continuous optimization › convex optimization
variational inequality
0.612022
Optimal Algorithms for Decentralized Stochastic Variational Inequalities · NeurIPS 2022
Machine learning › Optimization for machine learning
gradient clipping
0.522025
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits · NeurIPS 2025
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise · ICML 2024
Machine learning › Efficient and distributed learning
distributed training
0.312025
Error Feedback under (L0, L1)-Smoothness: Normalization and Momentum · NeurIPS 2025
Machine learning › Optimization for machine learning
robust optimization
0.312025
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.212024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.212024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Optimization for machine learning › convergence analysis
convergence bounds
0.212023
High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance · ICML 2023

Methods — techniques the papers use, named apart from their topics

sample complexity analysis · 0.9normalized stochastic gradient descent · 0.9normalization · 0.9momentum · 0.9hessian clipping · 0.9generalized smoothness · 0.9error feedback · 0.9stochastic gradient difference clipping · 0.8gradient clipping · 0.8control iterates · 0.8stochastic approximation · 0.6optimal algorithm design · 0.6
YearPublicationVenuePosition
2025 Error Feedback under (L0, L1)-Smoothness: Normalization and Momentum
abstract
We provide the first proof of convergence for normalized error feedback algorithms across a wide range of machine learning problems. Despite their popularity and efficiency in training deep neural networks, traditional analyses of error feedback algorithms rely on the smoothness assumption that does not capture the properties of objective functions in these problems. Rather, these problems have recently been shown to satisfy generalized smoothness assumptions, and the theoretical understanding of error feedback algorithms under these assumptions remains largely unexplored. Moreover, to the best of our knowledge, all existing analyses under generalized smoothness either i) focus on single-node settings or ii) make unrealistically strong assumptions for distributed settings, such as requiring data heterogeneity, and almost surely bounded stochastic gradient noise variance. In this paper, we propose distributed error feedback algorithms that utilize normalization to achieve the $\mathcal{O}(1/\sqrt{K})$ convergence rate for nonconvex problems under generalized smoothness. Our analyses apply for distributed settings without data heterogeneity conditions, and enable stepsize tuning that is independent of problem parameters. Additionally, we provide strong convergence guarantees of normalized error feedback algorithms for stochastic settings. Finally, we show that due to their larger allowable stepsizes, our new normalized error feedback algorithms outperform their non-normalized counterparts on various tasks, including the minimization of polynomial functions, logistic regression, and ResNet-20 training.
Sarit Khirirat, Abdurakhmon Sadiev, Artem Riabinin, Eduard Gorbunov, Peter Richtárik
NeurIPS2
2025 Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits
abstract
Heavy-tailed noise is pervasive in modern machine learning applications, arising from data heterogeneity, outliers, and non-stationary stochastic environments. While second-order methods can significantly accelerate convergence in light-tailed or bounded-noise settings, such algorithms are often brittle and lack guarantees under heavy-tailed noise—precisely the regimes where robustness is most critical. In this work, we take a first step toward a theoretical understanding of second-order optimization under heavy-tailed noise. We consider a setting where stochastic gradients and Hessians have only bounded $p$-th moments, for some $p\in (1,2]$, and establish tight lower bounds on the sample complexity of any second-order method. We then develop a variant of normalized stochastic gradient descent that leverages second-order information and provably matches these lower bounds. To address the instability caused by large deviations, we introduce a novel algorithm based on gradient and Hessian clipping, and prove high-probability upper bounds that nearly match the fundamental limits. Our results provide the first comprehensive sample complexity characterization for second-order optimization under heavy-tailed noise. This positions Hessian clipping as a robust and theoretically sound strategy for second-order algorithm design in heavy-tailed regimes.
Abdurakhmon Sadiev, Peter Richtárik, Ilyas Fatkhullin
NeurIPS1
2024 High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
abstract
High-probability analysis of stochastic first-order optimization methods under mild assumptions on the noise has been gaining a lot of attention in recent years. Typically, gradient clipping is one of the key algorithmic ingredients to derive good high-probability guarantees when the noise is heavy-tailed. However, if implemented naively, clipping can spoil the convergence of the popular methods for composite and distributed optimization (Prox-SGD/Parallel SGD) even in the absence of any noise. Due to this reason, many works on high-probability analysis consider only unconstrained non-distributed problems, and the existing results for composite/distributed problems do not include some important special cases (like strongly convex problems) and are not optimal. To address this issue, we propose new stochastic methods for composite and distributed optimization based on the clipping of stochastic gradient differences and prove tight high-probability convergence results (including nearly optimal ones) for the new methods. In addition, we also develop new methods for composite and distributed variational inequalities and analyze the high-probability convergence of these methods.
Eduard Gorbunov, Abdurakhmon Sadiev, Marina Danilova, Samuel Horváth, Gauthier Gidel, Pavel E. Dvurechensky, Alexander V. Gasnikov, Peter Richtárik
ICML2
2024 Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences
abstract
Gradient compression is a popular technique for improving communication complexity of stochastic first-order methods in distributed training of machine learning models. However, the existing works consider only with-replacement sampling of stochastic gradients. In contrast, it is well-known in practice and recently confirmed in theory that stochastic methods based on without-replacement sampling, e.g., Random Reshuffling (RR) method, perform better than ones that sample the gradients with-replacement. In this work, we close this gap in the literature and provide the first analysis of methods with gradient compression and without-replacement sampling. We first develop a distributed variant of random reshuffling with gradient compression (Q-RR), and show how to reduce the variance coming from gradient quantization through the use of control iterates. Next, to have a better fit to Federated Learning applications, we incorporate local computation and propose a variant of Q-RR called Q-NASTYA. Q-NASTYA uses local gradient steps and different local and global stepsizes. Next, we show how to reduce compression variance in this setting as well. Finally, we prove the convergence results for the proposed methods and outline several settings in which they improve upon existing algorithms.
Abdurakhmon Sadiev, Grigory Malinovsky, Eduard Gorbunov, Igor Sokolov 0001, Ahmed Khaled 0001, Konstantin Burlachenko, Peter Richtárik
NeurIPS1
2023 High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance
abstract
During the recent years the interest of optimization and machine learning communities in high-probability convergence of stochastic optimization methods has been growing. One of the main reasons for this is that high-probability complexity bounds are more accurate and less studied than in-expectation ones. However, SOTA high-probability non-asymptotic convergence results are derived under strong assumptions such as boundedness of the gradient noise variance or of the objective's gradient itself. In this paper, we propose several algorithms with high-probability convergence results under less restrictive assumptions. In particular, we derive new high-probability convergence results under the assumption that the gradient/operator noise has bounded central $\alpha$-th moment for $\alpha \in (1,2]$ in the following setups: (i) smooth non-convex / Polyak-Lojasiewicz / convex / strongly convex / quasi-strongly convex minimization problems, (ii) Lipschitz / star-cocoercive and monotone / quasi-strongly monotone variational inequalities. These results justify the usage of the considered methods for solving problems that do not fit standard functional classes studied in stochastic optimization.
Abdurakhmon Sadiev, Marina Danilova, Eduard Gorbunov, Samuel Horváth, Gauthier Gidel, Pavel E. Dvurechensky, Alexander V. Gasnikov, Peter Richtárik
ICML1
2022 Optimal Algorithms for Decentralized Stochastic Variational Inequalities
abstract
Variational inequalities are a formalism that includes games, minimization, saddle point, and equilibrium problems as special cases. Methods for variational inequalities are therefore universal approaches for many applied tasks, including machine learning problems. This work concentrates on the decentralized setting, which is increasingly important but not well understood. In particular, we consider decentralized stochastic (sum-type) variational inequalities over fixed and time-varying networks. We present lower complexity bounds for both communication and local iterations and construct optimal algorithms that match these lower bounds. Our algorithms are the best among the available literature not only in the decentralized stochastic case, but also in the decentralized deterministic and non-distributed stochastic cases. Experimental results confirm the effectiveness of the presented algorithms.
Dmitry Kovalev, Aleksandr Beznosikov, Abdurakhmon Sadiev, Michael Persiianov, Peter Richtárik, Alexander V. Gasnikov
NeurIPS3
2022 Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact Prox
abstract
Inspired by a recent breakthrough of Mishchenko et al. [2022], who for the first time showed that local gradient steps can lead to provable communication acceleration, we propose an alternative algorithm which obtains the same communication acceleration as their method (ProxSkip). Our approach is very different, however: it is based on the celebrated method of Chambolle and Pock [2011], with several nontrivial modifications: i) we allow for an inexact computation of the prox operator of a certain smooth strongly convex function via a suitable gradient-based method (e.g., GD or Fast GD), ii) we perform a careful modification of the dual update step in order to retain linear convergence. Our general results offer the new state-of-the-art rates for the class of strongly convex-concave saddle-point problems with bilinear coupling characterized by the absence of smoothness in the dual function. When applied to federated learning, we obtain a theoretically better alternative to ProxSkip: our method requires fewer local steps ($\mathcal{O}(\kappa^{1/3})$ or $\mathcal{O}(\kappa^{1/4})$, compared to $\mathcal{O}(\kappa^{1/2})$ of ProxSkip), and performs a deterministic number of local steps instead. Like ProxSkip, our method can be applied to optimization over a connected network, and we obtain theoretical improvements here as well.
Abdurakhmon Sadiev, Dmitry Kovalev, Peter Richtárik
NeurIPS1