Konstantin Burlachenko

dblp:285/5386 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 55% Optimization for machine learning · 39% Language models and text generation · 6%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
distributed optimization
2.032024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024
MARINA: Faster Non-Convex Distributed Learning with Compression · ICML 2021
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
1.532024
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024
MARINA: Faster Non-Convex Distributed Learning with Compression · ICML 2021
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Efficient and distributed learning › distributed training
gradient compression
1.322024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
MARINA: Faster Non-Convex Distributed Learning with Compression · ICML 2021
Machine learning › Efficient and distributed learning › model compression
quantization
1.022024
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Efficient and distributed learning
federated learning
0.922024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
MARINA: Faster Non-Convex Distributed Learning with Compression · ICML 2021
Machine learning › Optimization for machine learning
convergence analysis
0.812024
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024
Machine learning › Optimization for machine learning › distributed optimization
error feedback
0.812024
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024
Natural language and speech › Language models and text generation
large language model
0.812024
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware fine-tuning
0.812024
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Optimization for machine learning › stochastic gradient descent
random reshuffling
0.812024
Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences · NeurIPS 2024
Machine learning › Optimization for machine learning
stochastic optimization
0.812024
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024
Machine learning › Efficient and distributed learning
distributed training
0.722024
MARINA: Faster Non-Convex Distributed Learning with Compression · ICML 2021
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024
Machine learning › Efficient and distributed learning
communication compression
0.212024
Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants · ICLR 2024

Methods — techniques the papers use, named apart from their topics

vector quantization · 0.8variance reduction · 0.8top-k compression · 0.8straight-through estimator · 0.8stochastic gradient descent · 0.8random reshuffling · 0.8gradient quantization · 0.8fine-tuning · 0.8error feedback · 0.8control iterates · 0.8
YearPublicationVenuePosition
2024 Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants
abstract
Error feedback (EF) is a highly popular and immensely effective mechanism for fixing convergence issues which arise in distributed training methods (such as distributed GD or SGD) when these are enhanced with greedy communication compression techniques such as Top-k. While EF was proposed almost a decade ago (Seide et al, 2014), and despite concentrated effort by the community to advance the theoretical understanding of this mechanism, there is still a lot to explore. In this work we study a modern form of error feedback called EF21 (Richtárik et al, 2021) which offers the currently best-known theoretical guarantees, under the weakest assumptions, and also works well in practice. In particular, while the theoretical communication complexity of EF21 depends on the quadratic mean of certain smoothness parameters, we improve this dependence to their arithmetic mean, which is always smaller, and can be substantially smaller, especially in heterogeneous data regimes. We take the reader on a journey of our discovery process. Starting with the idea of applying EF21 to an equivalent reformulation of the underlying problem which (unfortunately) requires (often impractical) machine cloning, we continue to the discovery of a new weighted version of EF21 which can (fortunately) be executed without any cloning, and finally circle back to an improved analysis of the original EF21 method. While this development applies to the simplest form of EF21, our approach naturally extends to more elaborate variants involving stochastic gradients and partial participation. Further, our technique improves the best-known theory of EF21 in the rare features regime (Richtárik et al, 2023). Finally, we validate our theoretical findings with suitable experiments.
Peter Richtárik, Elnur Gasanov, Konstantin Burlachenko
ICLR3
2024 PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
abstract
There has been significant interest in "extreme" compression of large language models (LLMs), i.e. to 1-2 bits per parameter, which allows such models to be executed efficiently on resource-constrained devices. Existing work focused on improved one-shot quantization techniques and weight representations; yet, purely post-training approaches are reaching diminishing returns in terms of the accuracy-vs-bit-width trade-off. State-of-the-art quantization methods such as QuIP# and AQLM include fine-tuning (part of) the compressed parameters over a limited amount of calibration data; however, such fine-tuning techniques over compressed weights often make exclusive use of straight-through estimators (STE), whose performance is not well-understood in this setting. In this work, we question the use of STE for extreme LLM compression, showing that it can be sub-optimal, and perform a systematic study of quantization-aware fine-tuning strategies for LLMs. We propose PV-Tuning - a representation-agnostic framework that generalizes and improves upon existing fine-tuning strategies, and provides convergence guarantees in restricted cases. On the practical side, when used for 1-2 bit vector quantization, PV-Tuning outperforms prior techniques for highly-performant models such as Llama and Mistral. Using PV-Tuning, we achieve the first Pareto-optimal quantization for Llama-2 family models at 2 bits per parameter.
Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtárik
NeurIPS5
2024 Don't Compress Gradients in Random Reshuffling: Compress Gradient Differences
abstract
Gradient compression is a popular technique for improving communication complexity of stochastic first-order methods in distributed training of machine learning models. However, the existing works consider only with-replacement sampling of stochastic gradients. In contrast, it is well-known in practice and recently confirmed in theory that stochastic methods based on without-replacement sampling, e.g., Random Reshuffling (RR) method, perform better than ones that sample the gradients with-replacement. In this work, we close this gap in the literature and provide the first analysis of methods with gradient compression and without-replacement sampling. We first develop a distributed variant of random reshuffling with gradient compression (Q-RR), and show how to reduce the variance coming from gradient quantization through the use of control iterates. Next, to have a better fit to Federated Learning applications, we incorporate local computation and propose a variant of Q-RR called Q-NASTYA. Q-NASTYA uses local gradient steps and different local and global stepsizes. Next, we show how to reduce compression variance in this setting as well. Finally, we prove the convergence results for the proposed methods and outline several settings in which they improve upon existing algorithms.
Abdurakhmon Sadiev, Grigory Malinovsky, Eduard Gorbunov, Igor Sokolov 0001, Ahmed Khaled 0001, Konstantin Burlachenko, Peter Richtárik
NeurIPS6
2021 MARINA: Faster Non-Convex Distributed Learning with Compression
abstract
We develop and analyze MARINA: a new communication efficient method for non-convex distributed learning over heterogeneous datasets. MARINA employs a novel communication compression strategy based on the compression of gradient differences that is reminiscent of but different from the strategy employed in the DIANA method of Mishchenko et al. (2019). Unlike virtually all competing distributed first-order methods, including DIANA, ours is based on a carefully designed biased gradient estimator, which is the key to its superior theoretical and practical performance. The communication complexity bounds we prove for MARINA are evidently better than those of all previous first-order methods. Further, we develop and analyze two variants of MARINA: VR-MARINA and PP-MARINA. The first method is designed for the case when the local loss functions owned by clients are either of a finite sum or of an expectation form, and the second method allows for a partial participation of clients {–} a feature important in federated learning. All our methods are superior to previous state-of-the-art methods in terms of oracle/communication complexity. Finally, we provide a convergence analysis of all methods for problems satisfying the Polyak-{Ł}ojasiewicz condition.
Eduard Gorbunov, Konstantin Burlachenko, Zhize Li 0001, Peter Richtárik
ICML2