Yelysei Bondarenko

dblp:295/8514 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 65% Deep learning architectures and training · 26% Trustworthy machine learning · 9%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.732023
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing · NeurIPS 2023
Overcoming Oscillations in Quantization-Aware Training · ICML 2022
Understanding and Overcoming the Challenges of Efficient Transformer Quantization · EMNLP (1) 2021
Machine learning › Efficient and distributed learning › model compression
quantization
1.732023
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing · NeurIPS 2023
Overcoming Oscillations in Quantization-Aware Training · ICML 2022
Understanding and Overcoming the Challenges of Efficient Transformer Quantization · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
transformer
1.222023
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing · NeurIPS 2023
Understanding and Overcoming the Challenges of Efficient Transformer Quantization · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
attention mechanism
0.712023
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing · NeurIPS 2023
Machine learning › Trustworthy machine learning
outlier mitigation
0.712023
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization
0.612022
Overcoming Oscillations in Quantization-Aware Training · ICML 2022
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware training
0.612022
Overcoming Oscillations in Quantization-Aware Training · ICML 2022

Methods — techniques the papers use, named apart from their topics

gated attention · 0.7clipped softmax · 0.7INT8 quantization · 0.7oscillation dampening · 0.6iterative weight freezing · 0.6quantization-aware training · 0.5post-training quantization · 0.5
YearPublicationVenuePosition
2023 Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
abstract
Transformer models have been widely adopted in various domains over the last years and especially large language models have advanced the field of AI significantly. Due to their size, the capability of these networks has increased tremendously, but this has come at the cost of a significant increase in necessary compute. Quantization is one of the most effective ways for reducing the computational time and memory consumption of neural networks. Many studies have shown, however, that modern transformer models tend to learn strong outliers in their activations, making them difficult to quantize. To retain acceptable performance, the existence of these outliers requires activations to be in higher-bitwidth or the use of different numeric formats, extra fine-tuning, or other workarounds. We show that strong outliers are related to very specific behavior of attention heads that try to learn a "no-op", or just a partial update of the residual. To achieve the exact zeros needed in the attention matrix for a no-update, the input to the softmax is pushed to be larger and larger during training, causing outliers in other parts of the network. Based on these observations, we propose two simple (independent) modifications to the attention mechanism - _clipped softmax_ and _gated attention_. We empirically show that models pre-trained using our methods learn significantly smaller outliers while maintaining and sometimes even improving the floating-point task performance. This enables us to quantize transformers to full INT8 quantization of the activations without any additional effort. We demonstrate the effectiveness of our methods on both language models (BERT, OPT) and vision transformers.
Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort
NeurIPS1
2022 Overcoming Oscillations in Quantization-Aware Training
abstract
When training neural networks with simulated quantization, we observe that quantized weights can, rather unexpectedly, oscillate between two grid-points. The importance of this effect and its impact on quantization-aware training (QAT) are not well-understood or investigated in literature. In this paper, we delve deeper into the phenomenon of weight oscillations and show that it can lead to a significant accuracy degradation due to wrongly estimated batch-normalization statistics during inference and increased noise during training. These effects are particularly pronounced in low-bit ($\leq$ 4-bits) quantization of efficient networks with depth-wise separable layers, such as MobileNets and EfficientNets. In our analysis we investigate several previously proposed QAT algorithms and show that most of these are unable to overcome oscillations. Finally, we propose two novel QAT algorithms to overcome oscillations during training: oscillation dampening and iterative weight freezing. We demonstrate that our algorithms achieve state-of-the-art accuracy for low-bit (3 & 4 bits) weight and activation quantization of efficient architectures, such as MobileNetV2, MobileNetV3, and EfficentNet-lite on ImageNet. Our source code is available at https://github.com/qualcomm-ai-research/oscillations-qat.
Markus Nagel, Marios Fournarakis, Yelysei Bondarenko, Tijmen Blankevoort
ICML3
2021 Understanding and Overcoming the Challenges of Efficient Transformer Quantization
abstract
Transformer-based architectures have become the de-facto standard models for a wide range of Natural Language Processing tasks.However, their memory footprint and high latency are prohibitive for efficient deployment and inference on resource-limited devices.In this work, we explore quantization for transformers.We show that transformers have unique quantization challenges -namely, high dynamic activation ranges that are difficult to represent with a low bit fixed-point format.We establish that these activations contain structured outliers in the residual connections that encourage specific attention patterns, such as attending to the special separator token.To combat these challenges, we present three solutions based on post-training quantization and quantization-aware training, each with a different set of compromises for accuracy, model size, and ease of use.In particular, we introduce a novel quantization scheme -per-embedding-group quantization.We demonstrate the effectiveness of our methods on the GLUE benchmark using BERT, establishing state-of-the-art results for post-training quantization.Finally, we show that transformer weights and embeddings can be quantized to ultra-low bit-widths, leading to significant memory savings with a minimum accuracy loss.Our source code is available at https://github. com/qualcomm-ai-research/ transformer-quantization.
Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort
EMNLP (1)1