Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Denis Mazur

dblp:250/2970 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 84% Language models and text generation · 8% Representation and self-supervised learning · 4%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.622025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
quantization
1.622025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Machine learning › Efficient and distributed learning
KV cache
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.812024
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware fine-tuning
0.812024
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression · NeurIPS 2024
Machine learning › Efficient and distributed learning
collaborative learning
0.512021
Distributed Deep Learning In Open Collaborations · NeurIPS 2021
Machine learning › Efficient and distributed learning
distributed training
0.512021
Distributed Deep Learning In Open Collaborations · NeurIPS 2021
Machine learning › Graph learning
graph representation
0.412019
Beyond Vector Spaces: Compact Data Representation as Differentiable Weighted Graphs · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › geometric embedding
non-euclidean embedding
0.412019
Beyond Vector Spaces: Compact Data Representation as Differentiable Weighted Graphs · NeurIPS 2019
Machine learning › Efficient and distributed learning › model compression › quantization
adaptive quantization
0.312025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025

Methods — techniques the papers use, named apart from their topics

compact adapters · 0.9adaptive quantization · 0.9vector quantization · 0.8straight-through estimator · 0.8fine-tuning · 0.8SwAV pretraining · 0.5ALBERT pretraining · 0.5shortest path distance · 0.4gradient descent · 0.4
YearPublicationVenuePosition
2025 Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
abstract
Efficient real-world deployments of large language models (LLMs) rely on Key-Value (KV) caching for processing and generating long outputs, reducing the need for repetitive computation. For large contexts, Key-Value caches can take up tens of gigabytes of device memory, as they store vector representations for each token and layer. Recent work has shown that the cached vectors can be compressed through quantization, pruning or merging, but these techniques often compromise quality towards higher compression rates. In this work, we aim to improve Key \& Value compression by exploiting two observations: 1) the inherent dependencies between keys and values across different layers, and 2) the existence of high-compression methods for internal network states (e.g. attention Keys \& Values). We propose AQUA-KV, an adaptive quantization for Key-Value caches that relies on compact adapters to exploit existing dependencies between Keys and Values, and aims to "optimally" compress the information that cannot be predicted. AQUA-KV significantly improves compression rates, while maintaining high accuracy on state-of-the-art LLM families. On Llama 3.2 LLMs, we achieve near-lossless inference at 2-2.5 bits per value with under $1\%$ relative error in perplexity and LongBench scores. AQUA-KV is one-shot, simple, and efficient: it can be calibrated on a single GPU within 1-6 hours, even for 70B models.
Alina Shutova, Vladimir Malinovskii, Vage Egiazarian, Denis Kuznedelev, Denis Mazur, Nikita Surkov, Ivan Ermakov, Dan Alistarh
ICML5
2024 PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
abstract
There has been significant interest in "extreme" compression of large language models (LLMs), i.e. to 1-2 bits per parameter, which allows such models to be executed efficiently on resource-constrained devices. Existing work focused on improved one-shot quantization techniques and weight representations; yet, purely post-training approaches are reaching diminishing returns in terms of the accuracy-vs-bit-width trade-off. State-of-the-art quantization methods such as QuIP# and AQLM include fine-tuning (part of) the compressed parameters over a limited amount of calibration data; however, such fine-tuning techniques over compressed weights often make exclusive use of straight-through estimators (STE), whose performance is not well-understood in this setting. In this work, we question the use of STE for extreme LLM compression, showing that it can be sub-optimal, and perform a systematic study of quantization-aware fine-tuning strategies for LLMs. We propose PV-Tuning - a representation-agnostic framework that generalizes and improves upon existing fine-tuning strategies, and provides convergence guarantees in restricted cases. On the practical side, when used for 1-2 bit vector quantization, PV-Tuning outperforms prior techniques for highly-performant models such as Llama and Mistral. Using PV-Tuning, we achieve the first Pareto-optimal quantization for Llama-2 family models at 2 bits per parameter.
Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtárik
NeurIPS2
2021 Distributed Deep Learning In Open Collaborations
abstract
Modern deep learning applications require increasingly more compute to train state-of-the-art models. To address this demand, large corporations and institutions use dedicated High-Performance Computing clusters, whose construction and maintenance are both environmentally costly and well beyond the budget of most organizations. As a result, some research directions become the exclusive domain of a few large industrial and even fewer academic actors. To alleviate this disparity, smaller groups may pool their computational resources and run collaborative experiments that benefit all participants. This paradigm, known as grid- or volunteer computing, has seen successful applications in numerous scientific areas. However, using this approach for machine learning is difficult due to high latency, asymmetric bandwidth, and several challenges unique to volunteer computing. In this work, we carefully analyze these constraints and propose a novel algorithmic framework designed specifically for collaborative training. We demonstrate the effectiveness of our approach for SwAV and ALBERT pretraining in realistic conditions and achieve performance comparable to traditional setups at a fraction of the cost. Finally, we provide a detailed report of successful collaborative language model pretraining with nearly 50 participants.
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Quentin Lhoest, Anton Sinitsin, Dmitry Popov 0003, Dmitry V. Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, Denis Mazur, Ilia Kobelev, Yacine Jernite, Thomas Wolf 0008, Gennady Pekhimenko
NeurIPS12
2019 Beyond Vector Spaces: Compact Data Representation as Differentiable Weighted Graphs
abstract
Learning useful representations is a key ingredient to the success of modern machine learning. Currently, representation learning mostly relies on embedding data into Euclidean space. However, recent work has shown that data in some domains is better modeled by non-euclidean metric spaces, and inappropriate geometry can result in inferior performance. In this paper, we aim to eliminate the inductive bias imposed by the embedding space geometry. Namely, we propose to map data into more general non-vector metric spaces: a weighted graph with a shortest path distance. By design, such graphs can model arbitrary geometry with a proper configuration of edges and weights. Our main contribution is PRODIGE: a method that learns a weighted graph representation of data end-to-end by gradient descent. Greater generality and fewer model assumptions make PRODIGE more powerful than existing embedding-based approaches. We confirm the superiority of our method via extensive experiments on a wide range of tasks, including classification, compression, and collaborative filtering.
Denis Mazur, Vage Egiazarian, Stanislav Morozov, Artem Babenko
NeurIPS1