Alina Shutova

dblp:354/8941 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 73% Language models and text generation · 27%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Machine learning › Efficient and distributed learning
KV cache
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Natural language and speech › Language models and text generation
large language model inference
0.912025
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding
0.912025
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025
Parallel and multicore computing › parallel computing › parallel machine learning
parallel inference
0.912025
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › quantization
adaptive quantization
0.312025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025

Methods — techniques the papers use, named apart from their topics

concurrent attention cache · 1.7rotary position embeddings · 0.9rotary position embedding · 0.9compact adapters · 0.9adaptive quantization · 0.9
YearPublicationVenuePosition
2025 Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
abstract
Efficient real-world deployments of large language models (LLMs) rely on Key-Value (KV) caching for processing and generating long outputs, reducing the need for repetitive computation. For large contexts, Key-Value caches can take up tens of gigabytes of device memory, as they store vector representations for each token and layer. Recent work has shown that the cached vectors can be compressed through quantization, pruning or merging, but these techniques often compromise quality towards higher compression rates. In this work, we aim to improve Key \& Value compression by exploiting two observations: 1) the inherent dependencies between keys and values across different layers, and 2) the existence of high-compression methods for internal network states (e.g. attention Keys \& Values). We propose AQUA-KV, an adaptive quantization for Key-Value caches that relies on compact adapters to exploit existing dependencies between Keys and Values, and aims to "optimally" compress the information that cannot be predicted. AQUA-KV significantly improves compression rates, while maintaining high accuracy on state-of-the-art LLM families. On Llama 3.2 LLMs, we achieve near-lossless inference at 2-2.5 bits per value with under $1\%$ relative error in perplexity and LongBench scores. AQUA-KV is one-shot, simple, and efficient: it can be calibrated on a single GPU within 1-6 hours, even for 70B models.
Alina Shutova, Vladimir Malinovskii, Vage Egiazarian, Denis Kuznedelev, Denis Mazur, Nikita Surkov, Ivan Ermakov, Dan Alistarh
ICML1
2025 Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
abstract
Large Language Models (LLMs) have demonstrated the ability to tackle increasingly complex tasks through advanced reasoning, long-form content generation, and tool use. Solving these tasks often involves long inference-time computations. In human problem solving, a common strategy to expedite work is collaboration: by dividing the problem into sub-tasks, exploring different strategies concurrently, etc. Recent research has shown that LLMs can also operate in parallel by implementing explicit cooperation frameworks, such as voting mechanisms or the explicit creation of independent sub-tasks that can be executed in parallel. However, each of these frameworks may not be suitable for all types of tasks, which can hinder their applicability. In this work, we propose a different design approach: we run LLM "workers" in parallel , allowing them to synchronize via a concurrently-updated attention cache and prompt these workers to decide how best to collaborate. Our approach allows the instances to come up with their own collaboration strategy for the problem at hand, all the while "seeing" each other's partial progress in the concurrent cache. We implement this approach via Hogwild! Inference: a parallel LLM inference engine where multiple instances of the same LLM run in parallel with the same attention cache, with "instant" access to each other's generated tokens. Hogwild! inference takes advantage of Rotary Position Embeddings (RoPE) to avoid recomputation while improving parallel hardware utilization. We find that modern reasoning-capable LLMs can perform inference with shared Key-Value cache out of the box, without additional fine-tuning.
Gleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev, Erik Schultheis, Vage Egiazarian, Anton Sinitsin, Denis Kuznedelev, Dan Alistarh
NeurIPS3