EDBT 2026 Demo / reviewers in the wild / expert
Alina Shutova
dblp:354/8941
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 73% Language models and text generation · 27% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025 |
Machine learning › Efficient and distributed learning
KV cache |
0.9 | 1 | 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.9 | 1 | 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | Hogwild! Inference: Parallel LLM Generation via Concurrent Attention · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding |
0.9 | 1 | 2025 | Hogwild! Inference: Parallel LLM Generation via Concurrent Attention · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025 |
Parallel and multicore computing › parallel computing › parallel machine learning
parallel inference |
0.9 | 1 | 2025 | Hogwild! Inference: Parallel LLM Generation via Concurrent Attention · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
adaptive quantization |
0.3 | 1 | 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
concurrent attention cache · 1.7rotary position embeddings · 0.9rotary position embedding · 0.9compact adapters · 0.9adaptive quantization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cache Me If You Must: Adaptive Key-Value Quantization for Large Language ModelsabstractEfficient real-world deployments of large language models (LLMs) rely on Key-Value (KV) caching for processing and generating long outputs, reducing the need for repetitive computation. For large contexts, Key-Value caches can take up tens of gigabytes of device memory, as they store vector representations for each token and layer. Recent work has shown that the cached vectors can be compressed through quantization, pruning or merging, but these techniques often compromise quality towards higher compression rates. In this work, we aim to improve Key \& Value compression by exploiting two observations: 1) the inherent dependencies between keys and values across different layers, and 2) the existence of high-compression methods for internal network states (e.g. attention Keys \& Values). We propose AQUA-KV, an adaptive quantization for Key-Value caches that relies on compact adapters to exploit existing dependencies between Keys and Values, and aims to "optimally" compress the information that cannot be predicted. AQUA-KV significantly improves compression rates, while maintaining high accuracy on state-of-the-art LLM families. On Llama 3.2 LLMs, we achieve near-lossless inference at 2-2.5 bits per value with under $1\%$ relative error in perplexity and LongBench scores. AQUA-KV is one-shot, simple, and efficient: it can be calibrated on a single GPU within 1-6 hours, even for 70B models. Alina Shutova, Vladimir Malinovskii, Vage Egiazarian, Denis Kuznedelev, Denis Mazur, Nikita Surkov, Ivan Ermakov, Dan Alistarh |
ICML | 1 |
| 2025 | Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionabstractLarge Language Models (LLMs) have demonstrated the ability to tackle increasingly complex tasks through advanced reasoning, long-form content generation, and tool use. Solving these tasks often involves long inference-time computations. In human problem solving, a common strategy to expedite work is collaboration: by dividing the problem into sub-tasks, exploring different strategies concurrently, etc. Recent research has shown that LLMs can also operate in parallel by implementing explicit cooperation frameworks, such as voting mechanisms or the explicit creation of independent sub-tasks that can be executed in parallel. However, each of these frameworks may not be suitable for all types of tasks, which can hinder their applicability. In this work, we propose a different design approach: we run LLM "workers" in parallel , allowing them to synchronize via a concurrently-updated attention cache and prompt these workers to decide how best to collaborate. Our approach allows the instances to come up with their own collaboration strategy for the problem at hand, all the while "seeing" each other's partial progress in the concurrent cache. We implement this approach via Hogwild! Inference: a parallel LLM inference engine where multiple instances of the same LLM run in parallel with the same attention cache, with "instant" access to each other's generated tokens. Hogwild! inference takes advantage of Rotary Position Embeddings (RoPE) to avoid recomputation while improving parallel hardware utilization. We find that modern reasoning-capable LLMs can perform inference with shared Key-Value cache out of the box, without additional fine-tuning. Gleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev, Erik Schultheis, Vage Egiazarian, Anton Sinitsin, Denis Kuznedelev, Dan Alistarh |
NeurIPS | 3 |