EDBT 2026 Demo / reviewers in the wild / expert
Tu Yi
dblp:410/3971
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Language models and text generation · 50% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › decoding
efficient decoding |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Machine learning › Efficient and distributed learning
KV cache management |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model inference
KV cache sharing |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
hierarchical KV sharing · 0.9greedy algorithm · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value SharingabstractThe frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leading to methods that either employ fixed sparsity patterns or dynamically select critical tokens based on the query. While dynamic sparse patterns have proven to be more effective, they introduce significant computational overhead, as critical tokens must be reselected for each self-attention computation. In this paper, we reveal substantial similarities in KV cache token criticality across neighboring queries, layers, and heads. Motivated by this insight, we propose HShare, a hierarchical KV sharing framework. HShare facilitates the sharing of critical KV cache token indices across layers, heads, and queries, which significantly reduces the computational overhead associated with query-aware dynamic token sparsity. In addition, we introduce a greedy algorithm that dynamically determines the optimal layer-level and head-level sharing configuration for the decoding phase. We evaluate the effectiveness and efficiency of HShare across various tasks using three models: LLaMA2-7b, LLaMA3-70b, and Mistral-7b. Experimental results demonstrate that HShare achieves competitive accuracy with different sharing ratios, while delivering up to an $8.6\times$ speedup in self-attention operations and a $2.7\times$ improvement in end-to-end throughput compared with FlashAttention2 and GPT-fast respectively. The source code is publicly available at ~\url{https://github.com/wuhuaijin/HShare}. Huaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi, Jihang Zhang, Minghui Yu, Junchi Yan |
ICLR | 4 |