Tu Yi

dblp:410/3971 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 50% Language models and text generation · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › decoding
efficient decoding
0.912025
HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025
Machine learning › Efficient and distributed learning
KV cache management
0.912025
HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025
Natural language and speech › Language models and text generation › large language model inference
KV cache sharing
0.912025
HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025

Methods — techniques the papers use, named apart from their topics

hierarchical KV sharing · 0.9greedy algorithm · 0.9
YearPublicationVenuePosition
2025 HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing
abstract
The frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leading to methods that either employ fixed sparsity patterns or dynamically select critical tokens based on the query. While dynamic sparse patterns have proven to be more effective, they introduce significant computational overhead, as critical tokens must be reselected for each self-attention computation. In this paper, we reveal substantial similarities in KV cache token criticality across neighboring queries, layers, and heads. Motivated by this insight, we propose HShare, a hierarchical KV sharing framework. HShare facilitates the sharing of critical KV cache token indices across layers, heads, and queries, which significantly reduces the computational overhead associated with query-aware dynamic token sparsity. In addition, we introduce a greedy algorithm that dynamically determines the optimal layer-level and head-level sharing configuration for the decoding phase. We evaluate the effectiveness and efficiency of HShare across various tasks using three models: LLaMA2-7b, LLaMA3-70b, and Mistral-7b. Experimental results demonstrate that HShare achieves competitive accuracy with different sharing ratios, while delivering up to an $8.6\times$ speedup in self-attention operations and a $2.7\times$ improvement in end-to-end throughput compared with FlashAttention2 and GPT-fast respectively. The source code is publicly available at ~\url{https://github.com/wuhuaijin/HShare}.
Huaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi, Jihang Zhang, Minghui Yu, Junchi Yan
ICLR4