EDBT 2026 Demo / reviewers in the wild / expert
Jihang Zhang
dblp:153/8045
· DBLP profile ↗
4ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-1579-0661ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Language models and text generation · 50% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › decoding
efficient decoding |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Machine learning › Efficient and distributed learning
KV cache management |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model inference
KV cache sharing |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
hierarchical KV sharing · 0.9greedy algorithm · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value SharingabstractThe frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leading to methods that either employ fixed sparsity patterns or dynamically select critical tokens based on the query. While dynamic sparse patterns have proven to be more effective, they introduce significant computational overhead, as critical tokens must be reselected for each self-attention computation. In this paper, we reveal substantial similarities in KV cache token criticality across neighboring queries, layers, and heads. Motivated by this insight, we propose HShare, a hierarchical KV sharing framework. HShare facilitates the sharing of critical KV cache token indices across layers, heads, and queries, which significantly reduces the computational overhead associated with query-aware dynamic token sparsity. In addition, we introduce a greedy algorithm that dynamically determines the optimal layer-level and head-level sharing configuration for the decoding phase. We evaluate the effectiveness and efficiency of HShare across various tasks using three models: LLaMA2-7b, LLaMA3-70b, and Mistral-7b. Experimental results demonstrate that HShare achieves competitive accuracy with different sharing ratios, while delivering up to an $8.6\times$ speedup in self-attention operations and a $2.7\times$ improvement in end-to-end throughput compared with FlashAttention2 and GPT-fast respectively. The source code is publicly available at ~\url{https://github.com/wuhuaijin/HShare}. Huaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi, Jihang Zhang, Minghui Yu, Junchi Yan |
ICLR | 5 |
| 2025 | SALS: Sparse Attention in Latent Space for KV Cache CompressionabstractLarge Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics within the hidden dimension, suggesting the potential for effective compression. However, due to the widely adopted Rotary Position Embedding (RoPE) mechanism in modern LLMs, naive low‑-rank compression suffers severe accuracy degradation or creates a new speed bottleneck, as the low-rank cache must first be reconstructed in order to apply RoPE. In this paper, we introduce two key insights: first, the application of RoPE to the key vectors increases their variance, which in turn results in a higher rank; second, after the key vectors are transformed into the latent space, they largely maintain their representation across most layers. Based on these insights, we propose the Sparse Attention in Latent Space (SALS) framework. SALS projects the KV cache into a compact latent space via low-rank projection, and performs sparse token selection using RoPE-free query--key interactions in this space. By reconstructing only a small subset of important tokens, it avoids the overhead of full KV cache reconstruction. We comprehensively evaluate SALS on various tasks using two large-scale models: LLaMA2-7b-chat and Mistral-7b, and additionally verify its scalability on the RULER-128k benchmark with LLaMA3.1-8B-Instruct. Experimental results demonstrate that SALS achieves SOTA performance by maintaining competitive accuracy. Under different settings, SALS achieves 6.4-fold KV cache compression and 5.7-fold speed-up in the attention operator compared to FlashAttention2 on the 4K sequence. For the end-to-end throughput performance, we achieves 1.4-fold and 4.5-fold improvement compared to GPT-fast on 4k and 32K sequences, respectively. The source code will be publicly available in the future. Junlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu, Tao Wang 0011, Yidong Li |
NeurIPS | 3 |
| 2015 | Bayesian-based preference prediction in bilateral multi-issue negotiation between intelligent agents
Jihang Zhang, Fenghui Ren, Minjie Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2014 | An Innovative Approach for Predicting Both Negotiation Deadline and Utility in Multi-issue Negotiation
Jihang Zhang, Fenghui Ren, Minjie Zhang 0001 |
PRICAI | 1 |