EDBT 2026 Demo / reviewers in the wild / expert
Yubo Miao
dblp:48/6858
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 50% GPUs and heterogeneous computing · 50% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache management |
1.0 | 1 | 2026 | Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching · AAAI 2026 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
1.0 | 1 | 2026 | Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model inference |
0.3 | 1 | 2026 | Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
l2 cache · 2.0asynchronous prefetching · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating LLM Inference Throughput via Asynchronous KV Cache PrefetchingabstractLarge Language Models (LLMs) exhibit pronounced memory-bound characteristics during inference due to High Bandwidth Memory (HBM) bandwidth constraints. In this paper, we propose an L2 Cache-oriented asynchronous KV Cache prefetching method to break through the memory bandwidth bottleneck in LLM inference through computation-load overlap. By strategically scheduling idle memory bandwidth during active computation windows, our method proactively prefetches required KV Cache into GPU L2 cache, enabling high-speed L2 cache hits for subsequent accesses and effectively hiding HBM access latency within computational cycles. Extensive experiments on NVIDIA H20 GPUs demonstrate that the proposed method achieves 2.15× improvement in attention kernel efficiency and up to 1.97× end-to-end throughput enhancement, surpassing state-of-the-art baseline FlashAttention-3. Notably, our solution maintains orthogonality to existing optimization techniques and can be integrated with current inference frameworks, providing a scalable latency-hiding solution for next-generation LLM inference engines. Yanhao Dong, Yubo Miao, Weinan Li, Jiesheng Wu, Feng Lyu 0001 |
AAAI | 2 |