Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yubo Miao

dblp:48/6858 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 50% GPUs and heterogeneous computing · 50%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
1.012026
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching · AAAI 2026
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
1.012026
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching · AAAI 2026
Natural language and speech › Language models and text generation
large language model inference
0.312026
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching · AAAI 2026

Methods — techniques the papers use, named apart from their topics

l2 cache · 2.0asynchronous prefetching · 2.0
YearPublicationVenuePosition
2026 Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
abstract
Large Language Models (LLMs) exhibit pronounced memory-bound characteristics during inference due to High Bandwidth Memory (HBM) bandwidth constraints. In this paper, we propose an L2 Cache-oriented asynchronous KV Cache prefetching method to break through the memory bandwidth bottleneck in LLM inference through computation-load overlap. By strategically scheduling idle memory bandwidth during active computation windows, our method proactively prefetches required KV Cache into GPU L2 cache, enabling high-speed L2 cache hits for subsequent accesses and effectively hiding HBM access latency within computational cycles. Extensive experiments on NVIDIA H20 GPUs demonstrate that the proposed method achieves 2.15× improvement in attention kernel efficiency and up to 1.97× end-to-end throughput enhancement, surpassing state-of-the-art baseline FlashAttention-3. Notably, our solution maintains orthogonality to existing optimization techniques and can be integrated with current inference frameworks, providing a scalable latency-hiding solution for next-generation LLM inference engines.
Yanhao Dong, Yubo Miao, Weinan Li, Jiesheng Wu, Feng Lyu 0001
AAAI2