VLDB 2026 Research / reviewers in the wild / expert
Tianyu Guo 0009
dblp:389/7247
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0005-2979-4486ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 41% Parallel and multicore computing · 32% Cloud and datacenter computing · 14% | |
| Artificial intelligence
2 papers |
Language models and text generation · 77% Efficient and distributed learning · 23% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
pipeline parallelism |
1.7 | 2 | 2025 | gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token Throttling · SC 2025 DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism · NeurIPS 2025 |
Cloud and datacenter computing
inference serving |
0.9 | 1 | 2025 | DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism · NeurIPS 2025 |
Memory systems
cache |
0.8 | 1 | 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level Parallelism · DAC 2024 |
GPUs and heterogeneous computing
GPU architecture |
0.8 | 1 | 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level Parallelism · DAC 2024 |
Memory systems › memory hierarchy › cache hierarchy management
last-level cache management |
0.8 | 1 | 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level Parallelism · DAC 2024 |
Memory systems
shared memory |
0.8 | 1 | 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level Parallelism · DAC 2024 |
Machine learning › Efficient and distributed learning › inference serving
large language model serving |
0.3 | 1 | 2025 | gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token Throttling · SC 2025 |
Memory systems
cache management |
0.3 | 1 | 2025 | DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism · NeurIPS 2025 |
Parallel and multicore computing
thread-level parallelism |
0.2 | 1 | 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level Parallelism · DAC 2024 |
Methods — techniques the papers use, named apart from their topics
token throttling · 1.7pipeline parallelism · 1.7latency prediction · 1.7hybrid scheduling · 1.7chunked prefill · 1.7asynchronous cache migration · 1.7software-managed cache · 0.8online profiling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mpache: Interaction Aware Multi-level Cache Bypassing on GPUsabstractGraphics Processing Units (GPUs) are essential for general-purpose applications and are commonly leveraging multi-level caches to alleviate memory access pressure. However, the default cache management may lose opportunities for optimal performance in different applications. Although existing cache bypassing techniques tend to address this challenge, these methods predominantly concentrate on single-level cache, thus restricting their potential for further enhancements. To mitigate this issue, we propose Mpache, a novel software-based mechanism designed to bypass multi-level caches based on the characterization of load instructions. Mpache constructs an interaction graph and analyzes the cooperation and contention among instructions. Then, the profiling data of bypassing effectiveness guides Mpache to select the appropriate cache levels to bypass for each instruction. Finally, the design is integrated into the compiler to enable automatic bypassing for existing workloads. Evaluations on off-the-shelf GPUs show that Mpache achieves an average 1.15× speedup over the default cache policy, and effectively outperforms prior arts. Mengyue Xi, Tianyu Guo 0009, Xuanteng Huang, Zejia Lin 0001, Xianwei Zhang 0001 |
ASP-DAC | 2 |
| 2025 | EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
Tianyu Guo 0009, Hande Dong, Yichong Leng, Cheater Lin, Nong Xiao 0001, Xianwei Zhang 0001 |
Euro-Par (2) | 1 |
| 2025 | DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline ParallelismabstractTo accelerate large language model (LLM) inference, pipeline parallelism partitions model layers into sequential stages, each assigned to a different device for concurrent execution. However, this method often suffers from pipeline bubbles caused by imbalanced computation in the tail stage. While upstream stages focus solely on layer-forward operations, the final stage must also handle post-processing tasks like sampling, introducing significant latency. This uneven workload leads to pipeline misalignment, forcing upstream stages to idle and degrading overall performance. Existing frameworks typically distribute layers evenly across stages without accounting for computational load differences. To address this, we propose DynaPipe, a dynamic layer redistribution scheme that adaptively balances computation by predicting execution latency in real time. Moreover, we introduce an asynchronous key-value (KV) cache migration coordinator to enable
non-blocking layer redistribution during inference. Experiments on representative LLMs demonstrate that DynaPipe reduces average end-to-end request latency by 8% to 49% across diverse workloads, outperforming state-of-the-art pipeline parallelism systems. Hongxin Xu, Tianyu Guo 0009, Xianwei Zhang 0001 |
NeurIPS | 2 |
| 2025 | gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token ThrottlingabstractPipeline parallelism has emerged as a predominant approach for deploying large language models (LLMs) across distributed nodes, owing to its lower communication overhead compared to tensor parallelism. While demonstrating high throughput in request serving, pipeline parallelism often faces performance limitations caused by pipeline bubbles, which are primarily resulted from imbalanced computation delays across batches. Existing methods like Sarathi-Serve attempt to address this through hybrid scheduling of chunked prefill and decode tokens with a fixed token budget. However, such methods may still experience significant fluctuations, arising either from insufficient prefill tokens or uneven distribution of decode tokens, ultimately leading to computational imbalance. Tianyu Guo 0009, Xianwei Zhang 0001, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001, Yutong Lu |
SC | 1 |
| 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level ParallelismabstractWhile designed for massive parallelism, GPUs are frequently suffering from low thread occupancy and limited data throughput, which are typically attributed to constrained on-chip resources, such as shared memory and register file. To alleviate the pressure, last-level cache (LLC) is being substantially enlarged to support continuously growing computation and to shrink the off-chip data traffic. Nevertheless, applications can be challenging to fully utilize the excessive LLC spaces. Towards the issue, we propose to manage partial LLC in an software way instead to expand precious shared memory (SMEM), named as SMILE, helping to alleviate the low thread occupancy. SMILE splits the monolithic LLC into normal data cache and new software region, with the latter being to extend the limited SMEM. For adapting to diverse application characteristics, SMILE enables multiple splitting grades and determines the appropriate partition via online profiling. Experimental results show that SMILE achieves average performance improvements of 14.7% and 8.4% respectively, compared to the default baseline and prior state-of-the-art. Tianyu Guo 0009, Xuanteng Huang, Xianwei Zhang 0001, Nong Xiao 0001 |
DAC | 1 |