VLDB 2026 Research / reviewers in the wild / expert
Yilong Zhao 0002
dblp:152/7119-2
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0002-7523-7400ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BlendServe: Optimizing Offline Inference with Resource-Aware BatchingabstractOffline batch inference is gaining popularity as a cost-effective solution for latency-insensitive tasks, such as model evaluation and data curation. As the latency objective is highly relaxed, maximizing throughput is the primary goal in offline inference. Previous studies focused solely on throughput optimization within a batch. However, the diverse resource demands (compute-intensive vs. memory-intensive) across a wide range of applications make these approaches less effective, as imbalanced resource demands between batches restrict optimization opportunities. Yilong Zhao 0002, Shuo Yang 0011, Kan Zhu, Lianmin Zheng, Baris Kasikci, Yifan Qiao 0002, Yang Zhou 0008, Jiarong Xing, Ion Stoica |
ASPLOS (2) | 1 |
| 2025 | From Optimal to Practical: Efficient Micro-op Cache Replacement Policies for Data Center ApplicationsabstractOptimizing the CPU frontend has become crucial for modern processors with intricate instruction decoding logic, especially for efficiently running planet-scale data center applications. Micro-operation (micro-op) cache is a key unit to help improve the energy efficiency of the CPU frontend. Unfortunately, we find that data center applications suffer from frequent micro-op cache misses due to the lack of an effective micro-op cache replacement policy. Developing micro-op cache-specific replacement policies is challenging, as there currently does not exist an optimal theoretical solution akin to Belady’s algorithm for conventional caches. As a result, it is unknown by how much replacement policies can be improved and how to get there. To address these challenges, we introduce FLACK, a new near-optimal offline policy that considers the key features of the micro-op cache, such as variable and disproportional costs of micro-op cache misses and partial hits. We show that FLACK substantially outperforms Belady’s algorithm, thus establishing a new baseline for micro-op cache replacement policies. We then design FURBYS, a practical policy that mimics FLACK via profile-guided methods. FURBYS has three key components to perform cache replacement decisions: (1) it uses profiles of the whole-execution hit/miss behavior, (2) it detects locally (transiently) hot data, and (3) it selectively ignores data with profiled low hit rates. We evaluate FLACK and FURBYS using 11 data center applications and find that FLACK demonstrates an average bound of 30.21% miss reduction, achieving 4.46% greater miss reduction than Belady’s algorithm. Our practical policy, FURBYS, provides 14.34% average miss reduction compared to LRU, which is $1.84 \times$ greater than the current state-of-the-art replacement policy, contributing to 3.10% of performance-perwatt improvement for the CPU core. On average, in terms of miss reduction and IPC gain, FURBYS is equivalent to LRU policy on $1.5 \times$ micro-op cache sizes (up to $2 \times$), demonstrating the effectiveness of the proposed replacement policy. Kan Zhu, Yilong Zhao 0002, Peter Braun 0005, Tanvir Ahmed Khan 0001, Heiner Litz, Baris Kasikci, Shuwen Deng |
HPCA | 2 |
| 2025 | Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityabstractDiffusion Transformers (DiTs) dominate video generation but their high computational cost severely limits real-world applicability, usually requiring tens of minutes to generate a few seconds of video even on high-performance GPUs. This inefficiency primarily arises from the quadratic computational complexity of 3D full attention with respect to the context length. In this paper, we propose a training-free framework termed Sparse VideoGen (SVG) that leverages the inherent sparsity in 3D full attention to boost inference efficiency. We reveal that the attention heads can be dynamically classified into two groups depending on distinct sparse patterns: (1) Spatial Head, where only spatially-related tokens within each frame dominate the attention output, and (2) Temporal Head, where only temporally-related tokens across different frames dominate. Based on this insight, SVG proposes an online profiling strategy to capture the dynamic sparse patterns and predicts the type of attention head. Combined with a novel hardware-efficient tensor layout transformation and customized kernel implementations, SVG achieves up to 2.28$\times$ and 2.33$\times$ end-to-end speedup on CogVideoX-v1.5 and HunyuanVideo, respectively, while preserving generation quality. Our code will be open-sourced upon publication. Haocheng Xi, Shuo Yang 0011, Yilong Zhao 0002, Chenfeng Xu, Xiuyu Li, Yujun Lin 0001, Han Cai, Dacheng Li, Jianfei Chen 0001, Ion Stoica, Kurt Keutzer, Song Han 0003 |
ICML | 3 |
| 2025 | Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationabstractDiffusion Transformers (DiTs) are essential for video generation but suffer from significant latency due to the quadratic complexity of attention.
By computing only critical tokens, sparse attention reduces computational costs and offers a promising acceleration approach.
However, we identify that existing methods fail to approach optimal generation quality under the same computation budget for two reasons:
(1) Inaccurate critical token identification: current methods cluster tokens based on position rather than semantics, leading to imprecise aggregated representations.
(2) Excessive computation waste: critical tokens are scattered among non-critical ones, leading to wasted computation on GPUs, which are optimized for processing contiguous tokens.
In this paper, we propose SVG2, a training-free framework that maximizes identification accuracy and minimizes computation waste, achieving a Pareto frontier trade-off between generation quality and efficiency.
The core of SVG2 is semantic-aware permutation, which clusters and reorders tokens based on semantic similarity using k-means. This approach ensures both a precise cluster representation, improving identification accuracy, and a densified layout of critical tokens, enabling efficient computation without padding.
Additionally, SVG2 integrates Top-p dynamic budget control and customized kernel implementations, achieving up to $2.30\times$ and $1.89\times$ speedup while maintaining a PSNR of up to $30$ and $26$ on HunyuanVideo and Wan 2.1, respectively. Our code is open-sourced at https://github.com/svg-project/Sparse-VideoGen. Shuo Yang 0011, Haocheng Xi, Yilong Zhao 0002, Han Cai, Yujun Lin 0001, Xiuyu Li, Chenfeng Xu, Kelly Peng, Jianfei Chen 0001, Song Han 0003, Kurt Keutzer, Ion Stoica |
NeurIPS | 3 |
| 2025 | NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yilong Zhao 0002, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye 0001, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci |
OSDI | 3 |
| 2024 | QUEST: Query-Aware Sparsity for Efficient Long-Context LLM InferenceabstractAs the demand for long-context large language models (LLMs) increases, models with context windows of up to 128K or 1M tokens are becoming increasingly prevalent. However, long-context LLM inference is challenging since the inference speed decreases significantly as the sequence length grows. This slowdown is primarily caused by loading a large KV cache during self-attention. Previous works have shown that a small portion of critical tokens will dominate the attention outcomes. However, we observe the criticality of a token highly depends on the query. To this end, we propose Quest, a query-aware KV cache selection algorithm. Quest keeps track of the minimal and maximal Key values in KV cache pages and estimates the criticality of a given page using Query vectors. By only loading the Top-K critical KV cache pages for attention, Quest significantly speeds up self-attention without sacrificing accuracy. We show that Quest can achieve up to 2.23x self-attention speedup, which reduces inference latency by 7.03x while performing well on tasks with long dependencies with negligible accuracy loss. Code is available at https://github.com/mit-han-lab/quest. Yilong Zhao 0002, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han 0003 |
ICML | 2 |