VLDB 2026 Research / reviewers in the wild / expert
Yifan Qiao 0002
dblp:200/8215-2
· DBLP profile ↗
14ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0003-3651-6973ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 8 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BlendServe: Optimizing Offline Inference with Resource-Aware BatchingabstractOffline batch inference is gaining popularity as a cost-effective solution for latency-insensitive tasks, such as model evaluation and data curation. As the latency objective is highly relaxed, maximizing throughput is the primary goal in offline inference. Previous studies focused solely on throughput optimization within a batch. However, the diverse resource demands (compute-intensive vs. memory-intensive) across a wide range of applications make these approaches less effective, as imbalanced resource demands between batches restrict optimization opportunities. Yilong Zhao 0002, Shuo Yang 0011, Kan Zhu, Lianmin Zheng, Baris Kasikci, Yifan Qiao 0002, Yang Zhou 0008, Jiarong Xing, Ion Stoica |
ASPLOS (2) | 6 |
| 2025 | Lost in Translation: The Search for Meaning in Network-Attached AI Accelerator DisaggregationabstractDatacenters often underutilize expensive AI accelerators (GPUs, TPUs, etc). A natural solution is disaggregation, where servers borrow network-attached accelerators on demand. However, current approaches to disaggregation suffer from a semantic translation gap: as computation descends the software stack, critical application knowledge—like model structure or execution phases—is lost. This forces an undesirable choice between low-level, general-purpose systems that are semantically-blind and inefficient, and high-level, single-workload systems that are efficient but not general. Jaewan Hong, Yifan Qiao 0002, Soujanya Ponnapalli, Marcos K. Aguilera, Vincent Liu 0001, Christopher J. Rossbach, Ion Stoica |
HotNets | 2 |
| 2025 | PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model ApplicationsabstractBesides typical generative applications, like ChatGPT, GitHub Copilot, and Cursor, we observe an emerging trend that LLMs are increasingly used in traditional discriminative tasks, such as recommendation, credit verification, and data labeling. The key characteristic of these emerging use cases is that the LLM generates only a single output token, rather than an arbitrarily long sequence of tokens. We refer to this as a prefill-only workload. However, since existing LLM engines assume arbitrary output lengths, they fail to leverage the unique properties of prefill-only workloads. In this paper, we present PrefillOnly, the first LLM inference engine that improves the inference throughput and latency by fully embracing the properties of prefill-only workloads. First, since it generates only one token, PrefillOnly only needs to store the KV cache of only the last computed layer, rather than of all layers. This drastically reduces the GPU memory footprint of LLM inference and allows handling long inputs without using solutions that reduce throughput, such as cross-GPU KV cache parallelization. Second, because the output length is fixed, rather than arbitrary, PrefillOnly can precisely determine the job completion time (JCT) of each prefill-only request before it starts. This enables efficient JCT-aware scheduling policies such as shortest prefill first. PrefillOnly can process up to 4× larger queries per second without inflating the average and P99 latency. Kuntai Du, Bowen Wang 0016, Chen Zhang 0001, Qing Lan, Hejian Sang, Yihua Cheng, Yifan Qiao 0002, Ion Stoica, Junchen Jiang |
SOSP | 10 |
| 2025 | Orthrus: Efficient and Timely Detection of Silent User Data Corruption in the Cloud with Resource-Adaptive Computation ValidationabstractEven with substantial endeavors to test and validate processors, computational errors may still arise post-installation. One particular category of CPU errors transpires discreetly, without crashing applications or triggering hardware warnings. These elusive errors pose a significant threat by undermining user data, and their detection is challenging. This paper introduces Orthrus, a solution for the timely detection of silent user data corruption caused by post-installation CPU errors. Orthrus safeguards user data in cloud applications by providing simple annotations and compiler support for users to identify data operators and validating these operators asynchronously across cores while maintaining a low overhead (2%–6%), making it practical for production deployment. Our evaluation, using carefully injected errors, demonstrates that Orthrus can detect 87% of data corruptions with just a single core dedicated to validation, increasing to 91% and 96% when two and four cores are used, respectively. Chenxiao Liu, Zhenting Zhu, Quanxi Li, Yanwen Xia, Yifan Qiao 0002, Xiangyun Deng, Youyou Lu, Tao Xie 0001, Huimin Cui, Zidong Du, Guoqing Harry Xu, Chenxi Wang 0005 |
SOSP | 5 |
| 2024 | Harvesting Idle Memory for Application-managed Soft State with Midas
Yifan Qiao 0002, Zhenyuan Ruan, Adam Belay, Miryung Kim, Guoqing Harry Xu |
NSDI | 1 |
| 2024 | A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
Chenxi Wang 0005, Yifan Qiao 0002, Zhe Wang 0017, Chenggang Wu 0002, Youyou Lu, Xiaobing Feng 0002, Huimin Cui, Shan Lu 0001, Guoqing Harry Xu |
OSDI | 5 |
| 2024 | DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
Yifan Qiao 0002, Shan Yu 0001, Yuanjiang Ni, Qingda Lu, Jiesheng Wu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu |
OSDI | 2 |
| 2023 | Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony
Yifan Qiao 0002, Chenxi Wang 0005, Zhenyuan Ruan, Adam Belay, Qingda Lu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu |
NSDI | 1 |
| 2023 | Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
John Thorpe, Pengzhan Zhao, Jon Eyolfson, Yifan Qiao 0002, Minjia Zhang, Ravi Netravali, Guoqing Harry Xu |
NSDI | 4 |
| 2023 | Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory
Chenxi Wang 0005, Yifan Qiao 0002, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
NSDI | 2 |
| 2022 | MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime
Chenxi Wang 0005, Yifan Qiao 0002, Jon Eyolfson, Christian Navasca, Shan Lu 0001, Guoqing Harry Xu |
OSDI | 4 |
| 2022 | Mako: a low-pause, high-throughput evacuating collector for memory-disaggregated datacentersabstractResource disaggregation has gained much traction as an emerging datacenter architecture, as it improves resource utilization and simplifies hardware adoption. Under resource disaggregation, different types of resources (memory, CPUs, etc.) are disaggregated into dedicated servers connected by high-speed network fabrics. Memory disaggregation brings efficiency challenges to concurrent garbage collection (GC), which is widely used for latency-sensitive cloud applications, because GC and mutator threads simultaneously run and constantly compete for memory and swap resources. Chenxi Wang 0005, Yifan Qiao 0002, Michael D. Bond, Steve Blackburn, Miryung Kim, Guoqing Harry Xu |
PLDI | 4 |
| 2021 | Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
John Thorpe, Yifan Qiao 0002, Jon Eyolfson, Shen Teng, Guanzhou Hu, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
OSDI | 2 |
| 2017 | Algorithm-Directed Crash Consistence in Non-volatile Memory for HPCabstractFault tolerance is one of the major design goals for HPC. The emergence of non-volatile memories (NVM) provides a solution to build fault tolerant HPC. Data in NVM-based main memory are not lost when the system crashes because of the non-volatility nature of NVM. However, because of volatile caches, data must be logged and explicitly flushed from caches into NVM to ensure consistence and correctness before crashes, which can cause large runtime overhead. In this paper, we introduce an algorithm-based method to establish crash consistence in NVM for HPC applications. We slightly extend application data structures or sparsely flush cache blocks, which introduce ignorable runtime overhead. Such extension or cache flushing allows us to use algorithm knowledge to reason data consistence or correct inconsistent data when the application crashes. We demonstrate the effectiveness of our method for three algorithms, including an iterative solver, dense matrix multiplication, and Monte-Carlo simulation. Based on comprehensive performance evaluation on a variety of test environments, we demonstrate that our approach has very small runtime overhead (at most 8.2% and less than 3% in most cases), much smaller than that of traditional checkpoint, while having the same or less recomputation cost. Shuo Yang 0011, Kai Wu 0006, Yifan Qiao 0002, Dong Li 0001, Jidong Zhai |
CLUSTER | 3 |