Yifan Qiao 0002

dblp:200/8215-2 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0003-3651-6973ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 8 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 BlendServe: Optimizing Offline Inference with Resource-Aware Batching
abstract
Offline batch inference is gaining popularity as a cost-effective solution for latency-insensitive tasks, such as model evaluation and data curation. As the latency objective is highly relaxed, maximizing throughput is the primary goal in offline inference. Previous studies focused solely on throughput optimization within a batch. However, the diverse resource demands (compute-intensive vs. memory-intensive) across a wide range of applications make these approaches less effective, as imbalanced resource demands between batches restrict optimization opportunities.
Yilong Zhao 0002, Shuo Yang 0011, Kan Zhu, Lianmin Zheng, Baris Kasikci, Yifan Qiao 0002, Yang Zhou 0008, Jiarong Xing, Ion Stoica
ASPLOS (2)6
2025 Lost in Translation: The Search for Meaning in Network-Attached AI Accelerator Disaggregation
abstract
Datacenters often underutilize expensive AI accelerators (GPUs, TPUs, etc). A natural solution is disaggregation, where servers borrow network-attached accelerators on demand. However, current approaches to disaggregation suffer from a semantic translation gap: as computation descends the software stack, critical application knowledge—like model structure or execution phases—is lost. This forces an undesirable choice between low-level, general-purpose systems that are semantically-blind and inefficient, and high-level, single-workload systems that are efficient but not general.
Jaewan Hong, Yifan Qiao 0002, Soujanya Ponnapalli, Marcos K. Aguilera, Vincent Liu 0001, Christopher J. Rossbach, Ion Stoica
HotNets2
2025 PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
abstract
Besides typical generative applications, like ChatGPT, GitHub Copilot, and Cursor, we observe an emerging trend that LLMs are increasingly used in traditional discriminative tasks, such as recommendation, credit verification, and data labeling. The key characteristic of these emerging use cases is that the LLM generates only a single output token, rather than an arbitrarily long sequence of tokens. We refer to this as a prefill-only workload. However, since existing LLM engines assume arbitrary output lengths, they fail to leverage the unique properties of prefill-only workloads. In this paper, we present PrefillOnly, the first LLM inference engine that improves the inference throughput and latency by fully embracing the properties of prefill-only workloads. First, since it generates only one token, PrefillOnly only needs to store the KV cache of only the last computed layer, rather than of all layers. This drastically reduces the GPU memory footprint of LLM inference and allows handling long inputs without using solutions that reduce throughput, such as cross-GPU KV cache parallelization. Second, because the output length is fixed, rather than arbitrary, PrefillOnly can precisely determine the job completion time (JCT) of each prefill-only request before it starts. This enables efficient JCT-aware scheduling policies such as shortest prefill first. PrefillOnly can process up to 4× larger queries per second without inflating the average and P99 latency.
Kuntai Du, Bowen Wang 0016, Chen Zhang 0001, Qing Lan, Hejian Sang, Yihua Cheng, Yifan Qiao 0002, Ion Stoica, Junchen Jiang
SOSP10
2025 Orthrus: Efficient and Timely Detection of Silent User Data Corruption in the Cloud with Resource-Adaptive Computation Validation
abstract
Even with substantial endeavors to test and validate processors, computational errors may still arise post-installation. One particular category of CPU errors transpires discreetly, without crashing applications or triggering hardware warnings. These elusive errors pose a significant threat by undermining user data, and their detection is challenging. This paper introduces Orthrus, a solution for the timely detection of silent user data corruption caused by post-installation CPU errors. Orthrus safeguards user data in cloud applications by providing simple annotations and compiler support for users to identify data operators and validating these operators asynchronously across cores while maintaining a low overhead (2%–6%), making it practical for production deployment. Our evaluation, using carefully injected errors, demonstrates that Orthrus can detect 87% of data corruptions with just a single core dedicated to validation, increasing to 91% and 96% when two and four cores are used, respectively.
Chenxiao Liu, Zhenting Zhu, Quanxi Li, Yanwen Xia, Yifan Qiao 0002, Xiangyun Deng, Youyou Lu, Tao Xie 0001, Huimin Cui, Zidong Du, Guoqing Harry Xu, Chenxi Wang 0005
SOSP5
2024 Harvesting Idle Memory for Application-managed Soft State with Midas
Yifan Qiao 0002, Zhenyuan Ruan, Adam Belay, Miryung Kim, Guoqing Harry Xu
NSDI1
2024 A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
Chenxi Wang 0005, Yifan Qiao 0002, Zhe Wang 0017, Chenggang Wu 0002, Youyou Lu, Xiaobing Feng 0002, Huimin Cui, Shan Lu 0001, Guoqing Harry Xu
OSDI5
2024 DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
Yifan Qiao 0002, Shan Yu 0001, Yuanjiang Ni, Qingda Lu, Jiesheng Wu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu
OSDI2
2023 Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony
Yifan Qiao 0002, Chenxi Wang 0005, Zhenyuan Ruan, Adam Belay, Qingda Lu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu
NSDI1
2023 Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
John Thorpe, Pengzhan Zhao, Jon Eyolfson, Yifan Qiao 0002, Minjia Zhang, Ravi Netravali, Guoqing Harry Xu
NSDI4
2023 Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory
Chenxi Wang 0005, Yifan Qiao 0002, Ravi Netravali, Miryung Kim, Guoqing Harry Xu
NSDI2
2022 MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime
Chenxi Wang 0005, Yifan Qiao 0002, Jon Eyolfson, Christian Navasca, Shan Lu 0001, Guoqing Harry Xu
OSDI4
2022 Mako: a low-pause, high-throughput evacuating collector for memory-disaggregated datacenters
abstract
Resource disaggregation has gained much traction as an emerging datacenter architecture, as it improves resource utilization and simplifies hardware adoption. Under resource disaggregation, different types of resources (memory, CPUs, etc.) are disaggregated into dedicated servers connected by high-speed network fabrics. Memory disaggregation brings efficiency challenges to concurrent garbage collection (GC), which is widely used for latency-sensitive cloud applications, because GC and mutator threads simultaneously run and constantly compete for memory and swap resources.
Chenxi Wang 0005, Yifan Qiao 0002, Michael D. Bond, Steve Blackburn, Miryung Kim, Guoqing Harry Xu
PLDI4
2021 Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
John Thorpe, Yifan Qiao 0002, Jon Eyolfson, Shen Teng, Guanzhou Hu, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, Guoqing Harry Xu
OSDI2
2017 Algorithm-Directed Crash Consistence in Non-volatile Memory for HPC
abstract
Fault tolerance is one of the major design goals for HPC. The emergence of non-volatile memories (NVM) provides a solution to build fault tolerant HPC. Data in NVM-based main memory are not lost when the system crashes because of the non-volatility nature of NVM. However, because of volatile caches, data must be logged and explicitly flushed from caches into NVM to ensure consistence and correctness before crashes, which can cause large runtime overhead. In this paper, we introduce an algorithm-based method to establish crash consistence in NVM for HPC applications. We slightly extend application data structures or sparsely flush cache blocks, which introduce ignorable runtime overhead. Such extension or cache flushing allows us to use algorithm knowledge to reason data consistence or correct inconsistent data when the application crashes. We demonstrate the effectiveness of our method for three algorithms, including an iterative solver, dense matrix multiplication, and Monte-Carlo simulation. Based on comprehensive performance evaluation on a variety of test environments, we demonstrate that our approach has very small runtime overhead (at most 8.2% and less than 3% in most cases), much smaller than that of traditional checkpoint, while having the same or less recomputation cost.
Shuo Yang 0011, Kai Wu 0006, Yifan Qiao 0002, Dong Li 0001, Jidong Zhai
CLUSTER3