VLDB 2026 Research / reviewers in the wild / expert
Yizhou Shan
dblp:206/4531
· DBLP profile ↗
23ranked-venue papers
4as first author
17since 2021 · last 2026
0009-0007-9519-0546ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-author · 12 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingabstractMultiple Low-Rank Adapters (Multi-LoRA) are gaining popularity for task-specific Large Language Model (LLM) applications. For Multi-LoRA serving, caching hot LoRAs and KV caches in the GPU memory can improve inference performance. However, existing Multi-LoRA inference systems fail to optimize serving performance like Time-To-First-Token (TTFT), neglecting usage dependencies when caching LoRAs and KV caches. We therefore propose ELORA, a Multi-LoRA caching system to optimize the serving performance. ELORA comprises a dependency-aware cache manager and a performancedriven cache swapper. The cache manager maintains the usage dependencies between LoRAs and KV caches during inference with a unified caching pool. The cache swapper determines the swap-in or swap-out of LoRAs and KV caches based on a unified cost model, when the GPU memory is idle or busy, respectively. Experimental results show that ELORA reduces the TTFT by$\mathbf{4 5. 7 \%}$on average, compared to state-of-the-art works. Jiuchen Shi, Quan Chen 0002, Yizhou Shan, Kaihua Fu, Wei Wang 0030, Minyi Guo |
HPCA | 5 |
| 2026 | Hermes: Efficient Serving of LLM Applications with Probabilistic Demand ModelingabstractApplications based on Large Language Models (LLMs) contain a series of tasks to address real-world problems with boosted capability, which have dynamic demand volumes on diverse backends. Existing serving systems treat the resource demands of LLM applications as a blackbox, compromising end-to-end efficiency due to improper queuing order and backend warm up latency. We find that the resource demands of LLM applications can be modeled in a general and accurate manner with Probabilistic Demand Graph (PDGraph). We then propose Hermes, which leverages PDGraph for efficient serving of LLM applications. Confronting probabilistic demand description, Hermes applies the Gittins policy to determine the scheduling order that can minimize the average application completion time. It also uses the PDGraph model to help prewarm cold backends at proper moments. Experiments with diverse LLM applications confirm that Hermes can effectively improve the application serving efficiency, reducing the average completion time by over 70% and the P95 completion time by over 80%. Zuo Gan, Zhenghao Gan, Chen Chen 0067, Yizhou Shan, Xusheng Chen, Zhenhua Han, Yifei Zhu 0001, Shixuan Sun, Minyi Guo |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceabstractThe widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constrained scenarios (e.g., edge computing). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstAttention, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstAttention designs a dedicated flashaware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstAttention improves throughput for long-sequence inference by up to $11.1 \times$, compared to existing SSD-based solutions such as FlexGen. Xiurui Pan, Endian Li, Qiao Li 0001, Shengwen Liang, Yizhou Shan, Ke Zhou 0001, Yingwei Luo, Xiaolin Wang 0001, Jie Zhang 0048 |
HPCA | 5 |
| 2025 | EPIC: Efficient Position-Independent Caching for Serving Large Language ModelsabstractLarge Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become more complex. Context caching improves serving performance by reusing Key-Value (KV) vectors, the intermediate representations of tokens that are repeated across requests. However, existing context caching requires exact prefix matches across requests, limiting reuse cases in settings such as few-shot learning and retrieval-augmented generation, where immutable content (e.g., documents) remains unchanged across requests but is preceded by varying prefixes. Position-Independent Caching (PIC) addresses this issue by enabling modular reuse of the KV vectors regardless of prefixes. We formalize PIC and advance prior work by introducing EPIC, a serving system incorporating our new LegoLink algorithm, which mitigates the inappropriate “attention sink” effect at every document beginning, to maintain accuracy with minimal computation. Experiments show that EPIC achieves up to 8$\times$ improvements in Time-To-First-Token (TTFT) and 7$\times$ throughput gains over existing systems, with negligible or no accuracy loss. Wenrui Huang, Haoyi Wang, Tiancheng Hu, Xusheng Chen, Yizhou Shan, Tao Xie 0001 |
ICML | 9 |
| 2025 | Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded Programs
Quanxi Li, Ying Liu 0055, Yanwen Xia, Jie Zhang 0048, Mosong Zhou, Xiaobing Feng 0002, Huimin Cui, Yizhou Shan, Chenxi Wang 0005 |
NSDI | 10 |
| 2025 | BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching
Dingyan Zhang, Xingda Wei, Yizhou Shan, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 5 |
| 2025 | DHAP: Towards Efficient OLAP in a Disaggregated and Heterogeneous EnvironmentabstractDisaggregation of hardware resources and integration of heterogeneous accelerators are two emerging trends in datacenters. Existing data systems focus on either disaggregated systems with homogeneous CPU processors or incorporation of heterogeneous accelerators within traditional monolithic servers. None can adequately address the challenges posed by systems that are both disaggregated and heterogeneous. Guangda Liu, Chenqi Zhang 0002, Yizhou Shan, Zeke Wang, Shixuan Sun, Minyi Guo, Jieru Zhao |
SC | 3 |
| 2025 | Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference
Suyi Li 0002, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng 0001, Xusheng Chen, Yizhou Shan, Binhang Yuan, Wei Wang 0030 |
USENIX ATC | 7 |
| 2025 | DEEPSERVE: Serverless Large Language Model Serving at Scale
Zhixia Liu, Yuetao Chen, Baoquan Zhang, Shining Wan, Gengyuan Dan, Zhiyu Dong, Zhihao Ren, Changhong Liu, Tao Xie 0001, Dayun Lin, Xusheng Chen, Yizhou Shan |
USENIX ATC | 21 |
| 2025 | DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack Communication
Xu Zhang 0033, Ke Liu 0004, Yuan Hui 0001, Yisong Chang, Yizhou Shan, Ke Zhang 0017, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
USENIX ATC | 6 |
| 2025 | ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream WorkloadsabstractTransformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase. However, existing LLM deployment practices often overlook the distinct characteristics of these phases, leading to significant interference. To mitigate interference, our insight is to carefully schedule and group inference requests based on their characteristics. We realize this idea in ShuffleInfer through three pillars. First, it partitions prompts into fixed-size chunks so that the accelerator always runs close to its computation-saturated limit. Second, it disaggregates prefill and decode instances so each can run independently. Finally, it uses a smart two-level scheduling algorithm augmented with predicted resource usage to avoid decode scheduling hotspots. Results show that ShuffleInfer improves time-to-first-token (TTFT), job completion time (JCT), and inference efficiency in terms of performance per dollar by a large margin, e.g., it uses 38% less resources all the while lowering average TTFT and average JCT by 97% and 47%, respectively. Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Chenxi Wang 0005, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan |
ACM Trans. Archit. Code Optim. | 12 |
| 2024 | SuperNIC: An FPGA-Based, Cloud-Oriented SmartNICabstractWith CPU scaling slowing down in today's data centers, more functionalities are being offloaded from the CPU to auxiliary devices. One such device is the SmartNIC, which is being increasingly adopted in data centers. In today's cloud environment, VMs on the same server can each have their own network computation (or network tasks) or workflows of network tasks to offload to a SmartNIC. These network tasks can be dynamically added/removed as VMs come and go and can be shared across VMs. Such dynamism demands that a SmartNIC not only schedules and processes packets but also manages and executes offloaded network tasks for different users. Although software solutions like an OS exist for managing software-based network tasks, such software-based SmartNICs cannot keep up with the quickly increasing data-center network speed. This paper proposes a new SmartNIC platform called SuperNIC that allows multiple tenants to efficiently and safely offload FPGA-based network computation DAGs. For efficiency and scalability, our core idea is to group network tasks into virtual chains that are dynamically mapped to different forms of physical chains depending on load and FPGA space availability. We further propose techniques to automatically scale network task chains with different types of parallelism. Moreover, we propose a fair sharing mechanism that considers both fair space sharing and fair time sharing of different types of hardware resources. Our FPGA prototype of SuperNIC achieves high bandwidth and low latency performance whilst efficiently utilizing and fairly sharing resources. Will Lin, Yizhou Shan, Ryan Kosta, Arvind Krishnamurthy, Yiying Zhang 0005 |
FPGA | 2 |
| 2023 | MARB: Bridge the Semantic Gap between Operating System and Application Memory Access BehaviorabstractThe virtual memory subsystem (VMS) is a long-standing and integral part of an operating system (OS). It plays a vital role in enabling remote memory systems over fast data center networks and is promising in terms of transparency and generality. Specifically, these systems use three VMS mechanisms: demand paging, page swapping, and page prefetching. However, the VMS inherent data path is costly, which takes a huge toll on performance. Despite prior efforts to propose page swapping and prefetching algorithms to minimize the occurrences of the data path, they still fall short due to the semantic gap between the OS and applications - the VMS has limited knowledge of its running applications' memory access behaviors. In this paper, orthogonal to prior efforts, we take a fundamen-tally different approach by building an efficient framework to collect full memory access traces at the local bus, and make them available to the OS through CPU cache. Consequently, the page swapping and page prefetching can use this trace to make better decisions, thereby improving the overall performance of systems. We implement a proof-of-concept prototype on commodity x86 servers using a hardware-based memory tracking tool. To show-case our framework's benefits, we integrate it with a state-of-the-art remote memory system and the default kernel page eviction subsystem. Our evaluation shows promising improvements. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yisong Chang, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
DATE | 11 |
| 2023 | Skadi: Building a Distributed Runtime for Data Systems in Disaggregated Data CentersabstractData-intensive systems are the backbone of today's computing and are responsible for shaping data centers. Over the years, cloud providers have relied on three principles to maintain cost-effective data systems: use disaggregation to decouple scaling, use domain-specific computing to battle waning laws, and use serverless to lower costs. Although they work well individually, they fail to work in harmony: an issue amplified by emerging data system workloads. Cunchen Hu, Chenxi Wang 0005, Sa Wang, Ninghui Sun, Yungang Bao, Jieru Zhao, Sanidhya Kashyap, Pengfei Zuo, Xusheng Chen, Liangliang Xu, Yizhou Shan |
HotOS | 13 |
| 2023 | HoPP: Hardware-Software Co-Designed Page Prefetching for Disaggregated MemoryabstractMemory disaggregation is a promising direction to mitigate memory contention in datacenters. To make memory disaggregation practical, prior efforts expose remote memory to applications transparently via virtual memory subsystem’s swapping interface. However, due to the semantic gap between OS and applications – OS cannot know the memory accessing sequences of an application but via page faults. This approach has two limitations. First, it learns little from page faults’ access history, which leads to sub-optimal prefetching predictions. Second, a page fault can still occur even if there is a prefetch-hit which leads to a large kernel overhead.To address such limitations, our key insight is to decouple the address capturing from page faults by collecting full memory access traces in the memory controller. Using this idea, we buildHoPP– a hardware-software co-designed prefetching framework.HoPPadds hardware modules to the memory controller to feed sufficient hot pages to OS in real-time, which has three benefits inHoPP’s software design: 1) it improves existing prefetching algorithms with simple revamps, also offers more insights to build better policies; 2) the prefetch algorithm can run as a separate data path alongside the normal remote data path via page faults, potentially hiding the swap latency from applications, and enabling fine-grained control over prefetching behaviors; 3) the prefetch-hit overhead can be eliminated by early page table entry (PTE) injection, i.e., inject PTE for the prefetched page as soon as it returns. We implemented a proof-of-concept prototype using commodity servers along with a hardware-based memory tracking tool calledHMTTto emulate a modified memory controller. Results show that compared to Fastswap and Leap,HoPP-optimized prefetching algorithm achieves over 90% accuracy and coverage, which leads to up to 59% completion time improvement for various datacenter applications. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
HPCA | 10 |
| 2023 | Core slicing: closing the gap between leaky confidential VMs and bare-metal cloud
Ziqiao Zhou, Yizhou Shan, Weidong Cui, Xinyang Ge, Marcus Peinado, Andrew Baumann |
OSDI | 2 |
| 2022 | Clio: a hardware-software co-designed disaggregated memory systemabstractMemory disaggregation has attracted great attention recently because of its benefits in efficient memory utilization and ease of management. So far, memory disaggregation research has all taken one of two approaches: building/emulating memory nodes using regular servers or building them using raw memory devices with no processing power. The former incurs higher monetary cost and faces tail latency and scalability limitations, while the latter introduces performance, security, and management problems. Yizhou Shan, Xuhao Luo, Yutong Huang, Yiying Zhang 0005 |
ASPLOS | 2 |
| 2020 | Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value Stores
Shin-Yeh Tsai, Yizhou Shan, Yiying Zhang 0005 |
USENIX ATC | 2 |
| 2019 | Storm: a fast transactional dataplane for remote data structuresabstractRDMA technology enables a host to access the memory of a remote host without involving the remote CPU, improving the performance of distributed in-memory storage systems. Previous studies argued that RDMA suffers from scalability issues, because the NIC's limited resources are unable to simultaneously cache the state of all the concurrent network streams. These concerns led to various software-based proposals to reduce the size of this state by trading off performance. Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui, Yiying Zhang 0005, Haggai Eran, Boris Pismenny, Liran Liss, Michael Wei, Dan Tsafrir, Marcos K. Aguilera |
SYSTOR | 2 |
| 2019 | LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation
Yizhou Shan, Yutong Huang, Yiying Zhang 0005 |
USENIX ATC | 1 |
| 2018 | LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation
Yizhou Shan, Yutong Huang, Yiying Zhang 0005 |
OSDI | 1 |
| 2017 | Disaggregated operating systemabstractRecently, there is an emerging trend to move towards a disaggregated hardware architecture that breaks monolithic servers into independent hardware components that are connected to a fast, scalable network [1, 2]. The disaggregated architecture offers several benefits over traditional monolithic server model, including better resource utilization, ease of hardware deployment, and support for heterogeneity. Our vision of the future disaggregated architecture is that each component will have its own controller to manage its hardware and can communicate with other components through a fast network. Yizhou Shan, Sumukh Hallymysore, Yutong Huang, Yiying Zhang 0005 |
SoCC | 1 |
| 2017 | Distributed shared persistent memoryabstractNext-generation non-volatile memories (NVMs) will provide byte addressability, persistence, high density, and DRAM-like performance. They have the potential to benefit many datacenter applications. However, most previous research on NVMs has focused on using them in a single machine environment. It is still unclear how to best utilize them in distributed, datacenter environments. Yizhou Shan, Shin-Yeh Tsai, Yiying Zhang 0005 |
SoCC | 1 |