VLDB 2026 Research / reviewers in the wild / expert
Diyu Zhou
dblp:223/6823
· DBLP profile ↗
25ranked-venue papers
6as first author
22since 2021 · last 2026
0009-0003-8620-1064ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 9 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlowGPU: Transparent and Efficient GPU Checkpointing and Restore
Zehua Yang, Yonghao Zou, Junyang Zhang 0003, Zhisheng Ye 0002, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Diyu Zhou |
Euro-Par (2) | 10 |
| 2026 | Xerxes: Extensive Exploration of Scalable Hardware Systems with CXL-Based Simulation Framework
Yuda An, Shushu Yi, Bo Mao 0003, Qiao Li 0001, Mingzhe Zhang 0005, Diyu Zhou, Ke Zhou 0001, Nong Xiao 0001, Guangyu Sun 0003, Yingwei Luo, Jie Zhang 0048 |
FAST | 6 |
| 2026 | Latency-SLO-Aware Memory Offloading for Large Language Model InferenceabstractOffloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. Chenxiang Ma, Zhisheng Ye 0002, Zehua Yang, Tianhao Fu, Jiaxun Han, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yong Li 0045, Diyu Zhou |
ICS | 12 |
| 2026 | A Comprehensive Study on Solving Memory Bloat Under VirtualizationabstractHuge pages are effective in reducing address translation overhead under virtualization. However, huge pages can lead to the memory bloat problem, which manifests in two primary forms: hot bloat and usage bloat . Hot bloat occurs when accesses to a huge page are heavily skewed towards a small subset of base pages, leading the hypervisor to (mistakenly) classify the entire huge page as hot. Hot bloat undermines several critical virtualization techniques, including tiered memory and page sharing. Usage bloatrefers to the base pages within a huge page that has not yet been allocated, causing virtual machines (VMs) to demand excessive memory. Prior work addressing memory bloat either requires hardware modification or targets a specific scenario and is not applicable to a hypervisor. This article presents HugeScope , a lightweight, effective and generic system that addresses the memory bloat problem under virtualization based on commodity hardware. HugeScope includes an efficient and precise page tracking mechanism, leveraging the other level of indirect memory translation in the hypervisor. HugeScope provides a generic framework to support page splitting and coalescing policies, considering the memory pressure, as well as the recency, frequency, and skewness of page access. Moreover, HugeScope is general and modular. It can not only be easily applied to various scenarios concerning hot bloat , including tiered memory management ( HS-TMM ) and page sharing ( HS-Share ), but also seamlessly expose its capabilities to VMs to address the usage bloat problem ( HS-HP ). Evaluation shows that HugeScope incurs less than 4% overhead, by addressing hot bloat , HS-TMM improves performance by up to 61% over vTMM while HS-Share saves 41% more memory than Ingens while offering comparable performance, and By addressing usage bloat , HS-HP can eliminate excessive memory usage, and achieve performance improvements of up to 11% over HawkEye. Chuandong Li 0004, Dong Liu 0042, Zhihong Xue, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Diyu Zhou |
ACM Trans. Comput. Syst. | 8 |
| 2025 | SPDK+: Low Latency or High Power Efficiency? We Take BothabstractSPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK. Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048 |
HotStorage | 5 |
| 2025 | Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsabstractTo reduce operational costs, modern data centers co-locate high-priority latency-critical (LC) tasks and low-priority best-effort (BE) tasks on the same physical node to increase resource utilization. However, such co-location leads to contention for memory bandwidth, resulting in priority inversion, where BE tasks severely slow down LC tasks. This priority inversion often leads to violations of the quality of service (QoS) requirements for LC tasks, defeating the purpose of co-location. Prior approaches to this issue either fail to enforce the QoS requirements for LC tasks or underutilize memory bandwidth.We present Pivot, a novel bandwidth partitioning system that overcomes the limitations of prior approaches based on two key insights. First, memory accesses from LC tasks must be prioritized across all the components on the memory path rather than a single component, as done in prior work. Second, only the scheduling of a selective portion of performance-critical loads (i.e., those causing a long stall on the re-order buffer), instead of all memory accesses from LC tasks, should be prioritized. To leverage these insights, Pivot overcomes the key challenge of accurately identifying performance-critical loads while incurring minimal runtime overhead by proposing a two-phase profiling technique. Our extensive evaluation shows that Pivot improves effective machine utilization by up to $\mathbf{3 4. 5 \%}$ while increasing the throughput of the BE applications by up to $2.76 \times$ compared to state-of-the-art approaches. Liren Zhu, Liujia Li, Jie Zhang 0048, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Diyu Zhou |
HPCA | 10 |
| 2025 | Blackbox Fuzzing of Distributed Systems with Multi-Dimensional Inputs and Symmetry-Based Feedback Pruning
Yonghao Zou, Jia-Ju Bai, Zu-Ming Jiang, Diyu Zhou |
NDSS | 5 |
| 2025 | Aeolia: A Fast and Secure Userspace Interrupt-Based Storage StackabstractPolling-based userspace storage stacks achieve great I/O performance. However, they cannot efficiently and securely share disks and CPUs among multiple tasks. In contrast, interrupt-based kernel stacks inherently suffer from subpar I/O performance but achieve advantages in resource sharing. Chuandong Li 0004, Ran Yi 0004, Zonghao Zhang, Jing Liu 0074, Changwoo Min, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou |
SOSP | 10 |
| 2025 | Analyzing and Enhancing ArckFS: An Anecdotal Example of Benefits of Artifact EvaluationabstractWe analyze and enhance Trio and ArckFS by Zhou et al. (SOSP 2023), high-performance NVM file system architecture and file system. A group of authors from KAIST initiated this study through a careful review of the paper and the released artifact, seeking to enhance the Trio work. Their analysis identifies (1) insufficient clarity in the paper on the handling of multi-inode operations, and (2) several implementation bugs in ArckFS that cause occasional operation failures or potential crash inconsistencies during inode creation. Jonguk Jeon, Subeen Park, Sanidhya Kashyap, Sudarsun Kannan, Diyu Zhou, Jeehoon Kang |
SOSP | 5 |
| 2025 | CortenMM: Efficient Memory Management with Strong Correctness GuaranteesabstractModern memory management systems suffer from poor performance and subtle concurrency bugs, slowing down applications while introducing security vulnerabilities. We observe that both issues stem from the conventional design of memory management systems with two levels of abstraction: a software-level abstraction (e.g., VMA trees in Linux) and a hardware-level abstraction (typically, page tables). This design increases portability but requires correctly and efficiently synchronizing two drastically different and complex data structures, which is generally challenging. Junyang Zhang 0003, Xiangcan Xu, Yonghao Zou, Xinyi Wan 0001, Siyuan Wang 0026, Di Wang 0017, Hao Chen 0023, Lin Huang 0005, Shoumeng Yan, Yuval Tamir, Yingwei Luo, Xiaolin Wang 0001, Huashan Yu, Zhenlin Wang 0003, Hongliang Tian, Diyu Zhou |
SOSP | 19 |
| 2025 | ASTERINAS: A Linux ABI-Compatible, Rust-Based Framekernel OS with a Small and Sound TCB
Yuke Peng, Hongliang Tian, Junyang Zhang 0003, Jinyi Xian, Xiaolin Wang 0001, Chenren Xu, Diyu Zhou, Yingwei Luo, Shoumeng Yan, Yinqian Zhang |
USENIX ATC | 10 |
| 2025 | A Sound Static Analysis Approach to I/O API MigrationabstractThe advances in modern storage technologies necessitate the development of new input/output (I/O) APIs to maximize their performance benefits. However, migrating existing software to use different APIs poses significant challenges due to mismatches in computational models and complex code structures surrounding stateful, non-contiguous multi-API call sites. We present Sprout, a new system for automatically migrating programs across I/O APIs that guarantees behavioral equivalence. Sprout uses flow-sensitive pointer analysis to identify semantic variables, which enables the typestate analysis for matching API semantics and the synthesis of migrated programs. Experimental results with real-world c programs highlight the efficiency and effectiveness of our approach. We also show that Sprout can be adapted to other domains, such as databases. Sizhe Zhong, Diyu Zhou, Jiasi Shen 0001 |
Proc. ACM Program. Lang. | 4 |
| 2024 | EKRM: Efficient Key-Value Retrieval Method to Reduce Data Lookup Overhead for Redis
Xiaolin Wang 0001, Diyu Zhou, Liujia Li, Liren Zhu, Zhenlin Wang 0003, Yingwei Luo |
Euro-Par (1) | 3 |
| 2024 | Transparent Multicore Scaling of Single-Threaded Network FunctionsabstractThis paper presents NFOS, a programming model, runtime, and profiler for productively developing software network functions (NFs) that scale on multicore machines. Writing shared-state concurrent systems that are both correct and scalable is still a serious challenge, which is why NFOS insulates developers from writing concurrent code. Lei Yan 0003, Yueyang Pan, Diyu Zhou, George Candea, Sanidhya Kashyap |
EuroSys | 3 |
| 2024 | Practical Verification of System-Software Components Written in Standard CabstractSystems code is challenging to verify, because it uses constructs (like raw pointers, pointer arithmetic, and bit twiddling) that are hard for tools to reason about. Existing approaches either sacrifice programmer friendliness, by demanding significant manual effort and verification expertise, or generality, by restricting the programming language or requiring that the code adapt to the verification tool. Can Cebeci, Yonghao Zou, Diyu Zhou, George Candea, Clément Pit-Claudel |
SOSP | 3 |
| 2024 | Taming Hot Bloat Under Virtualization with HUGESCOPE
Chuandong Li 0004, Sai Sha, Yangqing Zeng, Xiran Yang, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou |
USENIX ATC | 8 |
| 2023 | TENET: Memory Safe and Fault Tolerant Persistent Transactional Memory
Madhava Krishnan Ramanathan, Diyu Zhou, Wook-Hee Kim, Sudarsun Kannan, Sanidhya Kashyap, Changwoo Min |
FAST | 2 |
| 2023 | Ship your Critical Section, Not Your Data: Enabling Transparent Delegation with TCLOCKS
Vishal Gupta 0006, Kumar Kartikeya Dwivedi, Yugesh Kothari, Yueyang Pan, Diyu Zhou, Sanidhya Kashyap |
OSDI | 5 |
| 2023 | Enabling High-Performance and Secure Userspace NVM File Systems with the Trio ArchitectureabstractUserspace library file systems (LibFSes) promise to unleash the performance potential of non-volatile memory (NVM) by directly accessing it and enabling unprivileged applications to customize their LibFSes to their workloads. Unfortunately, such benefits pose a significant challenge to ensuring metadata integrity. Existing works either underutilize NVM's performance or forgo critical file system security guarantees. Diyu Zhou, Vojtech Aschenbrenner, Tao Lyu 0004, Sudarsun Kannan, Sanidhya Kashyap |
SOSP | 1 |
| 2022 | Application-Informed Kernel Synchronization Primitives
Diyu Zhou, Yuchen Qian, Irina Calciu, Taesoo Kim, Sanidhya Kashyap |
OSDI | 2 |
| 2022 | ODINFS: Scaling PM Performance with Opportunistic Delegation
Diyu Zhou, Yuchen Qian, Vishal Gupta 0006, Changwoo Min, Sanidhya Kashyap |
OSDI | 1 |
| 2022 | RRC: Responsive Replicated Containers
Diyu Zhou, Yuval Tamir |
USENIX ATC | 1 |
| 2020 | Fault-Tolerant Containers Using NiLiConabstractMany services deployed in the cloud require high reliability and must thus survive machine failures. Providing such fault tolerance transparently, without requiring application modifications, has motivated extensive research on replicating virtual machines (VMs). Cloud computing typically relies on VMs or containers to provide an isolation and multitenancy layer. Containers have advantages over VMs in smaller size, faster startup, and avoiding the need to manage updates of multiple VMs. This paper reports on the design, implementation, and evaluation of NiLiCon — a transparent container replication mechanism for fault tolerance. To the best of our knowledge, NiLiCon is the first implementation of container replication, demonstrating that it can be used for transparent deployment of critical services in the cloud.NiLiCon is based on high-frequency asynchronous incremental checkpointing to a warm spare, as previously used for VMs. The challenge to accomplishing this is that, compared to VMs, there is much tighter coupling between the container state and the state of the underlying platform. NiLiCon meets this challenge, eliminating the need to deploy services in VMs, with performance overheads that are competitive with those of similar VM replication mechanisms. Specifically, with the seven benchmarks used in the evaluation, the performance overhead of NiLiCon is in the range of 19%-67%. For fail-stop faults, the recovery rate is 100%. Diyu Zhou, Yuval Tamir |
IPDPS | 1 |
| 2019 | PUSh: Data Race Detection Based on Hardware-Supported Prevention of Unintended SharingabstractSome of the most difficult to find bugs in multi-threaded programs are caused by unintended sharing, leading to data races. The detection of data races can be facilitated by requiring programmers to explicitly specify any intended sharing and then verifying compliance with these intentions. We present a novel dynamic checker based on this approach, called PUSh. Diyu Zhou, Yuval Tamir |
MICRO | 1 |
| 2018 | Fast Hypervisor Recovery Without RebootabstractSystem recovery latency is decreased by using microreboot to reboot only the failed component instead of the entire system. For large, complex components, such as hypervisors, even the latency of microreboot is unacceptably high in important deployment scenarios. We investigate an alternative component-level recovery mechanism, which we call microreset, that can achieve dramatically lower recovery latency for some such components. Instead of component reboot, microreset quickly resets the component to a quiescent state that is highly likely to be valid and where the component is ready to handle new or retried interactions with the rest of the system. We present a recovery mechanism for the Xen hypervisor, called NiLiHype, based on microreset. We show that, compared to microreboot-based hypervisor recovery, NiLiHype achieves nearly the same recovery success rate but with a recovery latency that is shorter by a factor of over 30. Diyu Zhou, Yuval Tamir |
DSN | 1 |