VLDB 2026 Research / reviewers in the wild / expert
Yirui Eric Zhou
dblp:359/0805
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0006-6195-2269ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ByteFS: System Support for (CXL-based) Memory-Semantic Solid-State DrivesabstractUnlike non-volatile memory that resides on the processor memory bus, memory-semantic solid-state drives (SSDs) support both byte and block access granularity via PCIe or CXL interconnects. They provide scalable memory capacity using NAND flash at a much lower cost. In addition, they have different performance characteristics for their dual byte/block interface respectively, while offering essential memory semantics for upper-level software. Such a byte-accessible storage device provides new implications on the software system design. Shaobo Li 0005, Yirui Eric Zhou, Hao Ren 0015, Jian Huang 0006 |
ASPLOS (1) | 2 |
| 2025 | SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-designabstractThe CXL-based solid-state drive (CXL-SSD) provides a promising approach towards scaling the main memory capacity at low cost. However, the CXL-based SSD faces performance challenges due to the long flash access latency and unpredictable events such as garbage collection in the SSD device, stalling the host processor and wasting compute cycles. Although the CXL interface enables the byte-granular data access to the SSD, accessing flash chips is still at page granularity due to physical limitations. The mismatch of access granularity causes significant unnecessary I/O traffic to flash chips, worsening the suboptimal end-to-end data access performance. In this paper, we present SkyByte, an efficient CXL-based SSD that employs a holistic approach to address the aforementioned challenges by co-designing the host operating system (OS) and SSD controller. To alleviate the long memory stall when accessing the CXL-SSD, SkyByte revisits the OS context switch mechanism and enables opportunistic context switches upon the detection of long access delays. To accommodate byte-granular data accesses, SkyByte architects the internal DRAM of the SSD controller into a cacheline-level write $\log$ and a page-level data cache, and enables data coalescing upon log cleaning to reduce the I/O traffic to flash chips. SkyByte also employs optimization techniques that include adaptive page migration for exploring the performance benefits of fast host memory by promoting hot pages in CXL-SSD to the host. We implement SkyByte with a CXL-SSD simulator and evaluate its efficiency with various data-intensive applications. Our experiments show that SkyByte outperforms current CXL-based SSD by $6.11 \times$, and reduces the I/O traffic to flash chips by $\mathbf{2 3. 0 8} \times$ on average. SkyByte also reaches $\mathbf{7 5 \%}$ of the performance of the ideal case that assumes unlimited DRAM capacity in the host, which offers an attractive cost-effective solution. Yuqi Xue, Yirui Eric Zhou, Shaobo Li 0005, Jian Huang 0006 |
HPCA | 3 |
| 2025 | Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect StorageabstractWe present the design and implementation of a new lifetime-aware tensor offloading
framework for GPU memory expansion using low-cost PCIe-based solid-state
drives (SSDs). Our framework, TERAIO, is developed explicitly for large language
model (LLM) training with multiple GPUs and multiple SSDs. Its design is driven
by our observation that the active tensors take only a small fraction (1.7% on
average) of allocated GPU memory in each LLM training iteration, the inactive
tensors are usually large and will not be used for a long period of time, creating
ample opportunities for offloading/prefetching tensors to/from slow SSDs without
stalling the GPU training process. TERAIO accurately estimates the lifetime (active
period of time in GPU memory) of each tensor with the profiling of the first few
iterations in the training process. With the tensor lifetime analysis, TERAIO will
generate an optimized tensor offloading/prefetching plan and integrate it into the
compiled LLM program via PyTorch. TERAIO has a runtime tensor migration
engine to execute the offloading/prefetching plan via GPUDirect storage, which
allows direct tensor migration between GPUs and SSDs for alleviating the CPU
bottleneck and maximizing the SSD bandwidth utilization. In comparison with
state-of-the-art studies such as ZeRO-Offload and ZeRO-Infinity, we show that
TERAIO improves the training performance of various LLMs by 1.47× on average,
and achieves 80.7% of the ideal performance assuming unlimited GPU memory. Yirui Eric Zhou, Apoorve Mohan, I-Hsin Chung, Seetharami Seelam, Jian Huang 0006 |
NeurIPS | 3 |
| 2025 | Managing Scalable Direct Storage Accesses for GPUs with GoFSabstractAs we shift from CPU-centric computing to GPU-accelerated computing for supporting intelligent data processing at scale, the storage bottleneck has been exacerbated. To bypass the host CPUand alleviate unnecessary data movements, modern GPUs enable direct storage access to SSDs (i.e., GPUDirect Storage). However, current GPUDirect Storage solutions still rely on the host file system to manage the storage device, direct storage accesses are still bottlenecked by the host. Shaobo Li 0005, Yirui Eric Zhou, Yuqi Xue, Yuan Xu 0020, Jian Huang 0006 |
SOSP | 2 |
| 2023 | G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor MigrationsabstractTo break the GPU memory wall for scaling deep learning workloads, a variety of architecture and system techniques have been proposed recently. Their typical approaches include memory extension with flash memory and direct storage access. However, these techniques still suffer from suboptimal performance and introduce complexity to the GPU memory management, making them hard to meet the scalability requirement of deep learning workloads today. Yirui Eric Zhou, Yuqi Xue, Jian Huang 0006 |
MICRO | 2 |