EDBT 2026 Demo / reviewers in the wild / expert
Kaiye Zhou
dblp:319/2931
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-3959-756XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZTree: Towards an efficient B+-tree on zoned namespace SSDs
Kaiye Zhou, Shucheng Wang |
Future Gener. Comput. Syst. | 2 |
| 2026 | The Design of Trillion-scale SSD-based Indexing with Deterministic Latency for Cloud Block StorageabstractCloud block storage (CBS) provides virtual disks with block-level accessibility. The petabyte-scale CBS systems maintain trillions of block-mapping key-value entries as metadata to track the storage location of each virtual block. Although SSD-based KV stores have been widely adopted in cloud systems for their high efficiency and durability, current SSD-based schemes face significant challenges in achieving deterministic access latency for latency-sensitive metadata services. Our experimental observations indicate that the substantial long-tail latency is primarily caused by (1) I/O blocking due to internal tasks of SSDs including modern Zone Namespace SSDs; and (2) additional disk I/Os when querying high-level indexes across memory and SSDs under memory-constrained environments. In this article, we propose an SSD-based SIndex to store trillions of block-mapping entries for latency-critical cloud block storage, which performs comprehensive latency optimization across storage I/O scheduling and high-level indexing. To prevent long-tail I/Os while avoiding intrusive device modifications, SIndex introduces an inter-SSD I/O scheduling mechanism based on read/write separation and SSD state transitions, which mitigates latency fluctuations induced by garbage collection on conventional SSDs and zone operations on Zone Namespace SSDs. Additionally, SIndex employs opportunistic I/O speculation and a concurrent request balancing mechanism to reduce read disturbance and I/O contention. To query the storage location of targeted block-mapping entries with bounded latency, SIndex proposes a memory-efficient high-level index incorporates with a static data layout, preventing time-consuming disk lookups by keeping the index in memory. We evaluate the SIndex prototype using a variety of benchmarks and real-world traces on commodity SSDs. The results demonstrate that SIndex outperforms RocksDB and other approaches by up to 11.4× in tail latency, keeping the 99.99th-percentile latency below 400 μs. Shucheng Wang, Zhandong Guo, Kaiye Zhou, Jun Xu 0037, Qiang Cao 0001 |
ACM Trans. Storage | 3 |
| 2025 | LCache: Log-Structured SSD Caching for Training Deep Learning ModelsabstractTraining deep learning models is computationally demanding and data-intensive. Existing approaches utilize local SSDs within training servers to cache datasets, thereby accelerating data loading during model training. However, we experimentally observe that data loading remains a performance bottleneck when randomly retrieving small-sized sample files on SSDs. In this paper, we introduce LCache, a log-structured dataset caching mechanism designed to fully leverage the I/O capabilities of SSDs and reduce I/O-induced training stalls. LCache determines the randomized dataset access order by extracting the pseudo-random seed from the training frameworks. It then aggregates small-sized sample files into larger chunks and stores them in a log file on SSDs, thus enabling sequential I/O requests on data retrieval and improving data loading throughput. Further-more, LCache proposes a real-time log reordering mechanism that strategically schedules cached data to organize logs across different epochs, which enhances cache utilization and minimizes data retrieval from low-performance remote storage systems. Additionally, LCache incorporates an MetaIndex to enable rapid log traversal and querying. We evaluate LCache with various real-world DL models and datasets. LCache outperforms the native PyTorch Dataloader and NoPFS by up to 9.4x and 7.8x in throughput,. respectively, Shucheng Wang, Zhandong Guo, Jian Sheng, Kaiye Zhou, Qiang Cao 0001 |
DATE | 5 |
| 2024 | ParaCkpt: Heterogeneous Multi-Path Checkpointing Mechanism for Training Deep Learning ModelsabstractTraining large deep learning models is extremely computationally intensive and time-consuming; therefore, it relies on checkpointing mechanisms to save snapshots promptly, ensuring rapid recovery from a myriad of failures. Existing checkpointing approaches save snapshots to either CPU memory or storage, overlooking their aggregated I/O capability. In this paper, we propose a heterogeneous multi-path checkpointing mechanism, ParaCkpt, to make full use of both PCle-bandwidth and I/O capability of memory and storage to accelerate check-pointing. ParaCkpt first identifies multiple paths for GPUs to CPU memory, local and remote storages, and determines their available bandwidths. Then, ParaCkpt strategically partitions the training model states into a set of path-based shards and drains them from GPUs to the memory and storage in parallel. Moreover, ParaCkpt employs a two-stage persistence strategy to flush in-memory shards to local SSDs in the background, and then stores local shards using compression to remote storage. Finally, ParaCkpt maintains global snapshots distributed across memory and storage, enabling rapid recovery via the multi-path way. We evaluate ParaCkpt with various real-world deep learning models. ParaCkpt outperforms native Pytorch and state-of-the-art asynchronous checkpointing approaches by up to 96 x and 2.3 x in throughput, respectively. Shucheng Wang, Qiang Cao 0001, Kaiye Zhou, Jun Xu 0037, Zhandong Guo, Jiannan Guo 0001 |
ICCD | 3 |
| 2024 | SIndex: An SSD-based Large-scale Indexing with Deterministic Latency for Cloud Block StorageabstractThe Solid State Drives (SSD) based key-value stores face significant challenges in achieving deterministic access latency. We experimentally observe the long-tail latency is mainly caused by I/O blocking induced by SSD’s internal tasks. In this paper, we propose an SSD-based SIndex to store hundreds of billions of block-mapping entries for the latency-critical cloud block storage. To hide the latency fluctuations induced by garbage collection and buffer flushing, SIndex proposes an inter-SSD I/O scheduling based on read/write separation and SSD state transition, while adopting opportunistic request speculation and balancing mechanism to mitigate read disturbance and I/O contention. Moreover, SIndex introduces a write-staging buffer cache and a two-stage sync mechanism to preferentially buffer updated data before synchronizing them to SSDs. We evaluate the SIndex prototype with a variety of benchmarks and real-world traces on commodity SSDs. SIndex is demonstrated to outperform RocksDB and other approaches by up to 11.2 × in tail latency without affecting the throughput performance. Shucheng Wang, Kaiye Zhou, Zhandong Guo, Qiang Cao 0001, Jun Xu 0037, Jie Yao 0001 |
ICPP | 2 |