VLDB 2026 Research / reviewers in the wild / expert
Keyun Cheng
dblp:245/2761
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-8301-9633ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LESS is More for I/O-Efficient Repairs in Erasure-Coded Storage
Keyun Cheng, Xiaolu Li 0002, Sihuang Hu, Patrick P. C. Lee |
FAST | 1 |
| 2025 | A Survey of the Past, Present, and Future of Erasure Coding for Storage SystemsabstractErasure coding is a known redundancy technique that has been popularly deployed in modern storage systems to protect against failures. By introducing a small portion of coded redundancy into data storage, erasure coding is shown to provide higher reliability guarantees than replication under the same storage overhead. Despite its storage efficiency, erasure coding incurs high-performance overhead in repair and updates, and its reliability also depends on the amount of redundancy. How to resolve the tensions among storage efficiency, performance, and reliability has been the major research direction in the literature for decades. In this article, we present an in-depth survey of the past, present, and future of erasure coding in storage systems. We conduct our survey from a systems perspective, with an emphasis on how erasure coding is deployed in practical storage systems. Specifically, we first review the use of erasure coding in storage systems from both academia and industry and state the challenges of deploying erasure coding in practice. We then review the topics of erasure coding in three aspects: (i) new erasure code constructions, (ii) algorithmic techniques for efficient erasure coding operations, and (iii) erasure coding for emerging architectures. Finally, we provide future research directions for erasure coding. Zhirong Shen, Yuhui Cai, Keyun Cheng, Patrick P. C. Lee, Xiaolu Li 0002, Yuchong Hu, Jiwu Shu |
ACM Trans. Storage | 3 |
| 2025 | Toward Load-Balanced Redundancy Transitioning for Erasure-Coded StorageabstractRedundancy transitioning enables erasure-coded storage to adapt to varying performance and reliability requirements by re-encoding data with new coding parameters on-the-fly. Existing studies focus on bandwidth-driven redundancy transitioning that reduces the transitioning bandwidth across storage nodes, yet the actual redundancy transitioning performance remains bottlenecked by the most loaded node. We present BART, a load-balanced redundancy transitioning scheme that aims to reduce the redundancy transitioning time via carefully scheduled parallelization. We show that finding an optimal load-balanced solution is difficult due to the large solution space. Given this challenge, BART decomposes the redundancy transitioning problem into multiple sub-problems and solves the sub-problems via efficient heuristics. We evaluate BART using both simulations for large-scale storage and HDFS prototype experiments on Alibaba Cloud. We show that BART significantly reduces the redundancy transitioning time compared with the bandwidth-driven approach. Keyun Cheng, Huancheng Puyang, Xiaolu Li 0002, Patrick P. C. Lee, Yuchong Hu, Jie Li 0019, Ting-Yi Wu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Harmonizing Repair and Maintenance in LRC-Coded StorageabstractModern storage systems not only introduce data redundancy for fault tolerance, but also conduct regular main- tenance operations on storage nodes for system robustness. Erasure coding provides storage-efficient redundancy and has been widely deployed in production, yet it also incurs substantial bandwidth and I/O overhead due to the repair of storage failures. In particular, maintenance operations make storage nodes temporarily unavailable and lead to data unavailability, thereby incurring repair overhead for erasure-coded storage. In this paper, we study Locally Repairable Codes (LRCs), a class of practical repair-efficient erasure codes, and show that there exists an inherent performance trade-off between the repair and maintenance operations of LRCs in data center settings, such that the repair performance in regular (i.e., no-maintenance) and maintenance modes cannot be simultaneously optimized. To this end, we design a configurable data placement scheme that operates along the trade-off subject to fault-tolerance constraints. We prototype our data placement scheme atop Hadoop HDFS and show how it balances the performance trade-off of repair and maintenance operations in real network environments. Keyun Cheng, Si Wu 0003, Xiaolu Li 0002, Patrick P. C. Lee |
SRDS | 1 |
| 2023 | ParaRC: Embracing Sub-Packetization for Repair Parallelization in MSR-Coded Storage
Xiaolu Li 0002, Keyun Cheng, Kaichen Tang, Patrick P. C. Lee, Yuchong Hu, Dan Feng 0001, Jie Li 0019, Ting-Yi Wu |
FAST | 2 |
| 2023 | Balancing Repair Bandwidth and Sub-Packetization in Erasure-Coded Storage via Elastic Transformation
Kaichen Tang, Keyun Cheng, Helen H. W. Chan, Xiaolu Li 0002, Patrick P. C. Lee, Yuchong Hu, Jie Li 0019, Ting-Yi Wu |
INFOCOM | 2 |
| 2022 | Fast Proactive Repair in Erasure-Coded Storage: Analysis, Design, and ImplementationabstractErasure coding offers a storage-efficient redundancy mechanism for maintaining data availability guarantees in large-scale storage clusters, yet it also incurs high performance overhead in failure repair. Recent developments in accurate disk failure prediction allow soon-to-fail (STF) nodes to be repaired in advance, thereby opening new opportunities for accelerating failure repair in erasure-coded storage. To this end, we present a fast proactive repair solution called${{\sf FastPR}}$, which carefully couples two repair methods, namely migration (i.e., relocating the chunks of an STF node) and reconstruction (i.e., decoding the chunks of an STF node through erasure coding), so as to fully parallelize the repair operation across the storage cluster.${{\sf FastPR}}$solves a bipartite maximum matching problem and schedules both migration and reconstruction in a parallel fashion. We show that${{\sf FastPR}}$significantly reduces the repair time over the baseline repair approaches for both Reed-Solomon codes and Azure's Local Reconstruction Codes via mathematical analysis, large-scale simulation, and Amazon EC2 experiments. Xiaolu Li 0002, Keyun Cheng, Zhirong Shen, Patrick P. C. Lee |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Incorporating Temporal Prior from Motion Flow for Instrument Segmentation in Minimally Invasive Surgery Video
Yueming Jin, Keyun Cheng, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 2 |