EDBT 2026 Demo / reviewers in the wild / expert
Yuhui Cai
dblp:337/4184
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0007-0760-9707ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ChameleonEC: Exploiting Tunability of Erasure Coding for Low-Interference RepairabstractErasure coding provides fault tolerance in a storage-efficient manner, yet it introduces a high repair penalty. We uncover via trace-driven experiments that the substantial repair traffic in erasure coding is prone to entangling with the foreground traffic, thereby slowing down repair progress and downgrading service quality. We present ChameleonEC, a general mechanism that can assist a variety of erasure codes in realizing low-interference repair. ChameleonEC comprises the following design techniques: (i) repair task assignment, which decomposes a repair plan into multiple repair tasks and makes them coexist harmoniously with the foreground traffic, so as to saturate unoccupied bandwidth and avoid bandwidth contentions; (ii) repair path establishment, which orchestrates elastic transmission routings over the dispatched repair tasks to instruct the repair; and (iii) straggler-aware re-scheduling, which timely re-tunes task transmissions and repair plans to bypass unexpected stragglers emerging in repair. We conduct extensive experiments on Amazon EC2, showing that ChameleonEC can accelerate the repair by 4.9–498.2% for various erasure codes under different real-world traces. ChameleonEC can also speed up the repair process by 25.4–73.5% under the storage-bottlenecked scenarios. Yuhui Cai, Shiyao Lin, Zhirong Shen, Jiwu Shu |
HPCA | 1 |
| 2025 | TPRepair: Tree-based Pipelined Repair in Clustered Storage SystemsabstractErasure coding is an effective technique for guaranteeing data reliability for storage systems, yet it incurs a high repair penalty with amplified repair traffic. The repair becomes more intricate in clustered storage systems with the bandwidth diversity property. We present TPRepair , a T ree-based P ipelined Repair approach, aiming to expedite the overall repair process with the tailored pipelined repair procedure. TPRepair first prioritizes selecting racks with the current minimum load to participate in the repair process. It subsequently formulates tree-based links, tailored to align seamlessly with the pipelined repair procedure. TPRepair further designs an optimization algorithm to reduce the bottleneck load when repairing multiple chunks. Large-scale simulations demonstrate that TPRepair can increase 13.8%–41.3% of the balance ratio without amplifying cross-rack traffic. Meanwhile, Alibaba Cloud ECS experiments indicate that TPRepair can increase repair throughput by 11.3% to 72.9%. Fulin Nan, Zhirong Shen, Zhisheng Chen 0002, Yuhui Cai, Dmitry I. Kaplun, Xiaoli Wang 0002, Quanqing Xu, Chuanhui Yang, Jiwu Shu |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | ElasticEC: Achieving Fast and Elastic Redundancy Transitioning in Erasure-Coded ClustersabstractErasure coding has been extensively deployed in today’s commodity HPC systems against unexpected failures. To adapt to the varying access characteristics and reliability demands, storage clusters have to perform redundancy transitioning via tuning the coding parameters, which unfortunately gives rise to substantial transitioning traffic. We present ElasticEC, a fast and elastic redundancy transitioning approach for erasure-coded clusters. ElasticEC first minimizes the transitioning traffic via proposing a relocation-aware stripe reorganization mechanism and a collecting-and-encoding algorithm. It further heuristically balances the transitioning traffic across nodes. We implement ElasticEC in Hadoop HDFS and conduct extensive experiments on a real-world cloud storage cluster, showing that ElasticEC can reduce 71.1-92.6% of the transitioning traffic and shorten 65.9-90.7% of the transitioning time. Yuhui Cai, Guowen Gong, Zhirong Shen, Jiwu Shu |
IEEE Trans. Computers | 1 |
| 2025 | A Survey of the Past, Present, and Future of Erasure Coding for Storage SystemsabstractErasure coding is a known redundancy technique that has been popularly deployed in modern storage systems to protect against failures. By introducing a small portion of coded redundancy into data storage, erasure coding is shown to provide higher reliability guarantees than replication under the same storage overhead. Despite its storage efficiency, erasure coding incurs high-performance overhead in repair and updates, and its reliability also depends on the amount of redundancy. How to resolve the tensions among storage efficiency, performance, and reliability has been the major research direction in the literature for decades. In this article, we present an in-depth survey of the past, present, and future of erasure coding in storage systems. We conduct our survey from a systems perspective, with an emphasis on how erasure coding is deployed in practical storage systems. Specifically, we first review the use of erasure coding in storage systems from both academia and industry and state the challenges of deploying erasure coding in practice. We then review the topics of erasure coding in three aspects: (i) new erasure code constructions, (ii) algorithmic techniques for efficient erasure coding operations, and (iii) erasure coding for emerging architectures. Finally, we provide future research directions for erasure coding. Zhirong Shen, Yuhui Cai, Keyun Cheng, Patrick P. C. Lee, Xiaolu Li 0002, Yuchong Hu, Jiwu Shu |
ACM Trans. Storage | 2 |
| 2022 | TCAM-Resnet: A convolutional neural network for screening DR and AMD based on OCT imagesabstractDiabetic retinopathy (DR) and age-related macular degeneration (AMD) are important causes of blindness and visual loss. Optical coherence tomography (OCT) is a non-invasive optical imaging method that can capture retinal vascular information and even pathological information. In order to improve the screening rate and accuracy of these two diseases, we propose a new network structure named TCAM-Resnet, which uses OCT three-dimensional images to screen and classify AMD and DR. TCAM-Resnet is based on the Resnet network. A three-dimensional convolution attention module (TCAM) is added. The attention module can extract the weight features of blood vessels from 3D images and uses the residual-like structure when interacting with the Resnet network, which makes the original data retain more information during attention. Experimental results on the OCTA-500 dataset show that a three-dimensional convolution network is superior to a two-dimensional convolution network in lesion feature extraction. With the addition of the new module TCAM, Resnet3D has achieved higher accuracy in disease classification t asks, with the accuracy of AMD, DR, and NORMAL reaching 83.3%, and the accuracy of AMD and NORMAL reaching 98%. Deshui Yu, Yuhui Cai, Bojun Li, Wei Li 0117 |
BIBM | 3 |