VLDB 2026 Research / reviewers in the wild / expert
Yuanyuan Dong 0002
dblp:92/5070-2
· DBLP profile ↗
13ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-2744-8652ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MDS array codes with low disk I/O and small repair bandwidth
Lei Li 0050, Chenhao Ying 0001, Yuanyuan Dong 0002, Jie Li 0002, Yuan Luo 0003 |
Frontiers Comput. Sci. | 4 |
| 2024 | Construction of Binary Cooperative MSR Codes with Multiple Repair Degrees
Lei Li 0050, Xinchun Yu, Yaqian Zhang 0002, Yuanyuan Dong 0002, Chenhao Ying 0001, Yuan Luo 0003 |
COCOON (2) | 5 |
| 2024 | Constructions of Binary MDS Array Codes with Optimal Cooperative Repair BandwidthabstractErasure codes are widely implemented in distributed storage systems to provide high fault tolerance with small storage overhead. Maximum distance separable codes are an common choice as they achieve the optimal tradeoff between fault tolerance and storage overhead. In this paper, we focus on the repair of multiple erasures of binary MDS array codes. Specifically, we present constructions of binary MDS array codes with optimal cooperative repair bandwidth by stacking multiple Blaum-Roth code instances whose “evaluation points” are judiciously designed. The constructed array codes with length$n$and dimension$k$can achieve the optimal cooperative repair bandwidth for$2\leq h\leq n-k$and$k+1\leq d\leq n-h$where$h$and$d$are the numbers of failed nodes and helper nodes, respectively. As the codes are constructed on a special polynomial ring over binary field, computation operations involved in nodes repair and file reconstruction for these codes are only XORs and cyclic shifts. Moreover, due to the inherent parallel structure of the codes, both the encoding and decoding procedures can be finished in parallel, speeding up the computing process. Lei Li 0050, Xinchun Yu, Chenhao Ying 0001, Yuanyuan Dong 0002, Yuan Luo 0003 |
ISIT | 5 |
| 2024 | MDS array codes with efficient repair and small sub-packetization level
Lei Li 0050, Xinchun Yu, Chenhao Ying 0001, Yuanyuan Dong 0002, Yuan Luo 0003 |
Des. Codes Cryptogr. | 5 |
| 2024 | Constructions of Binary MDS Array Codes With Optimal Repair/Access BandwidthabstractMaximum distance separable (MDS) codes are commonly deployed in distributed storage systems as they provide the maximum failure tolerance for some given redundancy. The repair problem of MDS codes has drawn much attention and various constructions of MDS array codes with optimal repair bandwidth have been proposed in the last decade. However, few of the existing codes are constructed over the binary field. In this paper, we propose new constructions of binary MDS array codes with optimal repair (or access) bandwidth for single-node failure. Specifically, by stacking multiple Blaum-Roth code instances of which the parity-check matrices are judiciously designed, we obtain three families of binary MDS array codes with optimal repair bandwidth; using the permutation matrices as building blocks, we also construct two families of binary MDS array codes with optimal access bandwidth. Moreover, error-resilient capability while achieving the lower bound on repair (or access) bandwidth is obtained when the number of helper nodes d < n - 1. All the codes in this paper are constructed over a particular ring of binary polynomials. Consequently, computation operations involved in the encoding, decoding and node repair procedures for these codes are only XORs and cyclic shifts, avoiding complex multiplications and divisions over large finite fields. Lei Li 0050, Xinchun Yu, Yuanyuan Dong 0002, Yuan Luo 0003 |
IEEE Trans. Commun. | 4 |
| 2023 | New Constructions of Binary MDS Array Codes with Optimal Repair BandwidthabstractMaximum distance separable codes are commonly used in large-scale distributed storage systems since they achieve the maximum fault tolerance for some given redundancy. In this paper, we focus on the repair problem of MDS codes. By stacking multiple Blaum-Roth code instances of which the parity-check matrices are judiciously designed, we construct three families of binary MDS array codes with optimal repair bandwidth for single-node failure. Specifically, codes in the first family achieve the lower bound on repair bandwidth where the number of helper nodes d is n − 1; The second family of codes are optimal-repair for any fixed d and the third are for multiple values of d simultaneously. Moreover, the last two also possess error-resilient capability and can achieve the corresponding optimal repair bandwidth. All the codes in this paper are constructed on a particular polynomial ring over binary field. Consequently, computation operations involved in node repair and file reconstruction for these codes are only XORs and cyclic shifts, avoiding complex multiplications and divisions over large finite fields. Lei Li 0050, Chenhao Ying 0001, Yuanyuan Dong 0002, Yuan Luo 0003 |
ISIT | 4 |
| 2022 | RCS: A Redirection Computational Scheduler to Accelerate Straggler Recovery for Erasure Coded Cloud Storage SystemabstractThe straggler problem is one of the most significant problems in cloud computing systems, in which a large number of parallel processes are blocked by a small set of straggler tasks with a long waiting time. This problem is crucial in erasure coded storage systems, where the recovery processes require to retrieve a set of multiple chunks among different nodes. With skewed data accesses from various applications, several nodes with a high workload could easily become stragglers during the recovery process, leading to unacceptable long tail latency. To address the above problems, we propose a Redirection Computational Scheduling method called RCS, to accelerate the data recovery under straggler scenarios. The key idea of RCS is transferring the computational and network workload from one node to another, which can avoid the adverse effects caused by the stragglers. To demonstrate the effectiveness of RCS, we conduct several experiments in a cluster. The results show that, compared to the state-of-the-art recovery methods, RCS saves the recovery time by up to 72.1%, and speeds up the recovery throughput by up to a factor of 1.4X, respectively. Xinzhe Cao, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo, Yuanyuan Dong 0002 |
ICCD | 6 |
| 2022 | XHR-Code: An Efficient Wide Stripe Erasure Code to Reduce Cross-Rack Overhead in Cloud Storage SystemsabstractNowadays wide stripe erasure codes (ECs) become popular as they can achieve low monetary cost and provide high reliability for cold data. Generally, wide stripe erasure codes can be generated by extending traditional erasure codes with a large stripe size, or designing new codes. However, although wide stripe erasure codes can decrease the storage cost significantly, the construction of lost data is extraordinary slow, which stems primarily from high cross-rack overhead. It is because a large number of racks participate in the construction of the lost data, which results in high cross-rack traffic. To address the above problems, we propose a novel erasure code called XOR-Hitchhiker-RS (XHR) code, to decrease the cross-rack overhead and still maintain low storage cost. The key idea of XHR is that it utilizes a triple dimensional framework to place more chunks within racks and reduce global repair triggers. To demonstrate the effectiveness of XHR-Code, we provide mathematical analysis and conduct comprehensive experiments. The results show that, compared to the state-of-the-art solutions such as ECWide under various failure conditions, XHR can effectively reduce cross-rack repair traffic and the repair time by up to 36.50%. Guofeng Yang, Huangzhen Xue, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo, Yuanyuan Dong 0002 |
SRDS | 9 |
| 2021 | EC-Scheduler: A Load-Balanced Scheduler to Accelerate the Straggler Recovery for Erasure Coded Storage SystemsabstractErasure codes (EC) have become a typical technology for distributed storage systems in place of data replication, providing similar data availability but lower storage cost. However, a great number of data computations and migrations during the EC recovery process bring high I/O and network latency penalties. Although several EC recovery methods have been designed to compromise the recovery penalty with high parallelism, the performance of these schemes was usually bounded by the straggler problems due to the various (I/O) performance among different nodes in the storage system. Moreover, the variation of the access popularity from the upper layer application causes the dynamic load fluctuation and asymmetry upon different nodes, which makes the scheduling more difficult during the recovery. To address the above problem, we propose a dynamic load-balanced scheduling algorithm for straggler recovery called EC-Scheduler. EC-Scheduler adjusts the recovery schedule dynamically with the awareness of continuous load fluctuation on the nodes, guaranteeing high parallelism and load balance ability simultaneously. To demonstrate the effectiveness of EC-Scheduler, we conduct several experiments in a cluster. The results show that, compared to typical recovery schemes such as Fast-PR and EC-Store, EC-Scheduler could achieve a 1.3X speed-up in the recovery process and 10X improvement in recovery load imbalance factor. Xinzhe Cao, Yunfei Gu, Chentao Wu, Jie Li 0002, Guangtao Xue, Minyi Guo, Yuanyuan Dong 0002 |
IWQoS | 8 |
| 2020 | EC-Fusion: An Efficient Hybrid Erasure Coding Framework to Improve Both Application and Recovery Performance in Cloud Storage SystemsabstractNowadays erasure coding is one of the most significant techniques in cloud storage systems, which provides both quick parallel I/O processing and high capabilities of fault tolerance on massive data accesses. In these systems, triple disk failure tolerant arrays (3DFTs) is a typical configuration, which is supported by several classic erasure codes like Reed-Solomon (RS) codes, Local Reconstruction Codes (LRC), Minimum Storage Regeneration (MSR) codes, etc. For an online recovery process, the foreground application workloads and the background recovery workloads are handled simultaneously, which requires a comprehensive understanding on both two types of workload characteristics. Although several techniques have been proposed to accelerate the I/O requests of online recovery processes, they are typically unilateral due to the fact that the above two workloads are not combined together to achieve high cost-effective performance.To address this problem, we propose Erasure Codes Fusion (EC-Fusion), an efficient hybrid erasure coding framework in cloud storage systems. EC-Fusion is a combination of RS and MSR codes, which dynamically selects the appropriate code based on its properties. On one hand, for write-intensive application workloads or low risk on data loss in recovery workloads, EC-Fusion uses RS code to decrease the computational overhead and storage cost concurrently. On the other hand, for read-intensive or frequent reconstruction in workloads, MSR code is a proper choice. Therefore, a better overall application and recovery performance can be achieved in a cost-effective fashion. To demonstrate the effectiveness of EC-Fusion, several experiments are conducted in hadoop systems. The results show that, compared with the traditional hybrid erasure coding techniques, EC-Fusion accelerates the response time for application by up to 1.77×, and reduces the reconstruction time by up to 69.10%. Han Qiu 0003, Chentao Wu, Jie Li 0002, Minyi Guo, Tong Liu 0030, Xubin He, Yuanyuan Dong 0002 |
IPDPS | 7 |
| 2020 | AZ-Recovery: An Efficient Crossing-AZ Recovery Scheme for Erasure Coded Cloud Storage SystemsabstractAs massive data in modern cloud storage systems grow dramatically, it is a common method to partition and store data in multiple Availability Zones (AZs). Multiple AZs not only provide high reliability, but also reduce the network latency. Erasure Codes (ECs) are widely used in multiple AZs to provide high reliability at low storage cost. However, the recovery cost of EC is extremely high in multiple AZs' environment, which is mainly because a normal EC needs to reconstruct the lost data via transferring the data/parities across AZs. Although existing fast recovery approaches can save the I/O cost or network bandwidth in an effective manner, they are not suitable for multiple AZs. The reasons include low flexibility on various complex network scenarios, less consideration on crossing-AZ bandwidth, low capabilities on multiple disk/node failures, etc. To address the above problem, in this paper, we propose a crossing $\underline{\mathrm{A}}$vailability Zone Recovery (AZ-Recovery) method to efficiently improve the recovery performance for multiple AZs. AZ-Recovery investigates the complex homogeneous/heterogeneous network topologies, and finds an optimal data transmission path. Using this method, AZ-Recovery can significantly reduce the recovery cost and save the crossing AZ bandwidth in various failure scenarios. To demonstrate the effectiveness of AZ-Recovery, we evaluate various erasure codes via mathematical analysis and simulations in Network Simulator-3. The results show that, compared to the traditional erasure coding methods, AZ-Recovery saves the recovery bandwidth by up to 77.47%. Chentao Wu, Zongxin Ye, Xubin He, Jie Li 0002, Minyi Guo, Guangtao Xue, Yuanyuan Dong 0002 |
SRDS | 9 |
| 2019 | Optimizing the Parity Check Matrix for Efficient Decoding of RS-Based Cloud Storage SystemsabstractIn large scale distributed systems such as cloud storage systems, erasure coding is a fundamental technique to provide high reliability at low monetary cost. Compared with the traditional disk arrays, cloud storage systems use an erasure coding scheme with both flexible fault tolerance and high scalability. Thus, Reed-Solomon (RS) Codes or RS-based codes are popular choices for cloud storage systems. However, the decoding performance for RS-based codes is not as good as XOR-based codes, which are optimized via investigating the relationships among different parity chains or reducing the computational complexity of matrix multiplications. Therefore, exploring an efficient decoding method is highly desired. To address the above problem, in this paper, we propose an Advanced Parity-Check Matrix (APCM) based approach, which is extended from the original Parity-Check Matrix based (PCM) approach. Instead of improving the decoding performance of XOR-based codes in PCM, APCM focuses on optimizing the decoding efficiency for RS-based codes. Furthermore, APCM avoids the matrix inversion computations and reduces the computational complexity of the decoding process. To demonstrate the effectiveness of the APCM, we conduct intensive experiments by using both RS-based and XOR-based codes under cloud storage environment. The results show that, compared to typical decoding methods, APCM improves the decoding speed by up to 32.31% in the Alibaba cloud storage system. Junqing Gu, Chentao Wu, Han Qiu 0003, Jie Li 0002, Minyi Guo, Xubin He, Yuanyuan Dong 0002 |
IPDPS | 8 |
| 2019 | AZ-Code: An Efficient Availability Zone Level Erasure Code to Provide High Fault Tolerance in Cloud Storage SystemsabstractAs data in modern cloud storage system grows dramatically, it's a common method to partition data and store them in different Availability Zones (AZs). Multiple AZs not only provide high fault tolerance (e.g., rack level tolerance or disaster tolerance), but also reduce the network latency. Replication and Erasure Codes (EC) are typical data redundancy methods to provide high reliability for storage systems. Compared with the replication approach, erasure codes can achieve much lower monetary cost with the same fault-tolerance capability. However, the recovery cost of EC is extremely high in multiple AZ environment, especially because of its high bandwidth consumption in data centers. LRC is a widely used EC to reduce the recovery cost, but the storage efficiency is sacrificed. MSR code is designed to decrease the recovery cost with high storage efficiency, but its computation is too complex. To address this problem, in this paper, we propose an erasure code for multiple availability zones (called AZ-Code), which is a hybrid code by taking advantages of both MSR code and LRC codes. AZ-Code utilizes a specific MSR code as the local parity layout, and a typical RS code is used to generate the global parities. In this way, AZ-Code can keep low recovery cost with high reliability. To demonstrate the effectiveness of AZ-Code, we evaluate various erasure codes via mathematical analysis and experiments in Hadoop systems. The results show that, compared to the traditional erasure coding methods, AZ-Code saves the recovery bandwidth by up to 78.24%. Chentao Wu, Junqing Gu, Han Qiu 0003, Jie Li 0002, Minyi Guo, Xubin He, Yuanyuan Dong 0002 |
MSST | 8 |