Chunxue Zuo

dblp:234/1598 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0002-2374-5772ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 3 since 2021
YearPublicationVenuePosition
2026 TC-Cache: Accelerating Restore Performance With Type-Aware Cooperative Cache in Erasure-Coded Deduplicated Storage Systems
abstract
To ensure reliability, modern deduplication-based storage systems exploit erasure coding to redundantly distribute objects that store post-deduplicated data across multiple storage nodes. In the case of node failures, the failure recovery of impacted objects performs degraded read operations by reading from these storage nodes. However, this degraded read scheme introduces the additional available objects into the restore cache, which poses challenges to state-of-the-art caching policies applied in erasure-coded deduplicated storage systems. On one hand, chunk-based caching method fails to cache the remaining available objects after the unavailable objects are rebuilt, resulting in multiple reads of accessed objects and severely degrading restore performance. On the other hand, object-based caching methods do not identify lost objects to which duplicate chunks belong, causing unused objects to unnecessarily occupy cache space and destroy cache locality. In this paper, we propose TC-Cache, a restore cache scheme for node failures, which combines object-based caching and chunk-based caching methods to improve restore performance. The main idea of TC-Cache is to identify objects to which redundant data belongs, and selectively handle degraded reads of two types of objects in caches of different granularities during failure recovery to enhance cache utilization and maintain cache locality. Extensive experiments based on three real-world datasets demonstrate that TC-Cache outperforms state-of-the-art cache algorithms in terms of restore performance by 21%-65% with low computing overhead.
Chunxue Zuo, Fang Wang 0001, Li Liu 0047
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 QuickScale: A Quick Scaling Scheme for Erasure-Coded Deduplicated Storage Systems
abstract
Modern deduplication-based storage systems employ erasure coding to post-deduplicate data to ensure reliability. In the erasure-coded deduplicated storage systems, existing coding schemes mainly perform erasure coding over a single object (intra-object coding) to address the problems of degraded read performance and storage inefficiency. However, due to intra-object coding, traditional storage scaling schemes need to reorganize and relocate all stored data between the client and storage nodes to support storage scaling, which inevitably incurs substantial data transfer overhead, thereby degrading system scalability. In this paper, we propose a fast scaling scheme called QuickScale for erasure-coded deduplicated storage systems. QuickScale first implements the intra-object coding in a realistic distributed storage environment and verifies that this coding scheme can significantly alleviate the degraded read performance by 35%-67% and save storage space by 24% 30%. Motivated by these results, QuickScale uses a node-to-node scaling approach that avoids repeatedly transferring all stored data between the client and storage nodes during scaling, thereby improving the scalability. Experimental evaluations using six real-world datasets demonstrate that QuickScale achieves better scaling performance (in terms of scaling time during scaling) over the state-of-the-art scaling schemes by up to 63.5%-76.4%.
Chunxue Zuo, Lanlan Cui, Fang Wang 0001
IEEE Trans. Cloud Comput.1
2022 Ensuring high reliability and performance with low space overhead for deduplicated and delta-compressed storage systems
abstract
Abstract Data deduplication is a widely used technique to remove duplicate data to reduce the storage overhead. However, deduplication typically cannot eliminate the redundancy among nonidentical but similar data chunks. To reduce the storage overhead further, delta compression is often applied to compress the post‐deduplication data. While the two techniques are effective in saving storage space, they introduce complex references among data chunks, which inevitably undermines the system reliability and introduces fragmentation that may degrade the restore performance. In this paper, we observe that the delta compressed chunks (DCCs) are much smaller than regular chunks (non‐DCCs). Also, most fragmentation caused by the base chunk of DCCs remain fragmented in consecutive backups. Based on these observations, we introduce a framework called , which combines replication and erasure coding and uses History‐aware Delta Selection to ensure high reliability and restore performance. Specifically, uses a delta‐utilization‐aware filter and a cooperative cache scheme (CCS) to maintain cache locality and avoid unnecessary container reads, respectively. Moreover, the system selectively performs delta compression by historical information to avoid cyclic fragmentation in consecutive backups. Experimental results based on four real‐world datasets demonstrate that significantly improves the restore performance by 58.3%–76.7% with a low storage overhead.
Chunxue Zuo, Fang Wang 0001, Mai Zheng, Yuchong Hu, Dan Feng 0001
Concurr. Comput. Pract. Exp.1
2019 RepEC-Duet: Ensure High Reliability and Performance for Deduplicated and Delta-Compressed Storage Systems
abstract
Data deduplication is a widely deployed technique to remove duplicate content to save storage space, which is however incapable of eliminating the redundancy between nonidentical but similar data blocks. To achieve further space savings in deduplicated storage systems, delta compression is employed to compress post-deduplication data. Both deduplication and delta compression introduce content references among blocks, which inevitably undermines the reliability of deduplicated and delta compressed storage systems. To ensure better reliability, existing approaches utilize either replication or erasure codes to redundantly distribute data across multiple nodes. In deduplicated and delta compressed storage systems, we observe that delta compressed chunks (DCCs) are far smaller than regular chunks called non-DCCs. Motivated by this observation, we suggest a straightforward approach in which replication is used to protect DCCs and erasure code is deployed to protect non-DCCs. However, we need to address two critical challenges to ensure this solution effective. First, the random placement of DCCs replicas destroys cache locality. Second, the separate and individual recovery and restore cache could cause storage containers to be accessed repeatedly. To address these two challenges, in this paper, we propose RepEC-Duet which employs both replication and erasure codes to ensure high reliability and performance for deduplicated and delta-compressed storage systems. RepEC-Duet introduces a delta-utilization-aware filter to select and replicate containers based on the percentage of DCCs in the containers to maintain cache locality. Moreover, to avoid unnecessary container reads, we design a cooperative cache scheme that is aware of both failure recovery and regular restore cache. Our experimental results based on three real-world datasets demonstrate that RepEC-Duet significantly improves the restore performance by 26%-59%, and reduces the storage overhead by 54%-98% than the existing approaches.
Chunxue Zuo, Fang Wang 0001, Yuchong Hu, Dan Feng 0001
ICCD1
2018 PFCG: Improving the Restore Performance of Package Datasets in Deduplication Systems
abstract
Data deduplication, a lossless data compression technique, has been widely deployed in backup systems to save storage space. However, data fragmentation introduced by data deduplication seriously degrades restore performance. To alleviate the fragmentation problem, rewriting algorithms such as CBR, CAP, and HAR have been proposed. In certain backup datasets containing packed files like UNIX tar files, we have observed that a large amount of rewritten fragmented chunks tend to remain fragmented in following backups and thus get rewritten repeatedly in consecutive backups. Such repeatedly rewritten chunks are referred to as persistent fragmented chunks (PFCs), and they severely impact the efficiency of rewriting algorithms and restore performance. In this paper, we propose PFCG, an efficient scheme to enhance the efficiency of rewriting algorithms and restore performance. The central idea of PFCG is to identify and differentiate PFCs from regular fragmented chunks, i.e., non-PFCs, and store these two types of fragmented chunks into separate containers during backups to prevent the proliferation of PFCs in following backups. Our experimental results based on six real-world datasets demonstrate that PFCG significantly improves the restore performance by 21% to 47% over the state-of-the-art rewriting algorithms, without sacrificing deduplication efficiency.
Chunxue Zuo, Fang Wang 0001, Yuchong Hu, Dan Feng 0001
ICCD1