Chunfeng Du

dblp:296/0805 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-0097-5003ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Gemina: A Coordinated and High-Performance Memory Deduplication Engine
abstract
Memory deduplication is widely used to effectively reduce the memory footprint in operating systems, while huge pages are employed to enhance access performance by increasing TLB hit rates. Unfortunately, duplicate huge pages are very rare in memory, leading to most huge pages being split into base pages during deduplication, which degrades access performance by up to $\mathbf{4 3 \%}$. Although asynchronous huge page promotion attempts to elevate multiple base pages to huge pages to maximize performance, deduplication disrupts the uniformity of page attributes, limiting the effectiveness of such promotions. To address this problem, we designed Gemina, a new memory deduplication scheme that coordinates huge page management. Gemina employs a fine-grained detector to distinguish pages, redundancy and hotness, a distributor to process pages conservatively and avoid frequent switches, and an adapter to organize memory for fast conversion. This approach balances the trade-off between performance and space efficiency in memory management. Experimental evaluations show that, compared to the naive KSM-based deduplication scheme, Gemina can achieve $96.1 \%$ of the memory savings of KSM while also improving memory access performance by $\mathbf{4 3. 1 \%}$ in random access.
Zhehua Zhang, Suzhen Wu, Wenyan You, Chunfeng Du, Bo Mao 0003
HPCA4
2025 BL-Tree: The Best of Both Worlds by Combining B+- Tree on Top and LSM - Tree on Bottom
abstract
The shattered and overlapped Level-0 data organization is the primary cause of write stall and read amplification problems in LSM-Tree-based Key-Value (KV) stores: (1) Level-0 to Level-1 compaction involves a large amount of data which induces write stalls, and (2) A point lookup needs to access multiple files in Level-0 which leads to significant read amplification. To address the problem, we propose BL-Tree by replacing the shattered Level-0 in LSM-Tree with a B+-Tree in byte-addressable Persistent Memory (PM). The sorted B+-Tree of Level-0 can accelerate the point lookup speed and reduce read/write amplification. BL-Tree further conducts the locality-aware and parallel compaction from the B+-Tree in PM (Level-0) to the lower levels of LSM-Tree in SSDs by only moving cold data downward, thus alleviating the write stalls and reducing the read/write amplification simultaneously. The extensive experiments on the prototype of BL- Tree show that it definitely avoids the write stalls and significantly reduces the read/write amplification. As a result, BL-Tree reduces the P99 tail latency by 65.2 × than LevelDB-PM and speeds up the throughput by more than 2 × under workloads with spatial locality than other KV stores.
Suzhen Wu, Zuocheng Wang, Shengzhe Wang 0001, Jiahong Chen, Chunfeng Du, Ke Zhou 0001, Jie Zhang 0048, Bo Mao 0003
ICDE5
2024 FSDedup: Feature-Aware and Selective Deduplication for Improving Performance of Encrypted Non-Volatile Main Memory
abstract
Enhancing the endurance, performance, and energy efficiency of encrypted Non-Volatile Main Memory (NVMM) can be achieved by minimizing written data through inline deduplication. However, existing approaches applying inline deduplication to encrypted NVMM suffer from substantial performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). In the preliminary ESD [ 14 ], we proposed the Error Correcting Code (ECC) assisted selective deduplication scheme, utilizing the ECC information as a fingerprint to identify similar data effectively and then leveraging the selective deduplication technique to eliminate a large amount of redundant data with high reference counts. In this article, we proposed FSDedup. Compared with ESD, FSDedup could leverage the prefetch cache to reduce the read overhead during similarity comparison and utilize the cache refresh mechanism to identify further and eliminate more redundant data. Extensive experimental evaluations demonstrate that FSDedup can enhance the performance of the NVMM system further than the ESD. Experimental results show that FSDedup can improve both write and read speed by up to 1.8×, enhance Instructions Per Cycle by up to 1.5×, and reduce energy consumption by up to 2.0×, compared to ESD.
Chunfeng Du, Zihang Lin, Suzhen Wu, Yifei Chen 0011, Jiapeng Wu, Shengzhe Wang 0001, Weichun Wang 0002, Bo Mao 0003
ACM Trans. Storage1
2023 ESD: An ECC-assisted and Selective Deduplication for Encrypted Non-Volatile Main Memory
abstract
Reducing write data to encrypted Non-Volatile Main Memory (NVMM) can directly improve NVMM’s endurance, performance, and energy efficiency. However, existing works that straightforwardly apply inline deduplication on encrypted NVMM can significantly lead to system performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). This paper proposes ESD, an ECC-assisted and Selective Deduplication for encrypted NVMM by exploiting both the device characteristics (ECC mechanism) and the workload characteristics (content locality). First, ESD utilizes the ECC information associated with each cache line evicted from the Last-Level Cache (LLC) as the fingerprint to identify data similarity and avoids the costly hash calculating overhead on the non-duplicate cache lines. Second, ESD leverages selective deduplication to exploit the content locality within cache lines by only storing the fingerprints with high reference counts in the memory cache to reduce the memory space overhead and avoid fingerprints NVMM_lookup operations. The experimental results show that ESD can significantly speed up the writes by up to 3.4x, 4.3x, and 2.6x, speed up the reads by up to 5.3x, 5.0x, and 2.0x, and reduce the energy consumption by up to 96.3%, 96.2%, and 56.6% than Baseline, Dedup SHA1, and DeWrite, respectively. Meanwhile, ESD also can significantly outperform other schemes in tail latency.
Chunfeng Du, Suzhen Wu, Jiapeng Wu, Bo Mao 0003, Shengzhe Wang 0001
HPCA1
2023 LearnedSync: A Learning-Based Sync Optimization for Cloud Storage
Suzhen Wu, Shengzhe Wang 0001, Chunfeng Du, Jiayang Guo, Yijie Pan, Naian Xiao, Bo Mao 0003
ICA3PP (2)4
2023 EaD: ECC-Assisted Deduplication With High Performance and Low Memory Overhead for Ultra-Low Latency Flash Storage
abstract
Data deduplication has become a commodity feature in flash storage products to effectively reduce redundant write data and improve space efficiency. However, it also introduces computing and memory overhead to generate and store the cryptographic hash (fingerprint) in face of the moderate data redundancy in primary storage. With the advent of 3D XPoint and Z-NAND technologies, and the stronger cryptographic hash functions in use, such as SHA-256, both the computing and memory overheads are increasingly serious performance bottlenecks for inline data deduplication in these ultra-low latency flash storage. To address these problems, we propose an ECC-assisted Deduplication approach, called EaD, which exploits the ECC property and the asymmetric read-write performance characteristics of modern flash storage. EaD first identifies data similarity by leveraging the device-generated ECC values of data chunks as their fingerprints, significantly reducing the costly MD5/SHA-based cryptographic hash computing and alleviating the memory space overhead. Based on the identification results, similar data chunks and their ECCs are read from the flash to perform a byte-by-byte comparison in memory to definitively identify and remove redundant data chunks. Our experiments show that the EaD approach significantly increases I/O performance by up to 4.2${\times }$, with an average of 2.5${\times }$, compared with the existing MD5/SHA- and sampling-based deduplication approaches.
Suzhen Wu, Chunfeng Du, Weidong Zhu 0002, Jindong Zhou, Hong Jiang 0001, Bo Mao 0003, Lingfang Zeng
IEEE Trans. Computers2
2022 DedupHR: Exploiting Content Locality to Alleviate Read/Write Interference in Deduplication-Based Flash Storage
abstract
The read/write interference problem of flash storage, where on-going write requests prevent read requests from accessing available data and significantly increase read latency, remains a critical concern with the rapid evolutions of flash chips from SLC through MLC to TLC and on to QLC, with PLC on the horizon, especially under workloads with a mixture of read and write requests. To significantly improve the read performance in face of read/write interference, which is often more critical than the write performance, we propose an enhanced Hot Read Data Replication scheme for deduplication-based flash storage, called DedupHR, to alleviate the read/write interference problem. DedupHR exploits the content locality in the deduplication-based flash storage and outsources the data blocks with high reference count to a surrogate space such as a dedicated spare flash chip or an over-provisioned space in an SSD. By servicing some conflicted read requests on the surrogate flash space, DedupHR can significantly alleviate, if not entirely eliminate, the contention between the read requests and the on-going write requests. The evaluation results show that DedupHR significantly improves the state-of-the-art schemes in terms of the system performance and cost efficiency. Consequently, the tail-latency of the deduplication-based flash storage is also reduced.
Suzhen Wu, Chunfeng Du, Bo Mao 0003, Hong Jiang 0001
IEEE Trans. Computers2
2021 CAGC: A Content-aware Garbage Collection Scheme for Ultra-Low Latency Flash-based SSDs
abstract
With the advent of new flash-based memory technologies with ultra-low latency, directly applying inline data deduplication in flash-based storage devices can degrade the system performance since key deduplication operations lie on the shortened critical write path of such devices. To address the problem, we propose a Content-Aware Garbage Collection scheme (CAGC), which embeds the data deduplication into the data movement workflow of the Garbage Collection (GC) process in ultra-low latency flash-based SSDs. By parallelizing the operations of valid data pages migration, hash computing and flash block erase, the deduplication-induced performance overhead is alleviated and redundant page writes during the GC period are eliminated. To further reduce data writes and write amplification during GC, CAGC separates and stores data pages in different regions based on their reference counts. The performance evaluation of our CAGC prototype implemented in FlashSim shows that CAGC significantly reduces the number of flash blocks erased and data pages migrated during GC, leading to improved user I/O performance and reliability of ultra-low latency flash-based SSDs.
Suzhen Wu, Chunfeng Du, Hong Jiang 0001, Zhirong Shen, Bo Mao 0003
IPDPS2