Shengzhe Wang 0001

dblp:275/0653-1 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0000-8269-8702ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BL-Tree: The Best of Both Worlds by Combining B+- Tree on Top and LSM - Tree on Bottom
abstract
The shattered and overlapped Level-0 data organization is the primary cause of write stall and read amplification problems in LSM-Tree-based Key-Value (KV) stores: (1) Level-0 to Level-1 compaction involves a large amount of data which induces write stalls, and (2) A point lookup needs to access multiple files in Level-0 which leads to significant read amplification. To address the problem, we propose BL-Tree by replacing the shattered Level-0 in LSM-Tree with a B+-Tree in byte-addressable Persistent Memory (PM). The sorted B+-Tree of Level-0 can accelerate the point lookup speed and reduce read/write amplification. BL-Tree further conducts the locality-aware and parallel compaction from the B+-Tree in PM (Level-0) to the lower levels of LSM-Tree in SSDs by only moving cold data downward, thus alleviating the write stalls and reducing the read/write amplification simultaneously. The extensive experiments on the prototype of BL- Tree show that it definitely avoids the write stalls and significantly reduces the read/write amplification. As a result, BL-Tree reduces the P99 tail latency by 65.2 × than LevelDB-PM and speeds up the throughput by more than 2 × under workloads with spatial locality than other KV stores.
Suzhen Wu, Zuocheng Wang, Shengzhe Wang 0001, Jiahong Chen, Chunfeng Du, Ke Zhou 0001, Jie Zhang 0048, Bo Mao 0003
ICDE3
2024 LearnedFTL: A Learning-Based Page-Level FTL for Reducing Double Reads in Flash-Based SSDs
abstract
We present LearnedFTL, a new on-demand pagelevel flash translation layer (FTL) design, which employs learned indexes to improve the address translation efficiency of flashbased SSDs. The first of its kind, it reduces the number of double reads induced by address translation in random read accesses. LearnedFTL proposes three key techniques: an in-place-update linear model to build learned indexes efficiently, a virtual PPN representation to obtain contiguous PPNs for sorted LPNs, and a group-based allocation and model training via GC/rewrite strategy to reduce the training overhead. By tightly integrating the aforementioned key techniques, LearnedFTL considerably speeds up address translation while reducing the number of flash read accesses caused by the address translation. Our extensive experiments on a FEMU-based prototype show that LearnedFTL can reduce up to 55.5% address translation-induced double reads. As a result, LearnedFTL reduces the P99 tail latency by 2.9× ∼ 12.2× with an average of 5.5× and 8.2× compared to the state-of-the-art TPFTL and LeaFTL schemes, respectively.
Shengzhe Wang 0001, Zihang Lin, Suzhen Wu, Hong Jiang 0001, Jie Zhang 0048, Bo Mao 0003
HPCA1
2024 FSDedup: Feature-Aware and Selective Deduplication for Improving Performance of Encrypted Non-Volatile Main Memory
abstract
Enhancing the endurance, performance, and energy efficiency of encrypted Non-Volatile Main Memory (NVMM) can be achieved by minimizing written data through inline deduplication. However, existing approaches applying inline deduplication to encrypted NVMM suffer from substantial performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). In the preliminary ESD [ 14 ], we proposed the Error Correcting Code (ECC) assisted selective deduplication scheme, utilizing the ECC information as a fingerprint to identify similar data effectively and then leveraging the selective deduplication technique to eliminate a large amount of redundant data with high reference counts. In this article, we proposed FSDedup. Compared with ESD, FSDedup could leverage the prefetch cache to reduce the read overhead during similarity comparison and utilize the cache refresh mechanism to identify further and eliminate more redundant data. Extensive experimental evaluations demonstrate that FSDedup can enhance the performance of the NVMM system further than the ESD. Experimental results show that FSDedup can improve both write and read speed by up to 1.8×, enhance Instructions Per Cycle by up to 1.5×, and reduce energy consumption by up to 2.0×, compared to ESD.
Chunfeng Du, Zihang Lin, Suzhen Wu, Yifei Chen 0011, Jiapeng Wu, Shengzhe Wang 0001, Weichun Wang 0002, Bo Mao 0003
ACM Trans. Storage6
2023 iKnowFirst: An Efficient DPU-Assisted Compaction for LSM-Tree-Based Key-Value Stores
abstract
In scenarios with write-intensive workloads, LSM-tree-based key-value stores, such as RocksDB, suffer from compaction-induced performance degradation. RocksDB provides configurable compaction options to mitigate the severe read/write amplification problems associated with compaction. The advent of the Data Processing Unit (DPU) allows us to better utilize the configurable options of RocksDB to guide the key-value store system in choosing a suitable compaction strategy with prior knowledge of the workload characteristics. This paper proposes iKnowFirst, an efficient DPU-assisted key-value store. iKnowFirst (1) sets a data buffer on the DPU and separates hot-cold data to relieve the pressure of subsequent LSM-tree compaction, (2) senses the characteristics of the workloads in advance, and dynamically guides RocksDB to choose different compaction modes or enable/disable compaction when the workloads change, to cope with the scenario of write outbreak, and (3) implements an auto-selecting interface for compaction strategies selection. Our prototype implementation and experimental results show that iKnowFirst achieves 3.2× improvement compared to the original RocksDB on write-intensive and highly skewed workloads while showing acceptable performance under read-intensive workloads.
Jiahong Chen, Shengzhe Wang 0001, Suzhen Wu, Bo Mao 0003
ASAP2
2023 ESD: An ECC-assisted and Selective Deduplication for Encrypted Non-Volatile Main Memory
abstract
Reducing write data to encrypted Non-Volatile Main Memory (NVMM) can directly improve NVMM’s endurance, performance, and energy efficiency. However, existing works that straightforwardly apply inline deduplication on encrypted NVMM can significantly lead to system performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). This paper proposes ESD, an ECC-assisted and Selective Deduplication for encrypted NVMM by exploiting both the device characteristics (ECC mechanism) and the workload characteristics (content locality). First, ESD utilizes the ECC information associated with each cache line evicted from the Last-Level Cache (LLC) as the fingerprint to identify data similarity and avoids the costly hash calculating overhead on the non-duplicate cache lines. Second, ESD leverages selective deduplication to exploit the content locality within cache lines by only storing the fingerprints with high reference counts in the memory cache to reduce the memory space overhead and avoid fingerprints NVMM_lookup operations. The experimental results show that ESD can significantly speed up the writes by up to 3.4x, 4.3x, and 2.6x, speed up the reads by up to 5.3x, 5.0x, and 2.0x, and reduce the energy consumption by up to 96.3%, 96.2%, and 56.6% than Baseline, Dedup SHA1, and DeWrite, respectively. Meanwhile, ESD also can significantly outperform other schemes in tail latency.
Chunfeng Du, Suzhen Wu, Jiapeng Wu, Bo Mao 0003, Shengzhe Wang 0001
HPCA5
2023 LearnedSync: A Learning-Based Sync Optimization for Cloud Storage
Suzhen Wu, Shengzhe Wang 0001, Chunfeng Du, Jiayang Guo, Yijie Pan, Naian Xiao, Bo Mao 0003
ICA3PP (2)3