EDBT 2026 Demo / reviewers in the wild / expert
Lanlan Cui
dblp:258/5708
· DBLP profile ↗
11ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-9509-741XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSALDPC: A Syndrome-Sum Based Adaptive LDPC Decoding Scheme for NAND Flash MemoryabstractThe continuous increase in storage density for 3D NAND flash memory, driven by multi-layer stacking and multilevel cell technology, leads to a significant overlap and shift in the threshold voltage distributions. This phenomenon significantly elevates the raw bit error rate (RBER) and poses serious challenges to data reliability. Although solutions based on low-density parity-check (LDPC) codes and read-retry schemes have become the standard approach to mitigate high RBER, the latency introduced by repeated read operations considerably degrades system read performance. This paper proposes a syndrome-sum based adaptive LDPC decoding scheme, named SSALDPC. After an initial hard decision decoding failure, our scheme utilizes the real-time syndrome sum (SS)—generated during the decoding process—to assess the severity of errors. Based on this assessment, it adaptively selects the most appropriate subsequent decoding strategy from three modes: EfficiencyMode (E-Mode), Balance-Mode (B-Mode), or Performance-Mode (P-Mode). Experimental results demonstrate that the proposed SSALDPC scheme reduces the number of read-retry operations and decreases decoding latency under various RBER conditions, while maintaining high error correction capability. Lanlan Cui, Fei Wu 0005, Kun Jiang 0001, Yeqiu Xiao, Renzhi Xiao, Changsheng Xie 0001 |
DATE | 1 |
| 2026 | Enhanced LDPC Coding for 3-D TLC NAND Flash Memory: Leveraging RBER Difference From Intralayer VariationabstractNAND flash memory employs high code rate low-density parity-check (LDPC) codes to reduce the amount of redundant data that must be added. When the code rate is high, although the redundancy space is small, the error correction capability is inferior to medium or low code rate LDPC. RBER varies among the storage layers for 3D triple-level cell (TLC) NAND flash memory, which increases the frequency of read retry operations. Repeatedly initiating read retry seriously increases the decoding latency and decreases the performance of the 3D TLC NAND flash memory. To alleviate this problem, this article proposes Intra-Layer Variation aware LDPC coding, called LVLDPC. The proposed LVLDPC scheme establishes the correlation between inter-layer interference and raw bit error rate (RBER) based on a neural network model. By analyzing and predicting RBER through the neural network model, we are able to categorize RBER into distinct levels. Then, we then select LDPC codes with appropriate error correction capabilities to decode data with varying levels of RBER. Through this scheme, we don’t need to start read retry when RBER <1.56×10-2. The iteration number is reduced by 67% in total. This scheme only causes 1.15% space overhead, which is negligible. For the stage with high RBER, the number of iterations of LVLDPC is still large, and the extended LVLDPC scheme (eLVLDPC) is further proposed to reduce the use of high code rate and reduce the number of iterations by 19.1%, expanding the correctable RBER threshold to 2.68×10-2. Lanlan Cui, Fei Wu 0005, Meng Zhang 0014, Zhanzhan Zhao, Kun Jiang 0001, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | QuickScale: A Quick Scaling Scheme for Erasure-Coded Deduplicated Storage SystemsabstractModern deduplication-based storage systems employ erasure coding to post-deduplicate data to ensure reliability. In the erasure-coded deduplicated storage systems, existing coding schemes mainly perform erasure coding over a single object (intra-object coding) to address the problems of degraded read performance and storage inefficiency. However, due to intra-object coding, traditional storage scaling schemes need to reorganize and relocate all stored data between the client and storage nodes to support storage scaling, which inevitably incurs substantial data transfer overhead, thereby degrading system scalability. In this paper, we propose a fast scaling scheme called QuickScale for erasure-coded deduplicated storage systems. QuickScale first implements the intra-object coding in a realistic distributed storage environment and verifies that this coding scheme can significantly alleviate the degraded read performance by 35%-67% and save storage space by 24% 30%. Motivated by these results, QuickScale uses a node-to-node scaling approach that avoids repeatedly transferring all stored data between the client and storage nodes during scaling, thereby improving the scalability. Experimental evaluations using six real-world datasets demonstrate that QuickScale achieves better scaling performance (in terms of scaling time during scaling) over the state-of-the-art scaling schemes by up to 63.5%-76.4%. Chunxue Zuo, Lanlan Cui, Fang Wang 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | DHD: Double Hard Decision Decoding Scheme for NAND Flash MemoryabstractWith the advancement of NAND flash technology, the increased storage density leads to intensified interference, which in turn raises the error rate during data retrieval. To ensure data reliability, low-density parity-check (LDPC) codes are extensively employed for error correction in NAND flash memory. Although LDPC soft decision decoding offers high error correction capability, it comes with a significant latency. Conversely, hard-decision decoding, although faster, lacks sufficient error correction strength. Consequently, flash memory typically initiates with hard-decision decoding and resorts to multiple soft decision decoding upon failure. To minimize decoding latency, this paper proposes a decoding mechanism based on the double hard decision, called DHD. This DHD scheme improves the Log-Likelihood Ratio (LLR) in the hard decision process. After the first hard decision fails, the read reference voltage (RRV) is adjusted to perform the second hard decision decoding. If the second hard decision also fails, soft decision decoding is then employed. Experimental results demonstrate that when the Raw Bit Error Rate (RBER) is$8.5 \times 10^{-3}$, DHD reduces the Frame Error Rate (FER) by 86.4% compared to the traditional method. Lanlan Cui, Yichuan Wang 0003, Renzhi Xiao, Xinhong Hei 0001 |
DATE | 1 |
| 2025 | Write-Optimized Persistent Hash Index for Non-Volatile MemoryabstractA hashing index provides rapid search performance by swiftly locating key-value items. Non-volatile memory (NVM) technologies have driven research into hashing indexes for NVM, combining hard disk persistence with DRAM-level performance. Nevertheless, current NVM-based hashing indexes must tackle data inconsistency challenges caused by NVM write reordering or partial writes, and mitigate rapid local wear due to frequent updates, considering NVM's limited endurance. The temporary allocation of buckets in NVM-based chained hashing to resolve hash collisions prolongs the critical path for writing, thus hampering write performance. This paper presents WOPHI, a write-optimized persistent hash index scheme for NVM. By utilizing log-free failure-atomic writes, WOPHI minimizes data consistency overhead and addresses hash conflicts with bucket pre-allocation. Experimental results underscore WOPHI's significant performance enhancements, with insertion latency slashed by up to 88.2% and deletion latency boosted by up to 82.6% compared to existing state-of-the-art schemes. Moreover, WOPHI substantially mitigates data consistency overhead, reducing cache line flushes by 59.3%, while maintaining robust write throughput for insert and delete operations. Renzhi Xiao, Dan Feng 0001, Yuchong Hu, Lanlan Cui |
DATE | 5 |
| 2024 | LVLDPC: Intra-Layer Variation Aware LDPC Coding for 3D TLC NAND Flash MemoryabstractNAND flash memory employs high code rate low-density parity-check (LDPC) code to minimize redundant data. High code rates reduce redundancy but compromise error cor-rection compared to medium/low code rates. Raw bit error rate (RBER) varies among the storage layers for 3D triple-level cell (TLC) NAND flash memory, which causes the number of using read retry to increase. Repeatedly initiating read retry seriously increases the decoding latency and decreases the performance of the 3D TLC NAND flash memory. To alleviate this problem, this article proposes Intra-Layer Variation aware LDPC coding, called LVLDPC. The LVLDPC scheme categorizes RBER into distinct levels. Then, we select LDPC codes with appropriate error correction capabilities to decode data with varying levels of RBER. Through this scheme, we don't need to start read-retry when RBER$< 1.56\times 10^{-2}$. The iteration number is reduced by 67% in total. This scheme only causes 1.15 % space overhead, which is negliaible, Lanlan Cui, Meng Zhang 0014, Fei Wu 0005 |
ICCD | 1 |
| 2024 | Read-Optimized Persistent Hash Index for Query Acceleration through Fingerprint Filtering and Lock-Free PrefetchingabstractHash indexes are widely used in key-value storage systems due to their ability to perform rapid single-point queries. The persistent memory (PM) technology has received significant attention in both academia and industry due to its high performance, non-volatility, and large capacity characteristics. Currently, hash indexes tailored for persistent memories have been extensively researched. However, through an in-depth experimental study, we have discovered that existing persistent hash indexes suffer from low query performance. This is primarily due to persistent memory's higher read latency than DRAM's, which reduces the performance of both positive and negative queries in persistent hash indexes. Additionally, the former's higher read lock overhead further diminishes query performance. To address the above problems, we propose in this paper a Read-Optimized Persistent Hash Index, referred to as ROPHI, based on fingerprint filtering and lock-free prefetching. By employing a fingerprint filtering method, ROPHI introduces a DRAM-based Cuckoo filter to store fingerprints of keys on top of the PM-based hash table, effectively mitigating the time-consuming access overhead of persistent memory hash tables by accessing only the DRAM-based filter. Additionally, ROPHI employs lock-free prefetching for positive query acceleration, utilizing lock-free optimistic concurrent read techniques to avoid read lock overhead and high-speed cache prefetching techniques to reduce access overhead to persistent memory. Experimental results on the Intel Optane DC Persistent Memory Module (DCPMM) platform demonstrate that ROPHI significantly improves query performance over existing persistent hash index schemes. Specifically, ROPHI achieves an improvement of 2.67×-13.59× in negative query performance and 1.72x-7.86x in positive query performance. ROPHI outperforms the state-of-the-art SmartHT in positive query throughput by 34.5%, and in insertion and deletion throughput by 9.20% and 19.87% respectively, while sacrificing only 1.93% of negative query throughput. Additionally, it achieves a 5.07x improvement in recovery efficiency. Renzhi Xiao, Dan Feng 0001, Yuchong Hu, Hong Jiang 0001, Lanlan Cui, Guanglei Xu, Fang Wang 0001 |
ICCD | 7 |
| 2022 | A Low Bit-Width LDPC Min-Sum Decoding Scheme for NAND FlashabstractFor NAND flash memory, designing a good low-density parity-check (LDPC) decoding algorithm could ensure data reliability. When the decoding algorithm is implemented in hardware, it is necessary to achieve an attractive tradeoff between implementation complexity and decoding performance. In this article, a novel low-bit-width decoding scheme is introduced. In this scheme, the quasi-cyclic LDPC (QC-LDPC) is used, and the row-layered normalized min-sum algorithm is improved by restricting the amplitude of minimum and second-minimum values in each check node (CN) updating. The simulation shows that our approach achieves a lower uncorrectable bit error rate (UBER) with a negligible increase in computational complexity, especially with low-precision input log-likelihood ratio (LLR). Lanlan Cui, Fei Wu 0005, Zhonghai Lu, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Improving LDPC Decoding Performance for 3D TLC NAND Flash by LLR Optimization Scheme for Hard and Soft DecisionabstractLow-density parity-check (LDPC) codes have been widely adopted in NAND flash in recent years to enhance data reliability. There are two types of decoding, hard-decision and soft-decision decoding. However, for the two types, their error correction capability degrades due to inaccurate log-likelihood ratio (LLR) . To improve the LLR accuracy of LDPC decoding, this article proposes LLR optimization schemes, which can be utilized for both hard-decision and soft-decision decoding. First, we build a threshold voltage distribution model for 3D floating gate (FG) triple level cell (TLC) NAND flash. Then, by exploiting the model, we introduce a scheme to quantize LLR during hard-decision and soft-decision decoding. And by amplifying a portion of small LLRs, which is essential in the layer min-sum decoder, more precise LLR can be obtained. For hard-decision decoding, the proposed new modes can significantly improve the decoder’s error correction capability compared with traditional solutions. Soft-decision decoding starts when hard-decision decoding fails. For this part, we study the influence of the reference voltage arrangement of LLR calculation and apply the quantization scheme. The simulation shows that the proposed approach can reduce frame error rate (FER) for several orders of magnitude. Lanlan Cui, Fei Wu 0005, Meng Zhang 0014, Renzhi Xiao, Changsheng Xie 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2020 | BeLDPC: Bit Errors Aware Adaptive Rate LDPC Codes for 3D TLC NAND Flash MemoryabstractThree-dimensional (3D) NAND flash memory has high capacity and cell storage density by using the multi-bit technology and vertical stack architecture, but degrading data reliability due to high raw bit error rates (RBER) caused by program/erase (P/E) cycles and retention periods. Low-density parity-check (LDPC) codes become more popular error-correcting technologies to improve data reliability due to strong error correction capability, but introducing more decoding iterations at higher RBER. To reduce decoding iterations, this paper proposes BeLDPC: bit errors aware adaptive rate LDPC codes for 3D triple-level cell (TLC) NAND flash memory. Firstly, bit error characteristics in 3D charge trap TLC NAND flash memory are studied on a real FPGA testing platform, including asymmetric bit flipping and temporal locality of bit errors. Then, based on these characteristics, a high-efficiency LDPC code is designed. Experimental results show BeLDPC can reduce decoding iterations under different P/E cycles and retention periods. Meng Zhang 0014, Fei Wu 0005, Lanlan Cui, Yahui Zhao, Changsheng Xie 0001 |
DATE | 5 |
| 2019 | VaLLR: Threshold Voltage Distribution Aware LLR Optimization to Improve LDPC Decoding Performance for 3D TLC NAND FlashabstractLow-density parity-check (LDPC) codes have been widely adopted in NAND flash in recent years to improve data reliability. However, their error-correction capability degrades due to inaccurate log-likelihood ratio (LLR). To improve LLR accuracy of LDPC decoding, this paper proposes a threshold voltage distribution aware LLR optimization scheme, called VaLLR. Firstly, we build a threshold voltage distribution model for 3D triple-level cell (TLC) NAND flash. Then, by exploiting the model, we introduce the VaLLR scheme to quantize LLR during soft-decision decoding. And by amplifying a portion of small LLRs, which is essential in the layer minsum decoder, more precise LLR can be obtained. Finally, we study the influence of the reference voltage arrangement on LLR calculation and apply the VaLLR scheme during decoding. The simulation shows that the proposed approach can improve the FER performance for several orders of magnitude. Lanlan Cui, Fei Wu 0005, Meng Zhang 0014, Changsheng Xie 0001 |
ICCD | 1 |