Zongliang Huo

dblp:160/4435 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-9845-5649ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Temperature Effects of Program Operation in 3-D nand Flash Memory: Observations, Analysis, and Solutions
abstract
As flash memory storage density continues to increase, it has become the mainstream storage medium for electronic devices. Writing data in low-temperature environment causes distortions in the flash memory threshold voltage distribution (TVD), which spikes the raw bit error rate and ultimately leads to degradation of the performance of flash-based electronic devices. To ameliorate the reliability problem caused by flash memory read and program temperature variations, this study proposes a flash memory programming temperature compensation algorithm based on read reference voltage (PTC-RRV) calibration. 3-D triple-level cell (TLC) flash memory is currently the mainstream storage medium for consumer electronics. Based on a large number of real tests on this type of chips, the relationship between the programming/reading temperature and the TVD of flash memory is fully characterized, and a programming temperature compensation model is constructed. The model evaluation results show that the PTC-RRV strategy can significantly reduce the average number of read-retry of low temperature written data and effectively improve the storage reliability and read performance of flash memory, whose optimization effect on electronic devices is better than the existing temperature compensation algorithms.
Debao Wei, Qi Wang 0041, Yongchao Wang 0001, Liyan Qiao, Zongliang Huo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Hybrid Neural Network Model for Raw Bit Error Rate Prediction of 3-D TLC NAND Flash Memory
abstract
3D TLC NAND flash memory achieves high storage density due to multi-bit technology and vertical stacked structure, while suffering from severe reliability issues, particularly a high raw bit error rate (RBER). Accurate prediction of RBER in 3D TLC NAND flash memory enables proactive data management, thereby enhancing performance, improving reliability, and extending the lifetime of the memory. This paper focuses on accurately estimating the RBER of flash memory to provide more detailed information for the precise control and optimization strategies. In this paper, we propose a Hybrid Neural Network model (HNN) for RBER prediction, which combines the advantages of multi-layer perceptron (MLP) and recurrent neural network (RNN). The model utilizes a multi-layer perceptron (MLP) for static feature extraction and the gated recurrent unit (GRU) for time series analysis. By combining these approaches, the HNN can effectively capture the complex relationships and temporal patterns that affect the RBER. Experimental results, based on actual chip data, demonstrate that the HNN can predict the RBER with an r-squared value greater than 0.96, significantly outperforming traditional machine learning algorithms and neural networks.
Jing He 0020, Qianqi Zhao, Xuhong Qiang, Qi Wang 0041, Qianhui Li, Zongliang Huo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 POFGSP: Priority-Based Out-of-Order Scheduling and Fine-Grain Status Polling for SSD Performance Improvement
abstract
With the development of flash technology, the increasing throughput gap betweennandflash memory (NFM) arrays and the I/O interface has become a performance bottleneck for NFM-based solid-state drives (SSDs). Multilevel parallelism techniques have been employed on modern SSDs to meet the challenge of increasing demands for bandwidth in I/O-intensive workloads. However, conventional parallel methods only monitor the status of ways, resulting in the “idle bubble”—idle time of the dies cannot execute subsequent operations until all the dies in the way complete command execution. This issue limits the resource utilization and performance of SSDs. To minimize the idle bubble, we propose priority-based out-of-order scheduling and fine-grain status polling (POFGSP). The priority-based out-of-order scheduling relaxes constraints on command execution order and schedules commands with the same execution time to be executed in parallel. Therefore, the scheduler reduces these idle bubbles caused by differences in command execution times. Moreover, the fine-grain status polling approach polls the die-level status during the interface’s idle time, reducing idle bubbles with accurate status. Compared to state-of-the-art schedulers, our POFGSP approach can reduce request response time by 35.6% under real-world cloud block storage workloads and improve the SSD system’s maximum bandwidth by 8.7%–74.9%.
Wentian Wu, Qianhui Li, Tong Qu, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 NV-APP: Invalid Programming Performance Improved No-Verify and Adaptive Pulse Programming Scheme for 3-D QLC nand Flash
abstract
Quad-level cell (QLC) has received significant attention recently due to its extremely high storage capacity. However, because of its poor reliability, QLC-based solid-state drives (SSDs) require a two-step programming to reduce the layer interference. But during the interval between two programming steps on the same wordline (WL), data could be invalidated from update operations, leading to invalid programming and degraded performance. To mitigate the performance loss, we propose the NV-APP scheme to minimize the program and verify pulses during the second-step programming. NV-APP integrates the no-verify (NV) scheme and the adaptive pulse programming scheme (APP). The NV scheme omits verify pulses of invalid verify voltages. The APP scheme adaptively increases the programming step voltage$(V_{\mathrm { step}})$to accelerate cells’ threshold voltage shift, reducing the number of both program and verify pulses. Device-level simulation results show that the NV-APP scheme reduces the total number of program pulses by an average of 27.03% and verify pulses by an average of 48.70% across various invalid cases during the second-step programming. Based on a modified 3-D QLC SSD simulator with typical traces, the experiments demonstrate that our scheme reduces two-step programming time by an average of 17% on partially invalid WLs, close to the 19.8% reduction achieved by the ideal scheme with no performance loss.
Qianqi Zhao, Jing He 0020, Tong Qu, Wentian Wu, Qianhui Li, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 LIAD: A Method for Extending the Effective Time of 3-D TLC NAND Flash Hard Decision
abstract
Triple-level cell NAND flash memory is widely used today due to its higher storage density and capacity. However, with the increase in the storage density, lower reliability results in more read times for flash memory and significantly reduces the read performance. In order to avoid unnecessary read operations, this article proposes a hard decision–soft decoding method called location information-assisted decoding (LIAD) method, which determines the additional information required for decoding by mutual information, and then transmits the required information to correct the log-likelihood ratio (LLR). Different from the conventional LLR correction algorithm, this method does not require additional read operations and correct data. Only using sensing results, our method can reduce uncorrectable error bit rate (UBER) by up to 99%, and the system read latency under SSDsim (Hu et al. 2011) simulation can be reduced by up to 53%.
Jing He 0020, Qianhui Li, Xianliang Wang, Tian-Chun Ye 0001, Qi Wang 0041, Zongliang Huo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2023 Interleaved LDPC Decoding Scheme Improves 3-D TLC NAND Flash Memory System Performance
abstract
Although NAND flash memory does a lot of work in effectively using error correcting code (ECC) to reduce uncorrectable bit error rate (UBER). However, if the frame error rate (FER) is not reduced, the lower UBER cannot effectively reduce the read latency of the flash memory system. This phenomenon is especially evident at the end of the flash memory lifetime, where conventional methods significantly reduce the UBER but not to zero, and the remaining error bits are still evenly distributed throughout the flash memory page, resulting in a significant increase in read latency. In this article, an interleaved LDPC decoding scheme is proposed. By re-evaluating the flash memory channel during the decoding process, the codewords in the flash memory page are corrected frame by frame, and the problem of high FER is solved at the end of the flash memory lifetime. Compared with the conventional algorithm, the proposed method can reduce the FER by up to 34%, reduce the average decoding iterations by 63.4%, and reduce the read latency by up to 65%.
Jing He 0020, Xianliang Wang, Qianhui Li, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 A Small Ripple and High-Efficiency Wordline Voltage Generator for 3-D nand Flash Memories
abstract
This article presents a small ripple and high-efficiency wordline (WL) voltage generator to supply voltage to the selected/unselected WLs for program and read operation in 3-D NAND Flash memories. With the proposed scheme of dynamic pump clock voltage and frequency scaling, the output ripple voltage can be minimized to reduce the variation of threshold voltage of the memory cells and the efficiency can be improved to save power consumption at the same time. What is more, with the minimized ripple, the requirement for power supply rejection ratio (PSRR) of the high-voltage regulator used to provide low noise WL voltage can be reduced. The proposed WL voltage generator has been fabricated in a 0.18-$\mu \text{m}$triple-well CMOS process and the core chip size is 0.53 mm2. While operating at a 1.8-V supply, the measurement results show that the ripple voltage is 5.5 mV at 12-V output voltage under the typical 50-pF cap load conditions of 3-D NAND Flash memories. Furthermore, the output ripple hardly varies with the increase of load current. In addition, the maximum power efficiency is 49% at 400$\mu \text{A}$and can be maintained over 30% in the 50–450-$\mu \text{A}$current load range.
Qianqian Wang 0012, Cece Huang, Qianhui Li, Zongliang Huo
IEEE Trans. Very Large Scale Integr. Syst.5
2020 Adaptive Pulse Program Scheme to Improve the Vth Distribution for 3D NAND Flash
abstract
As the demand of multi-bit/cell NAND flash devices is increasing rapidly, getting a narrow cell Vth distribution becomes more challenging and necessary. To overcome this challenge, an adaptive pulse program (APP) scheme is reported that can tighten the Vth distribution in this work. Compared with conventional incremental step pulse program (ISPP) scheme, this proposed scheme uses adaptive program pulse to the cells with different program speed, which targets to prevent the extension of Vth distribution's upper tail. Our experimental result demonstrates that APP scheme achieves ~15% improvement for reducing cell Vth distribution width. This comparison of APP scheme and general ISPP scheme is performed by the FPGA platform using 64-layer 3D charge-trapping NAND flash chip.
Zhichao Du, Yu Wang 0097, Qi Wang 0041, Zongliang Huo
ISCAS6
2015 A 1G-cell floating-gate NOR flash memory in 65 nm technology with 100 ns random access time
Zongliang Huo, Ming Liu 0022, Liyang Pan
Sci. China Inf. Sci.4