VLDB 2026 Research / reviewers in the wild / expert
Qi Wang 0041
dblp:19/1924-41
· DBLP profile ↗
8ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0002-2259-6063ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I-COR: Instruction-Level Fault Tolerance for Register File in 3-Stage Pipeline RISC-V ProcessorsabstractThe RISC-V architecture is increasingly deployed in safety-critical domains, such as automotive systems, yet the register file remains susceptible to Single Event Upsets (SEUs). Traditional fault-tolerance schemes like Triple Modular Redundancy (TMR) and Error Correction Codes (ECC) incur excessive area overhead (>200%) or significant performance penalties (15.3–20.5% timing degradation), and critically, cannot prevent error accumulation. To address this, we propose I-COR , an instruction-level correction scheme that targets register file data errors. When an instruction accesses a corrupted register, I-COR replaces it with a corrective instruction pair: the first instruction rectifies the erroneous value through the write-back path to prevent error accumulation, while the second re-executes the original operation. By reusing existing forwarding paths, this pair maintains pipeline efficiency, introducing just one-cycle correction latency. Implemented on the Ibex core and synthesized in 28 nm, 110 nm CMOS, and 180 nm SOI technologies, I-COR reduces area overhead by 50.92–62.83% compared to TMR while avoiding ECC’s timing degradation. Our automated two-stage fault-injection validation demonstrates I-COR’s effectiveness: it achieves complete elimination of critical faults (0% occurrence for both application output mismatches and system hangs) and reduces non-critical faults (internal architectural errors without system-level impact) down to 0.12–3.17%, outperforming existing schemes by 2.6–28.8×. By combining minimal area, zero timing penalty, single-cycle correction latency, and guaranteed prevention of error accumulation, I-COR provides an efficient architectural-level fault-tolerance solution specifically designed for register file protection in safety-critical three-stage pipeline RISC-V processors. Zewen Cao, Qi Wang 0041 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | Temperature Effects of Program Operation in 3-D nand Flash Memory: Observations, Analysis, and SolutionsabstractAs flash memory storage density continues to increase, it has become the mainstream storage medium for electronic devices. Writing data in low-temperature environment causes distortions in the flash memory threshold voltage distribution (TVD), which spikes the raw bit error rate and ultimately leads to degradation of the performance of flash-based electronic devices. To ameliorate the reliability problem caused by flash memory read and program temperature variations, this study proposes a flash memory programming temperature compensation algorithm based on read reference voltage (PTC-RRV) calibration. 3-D triple-level cell (TLC) flash memory is currently the mainstream storage medium for consumer electronics. Based on a large number of real tests on this type of chips, the relationship between the programming/reading temperature and the TVD of flash memory is fully characterized, and a programming temperature compensation model is constructed. The model evaluation results show that the PTC-RRV strategy can significantly reduce the average number of read-retry of low temperature written data and effectively improve the storage reliability and read performance of flash memory, whose optimization effect on electronic devices is better than the existing temperature compensation algorithms. Debao Wei, Qi Wang 0041, Yongchao Wang 0001, Liyan Qiao, Zongliang Huo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Hybrid Neural Network Model for Raw Bit Error Rate Prediction of 3-D TLC NAND Flash Memoryabstract3D TLC NAND flash memory achieves high storage density due to multi-bit technology and vertical stacked structure, while suffering from severe reliability issues, particularly a high raw bit error rate (RBER). Accurate prediction of RBER in 3D TLC NAND flash memory enables proactive data management, thereby enhancing performance, improving reliability, and extending the lifetime of the memory. This paper focuses on accurately estimating the RBER of flash memory to provide more detailed information for the precise control and optimization strategies. In this paper, we propose a Hybrid Neural Network model (HNN) for RBER prediction, which combines the advantages of multi-layer perceptron (MLP) and recurrent neural network (RNN). The model utilizes a multi-layer perceptron (MLP) for static feature extraction and the gated recurrent unit (GRU) for time series analysis. By combining these approaches, the HNN can effectively capture the complex relationships and temporal patterns that affect the RBER. Experimental results, based on actual chip data, demonstrate that the HNN can predict the RBER with an r-squared value greater than 0.96, significantly outperforming traditional machine learning algorithms and neural networks. Jing He 0020, Qianqi Zhao, Xuhong Qiang, Qi Wang 0041, Qianhui Li, Zongliang Huo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | POFGSP: Priority-Based Out-of-Order Scheduling and Fine-Grain Status Polling for SSD Performance ImprovementabstractWith the development of flash technology, the increasing throughput gap betweennandflash memory (NFM) arrays and the I/O interface has become a performance bottleneck for NFM-based solid-state drives (SSDs). Multilevel parallelism techniques have been employed on modern SSDs to meet the challenge of increasing demands for bandwidth in I/O-intensive workloads. However, conventional parallel methods only monitor the status of ways, resulting in the “idle bubble”—idle time of the dies cannot execute subsequent operations until all the dies in the way complete command execution. This issue limits the resource utilization and performance of SSDs. To minimize the idle bubble, we propose priority-based out-of-order scheduling and fine-grain status polling (POFGSP). The priority-based out-of-order scheduling relaxes constraints on command execution order and schedules commands with the same execution time to be executed in parallel. Therefore, the scheduler reduces these idle bubbles caused by differences in command execution times. Moreover, the fine-grain status polling approach polls the die-level status during the interface’s idle time, reducing idle bubbles with accurate status. Compared to state-of-the-art schedulers, our POFGSP approach can reduce request response time by 35.6% under real-world cloud block storage workloads and improve the SSD system’s maximum bandwidth by 8.7%–74.9%. Wentian Wu, Qianhui Li, Tong Qu, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | NV-APP: Invalid Programming Performance Improved No-Verify and Adaptive Pulse Programming Scheme for 3-D QLC nand FlashabstractQuad-level cell (QLC) has received significant attention recently due to its extremely high storage capacity. However, because of its poor reliability, QLC-based solid-state drives (SSDs) require a two-step programming to reduce the layer interference. But during the interval between two programming steps on the same wordline (WL), data could be invalidated from update operations, leading to invalid programming and degraded performance. To mitigate the performance loss, we propose the NV-APP scheme to minimize the program and verify pulses during the second-step programming. NV-APP integrates the no-verify (NV) scheme and the adaptive pulse programming scheme (APP). The NV scheme omits verify pulses of invalid verify voltages. The APP scheme adaptively increases the programming step voltage$(V_{\mathrm { step}})$to accelerate cells’ threshold voltage shift, reducing the number of both program and verify pulses. Device-level simulation results show that the NV-APP scheme reduces the total number of program pulses by an average of 27.03% and verify pulses by an average of 48.70% across various invalid cases during the second-step programming. Based on a modified 3-D QLC SSD simulator with typical traces, the experiments demonstrate that our scheme reduces two-step programming time by an average of 17% on partially invalid WLs, close to the 19.8% reduction achieved by the ideal scheme with no performance loss. Qianqi Zhao, Jing He 0020, Tong Qu, Wentian Wu, Qianhui Li, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | LIAD: A Method for Extending the Effective Time of 3-D TLC NAND Flash Hard DecisionabstractTriple-level cell NAND flash memory is widely used today due to its higher storage density and capacity. However, with the increase in the storage density, lower reliability results in more read times for flash memory and significantly reduces the read performance. In order to avoid unnecessary read operations, this article proposes a hard decision–soft decoding method called location information-assisted decoding (LIAD) method, which determines the additional information required for decoding by mutual information, and then transmits the required information to correct the log-likelihood ratio (LLR). Different from the conventional LLR correction algorithm, this method does not require additional read operations and correct data. Only using sensing results, our method can reduce uncorrectable error bit rate (UBER) by up to 99%, and the system read latency under SSDsim (Hu et al. 2011) simulation can be reduced by up to 53%. Jing He 0020, Qianhui Li, Xianliang Wang, Tian-Chun Ye 0001, Qi Wang 0041, Zongliang Huo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | Interleaved LDPC Decoding Scheme Improves 3-D TLC NAND Flash Memory System PerformanceabstractAlthough NAND flash memory does a lot of work in effectively using error correcting code (ECC) to reduce uncorrectable bit error rate (UBER). However, if the frame error rate (FER) is not reduced, the lower UBER cannot effectively reduce the read latency of the flash memory system. This phenomenon is especially evident at the end of the flash memory lifetime, where conventional methods significantly reduce the UBER but not to zero, and the remaining error bits are still evenly distributed throughout the flash memory page, resulting in a significant increase in read latency. In this article, an interleaved LDPC decoding scheme is proposed. By re-evaluating the flash memory channel during the decoding process, the codewords in the flash memory page are corrected frame by frame, and the problem of high FER is solved at the end of the flash memory lifetime. Compared with the conventional algorithm, the proposed method can reduce the FER by up to 34%, reduce the average decoding iterations by 63.4%, and reduce the read latency by up to 65%. Jing He 0020, Xianliang Wang, Qianhui Li, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Adaptive Pulse Program Scheme to Improve the Vth Distribution for 3D NAND FlashabstractAs the demand of multi-bit/cell NAND flash devices is increasing rapidly, getting a narrow cell Vth distribution becomes more challenging and necessary. To overcome this challenge, an adaptive pulse program (APP) scheme is reported that can tighten the Vth distribution in this work. Compared with conventional incremental step pulse program (ISPP) scheme, this proposed scheme uses adaptive program pulse to the cells with different program speed, which targets to prevent the extension of Vth distribution's upper tail. Our experimental result demonstrates that APP scheme achieves ~15% improvement for reducing cell Vth distribution width. This comparison of APP scheme and general ISPP scheme is performed by the FPGA platform using 64-layer 3D charge-trapping NAND flash chip. Zhichao Du, Yu Wang 0097, Qi Wang 0041, Zongliang Huo |
ISCAS | 5 |