Wei Zhao 0034

dblp:181/2852-34 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-2765-498XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 An Efficient Independent Read Scheme for Contemporary QLC SSDs
abstract
QLC solid-state disks (SSDs) are increasingly deployed in large-scale storage systems. While achieving remarkable storage density and cost-effectiveness, QLC NAND exhibits degraded performance. To alleviate the issue, Independent Multi-Plane (IMP) read has been proposed to leverage the plane-level parallelism under random read workloads. However, compared to the previous generations of chips, the variation in read latency of QLC chip has widened significantly, and the number of planes in a QLC chip has increased. As a result, the idle time in IMP commands has escalated dramatically. Moreover, the conventional layout of data and parities in redundant array of independent NAND (RAIN) increases the probability of high latency read occureneces, exacerbating the contribution to idle time. Consequently, the incorporation of IMP is inherently inefficient in contemporary QLC NAND flash chips. In this paper, we propose an efficient independent read (EFFIR) scheme to tackle this challenge. EFFIR features a read latency variation aware transaction service that properly combines read transactions to minimize idle time and proactively transfers read data to eliminate unnecessary delays. Moreover, EFFIR incorporates a read latency variation aware RAIN that reorganizes the layout of data and parity to mitigate the impact of high-latency data access on idle time. Our comprehensive experimental results demonstrate and elucidate how EFFIR significantly enhances SSD responsiveness while consistently delivering favorable performance across a diverse range of read-intensive workloads.
Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001, Feng Zhu 0024, Maojun Yuan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 CMD: A Cache-Assisted GPU Memory Deduplication Architecture
abstract
Massive off-chip accesses in graphics processing units (GPUs) are the main performance bottleneck. We find that many writes are duplicate, and the duplication can beinter-dupandintra-dup. Whileinter-dupmeans different memory blocks are identical, andintra-dupmeans all the 4B elements in a line are the same. In this work, we propose a cache-assisted GPU memory deduplication architecture named cache-assisted GPU memory deduplicated (CMD) to reduce the off-chip accesses via utilizing the data duplication in GPU applications. CMD includes three key design contributions which aim to reduce the three kinds of accesses: 1) a novel GPU memory deduplication architecture that removes theintra-dupandinter-duplines. We design several techniques to manage duplicate blocks, reducing massive off-chip writes; 2) we propose a cache-assisted read scheme to reduce the reads to duplicate data. When an L2 cache miss wants to read the duplicate block, if the reference block has been fetched to L2 and it is clean, we can copy it to the L2 missed block without accessing off-chip DRAM. As for the reads tointra-dupdata, CMD uses the on-chip metadata cache to get the data; and 3) when a cache line is evicted, the clean sectors in the line are invalidated while the dirty sectors are written back. However, most read-only victims are rereferenced from DRAM more than twice. Therefore, we add a full-associate FIFO to accommodate the read-only (it is also clean) victims to reduce the rereference counts. Experiments show that CMD can decrease the off-chip accesses by 31.01%, reduce the energy by 32.78% and improve performance by 42.53%. Besides, CMD can improve the performance of memory-intensive workloads by 57.56%.
Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Xueliang Wei, Bing Wu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 A Read Latency Variation Aware Independent Read Scheme for QLC SSDs
abstract
QLC flash-based SSDs has attracted growing interest and is expected to fit in read-intensive scenarios owing to its higher cost-effectiveness and shorter write endurance compared with the Triple-Level Cell (TLC) SSDs. Recently, new commands supporting independent reads such as Single Operation Multiple Locations (SOML) and Independent multi-plane (IMP) read are proposed to improve read performance. Unfortuantely, while independent read exhibits a significant performance improvement, we show in this paper that the exisiting approach fails to fully exploit its potential due to the larger read latency variation and more planes per die for current SSD architecture. Through a set of experiments, we demonstrate that the lack of a read latency variation aware machanism leads to low performance of independent read among a wide variety of workloads. To alleviate this issue, we propose LITA, which key idea is to combine read transactions with similar latency into one command. LITA includes (1) LIT, a latency variation aware transaction combination, (2) and TASAP, a latency variation aware transaction completion service. The experimental results show LITA can reduce read latency by 20.4% and 9.6% on average for 4-planes QLC SSDs under IMP and SOML, respectively.
Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001
DATE5
2023 ODLPIM: A Write-Optimized and Long-Lifetime ReRAM-Based Accelerator for Online Deep Learning
abstract
ReRAM-based Processing-In-Memory (PIM) architectures have demonstrated high energy efficiency and performance in deep neural network (DNN) acceleration. Most of the existing PIM accelerators for DNN focus on offline batch learning (OBL) which requires the whole dataset to be available before training. However, in the real world, data instances arrive in sequential settings, and even the data pattern may change, which calls concept drift. OBL requires expensive retraining to solve concept drift, whereas online deep learning (ODL) is evidenced to be a better solution to keep the model evolving over streaming data. Unfortunately, when ODL optimizes models over a large-scale data stream in the PIM system, unbalanced writes are more severe than OBL, due to the heavier weight updates, resulting in the amplification of unbalanced writes and lifetime deterioration. In this work, we propose ODLPIM, an online deep learning PIM accelerator that extends the system lifetime through algorithm-hardware co-optimization. ODLPIM adopts a novel write-optimized parameter update (WARP) scheme that reduces the non-critical weight updates in hidden layers. Besides, a table-based inter-crossbar wear-leveling (TIWL) scheme is proposed and applied to the hardware controller to achieve wear-leveling between crossbars for lifetime improvement. Experiments show that WARP reduces weight updates on average to 15.25% and up to 24% compared to that without WARP, and eventually prolongs system lifetime on average to 9.65% and up to 26.81%, with a negligible rise in cumulative error rate (up to 0.31%). By combining WARP with TIWL, the lifetime of ODLPIM is improved by an average of$\mathbf{12}.\mathbf{59}\times$and up to$\mathbf{17}.\mathbf{73}\times$.
Bing Wu 0001, Huan Cheng, Wei Zhao 0034, Xueliang Wei, Dan Feng 0001, Wei Tong 0001
DATE4
2023 SplitZNS: Towards an Efficient LSM-Tree on Zoned Namespace SSDs
abstract
The Zoned Namespace (ZNS) Solid State Drive (SSD) is a nascent form of storage device that offers novel prospects for the Log Structured Merge Tree (LSM-tree). ZNS exposes erase blocks in SSD as append-only zones, enabling the LSM-tree to gain awareness of the physical layout of data. Nevertheless, LSM-tree on ZNS SSDs necessitates Garbage Collection (GC) owing to the mismatch between the gigantic zones and relatively small Sorted String Tables (SSTables). Through extensive experiments, we observe that a smaller zone size can reduce data migration in GC at the cost of a significant performance decline owing to inadequate parallelism exploitation. In this article, we present SplitZNS, which introduces small zones by tweaking the zone-to-chip mapping to maximize GC efficiency for LSM-tree on ZNS SSDs. Following the multi-level peculiarity of LSM-tree and the inherent parallel architecture of ZNS SSDs, we propose a number of techniques to leverage and accelerate small zones to alleviate the performance impact due to underutilized parallelism. (1) First, we use small zones selectively to prevent exacerbating write slowdowns and stalls due to their suboptimal performance. (2) Second, to enhance parallelism utilization, we propose SubZone Ring, which employs a per-chip FIFO buffer to imitate a large zone writing style; (3) Read Prefetcher, which prefetches data concurrently through multiple chips during compactions; (4) and Read Scheduler, which assigns query requests the highest priority. We build a prototype integrated with SplitZNS to validate its efficiency and efficacy. Experimental results demonstrate that SplitZNS achieves up to 2.77× performance and reduces data migration considerably compared to the lifetime-based data placement. 1
Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001
ACM Trans. Archit. Code Optim.5
2023 APPcache+: An STT-MRAM-Based Approximate Cache System With Low Power and Long Lifetime
abstract
Due to high static power and low scalability, the traditional SRAM-based cache is not a good solution for image processing applications. Emerging spin transfer torque magnetic RAM (STT-MRAM) is a promising candidate for cache due to its low leakage power and high density. However, STT-MRAM suffers from high write energy. Therefore, by making use of the ability of tolerating minor errors in image processing applications, this work presents an STT-MRAM-basedAPProximatecachearchitecture (APPcache+) to write/read approximate data, which can largely reduce the cache energy and improve the STT-MRAM lifetime. APPcache+ includes three main designs. First, we find that there are many similar elements (e.g., pixels in images) in cache lines. Therefore, APPcache+ presents several lightweight similarity-based encoding techniques to remove redundant elements, thus, shortening the data size and reducing the energy of STT-MRAM cache. Second, we design a partial read scheme to reduce the read energy of the STT-MRAM cache. In the traditional decompression process, the whole line is fetched into the decompressor, leading to unnecessary read energy. The partial read scheme can largely reduce read energy while keeping the overhead low. Third, we observe the encoding schemes may lead to bit write imbalance. Therefore, we propose a lightweight Ping-Pong intraline wear-leveling scheme to improve the lifetime. Compared with the baseline, extensive evaluation results show that our APPcache+ can largely reduce the overall energy by 32.58%, improve lifetime by 40.7% with only 2.2% performance degradation, and 1.86% output quality loss.
Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zhangyu Chen, Bing Wu 0001, Chengning Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A Low-Latency and High-Endurance MLC STT-MRAM-Based Cache System
abstract
Spin-transfer torque magnetic random access memory (STT-MRAM) is a promising cache memory candidate due to its high density, low leakage power, and nonvolatility. Multilevel cell (MLC) STT-MRAM can further increase density by storing 2 bits in one cell’s hard and soft domain, respectively. However, MLC STT-MRAM suffers two-step write, leading to high write energy, long latency, and severe lifetime degradation. Current encoding techniques propose to encode the new data to reduce the two-step data writes. However, they have two weaknesses: 1) high area overhead, e.g., recent work TSE (Hsieh et al., 2020) needs extra 37.5% MLCs and 2) prolong the write latency due to an extra read. Therefore, we propose enhanced one-step write (EOSwrite) to write data in one step. EOSwrite includes line bypassing and four intraline encoding techniques. Line bypassing schemes can bypass the writes to zero or clean lines, leading to low write/read latency. As for the intraline techniques, we propose four write modes. They utilize the data patterns and the clean data in cache lines to write data in one step, therefore reducing the data write latency. The key idea of one-step write is to write as much data as possible in the soft domain of MLC STT-MRAM. EOSwrite can greatly relieve the weaknesses of the current encoding schemes. Evaluation results show that EOSwrite can improve the lifetime of MLC STT-MRAM by 56.96%, reduce dynamic energy by 33.95%, reduce access latency by 36.95%, and improve system performance of MLC STT-MRAM by 4.30%, respectively. While the area overhead of EOSwrite is only 7.27%.
Wei Zhao 0034, Jie Xu 0013, Xueliang Wei, Bing Wu 0001, Chengning Wang, Weilin Zhu, Wei Tong 0001, Dan Feng 0001, Jingning Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 ZNSKV: Reducing Data Migration in LSMT-Based KV Stores on ZNS SSDs
abstract
Zoned Namespace Solid State Drives (ZNS SSDs) delegate the data placement and garbage collection (GC) to the host It can provide predictable performance, and stable bandwidth for Log-Structured Merge Tree (LSMT) based key-value (KV) stores. We observe redundant data migration in GC and LSMT compaction, which exacerbates the write amplification (WA) of the KV store. In this paper, we propose our ZNS SSD based KV store ZNSKV to reduce the redundant data migration by two major components. First, we design the Compaction-GC scheme to reduce data migration by completing part of the GC in compaction. Second, the Compaction-GC scheme adjusts the file selection strategy according to the free capacity of storage devices, improving system performance while maintaining a high space utilization. Compared to mainstream LSMT-based KV stores with greedy GC, ZNSKV improves throughput by 32%, space utilization to 2.7 times, and lowers WA by 34% under write-intensive workloads.
Denghui Wu, Biyong Liu, Wei Zhao 0034, Wei Tong 0001
ICCD3
2021 Improving the energy efficiency of STT-MRAM based approximate cache
abstract
Approximate computing applications lead to large energy consumption and performance demand for the memory system. However, traditional SRAM based cache cannot satisfy these demands due to high leakage power and limited density. Spin Transfer Torque Magnetic RAM (STT-MRAM) is a promising candidate of cache due to low leakage power and high density. However, STT-MRAM suffers from high write energy. To leverage the ability of tolerating acceptable quality loss via approximations to data, we propose an STT-MRAM based APProximate cache architecture (APPcache) to write/read approximate data thus largely reducing energy. We find many similar elements (e.g. pixels in images) existing in cache lines while running approximate computing applications. Therefore, APPcache uses several lightweight similarity-based encoding schemes to eliminate the similar elements to reduce the data size thus reducing the write energy of STT-MRAM based cache. Besides, we design a software interface to manually control the output quality. APPcache can significantly eliminate similar elements, thus improving energy efficiency. Experimental results show that our scheme can reduce write energy and improve the image raw data compression ratio by 21.9% and 38.0% compared with the state-of-the-art scheme with 1 % error rate, respectively. As for the output quality, the losses of all benchmarks are within 5% with 1 % error rate.
Wei Zhao 0034, Wei Tong 0001, Dan Feng 0001, Jingning Liu, Zhangyu Chen, Jie Xu 0013, Bing Wu 0001, Chengning Wang, Bo Liu 0057
DATE1
2021 MORE2: Morphable Encryption and Encoding for Secure NVM
abstract
Memory encryption can enhance the security of Non-volatile memories (NVMs), but it significantly increases the data bits written to NVMs and leads to severe lifetime and performance degradation. Current encryption techniques aim to reduce the re-encryption to many existing clean words, which unfortunately suffer from high encryption overheads (i.e. latency and energy) and many unnecessary writes. In the meantime, compression techniques can reduce the writes of encrypted NVM. However, we find that they may destroy the data patterns and increase the modified words, resulting in many encryptions in secure NVM. In this paper, we propose the MORphable Encryption and Encoding (MORE2) scheme to address these problems. Our MORphable Encryption (MORE) technique aims to reduce the full-line re-encryption and avoid clean line encryption. Besides, MORE proposes a prediction-based write scheme to avoid the encryption of clean lines, and pre-encrypt the lines that are predicted as dirty. Therefore, MORE can remove the encryption from the critical path of NVM. Furthermore, MORE2proposes the Morphable Selective Encoding (MSE) scheme to compress the modified words while preserving clean words. MORE2encrypts all metadata with the line counter to guarantee high security. Experimental results show that MORE2reduces the bit flips of encrypted NVM by 53.5 %, decreases the access latency by 27.32%, improves the IPC performance by 12.1 %, and reduces the write energy by 29.1 % compared with the state-of-the-art design.
Wei Zhao 0034, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Jie Xu 0013, Gaoxiang Xu, Yiran Chen 0001
ICCAD1
2021 Improving Write Performance on Cross-Point RRAM Arrays by Leveraging Multidimensional Non-Uniformity of Cell Effective Voltage
abstract
Resistive cross-point memory arrays can be used to construct high-density storage-class memory. However, coupled IR drop and sneak currents cause multidimensional non-uniformity of cell effective voltage in cross-point arrays. The voltage non-uniformity significantly degrades write performance on cross-point memory if only adopting the worst-case write latency at partial dimensions. Furthermore, the non-uniformity of cell effective voltage in cross-point arrays depends on multidimensional dynamic write operation parameters: row, column as well as layer address, the number of selected cells, and the number of half-selected low-resistance state cells. In this article, we aim to improve the write performance by leveraging multidimensional non-uniformity of cell effective voltage. First, we analyze the impact of multidimensional write parameters on effective voltage and write latency. Then, we design the memory array write scheme that measures the write parameters and sets the write latency accordingly. We further analyze the features and effects of interlayer sneak currents and extend the scheme to 3D cross-point memory. The evaluation shows that the proposed memory array write scheme can reduce the memory access latency by 75.6 and 64.1 percent, and improve the system performance by 4.5 times and 3.4 times on average, compared with the baseline and the state-of-the-art approach, respectively.
Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Yang Zhang 0051, Yiran Chen 0001
IEEE Trans. Computers6
2021 Improving Multilevel Writes on Vertical 3-D Cross-Point Resistive Memory
abstract
Resistive memory is promising to be constructed as a high-density storage-class memory. Multilevel cell, access-transistor-free cross-point array structure, and 3-D array integration are three approaches to scale up the density of resistive memory. However, composing the three approaches together strengthens the interactions between array-level and cell-level nonidealities (interconnect resistance-induced IR drop, sneak current, and device variability) of resistive memory arrays during write operations and significantly degrades write performance and reliability. In this article, we analyze the dynamic voltage-dividing effect along a selected write current path in 3-D cross-point memory arrays. We propose a nonideality-tolerant high-density resistive memory (HD-RRAM) architecture, that can weaken the interactions between nonidealities and mitigate their degradation effects on the performance and reliability of array multilevel write operations. HD-RRAM is equipped with a double-transistor array architecture with two-transistor- n-resistor (2TnR) cell organization along pillars to reduce the current driving requirement and the large undesired voltage drop across each vertical pillar access transistor. Moreover, multiside asymmetric bias improves the resistive switching velocity by leveraging current-dividing effects. Variability-aware multilevel state partition reduces the worst-case write error rate by leveraging target state dependency of variability. Proportional-control multilevel state tuning reduces the average number of required write-and-verify iterations by leveraging pulse amplitude dependency of variability. Multilevel cell parallel writing improves the cell-level parallelism by leveraging the pass-through feature of intermediate resistance states. The evaluations show that HD-RRAM reduces both memory access latency and energy consumption over an aggressive baseline.
Chengning Wang, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Linghao Song, Yang Zhang 0051, Jie Xu 0013, Xueliang Wei, Yiran Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2020 A Low Power Reconfigurable Memory Architecture for Complementary Resistive Switches
abstract
Memristive crossbar array suffers from severe sneak currents that incur reliability issues and extra energy waste. Complementary resistive switches (CRSs) provide a new concept to address the sneak-current problem. But the destructive read of CRS results in an additional recovery write operation, which strongly restricts its further promotion. Exploiting the dual CRS/memristor mode of CRS devices, we propose Aliens, a novel reconfigurable architecture that introduces one alien cell (memristor mode) for each bitline in the crossbar. Aliens draws advantages from both modes: restrained sneak currents of the CRS mode and nondestructive read of the memristor mode. The simple and regular cell mode organization one bitline one memristor (OBOM) of Aliens enables an energy-saving read method. Further, by exploiting memory access locality, an effective mode switching strategy called Lazy-Switch is proposed to delay and merge the recovery write operations of the CRS mode. Moreover, an 1TnR crossbar structure is adopted to enable larger crossbar arrays as well as a higher ratio of memristor mode cells without going against the OBOM rule. The effects of the memristor mode cell ratio on the energy consumption, endurance, and access performance are studied. Also, we show the bank architecture of Aliens and analyze how to extend our designs to 3-D arrays. Due to fewer recovery write operations and negligible sneak currents, Aliens achieves improvements in energy, overall endurance, and access performance. The experimental results show that our design offers average energy savings of 19.1× compared with memristor-only memory, a memory lifetime 10.7× longer than CRS-only memory, and a competitive performance compared with memristor-only memory.
Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Wei Zhao 0034, Yang Zhang 0051
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2019 ReRAM Crossbar-Based Analog Computing Architecture for Naive Bayesian Engine
abstract
Recent advances in Resistive RAM (ReRAM) have explored the in-situ Matrix-Vector Multiplication (MVM) ability of crossbar arrays to achieve high energy-efficiency Process-In-Memory (PIM) architectures for Convolutional Neural Network (CNN), image processing, and so on. However, the existing ReRAM-based PIM architectures suffer from considerable additional auxiliary logic and device variations. In this work, we propose a novel analog computing architecture NB Engine for classification by implementing Naive Bayesian (NB) algorithm on ReRAM crossbar arrays. The two key steps of the NB algorithm, that is, probability calculation and electing the class that has the highest probability, are elaborately accomplished in our architecture. The ReRAM arrays are both used as storage and computation components. We store the pre-calculated prior probabilities and conditional probabilities of every class in crossbar arrays. Then the probability calculation step is completed in parallel through the MVM operation of the array. In general, the election step is a multiple-comparison procedure and is normally implemented by a comparison tree. Here, we reuse the max pooling module in a conventional CNN PIM architecture to realize a compatible comparison logic. However, neither of the two designs can avoid the overhead of costly high bit-precision Analog-to-Digital Converters (ADCs). So we introduce a novel analog parallel comparison design which does not need any ADCs or other computing logic with better energy-saving and area-efficiency. Our proposed NB Engine is tested by 11 various datasets. The influence of several non-ideal device properties is discussed and the NB Engine exhibits great tolerance to these variations. The experiment results show that our design offers a runtime speedup up to 2289.6x compared with the software-implemented NB classifier with negligible accuracy loss. In addition, the NB Engine saves 96.2% energy consumption and 45.2% array area compared with the CNN PIM compatible design.
Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Wei Zhao 0034, Mengye Peng
ICCD6
2019 Cross-point Resistive Memory: Nonideal Properties and Solutions
abstract
Emerging computational resistive memory is promising to overcome the challenges of scalability and energy efficiency that DRAM faces and also break through the memory wall bottleneck. However, cell-level and array-level nonideal properties of resistive memory significantly degrade the reliability, performance, accuracy, and energy efficiency during memory access and analog computation. Cell-level nonidealities include nonlinearity, asymmetry, and variability. Array-level nonidealities include interconnect resistance, parasitic capacitance, and sneak current. This review summarizes practical solutions that can mitigate the impact of nonideal device and circuit properties of resistive memory. First, we introduce several typical resistive memory devices with focus on their switching modes and characteristics. Second, we review resistive memory cells and memory array structures, including 1T1R, 1R, 1S1R, 1TnR, and CMOL. We also overview three-dimensional (3D) cross-point arrays and their structural properties. Third, we analyze the impact of nonideal device and circuit properties during memory access and analog arithmetic operations with focus on dot-product and matrix-vector multiplication. Fourth, we discuss the methods that can mitigate these nonideal properties by static parameter and dynamic runtime co-optimization from the viewpoint of device and circuit interaction. Here, dynamic runtime operation schemes include line connection, voltage bias, logical-to-physical mapping, read reference setting, and switching mode reconfiguration. Then, we highlight challenges on multilevel cell cross-point arrays and 3D cross-point arrays during these operations. Finally, we investigate design considerations of memory array peripheral circuits. We also portray an unified reconfigurable computational memory architecture.
Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zheng Li 0005, Jiayi Chang, Yang Zhang 0051, Bing Wu 0001, Jie Xu 0013, Wei Zhao 0034, Ruoxi Ren
ACM Trans. Design Autom. Electr. Syst.10