EDBT 2026 Demo / reviewers in the wild / expert
Bing Wu 0001
dblp:75/1956-1
· DBLP profile ↗
25ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-5828-8399ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 5 first-author · 16 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data Distribution-Aware Analog/Digital Conversion Strategy for Energy-Efficient Memristive In-Situ AcceleratorsabstractMemristive in-situ computing offers energy-efficient DNN acceleration, but faces ADC-induced energy bottlenecks. We observe that bitline outputs exhibit significant non-uniformity and cycle-to-cycle variation, rendering conventional A/D conversion schemes suboptimal. We thus propose a data distribution-aware A/D conversion strategy that predicts key bits of digital outputs and skips unnecessary steps, with a switching mechanism adapting the optimal conversion method across cycles. Implemented via a reconfigurable SAR-ADC, our approach significantly reduces the energy consumption of in-situ accelerators. Taoming Lei, Bing Wu 0001, Wei Tong 0001, Dan Feng 0001 |
DATE | 3 |
| 2026 | An IR drop-robust Mapping Method for Reliable Memristive AcceleratorsabstractMemristive accelerators (MAs) facilitate efficient matrix-vector multiplication (MVM) by performing in situ computation within memory crossbar arrays, thereby ensuring a fast and energy-efficient application acceleration. A significant challenge associated with the MA lies in the limited computing accuracy caused by the IR drop effect. However, existing IR drop mitigation works provide an approximate compensation, resulting in less accurate results. In this paper, we propose an IR drop-robust mapping method for reliable memristive accelerators. Firstly, the IR drop-robust mapping (IRM) method exploits the residuals between the equivalent matrix after the IR drop effect and the original matrix, and iteratively maps them to the crossbars for IR drop compensation. Based on the IRM method, a novel mechanism of the matrix-vector multiplication (MVM) operation is derived, ensuring that MVM is computed correctly. Secondly, the Calibrate-Shift-Reflect (CSR) strategy is developed to significantly reduce the number of arrays required by the IRM method to map the residuals. Thirdly, the hardware support for the IRM method is designed, and the overhead is reduced by sharing drivers/selectors between neighboring arrays. The experimental results indicate that the IRM-CSR method can effectively mitigate the IR drop effect, restoring inference accuracy by at most 80% (for neural network applications), and achieving a reduction in the relative root-mean-squared error by 103×~1010× (for scientific computing), compared with the state-of-the-art methods. Shiyi Song, Bing Wu 0001, Huan Cheng, Xueliang Wei, Wei Tong 0001, Dan Feng 0001 |
DATE | 3 |
| 2026 | pTree: Building Efficient B${}^{+}$+-Tree on Non-Volatile Memory With Processing-in-MemoryabstractB+-Trees are widely used in storage systems and diverse applications. However, their performance is hindered by frequent memory accesses required for key comparisons and structural maintenance. The limited parallelism of CPUs, along with their sensitivity to data order and volume, further restricts the efficiency of B+-Trees. Processing-in-memory (PIM) offers a promising alternative with its in-situ parallel computing capability. In this paper, we propose pTree, a novel set of PIM techniques and architecture specifically designed to accelerate B+-Tree operations. pTree introduces in-situ parallel comparison mechanisms that significantly reduce costly memory accesses and inefficient CPU-side comparisons during tree traversal. This parallelism eliminates the need to maintain intra-node key order and, together with our in-situ parallel node bisecting techniques, greatly minimizes tree structure maintenance overhead during insertions and deletions. Additionally, pTree decouples both comparison and modification efficiency from node size, enabling the use of larger nodes to reduce tree height and traversal complexity. Evaluation shows that pTree reduces the average latency to 35%, 26%, 52%, and 32% compared to state-of-theart B+-Trees forInsert, Search, Update,andDeleteoperations. Bing Wu 0001, Shiyi Song, Xueliang Wei, Huan Cheng, Wei Tong 0001, Dan Feng 0001 |
IEEE Trans. Computers | 2 |
| 2025 | COVER: Alleviating Crash-Consistency Error Amplification in Secure Persistent Memory SystemsabstractData security (including confidentiality, integrity, and availability) and crash consistency guarantees are essential for building trusted persistent memory (PM) systems. Security and consistency metadata are added to enable the guarantees. Recent studies show that errors in security metadata have the amplified effect, which significantly affects data availability. However, the impact of consistency metadata errors on data availability has rarely been discussed. We identify the crash-consistency error amplification (CCEA) problem, several errors in consistency metadata can make a large portion of data in PM possibly inconsistent. The error sensitivity of consistency metadata is higher than data and security metadata, thus requiring special attention. It is inefficient to address this problem by using the methods that are proposed to alleviate the amplified effect of security metadata errors, because security metadata are generally designed for a single purpose (e.g., integrity verification), while consistency metadata are designed for multiple purposes, including inconsistency locating and recovery. To effectively and efficiently alleviate the CCEA problem, we propose a c rash c o nsistency ver ification approach (COVER) that decouples inconsistency locating and recovery. COVER provides three design options that support different tradeoffs between effectiveness and efficiency. Experimental results show that COVER effectively alleviates the problem with only about 1.0% performance degradation on average compared with the state-of-the-art secure PM design. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Bing Wu 0001, Xu Jiang 0005 |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | SEED: Speculative Security Metadata Updates for Low-Latency Secure MemoryabstractSecuring systems’ main memory is important for building trusted data centers. To ensure memory security, encryption and integrity verification techniques update the security metadata (e.g., encryption counters and integrity trees) during memory data writes. Existing studies are optimistic about the effect of data writes on system performance since they regard all data writes as background operations. However, we show that security metadata updates significantly increase data write latency. High-latency data writes frequently fill up write buffers in the system, forcing the system to perform the writes in the critical path. As a result, performance-critical data reads need to wait for the execution of these writes, which increases data read latency and degrades system performance. In this paper, we propose SEED that improves the performance of secure memory systems by speculatively updating security metadata in the background before data writes arrive. To enable speculative updates, SEED predicts which dirty cache lines will be written to memory through natural evictions. We find that cache evictions depend on multiple factors. To decouple the dependencies for accurate predictions, we devise a two-step eviction prediction method based on our observation that the next eviction victim rarely changes in a set. The first step predicts which cache sets will evict cache lines, while the second step predicts which cache lines will be evicted by finding the next eviction victims in the sets. For predicted evictions, we develop a speculative updater to perform speculative updates. We analyze the invariants that must be followed by the updater to ensure the correctness of speculative updates. The updater rolls back the speculatively updated security metadata of inaccurate predictions. To reduce the rollback overhead, we devise a rollback batching and an update pausing optimization for the updater. Experimental results show that SEED reduces data write latency by 39.8%, data read latency by 44.9%, and improves performance by 40.0% on average compared with the state-of-the-art secure memory design. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Bing Wu 0001, Xu Jiang 0005 |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | FADESIM: Enable Fast and Accurate Design Exploration for Memristive Accelerators Considering NonidealitiesabstractMemristive accelerators (MAs), built with memristive crossbar arrays (MCAs), have gained significant attention for their ability to efficiently perform matrix-vector multiplication in diverse applications. Modeling and simulation are indispensable tools for exploring and evaluating architectural design. Specifically, for MAs, maintaining high computation accuracy has been challenging because of realistic nonidealities like wire resistance (i.e., IR drop), I–V nonlinearity, program variation, and so on. Thus, for the system architects, fast and exact analysis of the effects of nonidealities is highly desirable, especially during the vast early design space exploration of the architecture. However, the SPICE model and existing MCA compact models (CMs) do not offer acceptable speeds for accurate simulation purposes, making them less practical. Additionally, the existing MCA simplified circuit models and predictive models fail to produce effective results due to the complexity of IR drop, let alone the coupling of multiple nonidealities. To enable fast and accurate design exploration for MAs with nonidealities considered, in this article, we propose FADESIM which includes fast IR drop simulation methods and a processing chain for the joint simulation of multiple nonidealities. Starting with the analysis of the accurate CM for IR drop, we explore the special properties of the model to enable fast iterative methods as well as a method to skip invalid calculations. This significantly reduces the time complexity of the simulation from naive$O(n^{6})$to$O(n^{3})$. For less severe IR drop cases, a custom iterative update algorithm is presented for faster simulations as a supplement, specifically with a time complexity of near$O(n^{2})$and proven applicable conditions. To simulate multiple nonidealities, we introduce a processing chain to inject corresponding processing functions before, during, and after the proposed fast IR drop simulation process according to the stages at which nonidealities take effect. The array-level experimental results show that our method achieves accurate simulations, with$19.8 \times - 884.9 \times $faster and$276.7 \times - 8018.8 \times $reduced memory usage compared to SPICE simulation. Further experiments at the algorithm-level demonstrate the effectiveness of our method in assisting architects with evaluating their designs. Bing Wu 0001, Huan Cheng, Xueliang Wei, Wei Tong 0001, Dan Feng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | CMD: A Cache-Assisted GPU Memory Deduplication ArchitectureabstractMassive off-chip accesses in graphics processing units (GPUs) are the main performance bottleneck. We find that many writes are duplicate, and the duplication can beinter-dupandintra-dup. Whileinter-dupmeans different memory blocks are identical, andintra-dupmeans all the 4B elements in a line are the same. In this work, we propose a cache-assisted GPU memory deduplication architecture named cache-assisted GPU memory deduplicated (CMD) to reduce the off-chip accesses via utilizing the data duplication in GPU applications. CMD includes three key design contributions which aim to reduce the three kinds of accesses: 1) a novel GPU memory deduplication architecture that removes theintra-dupandinter-duplines. We design several techniques to manage duplicate blocks, reducing massive off-chip writes; 2) we propose a cache-assisted read scheme to reduce the reads to duplicate data. When an L2 cache miss wants to read the duplicate block, if the reference block has been fetched to L2 and it is clean, we can copy it to the L2 missed block without accessing off-chip DRAM. As for the reads tointra-dupdata, CMD uses the on-chip metadata cache to get the data; and 3) when a cache line is evicted, the clean sectors in the line are invalidated while the dirty sectors are written back. However, most read-only victims are rereferenced from DRAM more than twice. Therefore, we add a full-associate FIFO to accommodate the read-only (it is also clean) victims to reduce the rereference counts. Experiments show that CMD can decrease the off-chip accesses by 31.01%, reduce the energy by 32.78% and improve performance by 42.53%. Besides, CMD can improve the performance of memory-intensive workloads by 57.56%. Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Xueliang Wei, Bing Wu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | DRCTL: A Disorder-Resistant Computation Translation Layer Enhancing the Lifetime and Performance of Memristive CIM ArchitectureabstractThe memristive Computing-in-Memory (CIM) sys-tem can efficiently accelerate matrix-vector multiplication (MVM) operations through in-situ computing. The data layout has a significant impact on the communication performance of CIM systems. Existing software-level communication opti-mizations aim to reduce communication distance by carefully designing static data layouts, while wear-leveling (WL) and error mitigation methods use dynamic scheduling to enhance system reliability, resulting in randomized data layouts and increased communication overhead. Besides, existing CIM compilers di-rectly map data to physical crossbars and generate instructions, which causes inconvenience for dynamic scheduling. To address these challenges of balancing communication performance and reliability while coordinating existing CIM compilers and dy-namic scheduling, we propose a disorder-resistant computation translation layer (DRCTL), which improves system lifetime and communication performance through co-optimization of data layout and dynamic scheduling. It consists of three parts: (1) We propose an address conversion method for dynamic scheduling, which updates the addresses in the instruction stream after dynamic scheduling, thereby avoiding recompilation. (2) Dynamic scheduling strategy for reliability improvement. We propose a hierarchical wear-leveling (HWL) strategy, which reduces communication by increasing scheduling granularity. (3) Communication optimization for dynamic scheduling. We propose data layout-aware selective remapping (LASR), which helps dynamic scheduling methods improve communication lo-cality and reduce latency by exploiting data dependencies. The experiments demonstrate that HWL extends lifetime by 100.3-205.9 x compared to not using WL. Even with a slight lifetime decrease compared to the state-of-the-art WL (TIWL), it still supports continuous neural network training for 7 years. After applying LASR to HWL, the number of execution cycles, energy consumption, on-chip and off-chip NoC accesses decrease by an average of 26.91 %, 26.88%, 36.41 %, and 80.62%, respectively. Bing Wu 0001, Huan Cheng, Taoming Lei, Dan Feng 0001, Wei Tong 0001 |
MICRO | 2 |
| 2023 | ODLPIM: A Write-Optimized and Long-Lifetime ReRAM-Based Accelerator for Online Deep LearningabstractReRAM-based Processing-In-Memory (PIM) architectures have demonstrated high energy efficiency and performance in deep neural network (DNN) acceleration. Most of the existing PIM accelerators for DNN focus on offline batch learning (OBL) which requires the whole dataset to be available before training. However, in the real world, data instances arrive in sequential settings, and even the data pattern may change, which calls concept drift. OBL requires expensive retraining to solve concept drift, whereas online deep learning (ODL) is evidenced to be a better solution to keep the model evolving over streaming data. Unfortunately, when ODL optimizes models over a large-scale data stream in the PIM system, unbalanced writes are more severe than OBL, due to the heavier weight updates, resulting in the amplification of unbalanced writes and lifetime deterioration. In this work, we propose ODLPIM, an online deep learning PIM accelerator that extends the system lifetime through algorithm-hardware co-optimization. ODLPIM adopts a novel write-optimized parameter update (WARP) scheme that reduces the non-critical weight updates in hidden layers. Besides, a table-based inter-crossbar wear-leveling (TIWL) scheme is proposed and applied to the hardware controller to achieve wear-leveling between crossbars for lifetime improvement. Experiments show that WARP reduces weight updates on average to 15.25% and up to 24% compared to that without WARP, and eventually prolongs system lifetime on average to 9.65% and up to 26.81%, with a negligible rise in cumulative error rate (up to 0.31%). By combining WARP with TIWL, the lifetime of ODLPIM is improved by an average of$\mathbf{12}.\mathbf{59}\times$and up to$\mathbf{17}.\mathbf{73}\times$. Bing Wu 0001, Huan Cheng, Wei Zhao 0034, Xueliang Wei, Dan Feng 0001, Wei Tong 0001 |
DATE | 2 |
| 2023 | ICON: An IR Drop Compensation Method at OU Granularity with Low Overhead for eNVM-based AcceleratorsabstractProcessing at operating unit (OU) granularity can alleviate the program variation effect and ADC conversion overhead in emerging non-volatile memory (eNVM) based accelerators. However, our experiments show that the IR drop effect can severely decrease computing accuracy when processing at OU granularity. Moreover, the IR drop effect on the entire array differs from the IR drop effect on OUs, meaning compensating at OU granularity is necessary. We also notice that the IR drop effect differs among OUs, and previous IR drop mitigation methods introduce more latency, area, and power overhead to adapt to these differences. Compensation modules from their methods calibrated for one OU do not apply to other OUs and need to be configured for each OU compensation using configuration modules. This paper proposes ICON, an IR drop compensation method at OU granularity with low overhead for eNVM-based accelerators. In order to decrease compensation latency, area, and power overhead, we perform several optimizations. First, the designed compensation circuit is simplified and does not compensate for the IR drop effect caused by parasitic resistances inside an OU. This simplification is based on our observation that the parasitic wire resistances inside an OU can be ignored using the OU size mentioned in previous works. Second, the compensation circuit is designed without the help of configuration circuits. We take the IR drop differences among OUs as input parameters of the compensation circuit so that it can apply to all OUs. Furthermore, the compensation circuit is pipelined into six stages to increase throughput. Experiments show that our compensation method can overcome the IR drop problem when processing in the eNVM-based crossbar array at OU granularity, with 1.3× ~ 13× lower latency, 1.5× ~ 33.1× lower area, and 1.4× ~ 8.4× lower power overhead compared with state-of-the-art methods. Wei Tong 0001, Bing Wu 0001, Huan Cheng, Chengning Wang |
ICCD | 3 |
| 2023 | APPcache+: An STT-MRAM-Based Approximate Cache System With Low Power and Long LifetimeabstractDue to high static power and low scalability, the traditional SRAM-based cache is not a good solution for image processing applications. Emerging spin transfer torque magnetic RAM (STT-MRAM) is a promising candidate for cache due to its low leakage power and high density. However, STT-MRAM suffers from high write energy. Therefore, by making use of the ability of tolerating minor errors in image processing applications, this work presents an STT-MRAM-basedAPProximatecachearchitecture (APPcache+) to write/read approximate data, which can largely reduce the cache energy and improve the STT-MRAM lifetime. APPcache+ includes three main designs. First, we find that there are many similar elements (e.g., pixels in images) in cache lines. Therefore, APPcache+ presents several lightweight similarity-based encoding techniques to remove redundant elements, thus, shortening the data size and reducing the energy of STT-MRAM cache. Second, we design a partial read scheme to reduce the read energy of the STT-MRAM cache. In the traditional decompression process, the whole line is fetched into the decompressor, leading to unnecessary read energy. The partial read scheme can largely reduce read energy while keeping the overhead low. Third, we observe the encoding schemes may lead to bit write imbalance. Therefore, we propose a lightweight Ping-Pong intraline wear-leveling scheme to improve the lifetime. Compared with the baseline, extensive evaluation results show that our APPcache+ can largely reduce the overall energy by 32.58%, improve lifetime by 40.7% with only 2.2% performance degradation, and 1.86% output quality loss. Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zhangyu Chen, Bing Wu 0001, Chengning Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | A Low-Latency and High-Endurance MLC STT-MRAM-Based Cache SystemabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising cache memory candidate due to its high density, low leakage power, and nonvolatility. Multilevel cell (MLC) STT-MRAM can further increase density by storing 2 bits in one cell’s hard and soft domain, respectively. However, MLC STT-MRAM suffers two-step write, leading to high write energy, long latency, and severe lifetime degradation. Current encoding techniques propose to encode the new data to reduce the two-step data writes. However, they have two weaknesses: 1) high area overhead, e.g., recent work TSE (Hsieh et al., 2020) needs extra 37.5% MLCs and 2) prolong the write latency due to an extra read. Therefore, we propose enhanced one-step write (EOSwrite) to write data in one step. EOSwrite includes line bypassing and four intraline encoding techniques. Line bypassing schemes can bypass the writes to zero or clean lines, leading to low write/read latency. As for the intraline techniques, we propose four write modes. They utilize the data patterns and the clean data in cache lines to write data in one step, therefore reducing the data write latency. The key idea of one-step write is to write as much data as possible in the soft domain of MLC STT-MRAM. EOSwrite can greatly relieve the weaknesses of the current encoding schemes. Evaluation results show that EOSwrite can improve the lifetime of MLC STT-MRAM by 56.96%, reduce dynamic energy by 33.95%, reduce access latency by 36.95%, and improve system performance of MLC STT-MRAM by 4.30%, respectively. While the area overhead of EOSwrite is only 7.27%. Wei Zhao 0034, Jie Xu 0013, Xueliang Wei, Bing Wu 0001, Chengning Wang, Weilin Zhu, Wei Tong 0001, Dan Feng 0001, Jingning Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Space-Time-Efficient Modeling of Large-Scale 3-D Cross-Point Memory Arrays by Operation Adaption and Network CompactionabstractThree-dimensional (3-D) integrated cross-point memory arrays can be used to build high-density storage-class memory systems. However, the coupled network topology caused by sharing word lines or bit lines between adjacent memory layers significantly enlarges the memory space overhead and time cost of memory operation simulations on mega-scale 3-D cross-point memory arrays. We observe that different components of the 3-D cross-point array have different contribution significance to the key metrics of write or read operations, and the distribution patterns of principal components in the arrays are different for write and read operations. We propose an operation-adaptive array modeling framework that exploits the impact of the applied operation on the distribution of array principal components for 3-D cross-point memory arrays. Based on the modeling framework, we propose two array network compaction methods for efficient write and read operation simulations on 3-D cross-point arrays, respectively: 1) pruning zero-biased unselected cells and 2) group-merging neighboring half-selected cells with a specific granularity. Also, serially connected line segments are merged, and floating segments are deleted. Evaluations show that the proposed methods significantly reduce the memory space overhead and time cost of memory operation simulations for various layer sizes and different cell-level access parallelism in an array. Besides, the proposed methods can efficiently simulate memory operations on multilayered 3-D cross-point memory arrays with up to$4096\times 4096$layer size under the 16-GB memory space constraint, achieving a 64-times improvement in layer size that can be simulated compared with the conventional complete 3-D cross-point array network model. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Bing Wu 0001, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Improving the energy efficiency of STT-MRAM based approximate cacheabstractApproximate computing applications lead to large energy consumption and performance demand for the memory system. However, traditional SRAM based cache cannot satisfy these demands due to high leakage power and limited density. Spin Transfer Torque Magnetic RAM (STT-MRAM) is a promising candidate of cache due to low leakage power and high density. However, STT-MRAM suffers from high write energy. To leverage the ability of tolerating acceptable quality loss via approximations to data, we propose an STT-MRAM based APProximate cache architecture (APPcache) to write/read approximate data thus largely reducing energy. We find many similar elements (e.g. pixels in images) existing in cache lines while running approximate computing applications. Therefore, APPcache uses several lightweight similarity-based encoding schemes to eliminate the similar elements to reduce the data size thus reducing the write energy of STT-MRAM based cache. Besides, we design a software interface to manually control the output quality. APPcache can significantly eliminate similar elements, thus improving energy efficiency. Experimental results show that our scheme can reduce write energy and improve the image raw data compression ratio by 21.9% and 38.0% compared with the state-of-the-art scheme with 1 % error rate, respectively. As for the output quality, the losses of all benchmarks are within 5% with 1 % error rate. Wei Zhao 0034, Wei Tong 0001, Dan Feng 0001, Jingning Liu, Zhangyu Chen, Jie Xu 0013, Bing Wu 0001, Chengning Wang, Bo Liu 0057 |
DATE | 7 |
| 2021 | Improving Write Performance on Cross-Point RRAM Arrays by Leveraging Multidimensional Non-Uniformity of Cell Effective VoltageabstractResistive cross-point memory arrays can be used to construct high-density storage-class memory. However, coupled IR drop and sneak currents cause multidimensional non-uniformity of cell effective voltage in cross-point arrays. The voltage non-uniformity significantly degrades write performance on cross-point memory if only adopting the worst-case write latency at partial dimensions. Furthermore, the non-uniformity of cell effective voltage in cross-point arrays depends on multidimensional dynamic write operation parameters: row, column as well as layer address, the number of selected cells, and the number of half-selected low-resistance state cells. In this article, we aim to improve the write performance by leveraging multidimensional non-uniformity of cell effective voltage. First, we analyze the impact of multidimensional write parameters on effective voltage and write latency. Then, we design the memory array write scheme that measures the write parameters and sets the write latency accordingly. We further analyze the features and effects of interlayer sneak currents and extend the scheme to 3D cross-point memory. The evaluation shows that the proposed memory array write scheme can reduce the memory access latency by 75.6 and 64.1 percent, and improve the system performance by 4.5 times and 3.4 times on average, compared with the baseline and the state-of-the-art approach, respectively. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Yang Zhang 0051, Yiran Chen 0001 |
IEEE Trans. Computers | 5 |
| 2021 | Improving Multilevel Writes on Vertical 3-D Cross-Point Resistive MemoryabstractResistive memory is promising to be constructed as a high-density storage-class memory. Multilevel cell, access-transistor-free cross-point array structure, and 3-D array integration are three approaches to scale up the density of resistive memory. However, composing the three approaches together strengthens the interactions between array-level and cell-level nonidealities (interconnect resistance-induced IR drop, sneak current, and device variability) of resistive memory arrays during write operations and significantly degrades write performance and reliability. In this article, we analyze the dynamic voltage-dividing effect along a selected write current path in 3-D cross-point memory arrays. We propose a nonideality-tolerant high-density resistive memory (HD-RRAM) architecture, that can weaken the interactions between nonidealities and mitigate their degradation effects on the performance and reliability of array multilevel write operations. HD-RRAM is equipped with a double-transistor array architecture with two-transistor- n-resistor (2TnR) cell organization along pillars to reduce the current driving requirement and the large undesired voltage drop across each vertical pillar access transistor. Moreover, multiside asymmetric bias improves the resistive switching velocity by leveraging current-dividing effects. Variability-aware multilevel state partition reduces the worst-case write error rate by leveraging target state dependency of variability. Proportional-control multilevel state tuning reduces the average number of required write-and-verify iterations by leveraging pulse amplitude dependency of variability. Multilevel cell parallel writing improves the cell-level parallelism by leveraging the pass-through feature of intermediate resistance states. The evaluations show that HD-RRAM reduces both memory access latency and energy consumption over an aggressive baseline. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Linghao Song, Yang Zhang 0051, Jie Xu 0013, Xueliang Wei, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | DualFS: A Coordinative Flash File System with Flash Block Dual-mode SwitchingabstractAs users' demand for large-capacity storage continues to grow, the widely used NAND flash memory always adopts multi-bit per cell and 3D-stacking technology to improve the storage density however impairing flash performance. Modern flash chips allow flash blocks to switch between multi-bit per cell and one-bit per cell, which gives the potential to benefit from the high performance of Single-Level Cells (SLC, i.e. one-bit per cell). However, the semantic gap caused by the block I/O interface and flash translation layer (FTL) prevents the performance of the underlying flash memory from being fully utilized. In this paper, we propose a log-structured file system, called DualFS. DualFS allows flash blocks to switch between the original-mode and SLC-mode in a free manner and uses these SLC blocks to accelerate critical requests. DualFS dynamically adjusts the capacity of the SLC-mode area according to the total size of valid data to maintain the device a negligible capacity degradation. Besides, DualFS draws advantages from the Open-Channel solid-state disk (SSD) by using software semantic information to guide data placement and management. In particular, since writing the same amount of data to the SLC block will impair more endurance, DualFS proposes a novel life management scheme which throttles the ratio of writing to SLC-mode area within a certain window, thereby accurately control the endurance of SSD. Last, DualFS integrates the garbage collection (GC) procedure, which is directly driven by the file system and merges the GC of the SLC-mode area and the original-mode area. The experiment results show that DualFS offers an average 22.2% performance improvement than the state of art flash file system based on the open-channel SSD. Compared with other storage systems based on flash block dual-mode switching feature, the average performance of DualFS also exhibits considerable improvements by 13.3% to 24.6%. Moreover, DualFS effectively reduces the read response time under different read/write ratios and precisely manage the device endurance. Bing Wu 0001, Mengye Peng, Dan Feng 0001, Wei Tong 0001 |
ICCD | 1 |
| 2020 | A Low Power Reconfigurable Memory Architecture for Complementary Resistive SwitchesabstractMemristive crossbar array suffers from severe sneak currents that incur reliability issues and extra energy waste. Complementary resistive switches (CRSs) provide a new concept to address the sneak-current problem. But the destructive read of CRS results in an additional recovery write operation, which strongly restricts its further promotion. Exploiting the dual CRS/memristor mode of CRS devices, we propose Aliens, a novel reconfigurable architecture that introduces one alien cell (memristor mode) for each bitline in the crossbar. Aliens draws advantages from both modes: restrained sneak currents of the CRS mode and nondestructive read of the memristor mode. The simple and regular cell mode organization one bitline one memristor (OBOM) of Aliens enables an energy-saving read method. Further, by exploiting memory access locality, an effective mode switching strategy called Lazy-Switch is proposed to delay and merge the recovery write operations of the CRS mode. Moreover, an 1TnR crossbar structure is adopted to enable larger crossbar arrays as well as a higher ratio of memristor mode cells without going against the OBOM rule. The effects of the memristor mode cell ratio on the energy consumption, endurance, and access performance are studied. Also, we show the bank architecture of Aliens and analyze how to extend our designs to 3-D arrays. Due to fewer recovery write operations and negligible sneak currents, Aliens achieves improvements in energy, overall endurance, and access performance. The experimental results show that our design offers average energy savings of 19.1× compared with memristor-only memory, a memory lifetime 10.7× longer than CRS-only memory, and a competitive performance compared with memristor-only memory. Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Wei Zhao 0034, Yang Zhang 0051 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Accelerating garbage collection for 3D MLC flash memory with SLC blocksabstract3D MLC NAND Flash is more appreciated for its massive capacity and significant performance. It's common to configure a portion of flash blocks to SLC-mode to further shorten the requests latency at a small cost of capacity. However, the limited SLC-mode blocks provoke GC (Garbage Collection) procedures more frequently and the GC penalty is heavier for 3D Flash than that for 2D Flash. As the block of 3D Flash consists of much more pages, the increment in the block size prolongs the erase operation latency and increases the number of migrated pages during GC. Existing works focus on reducing the number of migrated pages with sub-block GC strategies, which deeply depends on the distribution of valid pages across the sub-blocks in the victim block. In this paper, we propose a set of schemes called DCD which consists of Dual-mode Handler, Compensated GC and Dynamic Data Distribution. The key idea is to exploit the shorter operation delays of SLC-mode to accelerate the valid pages' migrations during GC, which doesn't rely on the distribution of valid pages inside the victim block. Experimental results show that compared with state-of-the-art designs, DCD shortens the average read response time by 19.8% and 37.3% for block-level and page-level FTL, respectively. Wei Tong 0001, Jingning Liu, Bing Wu 0001, Yazhi Feng |
ICCAD | 4 |
| 2019 | ReRAM Crossbar-Based Analog Computing Architecture for Naive Bayesian EngineabstractRecent advances in Resistive RAM (ReRAM) have explored the in-situ Matrix-Vector Multiplication (MVM) ability of crossbar arrays to achieve high energy-efficiency Process-In-Memory (PIM) architectures for Convolutional Neural Network (CNN), image processing, and so on. However, the existing ReRAM-based PIM architectures suffer from considerable additional auxiliary logic and device variations. In this work, we propose a novel analog computing architecture NB Engine for classification by implementing Naive Bayesian (NB) algorithm on ReRAM crossbar arrays. The two key steps of the NB algorithm, that is, probability calculation and electing the class that has the highest probability, are elaborately accomplished in our architecture. The ReRAM arrays are both used as storage and computation components. We store the pre-calculated prior probabilities and conditional probabilities of every class in crossbar arrays. Then the probability calculation step is completed in parallel through the MVM operation of the array. In general, the election step is a multiple-comparison procedure and is normally implemented by a comparison tree. Here, we reuse the max pooling module in a conventional CNN PIM architecture to realize a compatible comparison logic. However, neither of the two designs can avoid the overhead of costly high bit-precision Analog-to-Digital Converters (ADCs). So we introduce a novel analog parallel comparison design which does not need any ADCs or other computing logic with better energy-saving and area-efficiency. Our proposed NB Engine is tested by 11 various datasets. The influence of several non-ideal device properties is discussed and the NB Engine exhibits great tolerance to these variations. The experiment results show that our design offers a runtime speedup up to 2289.6x compared with the software-implemented NB classifier with negligible accuracy loss. In addition, the NB Engine saves 96.2% energy consumption and 45.2% array area compared with the CNN PIM compatible design. Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Wei Zhao 0034, Mengye Peng |
ICCD | 1 |
| 2019 | Cross-point Resistive Memory: Nonideal Properties and SolutionsabstractEmerging computational resistive memory is promising to overcome the challenges of scalability and energy efficiency that DRAM faces and also break through the memory wall bottleneck. However, cell-level and array-level nonideal properties of resistive memory significantly degrade the reliability, performance, accuracy, and energy efficiency during memory access and analog computation. Cell-level nonidealities include nonlinearity, asymmetry, and variability. Array-level nonidealities include interconnect resistance, parasitic capacitance, and sneak current. This review summarizes practical solutions that can mitigate the impact of nonideal device and circuit properties of resistive memory. First, we introduce several typical resistive memory devices with focus on their switching modes and characteristics. Second, we review resistive memory cells and memory array structures, including 1T1R, 1R, 1S1R, 1TnR, and CMOL. We also overview three-dimensional (3D) cross-point arrays and their structural properties. Third, we analyze the impact of nonideal device and circuit properties during memory access and analog arithmetic operations with focus on dot-product and matrix-vector multiplication. Fourth, we discuss the methods that can mitigate these nonideal properties by static parameter and dynamic runtime co-optimization from the viewpoint of device and circuit interaction. Here, dynamic runtime operation schemes include line connection, voltage bias, logical-to-physical mapping, read reference setting, and switching mode reconfiguration. Then, we highlight challenges on multilevel cell cross-point arrays and 3D cross-point arrays during these operations. Finally, we investigate design considerations of memory array peripheral circuits. We also portray an unified reconfigurable computational memory architecture. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zheng Li 0005, Jiayi Chang, Yang Zhang 0051, Bing Wu 0001, Jie Xu 0013, Wei Zhao 0034, Ruoxi Ren |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2018 | Aliens: a novel hybrid architecture for resistive random-access memoryabstractPassive crossbar arrays of resistive random-access memory (RRAM) have shown great potential to meet the demands of future memory. By eliminating transistor per cell, the crossbar array possesses a higher memory density but introduces sneak currents which incur extra energy waste and reliability issues. The complementary resistive switch (CRS), consisting of two anti-serially stacked memristors, is considered as a promising solution to the sneak current problem. However, the destructive read of the CRS results in an additional recovery write operation which strongly restricts its further promotion. Exploiting the dual CRS/memristor mode of CRS devices, we propose Aliens, a novel hybrid architecture for resistive random-access memory which introduces one alien cell (memristor mode) for each wordline in the crossbar to provide a practical hybrid memory without operating system's intervention. Aliens draws advantages from both modes: restrained sneak current of CRS mode and non-destructive read of memristor mode. The simple and regular cell mode organization of Aliens enables an energy-saving read method and an effective mode switching strategy called Lazy-Switch. By exploiting memory access locality, Lazy-Switch delays and merges the recovery write operations of the CRS mode. Due to fewer recovery write operations and negligible sneak currents, Aliens achieves improvement in energy, overall endurance, and access performance. The experiment results show that our design offers average energy savings of 13.9× compared with memristor-only memory, a memory lifetime 5.3× longer than CRS-only memory, and a competitive performance compared with memristor-only memory. Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Mingshun Yang, Chengning Wang, Yang Zhang 0051 |
ICCAD | 1 |
| 2018 | CACF: A Novel Circuit Architecture Co-optimization Framework for Improving Performance, Reliability and Energy of ReRAM-based Main Memory SystemabstractEmerging Resistive Random Access Memory (ReRAM) is a promising candidate as the replacement for DRAM due to its low standby power, high density, high scalability, and nonvolatility. By employing the unique crossbar structure, ReRAM can be constructed with extremely high density. However, the crossbar ReRAM faces some serious challenges in terms of performance, reliability, and energy consumption. First, ReRAM’s crossbar structure causes an IR drop problem due to wire resistance and sneak currents, which results in nonuniform access latency in ReRAM banks and reduces its reliability. Second, without access transistors in the crossbar structure, write disturbance results in serious data reliability problem. Third, the access latency, reliability, and energy use of ReRAM arrays are significantly influenced by the data patterns involved in a write operation. To overcome the challenges of the crossbar ReRAM, we propose a novel circuit architecture co-optimization framework for improving the performance, reliability, and energy use of ReRAM-based main memory system, called CACF. The proposed CACF consists of three levels, including the circuit level, circuit architecture level, and architecture level. At the circuit level, to reduce the IR drops along bitlines, we propose a double-sided write driver design by applying write drivers along both sides of bitlines and selectively activating the write drivers. At the circuit architecture level, to address the write disturbance with low overheads, we propose a RESET disturbance detection scheme by adding disturbance reference cells and conditionally performing refresh operations. At the architecture level, a region partition with address remapping method is proposed to leverage the nonuniform access latency in ReRAM banks, and two flip schemes are proposed in different regions to optimize the data patterns involved in a write operation. The experimental results show that CACF improves system performance by 26.1%, decreases memory access latency by 22.4%, shortens running time by 20.1%, and reduces energy consumption by 21.6% on average over an aggressive baseline. Meanwhile, CACF significantly improves the reliability of ReRAM-based memory systems. Yang Zhang 0051, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Chengning Wang, Bing Wu 0001, Zheng Li 0005, Gaoxiang Xu |
ACM Trans. Archit. Code Optim. | 8 |
| 2017 | A Novel ReRAM-based Main Memory Structure for Optimizing Access Latency and ReliabilityabstractEmerging Resistive Memory (ReRAM) is a promising candidate as the replacement for DRAM because of its low power consumption, high density and high endurance. Due to the unique crossbar structure, ReRAM can be constructed with a very high density. However, ReRAM's crossbar structure causes an IR drop problem which results in non-uniform access latency in ReRAM banks and reduces its reliability. Besides, the access latency and reliability of ReRAM arrays are greatly influenced by the data patterns involved in a write operation. In this paper, we propose a performance and reliability efficient ReRAM-based main memory structure. At the circuit level, we propose a double-sided write driver design to reduce the IR drops along bitlines. At the architecture level, a region partition with address remapping method and two flip schemes are proposed to reduce the access latency and improve the reliability of ReRAM arrays. The experimental results show that the proposed design can improve the system performance by 30.3% on average and reduce the memory access latency by 25.9% on average over an aggressive baseline, meanwhile the design improves the reliability of ReRAM-based memory system. Yang Zhang 0051, Dan Feng 0001, Jingning Liu, Wei Tong 0001, Bing Wu 0001, Caihua Fang |
DAC | 5 |
| 2017 | DAWS: Exploiting Crossbar Characteristics for Improving Write Performance of High Density Resistive MemoryabstractResistive random access memory (RRAM) is promising to be used as high density storage-class memory by employing crossbar structure. However, the wire resistance in crossbar array causes the IR drop problem, which makes nonuniformity of write latency throughout the array. In large crossbar array, the write latency differs greatly even in the same row. Since the write latency of a region is determined by its slowest write-unit, the conventional group-by-row region partition and addressing scheme is suboptimal for improving the overall performance of RRAM. In this work, we present DAWS, a novel RRAM architecture that exploits intrinsic features of crossbar structure. We first build a circuit model to analyze the voltage distribution and write latency distribution in a crossbar array. Then we propose a voltage bias scheme to optimize write latency via minimizing the IR drop path. We further present block diagonal partition to narrow the variance of write latency within each region, thus the write latency of each region is reduced. Moreover, we provide block diagonal addressing to make the write latency monotonically increase with the physical address, which is in favor of address mapping and memory allocation. We also design diagonal writing and diagonal swapping to overlap SET and RESET operations by applying a particular voltage bias pattern that can exploit row level parallelism, thus the number of write operations is halved. The experimental results show that DAWS can reduce memory access latency by 24.0% and improve system performance by 29.7% over an aggressive baseline. Chengning Wang, Dan Feng 0001, Jingning Liu, Wei Tong 0001, Bing Wu 0001, Yang Zhang 0051 |
ICCD | 5 |