VLDB 2026 Research / reviewers in the wild / expert
Jingning Liu
dblp:97/6737
· DBLP profile ↗
60ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0002-2680-7422ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 10 since 2021Software engineering, systems software and programming languages · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | STAGGER: Enabling All-in-One Subarray Sensing for Efficient Module-level Processing in Open-Bitline ReRAMabstractEmerging resistive RAM (ReRAM) devices can in-situ execute vector-matrix-multiplication (VMM) and is able to achieve energy-efficient in-memory scientific computing. However, the peripheral separated S&Hs and ADCs for row buffering and sensing in conventional designs are the system bottleneck. We propose an ADC-less all-in-one processing-in-ReRAM design that enables the precharge once, readout multiple-bits (PORM) functionality for overlapping the tRCD latencies of different-significance result bits sensed out on the same bitline. Specifically, to support PORM sensing, we propose a cascaded-feedback bitline sensing architecture for VMM and a buffering-and-sensing-collocated sense amplifier elemental design with the bitline and the storage node fully decoupled for enabling conflict-free column accesses. We further propose cross-level inter-leaving mechanism for successive column VMM accesses to reduce the overall latency through improving the hardware spatiotemporal utilization. Experimental results show that our proposed design achieves 297% overall performance improvement and 85.8% energy reduction, compared with an aggressive baseline. Chengning Wang, Dan Feng 0001, Yuchong Hu, Wei Tong 0001, Jingning Liu |
DAC | 5 |
| 2023 | CorcPUM: Efficient Processing Using Cross-Point Memory via Cooperative Row-Column Access Pipelining and Adaptive Timing Optimization in SubarraysabstractEmerging cross-point memory can in-situ perform vector-matrix multiplication (VMM) for energy-efficient scientific computation. However, parasitic-capacitance-induced row charging and discharging latency is a major performance bottleneck of subarray VMM. We propose a memory-timing-compliant bulk VMM processing-using-memory design with row access and column access co-optimization from rethinking of read access commands and µ-op timing. We propose row-level-parallelism-adaptive timing termination mechanism to reduce tail latency of tRCD and tRP by exploiting row nonlinear charging and bulk-interleaved row-column-cooperative VMM access mechanism to reduce tRAS and overlap CL without increasing column ADC precision. Evaluations show that our design can achieve 5.03× performance speedup compared with an aggressive baseline. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu |
DAC | 4 |
| 2023 | APPcache+: An STT-MRAM-Based Approximate Cache System With Low Power and Long LifetimeabstractDue to high static power and low scalability, the traditional SRAM-based cache is not a good solution for image processing applications. Emerging spin transfer torque magnetic RAM (STT-MRAM) is a promising candidate for cache due to its low leakage power and high density. However, STT-MRAM suffers from high write energy. Therefore, by making use of the ability of tolerating minor errors in image processing applications, this work presents an STT-MRAM-basedAPProximatecachearchitecture (APPcache+) to write/read approximate data, which can largely reduce the cache energy and improve the STT-MRAM lifetime. APPcache+ includes three main designs. First, we find that there are many similar elements (e.g., pixels in images) in cache lines. Therefore, APPcache+ presents several lightweight similarity-based encoding techniques to remove redundant elements, thus, shortening the data size and reducing the energy of STT-MRAM cache. Second, we design a partial read scheme to reduce the read energy of the STT-MRAM cache. In the traditional decompression process, the whole line is fetched into the decompressor, leading to unnecessary read energy. The partial read scheme can largely reduce read energy while keeping the overhead low. Third, we observe the encoding schemes may lead to bit write imbalance. Therefore, we propose a lightweight Ping-Pong intraline wear-leveling scheme to improve the lifetime. Compared with the baseline, extensive evaluation results show that our APPcache+ can largely reduce the overall energy by 32.58%, improve lifetime by 40.7% with only 2.2% performance degradation, and 1.86% output quality loss. Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zhangyu Chen, Bing Wu 0001, Chengning Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | A Low-Latency and High-Endurance MLC STT-MRAM-Based Cache SystemabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising cache memory candidate due to its high density, low leakage power, and nonvolatility. Multilevel cell (MLC) STT-MRAM can further increase density by storing 2 bits in one cell’s hard and soft domain, respectively. However, MLC STT-MRAM suffers two-step write, leading to high write energy, long latency, and severe lifetime degradation. Current encoding techniques propose to encode the new data to reduce the two-step data writes. However, they have two weaknesses: 1) high area overhead, e.g., recent work TSE (Hsieh et al., 2020) needs extra 37.5% MLCs and 2) prolong the write latency due to an extra read. Therefore, we propose enhanced one-step write (EOSwrite) to write data in one step. EOSwrite includes line bypassing and four intraline encoding techniques. Line bypassing schemes can bypass the writes to zero or clean lines, leading to low write/read latency. As for the intraline techniques, we propose four write modes. They utilize the data patterns and the clean data in cache lines to write data in one step, therefore reducing the data write latency. The key idea of one-step write is to write as much data as possible in the soft domain of MLC STT-MRAM. EOSwrite can greatly relieve the weaknesses of the current encoding schemes. Evaluation results show that EOSwrite can improve the lifetime of MLC STT-MRAM by 56.96%, reduce dynamic energy by 33.95%, reduce access latency by 36.95%, and improve system performance of MLC STT-MRAM by 4.30%, respectively. While the area overhead of EOSwrite is only 7.27%. Wei Zhao 0034, Jie Xu 0013, Xueliang Wei, Bing Wu 0001, Chengning Wang, Weilin Zhu, Wei Tong 0001, Dan Feng 0001, Jingning Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2022 | Space-Time-Efficient Modeling of Large-Scale 3-D Cross-Point Memory Arrays by Operation Adaption and Network CompactionabstractThree-dimensional (3-D) integrated cross-point memory arrays can be used to build high-density storage-class memory systems. However, the coupled network topology caused by sharing word lines or bit lines between adjacent memory layers significantly enlarges the memory space overhead and time cost of memory operation simulations on mega-scale 3-D cross-point memory arrays. We observe that different components of the 3-D cross-point array have different contribution significance to the key metrics of write or read operations, and the distribution patterns of principal components in the arrays are different for write and read operations. We propose an operation-adaptive array modeling framework that exploits the impact of the applied operation on the distribution of array principal components for 3-D cross-point memory arrays. Based on the modeling framework, we propose two array network compaction methods for efficient write and read operation simulations on 3-D cross-point arrays, respectively: 1) pruning zero-biased unselected cells and 2) group-merging neighboring half-selected cells with a specific granularity. Also, serially connected line segments are merged, and floating segments are deleted. Evaluations show that the proposed methods significantly reduce the memory space overhead and time cost of memory operation simulations for various layer sizes and different cell-level access parallelism in an array. Besides, the proposed methods can efficiently simulate memory operations on multilayered 3-D cross-point memory arrays with up to$4096\times 4096$layer size under the 16-GB memory space constraint, achieving a 64-times improvement in layer size that can be simulated compared with the conventional complete 3-D cross-point array network model. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Bing Wu 0001, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Improving the energy efficiency of STT-MRAM based approximate cacheabstractApproximate computing applications lead to large energy consumption and performance demand for the memory system. However, traditional SRAM based cache cannot satisfy these demands due to high leakage power and limited density. Spin Transfer Torque Magnetic RAM (STT-MRAM) is a promising candidate of cache due to low leakage power and high density. However, STT-MRAM suffers from high write energy. To leverage the ability of tolerating acceptable quality loss via approximations to data, we propose an STT-MRAM based APProximate cache architecture (APPcache) to write/read approximate data thus largely reducing energy. We find many similar elements (e.g. pixels in images) existing in cache lines while running approximate computing applications. Therefore, APPcache uses several lightweight similarity-based encoding schemes to eliminate the similar elements to reduce the data size thus reducing the write energy of STT-MRAM based cache. Besides, we design a software interface to manually control the output quality. APPcache can significantly eliminate similar elements, thus improving energy efficiency. Experimental results show that our scheme can reduce write energy and improve the image raw data compression ratio by 21.9% and 38.0% compared with the state-of-the-art scheme with 1 % error rate, respectively. As for the output quality, the losses of all benchmarks are within 5% with 1 % error rate. Wei Zhao 0034, Wei Tong 0001, Dan Feng 0001, Jingning Liu, Zhangyu Chen, Jie Xu 0013, Bing Wu 0001, Chengning Wang, Bo Liu 0057 |
DATE | 4 |
| 2021 | MORE2: Morphable Encryption and Encoding for Secure NVMabstractMemory encryption can enhance the security of Non-volatile memories (NVMs), but it significantly increases the data bits written to NVMs and leads to severe lifetime and performance degradation. Current encryption techniques aim to reduce the re-encryption to many existing clean words, which unfortunately suffer from high encryption overheads (i.e. latency and energy) and many unnecessary writes. In the meantime, compression techniques can reduce the writes of encrypted NVM. However, we find that they may destroy the data patterns and increase the modified words, resulting in many encryptions in secure NVM. In this paper, we propose the MORphable Encryption and Encoding (MORE2) scheme to address these problems. Our MORphable Encryption (MORE) technique aims to reduce the full-line re-encryption and avoid clean line encryption. Besides, MORE proposes a prediction-based write scheme to avoid the encryption of clean lines, and pre-encrypt the lines that are predicted as dirty. Therefore, MORE can remove the encryption from the critical path of NVM. Furthermore, MORE2proposes the Morphable Selective Encoding (MSE) scheme to compress the modified words while preserving clean words. MORE2encrypts all metadata with the line counter to guarantee high security. Experimental results show that MORE2reduces the bit flips of encrypted NVM by 53.5 %, decreases the access latency by 27.32%, improves the IPC performance by 12.1 %, and reduces the write energy by 29.1 % compared with the state-of-the-art design. Wei Zhao 0034, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Jie Xu 0013, Gaoxiang Xu, Yiran Chen 0001 |
ICCAD | 5 |
| 2021 | QBLKe: Host-side flash translation layer management for Open-Channel SSDs
Hongwei Qin, Dan Feng 0001, Wei Tong 0001, Mengye Peng, Jingning Liu |
J. Syst. Archit. | 6 |
| 2021 | Improving Write Performance on Cross-Point RRAM Arrays by Leveraging Multidimensional Non-Uniformity of Cell Effective VoltageabstractResistive cross-point memory arrays can be used to construct high-density storage-class memory. However, coupled IR drop and sneak currents cause multidimensional non-uniformity of cell effective voltage in cross-point arrays. The voltage non-uniformity significantly degrades write performance on cross-point memory if only adopting the worst-case write latency at partial dimensions. Furthermore, the non-uniformity of cell effective voltage in cross-point arrays depends on multidimensional dynamic write operation parameters: row, column as well as layer address, the number of selected cells, and the number of half-selected low-resistance state cells. In this article, we aim to improve the write performance by leveraging multidimensional non-uniformity of cell effective voltage. First, we analyze the impact of multidimensional write parameters on effective voltage and write latency. Then, we design the memory array write scheme that measures the write parameters and sets the write latency accordingly. We further analyze the features and effects of interlayer sneak currents and extend the scheme to 3D cross-point memory. The evaluation shows that the proposed memory array write scheme can reduce the memory access latency by 75.6 and 64.1 percent, and improve the system performance by 4.5 times and 3.4 times on average, compared with the baseline and the state-of-the-art approach, respectively. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Yang Zhang 0051, Yiran Chen 0001 |
IEEE Trans. Computers | 4 |
| 2021 | Improving Multilevel Writes on Vertical 3-D Cross-Point Resistive MemoryabstractResistive memory is promising to be constructed as a high-density storage-class memory. Multilevel cell, access-transistor-free cross-point array structure, and 3-D array integration are three approaches to scale up the density of resistive memory. However, composing the three approaches together strengthens the interactions between array-level and cell-level nonidealities (interconnect resistance-induced IR drop, sneak current, and device variability) of resistive memory arrays during write operations and significantly degrades write performance and reliability. In this article, we analyze the dynamic voltage-dividing effect along a selected write current path in 3-D cross-point memory arrays. We propose a nonideality-tolerant high-density resistive memory (HD-RRAM) architecture, that can weaken the interactions between nonidealities and mitigate their degradation effects on the performance and reliability of array multilevel write operations. HD-RRAM is equipped with a double-transistor array architecture with two-transistor- n-resistor (2TnR) cell organization along pillars to reduce the current driving requirement and the large undesired voltage drop across each vertical pillar access transistor. Moreover, multiside asymmetric bias improves the resistive switching velocity by leveraging current-dividing effects. Variability-aware multilevel state partition reduces the worst-case write error rate by leveraging target state dependency of variability. Proportional-control multilevel state tuning reduces the average number of required write-and-verify iterations by leveraging pulse amplitude dependency of variability. Multilevel cell parallel writing improves the cell-level parallelism by leveraging the pass-through feature of intermediate resistance states. The evaluations show that HD-RRAM reduces both memory access latency and energy consumption over an aggressive baseline. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Linghao Song, Yang Zhang 0051, Jie Xu 0013, Xueliang Wei, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | CCHL: Compression-Consolidation Hardware Logging for Efficient Failure-Atomic Persistent Memory UpdatesabstractNon-volatile memory (NVM) is emerging as a fast byte-addressable persistent memory (PM) that promises data persistence at the main memory level. One of the common choices for providing failure-atomic updates in PM is the write-ahead logging (WAL) technique. To mitigate logging overhead, recent studies propose WAL-based hardware logging designs that overlap log writes with transaction execution. However, existing hardware logging designs incur a large number of unnecessary log writes. Many log writes are still performed in the critical path, which causes high performance overhead, particularly for the multi-core systems with many threads. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Liuqing Ye |
ICPP | 4 |
| 2020 | MorLog: Morphable Hardware Logging for Atomic Persistence in Non-Volatile Main MemoryabstractByte-addressable non-volatile memory (NVM) is emerging as an alternative for main memory. Non-volatile main memory (NVMM) systems are required to support atomic persistence and deal with the high overhead of programming NVM cells. To this end, recent studies propose hardware logging and data encoding designs for NVMM systems. However, prior hardware logging designs incur either extra ordering constraints or redundant log data. Moreover, existing data encoding designs are unaware of the characteristics of log data, resulting in writing unnecessary log bits.In this paper, we propose a morphable hardware logging design (MorLog) that only logs the data necessary for recovery and dynamically selects encoding methods with least write overhead. We observe that (1) only the oldest undo and the newest redo data in each transaction are necessary for recovery, and (2) the log data for clean bits are clean. The first motivates our morphable logging mechanism. This mechanism logs both undo and redo data for the first update to the data in a transaction, and then logs only redo data. Undo data are eagerly written to NVMM to ensure atomicity, while redo data are buffered in a volatile log buffer and L1 caches to write only the newest redo data to NVMM. The second motivates our selective log data encoding mechanism. This mechanism simultaneously encodes log data with different methods, and writes the encoded log data with the least write cost to NVMM. We devise a differential log data compression method to exploit the characteristics of log data. This method directly discards clean bits from log data and compresses remained dirty bits. Our evaluation shows that MorLog improves performance by 72.5%, reduces NVMM write traffic by 41.1%, and decreases NVMM write energy by 49.9% compared with the state-of-the-art design. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Liuqing Ye |
ISCA | 4 |
| 2020 | Multiple Subpage Writing FTL in MLC by Exploiting Dual Mode OperationsabstractThe page size of NAND flash continuously grows as the manufacturing process advances. While larger pages can reduce the cost per bit and improve the throughput of NAND flash, it may waste the storage space and data transfer time, causing more frequent garbage collections when serving small write requests. The main methods solving the mismatch problem between the request size and the write unit are write buffer cache and flash page reprogramming. However, multi-level cell (MLC) chips impose additional constraints on page programming so reprogramming MLC pages is prohibited. We proposed a multiple subpage writing flash translation layer (MSPW-FTL) for MLC by exploiting single-level cell (SLC)/MLC dual mode and flash page reprogramming feature. By converting MLC mode blocks to SLC mode blocks, we store small data in subpages of the SLC mode block. Moreover, we proposed three management methods to improve system efficiency: 1) two-level mapping to serve requests of different sizes; 2) an allocation strategy determines how the subpages of different logical pages are mapped to physical pages; and 3) a data management module to deal with the data fragmentation caused by the subpage granularity allocation. We compared MSPW-FTL with some related state-of-the-art FTLs under different types of workloads. Experimental results show that in average, MSPW-FTL reduces the I/O response time by 57.2%, the write amplification by 52.1%, and the number of erasures by 34.1%. Yazhi Feng, Dan Feng 0001, Wei Tong 0001, Jingning Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | A Low-Overhead Encoding Scheme to Extend the Lifetime of Nonvolatile MemoriesabstractEmerging nonvolatile memories (NVMs) are promising to replace DRAM as main memory. However, NVMs suffer from limited write endurance and high write energy. Encoding method reduces the bit flips of NVMs by exploiting additional tag bits to encode the data. The effect of the encoding method is limited by the capacity overhead of the tag bits. In this article, we propose to exploit the space saved by compression to store the tag bits of the encoding method. We observe that the saved space size of each compressed cache line varies, and different encoding methods have different tradeoffs between capacity overhead and effect. To fully exploit the space saved by compression for improving lifetime, we select the proper encoding method according to the saved space size. To improve the compression coverage and compression ratio, we select an efficient compression scheme from two compression algorithms and provide more space for data encoding. Still, some data patterns cannot be compressed by any compression technique. We use the Flip-N-Write with 3.1% capacity overhead to encode uncompressible cache lines. The experimental results show that our scheme reduces the bit flips by 32.5%, decreases the energy consumption by 22.6% and improves the lifetime by 69.9% with 3.5% capacity overhead. Dan Feng 0001, Jie Xu 0013, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | A Low Power Reconfigurable Memory Architecture for Complementary Resistive SwitchesabstractMemristive crossbar array suffers from severe sneak currents that incur reliability issues and extra energy waste. Complementary resistive switches (CRSs) provide a new concept to address the sneak-current problem. But the destructive read of CRS results in an additional recovery write operation, which strongly restricts its further promotion. Exploiting the dual CRS/memristor mode of CRS devices, we propose Aliens, a novel reconfigurable architecture that introduces one alien cell (memristor mode) for each bitline in the crossbar. Aliens draws advantages from both modes: restrained sneak currents of the CRS mode and nondestructive read of the memristor mode. The simple and regular cell mode organization one bitline one memristor (OBOM) of Aliens enables an energy-saving read method. Further, by exploiting memory access locality, an effective mode switching strategy called Lazy-Switch is proposed to delay and merge the recovery write operations of the CRS mode. Moreover, an 1TnR crossbar structure is adopted to enable larger crossbar arrays as well as a higher ratio of memristor mode cells without going against the OBOM rule. The effects of the memristor mode cell ratio on the energy consumption, endurance, and access performance are studied. Also, we show the bank architecture of Aliens and analyze how to extend our designs to 3-D arrays. Due to fewer recovery write operations and negligible sneak currents, Aliens achieves improvements in energy, overall endurance, and access performance. The experimental results show that our design offers average energy savings of 19.1× compared with memristor-only memory, a memory lifetime 10.7× longer than CRS-only memory, and a competitive performance compared with memristor-only memory. Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Wei Zhao 0034, Yang Zhang 0051 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | An Efficient Spare-Line Replacement Scheme to Enhance NVM SecurityabstractNon-volatile memories (NVMs) are vulnerable to serious threat due to the endurance variation. We identify a new type of malicious attack, called Uniform Address Attack (UAA), which performs uniform and sequential writes to each line of the whole memory, and wears out the weaker lines (lines with lower endurance) early. Experimental results show that the lifetime of NVMs under UAA is reduced to 4.1% of the ideal lifetime. To address such attack, we propose a spare-line replacement scheme called Max-WE (Maximize the Weak lines' Endurance). By employing weak-priority and weak-strong-matching strategies for spare-line allocation, Max-WE is able to maximize the number of writes that the weakest lines can endure. Furthermore, Max-WE reduces the storage overhead of the mapping table by 85% through adopting a hybrid spare-line mapping scheme. Experimental results show that Max-WE can improve the lifetime by 9.5X with the spare-line overhead and mapping overhead as 10% and 0.016% of the total space respectively. Jie Xu 0013, Dan Feng 0001, Yu Hua 0001, Fangting Huang, Wen Zhou 0030, Wei Tong 0001, Jingning Liu |
DAC | 7 |
| 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main MemoryabstractData encoding methods have been proposed to alleviate the high write energy and limited write endurance disadvantages of Non-Volatile Memories (NVMs). Encoding methods are proved to be effective through theoretical analysis. Under the data patterns of workloads, existing encoding methods could become inefficient. We observe that the new cache line and the old cache line have many redundant (or unmodified) words. This makes the utilization ratio of the tag bits of data encoding methods become very low, and the efficiency of data encoding method decreases. To fully exploit the tag bits to reduce the bit flips of NVMs, we propose REdundant word Aware Data encoding (READ). The key idea of READ is to share the tag bits among all the words of the cache line and dynamically assign the tag bits to the modified words. The high utilization ratio of the tag bits in READ leads to heavy bit flips of the tag bits. To reduce the bit flips of the tag bits in READ, we further propose Sequential flips Aware Encoding (SAE). SAE is designed based on the observation that many sequential bits of the new data and the old data are opposite. For those writes, the bit flips of the tag bits will increase with the number of tag bits. SAE dynamically selects the encoding granularity which causes the minimum bit flips instead of using the minimum encoding granularity. Experimental results show that our schemes can reduce the energy consumption by 20.3%, decrease the bit flips by 25.0%, and improve the lifetime by 52.1%. Jie Xu 0013, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Gaoxiang Xu, Yiran Chen 0001 |
DAC | 5 |
| 2019 | QBLK: Towards Fully Exploiting the Parallelism of Open-Channel SSDsabstractBy exposing physical channels to host software, Open-Channel SSD shows great potential in future high performance storage systems. However, the existing scheme fails to achieve acceptable performance under heavy workloads. The main reasons reside not only in its single-buffer architecture, more importantly, but also in its line-based physical address management. Besides, the lock of address mapping table is also a performance burden under heavy workloads. We propose QBLK, an open source driver which tries to better exploit the parallelism of Open-Channel SSDs. Particularly, QBLK adopts four key techniques, namely (1) Multi-queue based buffering, (2) Per-channel based address management, (3) Lock-free address mapping, and (4) Fine-grained draining. Experimental results show that QBLK achieves up to 97.4% bandwidth improvement compared with the state-of-the-art PBLK scheme. Hongwei Qin, Dan Feng 0001, Wei Tong 0001, Jingning Liu |
DATE | 4 |
| 2019 | Accelerating garbage collection for 3D MLC flash memory with SLC blocksabstract3D MLC NAND Flash is more appreciated for its massive capacity and significant performance. It's common to configure a portion of flash blocks to SLC-mode to further shorten the requests latency at a small cost of capacity. However, the limited SLC-mode blocks provoke GC (Garbage Collection) procedures more frequently and the GC penalty is heavier for 3D Flash than that for 2D Flash. As the block of 3D Flash consists of much more pages, the increment in the block size prolongs the erase operation latency and increases the number of migrated pages during GC. Existing works focus on reducing the number of migrated pages with sub-block GC strategies, which deeply depends on the distribution of valid pages across the sub-blocks in the victim block. In this paper, we propose a set of schemes called DCD which consists of Dual-mode Handler, Compensated GC and Dynamic Data Distribution. The key idea is to exploit the shorter operation delays of SLC-mode to accelerate the valid pages' migrations during GC, which doesn't rely on the distribution of valid pages inside the victim block. Experimental results show that compared with state-of-the-art designs, DCD shortens the average read response time by 19.8% and 37.3% for block-level and page-level FTL, respectively. Wei Tong 0001, Jingning Liu, Bing Wu 0001, Yazhi Feng |
ICCAD | 3 |
| 2019 | ReRAM Crossbar-Based Analog Computing Architecture for Naive Bayesian EngineabstractRecent advances in Resistive RAM (ReRAM) have explored the in-situ Matrix-Vector Multiplication (MVM) ability of crossbar arrays to achieve high energy-efficiency Process-In-Memory (PIM) architectures for Convolutional Neural Network (CNN), image processing, and so on. However, the existing ReRAM-based PIM architectures suffer from considerable additional auxiliary logic and device variations. In this work, we propose a novel analog computing architecture NB Engine for classification by implementing Naive Bayesian (NB) algorithm on ReRAM crossbar arrays. The two key steps of the NB algorithm, that is, probability calculation and electing the class that has the highest probability, are elaborately accomplished in our architecture. The ReRAM arrays are both used as storage and computation components. We store the pre-calculated prior probabilities and conditional probabilities of every class in crossbar arrays. Then the probability calculation step is completed in parallel through the MVM operation of the array. In general, the election step is a multiple-comparison procedure and is normally implemented by a comparison tree. Here, we reuse the max pooling module in a conventional CNN PIM architecture to realize a compatible comparison logic. However, neither of the two designs can avoid the overhead of costly high bit-precision Analog-to-Digital Converters (ADCs). So we introduce a novel analog parallel comparison design which does not need any ADCs or other computing logic with better energy-saving and area-efficiency. Our proposed NB Engine is tested by 11 various datasets. The influence of several non-ideal device properties is discussed and the NB Engine exhibits great tolerance to these variations. The experiment results show that our design offers a runtime speedup up to 2289.6x compared with the software-implemented NB classifier with negligible accuracy loss. In addition, the NB Engine saves 96.2% energy consumption and 45.2% array area compared with the CNN PIM compatible design. Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Wei Zhao 0034, Mengye Peng |
ICCD | 4 |
| 2019 | Tiered-ReRAM: A Low Latency and Energy Efficient TLC Crossbar ReRAM ArchitectureabstractResistive Memory (ReRAM) is promising to be used as high density storage-class memory by employing Triple-Level Cell (TLC) and crossbar structures. However, TLC crossbar ReRAM suffers from high write latency and energy due to the IR drop issue and the iterative program-and-verify procedure. In this paper, we propose Tiered-ReRAM architecture to overcome the challenges of TLC crossbar ReRAM. The proposed Tiered-ReRAM consists of three components, namely Tiered-crossbar design, Compression-based Incomplete Data Mapping (CIDM), and Compression-based Flip Scheme (CFS). Specifically, based on the observation that the magnitude of IR drops is primarily determined by the long length of bitlines in Double-Sided Ground Biasing (DSGB) crossbar arrays, Tiered-crossbar design splits each long bitline into the near and far segments by an isolation transistor, allowing the near segment to be accessed with decreased latency and energy. Moreover, in the near segments, CIDM dynamically selects the most appropriate IDM for each cache line according to the saved space by compression, which further reduces the write latency and energy with insignificant space overhead. In addition, in the far segments, CFS dynamically selects the most appropriate flip scheme for each cache line, which ensures more high resistance cells written into crossbar arrays and effectively reduces the leakage energy. For each compressed cache line, the selected IDM or flip scheme is applied on the condition that the total encoded data size will never exceed the original cache line size. The experimental results show that, on average, Tiered-ReRAM can improve the system performance by 30.5%, reduce the write latency by 35.2%, decrease the read latency by 26.1%, and reduce the energy consumption by 35.6%, compared to an aggressive baseline. Yang Zhang 0051, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Jie Xu 0013 |
MSST | 4 |
| 2019 | CeSR: A Cell State Remapping Strategy to Reduce Raw Bit Error Rate of MLC NAND FlashabstractThe following topics are dealt with: storage management; flash memories; parallel processing; cache storage; cloud computing; meta data; data compression; data handling; optimisation; learning (artificial intelligence). Wei Tong 0001, Jingning Liu, Dan Feng 0001, Hongwei Qin |
MSST | 3 |
| 2019 | Per-File Secure Deletion for Flash-Based Solid State DrivesabstractFile update operations generate many invalid flash pages in Solid State Drives (SSDs) because of the-of-place update feature. If these invalid flash pages are not securely deleted, they will be left in the “missing” state, resulting in leakage of sensitive information. However, deleting these invalid pages in real time greatly reduces the performance of SSD. In this paper, we propose a Per-File Secure Deletion (PSD) scheme for SSD to achieve non-real-time secure deletion. PSD assigns a globally unique identifier (GUID) to each file to quickly locate the invalid data blocks and uses Security-TRIM command to securely delete these invalid data blocks. Moreover, we propose a PSD-MLC scheme for Multi-Level Cell (MLC) flash memory. PSD-MLC distributes the data blocks of a file in pairs of pages to avoid the influence of programming crosstalk between paired pages. We evaluate our schemes on different hardware platforms of flash media, and the results prove that PSD and PSD-MLC only have little impact on the performance of SSD. When the cache is disabled and enabled, compared with the system without the secure deletion, PSD decreases SSD throughput by 1.3% and 1.8%, respectively. PSD-MLC decreases SSD throughput by 9.5% and 10.0%, respectively. Tianran Xiao, Wei Tong 0001, Jingning Liu, Bo Liu 0057 |
NAS | 4 |
| 2019 | Exploiting flash memory characteristics to improve performance of RAIS storage systems
Linjun Mei, Dan Feng 0001, Lingfang Zeng, Jianxi Chen, Jingning Liu |
Frontiers Comput. Sci. | 5 |
| 2019 | NICO: Reducing Software-Transparent Crash Consistency Cost for Persistent MemoryabstractEmerging non-volatile byte-addressable memory (NVM) introduces many opportunities and challenges to memory system designs. As data become persistent at main memory level, persistent memory systems need to guarantee the consistent state of data in the event of system failures (i.e., crash consistency). Existing studies propose persistent memory designs with software-transparent crash consistency guarantee to reduce programmers' manual effort when taking advantage of persistent memory. However, these designs are suboptimal due to their performance overhead caused by creating checkpoints. In this paper, we propose a Non-Intrusive memory COntroller design (NICO) that uses backend operations for achieving software-transparent crash consistency with minimized checkpointing overhead. By moving data persist operations to the background, NICO fully decouples data persist operations from volatile execution and cache management. To efficiently enforce crash consistency, we design a lightweight checkpointing scheme which only needs to flush and modify a very small amount of data when creating a consistent snapshot of persistent memory data. Our results show that NICO reduces the percent of time spent on checkpointing to within 0.9 percent across different benchmarks, and improves performance by 2.04× compared with existing checkpoint-based designs on average. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Liuqing Ye |
IEEE Trans. Computers | 4 |
| 2019 | Cross-point Resistive Memory: Nonideal Properties and SolutionsabstractEmerging computational resistive memory is promising to overcome the challenges of scalability and energy efficiency that DRAM faces and also break through the memory wall bottleneck. However, cell-level and array-level nonideal properties of resistive memory significantly degrade the reliability, performance, accuracy, and energy efficiency during memory access and analog computation. Cell-level nonidealities include nonlinearity, asymmetry, and variability. Array-level nonidealities include interconnect resistance, parasitic capacitance, and sneak current. This review summarizes practical solutions that can mitigate the impact of nonideal device and circuit properties of resistive memory. First, we introduce several typical resistive memory devices with focus on their switching modes and characteristics. Second, we review resistive memory cells and memory array structures, including 1T1R, 1R, 1S1R, 1TnR, and CMOL. We also overview three-dimensional (3D) cross-point arrays and their structural properties. Third, we analyze the impact of nonideal device and circuit properties during memory access and analog arithmetic operations with focus on dot-product and matrix-vector multiplication. Fourth, we discuss the methods that can mitigate these nonideal properties by static parameter and dynamic runtime co-optimization from the viewpoint of device and circuit interaction. Here, dynamic runtime operation schemes include line connection, voltage bias, logical-to-physical mapping, read reference setting, and switching mode reconfiguration. Then, we highlight challenges on multilevel cell cross-point arrays and 3D cross-point arrays during these operations. Finally, we investigate design considerations of memory array peripheral circuits. We also portray an unified reconfigurable computational memory architecture. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zheng Li 0005, Jiayi Chang, Yang Zhang 0051, Bing Wu 0001, Jie Xu 0013, Wei Zhao 0034, Ruoxi Ren |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2018 | Extending the lifetime of NVMs with compressionabstractEmerging Non-Volatile Memories (NVMs) such as Phase Change Memory (PCM) and Resistive RAM (RRAM) are promising to replace traditional DRAM technology. However, they suffer from limited write endurance and high write energy consumption. Encoding methods such as Flip-N-Write, FlipMin and CAFO can reduce the bit flips of NVMs by exploiting additional capacity to store the tag bits of encoding methods. The effects of encoding methods are limited by the capacity overhead of the tag bits. In this paper, we propose COE to COmpress cacheline for Extending the lifetime of NVMs. COE exploits the space saved by compression to store the tag bits of data encoding methods. Through combining data compression techniques with data encoding methods, COE can reduce the bit flips with negligible capacity overhead. We further observe that the saved space size of each compressed cacheline varies, and different encoding methods have different tradeoffs between capacity overhead and effects. To fully exploit the space saved by compression for improving lifetime, we select the proper encoding methods according to the saved space size. Experimental results show that our scheme can reduce the bit flips by 14.2%, decrease the energy consumption by 11.8% and improve the lifetime by 27.5% with only 0.2% capacity overhead. Jie Xu 0013, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu |
DATE | 5 |
| 2018 | An efficient PCM-based main memory system via exploiting fine-grained dirtiness of cachelinesabstractPhase Change Memory (PCM) has the potential to replace traditional DRAM memory due to its better scalability and non-volatility. However, PCM also suffers from high write latency and energy consumption. To mitigate the write overhead of PCM-based main memory, we propose a Fine-grained Dirtiness Aware (FDA) last-level cache (LLC) victimization scheme. The key idea of FDA is to preferentially evict cachelines with fewer dirty words when victimizing dirty cachelines. The modified word is defined to be dirty. FDA exploits two key observations. First, the write service time of a cacheline is proportional to the number of dirty words. Second, a cacheline with fewer dirty words has the same or lower reference frequency compared with other dirty cachelines. Therefore, evicting cachelines with fewer dirty words can reduce the write service time of cachelines and will not increase the miss rate. To reduce the write service time of cachelines, FDA evicts the cacheline with the fewest dirty words when victimizing dirty cachelines. We also present FDARP to decrease the miss rate by further synergizing the number of dirty words with Re-reference Prediction Value. Experimental results show that FDA (FDARP) can improve the IPC performance by 8.3% (14.8%), decrease the write service time of cachelines by 37.0% (36.3%) and reduce write energy consumption of PCM by 27.0% (32.5%) under the mixed benchmarks. Jie Xu 0013, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Zheng Li 0005 |
DATE | 5 |
| 2018 | A High-Performance and High-Reliability RAIS5 Storage Architecture with Adaptive Stripe
Linjun Mei, Dan Feng 0001, Lingfang Zeng, Jianxi Chen, Jingning Liu |
ICA3PP (1) | 5 |
| 2018 | Aliens: a novel hybrid architecture for resistive random-access memoryabstractPassive crossbar arrays of resistive random-access memory (RRAM) have shown great potential to meet the demands of future memory. By eliminating transistor per cell, the crossbar array possesses a higher memory density but introduces sneak currents which incur extra energy waste and reliability issues. The complementary resistive switch (CRS), consisting of two anti-serially stacked memristors, is considered as a promising solution to the sneak current problem. However, the destructive read of the CRS results in an additional recovery write operation which strongly restricts its further promotion. Exploiting the dual CRS/memristor mode of CRS devices, we propose Aliens, a novel hybrid architecture for resistive random-access memory which introduces one alien cell (memristor mode) for each wordline in the crossbar to provide a practical hybrid memory without operating system's intervention. Aliens draws advantages from both modes: restrained sneak current of CRS mode and non-destructive read of memristor mode. The simple and regular cell mode organization of Aliens enables an energy-saving read method and an effective mode switching strategy called Lazy-Switch. By exploiting memory access locality, Lazy-Switch delays and merges the recovery write operations of the CRS mode. Due to fewer recovery write operations and negligible sneak currents, Aliens achieves improvement in energy, overall endurance, and access performance. The experiment results show that our design offers average energy savings of 13.9× compared with memristor-only memory, a memory lifetime 5.3× longer than CRS-only memory, and a competitive performance compared with memristor-only memory. Bing Wu 0001, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Mingshun Yang, Chengning Wang, Yang Zhang 0051 |
ICCAD | 4 |
| 2018 | CACF: A Novel Circuit Architecture Co-optimization Framework for Improving Performance, Reliability and Energy of ReRAM-based Main Memory SystemabstractEmerging Resistive Random Access Memory (ReRAM) is a promising candidate as the replacement for DRAM due to its low standby power, high density, high scalability, and nonvolatility. By employing the unique crossbar structure, ReRAM can be constructed with extremely high density. However, the crossbar ReRAM faces some serious challenges in terms of performance, reliability, and energy consumption. First, ReRAM’s crossbar structure causes an IR drop problem due to wire resistance and sneak currents, which results in nonuniform access latency in ReRAM banks and reduces its reliability. Second, without access transistors in the crossbar structure, write disturbance results in serious data reliability problem. Third, the access latency, reliability, and energy use of ReRAM arrays are significantly influenced by the data patterns involved in a write operation. To overcome the challenges of the crossbar ReRAM, we propose a novel circuit architecture co-optimization framework for improving the performance, reliability, and energy use of ReRAM-based main memory system, called CACF. The proposed CACF consists of three levels, including the circuit level, circuit architecture level, and architecture level. At the circuit level, to reduce the IR drops along bitlines, we propose a double-sided write driver design by applying write drivers along both sides of bitlines and selectively activating the write drivers. At the circuit architecture level, to address the write disturbance with low overheads, we propose a RESET disturbance detection scheme by adding disturbance reference cells and conditionally performing refresh operations. At the architecture level, a region partition with address remapping method is proposed to leverage the nonuniform access latency in ReRAM banks, and two flip schemes are proposed in different regions to optimize the data patterns involved in a write operation. The experimental results show that CACF improves system performance by 26.1%, decreases memory access latency by 22.4%, shortens running time by 20.1%, and reduces energy consumption by 21.6% on average over an aggressive baseline. Meanwhile, CACF significantly improves the reliability of ReRAM-based memory systems. Yang Zhang 0051, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Chengning Wang, Bing Wu 0001, Zheng Li 0005, Gaoxiang Xu |
ACM Trans. Archit. Code Optim. | 5 |
| 2017 | A Novel ReRAM-based Main Memory Structure for Optimizing Access Latency and ReliabilityabstractEmerging Resistive Memory (ReRAM) is a promising candidate as the replacement for DRAM because of its low power consumption, high density and high endurance. Due to the unique crossbar structure, ReRAM can be constructed with a very high density. However, ReRAM's crossbar structure causes an IR drop problem which results in non-uniform access latency in ReRAM banks and reduces its reliability. Besides, the access latency and reliability of ReRAM arrays are greatly influenced by the data patterns involved in a write operation. In this paper, we propose a performance and reliability efficient ReRAM-based main memory structure. At the circuit level, we propose a double-sided write driver design to reduce the IR drops along bitlines. At the architecture level, a region partition with address remapping method and two flip schemes are proposed to reduce the access latency and improve the reliability of ReRAM arrays. The experimental results show that the proposed design can improve the system performance by 30.3% on average and reduce the memory access latency by 25.9% on average over an aggressive baseline, meanwhile the design improves the reliability of ReRAM-based memory system. Yang Zhang 0051, Dan Feng 0001, Jingning Liu, Wei Tong 0001, Bing Wu 0001, Caihua Fang |
DAC | 3 |
| 2017 | Mapping granularity adaptive FTL based on flash page re-programmingabstractThe page size of NAND flash continuously grows as the manufacturing process advances. While larger page can reduce the cost per bit and improve the throughput of NAND flash, it may waste the storage space and data transfer time. Meanwhile, it causes more frequent garbage collections when serving small write requests. To address the issues, we proposed a Mapping Granularity Adaptive FTL (MGA-FTL) based on flash page re-programming feature. MGA-FTL enables a finer granularity NAND flash space management and exploits multiple subpage writes on a single flash page without erase. 2-Level Mapping is introduced to serve requests of different sizes in order to control the overhead of DRAM requirement. Meanwhile, the allocation strategy determines whether different logical pages can be mapped to a single physical page to balance the space utilization and performance. Subpage merging limits the number of associated physical pages to a logical page, which could reduce data fragmentation and improves the performance of read operations. We compared MGA-FTL with some typical FTLs, including page-level mapping FTL and sector-log mapping FTL. Experimental results show that MGA-FTL reduces the I/O response time, write amplification and the number of erasures by 53%, 30% and 40% respectively. Despite the overhead of finegrained management, MGA-FTL increases no more than 16.5% DRAM requirement compared with a page-level mapping FTL. Unlike the subpage-level mapping, MGA-FTL only needs one third of DRAM space for storing mapping tables. Yazhi Feng, Dan Feng 0001, Chenye Yu, Wei Tong 0001, Jingning Liu |
DATE | 5 |
| 2017 | DAWS: Exploiting Crossbar Characteristics for Improving Write Performance of High Density Resistive MemoryabstractResistive random access memory (RRAM) is promising to be used as high density storage-class memory by employing crossbar structure. However, the wire resistance in crossbar array causes the IR drop problem, which makes nonuniformity of write latency throughout the array. In large crossbar array, the write latency differs greatly even in the same row. Since the write latency of a region is determined by its slowest write-unit, the conventional group-by-row region partition and addressing scheme is suboptimal for improving the overall performance of RRAM. In this work, we present DAWS, a novel RRAM architecture that exploits intrinsic features of crossbar structure. We first build a circuit model to analyze the voltage distribution and write latency distribution in a crossbar array. Then we propose a voltage bias scheme to optimize write latency via minimizing the IR drop path. We further present block diagonal partition to narrow the variance of write latency within each region, thus the write latency of each region is reduced. Moreover, we provide block diagonal addressing to make the write latency monotonically increase with the physical address, which is in favor of address mapping and memory allocation. We also design diagonal writing and diagonal swapping to overlap SET and RESET operations by applying a particular voltage bias pattern that can exploit row level parallelism, thus the number of write operations is halved. The experimental results show that DAWS can reduce memory access latency by 24.0% and improve system performance by 29.7% over an aggressive baseline. Chengning Wang, Dan Feng 0001, Jingning Liu, Wei Tong 0001, Bing Wu 0001, Yang Zhang 0051 |
ICCD | 3 |
| 2017 | Improving Performance of TLC RRAM with Compression-Ratio-Aware Data EncodingabstractResistive Random Access Memory (RRAM) technology is proposed as a promising replacement candidate for DRAM-based main memory due to its good scalability, low standby power, and non-volatility. The structure of Triple-Level Cell (TLC) can offer higher data density over Single-Level Cell (SLC). However, TLC RRAM suffers from high write energy and latency. Data compression techniques can reduce the size of the data to store. In contrast, data encoding methods such as Incomplete Data Mapping (IDM) can 'expand' the size for latency and energy reduction. We observe that the compression ratio of each cacheline varies, and therefore the saved space of each compressed cacheline is different. On the other hand, we find that different IDMs have different tradeoffs in capacity and write latency/energy. To fully exploit the space saved by compression for reducing the write latency/energy, and improving the performance of TLC RRAM-based main memory system, Compression-Ratio-Aware Data Encoding (CRADE) is proposed. The key idea of CRADE is to dynamically select the best-performing IDM according to the compression ratio of each cacheline. The cacheline is compressed first, and then the compressed cacheline is encoded by IDM. For each compressed cacheline, the IDM which uses the fewest states to encode is applied on the condition that the encoded data size will not exceed the cacheline size. Experimental results show that CRADE can reduce the write energy by 15%, decrease the write latency by 19%, reduce the read latency by 4%, and improve the IPC performance by 2% compared with the state-of-the-art scheme. Jie Xu 0013, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Wen Zhou 0030 |
ICCD | 5 |
| 2017 | Encoding Separately: An Energy-Efficient Write Scheme for MLC STT-RAMabstractMulti Level Cell (MLC) Spin Transfer Torque RAM (STT-RAM) provides higher density than Single Level Cell (SLC) STT-RAM by storing two digital bits in a single cell, and is proposed as a promising candidate for on-chip cache. However, MLC STT-RAM suffers from high write energy. We observe that general encoding methods, which map the frequent data patterns to the energy-efficient resistance states, cannot reduce the write energy of MLC STT-RAM. To reduce the write energy of MLC STT-RAM, we propose a novel encoding method, i.e., Encoding Separately (ES). The key idea of ES is to encode the hard bits and soft bits of MLCs separately. The hard bits are encoded for fewer hard-bit writes (hard transitions) and soft bits are encoded for fewer soft-bit writes (soft transitions). Specifically, existing encoding methods commonly used in SLC can be applied to MLC STT-RAM when encoding the two bits separately. We further apply two encoding methods for SLC to MLC STT-RAM through encoding separately, and experimental results show that the proposed scheme can reduce the writes to hard bits and soft bits by 28% and 16%, and achieve an energy reduction of 25%. Jie Xu 0013, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Wen Zhou 0030 |
ICCD | 4 |
| 2017 | A Write-Through Cache Method to Improve Small Write Performance of SSD-Based RAIDabstractWith the development of technology and price decline, flash-based Solid state drives (SSDs) are rapidly used to construct RAIDs by storage vendors. SSD does not need to seek and rotate, therefore, its read performance is much better than that of HDD. However, the small write performance of SSD is limited by its inherent characteristics such as out- of-place updates and garbage collection. The traditional parity-based RAID also has small write problem because of parity updating. SSD-based RAID, which is called RAIS, is generally based on the traditional RAID design and implementation. Consequently, handling small write requests is a serious challenge when SSD is used to construct parity-based RAID. In RAIS storage system, small write requests not only result in poor performance, but also shorten the lifetime of each SSD. In this paper, we propose a novel write through cache method, called CRAIS5, which uses a RAM as the write cache of RAIS5, and adopts the write-through mode to delay the parity update. The write-through cache method makes full use of the flash characteristics, and removes the pre-read operation. CRAIS5 improves the small write performance and reduces the erase time. We have implemented the CRAIS5 prototype in Disksim simulator, and used the real traces to evaluate the performance. The evaluations demonstrate that our CRAIS5 outperforms RAIS5, and PPC, on average, by 42.82%, and 34.49% respectively. Linjun Mei, Dan Feng 0001, Jianxi Chen, Lingfang Zeng, Jingning Liu |
NAS | 5 |
| 2017 | Time and Space-Efficient Write Parallelism in PCM by Exploiting Data PatternsabstractThe size of write unit in PCM, namely the number of bits allowed to be written concurrently at one time, is restricted due to high write energy consumption. It typically needs several serially executed write units to finish a cache line service when using PCM as the main memory, which results in long write latency and high energy consumption. To address the poor write performance problem, we propose a novel PCM write scheme called Min-WU (Minimize the number of Write Units). We observe data access locality that some frequent zero-extended values dominate the write data patterns in typical multi-threaded applications (more than 40 and 44.9 percent of all memory accesses in PARSEC workloads and SPEC 2006 benchmarks, respectively). By leveraging carefully designed chip-level data redistribution method, the data amount is balanced and the data pattern is the same among all PCM chips. The key idea behind Min-WU is to minimize the number of serially executed write units in a cache line service after data redistribution through sFPC (simplified Frequent Pattern Compression), eRW (efficient Reordering Write operations method) and fWP (fine-tuned Write Parallelism circuits). Using Min-WU, the zero parts of write units can be indicated with predefined prefixes and the residues can be reordered and written simultaneously under power constraints. Our design can improve the performance, energy consumption and endurance of PCM-based main memory with low space and time overhead. Experimental results of 12 multi-threaded PARSEC 2.0 workloads show that Min-WU reduces 44 percent read latency, 28 percent write latency, 32.5 percent running time and 48 percent energy while receiving 32 percent IPC improvement compared with the conventional write scheme with few memory cycles and less than 3 percent storage space overhead. Evaluation results of 8 SPEC 2006 benchmarks demonstrate that Min-WU earns 57.8/46.0 percent read/write latency reduction, 28.7 percent IPC improvement, 28 percent running time reduction and 62.1 percent energy reduction compared with the baseline under realistic memory hierarchy configurations. Zheng Li 0005, Fang Wang 0001, Dan Feng 0001, Yu Hua 0001, Jingning Liu, Wei Tong 0001, Yu Chen 0085, Salah S. Harb |
IEEE Trans. Computers | 5 |
| 2017 | CDF-LDPC: A New Error Correction Method for SSD to Improve the Read PerformanceabstractThe raw error rate of a Solid-State drive (SSD) increases gradually with the increase of Program/Erase (P/E) cycles, retention time, and read cycles. Traditional approaches often use Error Correction Code (ECC) to ensure the reliability of SSDs. For error-free flash memory pages, time costs spent on ECC are redundant and make read performance suboptimal. This article presents a CRC-Detect-First LDPC (CDF-LDPC) algorithm to optimize the read performance of SSDs. The basic idea is to bypass Low-Density Parity-Check (LDPC) decoding of error-free flash memory pages, which can be found using a Cyclic Redundancy Check (CRC) code. Thus, error-free pages can be read directly without sacrificing the reliability of SSDs. Experiment results show that the read performance is improved more than 50% compared with traditional approaches. In particular, when idle time of benchmarks and SSD parallelism are exploited, CDF-LDPC can be performed more efficiently. In this case, the read performance of SSDs can be improved up to about 80% compared to that of the state-of-art. Shigui Qi, Dan Feng 0001, Linjun Mei, Jingning Liu |
ACM Trans. Storage | 5 |
| 2017 | I/O Stack Optimization for Efficient and Scalable Access in FCoE-Based SAN StorageabstractDue to the high complexity in software hierarchy and the shared queue & lock mechanism for synchronized access, existing I/O stack for accessing the FCoE based SAN storage becomes a performance bottleneck, thus leading to a high I/O overhead and limited scalability in multi-core servers. In order to address this performance bottleneck, we propose a synergetic and efficient solution that consists of three optimization strategies for accessing the FCoE based SAN storage: (1) We use private per-CPU structures and disabling kernel preemption method to process I/Os, which significantly improves the performance of parallel I/O in multi-core servers; (2) We directly map the requests from the block-layer to the FCoE frames, which efficiently translates I/O requests into network messages; (3) We adopt a low latency I/O completion scheme, which substantially reduces the I/O completion latency. We have implemented a prototype (called FastFCoE, a protocol stack for accessing the FCoE based SAN storage). Experimental results demonstrate that FastFCoE achieves efficient and scalable I/O throughput, obtaining 1132.1K/836K IOPS (6.6/5.4 times as much as original Linux Open-FCoE stack) for read/write requests. Yunxiang Wu, Fang Wang 0001, Yu Hua 0001, Dan Feng 0001, Yuchong Hu, Wei Tong 0001, Jingning Liu |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2016 | Exploiting more parallelism from write operations on PCM
Zheng Li 0005, Fang Wang 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Yu Chen 0085, Dan Feng 0001 |
DATE | 5 |
| 2016 | Application-Aware and Software-Defined SSD Scheme for Tencent Large-Scale Storage SystemabstractTencent, one of the biggest Internet companies in China, contains billions of users and over 600-PB data, and leverages thousands of SSDs in the storage system to improve system performance and obtain energy savings. Existing commercial SSDs however fail to meet the needs of the ultra largescale applications due to not matching the service patterns. In order to address this problem and deliver high performance, we propose an application-aware and software-defined SSD scheme for Tencent applications, called TSSD. TSSD explores and exploits the business characteristics of Tencent, which facilitates the efficient use of SSDs. TSSD is software-defined by packaging each flash chip as a fully independent and concurrent storage unit. Each concurrent unit can be mounted as a character device, which allows the application layer to manage the flash chips in a more efficient manner, while optimizing the data layout. TSSD further employs a host-target FTL (TFTL) that uses a dedicated interface in the application layer, which efficiently connects the application layer with flash chips. Application layer hence becomes more accurately by using the flash memory chip-level information from TFTL, including the storage utilization, the degree of wear, etc. Moreover, TFTL is a programmable FTL and provides a programmable interface to the application layer. According to the running states of SSDs and workload information, TSSD makes use of the programmable interface to efficiently improve the performance of the FTL, wear leveling, and garbage collection for the specified applications. Extensive experiments use the real-world datasets from the commercial storage systems of Tencent. The results demonstrate that TSSD significantly improves the storage system performance and meets the needs of the Tencent's large-scale business applications. Jianquan Zhang, Dan Feng 0001, Jianlin Gao, Wei Tong 0001, Jingning Liu, Yu Hua 0001, Caihua Fang, Wen Xia, Feiling Fu, Yaqing Li |
ICPADS | 5 |
| 2016 | Increasing Lifetime and Security of Phase-Change Memory with Endurance VariationabstractPhase Change Memory (PCM) has emerged as a promising candidate for building the future main memory systems. However, the limited write endurance is one of the major obstacles for PCM to be practically applied. Traditional wear-leveling techniques try to uniformly balance the write traffics under both general applications and malicious attacks to enhance the PCM lifetime. However, these techniques fail to consider the endurance variation in PCM chips, and result in severe lifespan degradation since uniform write distribution leads to the weakest cell to be worn out much earlier. In this paper, we propose a weight-based algebraic wear-leveling (WAWL) scheme to balance wear rates (i.e., write traffics/endurance) in a secure manner according to the endurance distribution. In WAWL, the entire memory space is divided into multiple regions. When the number of the writes to a region reaches a threshold (i.e., swapping interval), the region is swapped with a randomly chosen region. The basic idea behind WAWL is that the swapping interval and the chosen probability of each region are variable and associated with the endurance metric of the region. By deploying suitable swapping interval and chosen probability, WAWL achieves uniform wear-rate distribution across the entire memory in an undetectable way. In addition, to reduce space consumption and alleviate performance degradation during region swapping, we propose a fine-grained swapping scheme which migrates the lines one-by-one between the candidate regions. Experimental evaluation driven by the various attacks demonstrates that WAWL significantly increases the PCM lifespan and improves security with slight performance degradation and affordable hardware overhead. Wen Zhou 0030, Dan Feng 0001, Yu Hua 0001, Jingning Liu, Fangting Huang, Pengfei Zuo |
ICPADS | 4 |
| 2016 | Tetris Write: Exploring More Write Parallelism Considering PCM AsymmetriesabstractThe noises at the power lines limit the charge pump to provide large instantaneous current to PCM cells, which results in the number of bits can be written concurrently, i.e. the size of write unit, is restricted in PCM. When implementing PCM as the main memory, the inequality of cache line's size and write unit's size may result in many consecutive executed write units, which greatly decreases the system performance. Existing PCM write schemes, however, consider the worst power and time cases of written data, and ignore the actual current consumption. It is assumed that all data bits are changed and the electric current of each data unit is under fully utilized. The write performance is blocked due to pessimistic estimates, i.e. the current is often excessively supplied but is not used effectively, which leads to huge energy consumption. As a result, the write parallelism is limited and therefore restricts the overall system performance. To address this problem, this paper proposes a novel PCM write scheme named Tetris Write to explore more write parallelism and reduce the critical number of write units in PCM chip. The key idea behind Tetris Write is to monitor the number of '1' and '0' changed in each data unit, and schedule the order of data units' write-1 and write-0 execution considering not only the time and power asymmetries, but also the number asymmetry between RET and SET operations, to allow a larger number of concurrent bit-writes and make the best use of power supply. Tetris Write tries to schedule the dominating long term write-1s first and attempts to steal interspaces remained by write-1s to put the extraessential short write-0s. 4-core PARSEC benchmarks' results show that Tetris Write can get 65% read latency reduction, 40% write latency reduction, 46% running time reduction and 2X IPC improvement compared with the baseline on average. In addition, Tetris Write earns 26%, 15% and 10% more read latency reduction, 15%, 7% and 5% more write latency reduction, and outperforms 22%, 12% and 7% more running time reduction, compared with the state-of-the-art Flip-N-Write, 2-Stage-Write and Three-Stage-Write schemes, whose IPC improvements are 1.4X, 1.6X and 1.8X, respectively. Zheng Li 0005, Fang Wang 0001, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu |
ICPP | 6 |
| 2016 | An Efficient Parallel Scheduling Scheme on Multi-partition PCM ArchitectureabstractPhase Change Memory (PCM) is an emerging non-volatile memory with the salient features of large-scale, high-speed, low-power and radiation resistance. It hence becomes an ideal candidate for the next-generation storage media of main memory. However, PCM suffers from inefficient I/O performance due to long write latency. Recent studies propose a multi-partition (or multi-subarray) architecture within each bank to enhance internal parallelism. However, conventional scheduling schemes fail to exploit the advantage of multiple partitions and incur inefficient bank utilization. In this paper, we propose a Write Priority overlap Read (WPoR) scheduling scheme which preferentially serves for a write request in one partition and allows other partitions to perform as many read requests as possible within this partition's program duration. Experimental results demonstrate that WPoR reduces the write latency by 24.7% (on average) compared with state-of-the-art scheduling algorithms. Meanwhile, the IPC indicator of WPoR scheduling increases respectively 6%, 7% and 26% (on average) compared with Read Priority, Write Pausing and Write Cancellation schemes. Wen Zhou 0030, Dan Feng 0001, Yu Hua 0001, Jingning Liu, Fangting Huang, Yu Chen 0085 |
ISLPED | 4 |
| 2016 | A Stripe-Oriented Write Performance Optimization for RAID-Structured Storage SystemsabstractIn modern RAID-structured storage systems, reliability is guaranteed by the use of parity blocks. But the parity-update overheads upon each write request have become a performance bottleneck of RAID systems. In some ways, an attached log disk is used to improve the write performance by delaying the parity blocks update. However, these methods are data-block-oriented and they need more time to rebuild or synchronize the RAID system when a data disk or the log disk fails. In this paper, we propose a novel optimization method, called SWO, which can improve RAID write performance and reconstruction performance. Moreover, when handling a write request, the SWO chooses reconstruction- write or read-modify-write combining with the log information to further minimize the number of pre- read data blocks. We have implemented the proposed SWO prototype and carried out some performance measurements using IOmeter and RAIDmeter. We have implemented the main idea of RAID6L in RAID5 and call it RAID5L. At the same time, we have evaluated the reconstruction time and the synchronization time of the SWO. Our experiments demonstrate that the SWO significantly improves write performance and saves more time than Data Logging and RAID5L when rebuilding and synchronizing. Linjun Mei, Dan Feng 0001, Lingfang Zeng, Jianxi Chen, Jingning Liu |
NAS | 5 |
| 2016 | Prober: exploiting sequential characteristics in buffer for improving SSDs write performance
Wen Zhou 0030, Dan Feng 0001, Yu Hua 0001, Jingning Liu, Fangting Huang, Yu Chen 0085, Shuangwu Zhang |
Frontiers Comput. Sci. | 4 |
| 2016 | A user-visible solid-state storage system with software-defined fusion methods for PCM and NAND flash
Zheng Li 0005, Fang Wang 0001, Jingning Liu, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Shuangwu Zhang |
J. Syst. Archit. | 3 |
| 2016 | MaxPB: Accelerating PCM Write by Maximizing the Power Budget UtilizationabstractPhase Change Memory (PCM) is one of the promising memory technologies but suffers from some critical problems such as poor write performance and high write energy consumption. Due to the high write energy consumption and limited power supply, the size of concurrent bit-write is restricted inside one PCM chip. Typically, the size of concurrent bit-write is much less than the cache line size and it is normal that many serially executed write units are consumed to write down the data block to PCM when using it as the main memory. Existing state-of-the-art PCM write schemes, such as FNW (Flip-N-Write) and two-stage-write, address the problem of poor performance by improving the write parallelism under the power constraints. The parallelism is obtained via reducing the data amount and leveraging power as well as time asymmetries, respectively. However, due to the extremely pessimistic assumptions of current utilization (FNW) and optimistic assumptions of asymmetries (two-stage-write), these schemes fail to maximize the power supply utilization and hence improve the write parallelism. In this article, we propose a novel PCM write scheme, called MaxPB (Maximize the Power Budget utilization) to maximize the power budget utilization with minimum changes about the circuits design. MaxPB is a “think before acting” method. The main idea of MaxPB is to monitor the actual power needs of all data units first and then effectively package them into the least number of write units under the power constraints. Experimental results show the efficiency and performance improvements on MaxPB. For example, four-core PARSEC and SPEC experimental results show that MaxPB gets 32.0% and 20.3% more read latency reduction, 26.5% and 16.1% more write latency reduction, 24.3% and 15.6% more running time decrease, 1.32× and 0.92× more speedup, as well as 30.6% and 18.4% more energy consumption reduction on average compared with the state-of-the-art FNW and two-stage-write write schemes, respectively. Zheng Li 0005, Fang Wang 0001, Dan Feng 0001, Yu Hua 0001, Jingning Liu, Wei Tong 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2016 | Reducing Fragmentation for In-line Deduplication Backup Storage via Exploiting Backup History and Cache KnowledgeabstractIn backup systems, the chunks of each backup are physically scattered after deduplication, which causes a challenging fragmentation problem. We observe that the fragmentation comes into sparse and out-of-order containers. The sparse container decreases restore performance and garbage collection efficiency, while the out-of-order container decreases restore performance if the restore cache is small. In order to reduce the fragmentation, we propose History-Aware Rewriting algorithm (HAR) and Cache-Aware Filter (CAF). HAR exploits historical information in backup systems to accurately identify and reduce sparse containers, and CAF exploits restore cache knowledge to identify the out-of-order containers that hurt restore performance. CAF efficiently complements HAR in datasets where out-of-order containers are dominant. To reduce the metadata overhead of the garbage collection, we further propose a Container-Marker Algorithm (CMA) to identify valid containers instead of valid chunks. Our extensive experimental results from real-world datasets show HAR significantly improves the restore performance by 2.84-175.36 × at a cost of only rewriting 0.5-2.03 percent data. Min Fu 0002, Dan Feng 0001, Yu Hua 0001, Xubin He, Zuoning Chen, Jingning Liu, Wen Xia, Fangting Huang, Qing Liu 0007 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2015 | 2QW-Clock: An Efficient SSD Buffer Management AlgorithmabstractModern solid state disk (SSD) has a buffer (SDRAM), which is used to store commonly used data and map in the near future. How to efficient management of this buffer is an important things of improving performance of SSD. Flash read and write speed have asymmetric characteristic. SSD buffer management algorithms must consider this characteristic of flash. Current page mapping SSD buffer management algorithms mainly use the Clean-First LRU (CFLRU) algorithm to first replace the clean buffer pages regardless of whether these pages will soon be used in the near future. At the same time, LRU buffer management algorithm of SSD does not consider file scanning. In order to solve these problems, we proposes a new SSD internal buffer management algorithm, called Two Queue Weight-Clock (2QW-Clock). This algorithm combines the advantages of 2Q and gives different weights to read page and write page to reflect the asymmetry of flash read and write speed. Therefore, it can get high write page hit ratios while maintaining high total page hit ratios. With the high write ratios, 2QW-Clock reduces the numbers of SSD write and erase operations. So it can greatly extend the life of the SSD. Conducting simulations with a variety of traces and a wide range of buffer sizes, we show that 2QW-Clock write hit ratios are significantly higher than CFLRU, LRU and 2Q in most cases while total hit ratios are almost as the 2Q. Simulation result shows that the numbers of 2QW-Clock write and erase counts reduced by up to 30% less than that of 2Q and CFLRU. Fang Wang 0001, Dan Feng 0001, Jingning Liu, Yunxiang Wu, Yang Hu 0007, Ying He 0012 |
HiPC | 4 |
| 2015 | Parallel Aware Hybrid Solid-State Storage
Fang Wang 0001, Dan Feng 0001, Jingning Liu, Yunxiang Wu, Ying He 0012, Yang Hu 0007 |
ICA3PP (4) | 4 |
| 2015 | Fast FCoE: An Efficient and Scale-Up Multi-core Framework for FCoE-Based SAN Storage SystemsabstractDue to the high complexity in software hierarchy and the shared queue & lock mechanism for synchronized access, existing I/O stack for remote target access in FCoE-based SAN storage becomes a performance bottleneck, thus leading to a high I/O overhead and limited I/O scalability in multi-core servers. For scalable performance, existing works focus on improving the efficiency of lock algorithm or reducing the number of synchronization points to decrease the synchronization overhead. However, the synchronization problem still exists and leads to a limited I/O scalability. In this paper, we propose Fast FCoE, a protocol stack framework for remote storage access in FCoE based SAN storage. Fast FCoE uses private per-CPU structures and disables the kernel preemption to process I/Os. This method avoids the synchronization overhead. For further I/O efficiency, Fast FCoE directly maps the requests from the block-layer to the FCoE frames. A salient feature of Fast FCoE is using the standard interfaces, thus supporting all upper softwares (such as existing file systems and applications) and offering flexible use in existing infrastructure (e.g., Adaptors, switches, storage devices). Our results demonstrate that Fast FCoE achieves efficient and scalable I/O throughput, obtaining 1107.3K/831.3K IOPS (5.43/4.88 times as much as Open-FCoE stack) for read/write requests. Yunxiang Wu, Fang Wang 0001, Yu Hua 0001, Dan Feng 0001, Yuchong Hu, Jingning Liu, Wei Tong 0001 |
ICPP | 6 |
| 2015 | Caching on dual-mode flash memoryabstractNAND flash memory has attracted wide attention in both academia and industry in recent years. Its high random access performance fills the gap between DRAM and hard disks. While MLC is endorsed for higher density and lower cost per bit, it suffers from poor performance and endurance. Dual-mode flash combines SLC and MLC in a single device and thus provides the opportunity to trade density for performance. In this paper, we propose the Scalable Flash Storage(SFS) abstraction layer to facilitate cache management on dual-mode flash. SFS exposes a virtualized address space to hide the variable density of the medium. A differentiated write interface is introduced, which allows the cache manager to explicitly send write requests to SLC for high performance. SFS dynamically scales the proportions of SLC and MLC to balance between cache capacity and performance. SFS provides partially persistent storage service. It allows the cache manager to manage the data persistence on flash so that critical data can be retained persistently. Non-persistent data are discarded during garbage collection to mitigate write amplification. Based on the SFS, a Dual-mode Flash Cache(DMFC) architecture is designed to utilize the configurable density and performance. Experimental results show that DMFC can significantly improve overall performance for various workloads. Sai Huang, Dan Feng 0001, Jianxi Chen, Jingning Liu |
NAS | 4 |
| 2014 | Improving Hybrid FTL by Fully Exploiting Internal SSD Parallelism with Virtual BlocksabstractCompared with either block or page-mapping Flash Translation Layer (FTL), hybrid-mapping FTL for flash Solid State Disks (SSDs), such as Fully Associative Section Translation (FAST), has relatively high space efficiency because of its smaller mapping table than the latter and higher flexibility than the former. As a result, hybrid-mapping FTL has become the most commonly used scheme in SSDs. But the hybrid-mapping FTL incurs a large number of costly full-merge operations. Thus, a critical challenge to hybrid-mapping FTL is how to reduce the cost of full-merge operations and improve partial merge operations and switch operations. In this article, we propose a novel FTL scheme, called Virtual Block-based Parallel FAST (VBP-FAST), that divides flash area into Virtual Blocks (VBlocks) and Physical Blocks (PBlocks) where VBlocks are used to fully exploit channel-level, die-level, and plane-level parallelism of flash. Leveraging these three levels of parallelism, the cost of full merge in VBP-FAST is significantly reduced from that of FAST. In the meantime, VBP-FAST uses PBlocks to retain the advantages of partial merge and switch operations. Our extensive trace-driven simulation results show that VBP-FAST speeds up FAST by a factor of 5.3--8.4 for random workloads and of 1.7 for sequential workloads with channel-level, die-level, and plane-level parallelism of 8, 2, and 2 (i.e., eight channels, two dies, and two planes). Fang Wang 0001, Hong Jiang 0001, Dan Feng 0001, Jingning Liu, Wei Tong 0001, Zheng Zhang 0013 |
ACM Trans. Archit. Code Optim. | 5 |
| 2012 | A Parity Scheme to Enhance Reliability for SSDsabstractRecent years, the application of solid-state disks (SSDs) increases explosively. All SSDs have to employ error correcting code (ECC) technique to ensure the reliability of flash memory at page level. However, data loss may be caused by bad block or chip failure of flash memory. To solve this problem, the article proposes a flash memory redundant array technique, which is similar to RAID-4. In this scheme, we utilize built-in NVRAM to cache the parity data update for minimal write to flash memory in parity channel. Dan Feng 0001, Jingning Liu, Wei Tong 0001, Yang Hu 0007, Zhiming Zhu |
NAS | 3 |
| 2010 | TRIP: Temporal Redundancy Integrated Performance Booster for Parity-Based RAID Storage SystemsabstractParity redundancy is widely employed in RAID-structured storage systems to protect against disk failures. However, the small-write problem has been a persistent root cause of the performance bottleneck of such parity-based RAID systems, due to the additional parity update overhead upon each write operation. In this paper, we propose a novel RAID architecture, TRIP, based on the conventional parity-based RAID systems. TRIP alleviates the small-write problem by integrating and exploiting the temporal redundancy (i.e., snapshots and logs) that commonly exists in storage systems to protect data from soft errors while boosting write performance. During the write-intensive periods, TRIP can reduce the penalty of each small-write request to as few as one device IO operation, at a minimal cost of maintaining the temporal redundant information. Reliability analysis, in terms of Mean Time to Data Loss (MTTDL), shows that the reliability of TRIP is only marginally affected. On the other hand, our prototype implementation and performance evaluation demonstrate that TRIP significantly outperforms the conventional parity-based RAID systems in data transfer rate and user response time, especially in write-intensive environments. Chao Jin 0002, Dan Feng 0001, Hong Jiang 0001, Lei Tian 0001, Jingning Liu, Xiongzi Ge |
ICPADS | 5 |
| 2010 | Achieving page-mapping FTL performance at block-mapping FTL cost by hiding address translationabstractFlash Translation Layer (FTL) is one of the most important components of SSD, whose main purpose is to perform logical to physical address translation in a way that is suitable to the unique physical characteristics of the Flash memory technology. The pure page-mapping FTL scheme, arguably the best FTL scheme due to its ability to map any logical page number (LPN) to any physical page number (PPN) to minimize erase operations, cannot be practically deployed since it consumes a prohibitively large RAM (SRAM or DRAM) space to store the page-mapping table for an SSD of moderate to large size. Alternatives to the pure page-mapping FTL, such as block-mapping FTLs, hybrid FTLs (e.g., FAST) and the latest demand-based page-mapping FTLs (e.g., DFTL), require significantly less RAM space but suffer from a few performance issues. Block-mapping FTLs perform poorly with higher erasure counts, particularly under random write workloads. Hybrid FTL schemes incur costly merge operations that hurt performance and increase the erasure counts. Performances of demand-based FTLs heavily depend on workload characteristics such as access locality, read/write ratio and request arrival interval time. This paper proposes a new FTL scheme, called HAT, to achieve the performance of a pure page-mapping FTL at the RAM cost of a block-mapping FTL while consuming lower energy, by hiding the address translation (HAT). The basic idea behind our scheme is to create a separate access path to read/write the address mapping information to significantly Hide the Address-Translation latency by incorporating a low energy-consuming solid-state memory device that stores the entire page mapping table. We implement an SSD simulator, SSDsim, to validate our HAT design and evaluate its performance. The extensive trace-driven simulation results show that the performance of HAT is within 0.8% of the pure page-mapping FTL, while consuming about 50% of the energy. Yang Hu 0007, Hong Jiang 0001, Dan Feng 0001, Lei Tian 0001, Shu Ping Zhang, Jingning Liu, Wei Tong 0001, Liuzheng Wang |
MSST | 6 |
| 2009 | MHPR: Multi-head Parallelism and Redundancy Disk ModelabstractPerformance and reliability are eternal topics in storage system. Nowadays, vast majority of storage system are made up of hard disk drive. The performance and reliability of hard disk drive are very important subclass of the former. So far, most researches of hard disk drive are single benefit. Some improve performance and some enhance the reliability. It is very little production which can heighten both performance and reliability at the same time. A new disk model we present fills this blank. The model not only improves disk drive performance, but also enhances reliability. The model is called MHPR (multi-head parallelism and redundancy). The models add more heads in one arm assembly, introduce redundancy and parallelism into disk. We analyze the performance and reliability in theory, and use different traces to evaluate the model in a simulator. We find that MHPR model can reduce seek time and transfer time effectively. The access time of single request is reduced. The model greatly shortens the queue time of single request in the case of two traces. So, the response of request is much faster; Response time of single disk drive is reduced by 98% and 80%, respectively in heavy workload and light workload. Using MHPR model to build RAID array, the MTTDL of this storage system can be improved 2-43 times. Yang Hu 0007, Dan Feng 0001, Shu Ping Zhang, Jingning Liu |
NAS | 4 |
| 2009 | 3DNBS: A Data De-duplication Disk-Based Network Backup SystemabstractTraditionally, backup and archiving have been performed on tapes. With the rapid advances in disk storage technology witnessed in recent years, it becomes practical to use disks other than tape libraries as backend storage device for a backup system. For such a disk-based system, storage space efficiency is essential. Since traditional backup method cannot eliminate redundancies during backup, a new data de-duplication backup technique should be developed to provide more efficient data storage at the system. This paper describes the design and performance evaluation of a data de-duplication disk-based network backup system,called 3DNBS. 3DNBS breaks files into variable sized chunks using content-defined chunking (CDC) for the purpose of duplication detection. Chunks are indexed and addressed by hashing their content, which leads to intrinsically single instance storage. Experimental results show that in comparison with traditional backup method such as Bacula, 3DNBS presents dramatic reduction in required storage space on various workloads. By eliminating duplicated data, 3DNBS also reduces the size of data to be transmitted, hence reducing time to perform backup in a bandwidth constraint environment. Tianming Yang, Dan Feng 0001, Jingning Liu, Yaping Wan, Zhongying Niu, Yuchang Ke |
NAS | 3 |