Erich F. Haratsch

dblp:31/1400 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
0since 2021 · last 2020
0009-0002-9152-9825ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-authorComputer networks · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3Theory of computation · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Storage systems · 83% Hardware reliability and fault tolerance · 12% Integrated circuit design · 4%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD
flash memory
1.462017
Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives · Proc. IEEE 2017
Vulnerabilities in MLC NAND Flash Memory Programming: Experimental Analysis, Exploits, and Mitigation Techniques · HPCA 2017
Enabling Accurate and Practical Online Flash Channel Modeling for Modern MLC NAND Flash Memory · IEEE J. Sel. Areas Commun. 2016
Storage systems › flash and SSD
flash memory reliability
0.622018
HeatWatch: Improving 3D NAND Flash Memory Device Reliability by Exploiting Self-Recovery and Temperature Awareness · HPCA 2018
Flash Memories: ISPP Renewal Theory and Flash Design Tradeoffs · IEEE J. Sel. Areas Commun. 2016
Storage systems
flash and SSD
0.522018
HeatWatch: Improving 3D NAND Flash Memory Device Reliability by Exploiting Self-Recovery and Temperature Awareness · HPCA 2018
Neighbor-cell assisted error correction for MLC NAND flash memories · SIGMETRICS 2014
Storage systems › flash and SSD › flash memory
multi-level cell flash
0.522017
Vulnerabilities in MLC NAND Flash Memory Programming: Experimental Analysis, Exploits, and Mitigation Techniques · HPCA 2017
Data retention in MLC NAND flash memory: Characterization, optimization, and recovery · HPCA 2015
Coding theory › error-correcting codes › storage coding
flash memory codes
0.412020
Syndrome-Coupled Rate-Compatible Error-Correcting Codes: Theory and Application · IEEE Trans. Inf. Theory 2020
Storage systems › flash and SSD › flash memory › NAND flash
3D NAND flash
0.312018
HeatWatch: Improving 3D NAND Flash Memory Device Reliability by Exploiting Self-Recovery and Temperature Awareness · HPCA 2018
Storage systems
storage reliability
0.312018
HeatWatch: Improving 3D NAND Flash Memory Device Reliability by Exploiting Self-Recovery and Temperature Awareness · HPCA 2018
Storage systems › storage reliability
data recovery
0.312017
Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives · Proc. IEEE 2017
Hardware reliability and fault tolerance
error correction
0.312017
Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives · Proc. IEEE 2017
Hardware reliability and fault tolerance
soft errors
0.312017
Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives · Proc. IEEE 2017
Storage systems › flash and SSD
SSD reliability
0.312017
Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives · Proc. IEEE 2017
Storage systems › magnetic recording
channel modeling
0.212016
Enabling Accurate and Practical Online Flash Channel Modeling for Modern MLC NAND Flash Memory · IEEE J. Sel. Areas Commun. 2016
Storage systems › flash and SSD › flash memory › flash storage
intercell interference
0.212016
Flash Memories: ISPP Renewal Theory and Flash Design Tradeoffs · IEEE J. Sel. Areas Commun. 2016
Integrated circuit design › semiconductor device modeling
threshold voltage modeling
0.212016
Enabling Accurate and Practical Online Flash Channel Modeling for Modern MLC NAND Flash Memory · IEEE J. Sel. Areas Commun. 2016
Hardware reliability and fault tolerance
error recovery
0.212015
Data retention in MLC NAND flash memory: Characterization, optimization, and recovery · HPCA 2015
Storage systems › storage reliability › durability
retention failure
0.212015
Data retention in MLC NAND flash memory: Characterization, optimization, and recovery · HPCA 2015
Storage systems › flash and SSD › flash memory reliability
NAND flash error correction
0.212014
Neighbor-cell assisted error correction for MLC NAND flash memories · SIGMETRICS 2014
Coding theory › error-correcting codes
capacity-achieving codes
0.112020
Syndrome-Coupled Rate-Compatible Error-Correcting Codes: Theory and Application · IEEE Trans. Inf. Theory 2020
Reconfigurable computing and FPGAs
FPGA prototyping
0.112011
FPGA-based nand flash memory error characterization and solid-state drive prototyping platform (abstract only) · FPGA 2011
Storage systems › storage reliability
data corruption
0.112017
Vulnerabilities in MLC NAND Flash Memory Programming: Experimental Analysis, Exploits, and Mitigation Techniques · HPCA 2017
Storage systems › flash and SSD
solid-state drive
0.112017
Vulnerabilities in MLC NAND Flash Memory Programming: Experimental Analysis, Exploits, and Mitigation Techniques · HPCA 2017

Methods — techniques the papers use, named apart from their topics

experimental characterization · 0.8nested linear codes · 0.4coset construction · 0.4LDPC codes · 0.4BCH codes · 0.4reliability modeling · 0.3mitigation techniques · 0.3student's t-distribution · 0.2renewal theory · 0.2power law · 0.2error correction · 0.2neighbor-cell assisted error correction · 0.2wear leveling · 0.1signal processing · 0.1
YearPublicationVenuePosition
2020 Syndrome-Coupled Rate-Compatible Error-Correcting Codes: Theory and Application
abstract
Rate-compatible error-correcting codes (ECCs), which consist of a set of extended codes, are of practical interest in both wireless communications and data storage. In this work, we first study the lower bounds for rate-compatible ECCs, thus proving the existence of good rate-compatible codes. Then, we propose a general framework for constructing rate-compatible ECCs based on cosets and syndromes of a set of nested linear codes. We evaluate our construction from two points of view. From a combinatorial perspective, we show that we can construct rate-compatible codes with increasing minimum distances, and we discuss decoding algorithms and correctable patterns of errors and erasures. From a probabilistic point of view, we prove that we are able to construct capacity-achieving rate-compatible codes, generalizing a recent construction of capacity-achieving rate-compatible polar codes. Applications of rate-compatible codes to data storage are considered. We design two-level rate-compatible codes based on Bose-Chaudhuri-Hocquenghem (BCH) and low-density parity-check (LDPC) codes which are two popular codes widely used in the data storage industry, and then we evaluate the performance of these codes in multi-level cell (MLC) flash memories. We also examine code performance on binary and $q$ -ary symmetric channels. Finally, we briefly discuss two variations of our main construction and their relative performance.
Pengfei Huang 0001, Yi Liu 0052, Paul H. Siegel, Erich F. Haratsch
IEEE Trans. Inf. Theory5
2018 HeatWatch: Improving 3D NAND Flash Memory Device Reliability by Exploiting Self-Recovery and Temperature Awareness
abstract
NAND flash memory density continues to scale to keep up with the increasing storage demands of data-intensive applications. Unfortunately, as a result of this scaling, the lifetime of NAND flash memory has been decreasing. Each cell in NAND flash memory can endure only a limited number of writes, due to the damage caused by each program and erase operation on the cell. This damage can be partially repaired on its own during the idle time between program or erase operations (known as the dwell time), via a phenomenon known as the self-recovery effect. Prior works study the self-recovery effect for planar (i.e., 2D) NAND flash memory, and propose to exploit it to improve flash lifetime, by applying high temperature to accelerate self-recovery. However, these findings may not be directly applicable to 3D NAND flash memory, due to significant changes in the design and manufacturing process that are required to enable practical 3D stacking for NAND flash memory. In this paper, we perform the first detailed experimental characterization of the effects of self-recovery and temperature on real, state-of-the-art 3D NAND flash memory devices. We show that these effects influence two major factors of NAND flash memory reliability: (1) retention loss speed (i.e., the speed at which a flash cell leaks charge), and (2) program variation (i.e., the difference in programming speed across flash cells). We find that self-recovery and temperature affect 3D NAND flash memory quite differently than they affect planar NAND flash memory, rendering prior models of self-recovery and temperature ineffective for 3D NAND flash memory. Using our characterization results, we develop a new model for 3D NAND flash memory reliability, which predicts how retention, wearout, self-recovery, and temperature affect raw bit error rates and cell threshold voltages. We show that our model is accurate, with an error of only 4.9%. Based on our experimental findings and our model, we propose HeatWatch, a new mechanism to improve 3D NAND flash memory reliability. The key idea of HeatWatch is to optimize the read reference voltage, i.e., the voltage applied to the cell during a read operation, by adapting it to the dwell time of the workload and the current operating temperature. HeatWatch (1) efficiently tracks flash memory temperature and dwell time online, (2) sends this information to our reliability model to predict the current voltages of flash cells, and (3) predicts the optimal read reference voltage based on the current cell voltages. Our detailed experimental evaluations show that HeatWatch improves flash lifetime by 3.85× over a baseline that uses a fixed read reference voltage, averaged across 28 real storage workload traces, and comes within 0.9% of the lifetime of an ideal read reference voltage selection mechanism.
Saugata Ghose, Yu Cai 0001, Erich F. Haratsch, Onur Mutlu
HPCA4
2017 Vulnerabilities in MLC NAND Flash Memory Programming: Experimental Analysis, Exploits, and Mitigation Techniques
abstract
Modern NAND flash memory chips provide high density by storing two bits of data in each flash cell, called a multi-level cell (MLC). An MLC partitions the threshold voltage range of a flash cell into four voltage states. When a flash cell is programmed, a high voltage is applied to the cell. Due to parasitic capacitance coupling between flash cells that are physically close to each other, flash cell programming can lead to cell-to-cell program interference, which introduces errors into neighboring flash cells. In order to reduce the impact of cell-to-cell interference on the reliability of MLC NAND flash memory, flash manufacturers adopt a two-step programming method, which programs the MLC in two separate steps. First, the flash memory partially programs the least significant bit of the MLC to some intermediate threshold voltage. Second, it programs the most significant bit to bring the MLC up to its full voltage state. In this paper, we demonstrate that two-step programming exposes new reliability and security vulnerabilities. We experimentally characterize the effects of two-step programming using contemporary 1X-nm (i.e., 15–19nm) flash memory chips. We find that a partially-programmed flash cell (i.e., a cell where the second programming step has not yet been performed) is much more vulnerable to cell-to-cell interference and read disturb than a fully-programmed cell. We show that it is possible to exploit these vulnerabilities on solid-state drives (SSDs) to alter the partially-programmed data, causing (potentially malicious) data corruption. Building on our experimental observations, we propose several new mechanisms for MLC NAND flash memory that eliminate or mitigate data corruption in partially-programmed cells, thereby removing or reducing the extent of the vulnerabilities, and at the same time increasing flash memory lifetime by 16%.
Yu Cai 0001, Saugata Ghose, Ken Mai, Onur Mutlu, Erich F. Haratsch
HPCA6
2017 Syndrome-coupled rate-compatible error-correcting codes
abstract
Rate-compatible error-correcting codes (ECCs), which consist of a set of extended codes, are of practical interest in both wireless communications and data storage. In this work, we first study the lower bounds for rate-compatible ECCs, thus proving the existence of good rate-compatible codes. Then, we propose a general framework for constructing rate-compatible ECCs based on cosets and syndromes of a set of nested linear codes. We evaluate our construction from two points of view. From a combinatorial perspective, we show that we can construct rate-compatible codes with increasing minimum distances. From a probabilistic point of view, we prove that we are able to construct capacity-achieving rate-compatible codes.
Pengfei Huang 0001, Yi Liu 0052, Paul H. Siegel, Erich F. Haratsch
ITW5
2017 Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives
abstract
NAND flash memory is ubiquitous in everyday life today because its capacity has continuously increased and cost has continuously decreased over decades. This positive growth is a result of two key trends: 1) effective process technology scaling; and 2) multi-level (e.g., MLC, TLC) cell data coding. Unfortunately, the reliability of raw data stored in flash memory has also continued to become more difficult to ensure, because these two trends lead to 1) fewer electrons in the flash memory cell floating gate to represent the data; and 2) larger cell-to-cell interference and disturbance effects. Without mitigation, worsening reliability can reduce the lifetime of NAND flash memory. As a result, flash memory controllers in solid-state drives (SSDs) have become much more sophisticated: they incorporate many effective techniques to ensure the correct interpretation of noisy data stored in flash memory cells. In this article, we review recent advances in SSD error characterization, mitigation, and data recovery techniques for reliability and lifetime improvement. We provide rigorous experimental data from state-of-the-art MLC and TLC NAND flash devices on various types of flash memory errors, to motivate the need for such techniques. Based on the understanding developed by the experimental characterization, we describe several mitigation and recovery techniques, including 1) cell-to-cell interference mitigation; 2) optimal multi-level cell sensing; 3) error correction using state-of-the-art algorithms and methods; and 4) data recovery when error correction fails. We quantify the reliability improvement provided by each of these techniques. Looking forward, we briefly discuss how flash memory and these techniques could evolve into the future.
Yu Cai 0001, Saugata Ghose, Erich F. Haratsch, Onur Mutlu
Proc. IEEE3
2016 Flash Memories: ISPP Renewal Theory and Flash Design Tradeoffs
abstract
In the write process of multilevel per cell (MLC) flash memories, an iterative approach is used to mitigate the monotonicity problem. The monotonicity in programming is considered to be the major restriction in MLC flash. To solve this issue, an iterative approach called incremental step pulse programming (ISPP) is used to concurrently program lots of cells in small steps. In this paper, we are mostly concerned with deriving a mathematical model for iterative programming using the framework of renewal theory. We obtain a closed-form approximation for the probability distribution of the number of steps required in the ISPP process. We also bound the maximal error between the true distribution and our approximation. Moreover, the results obtained help to accurately analyze the effect of inter-cell interference in this type of memory. Finally, we devise an adaptive step size approach for write process to strike a balance between latency and lifetime under fixed bit error rate constraints or information rate constraints.
Meysam Asadi, Erich F. Haratsch, Aleksandar Kavcic, Narayana P. Santhanam
IEEE J. Sel. Areas Commun.2
2016 Enabling Accurate and Practical Online Flash Channel Modeling for Modern MLC NAND Flash Memory
abstract
NAND flash memory is a widely used storage medium that can be treated as a noisy channel. Each flash memory cell stores data as the threshold voltage of a floating gate transistor. The threshold voltage can shift as a result of various types of circuit-level noise, introducing errors when data are read from the channel and ultimately reducing flash lifetime. An accurate model of the threshold voltage distribution across flash cells can enable mechanisms within the flash controller that improve channel reliability and device lifetime. Unfortunately, existing threshold voltage distribution models are either not accurate enough or have high computational complexity, which makes them unsuitable for online implementation within the controller. We propose a new low-complexity flash memory model, built upon a modified version of the Student's t-distribution and the power law, which captures the threshold voltage distribution and predicts future distribution shifts as wear increases. Using our experimental characterization of the state-of-the-art 1X-nm (i.e., 15-19 nm) multi-level cell NAND flash chips, we show that our model is highly accurate (with an average modeling error of 0.68%), and also simple to compute within the flash controller (requiring 4.41 times less computation time than the most accurate prior model, with negligible decrease in accuracy). Our model also predicts future threshold voltage distribution shifts with a 2.72% modeling error. We demonstrate several example applications of our model in the flash controller, which improve flash channel reliability significantly, including a new mechanism to predict the remaining lifetime of a flash device. Our evaluations for two of these applications show that our model: 1) helps improve flash memory lifetime by 48.9% and/or (2) enables the flash device to safely sustain 69.9% more write operations than manufacturer specifications. We hope and believe that the analyses and models developed in this paper can inspire other novel approaches to flash memory reliability and modeling.
Saugata Ghose, Yu Cai 0001, Erich F. Haratsch, Onur Mutlu
IEEE J. Sel. Areas Commun.4
2015 Data retention in MLC NAND flash memory: Characterization, optimization, and recovery
abstract
Retention errors, caused by charge leakage over time, are the dominant source of flash memory errors. Understanding, characterizing, and reducing retention errors can significantly improve NAND flash memory reliability and endurance. In this paper, we first characterize, with real 2y-nm MLC NAND flash chips, how the threshold voltage distribution of flash memory changes with different retention age - the length of time since a flash cell was programmed. We observe from our characterization results that 1) the optimal read reference voltage of a flash cell, using which the data can be read with the lowest raw bit error rate (RBER), systematically changes with its retention age, and 2) different regions of flash memory can have different retention ages, and hence different optimal read reference voltages. Based on our findings, we propose two new techniques. First, Retention Optimized Reading (ROR) adaptively learns and applies the optimal read reference voltage for each flash memory block online. The key idea of ROR is to periodically learn a tight upper bound, and from there approach the optimal read reference voltage. Our evaluations show that ROR can extend flash memory lifetime by 64% and reduce average error correction latency by 10.1%, with only 768 KB storage overhead in flash memory for a 512 GB flash-based SSD. Second, Retention Failure Recovery (RFR) recovers data with uncorrectable errors offline by identifying and probabilistically correcting flash cells with retention errors. Our evaluation shows that RFR reduces RBER by 50%, which essentially doubles the error correction capability, and thus can effectively recover data from otherwise uncorrectable flash errors.
Yu Cai 0001, Erich F. Haratsch, Ken Mai, Onur Mutlu
HPCA3
2015 Write process modeling in MLC flash memories using renewal theory
abstract
In the write process of multilevel per cell (MLC) flash memories, an iterative approach is used to mitigate the monotonicity problem. The monotonicity in programming is considered to be the major restriction in MLC flash. In this paper, we are mostly concerned with deriving a mathematical model for iterative programming using the framework of “renewal processes”. Then, we approximate the maximum number of steps in iterative programming, and obtain the voltage distribution in flash due to iterative programming. Moreover, the obtained results help us to accurately analyze the effect of inter-cell interference (ICI) in this type of memory. Finally, we obtain a more precise voltage distribution for the symbol states in flash memory. Simulation results show the effect of varying the step size in the iterative programming and the effect of ICI on the information rate.
Meysam Asadi, Erich F. Haratsch, Aleksandar Kavcic, Narayana P. Santhanam
ISIT2
2014 Noise modeling and capacity analysis for NAND flash memories
abstract
Flash memories have become a significant storage technology. However, they have various types of error mechanisms, which are drastically different from traditional communication channels. Understanding the error models is necessary for developing better coding schemes in the complex practical settings. This paper endeavors to survey the noise and disturbs in NAND flash memories, and construct channel models for them. The capacity of flash memory under these models is analyzed, particularly regarding capacity degradation with flash operations, the trade-off of sub-thresholds for soft cell-level information, and the importance of dynamic thresholds.
Qing Li 0002, Anxiao Jiang, Erich F. Haratsch
ISIT3
2014 Neighbor-cell assisted error correction for MLC NAND flash memories
abstract
Continued scaling of NAND flash memory to smaller process technology nodes decreases its reliability, necessitating more sophisticated mechanisms to correctly read stored data values. To distinguish between different potential stored values, conventional techniques to read data from flash memory employ a single set of reference voltage values, which are determined based on the overall threshold voltage distribution of flash cells. Unfortunately, the phenomenon of program interference, in which a cell's threshold voltage unintentionally changes when a neighboring cell is programmed, makes this conventional approach increasingly inaccurate in determining the values of cells.
Yu Cai 0001, Gulay Yalcin, Onur Mutlu, Erich F. Haratsch, Osman S. Unsal, Adrián Cristal, Ken Mai
SIGMETRICS4
2013 Threshold voltage distribution in MLC NAND flash memory: characterization, analysis, and modeling
abstract
With continued scaling of NAND flash memory process technology and multiple bits programmed per cell, NAND flash reliability and endurance are degrading. Understanding, characterizing, and modeling the distribution of the threshold voltages across different cells in a modern multi-level cell (MLC) flash memory can enable the design of more effective and efficient error correction mechanisms to combat this degradation. We show the first published experimental measurement-based characterization of the threshold voltage distribution of flash memory. To accomplish this, we develop a testing infrastructure that uses the read retry feature present in some 2Y-nm (i.e., 20–24nm) flash chips. We devise a model of the threshold voltage distributions taking into account program/erase (P/E) cycle effects, analyze the noise in the distributions, and evaluate the accuracy of our model. A key result is that the threshold voltage distribution can be modeled, with more than 95% accuracy, as a Gaussian distribution with additive white noise, which shifts to the right and widens as P/E cycles increase. The novel characterization and models provided in this paper can enable the design of more effective error tolerance mechanisms for future flash memories.
Yu Cai 0001, Erich F. Haratsch, Onur Mutlu, Ken Mai
DATE2
2013 Program interference in MLC NAND flash memory: Characterization, modeling, and mitigation
abstract
As NAND flash memory continues to scale down to smaller process technology nodes, its reliability and endurance are degrading. One important source of reduced reliability is the phenomenon of program interference: when a flash cell is programmed to a value, the programming operation affects the threshold voltage of not only that cell, but also the other cells surrounding it. This interference potentially causes a surrounding cell to move to a logical state (i.e., a threshold voltage range) that is different from its original state, leading to an error when the cell is read. Understanding, characterizing, and modeling of program interference, i.e., how much the threshold voltage of a cell shifts when another cell is programmed, can enable the design of mechanisms that can effectively and efficiently predict and/or tolerate such errors. In this paper, we provide the first experimental characterization of and a realistic model for program interference in modern MLC NAND flash memory. To this end, we utilize the read-retry mechanism present in some state-of-the-art 2Y-nm (i.e., 20-24nm) flash chips to measure the changes in threshold voltage distributions of cells when a particular cell is programmed. Our results show that the amount of program interference received by a cell depends on 1) the location of the programmed cells, 2) the order in which cells are programmed, and 3) the data values of the cell that is being programmed as well as the cells surrounding it. Based on our experimental characterization, we develop a new model that predicts the amount of program interference as a function of threshold voltage values and changes in neighboring cells. We devise and evaluate one application of this model that adjusts the read reference voltage to the predicted threshold voltage distribution with the goal of minimizing erroneous reads. Our analysis shows that this new technique can reduce the raw flash bit error rate by 64% and thereby improve flash lifetime by 30%. We hope that the understanding and models developed in this paper lead to other error tolerance mechanisms for future flash memories.
Yu Cai 0001, Onur Mutlu, Erich F. Haratsch, Ken Mai
ICCD3
2012 Error patterns in MLC NAND flash memory: Measurement, characterization, and analysis
abstract
As NAND flash memory manufacturers scale down to smaller process technology nodes and store more bits per cell, reliability and endurance of flash memory reduce. Wear-leveling and error correction coding can improve both reliability and endurance, but finding effective algorithms requires a strong understanding of flash memory error patterns. To enable such understanding, we have designed and implemented a framework for fast and accurate characterization of flash memory throughout its lifetime. This paper examines the complex flash errors that occur at 30-40nm flash technologies. We demonstrate distinct error patterns, such as cycle-dependency, location-dependency and value-dependency, for various types of flash operations. We analyze the discovered error patterns and explain why they exist from a circuit and device standpoint. Our hope is that the understanding developed from this characterization serves as a building block for new error tolerance algorithms for flash memory.
Yu Cai 0001, Erich F. Haratsch, Onur Mutlu, Ken Mai
DATE2
2012 Flash correct-and-refresh: Retention-aware error management for increased flash memory lifetime
abstract
With the continued scaling of NAND flash and multi-level cell technology, flash-based storage has gained widespread use in systems ranging from mobile platforms to enterprise servers. However, the robustness of NAND flash cells is an increasing concern, especially at nanometer-regime process geometries. NAND flash memory bit error rate increases exponentially with the number of program/erase cycles. Stronger error correcting codes (ECC) can be used to tolerate higher error rates, but these have diminishing returns with increasing P/E cycles and can have prohibitively high power, area, and latency overheads. The goal of this paper is to develop new techniques that can tolerate high bit error rates without requiring prohibitively strong ECC. Our techniques, called Flash Correct-and-Refresh (FCR) exploit the observation that the dominant error source in NAND flash memory is retention errors, caused by flash cells losing charge over time. The key idea is to periodically read, correct, and reprogram (in-place) or remap the stored data before it accumulates more retention errors than can be corrected by simple ECC. Detailed simulations of a solid-state drive (SSD) storage system driven by measured experimental data from error characterization on real flash memory chips show that our techniques provide 46× average lifetime improvement on a variety of workloads at no additional hardware cost. We also find that our techniques achieve lifetime improvements that cannot feasibly be achieved with stronger ECC.
Yu Cai 0001, Gulay Yalcin, Onur Mutlu, Erich F. Haratsch, Adrián Cristal, Osman S. Unsal, Ken Mai
ICCD4
2011 FPGA-Based Solid-State Drive Prototyping Platform
abstract
NAND flash memory has been widely used for data storage due to its high density, high throughput, low cost, and low power. However, as flash memory manufacturers scale to smaller process technologies and store more bits per cell, the reliability and endurance of flash memory are decreasing. Wear-leveling and error correction coding can significantly improve both reliability and endurance, but finding effective algorithms requires quick and accurate characterization of flash memory error patterns. To this end, we have designed and implemented an FPGA-based open framework for quick, accurate, and comprehensive characterization of flash memories. Using this framework, we characterized detailed error patterns of flash memory throughout its entire lifetime. Our implementation uses an error accelerator block to decrease the test time by 20x. Based on these results, we propose and evaluate a smart bad block management policy in flash translation layer to increase SSD lifetime by up to 51%.
Yu Cai 0001, Erich F. Haratsch, Mark P. McCartney, Ken Mai
FCCM2
2011 FPGA-based nand flash memory error characterization and solid-state drive prototyping platform (abstract only)
abstract
NAND Flash memory has been widely used for data storage due to its high density, high throughput, low cost, and low power. However, as the storage cells become smaller and with more bits programmed per cell, they are expected to suffer from reduced reliability and limited endurance. Wear-leveling and signal processing can significantly improve both reliability and endurance. However, finding optimal algorithms would require a quick and accurate characterization of Flash memory providing an insight into the error patterns. To this end, we have designed and implemented an FPGA-based framework for quick, accurate, and comprehensive characterization of Flash memories to allow efficient algorithm explorations.
Yu Cai 0001, Erich F. Haratsch, Mark P. McCartney, Mudit Bhargava, Ken Mai
FPGA2
2007 A Bit-Node Centric Architecture for Low-Density Parity-Check Decoders
abstract
A bit-node centric decoder architecture for low- density parity-check codes is proposed. This architecture performs the optimum sum-product algorithm. A bit node processing unit computes the bit-to-check node messages sequentially, while the computation of the check-to-bit node messages is broken up into several steps. A stand-alone decoder architecture, and a decoder architecture for a concatenated detector-decoder system are presented. The proposed stand-alone decoder architecture requires significantly less memory compared to other known serial architectures. The hardware requirements are reduced even further for the concatenated detector-decoder system.
Ruwan N. S. Ratnayake, Erich F. Haratsch, Gu-Yeon Wei
GLOBECOM2
2007 Serial Sum-Product Architecture for Low-Density Parity-Check Codes
abstract
A serial sum-product architecture for low-density parity-check (LDPC) codes is presented. In the proposed architecture, a standard bit node processing unit computes the bit to check node messages sequentially, while the check node computations are broken up into several steps and computed on the fly. This bit node centric architecture requires considerably less memory compared to other serial architectures, including the check node centric architecture.
Ruwan N. S. Ratnayake, Erich F. Haratsch, Gu-Yeon Wei
ICCCN2
2006 High-rate quasi-cyclic LDPC codes for magnetic recording channel with low error floor
abstract
By implementing an FPGA-based simulator, we investigate the performance of high-rate quasi-cyclic (QC) LDPC codes for the magnetic recording channel at very low sector error rates. Results show that error-floor-free performance can be realized by randomly constructed high-rate regular QC-LDPC codes with column weight 4 for sector error rates as low as 10/sup -9/. We also conjecture several rules for designing randomly constructed high-rate regular QC-LDPC codes with low error floor. We also present a decoder architecture that is well suited to achieving high decoding throughput for these high-rate QC-LDPC codes with low error floor.
Hao Zhong 0006, Tong Zhang 0002, Erich F. Haratsch
ISCAS3
2000 Pipelined reduced-state sequence estimation
abstract
The throughput of reduced-state sequence estimation (RSSE) is limited by a substantially longer recursive loop than the add-compare-select function, which is the bottleneck of Viterbi decoding. This paper reformulates the RSSE algorithm to allow pipelining of the branch metric and decision-feedback computation. With this approach the critical path is shortened to the order of the add-compare-select function with only modest increase in hardware.
Erich F. Haratsch, Kamran Azadet
GLOBECOM1
2000 Reduced-State Sequence Estimation with Tap-Selectable Decision-Feedback
abstract
Reduced-state sequence estimation (RSSE) with tap-selectable decision-feedback is presented which significantly reduces the computational complexity of RSSE for sparse postcursor impulse response channels. Prefiltering is considered for more general channels, where a few channel coefficients constitute a large portion of the overall channel energy. Tap-selectable RSSE is less computationally expensive and better suited for high-speed VLSI implementation than prior reduced complexity RSSE techniques.
Erich F. Haratsch, Andrew J. Blanksby, Kamran Azadet
ICC (1)1
2000 High-speed reduced-state sequence estimation
abstract
Reduced-state sequence estimation (RSSE) is a suboptimal modification of the Viterbi algorithm (VA) which reduces the computational complexity of maximum likelihood sequence estimation. However, the maximum achievable throughput of RSSE may be significantly lower than the VA, as in addition to the add-compare-select function the critical path comprises branch metric computation, survivor memory operations, and decision-feedback calculation, This paper shows that precomputation of branch metrics can shorten the critical path of RSSE such that its delay becomes of the same order as the VA.
Erich F. Haratsch, Kamran Azadet
ISCAS1