Ken Takeuchi

dblp:93/4906 · DBLP profile ↗
← Back
27ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 Transformer Hetero-CiM: Heterogeneous Integration of ReRAM CiM and SRAM CiM for Vision Transformer at Edge Devices
abstract
This paper proposes design of Transformer Hetero-Computation-in-Memory (CiM) for Vision Transformer (ViT) at edge devices. ViT achieves high inference accuracy using parallel processing. However, ViT is computationally intensive due to a huge number of multiply-accumulate (MAC) operations. Moreover, ViT has the diverse requirements of Read-MAC operation in linear & FC layers and Read/Write-MAC operation in self-attention. Thus, proposed Transformer Hetero-CiM overcomes the issues by heterogeneous integration system of MLC ReRAM CiM, SRAM CiM and digital processors. As a result, proposed Transformer Hetero-CiM achieves 97.9% inference accuracy with reducing 89.1% of array area.
Naoko Misawa, Tao Wang 0126, Chihiro Matsui, Ken Takeuchi
ASP-DAC4
2025 VaLI: Variability-aware Fine-tuning with Low-rank Adapter and Iterative Training for ReRAM Computation-in-Memory
abstract
This paper proposes a novel algorithm called Variability-aware fine-tuning with Low-rank adapter and Iterative training (VaLI). Previous research applies variability-aware training (VAT) to tackle the difficulty in deploying AI models on non-volatile Computation-in-Memory (CiM) because of the device error introduced from non-idealities of non-volatile memory, e.g. ReRAM. To enable large AI models on the edge device, VaLI introduces Variability-Aware Fine-Tuning (VAFT) which extends the conventional VAT and saves training time. Moreover, VaLI incorporates Low-Rank Adapter (LoRA) to further reduce the excessive computation resources in the training, while proposing iterative training to improve the instability of VAFT with LoRA due to low-rank matrices. The proposed VaLI is evaluated across several models and datasets to showcase its effectiveness by reducing trainable parameters by an average of 90% while maintaining competitive model accuracy against device error.
Naoko Misawa, Chihiro Matsui, Ken Takeuchi
ISCAS4
2024 FeFET Local Multiply and Global Accumulate Voltage-Sensing Computation-In-Memory Circuit Design for Neuromorphic Computing
abstract
This work presents a design of voltage-sensing Computation-in-Memory (CiM) using ferroelectric FET (FeFET) in the point of device and circuit for neuromorphic computing. FeFET CiM with Local Multiply and Global Accumulate (LM-GA) operation works for multiply–accumulate (MAC) of artificial neural networks (ANNs) and integrate operation of spiking neural networks (SNNs). The high scalability and high ON/OFF ratio of FeFETs contribute to large-capacity CiM for neural networks. For the device design, measured FeFET characteristics by source-follower read show small variation in read-disturb and data-retention. The circuit design of FeFET LM-GA CiM is discussed with evaluations of wiring capacitance and resistance, threshold voltage shift, and gate timing. In addition, integration operation of SNNs, which is not possible with current-sensing CiM, is possible with FeFET LM-GA CiM.
Chihiro Matsui, Kasidit Toprasertpong, Shinichi Takagi, Ken Takeuchi
IEEE Trans. Very Large Scale Integr. Syst.4
2023 LIORAT: NN Layer I/O Range Training for Area/Energy-Efficient Low-Bit A/D Conversion System Design in Error-Tolerant Computation-in-Memory
abstract
Analog Computation-in-Memory (CiM) with ReRAM accelerates the MAC operations of neural networks (NNs). A major issue of CiM is the area and power consumption of analog-to-digital converters (ADCs). This work proposes a low-bit A/D conversion system to improve area/energy efficiency. However, the application-level accuracy is degraded due to quantization error and the limited range of low-bit ADC. To determine the optimal ADC range systematically while maintaining application-level accuracy, Layer Input/Output (I/O) Range Training (LIORAT) is proposed. LIORAT simultaneously trains the weights of a NN and the I/O range of each NN layer. Additionally, a digital ReRAM look-up table (LUT) is placed just after the ADC in the proposed A/D conversion system. Digital ReRAM LUT is used for non-MAC operations in the NN, such as batch normalization (BN). The values of the LUT are uniquely determined by the BN parameters and I/O ranges obtained by LIORAT. The application-level accuracy degradation caused by the ADC non-linearity and ReRAM weight errors can be compensated only by retraining BN parameters. Hence, the accuracy is recovered by updating digital L UT with the retrained BN parameters. Weight error compensation by updating digital L UT requires lower write accuracy compared to analog weight rewriting. ResNet-32 trained with the proposed LIORAT achieves 87.1 % inference accuracy on the CIFAR-10 dataset with only 10% LUT area overhead, 4-bit weights, 2-bit DAC, and 4-bit ADC. By updating the L UT, the magnitude of the tolerable error is more than doubled compared to the case without compensation.
Ayumu Yamada, Naoko Misawa, Chihiro Matsui, Ken Takeuchi
ICCAD4
2022 Versatile FeFET Voltage-sensing Analog CiM for Fast & Small-area Hyperdimensional Computing
abstract
This paper proposes fast and small-area FeFET-based voltage-sensing analog Computation-in-Memory (CiM) for hyperdimensional computing (HDC) by eliminating large-scale digital circuit overhead. In both training and inference of HDC, MAP (bit-wise XOR, bit-wise majority rule, and 1-bit shift) of hypervectors (HVs) is operated by Partially added Text HV FeFET CiM and Text HV FeFET. In inference, Similarity search FeFET CiM obtains classification result. By taking an example of Language classification problem, the proposed voltage-sensing FeFET CiM for HDC encodes HV in training phase by 5,000 times faster and smaller area than the conventional method.
Chihiro Matsui, Eitaro Kobayashi, Kasidit Toprasertpong, Shinichi Takagi, Ken Takeuchi
ISCAS5
2022 Domain Specific ReRAM Computation-in-Memory Design Considering Bit Precision and Memory Errors for Simulated Annealing
abstract
In this paper, domain specific ReRAM-based Computation-in-Memory (CiM) design for simulated annealing (SA) is proposed. This paper reveals that the influence of bit precision and memory cell errors of ReRAM CiM on the accuracy for SA depends on the domains of combinatorial optimization problems, such as Max-Cut and Knapsack problems. It is found that Max-Cut problem has smaller circuit structure and is 3-bit higher tolerant of bit precision, but 4% lower bit-error rate (BER) tolerant, compared with Knapsack problem. In this paper, considering the requirements of bit precision and BER from each domain, examples of case studies and design strategy are presented such that ReRAM CiMs are best optimized in terms of reliability and array area of memory cells. Furthermore, the proposed best-optimized ReRAM CiM for Max-Cut problem improves the quality of SA by introducing approximate answers and avoiding local minimum.
Naoko Misawa, Kenta Taoka, Chihiro Matsui, Ken Takeuchi
ISCAS4
2022 Edge Computation-in-Memory for In-situ Class-incremental Learning with Knowledge Distillation
abstract
This paper proposes a Computation-in-Memory (CiM) architecture for in-situ class-incremental learning. Due to usage change or environmental change, neural network models implemented in edge devices need to be retrained. Retraining on edge devices improves the latency of the retraining and reduces communication traffic, power consumption, and risk of security and privacy issue. The proposed CiM updates only the final fully connected (fc) layer, not the convolution layers. CiM does not need backpropagation and the number of rewrites to nonvolatile memory is small. The proposed CiM realizes knowledge distillation by cooperation of digital processor and CiM and can be retrained even when the old class data are not available. As a result, the accuracy keeps above 80% on CIFAR-10 when the bit precision of convolution and fc layer are 6 bits and 3 bits, respectively and bit-error rate of convolution and fc layers are less than 0.001% and 1%, respectively. In the proposed CiM, cell program during the retraining concentrate on the final fc layer, which is consistent with the characteristics of proposed class-incremental learning where the fc layer tolerates higher BER of memory cells and low bit precision than the convolution layers.
Shinsei Yoshikiyo, Naoko Misawa, Chihiro Matsui, Ken Takeuchi
ISCAS4
2021 BER Evaluation System Considering Device Characteristics of TLC and QLC NAND Flash Memories in Hybrid SSDs with Real Storage Workloads
abstract
This paper proposes BER evaluation system that evaluates BER of TLC and QLC NAND flash memories with reliability information such as write and erase (W/E) cycle and data-retention time by combining SSD model emulator and device characteristics of NAND flash memories. Proposed system decides which ECC type should be used in TLC and QLC NAND flash in SCM/TLC/QLC NAND flash tri-hybrid SSD, corresponding to various applications and memory capacity ratio. For hm_0 (write- cold application), BCH ECC is enough to correct bit errors in TLC NAND flash. On the other hand, for prxy_0 (write-hot application), LDPC ECC must be applied to TLC NAND flash in case of small SCM capacity, large W/E cycles and high BER in TLC NAND flash. In contrast, this paper concludes that QLC NAND flash needs LDPC ECC regardless of application and memory capacity.
Mamoru Fukuchi, Shun Suzuki, Kyosuke Maeda, Chihiro Matsui, Ken Takeuchi
ISCAS5
2021 Error Suppression of Last-Programmed Word-Line for Real Usage of 3D-NAND Flash Memory
abstract
Blocks of 3D-NAND flash memory are programmed in order of WL. In real usage of 3D-NAND flash memory, a block contains both programmed WLs and not programmed WLs because the programmed data do not always fill the block. In the block, the last data are programmed at 'last-programmed WL' where the upper WL is not programmed. The last programmed data are supposed to have high reliability owing to shorter data-retention time. However, when upper WL is not programmed, BER of last-programmed WL largely increases because upper WL has less electrons and causes lateral charge migration. To suppress the errors at last-programmed word-line, this paper proposes Last-programmed Word-line Protection (LWLP). Proposed LWLP suppresses BER by 37% at last-programmed WL and extends data-retention time by more than 4 times.
Daiki Kojima, Ken Takeuchi
ISCAS2
2020 Workload-aware Data-eviction Self-adjusting System of Multi-SCM Storage to Resolve Trade-off between SCM Data-retention Error and Storage System Performance
abstract
Storage Class Memories (SCMs) are used as non-volatile (NV) cache memory as well as storage. Multi-SCM storage with two types of SCMs, M-SCM (fast but small capacity memory-type SCM) and S-SCM (slow but large capacity storage-type SCM), has been proposed. In Multi-SCM storage, M-SCM works as NV-cache of S-SCM based storage. M-SCM such as MRAM is fast but may suffer from thermal instabilities and cause data-retention errors at high temperature. Therefore, data in M-SCM should be evicted to S-SCM at short interval before exceeding acceptable data-retention time. However, in case of short interval eviction, frequent data eviction from M-SCM to S-SCM severely degrades the storage system performance. To resolve this trade-off between data-retention reliability and the storage system performance, this paper proposes workload-aware data-eviction self-adjusting system. Proposed system is composed of Access Frequency Monitor (Proposal 1) and Evict Interval Adjustment (Proposal 2). Proposal 1 observes the access frequency of evicted data that directly affects data-retention time of M-SCM. By referring to the results of Proposal 1, Proposal 2 automatically changes the data-eviction interval so that long retention data are moved immediately to S-SCM and the storage system performance can be improved. As a result, maximum data-retention time of M-SCM decreases by 83%, and the storage system performance increases by 5.9 times. Moreover, the acceptable endurance increases by 103times. Finally, measured data-retention errors and memory cell area decrease by 79% and 5.7%, respectively.
Reika Kinoshita, Chihiro Matsui, Atsuya Suzuki, Shouhei Fukuyama, Ken Takeuchi
ASP-DAC5
2019 Design of heterogeneously-integrated memory system with storage class memories and NAND flash memories
abstract
Heterogeneously-integrated memory system is configured with various types of storage class memories (SCMs) and NAND flash memories. SCMs are faster than NAND flash, and they are divided into memory and storage types with their characteristics. NAND flash memories are also classified by the number of stored bits per memory cell. These non-volatile memories have trade-offs among access speed, capacity and bit cost. Therefore, mix and match of various non-volatile memories are essential to simultaneously achieve the best speed and cost of the storage. This paper proposes a design methodology with unique interaction of device, circuit and system to achieve the appropriate configurations in the heterogeneously-integrated memory system for application.
Chihiro Matsui, Ken Takeuchi
ASP-DAC2
2019 Self-Determining Resource Control in Multi-Tenant Data Center Storage with Future NV Memories
abstract
Self-determining resource control (SDRC) of hierarchical non-volatile memories is proposed for the multi-tenant data center storage. Proposed SDRC adjusts the capacity of emerging storage class memories (SCMs) such as MRAM (memory-type SCM) and PRAM (storage-type SCM) by considering the tenant-specific Service Level Agreement (SLA). Proposed SDRC is a self-contained system and therefore is immune to the frequent generation transition of non-volatile memories with different bit cost, latencies, and I/O bandwidth. This paper also introduces the circuit design of future scaled and high capacity PRAM as S-SCM, and the future interface for SCMs. The vertically integrated total optimization is achieved from memory device, memory circuit, and application to service and memory hierarchy-level. Evaluation results show that the optimal balance of performance and costs are automatically adjusted by SDRC and the Quality of Service requirement is achieved.
Chihiro Matsui, Ken Takeuchi
ISCAS2
2019 Step-by-Step Design of memory hierarchy for heterogeneously-integrated SCM/NAND flash storage
Chihiro Matsui, Ken Takeuchi
Integr.2
2019 Dynamic Adjustment of Storage Class Memory Capacity in Memory-Resource Disaggregated Hybrid Storage With SCM and NAND Flash Memory
abstract
Using storage class memories (SCMs) as nonvolatile cache of NAND flash memory is a promising solution for the high-performance storage. However, the problem is the high SCM cost per bit which is about ten times higher than NAND flash and the optimal SCM capacity is application dependent. The optimum SCM capacity is conventionally determined manually for every application operating on the data centers' storages. To achieve high performance while reducing the overall storage cost, application-aware autonomous SCM capacity adjustment (3ASCA) method is utilized to observe the warm data in NAND flash by the ghost least recently used list. SCM capacity in the memory-resource disaggregated storage with hybrid use of SCM and NAND flash is autonomously adjusted and effectively utilized for different applications. By saving the SCM capacity with 3ASCA, multiple applications in the disaggregated storage can efficiently utilize the limited SCM capacity at the same time.
Chihiro Matsui, Ken Takeuchi
IEEE Trans. Very Large Scale Integr. Syst.2
2018 20% System-performance Gain of 3D Charge-trap TLC NAND Flash over 2D Floating-gate MLC NAND Flash for SCM/NAND Flash Hybrid SSD
abstract
This paper analyzes the system-level performance of Storage Class Memory (SCM) / NAND flash hybrid solid-state drive (SSD). Four types of NAND flash, 1) 3-dimentional (3D) charge-trap (CT) Triple-Level Cell (TLC) [1], 2) 3D floating-gate (FG) TLC [2], 3) 2-dimentional (2D) FG TLC, and 4) 2D FG Multi-Level Cell (MLC) NAND flash are compared for various applications. For both read-and write-intensive workloads, the low cost 3D CT TLC NAND flash realizes the best performance that is 20% higher than 2D FG MLC NAND flash. This result is very encouraging for NAND flash communities because the low cost high capacity TLC NAND flash can be used at many applications. Especially, low cost 3D CT TLC NAND flash can be extensively used at data centers that requires the high performance SSD, while in 2D FG NAND flash, high cost MLC NAND flash was necessary. The performance gain of 3D CT TLC NAND flash is obtained by the short read/write latency and the smaller write unit which is the word-lines, not the block. Disadvantage of 130% larger block size of 3D CT TLC NAND flash over 2D FG MLC NAND flash is overcome by storing frequently accessed hot data in SCM. On the other hand, in 3D FG TLC NAND flash, the large write unit (block) seriously degrades the performance by 54%. Thus, in 3D FG NAND flash, high cost MLC NAND flash is still necessary at data centers. Finally, this paper shows that in the future 3D CT TLC NAND flash, the stacked layers may increase from 48 to 512. Although the block size increases from 9.4 MBytes to 100 MBytes, the system-level performance of 3D CT TLC NAND flash-based SSD does not degraded by utilizing SCM.
Mamoru Fukuchi, Yukiya Sakaki, Chihiro Matsui, Ken Takeuchi
ISCAS4
2018 Data-Aware Partial ECC with Data Modulation of ReRAM with Non-volatile In-memory Computing for Image Recognition with Deep Neural Network
abstract
This paper proposes data-aware ECC of resistive random access memory (ReRAM) in the non-volatile in-memory computing for image recognition. Proposed Data-Aw are Partial ECC with Data Modulation (DAP-ECC w/ DM) is implemented in the memory controller without controller circuit area overhead. Because ReRAM has smaller capacity than NAND flash, proposed ECC efficiently reduces overhead of error-correcting code (ECC) parity for ReRAM. Proposed DAP-ECC considers the importance of the feature vector data of image recognition, decreases the data overhead, and ECC decode latency by each 75% compared with conventional ECC. Moreover, proposed data-aware ECC using DAP-ECC w/ DM, which fully utilizes the asymmetrical bit error rate (BER) characteristics of ReRAM, improves the acceptable BER by 1.91-times. Finally, the endurance of ReRAM for the image recognition is improved by 10-times.
Atsuna Hayakawa, Toshiki Nakamura, Yoshiaki Deguchi, Kazuki Maeda, Ken Takeuchi
ISCAS5
2018 3ASCA: Application-Aware Autonomous SCM Capacity Adjustment for SCM and NAND Flash Pooled Storage
abstract
Using storage class memories (SCMs) as non-volatile cache of NAND flash memory is a promising solution for the high performance storage. However, the problem is the high SCM cost per bit which is about 10 times higher than NAND flash and the optimal SCM capacity is application dependent. The optimum SCM capacity is conventionally determined manually for every application operating on the data centers' storages. To achieve high performance while reducing the overall storage cost, this paper proposes Application-Aware Autonomous SCM Capacity Adjustment (3ASCA) method. SCM capacity in hybrid memory pool of SCM and NAND flash is autonomously adjusted by observing the warm data in NAND flash. SCM-assisted eviction algorithm is also proposed to reduce the warm data and use SCM more efficiently. As a result, the total storage cost during storage operation is decreased by up to 42% while its IOPS performance is degraded by only 4.8%. If the access count in SCM is also considered by introducing SCM-assisted eviction algorithm, the total storage cost during one-week storage operation is further decreased by 6.4%.
Chihiro Matsui, Ken Takeuchi
ISCAS2
2018 Layer-by-layer Adaptively Optimized ECC of NAND flash-based SSD Storing Convolutional Neural Network Weight for Scene Recognition
abstract
Layer-by-layer Adaptively Optimized Error Correcting Code (ECC) is proposed to improve the reliability of triple-level cell (TLC) NAND flash-based SSD for the scene recognition using convolutional neural network (CNN) of IoT edge devices. Layer-by-layer Adaptively Optimized ECC is composed of Layer-by-layer Iteration-Optimized Low Density Parity-Check (LBL-LDPC) and Layer-by-layer Code-length Adjusted Asymmetric Coding (LBL-AC). The conventional techniques like LDPC ECC and Asymmetric Coding (AC) improve the reliability. However, they require large overheads of the ECC decoding time and the flag/parity cell. Proposed LBL-LDPC and LBL-AC decrease the ECC decoding time by 14% and the data overhead by 26%, respectively, without recognition accuracy degradation. In addition, the data-retention time extends by 230%.
Keita Mizushina, Toshiki Nakamura, Yoshiaki Deguchi, Ken Takeuchi
ISCAS4
2017 Design of Hybrid SSDs With Storage Class Memory and NAND Flash Memory
abstract
NAND flash memory-based solid-state drives (SSDs) are increasingly being used in both consumer and enterprise storage markets, due to their superior performance over hard disk drives (HDDs) and continuous bit cost reductions. With multiple-level cell technology memory device is capable of trading off the performance and endurance with bit density. The more bits per cell there are, the longer latency and shorter lifetime. On the other hand, the performance of such SSDs is limited due to NAND flash access speed as well as the need of garbage collection. Recently, storage class memories (SCMs) like resistive RAM (ReRAM) and phase change RAM (PRAM) have been developed to fill the bandwidth gap between DRAM and NAND flash memory. SCMs are nonvolatile and byte addressable, which are much faster and durable than NAND flash. Therefore, with SCMs, the storage performance would be significantly improved. Hybrid SSDs are promising cost-efficient storage solutions. Various types of memories like single-level cell (SLC), multiple-level cell (MLC), triple-level cell (TLC) NAND flash memories, and SCMs create lots of opportunities for new system architectures and algorithms. In this paper, the architecture and algorithm design overview of three types of hybrid drives including MLC/TLC NAND flash hybrid, SCM/MLC NAND flash hybrid, and SCM/MLC/TLC NAND flash tri-hybrid are presented. From the evaluation results, hybrid drives demonstrate better performance, endurance, and power consumption, compared to the MLC NAND flash only SSD. Furthermore, the relationship between device reliability and performance of the SCM/NAND flash hybrid SSD has been understood at a system level. There is a tradeoff between acceptable bit error rate of SCM and NAND flash. In addition, the decoding latency of SCM affects the performance of hybrid SSD more than that of NAND flash.
Chihiro Matsui, Chao Sun 0001, Ken Takeuchi
Proc. IEEE3
2017 Write Order-Based Garbage Collection Scheme for an LBA Scrambler Integrated SSD
abstract
Solid-state drives (SSDs) are rapidly replacing hard disk drives in enterprise data centers due to their higher throughput and reliability. However, the SSD's random write performance is limited by the NAND flash memories within the SSD, which require garbage collection (GC). To improve the write throughput, a logical block address (LBA) scrambler has been previously proposed. However, there are two issues associated with this solution. First, with the LBA scrambler, SSD throughput actually worsens for some types of workloads, such as prxy_0. Second, a large table size is needed. In this paper, the first problem is solved by a write order (WO)-based GC scheme. In order to choose the victim block, the parameters of valid page ratio, write order, and erase count of the NAND flash blocks are collectively considered according to a new formula. A key advantage of utilizing the relative write order of the blocks is that an internal timer is not needed to monitor the ages of the blocks. Second, a sector bundling scheme is proposed to reduce the table size of the LBA scrambler. Based on the experimental results, with the two proposals, SSD throughput is improved by 2.4 times, and the table size of the LBA scrambler is reduced by 45%.
Chihiro Matsui, Asuka Arakawa, Chao Sun 0001, Ken Takeuchi
IEEE Trans. Very Large Scale Integr. Syst.4
2016 LBA Scrambler: A NAND Flash Aware Data Management Scheme for High-Performance Solid-State Drives
abstract
There is an increasing demand for the solid-state drive (SSD) due to its high speed, low power, and high reliability. However, random write intensive workload is not good for the SSD performance due to the inherent characteristics of the NAND flash memory. As the garbage collection (GC) causes the bottleneck of the SSD write performance due to the page-copy overhead, a NAND flash aware system is proposed to improve the SSD performance with a scheme called logical block address (LBA) scrambler. In the proposed scheme, new data are actively written to the fragmented pages in the next erase block. As a result, the number of valid pages inside the block is reduced when the block is recycled. Considering that there are NAND flash blocks full of valid pages in the proposed scheme, a skipping full block round robin (SFB_RR) GC policy is proposed, showing 0%-58% performance improvement compared with the RR GC policy. Furthermore, certain valid pages in the SSD have obsolete data due to the logical address remapping of the LBA scrambler, which cannot be invalidated by the conventional TRIM command, thus a SWEEP command is introduced. With the SWEEP command, maximum 12% additional SSD performance gain is obtained. From the experimental results, 35%-394% performance improvement, 27%-56% energy consumption reduction, and 25%-55% endurance enhancement are achieved by the proposed LBA scrambler scheme + SFB_RR GC policy + SWEEP command support, compared with the conventional SSD.
Chao Sun 0001, Ayumi Soga, Chihiro Matsui, Asuka Arakawa, Ken Takeuchi
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Understanding the Relation Between the Performance and Reliability of nand Flash/SCM Hybrid Solid-State Drive
abstract
A NAND flash memory/storage-class memory (SCM) hybrid solid-state drive (SSD) can achieve higher performance than the conventional NAND flash-only SSD. Error-correcting codes (ECCs) are applied to the SSD to correct bit errors occurring inside the NAND flash and SCM. To correct more bit errors, the stronger ECC is required and the ECC latency increases. This paper evaluates the relation between the performance and the reliability of the NAND flash/SCM hybrid SSD. First, how the ECC latency impacts the SSD performance is analyzed. Then, the SSD performances are evaluated with various data-access patterns. The ECC effect is significantly different among the data-access patterns. Moreover, four scenarios of the SCM reliability are established and the performances are evaluated with the four data-access patterns. When the SCM reliability becomes high, the decrease in the throughput due to the ECC for SCM becomes significantly small. Finally, by setting the acceptable SSD performance, the acceptable bit-error rate (BER) of the SCM is evaluated. The SCM BER can be as high as around 0.9%.
Shuhei Tanakamaru, Shogo Hosaka, Koh Johguchi, Hirofumi Takishita, Ken Takeuchi
IEEE Trans. Very Large Scale Integr. Syst.5
2014 Hybrid solid-state storage system with storage class memory and NAND flash memory for big-data application
abstract
Big-data enterprise storage is escalating demand for SSD, because of SSD's high speed, low power and small form factor. As a high-speed, low power and highly reliable enterprise storage, this paper overviews the high reliability signal processing technologies and hybrid memory solution which is the best mix and match of the high capacity NAND flash memories and storage class memories with non-volatility, speed, page rewritability and high endurance.
Ken Takeuchi
ISCAS1
2013 Over 10-times high-speed, energy efficient 3D TSV-integrated hybrid ReRAM/MLC NAND SSD by intelligent data fragmentation suppression
abstract
A 3D through-silicon-via (TSV)-integrated hybrid ReRAM/multi-level-cell (MLC) NAND solid-state drive's (SSD's) architecture is proposed with NAND-like interface (I/F) and sector-access overwrite policy for ReRAM. Furthermore, intelligent data management algorithms are proposed to suppress data fragmentation and excess usage of MLC NAND. As a result, 11-times performance increase, 6.9-times endurance enhancement and 93% write energy reduction are achieved. Both ReRAM write and read latency should be less than 3 μs to obtain these improvements. The required endurance for ReRAM is 105.
Chao Sun 0001, Hiroki Fujii, Kousuke Miyaji, Koh Johguchi, Kazuhide Higuchi, Ken Takeuchi
ASP-DAC6
2013 Highly reliable solid-state drives (SSDs) with error-prediction LDPC (EP-LDPC) architecture and error-recovery scheme
abstract
11-times extended lifetime, 76% reduced error SSD is proposed. The error-prediction LDPC realizes both 7-times faster read and high reliability. Errors are most efficiently corrected by calibrating memory data based on the VTH, inter-cell coupling, write/erase cycles and data-retention time. The error-recovery scheme with a program-disturb error-recovery pulse and a data-retention error-recovery pulse is also proposed to reduce the program-disturb error and the data-retention error by 76% and 56%, respectively.
Shuhei Tanakamaru, Yuki Yanagihara, Ken Takeuchi
ASP-DAC3
2011 Green high performance storage class memory & NAND flash memory hybrid SSD system
Ken Takeuchi
ISLPED1
2009 Inductor design of 20-V boost converter for low power 3D solid state drive with NAND flash memories
abstract
A 3D-integrated Solid State Drive (SSD) with the boost converter can achieve both the low power and the fast write-operation at the small die area of the NAND flash memory. The performance of the boost converter, however, is critically affected by the inductor, because the output voltage of the boost converter, the rising time, and the energy consumption during the boost are determined by the inductor. Therefore, this paper proposes a design methodology of the inductor of the boost converter for the 3D SSD. By using the boost converter with the optimized inductor, the energy during write-operation of the proposed 1.8-V 3D-SSD is decreased by 68% compared with the conventional 3.3-V 3D-SSD with the charge pump.
Tadashi Yasufuku, Koichi Ishida, Shinji Miyamoto, Hiroto Nakai, Makoto Takamiya, Takayasu Sakurai, Ken Takeuchi
ISLPED7