EDBT 2026 Demo / reviewers in the wild / expert
Ren-Shuo Liu
dblp:124/2061
· DBLP profile ↗
28ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-5311-4955ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 6 first-author · 18 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Input Reuse, Weight-Stationary Dataflow and Mapping Strategy for Depthwise Convolution in Computing-in-Memory Neural Network AcceleratorsabstractComputing-in-Memory (CIM) is a promising solution to address the bottleneck of data movement in traditional Von Neumann architecture by performing in-situ computation in the memory. However, naively mapping depthwise convolution onto general CIM results in high computation latency or storage overhead due to poor input reuse. Unlike previous works focusing on designing CIM to support depthwise convolution effectively, we propose an Input Reuse, Weight-Stationary dataflow and mapping strategy to support depthwise convolution without modifying the macro, achieving a good balance between computation latency and storage size. The key aspect of our dataflow and mapping is to maximize the reuse of input features by allowing CIM to produce partial sums of output features in the same cycle. Additionally, we propose a detailed hardware design of a partial sum processing unit to handle the reconstruction of output features from the partial sums. The experimental results show that our approach achieves up to $2.87 \times$ model-wise energy efficiency compared to the baseline design with only 6.3% area overhead of the CIM macro. Chia-Chun Wang, Yu-Chih Tsai, Ren-Shuo Liu |
ASP-DAC | 3 |
| 2026 | Clipping Error Compensation for Accuracy Recovery and Throughput Improvement in Computing-in-Memory
Yu-Chih Tsai, Hsuan-Hung Shen, Ren-Shuo Liu |
ISCAS | 3 |
| 2026 | Adaptive Row-wise Attention Score Pruning and Compensation for Attention Acceleration
Chia-Chun Wang, Yu-Lin Lin, Ren-Shuo Liu |
ISCAS | 3 |
| 2026 | Reliability and Optimization for Neural Network Accelerators using Value-Aware Error-Marking Pattern with Sequential Access Error Correction
Jun-Shen Wu, Ren-Shuo Liu |
ISCAS | 2 |
| 2025 | DMP-BFP: Dynamic Mixed-Precision Block Floating-Point and Exponent-Guided Precision AdjustmentabstractBlock Floating-Point (BFP), an emerging datatype, has demonstrated significant potential in model accuracy and hardware efficiency. This paper presents a dynamic mixedprecision BFP processing engine (PE), an accompanying framework, and optimization techniques to improve hardware efficiency. First, we propose a strategy for identifying accuracysensitive inner products within BFP models by comparing exponent values against predefined thresholds. This enables precision adjustments according to the sensitivity of the calculations at runtime. Second, we observe that only a small subset of inner products require full-precision (i.e., accuracy-sensitive inner products). Furthermore, a full-precision multiplication can be decomposed into four low-precision multiplications. Based on this, we propose the design of a low-precision PE capable of supporting full-precision mode, thereby reducing area overhead. Third, we optimize the BFP quantization scheme and datatype representation within the PE, significantly mitigating quantization errors in low-precision mode and reducing power consumption during datatype conversion. Finally, experimental results demonstrate that our dynamic mixed-precision BFP approach maintains accuracy while employing over 80% lowprecision operations and increases this ratio to 95% through retraining. Compared to state-of-the-art BFP architectures, our design improves inference speed, area efficiency, and energy efficiency by up to$1.64 \times, 1.42 \times$, and$1.47 \times$, respectively. Yu-Chih Tsai, Chia-Cheng Chang, Ren-Shuo Liu |
ICCD | 3 |
| 2025 | Access Frequency-Aware Storage Reduction for Deep Learning Recommendation ModelabstractPersonalized recommendation is one of the flourishing AI applications in recent years. However, its powerful functionality comes with the need of a significant amount of embeddings, which hinders the model performance by causing large memory footprints. In addition, each inference request requires multiple memory access to these embeddings, which puts even more stress on the already-suffering memory bandwidth. In this paper, we aim to reduce the storage size of the embeddings required by personalized recommendation while inducing minimal impact on the accuracy. To achieve such goals, we take advantage of the fact that not all data contribute equally towards the final result. Small values and less frequently accessed embeddings have little impact on the accuracy. In addition, we opt to avoid any extra steps of training to restore the accuracy since training recommendation models can take hours, consuming a huge amount of computing power, which is undesirable. Instead, to restore accuracy, we further identify the insight that some data, although less frequently accessed, do not offer good storage reduction-accuracy trade-offs. Following these key guidelines, we judiciously choose certain embeddings to apply element-wise pruning, leaving the rest untouched. We are able to prune the embedding data by more than 99% while inducing less than 1% accuracy drop. Chia-Chun Wang, Chuan-Yao Lai, Ren-Shuo Liu |
ICCD | 3 |
| 2024 | ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-MemoryabstractAmong several emerging architectures, computing in memory (CIM), which featuresin-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out thatmutually stationary vectors (MSVs), which can be maximized by introducingassociativityto CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. We have designed and realized an SA-CIM silicon prototype and corresponding architecture and acceleration schemes in the TSMC 28 nm process. More specifically, the contributions of this paper are fivefold: 1) We identify MSVs as new features that can be exploited to improve the current performance and energy challenges of the CIM-based hardware. 2) We propose SA-CIM to enhance MSVs (input-reordering flexibility) for skipping the zeros, small values, and sparse vectors. 3) We propose channel swapping to enhance the zero-skipping technique. 4) We propose a transposed systolic dataflow to efficiently conduct conv3×3 while being capable of exploiting input-skipping schemes. 5) We propose a design flow to search for optimal aggressive skipping scheme setups while satisfying the accuracy loss constraint. The proposed ISSA architecture improves the throughput by 1.91× to 2.97× speedup and the energy efficiency by 2.5× to 4.2×. Yun-Chen Lo, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Chih-Chen Yeh, Wen-Chien Ting, Ren-Shuo Liu |
IEEE Trans. Computers | 7 |
| 2023 | Bit-Serial Cache: Exploiting Input Bit Vector Repetition to Accelerate Bit-Serial InferenceabstractBit-serial computation has demonstrated superiority in processing precision-varying DNNs by slicing multi-bit vectors into multiple single-bit vectors and computing the inner product using multiple steps of shift-and-adds. In this paper, we identify that performing real-world DNNs inference with bit-serial computation exhibits high input bit vector locality, where up to 85.7% of non-zero input bit vectors, as well as their associated computation, are previously-seen and previously-done ones. We propose Bit-Serial Cache to transfer the identified locality into performance and energy efficiency gains. The key design strategy is to store recently-computed partial sums of input bit vectors to a cache and utilize cache accesses to replace redundant computations. In addition to the bit-serial computation architecture, we also present: 1) request clustering and 2) interleaved scheduling, to further enhance the performance and energy efficiency.Our experiments using six popular DNNs (in both 8-b and 4-b) show that Bit-Serial Cache speeds up DNN inference by up to 2.72×, 1.82×, and 4.03×, energy efficiency by 3.19×, 3.29×, and 2.82×, area efficiency by 1.35×, 1.24×, and 2.76× over state-of-the-art Loom, DPRed Loom, and Laconic. Yun-Chen Lo, Ren-Shuo Liu |
DAC | 2 |
| 2023 | Morphable CIM: Improving Operation Intensity and Depthwise Capability for SRAM-CIM ArchitectureabstractSRAM-based computing in memory (SRAM-CIM) supports weight-stationary dataflows and in-situ computation, successfully achieving lower weight SRAM traffic and hence better operation intensity (OI) than Von Neumann architecture. However, sticking to weight-stationary dataflows is sub-optimal since some layers of CNNs, e.g., depth-wise and input-dominant, suffer from large activation traffic and underutilization.We propose Morphable CIM to address the above challenges, in which the key contributions are: 1) We propose a dataflow-morphable SRAM-CIM architecture that adaptively switches between weight-stationary and input-stationary dataflow to enhance the overall operation intensity (OI) and computation utilization. 2) We propose an input-stationary systolic dataflow and a word-wise mapping to efficiently achieve dataflow reconfigurability for SRAM-CIM. 3) We propose a depthwise-capable CIM macro to improve the utilization of processing depth-wise layers.The experimental results on ten workloads show that our proposed SRAM-CIM architecture successfully outperforms the traffic-optimized and the performance-optimized baselines by up to 16.7× performance speedup, 3.6× higher energy efficiency, and 94.8% memory traffic reduction. Yun-Chen Lo, Ren-Shuo Liu |
DAC | 2 |
| 2023 | Built-in Self-Test and Built-in Self-Repair Strategies Without Golden Signature for Computing in MemoryabstractThis paper proposes built-in self-test (BIST) and built-in self-repair (BISR) strategies for computing in memory (CIM), including a novel test method and two repair schemes. They all focus on mitigating the impacts of inherent and in-evitable CIM inaccuracy on convolution neural networks (CNNs). Regarding the proposed BIST strategy, it exploits the distributive law to achieve at-speed CIM tests without storing testing vectors or golden results. Besides, it can assess the severity of the inherent inaccuracies among CIM bitlines instead of only offering a pass/fail outcome. In addition to BIST, we propose two BISR strategies. First, we propose to slightly offset the dynamic range of CIM outputs toward the negative side to create a margin for negative noises. By not cutting CIM outputs off at zero, negative noises are preserved to cancel out positive noises statistically, and accuracy impacts are mitigated. Second, we propose to remap the bitlines of CIM according to our BIST outcomes. Briefly speaking, we propose to map the least noisy bitlines to be the MSBs. This remapping can be done in the digital domain without touching the CIM internals. Experiments show that our proposed BIST and BISR strategies can restore CIM to less than 1% Top-1 accuracy loss with slight hardware overhead. Yu-Chih Tsai, Wen-Chien Ting, Chia-Chun Wang, Chia-Cheng Chang, Ren-Shuo Liu |
DATE | 5 |
| 2023 | BICEP: Exploiting Bitline Inversion for Efficient Operation-Unit-Based Compute-in-Memory Architecture: No Retraining Needed!abstractCompute-in-memory (CIM) architecture is promising for its in-situ analog computing ability. However, one practical constraint for CIM architectures is the limited number of activated rows in an operation Unit (OU). OU-based CIM architecture only activates a subgroup of memory cells to ensure a large signal margin and enough consideration of non-ideal device/circuit effect, which pays the cost of lowered computing throughput. In short, the OU-based CIM architectures suffer from array underutilization to ensure high accuracy.This work proposes a novel architecture, BICEP, which exploits bitline inversion technique to enlarge the OU size without the need to prune, approximate, and retrain. More specifically, the key contributions of this work are threefold: 1) We propose a bitline inversion scheme, which guarantees more than 2× larger OU size without affecting the numerical results and the ADC resolution. The key insight is to selectively apply code inversion on heavy bitlines to constrain their MAC outputs and compensate using low-cost compensation units. We mathematically prove that the proposal can be applied to both single- and multi-level cells (SLC and MLC). 2) We propose an inversion-aware weight swapping scheme, which swaps the weight order to maximize the OU size exploiting bitline inversion. 3) We propose weight order propagation to enable inversion-aware weight swapping without storage overheads. The extensive experiments on ImageNet classification tasks demonstrate that this work outperforms state-of-the-art OU-based CIM architecture (DL-RSIM) by up to 2.06× speedup and 1.97× energy efficiency. Yun-Chen Lo, Chia-Chun Wang, Ren-Shuo Liu |
ICCD | 3 |
| 2023 | CNN Inference Accelerators with Adjustable Feature Map Compression RatiosabstractRecently, an increasing interest has been in developing a convolution neural network (CNN) with adjustable configurations, enabling instant adaption to different resource constraints during inference. The trained CNN in run-time can switch to different modes to achieve a certain accuracy-energy trade-off point, similar to DVFS (dynamic voltage and frequency scaling) and turbo boost, which are widely adopted in CPUs. In this paper, we propose strategies to enable CNN inference accelerators to have an adjustable feature map compression ratio, making them tunable regarding their external memory access amount. We resort to the mature JPEG technique to compress those intermediate feature maps. The critical challenge is to support such adjustable compression ratios using one single CNN instead of multiple CNNs corresponding to multiple ratios. In response, we propose compression-aware joint-training and switchable batch normalization.We use ResNet18, ResNet50, and MobileNetV2 on ImageNet to demonstrate our design, achieve inference-time compression ratio adjustability, and reduce external memory access bandwidth requirements. The result shows that our proposed strategies can maintain the Top-1 accuracy and reduce external memory access by at most 22.7× ∼ 28.3× only using a single CNN model with sets of BN parameters corresponding to multiple compression ratios. Yu-Chih Tsai, Chung-Yueh Liu, Chia-Chun Wang, Tsen-Wei Hsu, Ren-Shuo Liu |
ICCD | 5 |
| 2023 | Exploiting and Enhancing Computation Latency Variability for High-Performance Time-Domain Computing-in-Memory Neural Network AcceleratorsabstractTo address the inefficiency resulting from data movement in Von Neumann architecture, computing-in-memory (CIM) is a promising solution due to its in-situ analog computation. Among the various types of CIMs, time-domain CIM stands out as a promising solution for achieving high energy efficiency and high readout resolution by employing time-to-digital converters (TDC) instead of analog-to-digital converters (ADC) to convert time-domain delays into digital values. However, the performance of the accelerator may be constrained by the maximum operating frequency of time-domain CIM, which is significantly lower than that of digital circuits.This paper proposes an architecture for a time-domain CIM-based neural network accelerator that leverages the varying output time of the TDC. The key contributions of this work are as follows: 1) We introduce an early-termination scheme for time-domain CIM, which dynamically determines the length of the CIM clock period by deriving the maximum possible multiply-accumulate (MAC) value based on the current input. This approach reduces computation time for low-MAC results. 2) We propose an input-inversion scheme to decrease the computation time for high-MAC results. By employing linear combination, we perform bit-inversion on large inputs and compensate for the results using a low-cost digital circuit. 3) We propose a hardware optimization on the compensation circuit by combining it with shift-adders in traditional neural network accelerators.Experiments show that our schemes could gain 2× ∼ 2.9× speedup under different clock period specifications with 5.82% area overhead compared to the CIM macro. Chia-Chun Wang, Yun-Chen Lo, Jun-Shen Wu, Yu-Chih Tsai, Chia-Cheng Chang, Tsen-Wei Hsu, Min-Wei Chu, Chuan-Yao Lai, Ren-Shuo Liu |
ICCD | 9 |
| 2023 | FM-P2L: An Algorithm Hardware Co-design of Fixed-Point MSBs with Power-of-2 LSBs in CNN AcceleratorsabstractConvolutional neural networks (CNNs) are a focal point for advancing the field of artificial intelligence, enabling significant advances in image and speech recognition, natural language processing, and other complex tasks. However, due to the high requirements of memory storage and computational resources, implementing CNNs can be challenging.To alleviate these deficiencies, this work presents a novel algorithm hardware co-design centered on a new number format, fixed-point MSBs with power-of-2 LSBs (FM-P2L), which can significantly reduce the requirement of computational resources of the CNN accelerators by trading negligible accuracy loss. First, we propose the novel FM-P2L number format that uses fixed-point to represent MSBs and power-of-2 to represent LSBs, which can reduce the computation complexity of multiplications. The optimal bitwidths of MSBs and LSBs are determined from our proposed algorithm with the capability to preserve accuracy. Second, we propose a novel multiplier that best matches FM-P2L with CNN accelerators, which can significantly reduce the area and power of the computational units. Finally, to evaluate the benefits of FM-P2L, we compared FM-P2L with fixed-point and low-bitwidith floating-point on a weight-stationary based systolic array accelerator with vector-vector multiplication PEs, which is adopted by Google TPU and NVDLA, in TSMC 40nm technology. Our evaluation results demonstrate that FM-P2L can achieve up to 44% computing power and 50% area reduction compared with fixed-point and up to 55% computing power and 65% area reduction compared with floating-point, while maintaining negligible inference accuracy loss on the state-of-the-art CNN models. Jun-Shen Wu, Ren-Shuo Liu |
ICCD | 2 |
| 2023 | Block and Subword-Scaling Floating-Point (BSFP) : An Efficient Non-Uniform Quantization For Low Precision Inference
Yun-Chen Lo, Tse-Kuang Lee, Ren-Shuo Liu |
ICLR | 3 |
| 2023 | Bucket Getter: A Bucket-based Processing Engine for Low-bit Block Floating Point (BFP) DNNsabstractBlock floating point (BFP), an efficient numerical system for deep neural networks (DNNs), achieves a good trade-off between dynamic range and hardware costs. Specifically, prior works have demonstrated that BFP format with 3 ∼ 5-bit mantissa can achieve FP32-comparable accuracy for various DNN workloads. We find that the floating-point adder (FP-Acc), which contains modules for normalization, alignment, addition, and fixed-point-to-floating-point (FXP2FP) conversion, dominates the power and area overheads, hence hindering the hardware efficiency of state-of-the-art low-bit BFP processing engines (BFP-PE). Yun-Chen Lo, Ren-Shuo Liu |
MICRO | 2 |
| 2023 | SG-Float: Achieving Memory Access and Computing Power Reduction Using Self-Gating Float in CNNsabstractConvolutional neural networks (CNNs) are essential for advancing the field of artificial intelligence. However, since these networks are highly demanding in terms of memory and computation, implementing CNNs can be challenging. To make CNNs more accessible to energy-constrained devices, researchers are exploring new algorithmic techniques and hardware designs that can reduce memory and computation requirements. In this work, we present self-gating float (SG-Float), algorithm hardware co-design of a novel binary number format, which can significantly reduce memory access and computing power requirements in CNNs. SG-Float is a self-gating format that uses the exponent to self-gate the mantissa to zero, exploiting the characteristic of floating-point that the exponent determines the magnitude of a floating-point value and the error tolerance property of CNNs. SG-Float represents relatively small values using only the exponent, which increases the proportion of ineffective mantissas, corresponding to reducing mantissa multiplications of floating-point numbers. To minimize the accuracy loss caused by the approximation error introduced by SG-Float, we propose a fine-tuning process to determine the exponent thresholds of SG-Float and reclaim the accuracy loss. We also develop a hardware optimization technique, called the SG-Float buffering strategy, to best match SG-Float with CNN accelerators and further reduce memory access. We apply the SG-Float buffering strategy to vector-vector multiplication processing elements (PEs), which NVDLA adopts, in TSMC 40nm technology. Our evaluation results demonstrate that SG-Float can achieve up to 35% reduction in memory access power and up to 54% reduction in computing power compared with AdaptivFloat, a state-of-the-art format, with negligible power and area overhead. Additionally, we show that SG-Float can be combined with neural network pruning methods to further reduce memory access and mantissa multiplications in pruned CNN models. Overall, our work shows that SG-Float is a promising solution to the problem of CNN memory access and computing power. Jun-Shen Wu, Tsen-Wei Hsu, Ren-Shuo Liu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | ISSA: Input-Skippable, Set-Associative Computing-in-Memory (SA-CIM) Architecture for Neural Network AcceleratorsabstractAmong several emerging architectures, computing in memory (CIM), which features in-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out that mutually stationary vectors (MSVs), which can be maximized by introducing associativity to CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. Yun-Chen Lo, Chih-Chen Yeh, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Wen-Chien Ting, Ren-Shuo Liu |
ICCAD | 7 |
| 2021 | Value-Aware Error Detection and Correction for SRAM Buffers in Low-Bitwidth, Floating-Point CNN AcceleratorsabstractLow-power CNN accelerators are a key technique to enable the future artificial intelligence world. Dynamic voltage scaling is an essential low-power strategy, but it is bottlenecked by on-chip SRAM. More specifically, SRAM can exhibit stuck-at (SA) faults at a rate as high as 0.1% when the supply voltage is lowered to, e.g., 0.5 V. Although this issue has been studied in CPU cache design, since their solutions are tailored for CPUs instead of CNN accelerators, they inevitably incur unnecessary design complexity and SRAM capacity overhead. Jun-Shen Wu, Chi-En Wang, Ren-Shuo Liu |
ASP-DAC | 3 |
| 2019 | Long-Term JPEG Data Protection and Recovery for NAND Flash-Based Solid-State StorageabstractNAND flash memory is widely used in solid-state storage including SD cards and eMMC chips, in which JPEG pictures are one of the most valuable data. In this work, we study NAND flash memory-aware, long-term JPEG data protection and recovery. Our goal is to increase the robustness of JPEG files stored in flash-based storage and rescue JPEG files that are corrupted due to long-term retention. JPEG files with our proposed protection techniques are compatible with existing JPEG viewers. We conduct real-system experiments by storing JPEG files on 16 nm, 3-bit-per-cell flash chips and let the JPEG files undergo a retention process equivalent to ten years at 25 degrees Celsius. Experimental results show that the proposed techniques can rescue corrupted JPEG files to achieve a PSNR improvement of up to 23.5 dB. Yu-Chun Kuo, Ruei-Fong Chiu, Ren-Shuo Liu |
MSST | 3 |
| 2019 | OPTR: Order-Preserving Translation and Recovery Design for SSDs with a Standard Block Device Interface
Yun-Sheng Chang, Ren-Shuo Liu |
USENIX ATC | 2 |
| 2018 | DI-SSD: Desymmetrized interconnection architecture and dynamic timing calibration for solid-state drivesabstractNAND flash-based solid-state drives (SSDs) have long been architected in the way that the interconnections between a flash controller and the associated flash memory chips operate at a symmetric speed in both directions. However, this commonly accepted and widely used architecture is suboptimal to SSDs because reading flash cells is 10 to 20x faster than writing them. In response, we propose desymmetrized interconnection SSD architecture (DI-SSD) and dynamic timing calibration (DTC), which selectively push the flash-to-controller speed to the limit. We conduct comprehensive experiments including characterizing real SSD products, using industrial-strength IC test equipment to emulate a flash controller that adopts DTC, and simulating DI-SSD using simulators to demonstrate the benefits of our proposals. Ren-Shuo Liu |
ASP-DAC | 2 |
| 2017 | VST: A virtual stress testing framework for discovering bugs in SSD flash-translation layersabstractFlash translation layers (FTLs) are the core embedded software (also known as firmware) of NAND flash-based solid-state drives (SSDs). The relentless pursuit of high-performance SSDs renders FTLs increasingly complex and intricate. Therefore, testing and validating FTLs are crucial and challenging tasks. Directly testing and validating FTLs on SSD hardware are common practices though, they are time-consuming and cumbersome because 1) the testing speed is limited by the hardware speed of SSDs and 2) just reproducing bugs can be challenging, let alone locating and root causing the bugs. This work presents virtual stress testing (VST), a simulation framework to enable executing SSD FTLs on PCs or servers against virtual SRAM, DRAM, and flash emulated by host-side main memory. FTL function calls, such as moving data from flash to DRAM, are served by the VST framework. Therefore, VST can test FTLs without SSD hardware requirements nor SSD speed limitations, and root causing bugs becomes manageable tasks. We apply VST to representative SSD design, OpenSSD, which is actively utilized and maintained by SSD and FTL communities. Experimental results show that VST can test FTLs at a speed up to 375 GB/s, which is several hundred times faster than directly testing FTLs on SSD hardware. Moreover, we successfully discover seven new FTL bugs in the OpenSSD design using VST, which is a solid evidence of VST's bug-discovering effectiveness. Ren-Shuo Liu, Yun-Sheng Chang, Chih-Wen Hung |
ICCAD | 1 |
| 2016 | Improving Read Performance of NAND Flash SSDs by Exploiting Error LocalityabstractNAND flash-based solid-state drives (SSDs), which can serve as the caches of hard disk drives, have gained popularity in large-scale, high-performance storage. A type of advanced error correction code for SSDs, low-density parity-check (LDPC), is required to mitigate a considerable number of errors in the raw data of NAND flash. However, LDPC imposes read performanceoverhead due to the complex decoding procedure of LDPC. In this paper, we propose an error-correcting cache (EC-Cache) that exploits “error locality”, a characteristic of NAND flash memory, to improve the read performance of SSDs. We use the term “error locality” to refer to the property that the majority of errors in reads to the same flash page appear at the same positions until thepage is erased. By caching detected errors, we can correct a significant portion of errors in the requested flash page prior to the LDPC decoding process. This design significantly reduces LDPC decoding overhead because the latency of LDPC is correlated with thenumber of errors in the input data. We conduct experiments, including flash characterization, LDPC simulation, and SSD simulation,to evaluate EC-Cache. The experimental results demonstrate that EC-Cache can improve the read performance of LDPC-based SSDs by up to$2.6\times$. Ren-Shuo Liu, Meng-Yen Chuang, Chia-Lin Yang, Cheng-Hsuan Li, Kin-Chu Ho, Hsiang-Pang Li |
IEEE Trans. Computers | 1 |
| 2014 | NVM duet: unified working memory and persistent store architectureabstractEmerging non-volatile memory (NVM) technologies have gained a lot of attention recently. The byte-addressability and high density of NVM enable computer architects to build large-scale main memory systems. NVM has also been shown to be a promising alternative to conventional persistent store. With NVM, programmers can persistently retain in-memory data structures without writing them to disk. Therefore, one can envision that in the future, NVM will play the role of both working memory and persistent store at the same time. Ren-Shuo Liu, De-Yu Shen, Chia-Lin Yang, Shun-Chih Yu, Cheng-Yuan Michael Wang |
ASPLOS | 1 |
| 2014 | EC-Cache: Exploiting Error Locality to Optimize LDPC in NAND Flash-Based SSDsabstractLow-density parity-check (LDPC) is widely accepted as the baseline error-correction codes offering strong error-correcting capability for future NAND flash-based SSDs. However, LDPC incurs read performance overhead because of its complex decoding procedure. To mitigate such overhead, we propose the error-correcting cache (EC-Cache) that exploits the "error locality" of NAND flash. Error locality means that the majority of errors in reads to the same NAND flash page appear in the same positions until the page is erased. By caching detected errors, EC-Cache can correct a significant portion of errors present in a requested flash page before the associated LDPC decoding process begins. EC-Cache can greatly speed up LDPC decoding because LDPC's latency is directly correlated to the number of errors present in the input data. Experimental results show that EC-Cache achieves up to 2.6× SSD read performance gain. Ren-Shuo Liu, Meng-Yen Chuang, Chia-Lin Yang, Cheng-Hsuan Li, Kin-Chu Ho, Hsiang-Pang Li |
DAC | 1 |
| 2013 | DuraCache: a durable SSD cache using MLC NAND flashabstractAdopting SSDs as caches for HDD arrays has gained popularity in datacenters because SSDs are superior in handling random reads that HDDs cannot efficiently deal with. Two types of flash memory cells are available for building SSD caches, single-level cells (SLC) and multi-level cells (MLC). MLC is more appealing than SLC because it can achieve higher cache capacity at the same cost. However, we see a critical issue for SSD caches to adopt MLC NAND flash: the endurance of modern MLC NAND flash is too low to sustain datacenter workloads. In this paper, we propose DuraCache that addresses the durability issue of SSD caches. DuraCache exploits the fact that SSD caches are write-through caches in datacenters. Therefore, uncorrectable errors in SSD caches can be handled like cache misses which bring in correct data from HDD arrays. In addition, DuraCache gradually allocates more ECC parities associated with data when NAND flash reaches wearout thresholds. This allows SSD caches to continue operating by sacrificing available capacity. We conduct empirical experiments and demonstrate that DuraCache enables MLC SSD caches to achieve 4.1 years of service life assuming a TPC-C workload. Ren-Shuo Liu, Chia-Lin Yang, Cheng-Hsuan Li, Geng-You Chen |
DAC | 1 |
| 2012 | Optimizing NAND flash-based SSDs via retention relaxation
Ren-Shuo Liu, Chia-Lin Yang |
FAST | 1 |