Hsiang-Pang Li

dblp:127/7936 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-6805-4767ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 In-3-D nand Flash Computing for Vector Similarity Search Acceleration on Edge Devices
abstract
Vector Similarity Search (VSS) on edge devices is increasingly essential for data privacy but faces substantial latency and energy overhead due to frequent data transfers between storage and DRAM. To address these challenges, we propose the Intelligent Cognition Engine (ICE), a fully digital non-volatile in-memory computing (nvIMC) framework designed for integration with commercial 3DnandFlash. ICE avoids the use of analog-digital conversion (ADC/DAC) and reduces data movement by performing vector similarity computation directly withinnandstorage. Key architectural components include a digital page multiplier, a two’s-complement accumulator that supports signed computation with minimal circuit modification, and a hierarchical Top-N search strategy to reduce unnecessary data accesses. The proposed framework is evaluated through a combination of post-layout circuit simulations, representative silicon measurements, and system-level modeling on edge platforms. Experimental results indicate that ICE can achieve 17.8–$122.6\times $speedup and 11.5–$162\times $improvement in energy efficiency compared to conventional von Neumann-based approaches, demonstrating the feasibility and scalability of fully digital 3Dnand-based nvIMC for edge AI workloads.
Han-Wen Hu, Yuan-Hao Chang 0001, Bo-Rong Lin, Huai-Mu Wang, Yung-Chun Lee, Hsiang-Pang Li, Chung Kuang Chen, Tei-Wei Kuo, Meng-Fan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 In-Storage Read-Centric Seed Location Filtering Using 3D-NAND Flash for Genome Sequence Analysis
abstract
Read mapping is a critical bottleneck in genome sequence analysis, requiring costly approximate string matching to identify potential matches between reads and a reference genome. Pre-alignment filtering methods aim to mitigate this issue by filtering out unnecessary mapping locations, and implementing them with processing-in-memory (PIM) approaches offers potential benefits by offloading filtering from the computing unit. However, the sparse number of potential mapping locations for each read limits the utilization of PIM's parallel computing capabilities, thereby hindering the overlapping of filtering and sequence alignment to hide filtering latency overheads. In this paper, we propose a 3D NAND-based in-storage pre-alignment filtering approach. Leveraging the read depth property, we introduce a read-centric pre-alignment filtering method that enables parallel comparison of multiple reads. We co-design software and hardware for in-situ processing of read-centric pre-alignment filtering within the storage, capitalizing on 3D NAND Flash's approximate parallel search capability. When integrating with a representative read mapping accelerator, our design achieves an average 1.36x performance improvement with comparable energy consumption. Compared to the state-of-the-art (SOTA) PIM solution, our design results in 123.8x and 53.3x performance gain and energy efficiency improvement.
You-Kai Zheng, Ming-Liang Wei, Hsiang-Yun Cheng, Chia-Lin Yang, Ming-Hsiang Tsai, Chia-Chun Chien, Yuan-Hao Zhong, Po-Hao Tseng, Hsiang-Pang Li
ASP-DAC9
2025 Accelerating Genome Alignment Pipeline with In-NAND Search Technology and Group Testing Techniques
abstract
Genomic sequence analysis deciphers and interprets an organism’s DNA, offering crucial insights into personalized medicine, disease diagnosis, evolutionary biology, and agricultural biotechnology. While Next-Generation Sequencing (NGS) has revolutionized genomics by providing a fast and cost-effective method for generating genomic sequences, the computational complexity of aligning short reads back to a reference genome remains a significant bottleneck. The exact-match-based preseeding filter has emerged as an effective and general methodology to address this issue, capable of removing 70% to 80% of exact-matched genomic reads at the source and applicable to a wide range of alignment tools. However, the state-of-the-art exact-match filter architecture, GenStore, encounters performance limitations due to the need to load reference sequences from NAND flash memory to the controller page by page.In this work, we propose a novel Solid-State Drive (SSD) architecture that leverages computing-in-NAND-flash techniques to perform match detection directly within memory. By harnessing the two-dimensional input capability of 3D NAND flash memory and integrating group testing methods, our design enables comparisons across hundreds of pages in a single read cycle and supports simultaneous multi-query searches. Combined with a Bloom filter for in-NAND search, our architecture significantly reduces data movement by 48% to 96%, achieves a speedup of 1.60× to 4.99× over GenStore, and delivers 30% higher energy efficiency with only a 4.5% circuit overhead.
Ming-Hsiang Tsai, Ming-Liang Wei, Chia-Chun Chien, Po-Hao Tseng, Yung-Chun Lee, Hsiang-Pang Li, Chia-Lin Yang
ICCAD6
2024 LUTIN: Efficient Neural Network Inference with Table Lookup
abstract
DNN models are becoming increasingly large and complex, but they are also being deployed on commodity devices that require low power and latency but lack specialized accelerators. We introduce LUTIN (LUT-based INference), which reduces the amount of matrix multiplication in DNN inference by converting it into table lookups. LUTIN's innovation is its use of hyperparameter optimization to refine the quantization process and vector partitioning, allowing it to run efficiently on a variety of hardware. By reducing off-chip memory lookups and designing a cache-efficient data layout, LUTIN reduces energy consumption while increasing the use of available CPU cache, even on devices with limited processing power. Our approach goes beyond the traditional limitations of 8-bit quantization, investigating lower bit-widths to further reduce LUT size while meeting accuracy requirements. Experimental results show that LUTIN achieves up to a 2.34x speedup in latency and a 2.04x improvement in energy efficiency over full-precision models.
Shi-Zhe Lin, Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo, Hsiang-Pang Li
ISLPED5
2023 A digital 3D TCAM accelerator for the inference phase of Random Forest
abstract
Random forest is a popular ensemble machine-learning algorithm for classification and regression tasks. However, the irregular tree shapes and non-deterministic memory access patterns make it hard for the current von Neumann architecture to handle random forest efficiently. This paper proposes a digital 3D TCAM-based accelerator for the random forest, adopting the idea of processing-in-memory (PIM) to reduce data movement. By utilizing this accelerator, we propose a TCAM-based approach to provide real-time inference with low energy consumption, making it suitable for edge or embedded environments. In the experiments, the proposed approach achieves an average of 3.13 times higher throughput with 22 times more energy saving than the GPU approach.
Chieh-Lin Tsai, Chun-Feng Wu, Yuan-Hao Chang 0001, Han-Wen Hu, Yung-Chun Lee, Hsiang-Pang Li, Tei-Wei Kuo
DAC6
2022 ICE: An Intelligent Cognition Engine with 3D NAND-based In-Memory Computing for Vector Similarity Search Acceleration
abstract
Vector similarity search (VSS) for unstructured vectors generated via machine learning methods is a promising solution for many applications, such as face search. With increasing awareness and concern about data security requirements, there is a compelling need to store data and process VSS applications locally on edge devices rather than send data to servers for computation. However, the explosive amount of data movement from NAND storage to DRAM across memory hierarchy and data processing of the entire dataset consume enormous energy and require long latency for VSS applications. Specifically, edge devices with insufficient DRAM capacity will trigger data swap and deteriorate the execution performance. To overcome this crucial hurdle, we propose an intelligent cognition engine (ICE) with cognitive 3D NAND, featuring non-volatile in-memory computing (nvIMC) to accelerate the processing, suppress the data movement, and reduce data swap between the processor and storage. This cognitive 3D NAND features digital nvIMC techniques (i. e., ADClDAC-free approach), high-density 3D NAND, and compatibility with standard 3D NAND products with minor modifications. To facilitate parallel INT8/INT4 vector-vector multiplication (VVM) and mitigate the reliability issue of 3D NAND, we develop a bit-error-tolerance data encoding and a two’s complement-based digital accumulator. VVM can support similarity computations (e.g., cosine similarity and Euclidean distance), which are required to search “the most similar data” right where they are stored. In addition, the proposed solution can be realized on edge storage products, e.g., embedded Multi-Media Card (eMMC). The measured and simulated results on real 3D NAND chips show that ICE enhances the system execution time by $17\times to 95\times$ and energy efficiency by $11\times to 140\times$, compared to traditional von Neumann approaches using state-of-the-art edge systems with MobileFaceNet on CASIA-WebFace dataset. To the best of our knowledge, this work demonstrates the first 3D NAND-based digital nvIMC technique with measured silicon data.
Han-Wen Hu, Wei-Chen Wang 0002, Yuan-Hao Chang 0001, Yung-Chun Lee, Bo-Rong Lin, Huai-Mu Wang, Yen-Po Lin, Chong-Ying Lee, Tzu-Hsiang Su, Chih-Chang Hsieh, Chia-Ming Hu, Yi-Ting Lai, Chung Kuang Chen, Han-Sung Chen, Hsiang-Pang Li, Tei-Wei Kuo, Meng-Fan Chang, Keh-Chung Wang, Chun-Hsiung Hung, Chih-Yuan Lu
MICRO16
2022 DL-RSIM: A Reliability and Deployment Strategy Simulation Framework for ReRAM-based CNN Accelerators
abstract
Memristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. In addition, due to the hardware constraints, the way to deploy neural network models on memristor crossbar arrays affects the computation parallelism and communication overheads. To enable reliable and energy-efficient memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit/device properties on the inference accuracy and the influence of different deployment strategies on performance and energy consumption. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. A rich set of reliability impact factors and deployment strategies are explored by DL-RSIM, and it can be incorporated with any deep learning neural networks implemented by TensorFlow. Using several representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and energy-efficient deployment strategies and develop optimization techniques accordingly.
Hsiang-Yun Cheng, Chia-Lin Yang, Meng-Yao Lin, Kai Lien, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang, Yen-Ting Tsou, Chin-Fu Nien
ACM Trans. Embed. Comput. Syst.8
2019 Sparse ReRAM engine: joint exploration of activation and weight sparsity in compressed neural networks
abstract
Exploiting model sparsity to reduce ineffectual computation is a commonly used approach to achieve energy efficiency for DNN inference accelerators. However, due to the tightly coupled crossbar structure, exploiting sparsity for ReRAM-based NN accelerator is a less explored area. Existing architectural studies on ReRAM-based NN accelerators assume that an entire crossbar array can be activated in a single cycle. However, due to inference accuracy considerations, matrix-vector computation must be conducted in a smaller granularity in practice, called Operation Unit (OU). An OU-based architecture creates a new opportunity to exploit DNN sparsity. In this paper, we propose the first practical Sparse ReRAM Engine that exploits both weight and activation sparsity. Our evaluation shows that the proposed method is effective in eliminating ineffectual computation, and delivers significant performance improvement and energy savings.
Tzu-Hsien Yang, Hsiang-Yun Cheng, Chia-Lin Yang, I-Ching Tseng, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li
ISCA7
2018 DL-RSIM: a simulation framework to enable reliable ReRAM-based accelerators for deep learning
abstract
Memristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. To enable reliable memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit and device properties on the inference accuracy. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. DL-RSIM simulates the error rates of every sum-of-products computation in the memristor-based accelerator and injects the errors in the targeted TensorFlow-based neural network model. A rich set of reliability impact factors are explored by DL-RSIM, and it can be incorporated with any deep learning neural network implemented by TensorFlow. Using three representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and develop reliability optimization techniques.
Meng-Yao Lin, Hsiang-Yun Cheng, Tzu-Hsien Yang, I-Ching Tseng, Chia-Lin Yang, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang
ICCAD9
2017 A Hybrid DRAM/PCM Buffer Cache Architecture for Smartphones with QoS Consideration
abstract
Flash memory is widely used in mobile phones to store contact information, application files, and other types of data. In an operating system, the buffer cache keeps the I/O blocks in dynamic random access memory (DRAM) to reduce the slow flash accesses. However, in smartphones, we observed two issues which reduce the benefits of the buffer cache. First, a large number of synchronous writes force writing the data from the buffer cache to flash frequently. Second, the large amount of I/O accesses from background applications diminishes the buffer cache efficiency of the foreground application, which degrades the quality-of-service (QoS). In this article, we propose a buffer cache architecture with hybrid DRAM and phase change memory (PCM) memory, which improves the I/O performance and QoS for smartphones. We use a DRAM first-level buffer cache to provide high buffer cache performance and a PCM last-level buffer cache to reduce the impact of frequent synchronous writes. Based on the proposed hierarchical buffer cache architecture, we propose a sub-block management and background flush to reduce the impact of the PCM write limitation and the dirty block write-back overhead, respectively. To improve the QoS, we propose a least-recently-activated first replacement policy (LRA) to keep the data from the applications that are most likely to become the foreground one. The experimental results show that with the proposed mechanisms, our hierarchical buffer cache can improve the I/O response time by 20% compared to the conventional buffer cache. The proposed LRA can improve the foreground application performance by 1.74x compared to the conventional CLOCK policy.
Ye-Jyun Lin, Chia-Lin Yang, Hsiang-Pang Li, Cheng-Yuan Michael Wang
ACM Trans. Design Autom. Electr. Syst.3
2016 Disturbance Relaxation for 3D Flash Memory
abstract
Even though 3D flash memory presents a grand opportunity to huge-capacity non-volatile memory, it suffers from serious program disturbance problems. In contrast to the past efforts in error correction codes and the work in trading the space utilization for reliability, we propose a disturbance-relaxation scheme that can alleviate the negative effects caused by program disturbance inside a physical block. This scheme does not introduce any extra overheads on encoding or storing of extra redundant data. In particular, a methodology is proposed to reduce the data error rate by distributing unavoidable disturbance errors to the flash-memory space of invalid data, with the considerations of the physical organization of 3D flash memory. A series of experiments was conducted based on real multi-layer 3D flash chips, and it showed that the proposed scheme could significantly enhance the reliability of 3D flash memory.
Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Yung-Chun Li, Hsiang-Pang Li
IEEE Trans. Computers5
2016 Improving Read Performance of NAND Flash SSDs by Exploiting Error Locality
abstract
NAND flash-based solid-state drives (SSDs), which can serve as the caches of hard disk drives, have gained popularity in large-scale, high-performance storage. A type of advanced error correction code for SSDs, low-density parity-check (LDPC), is required to mitigate a considerable number of errors in the raw data of NAND flash. However, LDPC imposes read performanceoverhead due to the complex decoding procedure of LDPC. In this paper, we propose an error-correcting cache (EC-Cache) that exploits “error locality”, a characteristic of NAND flash memory, to improve the read performance of SSDs. We use the term “error locality” to refer to the property that the majority of errors in reads to the same flash page appear at the same positions until thepage is erased. By caching detected errors, we can correct a significant portion of errors in the requested flash page prior to the LDPC decoding process. This design significantly reduces LDPC decoding overhead because the latency of LDPC is correlated with thenumber of errors in the input data. We conduct experiments, including flash characterization, LDPC simulation, and SSD simulation,to evaluate EC-Cache. The experimental results demonstrate that EC-Cache can improve the read performance of LDPC-based SSDs by up to$2.6\times$.
Ren-Shuo Liu, Meng-Yen Chuang, Chia-Lin Yang, Cheng-Hsuan Li, Kin-Chu Ho, Hsiang-Pang Li
IEEE Trans. Computers6
2015 Achieving SLC performance with MLC flash memory
abstract
Although the Multi-Level-Cell technique is widely adopted by flash-memory vendors to boost the chip density and to lower the cost, it results in serious performance and reliability problems. Different from the past work, a new cell programming method is proposed to not only significantly improve the chip performance but also reduce the potential bit error rate. In particular, a Single-Level-Cell-like programming style is proposed to better explore the threshold-voltage relationship to denote different Multi-Level-Cell bit information, which in turn drastically provides a larger window of threshold voltage similar to that found in Single-Level-Cell chips. It could result in less programming iterations and simultaneously a much less reliability problem in programming flash-memory cells. In the experiments, the new programming style could accelerate the programming speed up to 742% and even reduce the bit error rate up to 471% for Multi-Level-Cell pages.
Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Yung-Chun Li, Hsiang-Pang Li
DAC5
2015 A Light-Weighted Software-Controlled Cache for PCM-based Main Memory Systems
abstract
The replacement of DRAM with non-volatile memory relies on solutions to resolve the wear leveling and slow write problems. Different from the past work in compiler-assisted optimization or joint DRAM-PCM management strategies, we explore a light-weighted software-controlled DRAM cache design for the non-volatile-memory-based main memory. The run-time overheads in the management of the DRAM cache is minimized by utilizing the information from a miss of the translation lookaside buffer (TLB) or the cache. Experiments were conducted based on a series of the well-known benchmarks to evaluate the effectiveness of the proposed design, for which the results are very encouraging.
Hung-Sheng Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Hsiang-Pang Li
ICCAD4
2015 On Relaxing Page Program Disturbance over 3D MLC Flash Memory
abstract
With the rapidly-increasing capacity demand over flash memory, 3D NAND flash memory has drawn tremendous attention as a promising solution to further reduce the bit cost and to increase the bit density. However, such advanced 3D devices will suffer more intensive program disturbance, compared to 2D NAND flash memory. Especially when multi-level-cell (MLC) technology is adopted, the deteriorated disturbance due to the program operations of intra and inter pages will become even more critical for reliability. In contrast to the past efforts that try to resolve the reliability issue with error correction codes or hardware designs, this work seeks for the redesign of the program operation. A disturb-aware programming scheme is proposed to not only relax the disturbance induced by slow cells as much as possible but also reduce the possibility in requiring a high voltage to program the slow cells. A series of experiments was conducted based on real 3D MLC flash chips, and the results demonstrate that the proposed scheme is extremely effective on reducing the disturbance as well as the bit error rate.
Yu-Ming Chang, Yung-Chun Li, Yuan-Hao Chang 0001, Tei-Wei Kuo, Chih-Chang Hsieh, Hsiang-Pang Li
ICCAD6
2015 Fine-grained write scheduling for PCM performance improvement under write power budget
abstract
Phase-change memory (PCM) has gained much attention recently since it offers several advantages over DRAM, such as high cell density and low leakage power. PCM has similar read power and latency as DRAM; however, its write power and latency are significantly higher than DRAM. Therefore, one challenge with PCM is how to increase write throughput under write power budget constraints. To increase write concurrency, PCM often adopts division programming, where a write occurs in a series of divisions, so that writes to different banks proceed concurrently. In this study, we observe that since the write scheduling granularity in the memory controller differs from the actual write granularity in PCM chips, i.e., requests vs. divisions, the available power budget cannot be fully utilized. We therefore propose enhancing the interface between the memory controller and PCM chips to allow the memory controller to schedule writes in the division granularity. To further increase power budget utilization, we design a variable-length division mechanism to allow the division granularity to be adjusted at runtime according to the available write power budget. Our experimental results show that these techniques improve system performance by up to 65%.
Chun-Hao Lai, Shun-Chih Yu, Chia-Lin Yang, Hsiang-Pang Li
ISLPED4
2015 Marching-Based Wear-Leveling for PCM-Based Storage Systems
abstract
Improving the performance of storage systems without losing the reliability and sanity/integrity of file systems is a major issue in storage system designs. In contrast to existing storage architectures, we consider a PCM-based storage architecture to enhance the reliability of storage systems. In PCM-based storage systems, the major challenge falls on how to prevent the frequently updated (meta)data from wearing out their residing PCM cells without excessively searching and moving metadata around the PCM space and without extensively updating the index structures of file systems. In this work, we propose an adaptive wear-leveling mechanism to prevent any PCM cell from being worn out prematurely by selecting appropriate data for swapping with constant search/sort cost. Meanwhile, the concept of indirect pointers is designed in the proposed mechanism to swap data without any modification to the file system's indexes. Experiments were conducted based on well-known benchmarks and realistic workloads to evaluate the effectiveness of the proposed design, for which the results are encouraging.
Hung-Sheng Chang, Yuan-Hao Chang 0001, Pi-Cheng Hsiu, Tei-Wei Kuo, Hsiang-Pang Li
ACM Trans. Design Autom. Electr. Syst.5
2014 On Trading Wear-leveling with Heal-leveling
abstract
Manufacturers are constantly seeking to increase flash memory density in order to fulfill the ever growing demand for storage capacity. However, this trend significantly reduces the reliability and endurance of flash memory chips. The lifetime degradation worsens as the number of erase cycles grows, even with wear leveling technology being adopted to extend flash memory lifetime by evenly distributing erase cycles to every flash block. To address this issue, self-healing technology is proposed to recover a flash block before the flash block is worn out, but such a technology still has its limitation when recovering flash blocks. In contrast to the existing wear leveling designs, we adopt the self-healing technology to propose a heal-leveling design that evenly distributes healing cycles to flash blocks. Ultimately, heal-leveling aims to extend the lifetime of flash memory without introducing a large amount of live-data copying overheads. We conducted a series of experiments to evaluate the capability of the proposed design. The results show that our design can significantly improve the access performance and the effective lifetime of flash memory without the unnecessary overheads caused by wear leveling technology.
Yu-Ming Chang, Yuan-Hao Chang 0001, Jian-Jia Chen, Tei-Wei Kuo, Hsiang-Pang Li, Hang-Ting Lue
DAC5
2014 EC-Cache: Exploiting Error Locality to Optimize LDPC in NAND Flash-Based SSDs
abstract
Low-density parity-check (LDPC) is widely accepted as the baseline error-correction codes offering strong error-correcting capability for future NAND flash-based SSDs. However, LDPC incurs read performance overhead because of its complex decoding procedure. To mitigate such overhead, we propose the error-correcting cache (EC-Cache) that exploits the "error locality" of NAND flash. Error locality means that the majority of errors in reads to the same NAND flash page appear in the same positions until the page is erased. By caching detected errors, EC-Cache can correct a significant portion of errors present in a requested flash page before the associated LDPC decoding process begins. EC-Cache can greatly speed up LDPC decoding because LDPC's latency is directly correlated to the number of errors present in the input data. Experimental results show that EC-Cache achieves up to 2.6× SSD read performance gain.
Ren-Shuo Liu, Meng-Yen Chuang, Chia-Lin Yang, Cheng-Hsuan Li, Kin-Chu Ho, Hsiang-Pang Li
DAC6
2013 A disturb-alleviation scheme for 3D flash memory
abstract
Even though 3D flash memory presents a grand opportunity for huge-capacity non-volatile memory, it suffers from serious program disturb problems. Different from the past efforts in error correction codes or the work in trading the space utilization with reliability, we propose a disturb-alleviation scheme that can alleviate the negative effects caused by program disturb, especially inside a block, without introducing extra overheads on encoding or storing of extra redundant data. In particular, a methodology is proposed to reduce the data error rate by distributing unavoidable disturb errors over the flash-memory space of invalid data, with the considerations of the physical organization of 3D flash memory. A series of experiments was conducted based on real multi-layer 3D flash chips, and it showed that the proposed scheme could significantly enhance the reliability of 3D flash memory.
Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Hsiang-Pang Li, Yung-Chun Li
ICCAD4