Hsiang-Yun Cheng

dblp:19/9032 · DBLP profile ↗
← Back
34ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-8983-9835ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 7 first-author · 20 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Tris-GCN: A 3D NAND Flash-based In-Storage Processing Architecture for GCN Acceleration
abstract
Graph convolutional networks (GCNs) excel in many applications, but scaling to large graphs is bottlenecked by heavy data movement. Existing in-storage processing (ISP) solutions offload I/O-intensive operations to the SSD controller to reduce PCIe traffic, but limited parallelism and flash bandwidth still constrain energy efficiency, even with accuracy-degrading neighbor sampling. We propose Tris-GCN, a 3D NAND-based ISP design that executes core GCN computations in situ by leveraging inherent flash computing capabilities, reducing channel traffic without accuracy loss. It incorporates mapping and scheduling optimizations for energy efficiency, alongside wear-leveling with selective recomputation for reliability. Results show that Tris-GCN achieves average 11.7× speedup (up to 41.0×) and 99.7% energy savings over CPU baselines, and 2.45× speedup and 63.8% energy savings over the SOTA ISP on the Amazon dataset.
Yi-Wa Wu, Jia-You Li, Chi-Jung Chen, Ching (Ryan) Cheng, Chia-Chun Wang, Chin-Fu Nien, Hsiang-Yun Cheng
ISLPED7
2026 Energy-Efficient In-Memory Vector Similarity Search With Fidelity-Aware Range Encoding
abstract
Vector similarity search (VSS) is a crucial operation in applications of machine learning, but it often incurs high energy consumption due to frequent memory accesses. Previous works have adopted ternary content-addressable memories (TCAMs) to perform parallel VSS within memory. Among these approaches, Exact-Match TCAM (EX-TCAM) combined with range encoding has shown promise for executing in-memory VSS under theL∞norm distance. Existing EX-TCAM-based approaches initiate the search from the positions of query vectors and iteratively expand the search ranges to identify the closest stored vectors. However, this initialization strategy leads to excessive search iterations and long codewords. To address these challenges, we present FORE, an EX-TCAM-based framework that significantly improves latency, energy efficiency, and accuracy for in-memory VSS. To reduce redundant search iterations, we first initialize the search from ranges based on theL∞norm distance. To shorten the codeword length, we then develop a range encoding scheme that supports range-to-range matching. In addition, we introduce a novel metric, “range fidelity,” to evaluate the quality of range encoding. Building on the insight that a certain degree of loss in range fidelity is tolerable in EX-TCAM-based VSS, we further propose a lossy range encoding scheme that yields more compact codewords without significantly compromising accuracy. Finally, FORE incorporates aL∞norm distance-based training mechanism that further reduces search iterations and enhances classification accuracy in EX-TCAM-based VSS. Experimental results demonstrate that FORE improves energy efficiency by 45.25× to 59.72× and reduces latency by 4.47× to 5.66× compared to previous EX-TCAM-based methods within 1% accuracy loss. Moreover, FORE outperforms Best-Match TCAM (Best-TCAM)-based approaches by 6.9× in energy efficiency under non-ideal conditions.
Chi-Tse Huang, Hsiang-Yun Cheng, An-Yeu Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Efficient and Reliable Vector Similarity Search Using Asymmetric Encoding with NAND-Flash for Many-Class Few-Shot Learning
abstract
While memory-augmented neural networks (MANNs) offer an effective solution for few-shot learning (FSL) by integrating deep neural networks with external memory, the capacity requirements and energy overhead of data movement become enormous due to the large number of support vectors in many-class FSL scenarios. Various in-memory search solutions have emerged to improve the energy efficiency of MANNs. NAND-based multi-bit content addressable memory (MCAM) is a promising option due to its high density and large capacity. Despite its potential, MCAM faces limitations such as a restricted number of word lines, limited quantization levels, and non-ideal effects like varying string currents and bottleneck effects, which lead to significant accuracy drops. To address these issues, we propose several innovative methods. First, the Multi-bit Thermometer Code (MTMC) leverages the extensive capacity of MCAM to enhance vector precision using cumulative encoding rules, thereby mitigating the bottleneck effect. Second, the Asymmetric Vector Similarity Search (AVSS) reduces the precision of the query vector while maintaining that of the support vectors, thereby minimizing the search iterations and improving efficiency in many-class scenarios. Finally, the Hardware-Aware Training (HAT) method optimizes controller training by modeling the hardware characteristics of MCAM, thus enhancing the reliability of the system. Our integrated framework reduces search iterations by up to 32×, and increases overall accuracy by 1.58% to 6.94%.
Hao-Wei Chiang, Chi-Tse Huang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Wu
ASP-DAC3
2025 ReTAP: Processing-in-ReRAM Bitap Approximate String Matching Accelerator for Genomic Analysis
abstract
Read mapping, which involves computationally intensive approximate string matching (ASM) on large datasets, is the primary performance bottleneck in genome sequence analysis. To accelerate read mapping, a processing-in-memory (PIM) architecture that conducts highly parallel computations within the memory to reduce energy-inefficient data movements can be a promising solution. In this paper, we present ReTAP, a processing-in-ReRAM Bitap accelerator for genomic analysis. Instead of using the intricate dynamic programming algorithm, our design incorporates the Bitap algorithm, which uses only simple bitwise operations to perform ASM. Additionally, we explore the opportunity to reduce redundant computations by dynamically adjusting the error tolerance of Bitap and co-design the hardware to enhance computation parallelism. Our evaluation demonstrates that ReTAP outperforms GenASM, the state-of-the-art Bitap accelerator, with a 5.74× throughput and 1.79× energy efficiency.
Tsung-Yu Liu, Yen-An Lu, James Yu, Chin-Fu Nien, Hsiang-Yun Cheng
ASP-DAC5
2025 In-Storage Read-Centric Seed Location Filtering Using 3D-NAND Flash for Genome Sequence Analysis
abstract
Read mapping is a critical bottleneck in genome sequence analysis, requiring costly approximate string matching to identify potential matches between reads and a reference genome. Pre-alignment filtering methods aim to mitigate this issue by filtering out unnecessary mapping locations, and implementing them with processing-in-memory (PIM) approaches offers potential benefits by offloading filtering from the computing unit. However, the sparse number of potential mapping locations for each read limits the utilization of PIM's parallel computing capabilities, thereby hindering the overlapping of filtering and sequence alignment to hide filtering latency overheads. In this paper, we propose a 3D NAND-based in-storage pre-alignment filtering approach. Leveraging the read depth property, we introduce a read-centric pre-alignment filtering method that enables parallel comparison of multiple reads. We co-design software and hardware for in-situ processing of read-centric pre-alignment filtering within the storage, capitalizing on 3D NAND Flash's approximate parallel search capability. When integrating with a representative read mapping accelerator, our design achieves an average 1.36x performance improvement with comparable energy consumption. Compared to the state-of-the-art (SOTA) PIM solution, our design results in 123.8x and 53.3x performance gain and energy efficiency improvement.
You-Kai Zheng, Ming-Liang Wei, Hsiang-Yun Cheng, Chia-Lin Yang, Ming-Hsiang Tsai, Chia-Chun Chien, Yuan-Hao Zhong, Po-Hao Tseng, Hsiang-Pang Li
ASP-DAC3
2025 Energy-Efficient Large-Scale Vector Similarity Search in NAND-Flash via Hybrid Matching
abstract
Vector similarity search (VSS) is crucial in many AI applications, such as few-shot learning (FSL) and approximate nearest-neighbor search (ANNS), but it demands significant memory capacity and incurs substantial energy costs for data transfers during large-scale comparisons. Various in-memory search technologies have been developed to improve energy efficiency, with NAND-based multi-bit content-addressable memory (MCAM) standing out as a promising solution for its high density and large capacity. MCAM can operate in exact-search (ES) mode, supporting only perfect matches with low energy cost, or in approximate-search (AS) mode, enabling flexible VSS. However, AS mode incurs significant energy waste when comparing queries with non-target stored vectors. To address this issue, we propose Hybrid-M, a 3D NAND-based in-memory VSS architecture that integrates both modes into a single hybrid matching process, using ES mode as a filter to reduce redundant searches for AS mode. We apply three techniques to optimize this integration: range encoding for multi-level cells (MLC) to enhance filtering, search voltage shifts to mitigate the impact on AS accuracy and reduce matching currents, and a filtering-aware training method to further improve reliability and energy efficiency. Results show that Hybrid-M achieves comparable accuracy while reducing energy consumption by 67% to 83% compared to MACM-based VSS using only AS mode, across various many-class FSL and ANNS workloads.
Chih-Yu Hu, Chi-Tse Huang, Hao-Wei Chiang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Wu
DAC4
2025 Segmented Angular Pre-Processing for Accurate and Efficient In-Memory Vector Similarity Search
abstract
Vector similarity search (VSS) is a fundamental operation in modern AI applications, including few-shot learning (FSL) and approximate nearest neighbor search (ANNS). Cosine similarity is widely regarded as the optimal metric for VSS. However, VSS incurs substantial energy and computational overhead, primarily due to frequent vector transfers and the complexity of cosine similarity calculations in high-dimensional spaces. Prior research has explored the use of ternary content addressable memories (TCAMs) for parallel in-memory VSS to reduce vector movement. Exact-Match TCAM (EX-TCAM) enables exact bitmatching, and Best-Match TCAM (Best-TCAM) supports Hamming distance calculations, both of which are spatial metrics and computationally efficient. As a result, existing TCAM-based VSS approaches have focused on developing frameworks to efficiently support more complex spatial metrics such as the $L_{\infty}$ and $L_{1}$ norms. However, these spatial metrics exhibit notable discrepancies compared to angular metrics like cosine similarity. To overcome this limitation, we propose Seg-Cos, a TCAM-based framework that directly approximates cosine similarity within TCAM for angular VSS. Seg-Cos introduces a dedicated preprocessing technique and encoding scheme that segments vectors and encodes them as circular ranges based on their angles and magnitudes. Seg-Cos is the first angular VSS framework compatible with both EX-TCAM and Best-TCAM, enabling accurate and energy-efficient VSS in the angular domain. Simulation results demonstrate that Seg-Cos improves energy efficiency by $1.41 \times$ and achieves up to $2.2 \%$ higher accuracy over prior EX-TCAMbased methods in FSL. In ANNS, Seg-Cos enhances recall rate by $10 \%$ to $52 \%$ and improves energy efficiency by $2 \times$ compared to previous Best-TCAM approaches with $L_{1}$ norm.
Chi-Tse Huang, Jen-Chieh Wang, Hsiang-Yun Cheng, An-Yeu Wu
DAC3
2025 Filter-Based Adaptive Model Pruning for Efficient Incremental Learning on Edge Devices
abstract
Incremental Learning (IL) enhances Machine Learning (ML) models over time with new data, ideal for edge devices at the forefront of data collection. However, executing IL on edges faces challenges due to limited resources. Common methods involve IL followed by model pruning or specialized IL methods for edges. However, the former increases training time due to fine-tuning and compromises accuracy for past classes due to limited retained samples or features. Meanwhile, existing edge-specific IL methods utilize weight pruning, which requires specialized hardware or compilers to speed up and cannot reduce computations on general embedded platforms. In this paper, we propose Filter-based Adaptive Model Pruning (FAMP), the first pruning method designed specifically for IL. FAMP prunes the model before the IL process, allowing fine-tuning to occur concurrently with IL, thereby avoiding extended training time. To maintain high accuracy for both new and past data classes, FAMP adapts the compressed model based on observed data classes and retains filter settings from the previous IL iteration to mitigate forgetting. Across all tests, FAMP achieves the best average accuracy, with only a 2.78% accuracy drop over full ML models with IL. Moreover, unlike the common methods that prolong training time, FAMP takes 35% shorter training time on average than using the full ML models for IL.
Jing-Jia Hung, Yi-Jung Chen, Hsiang-Yun Cheng, Hsu Kao, Chia-Lin Yang
DATE3
2024 BORE: Energy-Efficient Banded Vector Similarity Search with Optimized Range Encoding for Memory-Augmented Neural Network
abstract
Memory-augmented neural networks (MANNs) in-corporate external memories to address the significant issue of catastrophic forgetting in few-shot learning applications. MANNs rely on vector similarity search (VSS), which incurs substantial energy and computational overhead due to frequent data transfers and complex cosine similarity calculations. To tackle these challenges, prior research has proposed adopting ternary content addressable memories (TCAMs) for parallel VSS within memory. One promising approach is to use Exact-Match TCAM (EX-TCAM) with range encoding to find the vector with the minimum$L_{\infty}$distance, avoiding the need for sensing circuit modifications as required by Best-Match TCAM (Best-TCAM), However, this method demands multiple search iterations and longer code words, limiting its practicality. In this paper, we propose an energy-efficient EX-TCAM-based design called BORE. BORE skips redundant search iterations and reduces code word length through performing Banded$L_{\infty}$distance search with Optimized Range Encoding. Additionally, we consider the characteristics of the similarity metric and develop a distance-based training mechanism aimed at improving classification accuracy. Simulation results demonstrate that BORE enhances energy efficiency by$9.35\times$to$11.84\times$and accuracy by 2.95 % to 4.69 % compared to previous EX-TCAM-based approaches. Furthermore, BORE improves energy efficiency by$1.04\times$to$1.63\times$over prior works of Best-TCAM-based VSS.
Chi-Tse Huang, Cheng-Yang Chang, Hsiang-Yun Cheng, An-Yeu Wu
DATE3
2024 ReTAP: Processing-in-ReRAM Bitap Approximate String Matching Accelerator for Genomic Analysis
abstract
Read mapping, which involves computationally in-tensive approximate string matching (ASM) on large datasets, is the primary performance bottleneck in genome sequence analysis. To accelerate read mapping, a processing-in-memory (PIM) architecture that conducts highly parallel computations within the memory to reduce energy-inefficient data movements can be a promising solution. In this paper, we present ReTAP, a processing-in-ReRAM Bitap accelerator for genomic analysis. Instead of using the intricate dynamic programming algorithm, our design incorporates the Bitap algorithm, which uses only simple bitwise operations to perform ASM. Additionally, we explore the opportunity to reduce redundant computations by dynamically adjusting the error tolerance of Bitap and co-design the hardware to enhance computation parallelism. Our evaluation demonstrates that ReTAP outperforms GenASM, the state-of-the-art Bitap accelerator, with a 153.7 x higher throughput.
Tsung-Yu Liu, Yen-An Lu, James Yu, Chin-Fu Nien, Hsiang-Yun Cheng
DATE5
2024 ReAIM: A ReRAM-based Adaptive Ising Machine for Solving Combinatorial Optimization Problems
abstract
Recently, in light of the success of quantum computers, research teams have actively developed quantum-inspired computers using classical computing technology. One notable success story is the Ising Machine (IM), which excels in efficiently solving NP-hard combinatorial optimization problems (COPs) in various domains such as finance, drug development, and logistics. However, IMs may encounter the von Neumann bottleneck due to significant data transfers between computing and memory units. To tackle this issue, processing-in-memory (PiM) leveraging resistive random-access memory (ReRAM) has emerged as a potential solution. Nonetheless, existing ReRAM-based IM accelerators face challenges stemming from non-ideal ReRAM devices and the limited flexibility in selecting solver algorithms. To overcome these limitations, we propose ReAIM, a ReRAMbased Adaptive Ising Machine, co-designed with an adaptive parameter search algorithm to dynamically select an appropriate solver algorithm and hardware/software parameters for the target COP. Based on our in-depth analysis, ReAIM accounts for the impact of both software parameters, such as the choice of local search algorithm and the number of spin flips, and hardware characteristics, including ReRAM-induced errors. Its reconfigurable architecture and optimized mapping/scheduling strategies enable efficient execution of the adaptive algorithm across various COPs, even for large-scale problems. Compared to the state-of-the-art SRAM-based IM accelerator, ReAIM achieves a $2.2 \times$ shorter time-to-solution (TTS), showcasing superior hardware performance without compromising solution quality.
Hao-Wei Chiang, Chin-Fu Nien, Hsiang-Yun Cheng, Kuei-Po Huang
ISCA3
2023 Special Session - Non-Volatile Memories: Challenges and Opportunities for Embedded System Architectures with Focus on Machine Learning Applications
abstract
This paper explores the challenges and opportunities of integrating non-volatile memories (NVMs) into embedded systems for machine learning. NVMs offer advantages such as increased memory density, lower power consumption, non-volatility, and compute-in-memory capabilities. The paper focuses on integrating NVMs into embedded systems, particularly in intermittent computing, where systems operate during periods of available energy. NVM technologies bring persistence closer to the CPU core, enabling efficient designs for energy-constrained scenarios. Next, computation in resistive NVMs is explored, highlighting its potential for accelerating machine learning algorithms. However, challenges related to reliability and device non-idealities need to be addressed. The paper also discusses memory-centric machine learning, leveraging NVMs to overcome the memory wall challenge. By optimizing memory layouts and utilizing probabilistic decision tree execution and neural network sparsity, NVM-based systems can improve cache behavior and reduce unnecessary computations. In conclusion, the paper emphasizes the need for further research and optimization for the widespread adoption of NVMs in embedded systems presenting relevant challenges, especially for machine learning applications.
Jörg Henkel, Lokesh Siddhu, Lars Bauer, Jürgen Teich, Stefan Wildermann, Mehdi Baradaran Tahoori, Mahta Mayahinia, Jerónimo Castrillón, Asif Ali Khan, Hamid Farzaneh, João Paulo C. de Lima, Jian-Jia Chen, Christian Hakert, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng
CASES16
2023 Tensor Movement Orchestration in Multi-GPU Training Systems
abstract
As deep neural network (DNN) models grow deeper and wider, one of the main challenges for training large-scale neural networks is overcoming limited GPU memory capacity. One common solution is to utilize the host memory as the external memory for swapping tensors in and out of GPU memory. However, the effectiveness of such tensor swapping can be impaired in data-parallel training systems due to contention on the shared PCIe channel to the host. In this paper, we propose the first large-model support framework that coordinates tensor movements among GPUs to alleviate PCIe channel contention. We design two types of coordination mechanisms. In the first mechanism, PCIe channel accesses from different GPUs are interleaved by selecting disjoint swapped-out tensors for each GPU. In the second method, swap commands are orchestrated to avoid contention. The effectiveness of these two methods depends on the model size and how often the GPUs synchronize on gradients. Experimental results show that compared to large-model support that is oblivious to channel contention, the proposed solution achieves average speedups of 38.3% to 31.8% when the memory footprint size is 1.33 to 2 times the GPU memory size.
Shao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin Yang
HPCA3
2022 This is SPATEM! A Spatial-Temporal Optimization Framework for Efficient Inference on ReRAM-based CNN Accelerator
abstract
Resistive memory-based computing-in-memory (CIM) has been considered as a promising solution to accelerate convolutional neural networks (CNN) inference, which stores the weights in crossbar memory arrays and performs in-situ matrix-vector multiplications (MVMs) in an analog manner. Several techniques assume that a whole crossbar can operate concurrently and discuss how to efficiently map the weights onto crossbar arrays. However, in practice, the accumulated effect of per-cell current deviation and Analog-to-Digital-Converter overhead may greatly degrade inference accuracy, which motivates the concept of Operation Unit (OU), by which an operation per cycle in a crossbar only involve limited wordlines and bitlines to preserve satisfactory inference accuracy. With OU-based operations, the mapping of weights and scheduling strategy for parallelizing CNN convolution operations should take the cost of communication overhead and resource utilization into consideration to optimize the inference acceleration. In this work, we propose the first optimization framework named SPATEM, that efficiently executes MVMs with OU-based operations on ReRAM-based CIM accelerators. It decouples the design space into tractable steps, models the expected inference latency, and derives an optimized spatial-temporal-aware scheduling strategy. By comparing with state-of-the-arts, the experimental result shows that the derived scheduling strategy of SPATEM achieves on average 29.24% inference latency reduction with 31.28% less communication overhead by exploiting more originally unused crossbar cells.
Yen-Ting Tsou, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng, Jian-Jia Chen, Der-Yu Tsai
ASP-DAC4
2022 RePAIR: A ReRAM-based Processing-in-Memory Accelerator for Indel Realignment
abstract
Genomic analysis has attracted a lot of interest recently since it is the key to realizing precision medicine for diseases such as cancer. Among all the genomic analysis pipeline stages, Indel Realignment is the most time-consuming and induces intensive data movements. Thus, we propose RePAIR, the first ReRAM-based processing-in-memory accelerator targeting the Indel Realignment algorithm. To further increase the computation parallelism, we design several mapping and scheduling optimization schemes. RePAIR achieves 7443× speedup and is 27211× more energy efficient over the GATK3.8 running on a CPU server, significantly outperforming the state-of-the-art.
Chin-Fu Nien, Kuang-Chao Chou, Hsiang-Yun Cheng
DATE4
2022 Efficient Bad Block Management with Cluster Similarity
abstract
Process variation in the 3D flash memory architecture raises the difficulty of bad block management. Since the error characteristics vary among different blocks, it is difficult for the existing P/E cycle-based bad block management policies to decide a suitable cycle threshold. This increases the possibility of data loss and decreases the SSD’s lifetime. In this work, we characterize the 3D flash memory and observe spatial correlation among flash blocks in the aspect of error behaviors. This phenomenon is referred to as cluster similarity. A novel cluster-based bad block management policy is proposed, which treats the failure of a block as an indicator of near-future failures of its neighboring blocks. Moreover, we provide quantitative methods to enable judicious selection of the cluster size to meet the desired tradeoff between the SSD lifetime and reliability. Compared with the commonly-used cycle-based bad block management policy, our cluster-based management policy has a lifetime improvement of 2x with comparable failure rates. And with comparable lifetime, the failure rate of the cycle-based policy is 9x higher than our method. To alleviate the I/O performance impact caused by the cluster retirement, we proposes a critical-block first reallocation scheduling. Our experiments show up to two times improvement of the 95th percentile latency compared to the naive scheduling of cluster reallocation.
Jui-Nan Yen, Tseng-Yi Chen, Chia-Lin Yang, Hsiang-Yun Cheng
HPCA6
2022 DL-RSIM: A Reliability and Deployment Strategy Simulation Framework for ReRAM-based CNN Accelerators
abstract
Memristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. In addition, due to the hardware constraints, the way to deploy neural network models on memristor crossbar arrays affects the computation parallelism and communication overheads. To enable reliable and energy-efficient memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit/device properties on the inference accuracy and the influence of different deployment strategies on performance and energy consumption. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. A rich set of reliability impact factors and deployment strategies are explored by DL-RSIM, and it can be incorporated with any deep learning neural networks implemented by TensorFlow. Using several representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and energy-efficient deployment strategies and develop optimization techniques accordingly.
Hsiang-Yun Cheng, Chia-Lin Yang, Meng-Yao Lin, Kai Lien, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang, Yen-Ting Tsou, Chin-Fu Nien
ACM Trans. Embed. Comput. Syst.2
2021 RePIM: Joint Exploitation of Activation and Weight Repetitions for In-ReRAM DNN Acceleration
abstract
Eliminating redundant computations is a common approach to improve the performance of ReRAM-based DNN accelerators. While existing practical ReRAM-based accelerators eliminate part of the redundant computations by exploiting sparsity in inputs and weights or utilizing weight patterns of DNN models, they fail to identify all the redundancy, resulting in many unnecessary computations. Thus, we propose a practical design, RePIM, that is the first to jointly exploit the repetition of both inputs and weights. Our evaluation shows that RePIM is effective in eliminating unnecessary computations, achieving an average of $ 15.24\times$ speedup and 96.07% energy savings over the state-of-the-art practical ReRAM-based accelerator.
Chen-Yang Tsai, Chin-Fu Nien, Tz-Ching Yu, Hung-Yu Yeh, Hsiang-Yun Cheng
DAC5
2021 Future Computing Platform Design: A Cross-Layer Design Approach
abstract
Future computing platforms are facing a paradigm shift with the emerging resistive memory technologies. First, they offer fast memory accesses and data persistence in a single large-capacity device deployed on the memory bus, blurring the boundary between memory and storage. Second, they enable computing-in-memory for neuromorphic computing to mitigate costly data movements. Due to the non-ideality of these resistive memory devices at the moment, we envision that cross-layer design is essential to bring such a system into practice. In this paper, we showcase a few examples to demonstrate how cross-layer design can be developed to fully exploit the potential of resistive memories and accelerate its adoption for future computing platforms.
Hsiang-Yun Cheng, Chun-Feng Wu, Christian Hakert, Kuan-Hsun Chen, Yuan-Hao Chang 0001, Jian-Jia Chen, Chia-Lin Yang, Tei-Wei Kuo
DATE1
2021 ReSpar: Reordering Algorithm for ReRAM-based Sparse Matrix-Vector Multiplication Accelerator
abstract
Sparse matrix-vector multiplication (SpMV) serves as a crucial operation for several key application domains, such as graph analytics and scientific computing, in the era of big data. The performance of SpMV is bounded by the data transmissions across memory channels in conventional von Neumann systems. Emerging metal-oxide resistive random access memory (ReRAM) has shown its potential to address this memory wall challenge through performing SpMV directly within its crossbar arrays. However, due to the tightly coupled crossbar structure, it is unlikely to skip all redundant data loading and computations with zero-valued entries of the sparse matrix in such ReRAM-based processing-in-memory architecture. These unnecessary ReRAM writes and computations hurt the energy efficiency. As only the crossbar-sized sub-matrices with full-zero entries can be skipped, prior studies have proposed some matrix reordering methods to aggregate non-zero entries to few crossbar arrays, such that more full-zero crossbar arrays can be skipped. Nevertheless, the effectiveness of prior reordering methods is constrained by the original ordering of matrix rows. In this paper, we show that the amount of full-zero sub-matrices derived by these prior studies are less than a theoretical lower bound in some cases, indicating that there are still rooms for improvement. Hence, we propose a novel reordering algorithm, ReSpar, that aims to aggregate matrix rows with similar non-zero column entries together and concentrates the non-zeros columns to increase the zero-skipping opportunities. Results show that ReSpar achieves 1.68× and 1.37× more energy savings, while reducing the required number of crossbar loads by 40.4% and 27.2% on average.
Yi-Jou Hsiao, Chin-Fu Nien, Hsiang-Yun Cheng
ICCD3
2021 Analyzing the Interplay Between Random Shuffling and Storage Devices for Efficient Machine Learning
abstract
Machine learning algorithms, such as Support Vector Machine (SVM) and Deep Neural Network (DNN), have gained a lot of interest recently. When training a machine learning algorithm, randomly shuffling all the training data can improve the testing accuracy and boost the convergence rate. Nevertheless, realizing training data random shuffling in a real system is not straightforward due to the slow random accesses in hard disk drives (HDDs). Common random shuffling implementations assume that HDD is used as storage, so they sacrifice the random degree of shuffling to reduce random storage accesses. Different from conventional HDD, emerging solid-state drive (SSD) based storage devices, such as Intel Optane SSD, offer fast random accesses. In this paper, we explore the opportunities to take advantage of the fast random access property in SSD to perform full-range random shuffling without taking up precious CPU memory and study the interplay between different shuffling methods and various types of storage devices. We use a lightweight implementation of random shuffling (LIRS) as an example of the SSD-aware shuffling method to conduct performance analysis. Evaluations show that, compared to conventional shuffling methods, LIRS can improve convergence rate and reduce the total training time of SVM and DNN by 67.1% and 33.9% on average.
Zhi-Lin Ke, Hsiang-Yun Cheng, Chia-Lin Yang, Han-Wei Huang
ISPASS2
2020 GraphRSim: A Joint Device-Algorithm Reliability Analysis for ReRAM-based Graph Processing
abstract
Graph processing has attracted a lot of interests in recent years as it plays a key role to analyze huge datasets. ReRAM-based accelerators provide a promising solution to accelerate graph processing. However, the intrinsic stochastic behavior of ReRAM devices makes its computation results unreliable. In this paper, we build a simulation platform to analyze the impact of non-ideal ReRAM devices on the error rates of various graph algorithms. We show that the characteristic of the targeted graph algorithm and the type of ReRAM computations employed greatly affect the error rates. Using representative graph algorithms as case studies, we demonstrate that our simulation platform can guide chip designers to select better design options and develop new techniques to improve reliability.
Chin-Fu Nien, Yi-Jou Hsiao, Hsiang-Yun Cheng, Cheng-Yu Wen, Ya-Cheng Ko, Che-Ching Lin
DATE3
2019 The Impact of Emerging Technologies on Architectures and System-level Management: Invited Paper
abstract
The goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management.
Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang
ICCAD11
2019 Sparse ReRAM engine: joint exploration of activation and weight sparsity in compressed neural networks
abstract
Exploiting model sparsity to reduce ineffectual computation is a commonly used approach to achieve energy efficiency for DNN inference accelerators. However, due to the tightly coupled crossbar structure, exploiting sparsity for ReRAM-based NN accelerator is a less explored area. Existing architectural studies on ReRAM-based NN accelerators assume that an entire crossbar array can be activated in a single cycle. However, due to inference accuracy considerations, matrix-vector computation must be conducted in a smaller granularity in practice, called Operation Unit (OU). An OU-based architecture creates a new opportunity to exploit DNN sparsity. In this paper, we propose the first practical Sparse ReRAM Engine that exploits both weight and activation sparsity. Our evaluation shows that the proposed method is effective in eliminating ineffectual computation, and delivers significant performance improvement and energy savings.
Tzu-Hsien Yang, Hsiang-Yun Cheng, Chia-Lin Yang, I-Ching Tseng, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li
ISCA2
2019 TAP: Reducing the Energy of Asymmetric Hybrid Last-Level Cache via Thrashing Aware Placement and Migration
abstract
Emerging non-volatile memories (NVMs) have favorable properties, such as low leakage and high density, and have attracted a lot of attention in recent years. Among them, spin-transfer torque magnetoresistive random access memory (STT-MRAM) with SRAM-comparable read speed is a good candidate to build large last-level caches (LLCs). However, STT-MRAM suffers from long write latency and high write energy. To mitigate the impact of asymmetric read/write energy and latency, hybrid cache designs have been proposed to combine the merits of STT-MRAM and SRAM. In such a hybrid SRAM/STT-MRAM LLC, intelligent block placement and migration policies are needed to improve the energy efficiency. Prior studies map write-intensive blocks to SRAM and keep read-intensive blocks in STT-MRAM for reducing the energy consumption of hybrid LLCs. The write-intensive/read-intensive blocks are usually captured by sampling the address (PC) of memory access instructions or adding simple access counters in each cache line. Nevertheless, these prior approaches cannot fully capture the energy-harmful access behavior in STT-MRAM, especially the writes caused by repetitive data transfer between the LLC and upper-level caches. In this paper, we find that conflict misses in L2 often generate thrashing blocks which move back and forth between L2 and LLC. If dirty thrashing blocks that incur extensive writes are placed in STT-MRAM, energy consumption would excessively increase, especially when running memory-bound workloads. Thus, we propose a thrashing aware placement and migration policy (TAP) to tackle the challenge. TAP places dirty thrashing blocks into SRAM and migrates clean thrashing blocks from SRAM to STT-MRAM. Evaluation results show that TAP can provide significant energy savings with minimal performance loss.
Jing-Yuan Luo, Hsiang-Yun Cheng, Ing-Chao Lin, Da-Wei Chang
IEEE Trans. Computers2
2018 DL-RSIM: a simulation framework to enable reliable ReRAM-based accelerators for deep learning
abstract
Memristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. To enable reliable memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit and device properties on the inference accuracy. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. DL-RSIM simulates the error rates of every sum-of-products computation in the memristor-based accelerator and injects the errors in the targeted TensorFlow-based neural network model. A rich set of reliability impact factors are explored by DL-RSIM, and it can be incorporated with any deep learning neural network implemented by TensorFlow. Using three representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and develop reliability optimization techniques.
Meng-Yao Lin, Hsiang-Yun Cheng, Tzu-Hsien Yang, I-Ching Tseng, Chia-Lin Yang, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang
ICCAD2
2017 Analyzing OpenCL 2.0 workloads using a heterogeneous CPU-GPU simulator
abstract
Heterogeneous CPU-GPU systems have recently emerged as an energy-efficient computing platform. A robust integrated CPU-GPU simulator is essential to facilitate researches in this direction. While few integrated CPU-GPU simulators are available, similar tools that support OpenCL 2.0, a widely used new standard with promising heterogeneous computing features, are currently missing. In this paper, we extend the existing integrated CPU-GPU simulator, gem5-gpu, to support OpenCL 2.0. In addition, we conduct experiments on the extended simulator to see the impact of new features introduced by OpenCL 2.0. Our OpenCL 2.0 compatible simulator is successfully validated against a state-of-the-art commercial product, and is expected to help boost future studies in heterogeneous CPU-GPU systems.
Ren-Wei Tsai, Shao-Chung Wang, Kun-Chih Chen, Po-Han Wang 0001, Hsiang-Yun Cheng, Yi-Chung Lee, Sheng-Jie Shu, Chun-Chieh Yang, Min-Yih Hsu, Li-Chen Kan, Chao-Lin Lee, Tzu-Chieh Yu, Rih-Ding Peng, Chia-Lin Yang, Yuan-Shin Hwang, Jenq Kuen Lee, Shiao-Li Tsao, Ouhyoung Ming
ISPASS6
2016 LAP: Loop-Block Aware Inclusion Properties for Energy-Efficient Asymmetric Last Level Caches
abstract
Emerging non-volatile memory (NVM) technologies, such as spin-transfer torque RAM (STT-RAM), are attractive options for replacing or augmenting SRAM in implementing last-level caches (LLCs). However, the asymmetric read/write energy and latency associated with NVM introduces new challenges in designing caches where, in contrast to SRAM, dynamic energy from write operations can be responsible for a larger fraction of total cache energy than leakage. These properties lead to the fact that no single traditional inclusion policy being dominant in terms of LLC energy consumption for asymmetric LLCs. We propose a novel selective inclusion policy, Loop-block-Aware Policy (LAP), to reduce energy consumption in LLCs with asymmetric read/write properties. In order to eliminate redundant writes to the LLC, LAP incorporates advantages from both non-inclusive and exclusive designs to selectively cache only part of upper-level data in the LLC. Results show that LAP outperforms other variants of selective inclusion policies and consumes 20% and 12% less energy than non-inclusive and exclusive STT-RAM-based LLCs, respectively. We extend LAP to a system with SRAM/STT-RAM hybrid LLCs to achieve energy-efficient data placement, reducing the energy consumption by 22% and 15% over non-inclusion and exclusion on average, with average-case performance improvements, small worst-case performance loss, and minimal hardware overheads.
Hsiang-Yun Cheng, Jishen Zhao, Jack Sampson, Mary Jane Irwin, Aamer Jaleel, Yuan Xie 0001
ISCA1
2016 Designs of emerging memory based non-volatile TCAM for Internet-of-Things (IoT) and big-data processing: A 5T2R universal cell
abstract
Many search engines or filters for the internet-of-things and big-data employ ternary content-addressable-memory (TCAM) to suppress power consumption in the transmission of data between end-devices and servers. Nonvolatile TCAMs (nvTCAM) are designed to achieve zero standby power with smaller area overhead and faster power off/on operations than those found in conventional TCAM+NVM 2-macro schemes. In this paper, we discuss the challenges involved in the design of nvTCAMs and propose a universal 5T2R nvTCAM cell with tolerance for the various R-ratios and write parameters associated with emerging memory devices. A 128×64b-nvTCAM macro was fabricated using HfO ReRAM and a 90nm-CMOS process for concept verification.
Meng-Fan Chang, Ching-Hao Chuang, Yen-Ning Chiang, Shyh-Shyuan Sheu, Chia-Chen Kuo, Hsiang-Yun Cheng, Jack Sampson, Mary Jane Irwin
ISCAS6
2015 Core vs. uncore: the heart of darkness
abstract
Even though Moore's Law continues to provide increasing transistor counts, the rise of the utilization wall limits the number of transistors that can be powered on and results in a large region of dark silicon. Prior studies have proposed energy-efficient core designs to address the "dark silico" problem. Nevertheless, the research for addressing dark silicon challenges in uncore components, such as shared cache, on-chip interconnect, etc, that contribute significant on-chip power consumption is largely unexplored. In this paper, we first illustrate that the power consumption of uncore components cannot be ignored to meet the chip's power constraint. We then introduce techniques to design energy-efficient uncore components, including shared cache and on-chip interconnect. The design challenges and opportunities to exploit 3D techniques and non-volatile memory (NVM) in dark-silicon-aware architecture are also discussed.
Hsiang-Yun Cheng, Jia Zhan, Jishen Zhao, Yuan Xie 0001, Jack Sampson, Mary Jane Irwin
DAC1
2015 EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors
abstract
Power management for large last-level caches (LLCs) is important in chip multiprocessors (CMPs), as the leakage power of LLCs accounts for a significant fraction of the limited on-chip power budget. Since not all workloads running on CMPs need the entire cache, portions of a large, shared LLC can be disabled to save energy. In this article, we explore different design choices, from circuit-level cache organization to microarchitectural management policies, to propose a low-overhead runtime mechanism for energy reduction in the large, shared LLC. We first introduce a slice-based cache organization that can shut down parts of the shared LLC with minimal circuit overhead. Based on this slice-based organization, part of the shared LLC can be turned off according to the spatial and temporal cache access behavior captured by low-overhead sampling-based hardware. In order to eliminate the performance penalties caused by flushing data before powering off a cache slice, we propose data migration policies to prevent the loss of useful data in the LLC. Results show that our energy-efficient cache design (EECache) provides 14.1% energy savings at only 1.2% performance degradation and consumes negligible hardware overhead compared to prior work.
Hsiang-Yun Cheng, Matthew Poremba, Narges Shahidi, Ivan Stalev, Mary Jane Irwin, Mahmut T. Kandemir, Jack Sampson, Yuan Xie 0001
ACM Trans. Archit. Code Optim.1
2015 Adaptive Burst-Writes (ABW): Memory Requests Scheduling to Reduce Write-Induced Interference
abstract
Main memory latencies have become a major performance bottleneck for chip-multiprocessors (CMPs). Since reads are on the critical path, existing memory controllers prioritize reads over writes. However, writes must be eventually processed when the write queue is full. These writes are serviced in a burst to reduce the bus turnaround delay and increase the row-buffer locality. Unfortunately, a large number of reads may suffer long queuing delay when the burst-writes are serviced. The long write latency of future nonvolatile memory will further exacerbate the long queuing delay of reads during burst-writes. In this article, we propose a run-time mechanism, Adaptive Burst-Writes (ABW), to reduce the queuing delay of reads. Based on the row-buffer hit rate of writes and the arrival rate of reads, we dynamically control the number of writes serviced in a burst to trade off the write service time and the queuing latency of reads. For prompt adjustment, our history-based mechanism further terminates the burst-writes earlier when the row-buffer hit rate of writes in the previous burst-writes is low. As a result, our policy improves system throughput by up to 28% (average 10%) and 43% (average 14%) in CMPs with DRAM-based and PCM-based main memory.
Hsiang-Yun Cheng, Mary Jane Irwin, Yuan Xie 0001
ACM Trans. Design Autom. Electr. Syst.1
2014 EECache: exploiting design choices in energy-efficient last-level caches for chip multiprocessors
abstract
Power management for large last-level caches (LLCs) is important in chip-multiprocessors (CMPs), as the leakage power of LLCs accounts for a significant fraction of the limited on-chip power budget. Since not all workloads need the entire cache, portions of a shared LLC can be disabled to save energy. In this paper, we explore different design choices, from circuit-level cache organization to micro-architectural management policies, to propose a low-overhead run-time mechanism for energy reduction in the shared LLC. Results show that our design (EECache) provides 14.1% energy saving at only 1.2% performance degradation on average, with negligible hardware overhead.
Hsiang-Yun Cheng, Matthew Poremba, Narges Shahidi, Ivan Stalev, Mary Jane Irwin, Mahmut T. Kandemir, Jack Sampson, Yuan Xie 0001
ISLPED1
2010 Memory Latency Reduction via Thread Throttling
abstract
Memory Wall is a well-known obstacle to processor performance improvement. The popularity of multi-core architecture will further exaggerate the problem since the memory resource is shared by all cores. Interferences among requests from different cores may prolong the latency of memory accesses thereby degrading the system performance. To tackle the problem, this paper proposes to decouple application threads into compute and memory tasks, and restrict the number of concurrent memory tasks to avoid the interference among memory requests. Yet with this scheduling restriction, a CPU core may unnecessarily stay idle, which incurs adverse impact on the overall performance. Therefore, we develop a memory thread throttling mechanism that tunes the allowable memory threads dynamically under workload variation to improve system performance. The proposed run-time mechanism monitors memory and computation ratios of a program for phase detection. It then decides the memory thread constraint for the next program phase based on an analytical model that can estimate system performance under different constraint values. To prove the concept, we prototype the mechanism in some real-world applications as well as synthetic workloads. We evaluate their performance on real machines. The experimental results demonstrate up to 20% speedup with a pool of synthetic workloads on an Intel i7 (Nehalem) machine and match with the speedup estimated by the proposed analytical model. Furthermore, the intelligent run-time scheduling leads to a geometric mean of 12% performance improvement for real-world applications on the same hardware.
Hsiang-Yun Cheng, Chung-Hsiang Lin, Chia-Lin Yang
MICRO1