Hyein Shin

dblp:258/6650 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2024
0000-0003-0382-4032ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 9 since 2021
YearPublicationVenuePosition
2024 ToEx: Accelerating Generation Stage of Transformer-Based Language Models via Token-Adaptive Early Exit
abstract
Transformer-based language models have recently gained popularity in numerous natural language processing (NLP) applications due to their superior performance compared to traditional algorithms. These models involve two execution stages: summarization and generation. The generation stage accounts for a significant portion of the total execution time due to its auto-regressive property, which necessitates considerable and repetitive off-chip accesses. Consequently, our objective is to minimize off-chip accesses during the generation stage to expedite transformer execution. To achieve the goal, we propose a token-adaptive early exit (ToEx) that generates output tokens using fewer decoders, thereby reducing off-chip accesses for loading weight parameters. Although our approach has the potential to minimize data communication, it brings two challenges: 1) inaccurate self-attention computation, and 2) significant overhead for exit decision. To overcome these challenges, we introduce a methodology that facilitates accurate self-attention by lazily performing computations for previously exited tokens. Moreover, we mitigate the overhead of exit decision by incorporating a lightweight output embedding layer. We also present a hardware design to efficiently support the proposed work. Evaluation results demonstrate that our work can reduce the number of decoders by 2.6× on average. Accordingly, it achieves 3.2× speedup on average compared to transformer execution without our work.
Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim
IEEE Trans. Computers3
2023 MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT
abstract
Recently, multiple transformer models, such as BERT, have been utilized together to support multiple natural language processing (NLP) tasks in a system, also known as multi-task BERT. Multi-task BERT with very high weight parameters increases the area requirement of a processing in resistive memory (ReRAM) architecture, and several works have attempted to address this model size issue. Despite the reduced parameters, the number of multi-task BERT computations remains the same, leading to massive energy consumption in ReRAM-based deep neural network (DNN) accelerators. Therefore, we suggest a framework for better energy efficiency during the ReRAM acceleration of multi-task BERT. First, we analyze the inherent redundancies of multi-task BERT and the computational properties of the ReRAM-based DNN accelerator, after which we propose what is termed the model generator, which produces optimal BERT models supporting multiple tasks. The model generator reduces multi-task BERT computations while maintaining the algorithmic performance. Furthermore, we present task scheduler, which adjusts the execution order of multiple tasks, to run the produced models efficiently. As a result, the proposed framework achieves maximally 4.4× higher energy efficiency over the baseline, and it can also be combined with the previous multi-task BERT works to achieve both a smaller area and higher energy efficiency.
Myeonggu Kang, Hyein Shin, Junkyum Kim, Lee-Sup Kim
IEEE Trans. Computers2
2023 Fault-Free: A Framework for Analysis and Mitigation of Stuck-at-Fault on Realistic ReRAM-Based DNN Accelerators
abstract
Resistive RAM(ReRAM) is gaining attention as a suitable memory platform for accelerating deep neural networks(DNNs) in an energy-efficient way. However, energy-efficient ReRAM-based DNN accelerators suffer from serious Stuck-At-Fault(SAF) issues that significantly degrade the inference accuracy. SAF is a device-level non-ideality, and the problems of SAF worsen in the realistic ReRAM with low cell resolution. To address the problem in the realistic ReRAM, we present a framework for mitigating SAF on ReRAM-based accelerators(Fault-free). We first analyze the impact of SAF on low-resolution cells. Based on the analysis, we present offline compilation, which drastically reduces the impact of SAF on inference accuracy. At the first stage, we extract indices of distorted weights due to SAF. For extracted weights, fault-aware weight decomposition and closest value mapping are applied to minimize the error of weights. In the online phase, the target DNN model is executed on the ReRAM-based accelerator along with lightweight compensation units. The online compensation is selectively performed for a small portion of weights to reduce the hardware overhead. With the proposed framework, the ReRAM-based accelerator successfully ensures the inference accuracy of various DNN models with an average area of 5% and an energy overhead of 0.8% from an ideal ReRAM-based accelerator.
Hyein Shin, Myeonggu Kang, Lee-Sup Kim
IEEE Trans. Computers1
2022 Re2fresh: A Framework for Mitigating Read Disturbance in ReRAM-Based DNN Accelerators
abstract
A severe read disturbance problem degrades the inference accuracy of a resistive RAM (ReRAM) based deep neural network (DNN) accelerator. Refresh, which reprograms the ReRAM cells, is the most obvious solution for the problem, but programming ReRAM consumes huge energy. To address the issue, we first analyze the resistance drift pattern of each conductance state and the actual read stress applied to the ReRAM array by considering the characteristics of ReRAM-based DNN accelerators. Based on the analysis, we cluster ReRAM cells into a few groups for each layer of DNN and generate a proper refresh cycle for each group in the offline phase. The individual refresh cycles reduce energy consumption by reducing the number of unnecessary refresh operations. In the online phase, the refresh controller selectively launches refresh operations according to the generated refresh cycles. ReRAM cells are selectively refreshed by minimally modifying the conventional structure of the ReRAM-based DNN accelerator. The proposed work successfully resolves the read disturbance problem by reducing 97% of the energy consumption for the refresh operation while preserving inference accuracy.
Hyein Shin, Myeonggu Kang, Lee-Sup Kim
ICCAD1
2022 S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level Sparsity
abstract
The processing in-memory (PIM) approach that combines memory and processor appears to solve the memory wall problem. NAND flash memory, which is widely adopted in edge devices, is one of the promising platforms for PIM with its high-density property and the intrinsic ability for analog vector-matrix multiplication. Despite its potential, the domain conversion process, which converts an analog current to a digital value, accounts for most energy consumption on the NAND flash-based accelerator. It restricts the NAND flash memory usage for PIM compared to the other platforms. In this paper, we propose a NAND flash-based DNN accelerator to achieve both large memory density and energy efficiency among various platforms. As the NAND flash memory already shows higher memory density than other memory platforms, we aim to enhance energy efficiency by reducing the domain conversion process burden. Firstly, we optimize the bit width of partial multiplication by considering the analog-to-digital converter (ADC) resource. For further optimization, we propose a methodology to exploit many zero partial multiplication results for enhancing both energy efficiency and throughput. The proposed work successfully exploits the bit-level sparsity of DNN, which results in achieving up to 8.6/8.2 larger energy efficiency/throughput over the provisioned baseline.
Myeonggu Kang, Hyeonuk Kim, Hyein Shin, Jaehyeong Sim, Kyeonghan Kim, Lee-Sup Kim
IEEE Trans. Computers3
2022 A Framework for Accelerating Transformer-Based Language Model on ReRAM-Based Architecture
abstract
Transformer-based language models have become thede-factostandard model for various natural language processing (NLP) applications given the superior algorithmic performances. Processing a transformer-based language model on a conventional accelerator induces the memory wall problem, and the ReRAM-based accelerator is a promising solution to this problem. However, due to the characteristics of the self-attention mechanism and the ReRAM-based accelerator, the pipeline hazard arises when processing the transformer-based language model on the ReRAM-based accelerator. This hazard issue greatly increases the overall execution time. In this article, we propose a framework to resolve the hazard issue. First, we propose the concept of window self-attention to reduce the attention computation scope by analyzing the properties of the self-attention mechanism. After that, we present a window-size search algorithm, which finds an optimal window size set according to the target application/algorithmic performance. We also suggest a hardware design that exploits the advantages of the proposed algorithm optimization on the general ReRAM-based accelerator. The proposed work successfully alleviates the hazard issue while maintaining the algorithmic performance, leading to a$5.8\times $speedup over the provisioned baseline. It also delivers up to$39.2\times /643.2\times $speedup/higher energy efficiency over GPU, respectively.
Myeonggu Kang, Hyein Shin, Lee-Sup Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Fault-free: A Fault-resilient Deep Neural Network Accelerator based on Realistic ReRAM Devices
abstract
Energy-efficient Resistive RAM (ReRAM) based deep neural network (DNN) accelerator suffers from severe Stuck-At-Fault (SAF) problem that drastically degrades the inference accuracy. The SAF problem gets even worse in realistic ReRAM devices with low cell resolution. To address the issue, we propose a fault-resilient DNN accelerator based on realistic ReRAM devices. We first analyze the SAF problem in a realistic ReRAM device and propose a 3-stage offline fault-resilient compilation and lightweight online compensation. The proposed work enables the reliable execution of DNN with only 5% area and 0.8% energy overhead from the ideal ReRAM-based DNN accelerator.
Hyein Shin, Myeonggu Kang, Lee-Sup Kim
DAC1
2021 Optimizing ADC Utilization through Value-Aware Bypass in ReRAM-based DNN Accelerator
abstract
ReRAM-based Processing-In-Memory (PIM) has been widely studied as a promising approach for Deep Neural Networks (DNN) accelerator with its energy-efficient analog operations. However, the domain conversion process for the analog operation requires frequent accesses to power-hungry Analog-to-Digital Converter (ADC), hindering the overall energy efficiency. Although previous research has been suggested to address this problem, the ADC cost has not been sufficiently reduced because of its unsuitable approach for ReRAM. In this paper, we propose mixed-signal-based value-aware bypass techniques to optimize the ADC utilization of the ReRAM-based PIM. By utilizing the property of bit-line (BL) level value distribution, the proposed work bypasses the redundant ADC operations depending on the magnitude of value. Evaluation results show that our techniques successfully reduce ADC access and improve overall energy efficiency by 2.48 × -3.07 × compared to ISAAC.
HanCheon Yun, Hyein Shin, Myeonggu Kang, Lee-Sup Kim
DAC2
2021 A Framework for Area-efficient Multi-task BERT Execution on ReRAM-based Accelerators
abstract
With the superior algorithmic performances, BERT has become the de-facto standard model for various NLP tasks. Accordingly, multiple BERT models have been adopted on a single system, which is also called multi-task BERT. Although the ReRAM-based accelerator shows the sufficient potential to execute a single BERT model by adopting in-memory computation, processing multi-task BERT on the ReRAM-based accelerator extremely increases the overall area due to multiple fine-tuned models. In this paper, we propose a framework for area-efficient multi-task BERT execution on the ReRAM-based accelerator. Firstly, we decompose the fine-tuned model of each task by utilizing the base-model. After that, we propose a two-stage weight compressor, which shrinks the decomposed models by analyzing the properties of the ReRAM-based accelerator. We also present a profiler to generate hyper-parameters for the proposed compressor. By sharing the base-model and compressing the decomposed models, the proposed framework successfully reduces the total area of the ReRAM-based accelerator without an additional training procedure. It achieves a 0.26 x area than baseline while maintaining the algorithmic performances.
Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim
ICCAD2
2020 A Thermal-aware Optimization Framework for ReRAM-based Deep Neural Network Acceleration
abstract
Resistive RAM (ReRAM) is widely regarded as a promising platform for deep neural network (DNN) acceleration. However, the ReRAM device suffers from severe thermal problems that degrade the lifetime and inference accuracy of the ReRAM-based DNN accelerator. To address the issues, we propose a thermal-aware optimization framework for accelerating DNN on ReRAM (TOPAR). TOPAR includes 3-stage offline thermal optimization and online thermal-aware error compensation. Offline thermal optimization consists of thermal-aware weight decomposition, thermal-aware column reordering, and fine-grained weight adjustment to reduce the temperature of the ReRAM-based DNN accelerator. For online thermal-aware error compensation, we compensate conductance change according to the temperature variation. With TOPAR, the endurance degradation due to temperature rise improves up to 2.39×, and inference accuracy is preserved without harming the performance of the ReRAM-based DNN accelerator.
Hyein Shin, Myeonggu Kang, Lee-Sup Kim
ICCAD1
2019 An Energy-efficient Processing-in-memory Architecture for Long Short Term Memory in Spin Orbit Torque MRAM
abstract
Many recent studies have focused on Processing-in-memory (PIM) architectures for neural networks to resolve the memory bottleneck problem. Especially, an increased interest in Spin Orbit Torque (SOT)-MRAMs has emerged due to its low latency, high energy efficiency, and non-volatility. However, the previous work added extra computing circuits to support complicated computations, which results in large energy overheads. In this work, we propose a new PIM architecture with relatively small peripheral circuit, which produces the highest energy efficiency for processing a Long Short Term Memory (LSTM) among the PIM architectures. We improve the efficiency with a new computing method for logical operations, which exploits characteristics of SOT-MRAMs. We reduce the number of word lines (WLs) activated concurrently to one from two in the previous works. As a result, the energy for driving WLs is saved, and the sensing current for computation is reduced. Moreover, we propose efficient methods for additions, multiplications and non-linear activation functions in memory to process an LSTM. Accordingly, we achieve 1.26x energy efficiency with the proposed computing method for logical operations compared to the previous study based on SOT-MRAMs and up to 5.54x energy efficiency over the previous PIM architectures based on other memories.
Kyeonghan Kim, Hyein Shin, Jaehyeong Sim, Myeonggu Kang, Lee-Sup Kim
ICCAD2
2019 A PVT-robust Customized 4T Embedded DRAM Cell Array for Accelerating Binary Neural Networks
abstract
Deep neural networks (DNNs) are widely used for real-world applications. However, large amount of kernel and intermediate data incur a memory wall problem in resource-limited edge devices. The recent advances of a binary deep neural network (BNN) and a computing in-memory (CIM) have effectively alleviated this bottleneck especially when they are combined together. However, previous CIM-based accelerators for BNN are highly vulnerable to process/supply voltage/temperature (PVT) variation, resulting in severe accuracy degradation which makes them impractical to be employed in real-world edge devices. To address this vulnerability, we propose a PVT-robust accelerator architecture for BNN with a computable 4T embedded DRAM (eDRAM) cell array. First, we implement the XNOR operation of BNN in a time-multiplexed manner by utilizing the fundamental read operation of the conventional eDRAM cell. Next, a PVT-robust bit-count based on charge sharing is proposed with a computable 4T eDRAM cell array. In result, the proposed architecture achieves 6.9× less variation in PVT-variant environments which guarantees a stable accuracy and 2.03-49.4× improvement of energy efficiency over previous CIM-based accelerators.
Hyein Shin, Jaehyeong Sim, Daewoong Lee, Lee-Sup Kim
ICCAD1