Changjae Yi

dblp:303/8246 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-9752-303XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enabling Decoder-only Language Model Inference on a CNN Accelerator
abstract
The remarkable success of Transformer architectures in Natural Language Processing (NLP) has led to increased demand for embedded systems capable of efficiently handling NLP tasks along with traditional vision tasks based on Convolutional Neural Networks (CNNs) or Vision Transformers. Given that CNN accelerators are already widely adopted commercially, this paper investigates the feasibility of leveraging existing CNN accelerators to handle NLP workloads—particularly decoder-only language models—thus avoiding the high cost of developing dedicated NLP hardware. However, direct use of CNN accelerators for language model poses several challenges due to the distinct characteristics of this network. These include non-computational operations in attention layers, such as head merging/splitting and tensor transpositions, as well as the need for floating point precision in nonlinear operations such as Softmax and RMSNorm, which CNN accelerators typically do not support. Furthermore, inference throughput drops significantly during the generation phase due to memory-bound operations, in contrast to the compute-bound nature of convolution. To address these challenges, we propose a set of minimal hardware extensions to existing CNN accelerators to enable efficient decoder-only Transformer inference. We propose a method to execute multi-head attention layers without relying on dedicated reshaping hardware and augment the architecture with a lightweight SIMD coprocessor to handle non-linear operations. Additionally, we present techniques to improve power efficiency during the generation stage. Our experimental results demonstrate that even low-power CNN accelerators can achieve NLP inference throughput comparable to GPUs, opening a promising path toward versatile and cost-effective embedded AI hardware.
Seongwoo Choi, Hyunsu Moh, Changjae Yi, Joon Choi, Soonhoi Ha
ICCAD3
2024 Vision Transformer Inference on a CNN Accelerator
abstract
Following the remarkable performance demonstrated by the Transformer architecture in the field of computer vision as well as natural language processing (NLP), there is a growing demand for embedded systems capable of executing Vision Transformer (ViT) applications as well as Convolutional Neural Network (CNN) applications efficiently. Since CNN accelerators are already widely used commercially, this paper explores the possibility of using existing CNN accelerators to support ViT rather than developing separate accelerators for each. CNN accelerators inherently have some limitations in efficiently handling operations in transformers: matrix multiplication (MM) operations with two non-constant matrices and nonlinear operations. To overcome these limitations, we first propose a novel technique to efficiently handle MM operations without special reshaping hardware in an adder-tree type CNN accelerator. And we propose an optimal scheduling method to minimize the idle time caused by offloading computation of nonlinear operations of the Transformer. Additionally, we investigate the possibility of executing layer normalization and GELU operations on the accelerator with minor extensions. The experimental results validate the effectiveness of the proposed methods.
Changjae Yi, Hyunsu Moh, Soonhoi Ha
ICCD1
2023 Fast and Accurate Virtual Prototyping of an NPU with Analytical Memory Modeling
abstract
As the application area of convolutional neural networks (CNNs) is fast expanding, the demand for a customized hardware accelerator called a neural processing unit (NPU), is increasing to process them efficiently in terms of execution time and energy consumption. In the design of an NPU, building a fast and accurate virtual prototype enables us to develop a compiler concurrently with the hardware and to explore the micro-architectural design space. Since the memory access latency has a great effect on performance, it is necessary to model the memory access overhead accurately in the virtual prototype. In this work, we propose a novel analytical model for memory access latency, improving the performance estimation accuracy significantly compared with the previous state-of-the-art analytical model by considering the effect of memory access patterns of the NPU on the latency. The proposed high-level virtual prototype achieves an estimated execution time gap within a 6.8% difference from the RTL simulation result. To demonstrate the usefulness of a fast and accurate prototype, we propose a compiler optimization technique and a new DMA logic tailored for the NPU for further performance improvement.
Choonghoon Park, Hyunsu Moh, Changjae Yi, Soonhoi Ha
RSP4
2022 Hardware-Software Codesign of a CNN Accelerator
abstract
The explosive growth of deep learning applications based on convolutional neural network (CNN) in embedded sys-tems is spurring the development of a hardware CNN accelerator, called a neural processing unit (NPU). In this work, we present how the hardware-software codesign methodology could be applied to the design of a novel adder-type NPU. After devising a baseline datapath that enables fully-pipelined execution of layers, we define a high-level behavior model based on which a high-level compiler and a virtual prototyping system are built concurrently. Since it is easy to change the microarchitecture of an NPU by modifying the simulation models of the hardware modules, we could explore the design space of NPU microarchitecture easily. In addition, we could evaluate the effect of hardware extensions to support various types of non-convolutional operations that recent CNN models use widely. After the final datapath is determined, we design the control structure and low-level compiler and implement the NPU prototype. Implementation results on an FPGA prototype show the viability of the proposed methodology and its outcome.
Changjae Yi, Soonhoi Ha
DSD1
2021 Fast Simulation of a Many-NPU Network-on-Chip for Microarchitectural Design Space Exploration
abstract
A viable solution to cope with the ever-increasing computation complexity of deep learning applications is to integrate many neural processing units (NPUs) in a chip where a network-on-chip (NoC) is used as the communication fabric. Since the design space of an NoC is huge, the network topology is first selected based on the communication patterns of applications with a high-level performance estimation method. After the network topology is selected, the microarchitectural design space exploration is performed with a cycle-level NoC simulator. However, the existing NoC simulator is so slow that design space exploration of the microarchitecture is usually conducted manually in a narrow space. Since a synthetic trace is used, the simulation accuracy is also limited. To overcome these weak-nesses, we present a simulation technique that is fast and accurate enough for microarchitectural design space of an NoC. In the proposed technique, we use the real communication trace from the many-NPU simulation without NoC consideration. To this end, we define the trace format that defines the interface between a many-NPU simulator and the NoC simulator. To accelerate simulation speed, we propose a parallelization technique at the cluster level in the simulation of the hierarchical NoC. The key technique is to manage the timestamps of events at the cluster boundary to do without time synchronization error. And, we adjust the abstraction level of simulation models to reduce the number of modules in the SystemC NoC simulation. With the proposed technique, we could achieve up to 40 times speed-up for 32 NPU system, compared with the FlexNoC simulator.
Jintaek Kang, Changjae Yi, Keonjoo Lee, Seungwook Lee, Soojung Ryu, Soonhoi Ha
DSD2