Jaewan Choi

dblp:38/9893 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2024 AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model Inference
abstract
The Transformer-based generative model (TbGM), comprising summarization (Sum) and generation (Gen) stages, has demonstrated unprecedented generative performance across a wide range of applications. However, it also demands immense amounts of compute and memory resources. Especially, the Gen stages, consisting of the attention and fully-connected (FC) layers, dominate the overall execution time. Meanwhile, we reveal that the conventional system with GPUs used for TbGM inference cannot efficiently execute the attention layer, even with batching, due to various constraints. To address this inefficiency, we first propose AttAcc, a processing-in-memory (PIM) architecture for efficient execution of the attention layer. Subsequently, for the end-to-end acceleration of TbGM inference, we propose a novel heterogeneous system architecture and optimizations that strategically use xPU and PIM together. It leverages the high memory bandwidth of AttAcc for the attention layer and the powerful compute capability of the conventional system for the FC layer. Lastly, we demonstrate that our GPU-PIM system outperforms the conventional system with the same memory capacity, improving performance and energy efficiency of running a 175B TbGM by up to 2.81× and 2.67×, respectively.
Jaehyun Park 0006, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, Jung Ho Ahn
ASPLOS (2)2
2024 Unleashing the Potential of PIM: Accelerating Large Batched Inference of Transformer-Based Generative Models
abstract
Transformer-based generative models, such as GPT, summarize an input sequence by generating key/value (KV) matrices through attention and generate the corresponding output sequence by utilizing these matrices once per token of the sequence. Both input and output sequences tend to get longer, which improves the understanding of contexts and conversation quality. These models are also typically batched for inference to improve the serving throughput. All these trends enable the models' weights to be reused effectively, increasing the relative importance of sequence generation, especially in processing KV matrices through attention. We identify that the conventional computing platforms (e.g., GPUs) are not efficient at handling this attention part for inference because each request generates different KV matrices, it has a low operation per byte ratio regardless of the batch size, and the aggregate size of the KV matrices can even surpass that of the entire model weights. This motivates us to propose AttAcc, which exploits the fact that the KV matrices are written once during summarization but used many times (proportional to the output sequence length), each multiplied by the embedding vector corresponding to an output token. The volume of data entering/leaving AttAcc could be more than orders of magnitude smaller than what should be read internally for attention. We design AttAcc with multiple processing-in-memory devices, each multiplying the embedding vector with the portion of the KV matrices within the devices, saving external (inter-device) bandwidth and energy consumption.
Jaewan Choi, Jaehyun Park 0006, Kwanhee Kyung, Nam Sung Kim, Jung Ho Ahn
HPCA1
2024 Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
abstract
Large language models (LLMs) have emerged due to their capability to generate high-quality content across diverse contexts. To reduce their explosively increasing demands for computing resources, a mixture of experts (MoE) has emerged. The MoE layer enables exploiting a huge number of parameters with less computation. Applying state-of-the-art continuous batching increases throughput; however, it leads to frequent DRAM access in the MoE and attention layers. We observe that conventional computing devices have limitations when processing the MoE and attention layers, which dominate the total execution time and exhibit low arithmetic intensity (Op/B). Processing MoE layers only with devices targeting low-Op/B such as processing-in-memory (PIM) architectures is challenging due to the fluctuating Op/B in the MoE layer caused by continuous batching, To address these challenges, we propose Duplex, which comprises xPU tailored for high-Op/B and Logic-PIM to effectively perform low-Op/B operation within a single device. Duplex selects the most suitable processor based on the Op/B of each layer within LLMs. As the Op/B of the MoE layer is at least 1 and that of the attention layer has a value of 4–8 for grouped query attention, prior PIM architectures are not efficient, which place processing units inside DRAM dies and only target extremely low-Op/B (under one) operations. Based on recent trends, Logic-Pimadds more through-silicon vias (TSVs) to enable high-bandwidth communication between the DRAM die and the logic die and place powerful processing units on the logic die, which is best suited for handling low-Op/B operations ranging from few to a few dozens. To maximally utilize the xPU and Logic-Pim,we propose expert and attention co-processing. By exploiting proper processing units for MoE and attention layers, Duplex shows up to 2.67 × higher throughput and consumes 42.0% less energy compared to GPU systems for LLM inference.
Sungmin Yun 0001, Kwanhee Kyung, Juhwan Cho, Jaewan Choi, Jongmin Kim 0007, Byeongho Kim, Sukhan Lee 0002, Kyomin Sohn, Jung Ho Ahn
MICRO4
2023 SHARP: A Short-Word Hierarchical Accelerator for Robust and Practical Fully Homomorphic Encryption
abstract
Fully homomorphic encryption (FHE) is an emerging cryptographic technology that guarantees the privacy of sensitive user data by enabling direct computations on encrypted data. Despite the security benefits of this approach, FHE is associated with prohibitively high levels of computational and memory overhead, preventing its widespread use in real-world services. Numerous domain-specific hardware designs have been proposed to address this issue, but most of them use excessive amounts of chip area and power, leaving room for further improvements in terms of practicality.
Jongmin Kim 0007, Sangpyo Kim, Jaewan Choi, Jaiyoung Park, Jung Ho Ahn
ISCA3
2022 Future Scaling of Memory Hierarchy for Tensor Cores and Eliminating Redundant Shared Memory Traffic Using Inter-Warp Multicasting
abstract
The CUDA core of NVIDIA GPUs had been one of the most efficient computation units for parallel computing. However, recent rapid developments in deep neural networks demand an even higher level of computational performance. To meet this requirement, NVIDIA has introduced the Tensor core in recent generations. However, their impressive enhancements in computational performance have newly brought high pressure on the memory hierarchy. In this paper, first we identify the required memory bandwidth in the memory hierarchy as the computational performance increases in actual GPU hardware. Through a comparison of the CUDA core and the Tensor core in V100, we find that the tremendous performance increase of the Tensor core requires much higher memory bandwidth than that in the CUDA core. Moreover, we thoroughly investigate memory bandwidth requirement over Tensor core generations of V100, RTX TITAN, and A100. Lastly, we analyze a hypothetical next-generation Tensor core introduced by NVIDIA through a GPU simulation, through which we propose an inter-warp multicasting microarchitecture that reduces redundant shared memory (SMEM) traffic during the GEMM process. Our evaluation shows that inter-warp multicasting reduces the SMEM bandwidth pressure by 33% and improves the performance by 19% on average in all layers of ResNet-152 and BERT-Large.
Sunjung Lee, Seunghwan Hwang, Michael Jaemin Kim, Jaewan Choi, Jung Ho Ahn
IEEE Trans. Computers4
2022 MVP: An Efficient CNN Accelerator with Matrix, Vector, and Processing-Near-Memory Units
abstract
Mobile and edge devices become common platforms for inferring convolutional neural networks (CNNs) due to superior privacy and service quality. To reduce the computational costs of convolution (CONV) , recent CNN models adopt depth-wise CONV (DW-CONV) and Squeeze-and-Excitation (SE) . However, existing area-efficient CNN accelerators are sub-optimal for these latest CNN models because they were mainly optimized for compute-intensive standard CONV layers with abundant data reuse that can be pipelined with activation and normalization operations. In contrast, DW-CONV and SE are memory-intensive with limited data reuse. The latter also strongly depends on the nearby CONV layers, making an effective pipelining a daunting task. Therefore, DW-CONV and SE only occupy 10% of entire operations but become memory bandwidth bound, spending more than 60% of the processing time in systolic-array-based accelerators. We propose a CNN acceleration architecture called MVP, which efficiently processes both compute- and memory-intensive operations with a small area overhead on top of the baseline systolic-array-based architecture. We suggest a specialized vector unit tailored for processing DW-CONV, including multipliers, adder trees, and multi-banked buffers to meet the high memory bandwidth requirement. We augment the unified buffer with tiny processing elements to smoothly pipeline SE with the subsequent CONV, enabling concurrent processing of DW-CONV with standard CONV, thereby achieving the maximum utilization of arithmetic units. Our evaluation shows that MVP improves performance by 2.6 \( \times \) and reduces energy by 47% on average for EfficientNet-B0/B4/B7, MnasNet, and MobileNet-V1/V2 with only a 9% area overhead compared to the baseline.
Sunjung Lee, Jaewan Choi, Wonkyung Jung, Byeongho Kim, Jaehyun Park 0006, Hweesoo Kim, Jung Ho Ahn
ACM Trans. Design Autom. Electr. Syst.2
2020 MViD: Sparse Matrix-Vector Multiplication in Mobile DRAM for Accelerating Recurrent Neural Networks
abstract
Recurrent Neural Networks (RNNs) spend most of their execution time performing matrix-vector multiplication (MV-mul). Because the matrices in RNNs have poor reusability and the ever-increasing size of the matrices becomes too large to fit in the on-chip storage of mobile/IoT devices, the performance and energy efficiency of MV-mul is determined by those of main-memory DRAM. Therefore, computing MV-mul within DRAM draws much attention. However, previous studies lacked consideration for the matrix sparsity, the power constraints of DRAM devices, and concurrency in accessing DRAM from processors while performing MV-mul. We propose a main-memory architecture called MViD, which performs MV-mul by placing MAC units inside DRAM banks. For higher computational efficiency, we use a sparse matrix format and exploit quantization. Because of the limited power budget for DRAM devices, we implement the MAC units only on a portion of the DRAM banks. We architect MViD to slow down or pause MV-mul for concurrently processing memory requests from processors while satisfying the limited power budget. Our results show that MViD provides 7.2× higher throughput compared to the baseline system with four DRAM ranks (performing MV-mul in a chip-multiprocessor) while running inference of Deep Speech 2 with a memory-intensive workload.
Byeongho Kim, Jongwook Chung, Eojin Lee, Wonkyung Jung, Sunjung Lee, Jaewan Choi, Jaehyun Park 0006, Minbok Wi, Sukhan Lee 0002, Jung Ho Ahn
IEEE Trans. Computers6
2018 A Comparison of Hyper-Sharpening Algorithms for Fusing VNIR and SWIR Bands of WorldView-3 Satellite Imagery
abstract
In this study, visible and near-infrared (VNIR) and shortwave infrared (SWIR) bands of WorldView-3 imagery were sharpened to the spatial resolution of a panchromatic image. We performed experiments on three band schemes according to a method for generating an optimal panchromatic image. Various pan-sharpening algorithms based on component-substitution (CS) and multi-resolution analysis (MRA) were applied to each band scheme. Quantitative and qualitative assessments were performed using the quality indices and visual inspection.
Honglyun Park, Jaewan Choi
IGARSS2
2017 Image Fusion of Spectrally Nonoverlapping Imagery Using SPCA and MTF-Based Filters
abstract
Most spaceborne sensors have an inevitable tradeoff between spatial and spectral resolutions. This is a typical ill-posed inverse problem in the field of image fusion. To solve this problem, this letter proposes an image fusion method using spatial principal component analysis and modulation transfer function-based filters. The key behind the proposed fusion method is to efficiently estimate the missing spatial details by considering the spatial structures of the low-resolution multispectral (MS) imagery. Also, this letter proposes a newly developed injection gain model to resolve the local and global dissimilarity between panchromatic and MS imageries, which could prevent over- and under-injections. Finally, spatial details, optimized to be injected into the MS images, were constructed and paired with the developed injection gain model to produce high-resolution MS images. Two data sets acquired by WorldView-2 are employed for validation. The experimental results demonstrate that the proposed fusion method generates high-quality imagery in terms of both qualitative and quantitative standards.
Jaewan Choi, Yongil Kim
IEEE Geosci. Remote. Sens. Lett.3
2015 Hyperspectral change detection by using IR-MAD and synthetic image fusion
abstract
We propose a modified IR-MAD based on the generation of synthetically fused images in order to minimize the effect of change detection results corresponding to noise and feature reduction. Synthetically fused hyperspectral images were first generated using a cross-sharpening algorithm. MAD variates according to each pair of synthetically fused images were then calculated to reduce the influence of data noise in the hyperspectral image. In particular, we applied the integration of MAD variates in this study. To evaluate the performance of our algorithm, we constructed a hyperspectral dataset using the Hyperion sensor and analyzed the data noise and bands of principal components.
Jaewan Choi, Guhyeok Kim, Youkyung Han
IGARSS1
2015 Object-Based Change Detection of Very High Resolution Satellite Imagery Using the Cross-Sharpening of Multitemporal Data
abstract
In this letter, we present a method for unsupervised change detection based on the cross-sharpening of multitemporal images and image segmentation. Our method effectively reduces the change detection errors caused by relief or spatial displacement between multitemporal images with different acquisition angles. A total of four cross-sharpened images, including two general pansharpened images, were generated. Then, two pairs of cross-sharpened images were analyzed using change detection indexes. The effectiveness of the proposed method compared with other unsupervised change detection methods is demonstrated through experimentation.
Seokkeun Choi, Younggi Byun, Soungki Lee, Jaewan Choi
IEEE Geosci. Remote. Sens. Lett.5
2014 Parameter Optimization for the Extraction of Matching Points Between High-Resolution Multisensor Images in Urban Areas
abstract
The objective of this paper is to extract a suitable number of evenly distributed matched points, given the characteristics of the site and the sensors involved. The intent is to increase the accuracy of automatic image-to-image registration for high-resolution multisensor data. The initial set of matching points is extracted using a scale-invariant feature transform (SIFT)-based method, which is further used to evaluate the initial geometric relationship between the features of the reference and sensed images. The precise matching points are extracted considering location differences and local properties of features. The values of the parameters used in the precise matching are optimized using an objective function that considers both the distribution of the matching points and the reliability of the transformation model. In case studies, the proposed algorithm extracts an appropriate number of well-distributed matching points and achieves a higher correct-match rate than the SIFT method. The registration results for all sensors are acceptably accurate, with a root-mean-square error of less than 1.5 m.
Youkyung Han, Jaewan Choi, Younggi Byun, Yongil Kim
IEEE Trans. Geosci. Remote. Sens.2
2013 Hybrid Pansharpening Algorithm for High Spatial Resolution Satellite Imagery to Improve Spatial Quality
abstract
Most pansharpened images from existing algorithms are apt to present a tradeoff relationship between the spectral preservation and the spatial enhancement. In this letter, we developed a hybrid pansharpening algorithm based on primary and secondary high-frequency information injection to efficiently improve the spatial quality of the pansharpened image. The injected high-frequency information in our algorithm is composed of two types of data, i.e., the difference between panchromatic and intensity images, and the Laplacian filtered image of high-frequency information. The extracted high frequencies are injected by the multispectral image using the local adaptive fusion parameter and postprocessing of the fusion parameter. In the experiments using various satellite images, our results show better spatial quality than those of other fusion algorithms while maintaining as much spectral information as possible.
Jaewan Choi, Junho Yeom, Anjin Chang, Younggi Byun, Yongil Kim
IEEE Geosci. Remote. Sens. Lett.1
2011 A New Adaptive Component-Substitution-Based Satellite Image Fusion by Using Partial Replacement
abstract
Preservation of spectral information and enhancement of spatial resolution are regarded as important issues in remote sensing satellite image fusion. In previous research, various algorithms have been proposed. Although they have been successful, there are still some margins of spatial and spectral quality that can be improved. In addition, a new method that can be used for various types of sensors is required. In this paper, a new adaptive fusion method based on component substitution is proposed to merge a high-spatial-resolution panchromatic (PAN) image with a multispectral image. This method generates high-/low-resolution synthetic component images by partial replacement and uses statistical ratio-based high-frequency injection. Various remote sensing satellite images, such as IKONOS-2, QuickBird, LANDSAT ETM+, and SPOT-5, were employed in the evaluation. Experiments showed that this approach can resolve spectral distortion problems and successfully conserve the spatial information of a PAN image. Thus, the fused image obtained from the proposed method gave higher fusion quality than the images from some other methods. In addition, the proposed method worked efficiently with the different sensors considered in the evaluation.
Jaewan Choi, Kiyun Yu, Yongil Kim
IEEE Trans. Geosci. Remote. Sens.1