Jinkwon Kim

dblp:211/1062 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-1744-1393ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Avalanche: Optimizing Cache Utilization via Matrix Reordering for Sparse Matrix Multiplication Accelerator
abstract
Sparse Matrix Multiplication (SpMM) is essential in various scientific and engineering applications but poses significant challenges due to irregular memory access patterns.Many hardware accelerators have been proposed to accelerate SpMM.However, they have yet to focus on on-chip memory utilization.In this paper, we highlight the underutilization of the on-chip memory in the SpMM accelerators.Then we propose Avalanche, a novel hardware accelerator that optimally utilizes the on-chip memory to efficiently cache both matrices 𝐵 and 𝐶.Avalanche incorporates three key techniques: Matrix Reordering (Mat-Reorder), Dead-Product Early Eviction (DP-Evict), and Reuse Distance-Aware Matrix Caching (RM-Caching).Mat-Reorder enhances data locality by reordering the columns of matrix 𝐴, ensuring early completion of computations for matrix 𝐶.DP-Evict optimizes on-chip memory usage by promptly evicting fully computed (dead) products from on-chip memory.RM-Caching maximizes data reuse by caching frequently accessed elements of matrix 𝐵 based on their reuse distance.Experimental results demonstrate that Avalanche achieves an average performance improvement of 1.97× compared to the state-of-theart SpMM accelerator, with a chip area of 6.15 mm 2 CCS Concepts• Computer systems organization → Special purpose systems.
Gwangeun Byeon, Seongwook Kim, Sukhyun Han, Jinkwon Kim, Prashant J. Nair, Seokin Hong
ISCA5
2025 LongSight: Compute-Enabled Memory to Accelerate Large-Context LLMs via Sparse Attention
Derrick Quinn, E. Ezgi Yücel, Jinkwon Kim, José F. Martínez, Mohammad Alian
MICRO3
2024 SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference
abstract
•Recap Accelerator-in-Memory (AiM) & AiMX •System Extensions of AiMX Card for Datacenter •AiM & AiMX for On-device AI •Design Choices for Future AiM/AiMX •Conclusion
Guhyun Kim, Jinkwon Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim 0001, Ilkon Kim, Jaehan Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Yonghoon Park, Yousub Jung, Gi-Ho Park, Eui-Cheol Lim
HCS2
2024 Zero and Narrow-Width Value-Aware Compression for Quantized Convolutional Neural Networks
abstract
Convolutional neural networks are normally used in systems with dedicated neural processing units for CNN-related computations. For high performance and low hardware overheads, CNN datatype quantization is applied. As an additional optimization, to further reduce DRAM accesses, compression algorithms have been used for CNN data. However, conventional zero value-aware compression algorithms suffer from a reduction in compression ratio with the latest quantized CNNs, owing to the small number of zero values. Moreover, the appropriate zero run-length code width can be changed dynamically based on the CNNs, layers, and quantization datatypes. As another compressible data value for increasing the compression ratio, the latest quantized CNNs have many narrow-width values. Because low-precision quantization reduces the data bit width, CNN data are gathered into a few discrete values and incur a biased data distribution. These discrete values become narrow-width values, and constitute a large proportion of the biased distribution. In this article, we propose an efficient compression algorithm for quantized CNNs, ENCORE, which utilizes variable zero run-length encoding and compresses narrow-width values. With the latest quantized CNNs, ENCORE shows higher compression ratios, 93.55% and 50.85% in Mobilenet v1 and Tiny YOLO v3, respectively, than conventional zero value-aware CNN data compression algorithms.
Myeongjae Jang, Jinkwon Kim, Haejin Nam, Soontae Kim
IEEE Trans. Computers2
2023 HARP: Hardware-Based Pseudo-Tiling for Sparse Matrix Multiplication Accelerator
abstract
General sparse matrix-matrix multiplication (SpGEMM) is a memory-bound workload, due to the compression format used. To minimize data movements for input matrices, outer product accelerators have been proposed. Since these accelerators access input matrices only once and then generate numerous partial products, managing the generated partial products is the key optimization factor. To reduce the number of partial products handled, the state-of-the-art accelerator uses software to tile an input matrix. However, the software-based tiling has three limitations. First, a user manually executes the tiling software and manages the tiles. Second, generating a compression format for each tile incurs memory-intensive operations. Third, an accelerator that uses the compression format cannot skip ineffectual accesses for input matrices.
Jinkwon Kim, Myeongjae Jang, Haejin Nam, Soontae Kim
MICRO1
2023 PR-SSD: Maximizing Partial Read Potential by Exploiting Compression and Channel-Level Parallelism
abstract
Recent NAND flash memories provide a partial read operation that can read a page partially and has lower latency than a normal read operation. In order to maximize the benefit of the partial read operation, compression techniques can be applied to improve performance by generating additional partial page requests by compressing pages into smaller ones. Unfortunately, existing compression support SSDs suffer from a huge decompression latency that eventually cancels the benefit of the partial read operation. In this paper, we propose Partial Read-aware SSD (PR-SSD) for fully exploiting partial read operations. In order to mitigate the decompression latency, we propose a new compression algorithm, called Dominant Pattern Compression (DPC), which has extremely low decompression latency. Because uncompressed page requests cannot exploit the partial read operation, we propose split Flash Translation Layer (FTL) that can split the requests into smaller ones and allocate them to different channels for exploiting channel-level parallelism in SSD. Experimental results reveal that PR-SSD can reduce the read response time by 18% on average and also the number of writes and write response time by 29% and 24% on average, respectively
Mincheol Kang, Wonyoung Lee 0001, Jinkwon Kim, Soontae Kim
IEEE Trans. Computers3
2022 ENCORE Compression: Exploiting Narrow-width Values for Quantized Deep Neural Networks
abstract
Deep Neural Networks (DNNs) become a practical machine learning algorithm running on various Neural Processing Units (NPUs). For higher performance and lower hardware overheads, DNN datatype reduction through quantization is proposed. Moreover, to solve the memory bottleneck caused by large data size in DNNs, several zero value-aware compression algorithms are used. However, these compression algorithms do not compress modern quantized DNNs well because of decreased zero values. We find that the latest quantized DNNs have data redundancy due to frequent narrow-width values. Because low-precision quantization reduces DNN datatypes to a simple datatype with less bits, scattered DNN data are gathered to a small number of discrete values and incur a biased data distribution. Narrow-width values occupy a large proportion of the biased distribution. Moreover, an appropriate zero run-length bits can be dynamically changed according to DNN sparsity. Based on this observation, we propose a compression algorithm that exploits narrow-width values and variable zero run-length for quantized DNNs. In experiments with three quantized DNNs, our proposed scheme yields an average compression ratio of 2.99.
Myeongjae Jang, Jinkwon Kim, Jesung Kim, Soontae Kim
DATE2
2022 Exploiting Inter-block Entropy to Enhance the Compressibility of Blocks with Diverse Data
abstract
As higher memory bandwidth is required for data-intensive environments, memory compression can be a simple but effective solution to increase memory bandwidth. However, previous intra-block compression techniques do not provide sufficient bandwidth improvement owing to the incompressibility of blocks with diverse data while previous inter-block compression techniques suffer from huge additional memory access overheads or low compression coverages. To overcome the limitations of the previous intra-and inter-block compression techniques, we leverage both the naturally observed low-entropy among blocks and the artificially generated low-entropy resulting from our optimization techniques. Based on these two low-entropies, we propose an Entropy-based Pattern Compression (EPC), which generates an inter-block pattern from the same low-entropy region in numerous blocks and then compresses these blocks by using the selected pattern. Our evaluations show that EPC achieves up to 13% (3% on average) higher speedup and 13% (4% on average) DRAM energy consumption reduction with 160x (20x on average) fewer patterns(groups) compared to the state-of-the-art inter-block compression technique.
Jinkwon Kim, Mincheol Kang, Jeongkyu Hong, Soontae Kim
HPCA1
2021 CID: Co-Architecting Instruction Cache and Decompression System for Embedded Systems
abstract
Code compression is widely used to reduce the footprint of code memory in cost-sensitive embedded systems. However, despite the small code size, the decompressor and the address translator required to support the code compression incur energy and area overheads. To reduce such overheads while still supporting code compression, we co-architect the instruction cache and decompression system (CID). In CID, each component is placed at the optimal location and the instruction cache is redesigned to recognize the compression state and retain the original address, through the cache division and address space decompression process. As a result of the cache division, the energy consumption and area overheads of the CID instruction cache are reduced. Since the decompressor overhead depends on the code compression technique, we propose a new code compression technique called entropy-based pattern code compression, which reduces overheads of the decompressor. Our experimental results show that the total energy consumption of the instruction cache and decompression system is reduced by up to 29.7 percent and their area is reduced by up to 15.4 percent compared to the post-cache architecture with almost no performance degradation, while achieving an 18.8 percent improvement in the compression ratio compared to the state-of-the-art code compression technique.
Jinkwon Kim, Seokin Hong, Jeongkyu Hong, Soontae Kim
IEEE Trans. Computers1