Byeongcheol Kim

dblp:372/0025 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM Inference
abstract
Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often leads to accuracy degradation or hardware overhead due to a mismatch between the global nature of rotation and the localized behavior of group scaling. We propose GyRot, a quantization framework and hardware accelerator that bridges this gap through algorithm-hardware co-design. GyRot introduces Coarse Rotation, Fine Grouping (CoRFiG) and Harmonic-Aligned Permutation (HAP) to enable cooperative integration of rotation and group quantization, enhancing quantizability while relaxing scaling factor precision. To further reduce hardware cost, we reformulate asymmetric quantization and introduce a zero-point rounding strategy that enables fully integer dequantization. Implemented on an INT4-based tensor PE architecture, GyRot achieves state-of-the-art 4-bit accuracy across LLaMA-family models, while delivering up to 3.4× speedup and 3.6× energy efficiency over baseline LLM accelerators. These results validate GyRot's practical effectiveness for scalable and energy-efficient LLM deployment.
Yuseon Chou, Byeongcheol Kim, Jungjun Oh, Hoi-Jun Yoo
HPCA3
2026 SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
abstract
Low-bit quantization is a promising technique for efficient transformer inference by reducing computational and memory overhead. However, aggressive bitwidth reduction remains challenging due to activation outliers, leading to accuracy degradation. Existing methods, such as outlier-handling and group quantization, achieve high accuracy but incur substantial energy consumption. To address this, we propose SeVeDo, an energy-efficient SVD-based heterogeneous accelerator that structurally separates outlier-sensitive components into a high-precision low-rank path, while the remaining computations are executed in a low-bit residual datapath with group quantization. To further enhance efficiency, Hierarchical Group Quantization (HGQ) combines coarse-grained floating-point scaling with fine-grained shifting, effectively reducing dequantization cost. Also, SVD-guided mixed precision (SVD-MP) statically allocates higher bitwidths to precision-sensitive components identified through low-rank decomposition, thereby minimizing floating-point operation cost. Experimental results show that SeVeDo achieves a peak energy efficiency of 13.8TOPS/W, surpassing conventional designs, with 12.7TOPS/W on ViT-Base and 13.4TOPS/W on Llama2-7B benchmarks.
Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo
ISCAS4
2026 An Energy-Efficient High Resolution Vision Transformer Processor Exploiting Token Similarity Beyond Token Merging
Jungjun Oh, Junha Ryu, Byeongcheol Kim, Yuseon Choi, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.5
2025 A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform Quantization
abstract
Non-uniform Quantization (NUQ) is widely used in LLM accelerators due to its high accuracy. However, employing NUQ with models of varying sizes can substantially increase storage requirements on mobile devices. This paper presents a bit-layered NUQ accelerator architecture that supports multiple bit-width configurations while minimizing memory usage. Key features include Reconfigurable Condensed Look-up Accumulator (RCLA), Dual-Sign Path Accumulation (DSPA), and MSB-Sparse Encoding Compression (MSEC). RCLA enables the use of multiple weight precisions within a uniform PE array. In particular, it optimizes PE utilization in high-bit-width NUQ modes, reducing accumulation cycles by 63.2 %. DSPA facilitates energy-efficient computation, resulting in an average power reduction of 40.7 % across various weight modes. MSEC enhances weight compression, reducing the data storage size of each bit plane by up to 47.7 %. The proposed design supports models of different sizes, improves energy efficiency, and reduces memory capacity requirements, rendering it ideal for mobile LLM inference.
Byeongcheol Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo
ISCAS1
2025 A 9.6 TOPS/W Vision Transformer Processor with Hierarchical Token Merging for Similarity-Driven Difference Computing
abstract
Token merging is widely used in Vision Transformers(ViT) as an effective method to reduce computation with minimal accuracy drop. However, aggressive token elimination methods like token merging have limitations in applications requiring fine-grained output, such as text-unified object recognition. These methods reach an upper limit on token reduction due to an accuracy loss. This paper proposes exploiting token similarity beyond the merging upper bound to further reduce power consumption. The key features include Hierarchical Token Merging (HTM), Sparsity Separated Accumulation with Sign Magnitude Data Representation(SSA-SM), and Bidirectional Dynamic Allocation (BDA). HTM begins by performing token merging in the first iteration of the transformer block to reduce 34 % of tokens. Then, during the subsequent iterations, similar token difference computing is applied to the remaining tokens and reduces the effective bit by 31 %. SSA-SM then optimizes PE operations to fully exploit this sparsity by reducing the computation logic toggle rate, resulting in a 45 % higher energy efficiency. Lastly, BDA introduces a memory storage technique that ensures sparse and dense inputs are fed separately into the PE, resulting in an additional 14 % power savings. This proposed method improves energy efficiency by 2.39 times and achieves 9.6 TOPS/W for the text-unified object recognition application on the MS-COCO dataset.
Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo
ISCAS4