EDBT 2026 Demo / reviewers in the wild / expert
Gunho Park
dblp:294/3119
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-8078-4356ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up TablesabstractWeight-only quantization has emerged as a promising solution to the deployment challenges of large language models (LLMs). However, it necessitates FP-INT operations, which make implementation on general-purpose hardware like GPUs difficult. In this paper, we propose FIGLUT, an efficient look-up table (LUT)-based GEMM accelerator architecture. Instead of performing traditional arithmetic operations, FIGLUT retrieves precomputed values from an LUT based on weight patterns, significantly reducing the computational complexity. We also introduce a novel LUT design that addresses the limitations of conventional memory architectures. To further improve LUT-based operations, we propose a half-size LUT combined with a dedicated decoding and multiplexing unit. FIGLUT efficiently supports different bit precisions and quantization methods using a single fixed hardware configuration. For the same 3-bit weight precision, FIGLUT demonstrates 59% higher TOPS/W and 20% lower perplexity than state-of-the-art accelerator design. When targeting the same perplexity, FIGLUT achieves $98 \%$ higher TOPS/W by performing 2.4-bit operations. Gunho Park, Hyeokjun Kwon, Jeongin Bae, Baeseong Park, Dongsoo Lee |
HPCA | 1 |
| 2025 | CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMsabstractWeight-only quantization is widely used to mitigate the memory-bound nature of LLM inference. Codebook-based methods extend this trend by achieving strong accuracy in the extremely low-bit regime (e.g., 2-bit). However, current kernels rely on dequantization, which repeatedly fetches centroids and reconstructs weights, incurring substantial latency and cache pressure. We present CodeGEMM, a codebook-centric GEMM kernel that replaces dequantization with precomputed inner products between centroids and activations stored in a lightweight Psumbook. At inference, code indices directly gather these partial sums, eliminating per-element lookups and reducing the on-chip footprint. The kernel supports the systematic exploration of latency–memory–accuracy trade-offs under a unified implementation.
On Llama-3 models, CodeGEMM delivers 1.83x (8B) and 8.93x (70B) speedups in the 2-bit configuration compared to state-of-the-art codebook-based quantization at comparable accuracy and further improves computing efficiency and memory subsystem utilization. Gunho Park, Jeongin Bae, Byeongwook Kim, Baeseong Park, Jiwon Ryu, Hoseung Kim, Se Jung Kwon, Dongsoo Lee |
NeurIPS | 1 |
| 2024 | LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsabstractRecent advances in self-supervised learning and the Transformer architecture have significantly improved natural language processing (NLP), achieving remarkably low perplexity.
However, the growing size of NLP models introduces a memory wall problem during the generation phase.
To mitigate this issue, recent efforts have focused on quantizing model weights to sub-4-bit precision while preserving full precision for activations, resulting in practical speed-ups during inference on a single GPU.
However, these improvements primarily stem from reduced memory movement, which necessitates a resource-intensive dequantization process rather than actual computational reduction.
In this paper, we introduce LUT-GEMM, an efficient kernel for quantized matrix multiplication, which not only eliminates the resource-intensive dequantization process but also reduces computational costs compared to previous kernels for weight-only quantization.
Furthermore, we proposed group-wise quantization to offer a flexible trade-off between compression ratio and accuracy.
The impact of LUT-GEMM is facilitated by implementing high compression ratios through low-bit quantization and efficient LUT-based operations.
We show experimentally that when applied to the OPT-175B model with 3-bit quantization, LUT-GEMM substantially accelerates token generation latency, achieving a remarkable 2.1x improvement on a single GPU when compared to OPTQ, which relies on the costly dequantization process. Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee 0002, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Dongsoo Lee |
ICLR | 1 |
| 2024 | Low-Power Encoder and Compressor Design for Approximate Radix-8 Booth MultiplierabstractIn this paper, we propose an innovative design for low-power approximate radix-8 Booth multipliers, presenting a promising solution to significantly reduce power consumption in error-resilient signal processing applications. The approach simultaneously approximates both the partial product generation and accumulation stages using an approximate Booth encoder and a 4-2 compressor, achieving substantial energy savings compared to previous designs. Extensive simulations on FIR filtering and image classification validate the method, demonstrating the proposed approximate Booth multiplier’s attractive trade-offs between energy efficiency and accuracy. Experimental results show a remarkable 20% energy reduction compared to the traditional exact Booth multiplier in FIR filtering and image classification, with negligible accuracy loss. Gunho Park |
ISCAS | 2 |
| 2023 | TF-MVP: Novel Sparsity-Aware Transformer Accelerator with Mixed-Length Vector PruningabstractWe present the energy-efficient TF-MVP architecture, a sparsity-aware transformer accelerator, by introducing novel algorithm-hardware co-optimization techniques. From the previous fine-grained pruning map, for the first time, the direction strength is developed to analyze the pruning patterns quantitatively, indicating the major pruning direction and size of each layer. Then, the mixed-length vector pruning (MVP) is proposed to generate the hardware-friendly pruned-transformer model, which is fully supported by our TF-MVP accelerator with the reconfigurable PE structure. Implemented in a 28nm CMOS technology, as a result, TF-MVP achieves 377 GOPs/W for accelerating GPT-2 small model by realizing 4096 multiply-accumulate operators, which is 2.09 times better than the state-of-the-art sparsity-aware transformer accelerator. Eunji Yoo, Gunho Park, Jung Gyu Min, Se Jung Kwon, Baeseong Park, Dongsoo Lee |
DAC | 2 |
| 2023 | Energy-Efficient RISC-V-Based Vector Processor for Cache-Aware Structurally-Pruned TransformersabstractBased on recent RISC-V designs, we present in this paper a low-power vector processor architecture for efficiently deploying vision transformer (ViT) models. To fairly measure the processing efficiency of different processor designs with instruction/data cache memories, we first develop the evaluation framework based on numerous design tools for jointly considering the algorithm, architecture, and circuit performances together, numerically revealing that the previous CSR-based data compression cannot accelerate pruned transformer models at all due to under-utilization of the vector-extended processing units. We then introduce a series of algorithm-hardware co-optimization approaches to greatly minimize cache misses by applying 1) the accuracy-preserved structured ViT pruning, 2) the vertical-CSR (vCSR) data storing format, and 3) vCSR-aware custom memory-accessing instructions. Experimental results show that the proposed optimization schemes eventually improve the processing efficiency of pruned transformers in resource-limited computing platforms, e.g., achieving 11 times lower energy consumption for handling the 0.7-pruned ViT model. Jung Gyu Min, Dongyun Kam, Younghoon Byun, Gunho Park, Youngjoo Lee 0002 |
ISLPED | 4 |
| 2021 | Design and Analysis of Approximate Compressors for Balanced Error Accumulation in MAC OperatorabstractIn this paper, we present a novel approximate computing scheme suitable for realizing the energy-efficient multiply-accumulate (MAC) processing. In contrast to the prior works that suffer from the error accumulation limiting the approximate range, we utilize different approximate multipliers in an interleaved way to compensate errors in the opposite direction during accumulate operations. For the balanced error accumulation, we first design the approximate 4-2 compressors generating errors in the opposite direction while minimizing the computational costs. Based on the probabilistic analysis, positive and negative multipliers are then carefully developed to provide a similar error distance. Simulation results on various practical applications reveal that the proposed MAC processing offers the energy-efficient computing scenario by extending the range of approximate parts. Even compared to the state-of-the-art solutions, for example, the proposed interleaving scheme relaxes the core-level energy consumption of the recent CNN accelerator by more than 35% without degrading the recognition accuracy. Gunho Park, Jaeha Kung 0001, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |