Junsoo Kim 0002

dblp:83/8851-2 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6680-2602ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
abstract
Modern Large Language Model (LLM) serving system batches multiple requests to achieve high throughput, while batching attention operations is challenging, rendering memory bandwidth a critical bottleneck.Today, to mitigate this issue, the community relies on high-end GPUs with multiple high-bandwidth memory (HBM) channels.Unfortunately, HBM's high bandwidth often comes at the expense of limited memory capacity, necessitating systems to scale, which reduces core utilization and increases costs.Moreover, recent advancements enabling longer contexts for LLMs have substantially increased the key-value (KV) cache size, further intensifying the pressures on memory capacity.To lower the pressure, the literature has explored KV cache quantization techniques, which commonly use low bitwidth (e.g., INT4) for most values, selectively using higher bitwidth (e.g., FP16) for outlier values.While this approach helps achieve high accuracy and low bitwidth simultaneously, it comes with the limitation that the cost for online outlier detection is excessively high, negating the advantages of quantization.Inspired by these insights, we propose Oaken, an acceleration solution that achieves high accuracy and high performance simultaneously through co-designing algorithm and hardware.To effectively find a sweet spot in the accuracy-performance trade-off space of KV cache quantization, Oaken employs an online-offline hybrid approach, setting outlier thresholds offline, which are then used to determine the quantization scale online.To translate the proposed algorithmic technique into tangible performance gains, Oaken also comes with custom quantization/dequantization engines and memory management units that can be integrated with any LLM accelerators.We built an Oaken accelerator on top of
Minsu Kim 0004, Seongmin Hong, Ryeowook Ko, Soongyu Choi, Hunjong Lee, Junsoo Kim 0002, Joo-Young Kim 0001, Jongse Park
ISCA6
2025 ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
abstract
The growing adoption of Large Language Models (LLMs) across various domains has driven the demand for efficient and scalable AI-serving solutions. Deploying LLMs requires optimizations to manage their significant computational and data demands. The prefill stage processes large numbers of input tokens in parallel, increasing computational load, while the decoding stage relies heavily on memory bandwidth due to the auto-regressive nature of LLMs. Current hardware, such as GPUs, often fails to balance these demands, leading to inefficient utilization. While batching improves hardware efficiency, it delays response times, degrading Quality-of-Service (QoS). This disconnect between vendors, who aim to maximize resource efficiency, and users, who prioritize low latency, highlights the need for a better solution. To address this, we propose ADOR, a framework that automatically identifies and recommends hardware architectures tailored to LLM serving. By lever-aging predefined architecture templates specialized for heterogeneous dataflows, ADOR optimally balances throughput and latency. It efficiently explores design spaces to suggest architectures that meet the requirements of both vendors and users. ADOR demonstrates substantial performance improvements, achieving$2.51 \times$higher QoS and$4.01 \times$better area efficiency compared to the A100 at high batch sizes, making it a robust solution for scalable and cost-effective LLM serving.
Junsoo Kim 0002, Hunjong Lee, Geonwoo Ko, Gyubin Choi, Seri Ham, Seongmin Hong, Joo-Young Kim 0001
ISPASS1
2023 HyperAccel Latency Processing Unit (LPUTM) Accelerating Hyperscale Models for Generative AI
Seungjae Moon, Junsoo Kim 0002, Junseo Cha, Gyubin Choi, Seongmin Hong, Joo-Young Kim 0006
HCS2
2022 OpenMDS: An Open-Source Shell Generation Framework for High-Performance Design on Multi-Die FPGAs
abstract
FPGA is a promising platform in designing a hardware accelerator due to its design flexibility and fast development cycle, despite the device's limited hardware resources. To address this, latest FPGAs have adopted a multi-die architecture providing abundant hardware resources with high yield and cost-benefit. However, the multi-die architecture causes critical timing issues when signal paths cross the die-to-die boundaries, adding another design challenge in using FPGA. We propose OpenMDS, an open-source shell generation framework for high-performance design on multi-die FPGAs. Based on the user's design requirements, it generates an optimized shell for the target FPGA via automated bus pipelining, customized floorplanning, and scalable clocking scheme.
Gyeongcheol Shin, Junsoo Kim 0002, Joo-Young Kim 0001
FCCM2
2022 DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
abstract
•DFX: a low-latency multi-FPGA appliance for accelerating transformer-based text generation–DFX is a multi-FPGA appliance that accelerates transformer-based text generation–DFX adopts model parallelism to efficiently process the large-scale language model–Xilinx Alveo U280 data center accelerator card provides high performance with low-cost–FPGA-to-FPGA communication is enabled by QSFP cable at 100 Gb/s
Seongmin Hong, Seungjae Moon, Junsoo Kim 0002, Sungjae Lee 0002, Minsub Kim, Dongsoo Lee, Joo-Young Kim 0001
HCS3
2022 DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
abstract
Transformer is a deep learning language model widely used for natural language processing (NLP) services in datacenters. Among transformer models, Generative Pretrained Transformer (GPT) has achieved remarkable performance in text generation, or natural language generation (NLG), which needs the processing of a large input context in the summarization stage, followed by the generation stage that produces a single word at a time. The conventional platforms such as GPU are specialized for the parallel processing of large inputs in the summarization stage, but their performance significantly degrades in the generation stage due to its sequential characteristic. Therefore, an efficient hardware platform is required to address the high latency caused by the sequential characteristic of text generation. In this paper, we present DFX, a multi-FPGA acceleration appliance that executes GPT-2 model inference end-to-end with low latency and high throughput in both summarization and generation stages. DFX uses model parallelism and optimized dataflow that is model-and-hardware-aware for fast simultaneous workload execution among devices. Its compute cores operate on custom instructions and provide GPT-2 operations end-to-end. We implement the proposed hardware architecture on four Xilinx Alveo U280 FPGAs and utilize all of the channels of the high bandwidth memory (HBM) and the maximum number of compute resources for high hardware efficiency. DFX achieves 5.58$\times$ speedup and 3.99$\times$ energy efficiency over four NVIDIA V100 GPUs on the modern GPT-2 model. DFX is also 8.21$\times$ more cost-effective than the GPU appliance, suggesting that it is a promising solution for text generation workloads in cloud datacenters.
Seongmin Hong, Seungjae Moon, Junsoo Kim 0002, Sungjae Lee 0002, Minsub Kim, Dongsoo Lee, Joo-Young Kim 0001
MICRO3