Seungjae Moon

dblp:322/1163 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0002-5924-7000ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Adelia: A 4nm LLM Processor for Efficient Generative Al Inference
Seungjae Moon, Juntaek Oh, Jay Kim, Joo-Young Kim 0001
HCS1
2025 Hybe: GPU-NPU Hybrid System for Efficient LLM Inference with Million-Token Context Window
abstract
The growth of context window size in large language model (LLM) inference poses a very distinct computational challenge of hardware inefficiency.The inefficiency arises from the computational imbalance during LLM inference between the compute-intensive prefill stage, and memory-intensive decode stage.The predominant inference hardware, GPU, boasts large number of cores to excel in the prefill stage, which processes the entire input context at once, but suffers from hardware underutilization in the decode stage, which iteratively generates one output token at a time.In conventional LLM, batching has been able to alleviate the underutilization by generating multiple tokens of different requests.However, batching becomes infeasible in models with large context windows over 100K tokens because the Key-Value (KV) activations dominate the physical memory capacity, surpassing the entire model size.In this paper, we propose Hybe, a GPU-NPU hybrid system for efficient LLM inference with a million-token context window.Hybe utilizes the preexisting GPU for the prefill stage and employs lightweight NPUs during the decode stage.Each NPU includes only the necessary computing resources to fully utilize the given memory bandwidth, thereby achieving maximum hardware efficiency.Furthermore, Hybe introduces fine-grained KV transmission, a kernel scheduling method that immediately offloads partial KV produced from the GPU to the NPU, which significantly reduces the KV memory required in the GPU.Lastly, Hybe scheduler applies stage-wise pipelining that dynamically assigns queued requests to idle hardware to minimize stalls.Hybe utilizes NVIDIA H100 GPU with inference-optimized vLLM library and implement Hybe NPU in 4nm process with equal HBM specification.Hybe achieves 2.1× speedup for Phi-3 with 100K-token context window and 3.9× energy efficiency for Llama-3 with 1M-token context window, over H100 GPUs with equal total device count.
Seungjae Moon, Junseo Cha, Joo-Young Kim 0001
ISCA1
2024 BLESS: Bandwidth and Locality Enhanced SMEM Seeding Acceleration for DNA Sequencing
abstract
In an era marked by the pervasive spread of harmful viruses like COVID-19, the importance of DNA sequencing has grown significantly, given its crucial role in devising effective countermeasures. The seeding process, which aims to find locations of super-maximal exact matches (SMEM) between the DNA samples and reference genome for comparative analysis, has emerged as a major bottleneck due to its memory-intensive characteristics. The learned index approach has been developed that uses machine learning model to partially predict the location of the exact matches, which has effectively reduced the memory access. However, the lack of locality in the current in dexing structure and randomness at runtime of the seeding workload have constrained the memory bandwidth usage and have limited further performance advantage. In this paper, we propose BLESS, a bandwidth and locality enhanced SMEM seeding accelerator for learned-index-based DNA sequence alignment. BLESS is the first domain-specific seeding accelerator to maximize the potential hardware advantage of the learned index approach. We introduce coarse-fine (CF) block data structure, a novel memory mapping of seeding parameters to exploit spatial locality and increase effective bandwidth usage for any memory type, including high bandwidth memory (HBM). We also develop guaranteed search range update (GSRU) algorithm, a method that exploits caching in the search procedure to enable temporal locality and data reuse. Utilizing the CF block and GSRU algorithm, we develop a multi-core seeding accelerator using HBM with context switching and runtime scheduling for maximum core and memory bandwidth utilization. With these improvements, BLESS achieves $35.65 \times$ and $15.49 \times$ speedup over the state-of-the-art seeding system BWA-MEME and ERT-ASIC, respectively, in raw system performance.
Seungjae Moon, Teokkyu Suh, Jaehoon Heo, Joo-Young Kim 0001
ISCA2
2023 HyperAccel Latency Processing Unit (LPUTM) Accelerating Hyperscale Models for Generative AI
Seungjae Moon, Junsoo Kim 0002, Junseo Cha, Gyubin Choi, Seongmin Hong, Joo-Young Kim 0006
HCS1
2022 DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
abstract
•DFX: a low-latency multi-FPGA appliance for accelerating transformer-based text generation–DFX is a multi-FPGA appliance that accelerates transformer-based text generation–DFX adopts model parallelism to efficiently process the large-scale language model–Xilinx Alveo U280 data center accelerator card provides high performance with low-cost–FPGA-to-FPGA communication is enabled by QSFP cable at 100 Gb/s
Seongmin Hong, Seungjae Moon, Junsoo Kim 0002, Sungjae Lee 0002, Minsub Kim, Dongsoo Lee, Joo-Young Kim 0001
HCS2
2022 DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
abstract
Transformer is a deep learning language model widely used for natural language processing (NLP) services in datacenters. Among transformer models, Generative Pretrained Transformer (GPT) has achieved remarkable performance in text generation, or natural language generation (NLG), which needs the processing of a large input context in the summarization stage, followed by the generation stage that produces a single word at a time. The conventional platforms such as GPU are specialized for the parallel processing of large inputs in the summarization stage, but their performance significantly degrades in the generation stage due to its sequential characteristic. Therefore, an efficient hardware platform is required to address the high latency caused by the sequential characteristic of text generation. In this paper, we present DFX, a multi-FPGA acceleration appliance that executes GPT-2 model inference end-to-end with low latency and high throughput in both summarization and generation stages. DFX uses model parallelism and optimized dataflow that is model-and-hardware-aware for fast simultaneous workload execution among devices. Its compute cores operate on custom instructions and provide GPT-2 operations end-to-end. We implement the proposed hardware architecture on four Xilinx Alveo U280 FPGAs and utilize all of the channels of the high bandwidth memory (HBM) and the maximum number of compute resources for high hardware efficiency. DFX achieves 5.58$\times$ speedup and 3.99$\times$ energy efficiency over four NVIDIA V100 GPUs on the modern GPT-2 model. DFX is also 8.21$\times$ more cost-effective than the GPU appliance, suggesting that it is a promising solution for text generation workloads in cloud datacenters.
Seongmin Hong, Seungjae Moon, Junsoo Kim 0002, Sungjae Lee 0002, Minsub Kim, Dongsoo Lee, Joo-Young Kim 0001
MICRO2