Byungkuk Yoon

dblp:406/1781 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 86% Hardware accelerators and domain-specific architectures · 14%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
processing-in-memory
1.722025
SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIM · DAC 2025
Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid Bonding · DAC 2025
Natural language and speech › Language models and text generation
large language model inference
0.912025
Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid Bonding · DAC 2025
Memory systems › DRAM › DRAM microarchitecture
bank-level parallelism
0.912025
SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIM · DAC 2025
Memory systems
DRAM
0.912025
SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIM · DAC 2025
Memory systems › processing-in-memory
DRAM-PIM
0.912025
SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIM · DAC 2025
Memory systems › processing-in-memory › memory-centric computing
near-memory accelerator
0.912025
Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid Bonding · DAC 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.912025
Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid Bonding · DAC 2025

Methods — techniques the papers use, named apart from their topics

dual-IO scheme · 1.7centralized controller · 1.73d hybrid bonding · 1.7asynchronous execution · 0.9SIMD · 0.9
YearPublicationVenuePosition
2025 Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid Bonding
abstract
Large language model (LLM) inference poses dual challenges, demanding substantial memory bandwidth and computing resources. Recent advancements in near-memory accelerators leveraging 3D DRAM-to-logic hybrid-bonding (HB) interconnects have gained attention due to their highly parallel data transfer capabilities. We address limitations in previous HB-DRAM accelerators, such as those stemming from distributed controller designs, by introducing an architecture with a centralized controller and dual-IO scheme. This approach not only reduces the chip area overhead but also enables reconfigurable GEMV/GEMM operations, boosting the performance. Simulations for the OPT 66B model show that our proposed accelerator achieves 2.9X, 3.5X, and 2.5X higher performance compared to NPU, DRAM-PIM, and heterogeneous designs (DRAM-PIM + NPU), respectively.
Sanghyeok Han, Byungkuk Yoon, Gyeonghwan Park, Choungki Song, Dongkyun Kim, Jae-Joon Kim
DAC2
2025 SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIM
abstract
Processing in Memory (PIM) architectures enhance memory bandwidth by utilizing bank-level parallelism, typically implemented with a SIMD structure where all banks operate simultaneously under a single command. However, this synchronous approach requires the activation of all banks before computation, leading to activation times that exceed computation times, limiting performance gain. Recently, asynchronous execution PIM has been proposed as an alternative, allowing banks to operate asynchronously and overlap activation with processing to hide the row activation overhead. While effective at reducing row activation overhead, the independent operation requires large shared accumulators for each bank group, increasing area overhead. To address the issues, we propose bank group (BG)-level split synchronization DRAM PIM, where each bank group operates asynchronously to hide row activation overhead while operating synchronously within the bank group to eliminate the need for shared accumulators. Evaluation results show that our proposed design achieves an average throughput improvement of $1.70 x$ and $1.06 x$ compared to conventional PIM and asynchronous execution PIM. Furthermore, the area overhead per processing unit (PU) increases by only $1.5 \%$ compared to conventional PIM and is significantly lower than that of asynchronous execution PIM.
Byungkuk Yoon, Sanghyeok Han, Gyeonghwan Park, Jae-Joon Kim
DAC1
2025 DOTS: DRAM-PIM Optimization for Tall and Skinny GEMM Operations in LLM Inference
abstract
For large language models (LLMs), increasing token lengths require smaller batch sizes due to increase in memory requirement for KV caching, leading to under-utilization of processing units and memory bandwidth bottleneck in NPUs. To address the challenge, we propose DOTS, a new DRAM-PIM architecture that can handle both GEMV and GEMM efficiently, even outperforming NPUs in GEMM operations when batch sizes are small. The proposed DRAM-PIM reduces power consumption and latency caused by frequent DRAM row activation switching in conventional DRAM PIMs with negligible hardware overhead. Simulation results show that our proposed design achieves throughput improvements of 1.83x, 1.92x, and 1.7x over GPU, NPU, and heterogeneous NPU/PIM systems, respectively, for models as large as or larger than OPT-175B.
Gyeonghwan Park, Sanghyeok Han, Byungkuk Yoon, Jae-Joon Kim
DATE3