Sanghyeok Han

dblp:406/1204 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid Bonding
abstract
Large language model (LLM) inference poses dual challenges, demanding substantial memory bandwidth and computing resources. Recent advancements in near-memory accelerators leveraging 3D DRAM-to-logic hybrid-bonding (HB) interconnects have gained attention due to their highly parallel data transfer capabilities. We address limitations in previous HB-DRAM accelerators, such as those stemming from distributed controller designs, by introducing an architecture with a centralized controller and dual-IO scheme. This approach not only reduces the chip area overhead but also enables reconfigurable GEMV/GEMM operations, boosting the performance. Simulations for the OPT 66B model show that our proposed accelerator achieves 2.9X, 3.5X, and 2.5X higher performance compared to NPU, DRAM-PIM, and heterogeneous designs (DRAM-PIM + NPU), respectively.
Sanghyeok Han, Byungkuk Yoon, Gyeonghwan Park, Choungki Song, Dongkyun Kim, Jae-Joon Kim
DAC1
2025 SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIM
abstract
Processing in Memory (PIM) architectures enhance memory bandwidth by utilizing bank-level parallelism, typically implemented with a SIMD structure where all banks operate simultaneously under a single command. However, this synchronous approach requires the activation of all banks before computation, leading to activation times that exceed computation times, limiting performance gain. Recently, asynchronous execution PIM has been proposed as an alternative, allowing banks to operate asynchronously and overlap activation with processing to hide the row activation overhead. While effective at reducing row activation overhead, the independent operation requires large shared accumulators for each bank group, increasing area overhead. To address the issues, we propose bank group (BG)-level split synchronization DRAM PIM, where each bank group operates asynchronously to hide row activation overhead while operating synchronously within the bank group to eliminate the need for shared accumulators. Evaluation results show that our proposed design achieves an average throughput improvement of $1.70 x$ and $1.06 x$ compared to conventional PIM and asynchronous execution PIM. Furthermore, the area overhead per processing unit (PU) increases by only $1.5 \%$ compared to conventional PIM and is significantly lower than that of asynchronous execution PIM.
Byungkuk Yoon, Sanghyeok Han, Gyeonghwan Park, Jae-Joon Kim
DAC2
2025 DOTS: DRAM-PIM Optimization for Tall and Skinny GEMM Operations in LLM Inference
abstract
For large language models (LLMs), increasing token lengths require smaller batch sizes due to increase in memory requirement for KV caching, leading to under-utilization of processing units and memory bandwidth bottleneck in NPUs. To address the challenge, we propose DOTS, a new DRAM-PIM architecture that can handle both GEMV and GEMM efficiently, even outperforming NPUs in GEMM operations when batch sizes are small. The proposed DRAM-PIM reduces power consumption and latency caused by frequent DRAM row activation switching in conventional DRAM PIMs with negligible hardware overhead. Simulation results show that our proposed design achieves throughput improvements of 1.83x, 1.92x, and 1.7x over GPU, NPU, and heterogeneous NPU/PIM systems, respectively, for models as large as or larger than OPT-175B.
Gyeonghwan Park, Sanghyeok Han, Byungkuk Yoon, Jae-Joon Kim
DATE2
2025 LLM-on-the-Palm: Mobile LLM Inference with PIM-Enhanced NAND Flash Memory
abstract
Large Language Model (LLM) inference on edge devices is crucial for democratizing AI and addressing privacy and security concerns associated with cloud services. However, the large parameter sizes of LLMs pose significant challenges in deployment on resource-constrained devices, affecting user experience. Mobile DRAM typically cannot accommodate these parameters, necessitating the use of NAND flash memory, but it has considerably slower access speeds. To overcome the limitations, we propose LLM-on-the-Palm, an LLM inference system for smartphones that leverages Processing-in-Memory (PIM)-enhanced NAND flash memory to enable fast inference of billion-parameter LLMs with minimal hardware modifications. Our approach requires only a single multiply-accumulate (MAC) unit per plane in a NAND flash memory - 256 MAC units for 256 GB SSD - without compromising cell density. Simulation results show that our design achieves ∼100 ms per token generation on a 6.7B model under specifications similar to an iPhone-15.
Sanghyeok Han, Jae-Joon Kim
ICCAD2
2025 CrossBit: Bitwise Computing in NAND Flash Memory with Inter-Bitline Data Communication
abstract
In-flash processing (IFP), which involves performing data computation inside NAND flash memory, holds high potential for improving the performance and energy efficiency of data-intensive application by minimizing data movement.Recent research has introduced several IFP architectures enabling bulk bitwise operations inside NAND flash chips to demonstrate this potential.However, previous IFP designs were limited to performing bitwise operations on data sensed within the same bitline, thus lacking the capability to handle more complex functions requiring interactions between data from different bitlines.This paper presents CrossBit, a new IFP architecture that enables both intra-bitline and inter-bitline operations with minimal additional circuitry integrated into commodity NAND flash memory.With the capability for inter-bitline operations, CrossBit facilitates in-flash error correction code (IF-ECC) operations, thereby enabling reliable multi-level cell (MLC) IFP.Moreover, CrossBit efficiently processes fundamental database queries that were previously inefficient with existing work that only supports intra-bitline operations.Experimental results show that the implementation of IF-ECC in CrossBit results in a substantial reduction in bit error rate (BER) for MLC operations, leading to 1.8× increase in bit-density by using MLC compared to previous IFP designs which uses SLC only.When used for accelerating fundamental database queries, CrossBit achieves notable average speedup and energy efficiency improvement of 2.1× and 2.5× compared to the state-of-the-art (SOTA) IFP architecture.We further demonstrate the practicality by processing the full set of end-toend database queries from the widely used Star schema benchmark, where CrossBit achieves 1.7× speedup.
Seunghwan Song, Sukhyun Choi, Jeongin Choe, Sanghyeok Han, Jisung Park 0001, Jinho Lee 0001, Jae-Joon Kim
MICRO5