Yoonho Jang

dblp:177/3183 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 LibraPIM: Dynamic Load Rebalancing to Maximize Utilization in PIM-Assisted LLM Inference Systems
abstract
Large language models (LLMs) require inference systems that can handle both compute- and memory-intensive workloads. GPUs and NPUs (referred to as xPUs) efficiently process compute-intensive layers, while Processing-In-Memory (PIM) architectures are well suited for memory-bound stages. To exploit this complementary relationship, recent LLM inference systems have adopted heterogeneous architectures integrating xPUs with PIM units. However, this integration poses several challenges. Tight execution dependencies between the two devices limit concurrency, as PIM often must wait for data produced by the xPUs. Moreover, batch size and sequence length affect the computational load on PIM and the accelerator differently, leading to execution imbalance across devices and degrading their utilization. To address these challenges, we propose LibraPIM, a novel PIM framework that orchestrates workload rebalancing and concurrent execution between compute accelerators and in-memory compute units, enabling efficient and scalable LLM inference. LibraPIM addresses the above challenges through two key techniques: Dynamic Batch Offloading (DBO), which adaptively redirects portions of the workload to the underutilized device based on runtime profiling, and Dual-Path Execution (DEX), which enables concurrent PIM operations and memory accesses through sub-bank partitioning with minimal hardware overhead. Together, these techniques improve resource utilization across heterogeneous devices, thereby increasing system throughput in LLM inference. Our experimental results demonstrate that LibraPIM achieves $6.2 \times$ average speedup over the baseline PIM-enabled system and delivers a $2.1 \times$ average speedup compared to a state-of-the-art approach.
Hyeongjun Cho, Yoonho Jang, Hyungi Kim, Seongwook Kim, Keewon Kwon, Gwangsun Kim, Seokin Hong
PACT2
2025 PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic Lookup
abstract
Deploying Large Language Models (LLMs) on edge devices poses significant challenges due to their high computational and memory demands. In particular, General MatrixVector Multiplication (GEMV), a key operation in LLM inference, is highly memory-intensive, making it difficult to accelerate using conventional edge computing systems. While Processing-in-memory (PIM) architectures have emerged as a promising solution to this challenge, they often suffer from high area overhead or restricted computational precision. This paper proposes PIMPAL (Processing-In-Memory architecture with Parallel Arithmetic Lookup), a cost-effective PIM architecture leveraging LookUp Table (LUT)-based computation for GEMV acceleration in sLLMs (small LLMs). By replacing traditional arithmetic operations with parallel in-DRAM LUT lookups, PIMPAL significantly reduces area overhead while maintaining high performance. PIMPAL introduces three key innovations: (1) it divides DRAM bank subarrays into compute blocks for parallel LUT processing; (2) it employs Localityaware Compute Mapping (LCM) to reduce row activations by maximizing LUT access locality; and (3) it enables multi-precision computations through a LUT Aggregation (LAG) mechanism that combines results from multiple small LUTs. Experimental results show that PIMPAL achieves up to $17.8 x$ higher performance than previous LUT-based PIM designs and reduces area overhead by $40 \%$ compared to conventional processing unit-based PIM designs.
Yoonho Jang, Hyeongjun Cho, Yesin Ryu, Jungrae Kim, Seokin Hong
DAC1
2025 Minimizing Read Disturb via Localized Page Allocation for Modern NAND Flash-Based SSDs
abstract
To meet the increasing demand for higher storage density, modern NAND flash-based SSDs employ Quad-Level Cell (QLC) technology, which stores four bits per memory cell. However, it exacerbates the read disturb issue, where repeated read operations gradually shift the threshold voltages of unselected NAND flash cells. This voltage drift leads to frequent read retries and accelerates device wear, ultimately degrading performance and endurance. To address this issue, we propose Localized Page Allocation (LPA), a novel data mapping scheme that allocates each page to a restricted region of the wordline. LPA assigns four consecutive bits in a single page to a single cell. This approach confines read operations to a limited portion of the bitlines, thereby reducing the number of NAND flash cells exposed to read disturb. However, this localized access increases the complexity of sensing operations. To address this challenge, we introduce Progressive Voltage Convergence (PVC) method that adaptively determines the next sensing voltage for each bitline based on its previous sensing result. This fine-grained control reduces the number of sensing steps during read operations. For additional optimization, we employ lossless compression, which reduces the number of sensing steps for the compressed pages. Experimental evaluations using real-world workloads demonstrate that our design improves I/O performance by up to$\mathbf{1 1 \%}$and extends endurance by 79% compared to the state-of-the-art techniques. Our design requires minor hardware modifications to the conventional NAND flash chips and the SSD controller, which leads to small area and power overhead.
Joonseong Hwang, Minjin Park, Jihun Yoon, Yoonho Jang, Seokin Hong
ICCD5