Bongjoon Hyun

dblp:254/0959 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-2529-7264ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
abstract
The ability to dynamically allocate memory is fundamental in modern programming languages. However, this feature is not adequately supported in current general-purpose PIM devices. To identify key design principles that PIM must consider, we conduct a design space exploration of PIM memory allocators, examining various strategies for metadata placement and management of the allocator. Based on this exploration, we introduce PIM-malloc, a fast and scalable memory allocator for general-purpose PIM that operates on real PIM hardware, achieving a$66 \times$improvement in memory allocation performance. This design is further enhanced with a lightweight, per-PIM core hardware cache, specifically designed for dynamic memory allocation, achieving an additional 31% performance improvement. Finally, we demonstrate the applicability of PIM-malloc by developing several representative PIM workloads, demonstrating its effectiveness in enhancing programmability.
Bongjoon Hyun, Youngjin Kwon, Minsoo Rhu
HPCA2
2025 PIM-CCA: An Efficient PIM Architecture with Optimized Integration of Configurable Functional Units
abstract
Processing-in-Memory (PIM) is a promising architecture for alleviating data movement bottlenecks by performing computations closer to memory.However, PIM workloads often encounter computational bottlenecks within the PIM itself.As these workloads become more compute-intensive by leveraging PIM's high internal bandwidth, a small set of hot code regions emerges as the primary performance bottleneck.Unfortunately, increasing the complexity of the PIM processor is difficult due to inherent memory constraints, such as area and power.Therefore, enhancing the computational capability of PIM while maintaining a lightweight design within limited silicon budgets remains highly challenging.In this paper, we propose PIM-CCA, a novel PIM architecture that integrates a Configurable Compute Accelerator (CCA) to mitigate computational bottlenecks with minimal hardware overhead.The CCA-enabled PIM design allows for the flexible configuration of compute logic, enabling acceleration across diverse workloads.The PIM-CCA compiler constructs an instruction-level dataflow graph to identify hot and compute-bound regions and offload them to the CCA.Furthermore, we analyze the interaction between the PIM threading model and resource utilization to derive the optimal thread count for efficient CCA-enabled PIM usage.We implement PIM-CCA in a cycle-accurate simulator based on a commercially available PIM system, and evaluate it using 14 representative benchmarks.The experimental results show that PIM-CCA achieves up to 1.55× performance improvement over baseline PIM systems, with only 0.036% additional area overhead, based on P&R results with limited metal layers.
Jeehyun Kim, Donghyeon Kim 0001, Seokwon Kang, Bongjoon Hyun, Inho Lee 0002, Yongjun Park 0001
MICRO4
2024 Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
abstract
Processing-in-memory (PIM) has been explored for decades by computer architects, yet it has never seen the light of day in real-world products due to its high design overheads and lack of a killer application. With the advent of critical memoryintensive workloads, several commercial PIM technologies have been introduced to the market, ranging from domain-specific PIM architectures to more general-purpose PIM architectures. In this work, we deepdive into UPMEM's commercial PIM technology, a general-purpose PIM-enabled parallel computing architecture that is highly programmable. Our first key contribution is the development of a flexible simulation framework for PIM. The simulator we developed (aka uPIMulator) enables the compilation of UPMEM-PIM source codes into its compiled machine-level instructions, which are subsequently consumed by our cycle-level performance simulator. Using uPIMulator, we demystify UPMEM's PIM design through a detailed characterization study. Finally, we identify some key limitations of the current UPMEM-PIM system through our case studies and present some important architectural features that will become critical for future PIM architectures to support.
Bongjoon Hyun, Minsoo Rhu
HPCA1
2024 PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
abstract
Processing-in-memory (PIM) has emerged as a promising solution for accelerating memory-intensive workloads as they provide high memory bandwidth to the processing units. This approach has drawn attention not only from the academic community but also from the industry, leading to the development of real-world commercial PIM devices. In this work, we first conduct an in-depth characterization on UPMEM's general-purpose PIM system and analyze the bottlenecks caused by the data transfers across the DRAM and PIM address space. Our characterization study reveals several critical challenges associated with DRAM↔PIM data transfers in memory bus integrated PIM systems, for instance, its high CPU core utilization, high power consumption, and low read/write throughput for both DRAM and PIM. Driven by our key findings, we introduce the PIM-MMU architecture which is a hardware/software co-design that enables energy-efficient DRAM↔PIM transfers for PIM systems. PIM-MMU synergistically combines a hardware-based data copy engine, a PIM-optimized memory scheduler, and a heterogeneity-aware memory mapping function, the utilization of which is supported by our PIM-MMU software stack, significantly improving the efficiency of DRAM↔PIM data transfers. Experimental results show that PIM-MMU improves the DRAM↔PIM data transfer throughput by an average$4.1\times$and enhances its energy-efficiency by$4.1\times$, leading to a$2.2\times$end-to-end speedup for real-world PIM workloads.
Bongjoon Hyun, Minsoo Rhu
MICRO2
2020 NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing Units
abstract
To satisfy the compute and memory demands of deep neural networks (DNNs), neural processing units (NPUs) are widely being utilized for accelerating DNNs. Similar to how GPUs have evolved from a slave device into a mainstream processor architecture, it is likely that NPUs will become first-class citizens in this fast-evolving heterogeneous architecture space. This paper makes a case for enabling address translation in NPUs to decouple the virtual and physical memory address space. Through a careful data-driven application characterization study, we root-cause several limitations of prior GPU-centric address translation schemes and propose a memory management unit (MMU) that is tailored for NPUs. Compared to an oracular MMU design point, our proposal incurs only an average 0.06% performance overhead.
Bongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim 0001, Minsoo Rhu
ASPLOS1