Chunmyung Park

dblp:342/8735 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0005-0033-7302ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Communication-Aware Hybrid Parallelism Mapping for Low-Cost MCM-based DNN Accelerators
abstract
The growing scale of deep neural networks has surpassed the capacity of single-chip accelerators, particularly pin cost-sensitive edge devices. Multi-Chip-Module (MCM) architectures enable scalability but rely on bandwidth-limited chip-to-chip (C2C) interfaces, causing substantial inter-chip communication overhead. Among model-parallel strategies, tensor parallelism (TP) offers high concurrency at the cost of communication overhead, while pipeline parallelism (PP) reduces it at the cost of lower compute utilization inherent to pipeline execution. This work presents Stitch, a two-phase rebalancing framework for hybrid model-parallel mapping in low-cost MCM-based CNN accelerators. Phase I mitigates TP’s C2C-induced communication overhead by jointly optimizing partitioning and datapath through a layer-wise C2C–DRAM selection solved via dynamic programming. Since TP alone cannot fully minimize communication, Phase II extends the design space by combining TP and PP at the package level. Guided by simulated annealing, Stitch selects layer groups, tunes pipeline stages, and balances communication–utilization trade-offs. Evaluation on a cycle-accurate simulator shows that Stitch reduces the energy–delay product by up to 42.8% compared to prior TP-based methods, demonstrating its effectiveness under practical C2C bandwidth constraints.
Jicheon Kim, Chunmyung Park, Xuan Truong Nguyen
DATE2
2025 Leveraging Hot Data in a Multi-Tenant Accelerator for Effective Shared Memory Management
abstract
Multi-tenant neural networks (MTNN) have been emerging in various domains. To effectively handle multi-tenant workloads, modern hardware systems typically incorporate multiple compute cores with shared memory systems. While prior works have intensively studied compute- and bandwidth-aware allocation, on-chip memory allocation for MTNN accelerators has not been well studied. This work identifies two key challenges of on-chip memory allocation in MTNN accelerators: on-chip memory shortages, which force data eviction to off-chip memory, and on-chip memory underutilization, where memory remains idle due to coarse-grained allocation. Both issues lead to increased external memory accesses (EMAs), significantly degrading system performance. To address these challenges, we propose HotPot, a novel multi-tenant accelerator with a runtime temperature-aware memory allocator. HotPot prioritizes hot data for global on-chip memory allocation, reducing unnecessary EMAs and optimizing memory utilization. Specifically, HotPot introduces a temperature score that quantifies reuse potential and guides runtime memory allocation decisions. Experimental results demonstrate that HotPot improves system throughput (STP) by up to 1.88 × and average normalized turnaround time (ANTT) by 1.52 × compared to baseline methods.
Chunmyung Park, Jicheon Kim, Eunjae Hyun, Xuan Truong Nguyen
DATE1
2025 Live Demonstration: A Scalable CNN Accelerator SoC With a Cost-Effective Chip-to-Chip Adapter
abstract
In this demonstration, we present a system-on-chip (SoC) designed to support scalable CNN acceleration at a low cost. The SoC features a cost-effective chip-to-chip adapter that enables scalable performance improvements with minimal design costs. This adapter manages chip-to-chip synchronization and data scheduling, effectively reducing chip-to-chip latency overhead. The SoC is implemented using the Silterra 130-nm CMOS technology with a chip size of 6.3×6.3 mm2.
Jicheon Kim, Chunmyung Park, Eunjae Hyun, Xuan Truong Nguyen
ISCAS2
2025 NPC: A Non-Conflicting Processing-in-Memory Controller in DDR Memory Systems
abstract
Processing-in-Memory (PIM) has emerged as a promising solution to address the memory wall problem. Existing memory interfaces must support new PIM commands to utilize PIM, making the definition of PIM commands according to memory modes a major issue in the development of practical PIM products. For performance and OS-transparency, the memory controller is responsible for changing the memory mode, which requires modifying the controller and resolving conflicts with existing functionalities. Additionally, it must operate to minimize mode transition overhead, which can cause significant performance degradation. In this study, we present NPC, a memory controller designed for mode transition PIM that delivers PIM commands via the DDR interface. NPC issues PIM commands while transparently changing the memory mode with a dedicated scheduling policy that reduces the number of mode transitions with aggregative issuing. Moreover, existing functions, such as refresh, are optimized for PIM operation. We implement NPC in hardware and develop a PIM emulation system to validate it on FPGA platforms. Experimental results reveal that NPC is compatible with existing interfaces and functionality, and the proposed scheduling policy improves performance by 2.2$\boldsymbol{\times}$with balanced fairness, achieving up to 97% of the ideal performance. These findings have the potential to aid the application of PIM in real systems and contribute to the commercialization of mode transition PIM.
Seungyong Lee 0003, Chunmyung Park, Woojae Shin, Hyun Kim 0001
IEEE Trans. Computers4
2024 A Scalable Multi-Chip YOLO Accelerator With a Lightweight Inter-Chip Adapter
abstract
Multi-chip-module (MCM) technology offers a promising solution for designing large-scale deep-learning inference systems while concurrently minimizing fabrication and design costs. Nevertheless, compared to monolithic dies, MCMs often incur resource and performance overhead due to interchip communication. Addressing these challenges, this paper introduces a scalable MCM-based DNN accelerator that incorporates a lightweight chip-to-chip adapter (C2CA) and an effective multi-chip dataflow. Inspired by the on-chip bus architecture, the C2CA efficiently shares data and address channels, achieving nearly optimal throughput with significantly reduced pin count and hardware costs. Additionally, the proposed design adopts a layer-wise dataflow within a ring-based C2C topology to fully utilize the constrained C2C bandwidth, mitigating both performance and communication overhead. When implemented on the Xilinx ZCU104 FPGA board, the system demonstrates significant throughput improvements compared to a single-chip configuration, yielding 1.92x and 3.57x enhancements for 2-chip and 4-chip configurations, respectively, on the YOLOv3-Tiny.
Jicheon Kim, Chunmyung Park, Eunjae Hyun, Xuan Truong Nguyen
ISCAS2
2023 Live Demonstration: Layer-wise Configurable CNN Accelerator with High PE Utilization
abstract
We demonstrate two end-to-end frameworks, ShortcutFusion [1] and ShortcutFusion++ [2], that effectively map many well-known deep neural networks, such as YOLO-v3, MobileNet-v2, EfficientNet-B0, and Resnet-50, to a generic CNN accelerator on FPGA. The experimental results show that ShortcutFusion++ achieves a processing element utilization of 80.95% for the well-known object detector YOLO-v3.
Chunmyung Park, Eunjae Hyun, Jicheon Kim, Xuan Truong Nguyen
ISCAS1