Taehyeong Park 0001

dblp:292/5806-1 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Flow-Graph-Aware Tiling and Rescheduling for Memory-Efficient On-Device Inference
abstract
With the increasing popularity of artificial intelligence (AI) applications, deep neural networks (DNNs) are in demand for on-device serving in various real-life fields. Running DNN inference on a resource-constrained edge device requires aggressive memory optimization. While several recent tiling-based techniques reduce peak memory usage by partitioning large tensors into micro-tensors, they are specialized for the MCU environment and do not provide scalability to various edge platforms. Moreover, they greedily search for the target of tiling without considering the memory flow across the model while partitioning. In this paper, we propose OKO, a compiler-based optimization technique that minimizes peak memory usage by considering both tiling and the corresponding operation rescheduling. OKO estimates the memory savings from the tiling method based on the lifetime and dependencies of the tensor, and then algorithmically selects the optimal tiling strategy. It further maximizes the reuse of memory spaces by efficiently reordering operations and immediately releasing unnecessary tensors. Evaluations on various edge devices show that OKO achieves effective memory savings of up to 80% and an average of 59% with no loss of accuracy and negligible overhead, supporting memory-efficient inference across a broad range of target devices.
Yeonoh Jeong, Taehyeong Park 0001, Yongjun Park 0001
CGO2
2026 Efficient Data Processing using On-the-Fly Host-PIM Interactions in a Commodity PIM System
Hyojune Kim, Jeonghyeon Joo, Taehyeong Park 0001, Yongjun Park, Hyuck Han, Sooyong Kang
ICDE3
2025 Accelerating LLMs using an Efficient GEMM Library and Target-Aware Optimizations on Real-World PIM Devices
abstract
Real-time processing of deep learning models on conventional systems, such as CPUs and GPUs, is highly challenging due to memory bottlenecks. This is exacerbated in Large Language Models (LLMs), where the majority of executions are dominated by General Matrix Multiplication (GEMM) operations, which are relatively more memory-intensive than convolution operations. Processing-in-Memory (PIM), which provides high internal bandwidth, can be a promising alternative for LLM serving. However, since current PIM systems do not fully replace traditional memory, data transfer between the host and PIM-side memory is essential. Therefore, minimizing the transfer cost between the host and PIM is crucial for serving LLMs efficiently on the PIM. In this paper, we propose PIM-LLM, an end-to-end framework that accelerates LLMs using an efficient tiled GEMM library and several key target-aware optimizations on real-world PIM systems. We first propose PGEMMlib, which provides optimized tiling techniques for PIM, considering architecture specific characteristics to minimize unnecessary data transfer overhead and maximize parallelism. In addition, Tile-Selector explores optimized parameters and techniques for different GEMM shapes and available resources of PIM systems using an analytical model. To accelerate LLMs using PGEMMlib, we integrate it into the TVM deep learning compiler framework. We further optimize the LLM execution by applying several key optimizations: Build-time memory layout adjustment, PIM resource pooling, CPU/PIM cooperation support, and QKV generation fusion. Evaluation shows that PIM-LLM achieves significant performance gains of up to 45.75x over the TVM baseline for several well-known LLMs. We strongly believe that this work provides key insights for efficient LLM serving on real PIM devices.
Hyeoncheol Kim, Taehoon Kim 0001, Taehyeong Park 0001, Donghyeon Kim 0001, Yongseung Yu, Hanjun Kim 0001, Yongjun Park 0001
CGO3
2025 An Efficient PIM-Based Graph Engine on a Single Machine
abstract
With the increasing size of real-world networks, efficient analysis of large-scale graphs has become an important research area. To this end, we can consider Processing-in-Memory (PIM), which integrates processing units and main memory into a single chip, as a promising solution. Many studies have focused on enabling highly efficient processing of memory-intensive tasks by using PIM's high internal bandwidth. To the best of our knowledge, however, there have been no studies related to the scenarios where the entire graph does not fit in main memory and data movement across storage, memory, and cache should be considered. Motivated by this, we propose RealGraph PIM, a new PIM-based graph engine, that processes large-scale real-world graphs efficiently on top of the original RealGraph, a state-of-the-art CPU-based graph engine. RealGraph PIM employs (1) asynchronous I/O to reduce wasting time in an idle state and (2) column-wise partitioning to reduce CPU workloads, thereby issuing I/O requests more frequently. Experimental results on real-world datasets show that RealGraph PIM outperforms dramatically state-of-the-art graph engines including a naive version of RealGraphPIM.
Myung-Hwan Jang, Min-Kyeong Shin, Taehyeong Park 0001, Yongjun Park 0001, Sang-Wook Kim
CIKM3
2025 PIM-CARE: A Compiler-Assisted Dynamic Resource Allocation Framework for Real-world DRAM PIM
abstract
Processing-In-Memory (PIM) has recently emerged as a promising solution to alleviate the memory bottleneck by integrating computing capabilities into memory chips. Since PIM provides numerous Processing Elements (PEs) and high-bandwidth on-chip data transfers, full utilization of the PEs becomes a critical mission to maximize the performance of PIM applications. However, due to the diverse and complex characteristics of PIM applications, using more resources does not always improve performance. It is therefore important to find the suitable amount of resources to achieve the best performance and to fully utilize the PIM resources.To address this, we introduce PIM-CARE, a framework for dynamic resource allocation across multiple applications with compiler support on real-world PIM systems. PIM-CARE first determines the best amount of PIM resources to allocate for each application. To enable spatial multitasking, the PIM-CARE daemon monitors resource allocation and deallocation requests and estimates total PIM resource utilization at runtime. It then dynamically schedules applications using a priority-based out-of-order policy, considering both available PIM resources and resource requirements for best performance. Evaluation on real-world PIM systems shows that PIM-CARE improves throughput by 5.49x and average turnaround time by 5.71x compared to the baseline.
Inyong Hwang, Donghyeon Kim 0001, Seokwon Kang, Taehyeong Park 0001, Taehoon Kim 0001, Jiwon Seo 0002, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001
ICS4
2025 SortingHat: System Topology-aware Scheduling of Deep Neural Network Models on Multi-GPU Systems
abstract
The advent of cutting-edge AI applications has emphasized the importance of reducing inference latency.Consequently, efficient model-parallel execution on multiple GPUs represents a key challenge in achieving high performance through the partitioning of the target neural network.Nevertheless, in recent complex deep learning models, as the size of parameters continues to increase and overall inference latency is no longer solely dominated by kernel execution, performance improvements using multiple GPUs cannot be achieved by simply exploiting model parallelism without considering data transfer parallelism and the system topology.To address this challenge, this paper proposes SortingHat, which generates an efficient schedule of target neural network models on multi-GPU systems to minimize inference latency.Initially, SortingHat partitions a target model into multiple submodels based on dominator analysis to find the best solution within a reasonable time.Subsequently, SortingHat finds the best schedule for each submodel using Mixed Integer Linear Programming, taking system topology into account to exploit both model parallelism and data transfer parallelism.Once the schedules of all submodels are found, they are merged and executed on the ready queue-based executor.Evaluations on diverse multi-GPU environments with various large language models show that SortingHat achieves an average speedup of 2.28× and up to 2.96× over the single GPU on TVM baseline.
Seok Namkoong, Taehyeong Park 0001, Kiung Jung, Yongjun Park 0001
ICS2
2023 Virtual PIM: Resource-Aware Dynamic DPU Allocation and Workload Scheduling Framework for Multi-DPU PIM Architecture
abstract
Processing-in-Memory (PIM) is an attractive device that can effectively satisfy the rapidly increasing demands for memory-intensive workloads in emerging application domains, such as deep learning and big data processing. Thanks to the integrated design of the main memory (MRAM) and multiple data processing units (DPUs) on a single chip, the PIM devices can provide massive parallelism from numerous DPUs and the substantial bandwidth between the MRAM and DPUs, thus achieving the high performance for the memory-intensive workloads. However, although the recent PIM architectures, including UPMEM, can efficiently execute a single memory-intensive application, they fail to efficiently orchestrate multiple applications on the multiple DPU resources due to the conservative resource allocation, without a resource monitoring system, and large scheduling granularity. To solve these problems, we propose a novel resource-aware dynamic DPU allocation and workload scheduling framework, called Virtual PIM, for multi-DPU PIM architectures such as UPMEM. The framework initially virtualizes the DPU and MRAM to ensure data consistency in multi-application environments. For dynamic DPU allocation, the Virtual PIM framework continuously gathers resource requests from multiple processes and current DPU occupancy information to estimate the dynamic DPU resource status, irrespective of PIM hardware support. Based on this information, the framework dynamically allocates DPUs and schedules workloads in fine-grained levels with minimum occupancy to maximize total DPU utilization. Our evaluations in real PIM environments demonstrate that Virtual PIM significantly improves system throughput and average normalized turnaround time by up to 4.83x and 3.45x, respectively, compared to the SLURM-based baseline.
Donghyeon Kim 0001, Taehoon Kim 0001, Inyong Hwang, Taehyeong Park 0001, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001
PACT4
2023 Orchestrating Large-Scale SpGEMMs using Dynamic Block Distribution and Data Transfer Minimization on Heterogeneous Systems
abstract
Sparse general matrix-matrix multiplication (SpGEMM) is a major kernel in various emerging applications, such as database management systems, deep learning, graph analysis, and recommendation systems. Since SpGEMM requires extensive computation, many SpGEMM techniques have been implemented based on graphics processing units (GPUs) to exploit massive data parallelism completely. However, traditional SpGEMM techniques usually do not fully utilize the GPU because most non-zero elements of the target sparse matrices exist in a few hub nodes, and non-hub nodes barely have non-zero elements. The data-related characteristics (power law) result in a significant degradation in performance because of the load imbalance between the GPU cores and the low utilization of each core. Many attempts have been made through recent implementations to solve this problem using smart pre-/post-processing. However, the net performance hardly improves and sometimes even deteriorates owing to the large overheads. Additionally, non-hub nodes are inherently not suitable for GPU computing, even after optimization. Furthermore, the performance is no longer dominated by kernel execution, but by data transfers such as device-to-host (D2H) data transfers and file I/Os, owing to the rapid growth in the computing power of GPUs and input data size.Therefore, this work proposes a Dynamic Block Distributor (DBD), a novel full-system-level SpGEMM orchestration framework for heterogeneous systems, improving the overall performance by enabling an efficient CPU-GPU collaboration and further minimizing the overhead in data transfer between all the system elements. This framework first divides the target matrix into smaller blocks and then offloads the computation of each block to an appropriate computing unit between a GPU and CPU based on its workload type and the status of resource utilization at runtime. It also minimizes the overhead in data transfer with simple but suitable techniques, such as Row Collecting, I/O Overlapping, and I/O Binding. Our experiments showed that this framework increased the execution latency of SpGEMM, which included both the kernel execution and D2H transfers, by 3.24x on average, and the overall execution time by 2.07x on average, compared to that of the baseline cuSPARSE library.
Taehyeong Park 0001, Seokwon Kang, Myung-Hwan Jang, Sang-Wook Kim, Yongjun Park 0001
ICDE1