EDBT 2026 Demo / reviewers in the wild / expert
Taehoon Kim 0001
dblp:49/6795-1
· DBLP profile ↗
15ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-1819-968XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating LLMs using an Efficient GEMM Library and Target-Aware Optimizations on Real-World PIM DevicesabstractReal-time processing of deep learning models on conventional systems, such as CPUs and GPUs, is highly challenging due to memory bottlenecks. This is exacerbated in Large Language Models (LLMs), where the majority of executions are dominated by General Matrix Multiplication (GEMM) operations, which are relatively more memory-intensive than convolution operations. Processing-in-Memory (PIM), which provides high internal bandwidth, can be a promising alternative for LLM serving. However, since current PIM systems do not fully replace traditional memory, data transfer between the host and PIM-side memory is essential. Therefore, minimizing the transfer cost between the host and PIM is crucial for serving LLMs efficiently on the PIM. In this paper, we propose PIM-LLM, an end-to-end framework that accelerates LLMs using an efficient tiled GEMM library and several key target-aware optimizations on real-world PIM systems. We first propose PGEMMlib, which provides optimized tiling techniques for PIM, considering architecture specific characteristics to minimize unnecessary data transfer overhead and maximize parallelism. In addition, Tile-Selector explores optimized parameters and techniques for different GEMM shapes and available resources of PIM systems using an analytical model. To accelerate LLMs using PGEMMlib, we integrate it into the TVM deep learning compiler framework. We further optimize the LLM execution by applying several key optimizations: Build-time memory layout adjustment, PIM resource pooling, CPU/PIM cooperation support, and QKV generation fusion. Evaluation shows that PIM-LLM achieves significant performance gains of up to 45.75x over the TVM baseline for several well-known LLMs. We strongly believe that this work provides key insights for efficient LLM serving on real PIM devices. Hyeoncheol Kim, Taehoon Kim 0001, Taehyeong Park 0001, Donghyeon Kim 0001, Yongseung Yu, Hanjun Kim 0001, Yongjun Park 0001 |
CGO | 2 |
| 2025 | PIM-CARE: A Compiler-Assisted Dynamic Resource Allocation Framework for Real-world DRAM PIMabstractProcessing-In-Memory (PIM) has recently emerged as a promising solution to alleviate the memory bottleneck by integrating computing capabilities into memory chips. Since PIM provides numerous Processing Elements (PEs) and high-bandwidth on-chip data transfers, full utilization of the PEs becomes a critical mission to maximize the performance of PIM applications. However, due to the diverse and complex characteristics of PIM applications, using more resources does not always improve performance. It is therefore important to find the suitable amount of resources to achieve the best performance and to fully utilize the PIM resources.To address this, we introduce PIM-CARE, a framework for dynamic resource allocation across multiple applications with compiler support on real-world PIM systems. PIM-CARE first determines the best amount of PIM resources to allocate for each application. To enable spatial multitasking, the PIM-CARE daemon monitors resource allocation and deallocation requests and estimates total PIM resource utilization at runtime. It then dynamically schedules applications using a priority-based out-of-order policy, considering both available PIM resources and resource requirements for best performance. Evaluation on real-world PIM systems shows that PIM-CARE improves throughput by 5.49x and average turnaround time by 5.71x compared to the baseline. Inyong Hwang, Donghyeon Kim 0001, Seokwon Kang, Taehyeong Park 0001, Taehoon Kim 0001, Jiwon Seo 0002, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001 |
ICS | 5 |
| 2025 | Disaggregated Memory for File-backed PagesabstractTo explore the opportunity of expanding the page cache using disaggregated memory for file-backed pages, this study presents BalloonStasher, an RDMA-based disaggregated memory for data-intensive applications. Utilizing the ephemeral nature of the page cache, BalloonStasher dynamically adapts to the changing page cache demands of multiple clients. BalloonStasher supports one-sided RDMA-based memory pooling or two-sided RDMA-based memory sharing when using memory nodes. Additionally, it also supports peer memory mode, which utilizes the idle memory of peer nodes. Our extensive performance study compares the benefits and limitations of the three cache modes, and shows that BalloonStasher can mitigate the memory underutilization problem and improve the performance of data-intensive applications by a large margin. Daegyu Han, Jaeyoon Nam, Hokeun Cha, Changdae Kim 0001, Kwangwon Koh, Taehoon Kim 0001, Sang-Hoon Kim, Beomseok Nam |
ACM Trans. Storage | 6 |
| 2024 | Hardware-hardened Sandbox Enclaves for Trusted Serverless ComputingabstractIn cloud-based serverless computing, an application consists of multiple functions provided by mutually distrusting parties. For secure serverless computing, the hardware-based trusted execution environment (TEE) can provide strong isolation among functions. However, not only protecting each function from the host OS and other functions, but also protecting the host system from the functions, is critical for the security of the cloud servers. Such an emerging trusted serverless computing poses new challenges: Each TEE must be isolated from the host system bi-directionally, and the system calls from it must be validated. In addition, the resource utilization of each TEE must be accountable in a mutually trusted way. However, the current TEE model cannot efficiently represent such trusted serverless applications. To overcome the lack of such hardware support, this article proposes an extended TEE model called Cloister , designed for trusted serverless computing. Cloister proposes four new key techniques. First, it extends the hardware-based memory isolation in SGX to confine a deployed function only within its TEE (enclave). Second, it proposes a trusted monitor enclave that filters and validates system calls from enclaves. Third, it provides a trusted resource accounting mechanism for enclaves that is agreeable to both service developers and cloud providers. Finally, Cloister accelerates enclave loading by redesigning its memory verification for fast function deployment. Using an emulated Intel SGX platform with the proposed extensions, this article shows that trusted serverless applications can be effectively supported with small changes in the SGX hardware. Joongun Park, Seunghyo Kang, Taehoon Kim 0001, Jongse Park, Youngjin Kwon, Jaehyuk Huh 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | Virtual PIM: Resource-Aware Dynamic DPU Allocation and Workload Scheduling Framework for Multi-DPU PIM ArchitectureabstractProcessing-in-Memory (PIM) is an attractive device that can effectively satisfy the rapidly increasing demands for memory-intensive workloads in emerging application domains, such as deep learning and big data processing. Thanks to the integrated design of the main memory (MRAM) and multiple data processing units (DPUs) on a single chip, the PIM devices can provide massive parallelism from numerous DPUs and the substantial bandwidth between the MRAM and DPUs, thus achieving the high performance for the memory-intensive workloads. However, although the recent PIM architectures, including UPMEM, can efficiently execute a single memory-intensive application, they fail to efficiently orchestrate multiple applications on the multiple DPU resources due to the conservative resource allocation, without a resource monitoring system, and large scheduling granularity. To solve these problems, we propose a novel resource-aware dynamic DPU allocation and workload scheduling framework, called Virtual PIM, for multi-DPU PIM architectures such as UPMEM. The framework initially virtualizes the DPU and MRAM to ensure data consistency in multi-application environments. For dynamic DPU allocation, the Virtual PIM framework continuously gathers resource requests from multiple processes and current DPU occupancy information to estimate the dynamic DPU resource status, irrespective of PIM hardware support. Based on this information, the framework dynamically allocates DPUs and schedules workloads in fine-grained levels with minimum occupancy to maximize total DPU utilization. Our evaluations in real PIM environments demonstrate that Virtual PIM significantly improves system throughput and average normalized turnaround time by up to 4.83x and 3.45x, respectively, compared to the SLURM-based baseline. Donghyeon Kim 0001, Taehoon Kim 0001, Inyong Hwang, Taehyeong Park 0001, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001 |
PACT | 2 |
| 2023 | DEHype: Retrofitting Hypervisors for a Resource-Disaggregated EnvironmentabstractResource disaggregation has been proposed as a solution for resource under-utilization in data centers. However, host virtualization technologies, which are the basic building blocks for constructing data centers, are implemented without considering the disaggregated resources. In addition, we discover that a RDMA I/O unit plays a significant role in the performance of a disaggregated resource environment. In this study, we propose DEHype, which alleviates the inefficiency of the hypervisors utilized in a disaggregated environment by investigating host virtualization technologies that are suitable for disaggregated memory systems. Specifically, DEHype aims to identify and improve the performance issues associated with virtual machines through KVM/QEMU in a disaggregated resource environment. The results demonstrate the effectiveness of the proposed optimizations in improving the performance of disaggregated memory systems. DEHype achieves up to a 351% improvement over the state-of-the-art disaggregated memory system. Taehoon Kim 0001, Kwangwon Koh, Changdae Kim 0001, Eunji Pak, Yeonjeong Jeong, Sang-Hoon Kim |
CLUSTER | 1 |
| 2023 | Block Group Scheduling: A General Precision-scalable NPU Scheduling Technique with Capacity-aware Memory AllocationabstractPrecision-scalable neural processing units (PSNPUs) efficiently provide native support for quantized neural networks. However, with the recent advancements of deep neural networks, PSNPUs are affected by a severe memory bottleneck owing to the need to perform an extreme number of simple computations simultaneously. In this study, we first analyze whether the memory bottleneck issue can be solved using conventional neural processing unit scheduling techniques. Subsequently, we introduce new capacity-aware memory allocation and block-level scheduling techniques to minimize the memory bottleneck. Compared with the baseline, the new method achieves up to 2.26× performance improvements by substantially relieving the memory pressure of low-precision computations without hardware overhead. Seokho Lee, Younghyun Lee, Hyejun Kim, Taehoon Kim 0001, Yongjun Park 0001 |
DATE | 4 |
| 2022 | BWA-MEM-SCALE: Accelerating Genome Sequence Mapping on Commodity ServersabstractAs advances in Next-Generation Sequencing have made genome sequence data generation faster and cheaper, the acceleration of genome sequence mapping to the reference genome becomes an increasingly important problem. Much effort has been made to improve the performance of the sequence mapping process. Changdae Kim 0001, Kwangwon Koh, Taehoon Kim 0001, Daegyu Han, Jiwon Seo 0002 |
ICPP | 3 |
| 2020 | Nested Enclave: Supporting Fine-grained Hierarchical Isolation with SGXabstractAlthough hardware-based trusted execution environments (TEEs) have evolved to provide strong isolation with efficient hardware supports, their current monolithic model poses challenges in representing common software structures with modules produced from potentially untrusted 3rd parties. For better mapping of such modular software designs to trusted execution environments, it is necessary to extend the current monolithic model to a hierarchical one, which provides multiple inner TEEs within a TEE. For such hierarchical compartmentalization within a TEE, this paper proposes a novel hierarchical TEE called nested enclave, which extends the enclave support from Intel SGX. Inspired by the multi-level security model, nested enclave provides multiple inner enclaves sharing the same outer enclave. Inner enclaves can access the context of the outer enclave, but they are protected from the outer enclave and non-enclave execution. Peer inner enclaves are isolated from each other while accessing the execution environment of the shared outer enclave. Both of the inner and outer enclaves are protected from vulnerable privileged software and physical attacks. Such fine-grained nested enclaves allow secure multitiered environments using software modules from untrusted 3rd parties. The security-sensitive modules run on the inner enclave with the higher security level, while the 3rd party modules on the outer enclave. It can be further extended to provide a separate inner module for each user to process privacy-sensitive data while sharing the same library with efficient hardwareprotected communication channels. This study investigates three case scenarios implemented with an emulated nested enclave support, proving the feasibility and security improvement of the nested enclave model. Joongun Park, Naegyeong Kang, Taehoon Kim 0001, Youngjin Kwon, Jaehyuk Huh 0001 |
ISCA | 3 |
| 2019 | Heterogeneous Isolated Execution for Commodity GPUsabstractTraditional CPUs and cloud systems based on them have embraced the hardware-based trusted execution environments to securely isolate computation from malicious OS or hardware attacks. However, GPUs and their cloud deployments have yet to include such support for hardware-based trusted computing. As large amounts of sensitive data are offloaded to GPU acceleration in cloud environments, ensuring the security of the data is a current and pressing need. As deployed today, the outsourced GPU model is vulnerable to attacks from compromised privileged software. To support isolated remote execution on GPUs even under vulnerable operating systems, this paper proposes a novel hardware and software architecture, called HIX (Heterogeneous Isolated eXecution). HIX does not require modifications to the GPU architecture to offer protections: Instead, it offers security by modifying the I/O interconnect between the CPU and GPU, and by refactoring the GPU device driver to work from within the CPU trusted environment. A result of the architectural choices behind HIX is that the concept can be applied to other offload accelerators besides GPUs. This work implements the proposed HIX architecture on an emulated machine with KVM and QEMU. Experimental results from the emulated security support with a real GPU show that the performance overhead for security is curtailed to 26% on average for the Rodinia benchmark, while providing secure isolated GPU computing. Insu Jang, Taehoon Kim 0001, Simha Sethumadhavan, Jaehyuk Huh 0001 |
ASPLOS | 3 |
| 2019 | ShieldStore: Shielded In-memory Key-value Storage with SGXabstractThe shielded computation of hardware-based trusted execution environments such as Intel Software Guard Extensions (SGX) can provide secure cloud computing on remote systems under untrusted privileged system software. However, hardware overheads for securing protected memory restrict its capacity to a modest size of several tens of megabytes, and more demands for protected memory beyond the limit cause costly demand paging. Although one of the widely used applications benefiting from the enhanced security of SGX, is the in-memory key-value store, its memory requirements are far larger than the protected memory limit. Furthermore, the main data structures commonly use fine-grained data items such as pointers and keys, which do not match well with the coarse-grained paging of the SGX memory extension technique. To overcome the memory restriction, this paper proposes a new in-memory key-value store designed for SGX with application-specific data security management. The proposed key-value store, called ShieldStore, maintains the main data structures in unprotected memory with each key-value pair individually encrypted and integrity-protected by its secure component running inside an enclave. Based on the enclave protection by SGX, ShieldStore provides secure data operations much more efficiently than the baseline SGX key-value store, achieving 8--11 times higher throughput with 1 thread, and 24--30 times higher throughput with 4 threads. Taehoon Kim 0001, Joongun Park, Jae-Woo Chang, Seungheun Jeon, Jaehyuk Huh 0001 |
EuroSys | 1 |
| 2018 | Secure In-memory Key-Value Storage with SGXabstractNo abstract available. Taehoon Kim 0001, Joongun Park, Jae-Woo Chang, Seungheun Jeon, Jaehyuk Huh 0001 |
SoCC | 1 |
| 2017 | Transparent Dual Memory Compression ArchitectureabstractThe increasing memory requirements of big data applications have been driving the precipitous growth of memory capacity in server systems. To maximize the efficiency of external memory, HW-based memory compression techniques have been proposed to increase effective memory capacity. Although such memory compression techniques can improve the memory efficiency significantly, a critical trade-off exists in the HW-based compression techniques. As the memory blocks need to be decompressed as quickly as possible to serve cache misses, latency-optimized techniques apply compression at the cacheline granularity, achieving the decompression latency of less than a few cycles. However, such latency-optimized techniques can lose the potential high compression ratios of capacity-optimized techniques, which compress larger memory blocks with longer latency algorithms.Considering the fundamental trade-off in the memory compression, this paper proposes a transparent dual memory compression (DMC) architecture, which selectively uses two compression algorithms with distinct latency and compression characteristics. Exploiting the locality of memory accesses, the proposed architecture compresses less frequently accessed blocks with a capacity-optimized compression algorithm, while keeping recently accessed blocks compressed with a latency-optimized one. Furthermore, instead of relying on the support from the virtual memory system to locate compressed memory blocks, the study advocates a HW-based translation between the uncompressed address space and compressed physical space. This OS-transparent approach eliminates conflicts between compression efficiency and large page support adopted to reduce TLB misses. The proposed compression architecture is applied to the Hybrid Memory Cube (HMC) with a logic layer under the stacked DRAMs. The experimental results show that the proposed compression architecture provides 54% higher compression ratio than the state-of-the-art latency-optimized technique, with no performance degradation over the baseline system without compression. Seikwon Kim, Seonyoung Lee, Taehoon Kim 0001, Jaehyuk Huh 0001 |
PACT | 3 |
| 2016 | Dynamic prefetcher reconfiguration for diverse memory architecturesabstractWith the advent of stacked memory and new memory architectures, the heterogeneity of memory has been increasing. In the diverse memory technologies, each memory architecture has its own advantages and weaknesses. Considering the trade-offs, future systems are expected to support multiple memory architectures with a hybrid memory system. However, such diversity of memory architectures complicates the performance optimization of on-chip memory hierarchy. One of the key components affected by this trend is the hardware prefetcher. The available memory bandwidth highly affects the effectiveness of prefetchers, and the aggressiveness of prefetchers must be tuned for memory architectures as well as application behaviors. This paper investigates the effect of memory diversity on the prefetcher parameter selection, and proposes a dynamic parameter search mechanism to adjust the prefetch aggressiveness under various memory architectures. Using a general hill climbing scheme periodically, the mechanism adapts to the memory architectures and application behaviors effectively. In addition to such automatic tuning, the study improves the solution for cache pollution exacerbated by the increase of speculative data from more aggressive prefetchers in higher bandwidth memory. With the dynamic parameter search and pollution mitigation, the proposed framework improves the performance of applications by 12.4% on average compared to the prior scheme for tuning prefetch parameters. Taehoon Kim 0001, Jaehyuk Huh 0001 |
ICCD | 2 |
| 2016 | Reducing the Memory Bandwidth Overheads of Hardware Security Support for Multi-Core ProcessorsabstractTo prevent physical attacks on systems, secure processors have been proposed to reduce trusted computing base to the processor itself. In a secure processor, all off-chip data are encrypted and their integrity is protected. This paper investigates how the limited memory bandwidth of multi-core processors affects the design of secure processors. Although the performance of a single-core secure processor has improved significantly with the counter-mode encryption combined with Bonsai Merkle Tree, our results indicate that multi-core secure processors can suffer from significant performance degradation due to the limited memory bandwidth. To mitigate the performance overheads, this paper proposes three techniques for the multi-core design of secure processors. First, the paper advocates to use a combined cache for all normal and security-supporting data. Second, the paper proposes memory scheduling and mapping schemes for secure processors. Finally, the paper investigates a type-aware cache insertion scheme considering the distinct characteristics of normal and security-supporting data. Our simulation results show that the combined techniques reduce the performance degradation for supporting full confidentiality and integrity, from 25-34 percent to less than 8-14 percent in 8-core and 16-core secure processors, with minimal extra hardware costs. Taehoon Kim 0001, Jaehyuk Huh 0001 |
IEEE Trans. Computers | 2 |