Hong-Yeon Kim

dblp:91/1294 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-0452-6336ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021
YearPublicationVenuePosition
2026 BucketLSM: Breaking the Compaction Scalability Barrier in LSM-Based Key-Value Stores
Jaewan Park, Kyungwook Min, Sungjin Byeon, Taewan Noh, Hyungi Park, Xubin He, Hong-Yeon Kim, Youngjae Kim 0001
CCGrid7
2026 ColdMap: Compaction-Aware Cost-Benefit Zone Cleaning for ZNS-Based Key-Value Stores
Sungjin Byeon, Kyungwook Min, Jaewan Park, Hong-Yeon Kim, Junyoung Han, Jooyoung Hwang, Zhichao Cao 0002, Youngjae Kim 0001
ICS5
2025 Cost-Efficient VM Selection for Cloud-Based LLM Inference with KV Cache Offloading
abstract
LLM inference is essential for applications like text summarization, translation, and data analysis, but the high cost of GPU instances from Cloud Service Providers (CSPs) like AWS is a major burden. This paper proposes InferSave, a cost-efficient VM selection framework for cloud-based LLM inference. InferSave optimizes KV cache offloading based on Service Level Objectives (SLOs) and workload characteristics, estimating GPU memory needs, and recommending cost-effective VM instances. Additionally, the Compute Time Calibration Function (CTCF) improves instance selection accuracy by adjusting for discrepancies between theoretical and actual GPU performance. Experiments on AWS GPU instances show that selecting lower-cost instances without KV cache offloading improves cost efficiency by up to 73.7% for online workloads, while KV cache offloading saves up to 20.19% for offline workloads.
Hyunsun Chung, Myung-Hoon Cha, Hong-Yeon Kim, Youngjae Kim 0001
CLOUD5
2025 Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG Systems
abstract
Recent large language models (LLMs) face increasing inference latency as input context length and model size grow. Retrieval-augmented generation (RAG) exacerbates this by significantly increasing input tokens, leading to higher computational overhead during the prefill stage and prolonged time-to-first-token (TTFT). To address this, the paper proposes using a disk-based key-value (KV) cache to reduce the prefill computational burden, thereby shortening TTFT. We also introduce a disk-based shared KV cache management system, called Shared RAG-DCache, for multi-instance LLM RAG service environments. This system leverages the locality of documents related to user queries in RAG and queueing delays in LLM inference services to proactively generate and store disk KV caches for query-related documents, sharing them across multiple LLM instances to enhance inference performance. In experiments on a single host with 2 GPUs and 1 CPU, Shared RAG-DCache achieved a 15–71 % increase in throughput and up to a 12–65 % reduction in latency, depending on the resource configuration.
Hyungwoo Lee, Jungmin So, Myung-Hoon Cha, Hong-Yeon Kim, James J. Kim, Youngjae Kim 0001
CLOUD6
2025 MEMORYBRIDGE: Leveraging Cloud Resource Characteristics for Cost-Efficient Disk-Based GNN Training via Two-Level Architecture
abstract
Graph Neural Networks (GNNs) are machine learning models that process graph-structured data by learning relationships between vertices and edges, as well as graph-level characteristics. Recently, with the emergence of large graph datasets on a TB scale, dataset sizes have exceeded the memory capacity of single machines. As a result, traditional methods that load all graph data into memory have become unusable, leading to the emergence of disk-based GNN training that uses storage as a memory extension. Recent research has focused on reducing disk I/O bottlenecks in disk-based GNNs. However, disk-based GNNs face new challenges in cloud environments due to two main characteristics. First, compared to node-local environments, the significantly slower cloud storage I/O speed becomes the main bottleneck of the entire training process. Second, pre-defined virtual machines prevent users from freely utilizing desired memory sizes, bandwidth, and the latest GPU technologies. These limitations have made existing disk-based GNN research unusable in cloud environments. To overcome this, we propose MEMORYBRIDGE, a system that cost-effectively accelerates GNN training in cloud environments through a novel two-level architecture that utilizes affordable GPU resources as training nodes and remote memory resources without GPUs as memory nodes, instead of using a single expensive GPU resource. This architecture consists of two key components: (i) a mathematical solver that recommends the most cost-effective resource combination, and (ii) a cloud-specialized GNN framework that implements graphaware fixed caching and batch pipelining optimization. The experimental results show that MEMORYBRIDGE achieved a speed improvement of up to 32.7x compared to existing GNN training frameworks and a cost efficiency of 9.9x compared to alternative resource configuration strategies, effectively handling the unique problems that arise from the combination of cloud environments and GNN training.
Yoochan Kim, Weikuan Yu, Hong-Yeon Kim, Youngjae Kim 0001
CCGrid3
2024 DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
abstract
Distributed Deep Learning (DDL), as a paradigm, dictates the use of GPU-based clusters as the optimal infrastructure for training large-scale Deep Neural Networks (DNNs). However, the high cost of such resources makes them inaccessible to many users. Public cloud services, particularly Spot Virtual Machines (VMs), offer a cost-effective alternative, but their unpredictable availability poses a significant challenge to the crucial checkpointing process in DDL. To address this, we introduce DeepVM, a novel solution that recommends cost-effective cluster configurations by intelligently balancing the use of Spot and On-Demand VMs. DeepVM leverages a four-stage process that analyzes instance performance using the FLOPP (FLoating-point Operations Per Price) metric, performs architecture-level analysis with linear programming, and identifies the optimal configuration for the user-specific needs. Extensive simulations and real-world deployments in the AWS environment demonstrate that DeepVM consistently outperforms other policies, reducing training costs and overall makespan. By enabling cost-effective checkpointing with Spot VMs, DeepVM opens up DDL to a wider range of users and facilitates a more efficient training of complex DNNs.
Yoochan Kim, Yonghyeon Cho, Awais Khan 0002, Ki-Dong Kang, Baik-Song An, Myung-Hoon Cha, Hong-Yeon Kim, Youngjae Kim 0001
CCGrid9
2021 Fast and secure Global-Heap for memory-centric computing
Myung-Hoon Cha, Sangmin Lee 0003, Baik-Song An, Hong-Yeon Kim, Kang-Ho Kim
J. Supercomput.4
2019 Effective metadata management in exascale file system
Myung-Hoon Cha, Sangmin Lee 0003, Hong-Yeon Kim, Young-Kyun Kim
J. Supercomput.3
2019 Cost analysis of erasure coding for exa-scale storage
abstract
With the increasing demand for mass storage, research on exa-scale storage is actively underway. When the scale of storage grows to the exa-scale, the space efficiency becomes very important. To maintain the storage reliability and improve the space efficiency, we have begun to introduce erasure coding instead of replication. However, erasure coding has many I/O performance degradation factors such as Parity Calculation, degraded I/O, Data Distribution cost, etc., whereas the existing research mainly focuses on improving the performance of the Parity Calculation.In this study, we identified the issues and bottlenecks of using erasure coding in real storage. First, we measured the I/O performance of various erasure codes to find the suitable erasure codes for real storage. Next, we analyzed the execution time for each processing step when I/O was performed and the issues when erasure coding was used in storage. Finally, we predicted the cost of EC-based I/O processing in the exa-scale storage and identified the expected problems.
Dong-Oh Kim, Hong-Yeon Kim, Young-Kyun Kim, Jeong-Joon Kim
J. Supercomput.2
2018 APS: adaptable prefetching scheme to different running environments for concurrent read streams in distributed file systems
Sangmin Lee 0003, Soon J. Hyun, Hong-Yeon Kim, Young-Kyun Kim
J. Supercomput.3
2018 Fair bandwidth allocating and strip-aware prefetching for concurrent read streams and striped RAIDs in distributed file systems
Sangmin Lee 0003, Soon J. Hyun, Hong-Yeon Kim, Young-Kyun Kim
J. Supercomput.3
2017 Adaptive metadata rebalance in exascale file system
Myung-Hoon Cha, Dong-Oh Kim, Hong-Yeon Kim, Young-Kyun Kim
J. Supercomput.3
2007 Supporting Extended UNIX Remove Semantics in the OASIS Cluster Filesystem
Sangmin Lee 0003, Hong-Yeon Kim, Young-Kyun Kim, June Kim, Myoung-Joon Kim
ICCSA (1)2
2006 OASIS: Implementation of a Cluster File System Using Object-Based Storage Devices
Young-Kyun Kim, Hong-Yeon Kim, Sangmin Lee 0003, June Kim, Myoung-Joon Kim
ICCSA (1)2