Myoungwon Oh

dblp:150/5334 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0002-8132-4861ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 first-author · 3 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 NVMe-oF-R: Fast Recovery Design on Disaggregated Distributed Storage System
abstract
Failures in a large distributed storage system are often critical, leading to unexpected I/Os that are required to restore the system's health and ensure availability. With the advent of NVMe-oF, the disaggregation of compute and storage resources presents an opportunity to minimize the negative impact of the compute failure by reattaching the storage resources. However, despite advances in hardware, modern distributed storage systems have not yet fully adapted to the disaggregated architecture. There are four main reasons: (1) lack of awareness of recoverable failure events in the disaggregated architecture, (2) incorrect availability management with respect to the NVMe-oF fault domains, (3) unnecessary data rebalance I/Os for uniform distribution triggered even after the failure is recovered, (4) load imbalance caused by asymmetric deployment of compute resources after blind relocation for recovery. To address these challenges, we introduceNVMe-oF-R, a resilient disaggregated distributed storage architecture for fast recovery.NVMe-oF-Rcomprises three techniques: (1)NVMe-oF adapter, which detects recoverable failure events and orchestrates relocation; (2)DCRUSH, a data placement strategy that considers the NVMe-oF based disaggregation architecture; and (3)Relocater, which efficiently relocates failed compute resources and fixes stragglers that arise after recovery. We implementNVMe-oF-Ratop the storage orchestration layer in a CRUSH-based distributed storage system, Ceph. Our experimental results demonstrate thatNVMe-oF-Rcan eliminate unnecessary recovery traffic and reduce recovery time by more than 50%.
Myoungwon Oh, Cheolho Kang, Woojoong Kim, Yangwoo Roh, Jeong-Uk Kang, Silwan Chang
IEEE Trans. Parallel Distributed Syst.1
2023 TiDedup: A New Distributed Deduplication Architecture for Ceph
Myoungwon Oh, Samuel Just, Youngjin Yu, Duck-Ho Bae, Sage A. Weil, Sangyeun Cho, Heon Young Yeom
USENIX ATC1
2021 Re-architecting Distributed Block Storage System for Improving Random Write Performance
abstract
In cloud ecosystems, distributed block storage systems are used to provide a persistent block storage service, which is the fundamental building block for operating cloud native services. However, existing distributed storage systems performed poorly for random write workloads in an all-NVMe storage configuration, becoming CPU-bottlenecked. Our roofline-based approach to performance analysis on a conventional distributed block storage system with NVMe SSDs reveals that the bottleneck does not lie in one specific software module, but across the entire software stack; (1) tightly coupled I/O processing, (2) inefficient threading architecture, and (3) local backend data store causing excessive CPU usage. To this end, we re-architect a modern distributed block storage system for improving random write performance. The key ingredients of our system are (1) decoupled operation processing using non-volatile memory, (2) prioritized thread control, and (3) CPU-efficient backend data store. Our system emphasizes low CPU overhead and high CPU efficiency to efficiently utilize NVMe SSDs in a distributed storage environment. We implement our system in Ceph. Compared to the native Ceph, our prototype system delivers more than 3x performance improvement for small random write I/Os in terms of both IOPS and latency by efficiently utilizing CPU cores.
Myoungwon Oh, Jiwoong Park, Sung Kyu Park, Adel Choi, Jongyoul Lee, Jin-Hyeok Choi, Heon Young Yeom
ICDCS1
2018 Design of Global Data Deduplication for a Scale-Out Distributed Storage System
abstract
Scale-out distributed storage systems can uphold balanced data growth in terms of capacity and performance on an on-demand basis. However, it is a challenge to store and manage large sets of contents being generated by the explosion of data. One of the promising solutions to mitigate big data issues is data deduplication, which removes redundant data across many nodes of the storage system. Nevertheless, it is non-trivial to apply a conventional deduplication design to the scale-out storage due to the following root causes. First, chunk-lookup for deduplication is not as scalable and extendable as the underlying storage system supports. Second, managing the metadata associated to deduplication requires a huge amount of design and implementation modifications of the existing distributed storage system. Lastly, the data processing and additional I/O traffic imposed by deduplication can significantly degrade performance of the scale-out storage. To address these challenges, we propose a new deduplication method, which is highly scalable and compatible with the existing scale-out storage. Specifically, our deduplication method employs a double hashing algorithm that leverages hashes used by the underlying scale-out storage, which addresses the limits of current fingerprint hashing. In addition, our design integrates the meta-information of file system and deduplication into a single object, and it controls the deduplication ratio at online by being aware of system demands based on post-processing. We implemented the proposed deduplication method on an open source scale-out storage. The experimental results show that our design can save more than 90% of the total amount of storage space, under the execution of diverse standard storage workloads, while offering the same or similar performance, compared to the conventional scale-out storage.
Myoungwon Oh, Jungyeon Yoon, Sangjae Kim, Kang-Won Lee 0002, Sage A. Weil, Heon Young Yeom, Myoungsoo Jung
ICDCS1
2018 LALCA: Locality-Aware Lock Contention Avoidance for NVMe-Based Scale-out Storage System
abstract
Flash-based NVMe storage devices dramatically improve I/O latency as well as I/O throughput. However, existing scale-out storage systems are designed for targeting hard disk drives and this design limitation causes significant performance degradation when they are used with NVMe devices. In this paper, we analyzed the performance of an existing scale-out storage system and identified lack of data locality and excessive lock contentions. To mitigate these problems, we present a new design based on locality aware lock contention avoidance (LALCA). LALCA proposes two techniques for scale-out storage system on high speed storage device: (1) locality-aware thread control to minimize processor context switching overhead and relevant performance degradation and (2) lock contention avoidance to remove locking problems in existing scale-out distributed storage system. With evaluation, LALCA shows not only up to 9 times performance improvement but also up to half CPU usage reduction when servicing small random I/Os.
Myoungwon Oh, Jugwan Eom, Seungmin Kim, Sangjae Kim, Kang-Won Lee 0002, Heon Young Yeom
IPDPS1
2016 Performance Optimization for All Flash Scale-Out Storage
abstract
The proliferation of the big data analysis and the wide spread usage of public/private cloud services make it important to expand the storage capacity as the demand is increased. The scale-out storage is gaining more attention since it can inherently provide scalable storage capacity. The flash SSD, on the other hand, is getting popular as the drop-in replacement of the slow HDD, which seems to boost the system performance somewhat at least. However, the performance of traditional scale-out storage system does not get much better even though its HDD is replaced with the flash based high performance SSD since the whole system is designed based on HDD as its underlying storage device. In this paper, we identify performance problems of a representative scale-out storage system, Ceph, and analyze that these problems are caused by 1) Coarse-grained lock, 2) Throttling logic, 3) Batching based operation latency and 4) Transaction Overhead. We propose some optimization techniques for flash-based Ceph. First, we minimize coarse-grained locking. Second, we introduce throttle policy and system tuning. Third, we develop non-blocking logging and light-weight transaction processing. We found that our optimized Ceph shows up to 20 times improvement in the case of small random writes and it also shows more than two times better performance in the case of small random read through our experiments. We also show that the system exhibits linear performance increase as we add more nodes.
Myoungwon Oh, Jugwan Eom, Jungyeon Yoon, Jae Yeun Yun, Seungmin Kim, Heon Young Yeom
CLUSTER1
2014 Enhancing the I/O system for virtual machines using high performance SSDs
abstract
Storage I/O in VM (Virtual Machine) environments, which requires low latency, becomes problematic as the fast storage such as SSDs (Solid-State Drives) is currently in use. The low performance problem in the VM environment is caused by 1) the presence of additional software layer such as guest OS, 2) context switching between VM and host OS, and 3) scheduling delay for I/O process. These factors do not cause serious problems in the case of using HDD which leads to high latency batching. However, there will be significant performance degradation when fast storage devices are used. To address this problem, we have proposed the following methods to improve the performance of I/O stack in the VM environments by attempting to optimize the I/O stack: one is pipelined polling, and the other is multiple issues and multiple completions. We have found via experiments that our approach leads to increases in the performance of SSDs in a VM environment by up to 50% when multiple VM storage devices are used, and that it leads to improvements in the performance by more than 80% when a single VM storage device is used, with the CPU utilization reduced by up to 25%.
Myoungwon Oh, Hyeonsang Eom, Heon Young Yeom
IPCCC1
2014 OS I/O Path Optimizations for Flash Solid-state Drives
Woong Shin, Qichen Chen, Myoungwon Oh, Hyeonsang Eom, Heon Young Yeom
USENIX ATC3