Yongseok Son

dblp:137/0873 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-4512-0121ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ScaleSwap: A Scalable OS Swap System for All-Flash Swap Arrays
Taehwan Ahn, Chanhyeong Yu, Sangjin Lee 0003, Yongseok Son
FAST4
2026 AURORA-Q: Asynchronous Unified Resource Optimizer for Quantum Simulation on HPC System
Changjong Kim, Alex Sim, Kesheng Wu, Houjun Tang, Yongseok Son, Jisung Park 0001, Sunggon Kim
ICDCS5
2026 Design and implementation of a fast and predictable SSD liveness watchdog for storage systems
Jinyong Ha 0001, Yongseok Son
Future Gener. Comput. Syst.2
2026 Design and implementation of zoned namespace-tailored log-structured file system for commodity ZNS SSDs
Inhwi Hwang, Sangjin Lee 0003, Sunggon Kim, Hyeonsang Eom, Yongseok Son
Future Gener. Comput. Syst.5
2025 ScaleLFS: A Log-Structured File System with Scalable Garbage Collection for Commodity SSDs
Jinyong Ha 0001, Sangjin Lee 0003, Hyeonsang Eom, Yongseok Son
FAST4
2025 A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
abstract
In modern server CPUs, the Last-Level Cache (LLC) serves not only as a victim cache for higher-level private caches but also as a buffer for low-latency DMA transfers between CPU cores and I/O devices through Direct Cache Access (DCA).However, prior work has shown that high-bandwidth network-I/O devices can rapidly flood the LLC with packets, often causing significant contention with co-running workloads.One step further, this work explores hidden microarchitectural properties of the Intel Xeon CPUs, uncovering two previously unrecognized LLC contentions triggered by emerging high-bandwidth I/O devices.Specifically, (C1) DMAwritten cache lines in LLC ways designated for DCA (referred to as DCA ways) are migrated to certain LLC ways (denoted as inclusive ways) when accessed by CPU cores, unexpectedly contending with non-I/O cache lines within the inclusive ways.In addition, (C2) high-bandwidth storage-I/O devices, which are increasingly common in datacenter servers, benefit little from DCA while contending with (latency-sensitive) network-I/O devices within DCA ways.To this end, we present A4, a runtime LLC management framework designed to alleviate both (C1) and (C2) among diverse co-running workloads, using a hidden knob and other hardware features implemented in those CPUs.Additionally, we demonstrate that A4 can also alleviate other previously known network-I/Odriven LLC contentions.Overall, it improves the performance of latency-sensitive, high-priority workloads by 51% without notably compromising that of low-priority workloads.
Haneul Park, Jiaqi Lou, Sangjin Lee 0003, KyoungSoo Park, Yongseok Son, Ipoom Jeong, Nam Sung Kim
ISCA6
2025 Z-LFS: A Zoned Namespace-tailored Log-structured File System for Commodity Small-zone ZNS SSDs
Inhwi Hwang, Sangjin Lee 0003, Sunggon Kim, Hyeonsang Eom, Yongseok Son
USENIX ATC5
2025 Design and implementation of a decentralized document management system
Jongbeen Han, Yongseok Son
Expert Syst. Appl.2
2025 zCeph: Design and implementation of a ZNS-friendly distributed file system
Jinyong Ha 0001, Yongseok Son
Future Gener. Comput. Syst.2
2025 AS2: Adaptive sorting algorithm selection for heterogeneous workloads and systems
SangMyung Lee, Byungyoon Lee, Yongseok Son, Kiwook Sohn, Hwajung Kim, Sunggon Kim
Future Gener. Comput. Syst.3
2025 Regen: An object layout regenerator on large-scale production HPC systems
Dong Kyu Sung, Sunggon Kim, Sangjin Lee 0003, Houjun Tang, Alex Sim, Kesheng Wu, Surendra Byna, Yongseok Son
Future Gener. Comput. Syst.8
2024 EnC-IoT: An Efficient Encryption and Access Control Framework based on IPFS for Decentralized IoT
abstract
Recently, many decentralized IoT systems incorporating IPFS and blockchain have been widely used to share private data and overcome issues in a centralized IoT system such as a single point of failure and loss of data sovereignty. However, the constrained computing power of IoT devices can result in a performance overhead during encryption processes, which involves computationally intensive tasks. In addition, the blockchain involved in the access control mechanism can raise another performance issue and potentially expose sensitive information due to its transparent nature. To address these issues, in this paper, we propose an efficient IoT framework (EnC-IoT) to enhance performance and data protection (i.e., security and privacy) in a decentralized IoT environment. We first introduce an efficient data encryption scheme that shuffles the data order to enhance the security level. Then, it partitions the data into multiple regions and encrypts/decrypts each region in parallel to accelerate the cryptography process. Second, we provide an efficient access control scheme which adopts a tree-based authentication to perform the access control quickly. Moreover, it employs a proxy re-encryption method which allows the decryption of sensitive information without exposing the owner’s private key to improve both security and privacy levels. We implement EnC-IoT with two schemes in IPFS based on a decentralized IoT environment. The experimental results show that EnC-IoT provides a performance improvement of up to 100.7× for uploading and 9.6× for downloading compared with the existing systems.
Mansub Song, Sunggon Kim, Hyeonsang Eom, Yongseok Son
CCGrid5
2024 ScaleCache: A Scalable Page Cache for Multiple Solid-State Drives
abstract
This paper presents a scalable page cache called ScaleCache for improving SSD scalability. Specifically, we first propose a concurrent data structure of page cache based on XArray (ccXArray) to enable access and update the page cache concurrently. Second, we introduce a direct page flush (dflush) which directly flushes pages to storage devices in a parallel and opportunistic manner. We implement ScaleCache with two techniques in the Linux kernel and evaluate it on a 64-core machine with eight NVMe SSDs. Our evaluations show that ScaleCache improves the performance of Linux file systems by up to 6.81× and 4.50× compared with the existing scheme and scalable scheme for multiple SSDs, respectively.
Kiet Tuan Pham, Seokjoo Cho, Sangjin Lee 0003, Lan Anh Nguyen, Hyeongi Yeo, Ipoom Jeong, Sungjin Lee 0001, Nam Sung Kim, Yongseok Son
EuroSys9
2024 ScaleDFS: Accelerating Decentralized and Private File Sharing via Scaling Directed Acyclic Graph Processing
abstract
This paper introduces a novel file system, ScaleDFS, designed to accelerate decentralized file sharing for a private network, leveraging the potential of scaling file management based on directed acyclic graph (DAG) on modern hardware. Specifically, in ScaleDFS, we first design a DAG builder that parallelizes the construction of DAG nodes for a file while preserving critical orders to speed up the uploading process. Second, we introduce a DAG reader that retrieves leaf DAG nodes in parallel without graph traversal assisted by a devised DAG cache to accelerate the downloading process. Finally, we present a DAG remover that rapidly identifies obsolete DAG nodes/data and removes them in parallel to mitigate the garbage collection overhead without compromising consistency. We implement ScaleDFS based on IPFS and demonstrate that ScaleDFS outperforms IPFS by up to 3.7×, 1.8×, and 12.6× in realistic file, private blockchain, and gateway workloads, respectively.
Mansub Song, Lan Anh Nguyen, Sunggon Kim, Hyeonsang Eom, Yongseok Son
HPDC5
2024 A2FL: Autonomous and Adaptive File Layout in HPC through Real-time Access Pattern Analysis
abstract
Various scientific applications with different I/O characteristics are executed in HPC systems. However, underlying parallel file systems are unaware of these characteristics of applications, and using a single fixed file layout for all applications can degrade the performance of HPC systems. In this paper, we propose A2FL, an autonomous and adaptive file layout adjustment scheme that optimizes parallel file system configurations by analyzing the access pattern of the applications. The key steps of A2FL are as follows: (1) A2FL initially intercepts the I/O operations of the application, recording their access patterns in real-time. (2) The access patterns are then transformed into a graphical representation used for predicting I/O performance and providing adjustment recommendations. (3) A2FL autonomously adjusts the file layout based on the prediction results, delivering an optimal file layout within the parallel file system. Moreover, we propose A2FL-Compound which analyzes an access pattern by dividing it into smaller components to optimize the file layout in a fine-grained manner. Our evaluations demonstrate that A2FL significantly enhances I/O performance, with improvements of up to 65.9× compared to the default file layout.
Dong Kyu Sung, Yongseok Son, Alex Sim, Kesheng Wu, Surendra Byna, Houjun Tang, Hyeonsang Eom, Changjong Kim, Sunggon Kim
IPDPS2
2024 RL-Watchdog: A Fast and Predictable SSD Liveness Watchdog on Storage Systems
Jinyong Ha 0001, Sangjin Lee 0003, Heon Young Yeom, Yongseok Son
USENIX ATC4
2024 Empowering Cyberattack Identification in IoHT Networks With Neighborhood-Component-Based Improvised Long Short-Term Memory
abstract
Cybersecurity has become an inevitable concern in the healthcare industry due to the rapid growth of the Internet of Health Things (IoHT). The IoHT is revolutionizing healthcare by enabling remote access to hospital equipment, real-time patient monitoring, and urgent alerts to patients and hospitals. However, the convenience of these systems also makes them vulnerable to cyberattacks, with hackers seeking to disrupt health services or extort money through ransomware attacks. Efficiently detecting multiple threats is a challenging task because IoHT generates large temporal data and system log information. In this paper, we propose time series classification models for the identification of potential cyberattacks in IoHT networks. First, we introduce Neighborhood Component Analysis (NCA) with modifications of the regularization parameter to select the vital input features. With the selected features, we propose two LSTM-based models: Directed Acyclic Graph-based Long Short-Term Memory (DAG-LSTM) and Projected Layer-based Long Short-Term Memory (PL-LSTM) for detecting cyberattacks. We evaluate the existing time series classification models (i.e., GRU, LSTM, and Bi-LSTM) and proposed models (i.e., DAG-LSTM and PL-LSTM) using real-world IoHT data. We also validate the models by applying a non-parametric statistical test, Friedman test. Our evaluation results show that the proposed DAG-LSTM achieves the highest accuracy with 99.89% training and 92.04% an average testing accuracy.
Manish Kumar 0009, Changjong Kim, Yongseok Son, Sushil Kumar Singh 0001, Sunggon Kim
IEEE Internet Things J.3
2022 IFLustre: Towards Interference-Free and Efficient Storage Allocation in Distributed File System
abstract
Distributed file systems (DFSs) are widely used in large scale computing environments such as cloud computing and high performance computing systems (HPC) where thousands of applications are executed simultaneously. To support applications that produce and process a large amount of data, DFSs manage limited storage resources in the system and allocate the resources to the applications. To efficiently utilize limited storage resources, it is important to consider the I/O characteristics of applications as well as the effect of resource sharing among multiple applications in the system. In this paper, we first perform empirical performance analysis and investigate the effect of I/O interference caused by storage resource allocation. Based on our empirical performance analysis, we propose IFLustre towards interference-free and efficient storage allocation in Lustre file system. IFLustre first utilizes previous execution logs to determine the necessary storage resources for the application and predict the throughput and runtime of the application. Then, it allocates storage resources via file system configurations based on the prediction results and the DFS allocation status. This allows IFLustre to allocate sufficient resources to the application while mitigating interference from resource sharing. The experimental results show that IFLustre can improve the performance by up to 47% compared with Lustre.
Sunggon Kim, Dong Kyu Sung, Yongseok Son
MASCOTS3
2021 Improving I/O performance in distributed file systems for flash-based SSDs by access pattern reshaping
Sunggon Kim, Jaehyun Han, Hyeonsang Eom, Yongseok Son
Future Gener. Comput. Syst.4
2020 An Efficient Database Backup and Recovery Scheme using Write-Ahead Logging
abstract
Many cloud services perform periodic database backup to keep the data safe from failures such as sudden system crashes. In the database system, two techniques are widely used for data backup and recovery: a physical backup and a logical backup. The physical backup uses raw data by copying the files in the database, whereas the logical backup extracts data from the database and dumps it into separated files as a sequence of query statements. Both techniques support a full backup strategy that contains data of the entire database and incremental backup strategy that contains changed data since a previous backup. However, both strategies require additional I/O operations to perform the backup and need a long time to restore a backup. In this paper, we propose an efficient backup and recovery scheme by exploiting write-ahead logging (WAL) in database systems. In the proposed scheme, for backup, we devise a backup system to use log data generated by the existing WAL to eliminate the additional I/O operations. To restore a backup, we utilize and optimize the existing crash recovery procedure of WAL to reduce recovery time. For example, we divide the recovery range and applying the backup data for each range independently via multiple threads. We implement our scheme in MySQL, a popular database management system. The experimental result demonstrates that the proposed scheme provides instant backup while reducing recovery time compared with the existing schemes.
Hwajung Kim, Heon Young Yeom, Yongseok Son
CLOUD3
2020 FlexGPU: A Flexible and Efficient Scheduler for GPU Sharing Systems
abstract
The graphics processing unit (GPU) is extensively used in diverse domains, such as finance, machine learning, and image processing. The GPU can be underutilized as multiple applications may not share the same GPU concurrently owing to a memory oversubscription issue. For example, when applications that require fewer computational resources but a larger GPU memory are running instantaneously, the GPU memory may be insufficient; consequently, the number of GPU applications running simultaneously is restricted, decreasing GPU utilization. Further, it can even stop the execution of applications that are running on the GPU. To this end, we propose FlexGPU, which schedules the kernels of the GPU applications that run on the same GPU according to their features. This framework 1) schedules the kernel at the launching time according to its features to improve GPU utilization and 2) temporarily checkpoints and restores non-dependent content in the GPU memory to/from the host memory, which avoids oversubscription of the GPU when out-of-memory failure occurs and allows more kernels to run concurrently on the GPU. The experimental results show that compared to existing methods, our approach demonstrates a 7 times improvement in performance in terms of execution time and enables a 2.5 times increase in the concurrent execution of applications.
Qichen Chen, Heon Young Yeom, Yongseok Son
CCGRID4
2020 Towards HPC I/O Performance Prediction through Large-scale Log Analysis
abstract
Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units, while used by hundreds to thousands of users at the same time. Applications from these large numbers of users have diverse characteristics, such as varying compute, communication, memory, and I/O intensiveness. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, the I/O performance is difficult to predict because the I/O system software is complex, the I/O system is shared among all users, and the I/O operations also heavily rely on networking systems. To improve the prediction of the I/O performance on HPC systems, we propose to integrate information from a number of different system logs and develop a regression-based approach that dynamically selects the most relevant features from the most recent log entries, and automatically select the best regression algorithm for the prediction task. Evaluation results show that our proposed scheme can predict the I/O performance with up to 84% prediction accuracy in the case of the I/O-intensive applications using the logs from CORI supercomputer at NERSC.
Sunggon Kim, Alex Sim, Kesheng Wu, Surendra Byna, Yongseok Son, Hyeonsang Eom
HPDC5
2020 EdgeIso: Effective Performance Isolation for Edge Devices
abstract
Edges enable cloud services to be provided at low-latency and efficiently reduce the amount of transferred data by placing latency-critical tasks close to users. However, multi-tasking results in resource contention on edge devices, making it challenging to meet the service level objectives (SLOs) of tasks. Compared to the clouds, edges have relatively limited resources, but their tasks are required to meet a higher level of SLOs than clouds. Furthermore, modern edge devices equipped with additional accelerators (e.g., GPU) may worsen the resource contention due to the edge's integrated architecture, sharing the memory bandwidth between CPUs and accelerators. To address these challenges, we present EdgeIso, a light-weight scheduler that dynamically isolates the performance of tasks on edges. EdgeIso periodically monitors the resource contention and mitigates the contention to meet the SLOs of tasks by efficiently enforcing several isolation techniques (e.g., DVFS and core allocation) in an incremental manner. Moreover, it detects the changes of task executions or offered loads for tasks, thus handling high load fluctuations adaptively. We implement EdgeIso as a user-level scheduler on the Linux integrates into an NVIDIA Jetson TX2. Our experimental results show that EdgeIso improves the performance of the low-latency tasks significantly while improving resource efficiency compared with both the offloading and reservation scheme used in clouds.
Yoonsung Nam, Yongjun Choi, Byeonghun Yoo, Hyeonsang Eom, Yongseok Son
IPDPS5
2020 Towards Hybrid Isolation for Shared Multicore Systems
Yoonsung Nam, Byeonghun Yoo, Yongjun Choi, Yongseok Son, Hyeonsang Eom
JSSPP4
2020 Page Reusability-Based Cache Partitioning for Multi-Core Systems
abstract
Most modern multi-core processors provide a shared last level cache (LLC) where data from all cores are placed to improve performance. However, this opens a new challenge for cache management, owing to cache pollution. With cache pollution, data with weak temporal locality can evict other data with strong temporal locality when both are mapped into the same cache set. In this article, we propose page reusability-based cache partitioning (PRCP) for multi-core systems to maximize cache utilization by minimizing cache pollution. To achieve this, PRCP divides pages into two groups: (1) highly-reused pages and (2) lowly-reused pages. The reusability of each page is collected online via periodic page table scans. PRCP then dynamically partitions the shared cache into two corresponding areas using page coloring technique. We have implemented PRCP in Linux kernel and evaluated it using SPEC CPU2006 benchmarks. The results show that our scheme can achieve comparable performance to the optimal offline MRC-guided process-based cache partitioning scheme without a priori knowledge of workloads.
Jiwoong Park, Heon Young Yeom, Yongseok Son
IEEE Trans. Computers3
2020 Design and Implementation of SSD-Assisted Backup and Recovery for Database Systems
abstract
As flash-based solid-state drive (SSD) becomes more prevalent because of the rapid fall in price and the significant increase in capacity, customers expect better data services than traditional disk-based systems. However, the order of magnitude performance provided and new characteristics of flash require a rethinking of data services. For example, backup and recovery is an important service in a database system since it protects data against unexpected hardware and software failures. To provide backup and recovery, backup/recovery tools or backup/recovery methods by operating systems can be used. However, the tools perform time-consuming jobs, and the methods may negatively affect run-time performance during normal operation even though high-performance SSDs are used. To handle these issues, we propose an SSD-assisted backup/recovery scheme for database systems. Our scheme is to utilize the characteristics (e.g., out-of-place update) of flash-based SSD for backup/recovery operations. To this end, we exploit the resources (e.g., flash translation layer and DRAM cache with supercapacitors) inside SSD, and we call our SSD with new backup/ recovery functionality BR-SSD. We design and implement the functionality in the Samsung enterprise-class SSD (i.e., SM843Tn) for more realistic systems. Furthermore, we exploit and integrate BR-SSDs into database systems (i.e., MySQL) in replication and redundant array of independent disks (RAID) environments, as well as a database system in a single BR-SSD. The experimental result demonstrates that our scheme provides fast backup and recovery but does not negatively affect the run-time performance during normal operation.
Yongseok Son, Moonsub Kim, Sunggon Kim, Heon Young Yeom, Nam Sung Kim, Hyuck Han
IEEE Trans. Knowl. Data Eng.1
2019 z-READ: Towards Efficient and Transparent Zero-Copy Read
abstract
In cloud computing, I/O-intensive workloads can be co-located with other applications or virtual machines on a single physical machine. In this case, copy-based I/O (buffered I/O) can lead to severe performance interference to other memoryintensive workloads. It is because that the buffered I/O consumes memory bandwidth during memory copy even though it benefits from caching. To address this problem, many zero-copy I/O schemes have been proposed but none of them provides both 1) transparent copy avoidance through read/write system calls and 2) benefits of kernel-level caching at the same time. To this end, this paper presents z-READ, an efficient and transparent zero-copy read I/O scheme based on page remapping and copy-on-write techniques. In our scheme, we introduce several optimizations that minimize the overheads of page remapping by reducing the number of remote TLB shootdown.We implement z- READ prototype in memory management of Linux kernel 4.12.9. Our experimental results show that the performance of the colocated memory-intensive workloads can be negatively affected by I/O-intensive workloads in the case of copy-based I/O (up to 1.96x slowdown in-memory configurations) while z-READ incurs only up to 1.07x slowdown for the respective configuration.
Jiwoong Park, Cheolgi Min, Heon Young Yeom, Yongseok Son
CLOUD4
2019 DCA-IO: A Dynamic I/O Control Scheme for Parallel and Distributed File Systems
abstract
In high-performance computing, storage is a shared resource and used by all users with many different application requirements and knowledge of storage. Consequently, the optimal storage configuration varies according to the I/O behavior of each application. While system logs are helpful resources in understanding the storage behavior, it is non-trivial for each user to analyze the logs and adjust complex configurations. Even for experienced users, it is difficult to understand the full stack of I/O systems and find the optimal configuration for the specific application. In this work, we analyzed the I/O activities of CORI which is an HPC system in National Energy Research Scientific Computing Center (NERSC). The result of our analysis shows that most users do not adjust storage configurations and use the default settings. Also, it shows that only a few applications are executed repeatedly in the HPC environment. Based on this result, we have developed DCA-IO, a dynamic distributed file system configuration adjustment algorithm, which utilizes system log information and widely adapted rules to adjust storage configurations automatically without any user intervention. DCA-IO utilizes existing system logs and does not require any modifications in code or an additional library. To demonstrate the effectiveness of DCA-IO, we have performed experiments using I/O kernels of the real applications in both isolated small-sized Lustre environment and CORI. Our experimental result shows that the use of our scheme can lead to improvements in the performance of HPC applications by up to 75% in an isolated environment and 50% in a real HPC environment without user intervention.
Sunggon Kim, Alex Sim, Kesheng Wu, Surendra Byna, Yongseok Son, Hyeonsang Eom
CCGRID6
2019 IsoKV: An Isolation Scheme for Key-Value Stores by Exploiting Internal Parallelism in SSD
abstract
Modern data centers aim to take advantage of high parallelism in storage devices for I/O intensive applications such as storage servers, cache systems, and key-value stores. Key-value stores are the most typical applications that should provide a highly reliable service with high-performance. To increase the I/O performance of key-value stores, many data centers have actively adopted next-generation storage devices such as Non-Volatile Memory Express (NVMe) based Solid State Devices (SSDs). NVMe SSDs and its protocol are characterized to provide a high degree of parallelism. However, they may not guarantee predictable performance while providing high performance and parallelism. For example, heavily mixed read and write requests can result in performance degradation of throughput and response time due to the interference between the requests and internal operations (e.g., Garbage Collection (GC)). To minimize the interference and provide higher performance, this paper presents IsoKV, an isolation scheme for key-value stores by exploiting internal parallelism in SSDs. IsoKV manages the level of parallelism of SSD directly by running application-driven flash management scheme. By storing data with different characteristics in each dedicated internal parallel units of SSD, IsoKV reduces interference between I/O requests. Also, IsoKV synchronizes the LSM-tree logic and data management in SSD to eliminate GC. We implement IsoKV on RocksDB and evaluate it using Open-Channel SSD. Our extensive experiments have shown that IsoKV improves overall throughput and response time on average 1.20× and 43% compared with the existing scheme, respectively.
Heerak Lim, Hwajung Kim, Kihyeon Myung, Heon Young Yeom, Yongseok Son
HiPC5
2018 High-Performance Transaction Processing in Journaling File Systems
Yongseok Son, Sunggon Kim, Heon Young Yeom, Hyuck Han
FAST1
2017 SSD-Assisted Backup and Recovery for Database Systems
abstract
Backup and recovery is an important feature of database systems since it protects data against unexpected hardware and software failures. Database systems can provide data safety and reliability by creating a backup and restoring the backup from a failure. Database administrators can use backup/recovery tools that are provided with database systems or backup/recovery methods with operating systems. However, the existing tools perform time-consuming jobs and the existing methods may negatively affect run-time performance during normal operation even though high-performance SSDs are used. In this paper, we present an SSD-assisted backup/recovery scheme for database systems. In our scheme, we extend the out-of-place update characteristics of flash-based SSDs for backup/recovery operations. To this end, we exploit the resources (e.g., flash translation layer and DRAM cache with supercapacitors) inside SSDs, and we call our SSD with new backup/recovery features BR-SSD. We design and implement the backup/recovery functionality in the Samsung enterprise-class SSD (i.e., SM843Tn) for more realistic systems. Furthermore, we conduct a case study of BR-SSDs in replicated database systems and modify MySQL with replication to integrate BR-SSDs. The experimental result demonstrates that our scheme provides fast recovery while it does not negatively affect the run-time performance during normal operation.
Yongseok Son, Jaeyoon Choi, Jekyeom Jeon, Cheolgi Min, Sunggon Kim, Heon Young Yeom, Hyuck Han
ICDE1
2017 Optimizing I/O Operations in File Systems for Fast Storage Devices
abstract
Fast non-volatile memory (NVM) technologies (e.g., phase change memory, spin-transfer torque memory, and MRAM) provide high performance to legacy storage systems. These NVM technologies have attractive features, such as low latency and high throughput to satisfy application performance. Accordingly, fast storage devices based on fast NVM lead to a rapid increase in the demand for diverse computer systems and environments (e.g., cloud platforms, web servers, and database systems) where they are expected to be used as primary storage. Despite the promised benefits provided by fast storage devices, modern file systems do not take advantage of the storage's full performance. In this article, we analyze and explore existing I/O strategies in read, write, journal I/ O, and recovery paths between the file system and the storage device. The analysis shows that existing I/O strategies are an obstacle to get maximum performance of fast storage devices. To address this issue, we propose efficient I/O strategies that enable file systems to fully exploit the performance of fast storage devices. Our main idea is to transfer requests from discontiguous host memory buffers in the file systems to discontiguous storage segments in one I/O request to get maximize I/O performance. We implemented our scheme to read, write, journal I/O and recovery operations in the EXT4 file system and the JBD2 module. We demonstrate the implication of our idea in terms of application performance through well-known benchmarks. The experimental results show that our optimized file system achieves better performance than the existing file system, with improvements of up to 1.54 ×, 1.96×, and 2.28× on ordered mode, data journaling mode, and recovery, respectively.
Yongseok Son, Heon Young Yeom, Hyuck Han
IEEE Trans. Computers1
2016 An Empirical Evaluation of Enterprise and SATA-Based Transactional Solid-State Drives
abstract
In most file systems, performance is usually sacrificed in exchange for crash consistency, which ensures that data and metadata are restored consistently in the event of a system crash. To escape this trade-off between performance and crash consistency, recent researchers designed and implemented the transactional functionality inside Solid State Drives (SSDs). However, in order to investigate its benefit in a more realistic and standard fashion, this scheme should be re-evaluated in enterprise storage with standard interface. This paper explores the challenges and implications of a transactional SSD with extensive experiments. To evaluate the potential benefit of transactional SSD, we design and implement the transaction functionality in Samsung enterprise-class and SATA-based SSD (i.e., SM843TN) and name it TxSSD. We then modify the existing file systems (i.e., ext4 and btrfs) on topof TxSSD, making both file systems crash-consistent without redundant writes. We perform performance evaluation of two filesystems by using file I/O and OLTP benchmarks with a database. We also disclose and analyze the overhead of transactional functionality inside SSD. The experimental results show that TxSSD-aware file systems exhibit better performance compared to crash-consistent modes (i.e., data journaling mode of ext4 and cow mode of btrfs) but worse performance compared to weak consistent modes (i.e., ordered mode of ext4 and no datacow mode of btrfs).
Yongseok Son, Hara Kang, Jinyong Ha 0001, Jongsung Lee 0001, Hyuck Han, Hyungsoo Jung 0001, Heon Young Yeom
MASCOTS1
2016 Efficient Memory-Mapped I/O on Fast Storage Device
abstract
In modern operating systems, memory-mapped I/O ( mmio ) is an important access method that maps a file or file-like resource to a region of memory. The mapping allows applications to access data from files through memory semantics (i.e., load/store) and it provides ease of programming. The number of applications that use mmio are increasing because memory semantics can provide better performance than file semantics (i.e., read/write). As more data are located in the main memory, the performance of applications can be enhanced owing to the effect of a large cache. When mmio is used, hot data tend to reside in the main memory and cold data are located in storage devices such as HDD and SSD; data placement in the memory hierarchy depends on the virtual memory subsystem of the operating system. Generally, the performance of storage devices has a direct impact on the performance of mmio . It is widely expected that better storage devices will lead to better performance. However, the expectation is limited when fast storage devices are used since the virtual memory subsystem does not reflect the performance feature of those devices. In this article, we examine the Linux virtual memory subsystem and mmio path to determine the influence of fast storage on the existing Linux kernel. Throughout our investigation, we find that the overhead of the Linux virtual memory subsystem, negligible on the HDD, prevents applications from using the full performance of fast storage devices. To reduce the overheads and fully exploit the fast storage devices, we present several optimization techniques. We modify the Linux kernel to implement our optimization techniques and evaluate our prototyped system with low-latency storage devices. Experimental results show that our optimized mmio has up to 7x better performance than the original mmio . We also compare our system to a system that has enough memory to keep all data in the main memory. The system with insufficient memory and our mmio achieves 92% performance of the resource-rich system. This result implies that our virtual memory subsystem for mmap can effectively extend the main memory with fast storage devices.
Nae Young Song, Yongseok Son, Hyuck Han, Heon Young Yeom
ACM Trans. Storage2
2015 Optimizing file systems for fast storage devices
abstract
Emerging high-performance storage devices have attractive features such as low latency and high throughput. This leads to a rapid increase in the demand for fast storage devices in cloud platforms, social network services, etc. However, there are few block-based file systems that are capable of utilizing superior characteristics of fast storage devices. In this paper, we find that the I/O strategy of modern operating systems prevents file systems from exploiting fast storage devices. To address this problem, we propose several optimization techniques for block-based file systems. Then, we apply our techniques to two well-known file systems and evaluate them with multiple benchmarks. The experimental results show that our optimized file systems achieve 32% on average and up to 54% better performance than existing file systems.
Yongseok Son, Hyuck Han, Heon Young Yeom
SYSTOR1