Jiwoo Bang

dblp:243/6214 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2023
0000-0002-3556-2535ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2023 Towards Enhanced I/O Performance of NVM File Systems
abstract
Non-volatile memory (NVM) provides bulk storage capacity, like NAND flash, while providing low latency, like DRAM, at the same time. NVM enables high-performance, reliable, and cost-effective high performance systems by providing low-latency data access and high capacity storage compared to traditional disk-based system. As NVM becomes a novel tier in the memory hierarchy, efficiently utilizing NVM I/O capability is important. In this work, we evaluate the I/O performance of NVM in three aspects: the performance change with a varying number of concurrent accesses, the performance dif-ference between remote and local accesses, and the performance change with various access granularity. We also compare the performance of NVM file systems that handle the different I/O characteristics of NVM. Specifically, Odinfs is the state-of-the-art NVM file system that solves the performance degradation of NVM with large number of threads and remote NUMA node accesses. We further optimize Odinfs by solving the I/O performance degradation with a small number of threads. We evaluate the optimized version of Odinfs and show that the throughput of Odinfs is increased by 30.91 % with four or fewer threads.
Jiwoo Bang, Chungyong Kim, Eun-Kyu Byun, Hanul Sung, Jaehwan Lee 0001, Hyeonsang Eom
HiPC1
2023 Accelerating I/O performance of ZFS-based Lustre file system in HPC environment
Jiwoo Bang, Chungyong Kim, Eun-Kyu Byun, Hanul Sung, Jaehwan Lee 0001, Hyeonsang Eom
J. Supercomput.1
2021 An In-Depth I/O Pattern Analysis in HPC Systems
abstract
High-performance computing (HPC) systems consist of thousands of compute nodes, storage systems and high-speed networks, providing multiple layers of I/O stack with high complexity. By adjusting the diverse configuration settings that HPC systems provide, the I/O performance of applications can be improved. However, it is challenging to identify the optimal configuration settings without a thorough knowledge of the system, as each of the different I/O characteristics of applications can be an important factor for parameter decision. In this paper, we use multiple machine learning approaches to perform an in-depth analysis on I/O behaviors of HPC applications and to search for the optimal configuration settings for jobs sharing similar I/O characteristics. Improved by maximum 0.07 R-squared score, our results in overall show that jobs run on the HPC systems can obtain the predicted I/O performance for different configuration parameters with a high accuracy, using the proposed machine learning-based prediction models.
Jiwoo Bang, Chungyong Kim, Kesheng Wu, Alex Sim, Surendra Byna, Hanul Sung, Hyeonsang Eom
HiPC1
2021 MulConn: User-Transparent I/O Subsystem for High-Performance Parallel File Systems
abstract
Parallel file systems (PFS) are used to distribute data processing and establish shared access to large-scale data. Despite being able to provide high I/O bandwidth on each node, PFS has difficulty utilizing the I/O bandwidth due to a single connection between the client and server nodes. To mitigate the performance bottleneck, users increase the number of connections between the nodes by modifying PFS or applications. However, it is difficult to modify PFS itself due to its complicated internal structure. Thus, PFS users manually increase the number of connections between the nodes by employing several methods. In this paper, we propose a user-transparent I/O subsystem, MulConn, to make users exploit high I/O bandwidth between nodes. To avoid the modifications of PFS and user applications, we have developed a horizontal mount procedure and two I/O scheduling policies, TtoS and TtoM, in the virtual file system (VFS) layer. We expose a single mount point that has multiple connections by modifying the mount path of VFS from vertical hierarchy to horizontal hierarchy. We also introduce two I/O scheduling policies to distribute I/O requests evenly to multiple connections. The experimental results show that MulConn improves write and read performance by up to 2.6x and 2.8x, respectively, compared with those of PFS using the existing kernel. In addition, we provide the best I/O performance that PFS can provide in the given experimental environments.
Hwajung Kim, Jiwoo Bang, Dong Kyu Sung, Hyeonsang Eom, Heon Young Yeom, Hanul Sung
HiPC2
2021 Finer-LRU: A Scalable Page Management Scheme for HPC Manycore Architectures
abstract
In HPC systems, the increasing need for a higher level of concurrency has led to packing more cores within a single chip. However, since multiple processes share memory space, the frequent access to resources in critical sections where only atomic operation has to be executed can result in poor performance. In this paper, we focus on reducing lock contention on the memory management system of an HPC manycore architecture. One of the critical sections causing severe lock contention in the I/O path is in the page management system, which uses multiple Least Recently Used (LRU) lists with a single lock instance. To solve this problem, we propose a Finer-LRU scheme, which optimizes the page reclamation process by splitting LRU lists into multiple sub-lists, each having its own lock instance. Our evaluation result shows that the Finer-LRU scheme can improve sequential write throughput by 57.03% and reduce latency by 98.94% compared to the baseline Linux kernel version 5.2.8 in the Intel Knights Landing (KNL) architecture.
Jiwoo Bang, Chungyong Kim, Sunggon Kim, Qichen Chen, Cheongjun Lee, Eun-Kyu Byun, Jaehwan Lee 0001, Hyeonsang Eom
IPDPS1
2020 BBOS: Efficient HPC Storage Management via Burst Buffer Over-Subscription
abstract
To avoid access to PFS, dedicated BB allocation is preferred despite of severe BB underutilization. Recently, new all-flash HPC storage systems with integrated BB and PFS are proposed, which speed up access to PFS. For this reason, we adopt BB over-subscription allocation method by allowing HPC applications to use BB only for I/O phase for improving BB utilization. Unfortunately, BB over-subscription aggravates I/O interference and demotion overhead from BB to PFS, resulting in degraded performance. To minimize the performance degradation, we develop an I/O scheduler to prevent I/O congestion and a new transparent data management system based on checkpoint/restart characteristics of HPC applications. With the proposed approach, not only the BB utilization can be improved, but also high performance of applications is achieved. In our experiments, we find that BB utilization is improved at least 2.2×, and more stable and higher checkpoint performance is guaranteed compared to other approaches. Besides, we achieve up to 96.4% hit ratio of restart requests on BB and up to 3.1× higher restart performance than others.
Hanul Sung, Jiwoo Bang, Chungyong Kim, Hyung-Sin Kim, Alex Sim, Glenn K. Lockwood, Hyeonsang Eom
CCGRID2