EDBT 2026 Demo / reviewers in the wild / expert
Hyeonsang Eom
dblp:54/1476
· DBLP profile ↗
48ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-1902-6767ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design and implementation of zoned namespace-tailored log-structured file system for commodity ZNS SSDs
Inhwi Hwang, Sangjin Lee 0003, Sunggon Kim, Hyeonsang Eom, Yongseok Son |
Future Gener. Comput. Syst. | 4 |
| 2025 | HeteroScheduler: Dynamic Task Scheduling for CPU-GPU Optimization and Contention Mitigation in Cloud Data CentersabstractCloud data centers are increasingly integrating high-performance servers, shifting to heterogeneous environments where maintaining homogeneous infrastructures becomes inefficient. Thus, GPU-CPU heterogeneous cloud data centers have become essential. Optimal scheduling is essential to handle server performance differences and resource usage variations. Efficient resource utilization in heterogeneous clusters is a critical challenge for optimizing task execution performance and overall system efficiency. Traditional scheduling techniques primarily rely on offline static profiling, measuring task execution time in advance for each device, or classifying tasks based on CPU and memory utilization to colocate workloads with different characteristics, thereby mitigating resource contention and improving performance. However, these approaches fail to dynamically respond to fine-grained resource contention and environmental changes, leading to performance degradation and reduced resource efficiency. To address this issue, this study proposes a dynamic scheduling framework that utilizes real-time resource metrics. The proposed framework classifies tasks into Compute-intensive and Memory-intensive categories to determine optimal task placement in a heterogeneous cluster environment. Compute-intensive tasks require high IPC (Instructions Per Cycle) and primarily utilize CPU resources, while Memory-intensive tasks exhibit high LLC miss rates and consume substantial memory bandwidth. After defining task classification metrics through extensive experiments, the framework assigns tasks to either CPU or GPU servers based on their characteristics. Additionally, real-time resource monitoring detects contention and triggers task migration to optimize resource utilization, mitigate contention, and maximize performance. Experimental results demonstrate that the proposed method outperforms existing heterogeneous cluster scheduling techniques, achieving notable improvements in resource utilization, task execution time, and contention mitigation. Seokwon Choi, Hyeonsang Eom |
CLOUD | 2 |
| 2025 | ScaleLFS: A Log-Structured File System with Scalable Garbage Collection for Commodity SSDs
Jinyong Ha 0001, Sangjin Lee 0003, Hyeonsang Eom, Yongseok Son |
FAST | 3 |
| 2025 | Z-LFS: A Zoned Namespace-tailored Log-structured File System for Commodity Small-zone ZNS SSDs
Inhwi Hwang, Sangjin Lee 0003, Sunggon Kim, Hyeonsang Eom, Yongseok Son |
USENIX ATC | 4 |
| 2024 | EnC-IoT: An Efficient Encryption and Access Control Framework based on IPFS for Decentralized IoTabstractRecently, many decentralized IoT systems incorporating IPFS and blockchain have been widely used to share private data and overcome issues in a centralized IoT system such as a single point of failure and loss of data sovereignty. However, the constrained computing power of IoT devices can result in a performance overhead during encryption processes, which involves computationally intensive tasks. In addition, the blockchain involved in the access control mechanism can raise another performance issue and potentially expose sensitive information due to its transparent nature. To address these issues, in this paper, we propose an efficient IoT framework (EnC-IoT) to enhance performance and data protection (i.e., security and privacy) in a decentralized IoT environment. We first introduce an efficient data encryption scheme that shuffles the data order to enhance the security level. Then, it partitions the data into multiple regions and encrypts/decrypts each region in parallel to accelerate the cryptography process. Second, we provide an efficient access control scheme which adopts a tree-based authentication to perform the access control quickly. Moreover, it employs a proxy re-encryption method which allows the decryption of sensitive information without exposing the owner’s private key to improve both security and privacy levels. We implement EnC-IoT with two schemes in IPFS based on a decentralized IoT environment. The experimental results show that EnC-IoT provides a performance improvement of up to 100.7× for uploading and 9.6× for downloading compared with the existing systems. Mansub Song, Sunggon Kim, Hyeonsang Eom, Yongseok Son |
CCGrid | 4 |
| 2024 | ScaleDFS: Accelerating Decentralized and Private File Sharing via Scaling Directed Acyclic Graph ProcessingabstractThis paper introduces a novel file system, ScaleDFS, designed to accelerate decentralized file sharing for a private network, leveraging the potential of scaling file management based on directed acyclic graph (DAG) on modern hardware. Specifically, in ScaleDFS, we first design a DAG builder that parallelizes the construction of DAG nodes for a file while preserving critical orders to speed up the uploading process. Second, we introduce a DAG reader that retrieves leaf DAG nodes in parallel without graph traversal assisted by a devised DAG cache to accelerate the downloading process. Finally, we present a DAG remover that rapidly identifies obsolete DAG nodes/data and removes them in parallel to mitigate the garbage collection overhead without compromising consistency. We implement ScaleDFS based on IPFS and demonstrate that ScaleDFS outperforms IPFS by up to 3.7×, 1.8×, and 12.6× in realistic file, private blockchain, and gateway workloads, respectively. Mansub Song, Lan Anh Nguyen, Sunggon Kim, Hyeonsang Eom, Yongseok Son |
HPDC | 4 |
| 2024 | A2FL: Autonomous and Adaptive File Layout in HPC through Real-time Access Pattern AnalysisabstractVarious scientific applications with different I/O characteristics are executed in HPC systems. However, underlying parallel file systems are unaware of these characteristics of applications, and using a single fixed file layout for all applications can degrade the performance of HPC systems. In this paper, we propose A2FL, an autonomous and adaptive file layout adjustment scheme that optimizes parallel file system configurations by analyzing the access pattern of the applications. The key steps of A2FL are as follows: (1) A2FL initially intercepts the I/O operations of the application, recording their access patterns in real-time. (2) The access patterns are then transformed into a graphical representation used for predicting I/O performance and providing adjustment recommendations. (3) A2FL autonomously adjusts the file layout based on the prediction results, delivering an optimal file layout within the parallel file system. Moreover, we propose A2FL-Compound which analyzes an access pattern by dividing it into smaller components to optimize the file layout in a fine-grained manner. Our evaluations demonstrate that A2FL significantly enhances I/O performance, with improvements of up to 65.9× compared to the default file layout. Dong Kyu Sung, Yongseok Son, Alex Sim, Kesheng Wu, Surendra Byna, Houjun Tang, Hyeonsang Eom, Changjong Kim, Sunggon Kim |
IPDPS | 7 |
| 2024 | IoLens: Visual Analytics System for Exploring Storage I/O Tracking ProcessabstractAs we enter the era of big data, a substantial amount of Input/Output (I/O) requests to storage devices are generated, making the maintenance of I/O performance important. Furthermore, I/O performance directly affects the overall user experience in edge devices. However, with the increasing complexity of systems, numerous factors influencing I/O performance have emerged, making it challenging to analyze and explore the overall I/O processing workflow. To address this issue, we introduce IoLens, a visual analytics tool that helps users explore system I/O performance from kernel I/O stack up to virtual file system and storage device drivers. Our tool helps users analyze I/O performance by identifying the overall workload of I/O requests and intuitively identifying anomalies. The effectiveness and applicability of IoLens have been validated through a usage scenario following a system engineer working on system kernels. A user study with four domain experts is conducted to further validate the usability of the tool. Changmin Jeon, Jiwon Ha, Hyolim Hong, Hyeon Jeon, Hyeonsang Eom, Heonyoung Yeom, Jinwook Seo |
PacificVis | 5 |
| 2023 | Towards Enhanced I/O Performance of NVM File SystemsabstractNon-volatile memory (NVM) provides bulk storage capacity, like NAND flash, while providing low latency, like DRAM, at the same time. NVM enables high-performance, reliable, and cost-effective high performance systems by providing low-latency data access and high capacity storage compared to traditional disk-based system. As NVM becomes a novel tier in the memory hierarchy, efficiently utilizing NVM I/O capability is important. In this work, we evaluate the I/O performance of NVM in three aspects: the performance change with a varying number of concurrent accesses, the performance dif-ference between remote and local accesses, and the performance change with various access granularity. We also compare the performance of NVM file systems that handle the different I/O characteristics of NVM. Specifically, Odinfs is the state-of-the-art NVM file system that solves the performance degradation of NVM with large number of threads and remote NUMA node accesses. We further optimize Odinfs by solving the I/O performance degradation with a small number of threads. We evaluate the optimized version of Odinfs and show that the throughput of Odinfs is increased by 30.91 % with four or fewer threads. Jiwoo Bang, Chungyong Kim, Eun-Kyu Byun, Hanul Sung, Jaehwan Lee 0001, Hyeonsang Eom |
HiPC | 6 |
| 2023 | Accelerating I/O performance of ZFS-based Lustre file system in HPC environment
Jiwoo Bang, Chungyong Kim, Eun-Kyu Byun, Hanul Sung, Jaehwan Lee 0001, Hyeonsang Eom |
J. Supercomput. | 6 |
| 2021 | An In-Depth I/O Pattern Analysis in HPC SystemsabstractHigh-performance computing (HPC) systems consist of thousands of compute nodes, storage systems and high-speed networks, providing multiple layers of I/O stack with high complexity. By adjusting the diverse configuration settings that HPC systems provide, the I/O performance of applications can be improved. However, it is challenging to identify the optimal configuration settings without a thorough knowledge of the system, as each of the different I/O characteristics of applications can be an important factor for parameter decision. In this paper, we use multiple machine learning approaches to perform an in-depth analysis on I/O behaviors of HPC applications and to search for the optimal configuration settings for jobs sharing similar I/O characteristics. Improved by maximum 0.07 R-squared score, our results in overall show that jobs run on the HPC systems can obtain the predicted I/O performance for different configuration parameters with a high accuracy, using the proposed machine learning-based prediction models. Jiwoo Bang, Chungyong Kim, Kesheng Wu, Alex Sim, Surendra Byna, Hanul Sung, Hyeonsang Eom |
HiPC | 7 |
| 2021 | MulConn: User-Transparent I/O Subsystem for High-Performance Parallel File SystemsabstractParallel file systems (PFS) are used to distribute data processing and establish shared access to large-scale data. Despite being able to provide high I/O bandwidth on each node, PFS has difficulty utilizing the I/O bandwidth due to a single connection between the client and server nodes. To mitigate the performance bottleneck, users increase the number of connections between the nodes by modifying PFS or applications. However, it is difficult to modify PFS itself due to its complicated internal structure. Thus, PFS users manually increase the number of connections between the nodes by employing several methods. In this paper, we propose a user-transparent I/O subsystem, MulConn, to make users exploit high I/O bandwidth between nodes. To avoid the modifications of PFS and user applications, we have developed a horizontal mount procedure and two I/O scheduling policies, TtoS and TtoM, in the virtual file system (VFS) layer. We expose a single mount point that has multiple connections by modifying the mount path of VFS from vertical hierarchy to horizontal hierarchy. We also introduce two I/O scheduling policies to distribute I/O requests evenly to multiple connections. The experimental results show that MulConn improves write and read performance by up to 2.6x and 2.8x, respectively, compared with those of PFS using the existing kernel. In addition, we provide the best I/O performance that PFS can provide in the given experimental environments. Hwajung Kim, Jiwoo Bang, Dong Kyu Sung, Hyeonsang Eom, Heon Young Yeom, Hanul Sung |
HiPC | 4 |
| 2021 | Finer-LRU: A Scalable Page Management Scheme for HPC Manycore ArchitecturesabstractIn HPC systems, the increasing need for a higher level of concurrency has led to packing more cores within a single chip. However, since multiple processes share memory space, the frequent access to resources in critical sections where only atomic operation has to be executed can result in poor performance. In this paper, we focus on reducing lock contention on the memory management system of an HPC manycore architecture. One of the critical sections causing severe lock contention in the I/O path is in the page management system, which uses multiple Least Recently Used (LRU) lists with a single lock instance. To solve this problem, we propose a Finer-LRU scheme, which optimizes the page reclamation process by splitting LRU lists into multiple sub-lists, each having its own lock instance. Our evaluation result shows that the Finer-LRU scheme can improve sequential write throughput by 57.03% and reduce latency by 98.94% compared to the baseline Linux kernel version 5.2.8 in the Intel Knights Landing (KNL) architecture. Jiwoo Bang, Chungyong Kim, Sunggon Kim, Qichen Chen, Cheongjun Lee, Eun-Kyu Byun, Jaehwan Lee 0001, Hyeonsang Eom |
IPDPS | 8 |
| 2021 | Improving I/O performance in distributed file systems for flash-based SSDs by access pattern reshaping
Sunggon Kim, Jaehyun Han, Hyeonsang Eom, Yongseok Son |
Future Gener. Comput. Syst. | 3 |
| 2021 | An empirical study of I/O separation for burst buffers in HPC systems
Donghun Koo, Jaehwan Lee 0001, Jialin Liu 0002, Eun-Kyu Byun, Jae-Hyuck Kwak, Glenn K. Lockwood, Soonwook Hwang, Katie Antypas, Kesheng Wu, Hyeonsang Eom |
J. Parallel Distributed Comput. | 10 |
| 2020 | BBOS: Efficient HPC Storage Management via Burst Buffer Over-SubscriptionabstractTo avoid access to PFS, dedicated BB allocation is preferred despite of severe BB underutilization. Recently, new all-flash HPC storage systems with integrated BB and PFS are proposed, which speed up access to PFS. For this reason, we adopt BB over-subscription allocation method by allowing HPC applications to use BB only for I/O phase for improving BB utilization. Unfortunately, BB over-subscription aggravates I/O interference and demotion overhead from BB to PFS, resulting in degraded performance. To minimize the performance degradation, we develop an I/O scheduler to prevent I/O congestion and a new transparent data management system based on checkpoint/restart characteristics of HPC applications. With the proposed approach, not only the BB utilization can be improved, but also high performance of applications is achieved. In our experiments, we find that BB utilization is improved at least 2.2×, and more stable and higher checkpoint performance is guaranteed compared to other approaches. Besides, we achieve up to 96.4% hit ratio of restart requests on BB and up to 3.1× higher restart performance than others. Hanul Sung, Jiwoo Bang, Chungyong Kim, Hyung-Sin Kim, Alex Sim, Glenn K. Lockwood, Hyeonsang Eom |
CCGRID | 7 |
| 2020 | Towards HPC I/O Performance Prediction through Large-scale Log AnalysisabstractLarge-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units, while used by hundreds to thousands of users at the same time. Applications from these large numbers of users have diverse characteristics, such as varying compute, communication, memory, and I/O intensiveness. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, the I/O performance is difficult to predict because the I/O system software is complex, the I/O system is shared among all users, and the I/O operations also heavily rely on networking systems. To improve the prediction of the I/O performance on HPC systems, we propose to integrate information from a number of different system logs and develop a regression-based approach that dynamically selects the most relevant features from the most recent log entries, and automatically select the best regression algorithm for the prediction task. Evaluation results show that our proposed scheme can predict the I/O performance with up to 84% prediction accuracy in the case of the I/O-intensive applications using the logs from CORI supercomputer at NERSC. Sunggon Kim, Alex Sim, Kesheng Wu, Surendra Byna, Yongseok Son, Hyeonsang Eom |
HPDC | 6 |
| 2020 | EdgeIso: Effective Performance Isolation for Edge DevicesabstractEdges enable cloud services to be provided at low-latency and efficiently reduce the amount of transferred data by placing latency-critical tasks close to users. However, multi-tasking results in resource contention on edge devices, making it challenging to meet the service level objectives (SLOs) of tasks. Compared to the clouds, edges have relatively limited resources, but their tasks are required to meet a higher level of SLOs than clouds. Furthermore, modern edge devices equipped with additional accelerators (e.g., GPU) may worsen the resource contention due to the edge's integrated architecture, sharing the memory bandwidth between CPUs and accelerators. To address these challenges, we present EdgeIso, a light-weight scheduler that dynamically isolates the performance of tasks on edges. EdgeIso periodically monitors the resource contention and mitigates the contention to meet the SLOs of tasks by efficiently enforcing several isolation techniques (e.g., DVFS and core allocation) in an incremental manner. Moreover, it detects the changes of task executions or offered loads for tasks, thus handling high load fluctuations adaptively. We implement EdgeIso as a user-level scheduler on the Linux integrates into an NVIDIA Jetson TX2. Our experimental results show that EdgeIso improves the performance of the low-latency tasks significantly while improving resource efficiency compared with both the offloading and reservation scheme used in clouds. Yoonsung Nam, Yongjun Choi, Byeonghun Yoo, Hyeonsang Eom, Yongseok Son |
IPDPS | 4 |
| 2020 | Towards Hybrid Isolation for Shared Multicore Systems
Yoonsung Nam, Byeonghun Yoo, Yongjun Choi, Yongseok Son, Hyeonsang Eom |
JSSPP | 5 |
| 2019 | DCA-IO: A Dynamic I/O Control Scheme for Parallel and Distributed File SystemsabstractIn high-performance computing, storage is a shared resource and used by all users with many different application requirements and knowledge of storage. Consequently, the optimal storage configuration varies according to the I/O behavior of each application. While system logs are helpful resources in understanding the storage behavior, it is non-trivial for each user to analyze the logs and adjust complex configurations. Even for experienced users, it is difficult to understand the full stack of I/O systems and find the optimal configuration for the specific application. In this work, we analyzed the I/O activities of CORI which is an HPC system in National Energy Research Scientific Computing Center (NERSC). The result of our analysis shows that most users do not adjust storage configurations and use the default settings. Also, it shows that only a few applications are executed repeatedly in the HPC environment. Based on this result, we have developed DCA-IO, a dynamic distributed file system configuration adjustment algorithm, which utilizes system log information and widely adapted rules to adjust storage configurations automatically without any user intervention. DCA-IO utilizes existing system logs and does not require any modifications in code or an additional library. To demonstrate the effectiveness of DCA-IO, we have performed experiments using I/O kernels of the real applications in both isolated small-sized Lustre environment and CORI. Our experimental result shows that the use of our scheme can lead to improvements in the performance of HPC applications by up to 75% in an isolated environment and 50% in a real HPC environment without user intervention. Sunggon Kim, Alex Sim, Kesheng Wu, Surendra Byna, Yongseok Son, Hyeonsang Eom |
CCGRID | 7 |
| 2018 | OLM: online LLC management for container-based cloud service
Hanul Sung, Myungsun Kim, Jeesoo Min, Hyeonsang Eom |
J. Supercomput. | 4 |
| 2017 | Accelerating a Burst Buffer Via User-Level I/O IsolationabstractBurst buffers tolerate I/O spikes in High-Performance Computing environments by using a non-volatile flash technology. Burst buffers are commonly located between parallel file systems and compute nodes, handling bursty I/Os in the middle. In this architecture, burst buffers are shared resources. The performance of an SSD is significantly reduced when it is used excessively because of garbage collection, and we have observed that SSDs in a burst buffer become slow when many users simultaneously use the burst buffer. To mitigate the performance problem, we propose a new user-level I/O isolation framework in a High-Performance Computing environment using a multi-streamed SSD. The multi-streamed SSD allocates the same flash block for I/Os in the same stream. We assign a different stream to each user; thus, the user can use the stream exclusively. To evaluate the performance, we have used open-source supercomputing workloads and I/O traces from real workloads in the Cori supercomputer at the National Energy Research Scientific Computing Center. Via user-level I/O isolation, we have obtained up to a 125% performance improvement in terms of I/O throughput. In addition, our approach reduces the write amplification in the SSDs, leading to improved SSD endurance. This user-level I/O isolation framework could be applied to deployed burst buffers without having to make any user interface changes. Jaehyun Han, Donghun Koo, Glenn K. Lockwood, Jaehwan Lee 0001, Hyeonsang Eom, Soonwook Hwang |
CLUSTER | 5 |
| 2015 | Multi-source broadcast authentication with Combined Key Chains for wireless ad hoc networksabstractAbstract Multiple trust sources may be needed for broadcast in wireless ad hoc networks. For example, multiple base stations may be employed in some wireless sensor networks, or multiple trusts may be distributed among multiple routing nodes in multi‐hop routing protocol. Most of the previously proposed multicast/broadcast authentication protocols assume a single source of trust. With multiple trust sources, it becomes difficult to design resource‐efficient authentication protocols for multicast/broadcast services. Simply extending traditional approaches will result in increased bandwidth and memory consumptions in proportion to the number of trust sources. In this paper, we propose a new scheme utilizing Combined Key Chains. If there are m trust sources, our scheme generates m + 1 key chains, where m of them are distributed among the m source nodes and the last one is used as a Verification Key Chain in all the receiver nodes. The communication overhead is small and constant, and the memory requirement at a verifier node is also minimal. Copyright © 2014 John Wiley & Sons, Ltd. Seonho Choi, Hyeonsang Eom, Edward Jung |
Secur. Commun. Networks | 3 |
| 2014 | HIOPS-KV: Exploiting multiple flash solid-state drives for key value storesabstractCurrent key value stores rely on DRAM based inmemory architectures where scalability is limited by high power and low density of DRAM. As an alternative, flash SSDs has been explored because of the merits of low power, high density and high internal parallelism. However, the unpredictable latency caused by SSD internal resource conflicts challenges the use of flash SSDs. To address this issue, we present HIOPS-KV a storage I/O software stack for key value stores. HIOPS-KV exploits multiple solid-state drives (SSDs) to control the latencies. With replicas, HIOPS-KV avoids structural collisions which cause long latency operations by spreading colliding operations to distinct devices. For evaluation, we integrated HIOPS-KV into memcached on a low cost high IOPS SSD system built with PC components. At 32 YCSB clients, our system was capable of 117k ops/sec with 263 us average latency showing approximately 4ms at the 99th percentile latency. Woong Shin, Myeongcheol Kim, Hyeonsang Eom, Heon Young Yeom |
ICPADS | 4 |
| 2014 | Enhancing the I/O system for virtual machines using high performance SSDsabstractStorage I/O in VM (Virtual Machine) environments, which requires low latency, becomes problematic as the fast storage such as SSDs (Solid-State Drives) is currently in use. The low performance problem in the VM environment is caused by 1) the presence of additional software layer such as guest OS, 2) context switching between VM and host OS, and 3) scheduling delay for I/O process. These factors do not cause serious problems in the case of using HDD which leads to high latency batching. However, there will be significant performance degradation when fast storage devices are used. To address this problem, we have proposed the following methods to improve the performance of I/O stack in the VM environments by attempting to optimize the I/O stack: one is pipelined polling, and the other is multiple issues and multiple completions. We have found via experiments that our approach leads to increases in the performance of SSDs in a VM environment by up to 50% when multiple VM storage devices are used, and that it leads to improvements in the performance by more than 80% when a single VM storage device is used, with the CPU utilization reduced by up to 25%. Myoungwon Oh, Hyeonsang Eom, Heon Young Yeom |
IPCCC | 2 |
| 2014 | Bubble Task: A Dynamic Execution Throttling Method for Multi-core Resource Management
Dongyou Seo, Myungsun Kim, Hyeonsang Eom, Heon Young Yeom |
JSSPP | 3 |
| 2014 | OS I/O Path Optimizations for Flash Solid-state Drives
Woong Shin, Qichen Chen, Myoungwon Oh, Hyeonsang Eom, Heon Young Yeom |
USENIX ATC | 4 |
| 2014 | Optimizing the Block I/O Subsystem for Fast Storage DevicesabstractFast storage devices are an emerging solution to satisfy data-intensive applications. They provide high transaction rates for DBMS, low response times for Web servers, instant on-demand paging for applications with large memory footprints, and many similar advantages for performance-hungry applications. In spite of the benefits promised by fast hardware, modern operating systems are not yet structured to take advantage of the hardware’s full potential. The software overhead caused by an OS, negligible in the past, adversely impacts application performance, lessening the advantage of using such hardware. Our analysis demonstrates that the overheads from the traditional storage-stack design are significant and cannot easily be overcome without modifying the hardware interface and adding new capabilities to the operating system. In this article, we propose six optimizations that enable an OS to fully exploit the performance characteristics of fast storage devices. With the support of new hardware interfaces, our optimizations minimize per-request latency by streamlining the I/O path and amortize per-request latency by maximizing parallelism inside the device. We demonstrate the impact on application performance through well-known storage benchmarks run against a Linux kernel with a customized SSD. We find that eliminating context switches in the I/O path decreases the software overhead of an I/O request from 20 microseconds to 5 microseconds and a new request merge scheme called Temporal Merge enables the OS to achieve 87% to 100% of peak device performance, regardless of request access patterns or types. Although the performance improvement by these optimizations on a standard SATA-based SSD is marginal (because of its limited interface and relatively high response times), our sensitivity analysis suggests that future SSDs with lower response times will benefit from these changes. The effectiveness of our optimizations encourages discussion between the OS community and storage vendors about future device interfaces for fast storage devices. Youngjin Yu, Dongin Shin, Woong Shin, Nae Young Song, Jaewoo Choi 0004, Hyeong Seog Kim, Hyeonsang Eom, Heon Young Yeom |
ACM Trans. Comput. Syst. | 7 |
| 2014 | Towards High-Performance SAN with Fast Storage DevicesabstractStorage area network (SAN) is one of the most popular solutions for constructing server environments these days. In these kinds of server environments, HDD-based storage usually becomes the bottleneck of the overall system, but it is not enough to merely replace the devices with faster ones in order to exploit their high performance. In other words, proper optimizations are needed to fully utilize their performance gains. In this work, we first adopted a DRAM-based SSD as a fast backend-storage in the existing SAN environment, and found significant performance degradation compared to its own capabilities, especially in the case of small-sized random I/O pattern, even though a high-speed network was used. We have proposed three optimizations to solve this problem: (1) removing software overhead in the SAN I/O path; (2) increasing parallelism in the procedures for handling I/O requests; and (3) adopting the temporal merge mechanism to reduce network overheads. We have implemented them as a prototype and found that our approaches make substantial performance improvements by up to 39% and 280% in terms of both the latency and bandwidth, respectively. Jaewoo Choi 0004, Dongin Shin, Youngjin Yu, Hyeonsang Eom, Heon Young Yeom |
ACM Trans. Storage | 4 |
| 2013 | Virtual machine consolidation based on interference modeling
Shin Gyu Kim, Hyeonsang Eom, Heon Young Yeom |
J. Supercomput. | 2 |
| 2012 | Exploiting Peak Device Throughput from Random Access Workload
Youngjin Yu, Dongin Shin, Woong Shin, Nae Young Song, Hyeonsang Eom, Heon Young Yeom |
HotStorage | 5 |
| 2012 | Motion Object and Regional Detection Method Using Block-Based Background Difference Video FramesabstractSmart CCTV (Closed-Circuit Television) technology has increasingly been developed in the last few years to judge the situation and notify the administrator or take immediate action for security and surveillance reasons. Currently the methods to detect object motion typically include the Frame Difference Method (FDM) which can detect moving objects and the Background Subtraction Method (BSM) which is able to detect motionless objects. Those results can be obtained only if there were some background images ready in advance. The Adaptive Background Subtraction Method (ABSM) also could not recognize an object very well if there are rapid scene changes or an object does not move relatively for a long time. To resolve such a problem, in this research, a filmed image has been divided into the regular sized blocks and then, only the necessary parts of the previous frame image are updated in real-time and a background image was generated so that it is insensible to surrounding environment changes such as object motion, noise or light variations. We proposed a novel moving object detection method which showed high performance with regard to the MSE (Mean Squared Error) and the accuracy of detecting the moving object contours compared to other existing methods. We also evaluated quantitatively the detectability for a moving object region by quickly creating a background image even if it is difficult to shoot a background image or we do not have the baseline image prepared in advance. The proposed method could be used for cases that any background image does not exist or hard to be generated. It is also good for observation of many places at the same time with only a single CCTV system since it is especially robust to abrupt scene changes. Jiwoong Bang, Hyeonsang Eom |
RTCSA | 3 |
| 2012 | Asymmetry-aware load balancing for parallel applications in single-ISA multi-core systemsabstractContemporary operating systems for single-ISA (instruction set architecture) multi-core systems attempt to distribute tasks equally among all the CPUs. This approach works relatively well when there is no difference in CPU capability. However, there are cases in which CPU capability differs from one another. For instance, static capability asymmetry results from the advent of new asymmetric hardware, and dynamic capability asymmetry comes from the operating system (OS) outside noise caused from networking or I/O handling. These asymmetries can make it hard for the OS scheduler to evenly distribute the tasks, resulting in less efficient load balancing. In this paper, we propose a user-level load balancer for parallel applications, called the’ capability balancer’, which recognizes the difference of CPU capability and makes subtasks share the entire CPU capability fairly. The balancer can coexist with the existing kernel-level load balancer without detrimenting the behavior of the kernel balancer. The capability balancer can fairly distribute CPU capability to tasks with very little overhead. For real workloads like the NAS Parallel Benchmark (NPB), we have accomplished speedups of up to 9.8% and 8.5% in dynamic and static asymmetries, respectively. We have also experienced speedups of 13.3% for dynamic asymmetry and 24.1% for static asymmetry in a competitive environment. The impacts of our task selection policies, FIFO (first in, first out) and cache, were compared. The use of the cache policy led to a speedup of 5.3% in overall execution time and a decrease of 4.7% in the overall cache miss count, compared with the FIFO policy, which is used by default. Eunsung Kim, Hyeonsang Eom, Heon Young Yeom |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | Enhancing QoS and Energy Efficiency of Realtime Network Application on Smartphone Using Cloud ComputingabstractThis paper proposes a scheme to enhance energy efficiency and QoS of real time network applications on smart phone. The scheme reduces energy consumption and increases the successful interaction rate between the client at smart phone and the busy server of real time network application by deploying a surrogate of the client at smart phone in cloud computing environment. All interactions among the client at smart phone, the application server and the surrogate in the cloud are controlled by tokens. The proposed scheme considers security as well as energy waste in the cloud. Im Young Jung, Insoon Jo, Youngjin Yu, Hyeonsang Eom, Heon Young Yeom |
APSCC | 4 |
| 2011 | Modeling System Power Consumption Considering DVFS and Thermal Effect
Hyeong Seog Kim, Frank Yong-Kyung Oh, Hyeonsang Eom, Heon Young Yeom |
ICSOFT (1) | 3 |
| 2011 | Multi-layer Trust Reasoning on Open Provenance Model for E-Science EnvironmentabstractTrust for the data created and processed on e-Science environment can be estimated with provenance. The information to form provenance, which says how the data was created and reached its current state, increases as data evolves. It is a heavy burden to trace and verify the massive provenance along the history of data in order to trust data. On the other hand, it is another issue how to trust the verification of data with provenance assuming that the provenance is believable. This paper proposes the property-based trust reasoning which cuts down the overhead to track the history and the origin of data with provenance by semantic path on Open Provenance Model(OPM). Also, the domain-based trust reasoning is adopted, which uses the domain specialty of e-Science environment. The two trust reasonings form the multi-layer trust reasoning. The effectiveness of the proposal is shown by quantitative analysis of overhead reduction as well as by qualitative analysis of trust reasoning. Im Young Jung, Hyeonsang Eom, Heon Young Yeom |
ISPA | 2 |
| 2011 | An efficient skyline framework for matchmaking applications
Hyuck Han, Hyungsoo Jung 0001, Hyeonsang Eom, Heon Young Yeom |
J. Netw. Comput. Appl. | 3 |
| 2011 | Ozone (O3): An Out-of-Order Flash Memory Controller ArchitectureabstractOzone (O3) is a flash memory controller that increases the performance of a flash storage system by executing multiple flash operations out of order. In the O3 flash controller, data dependencies are the only ordering constraints on the execution of multiple flash operations. This allows O3 to exploit the multichip parallelism inherent in flash memory much more effectively than interleaving. The O3 controller also provides a prioritized handling of flash operations, equipping flash management software, such as the FTL (flash translation layer), with control knobs for managing flash operations of different time criticalities. Running a range of workloads on an FPGA implementation showed that the O3 flash controller achieves 3 to 100 percent more throughput than interleaving, with 46 to 88 percent lower response times. Eyee Hyun Nam, Bryan S. Kim, Hyeonsang Eom, Sang Lyul Min |
IEEE Trans. Computers | 3 |
| 2011 | Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk SubsystemsabstractWrite-through caching in modern disk drives enables the protection of data in the event of power failures as well as from certain disk errors when the write-back cache does not. Host system can achieve these benefits at the price of significant performance degradation, especially for small disk writes. We present new block-level techniques to address the performance problem of write-through caching disks. Our techniques are strongly motivated by some interesting results when the disk-level caching is turned off. By extending the conventional request merging, request bridging increases the request size and amortizes the inherent delays in the disk drive across more bytes of data. Like sector interleaving, request interleaving rearranges requests to prevent the disk head from missing the target sector position in close proximity, and thus reduces disk latency. We have evaluated our block-level approach using a variety of I/O workloads and shown that it increases disk I/O throughput by up to about 50%. For some real-world workloads, the disk performance is comparable or even superior to that of using the write-back disk cache. In practice, our simple yet effective solutions achieve better tradeoffs between data reliability and disk performance when applied to write-through caching disks. Dongin Shin, Youngjin Yu, Hyeong Seog Kim, Hyeonsang Eom, Heon Young Yeom |
ACM Trans. Storage | 4 |
| 2010 | Large Graph Processing Based on Remote Memory SystemabstractThis paper focuses on large graph processing based on the remote memory system. Using our remote memory system enables applications to deal with large data sets, especially graph data, which do not fit into the machines main memory. Although recent dramatic increases in DRAM capacity now allow us to build inexpensive computers with very large amounts of main memory, the rise in brand-new Internet services has resulted in rapid increases in data size. This is especially true for on-line social network services that generate various data sets that can be represented as graphs. On the other hand, high-speed networking technologies such as Infini Band, Myrinet and 10G Ethernet now enable us to transfer data with low latency and high throughput. The advanced networking technologies reduce the latency/bandwidth gap between main memory and remote memory. Thus, remote memory based processing could now be helpful in accelerating large-scale graph process when main memory space is insufficient to store application data. In this paper, we present our design and implementation of remote memory system that efficiently processes large graph data. We also evaluate a breadth-first search of various types of graphs using our system and show that our approach is good for large graph data processing. Kyungho Jeon, Hyuck Han, Shin Gyu Kim, Hyeonsang Eom, Heon Young Yeom |
HPCC | 4 |
| 2010 | NCQ vs. I/O scheduler: Preventing unexpected misbehaviorsabstractNative Command Queueing (NCQ) is an optimization technology to maximize throughput by reordering requests inside a disk drive. It has been so successful that NCQ has become the standard in SATA 2 protocol specification, and the great majority of disk vendors have adopted it for their recent disks. However, there is a possibility that the technology may lead to an information gap between the OS and a disk drive. A NCQ-enabled disk tries to optimize throughput without realizing the intention of an OS, whereas the OS does its best under the assumption that the disk will do as it is told without specific knowledge regarding the details of the disk mechanism. Let us call this expectation discord , which may cause serious problems such as request starvations or performance anomaly. In this article, we (1) confirm that expectation discord actually occurs in real systems; (2) propose software-level approaches to solve them; and (3) evaluate our mechanism. Experimental results show that our solution is simple, cheap (no special hardware required), portable, and effective. Youngjin Yu, Dongin Shin, Hyeonsang Eom, Heon Young Yeom |
ACM Trans. Storage | 3 |
| 2008 | Information-Dynamics-Conscious Development of Routing Software: A Case of Routing Software that Improves Link-State Routing Based on Future Link-Delay-Information EstimationabstractIn link-state routing, routes are determined based on the estimates of the current delays on the links, i.e. without considering the dynamics of the link-delay information. Ideally, a data packet should be routed based on the delays it will encounter at each link of the path at the time the packet gets to the link. To address this issue, we have designed a new routing software that improves link-state routing by estimating and using the future link delays encountered by data packets. In link-state routing, link-delay estimates are periodically flooded throughout the network. This flooding of link-delay estimates is done without considering the relevance of these estimates to routing quality, i.e. without taking into account the usefulness of the link-delay information. Our routing-software design also improves link-state routing by broadcasting these estimates only to the extent that they are relevant. In the design of routing software, we consider the temporal change of the link-delay information and its usefulness given the information at a time; we call this an information-dynamics-conscious approach. Simulation studies suggest that this design can lead to significant reductions in routing traffic with noticeable improvements of routing quality in high-load conditions, demonstrating the effectiveness of information-dynamics-conscious development of routing software. Hyeonsang Eom |
Comput. J. | 1 |
| 2008 | A compound framework for sports results prediction: A football case study
Byungho Min, Jinhyuck Kim, Chongyoun Choe, Hyeonsang Eom, Robert I. McKay |
Knowl. Based Syst. | 4 |
| 2006 | The robot software communications architecture (RSCA): embedded middleware for networked service robotsabstractIn this paper, we present a robot middleware technology named robot software communications architecture (RSCA) for its use in networked home service robots. The RSCA provides a standard operating environment for the robot applications together with a framework that expedites the development of such applications. The operating environment is comprised of a real-time operating system, a communication middleware, and a deployment middleware. Particularly, the deployment middleware supports the reconfiguration of component-based robot applications including installation, creation, start, stop, tear-down, and un-installation. In designing RSCA, we have adopted a middleware called SCA from the software defined radio domain and extend it since the original SCA lacks the real-time guarantees and appropriate event services. We have fully implemented RSCA and performed measurements to quantify its run-time performance. Our implementation clearly shows the viability of RSCA. Seongsoo Hong, Jaesoo Lee, Hyeonsang Eom, Gwangil Jeon |
IPDPS | 3 |
| 2001 | Achieving Efficiency and Accuracy in Simulation for I/O-Intensive Applications
Hyeonsang Eom, Jeffrey K. Hollingsworth |
J. Parallel Distributed Comput. | 1 |
| 2001 | A Tool to Help Tune where Computation Is PerformedabstractWe introduce a new performance metric, called load balancing factor (LBF), to assist programmers when evaluating different tuning alternatives. The LBF metric differs from traditional performance metrics since it is intended to measure the performance implications of a specific tuning alternative rather than quantifying where time is spent in the current version of the program. A second unique aspect of the metric is that it provides guidance about moving work within a distributed or parallel program rather than reducing it. A variation of the LBF metric can also be used to predict the performance impact of changing the underlying network. The LBF metric is computed incrementally and online during the execution of the program to be tuned. We also present a case study that shows that our metric can accurately predict the actual performance gains for a test suite of six programs. Hyeonsang Eom, Jeffrey K. Hollingsworth |
IEEE Trans. Software Eng. | 1 |
| 2000 | Speed vs. Accuracy in Simulation for I/O-Intensive ApplicationsabstractThis paper presents a family of simulators that have been developed for data-intensive applications, and a methodology to select the most efficient one based on a user-supplied requirement for accuracy. The methodology consists of a series of tests that select an appropriate simulation based on the attributes of the application. In addition, each simulator provides two estimates of application execution time: one for the minimum expected time and the other for the maximum. We present the results of applying the strategy to existing applications and show that we can accurately simulate applications tens to hundreds of times faster than application execution time. Hyeonsang Eom, Jeffrey K. Hollingsworth |
IPDPS | 1 |
| 1998 | LBF: A Performance Metric for Program ReorganizationabstractWe introduce a new performance metric, called Load Balancing Factor (LBF), to assist programmers with evaluating different tuning alternatives. The LBF metric differs from traditional performance metrics since it is intended to measure the performance implications of a specific tuning alternative rather than quantifying where time is spent in the current version of the program. A second unique aspect of the metric is that it provides guidance about moving work within a distributed or parallel program rather than reducing it. A variation of the LBF metric can also be used to predict the performance impact of changing the underlying network. The LBF metric can be computed incrementally and online during the execution of the program to be tuned. We also present a case study that shows that our metric can predict the actual performance gains accurately for a test suite of six programs. Hyeonsang Eom, Jeffrey K. Hollingsworth |
ICDCS | 1 |