Youngjin Yu

dblp:67/3578 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 99% Performance modeling and evaluation · 1%
Software engineering, system software, and programming languages
2 papers
Operating systems · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › data reduction
data deduplication
0.712023
TiDedup: A New Distributed Deduplication Architecture for Ceph · USENIX ATC 2023
Storage systems › data reduction › data deduplication
distributed deduplication
0.712023
TiDedup: A New Distributed Deduplication Architecture for Ceph · USENIX ATC 2023
Storage systems › flash and SSD
solid-state drive
0.422014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Storage systems › storage devices
fast storage devices
0.222014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Operating systems › i/o
i/o subsystem
0.212014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Storage systems › networked storage
storage area network
0.212014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Storage systems › i/o scheduling
disk scheduling
0.112011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Storage systems › magnetic storage
disk storage
0.112011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Operating systems › i/o › i/o subsystem
i/o scheduling
0.112010
NCQ vs. I/O scheduler: Preventing unexpected misbehaviors · ACM Trans. Storage 2010
Storage systems › magnetic storage
hard disk drive
0.112010
NCQ vs. I/O scheduler: Preventing unexpected misbehaviors · ACM Trans. Storage 2010
Storage systems
flash and SSD
0.112014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Storage systems › i/o optimization
i/o path optimization
0.112014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Storage systems
i/o workload
0.012011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Performance modeling and evaluation
workload characterization
0.012011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011

Methods — techniques the papers use, named apart from their topics

temporal merge · 0.6hardware interface modification · 0.4context switch elimination · 0.4request reordering · 0.2parallelism optimization · 0.2request merging · 0.1request interleaving · 0.1request bridging · 0.1
YearPublicationVenuePosition
2023 NVMe-Driven Lazy Cache Coherence for Immutable Data with NVMe over Fabrics
abstract
In this work, we explore opportunities to design shared storage systems that leverage the distance connectivity of NVMe over Fabrics (NVMe-oF). NVMe-oF enables the use of NVMe storage devices in a shared storage environment, where multiple servers can access the same storage device via RDMA. Leveraging the distance connectivity of NVMe-oF, we develop a shared file system called EXT4-oF by extending the EXT4 file system. EXT4-oF uses RDMA to enable a local file system to function as a shared file system without requiring remote daemon processes. EXT4-oF employs a novel NVMe-driven lazy cache coherence to maintain cache coherence of file system metadata across multiple compute nodes upon creating new files, all achieved without the need for any daemon processes. To ensure cache coherence, NVMe-driven lazy cache coherence mechanism requires compute nodes to perform a re-read of NVMe-oF to avoid false negative file open errors. Through our experiments, we demonstrate that EXT4-oF improves the performance of MinIO by minimizing network traffic between compute and storage nodes and eliminating the need for TCP/IP communication during remote reads.
Hyeongjun Jeon, Daegyu Han, Duck-Ho Bae, Youngjin Yu, Kyeungpyo Kim, Sung-Soon Park 0001, Jinkyu Jeong, Beomseok Nam
CLOUD5
2023 TiDedup: A New Distributed Deduplication Architecture for Ceph
Myoungwon Oh, Samuel Just, Youngjin Yu, Duck-Ho Bae, Sage A. Weil, Sangyeun Cho, Heon Young Yeom
USENIX ATC4
2014 Optimizing the Block I/O Subsystem for Fast Storage Devices
abstract
Fast storage devices are an emerging solution to satisfy data-intensive applications. They provide high transaction rates for DBMS, low response times for Web servers, instant on-demand paging for applications with large memory footprints, and many similar advantages for performance-hungry applications. In spite of the benefits promised by fast hardware, modern operating systems are not yet structured to take advantage of the hardware’s full potential. The software overhead caused by an OS, negligible in the past, adversely impacts application performance, lessening the advantage of using such hardware. Our analysis demonstrates that the overheads from the traditional storage-stack design are significant and cannot easily be overcome without modifying the hardware interface and adding new capabilities to the operating system. In this article, we propose six optimizations that enable an OS to fully exploit the performance characteristics of fast storage devices. With the support of new hardware interfaces, our optimizations minimize per-request latency by streamlining the I/O path and amortize per-request latency by maximizing parallelism inside the device. We demonstrate the impact on application performance through well-known storage benchmarks run against a Linux kernel with a customized SSD. We find that eliminating context switches in the I/O path decreases the software overhead of an I/O request from 20 microseconds to 5 microseconds and a new request merge scheme called Temporal Merge enables the OS to achieve 87% to 100% of peak device performance, regardless of request access patterns or types. Although the performance improvement by these optimizations on a standard SATA-based SSD is marginal (because of its limited interface and relatively high response times), our sensitivity analysis suggests that future SSDs with lower response times will benefit from these changes. The effectiveness of our optimizations encourages discussion between the OS community and storage vendors about future device interfaces for fast storage devices.
Youngjin Yu, Dongin Shin, Woong Shin, Nae Young Song, Jaewoo Choi 0004, Hyeong Seog Kim, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Comput. Syst.1
2014 Towards High-Performance SAN with Fast Storage Devices
abstract
Storage area network (SAN) is one of the most popular solutions for constructing server environments these days. In these kinds of server environments, HDD-based storage usually becomes the bottleneck of the overall system, but it is not enough to merely replace the devices with faster ones in order to exploit their high performance. In other words, proper optimizations are needed to fully utilize their performance gains. In this work, we first adopted a DRAM-based SSD as a fast backend-storage in the existing SAN environment, and found significant performance degradation compared to its own capabilities, especially in the case of small-sized random I/O pattern, even though a high-speed network was used. We have proposed three optimizations to solve this problem: (1) removing software overhead in the SAN I/O path; (2) increasing parallelism in the procedures for handling I/O requests; and (3) adopting the temporal merge mechanism to reduce network overheads. We have implemented them as a prototype and found that our approaches make substantial performance improvements by up to 39% and 280% in terms of both the latency and bandwidth, respectively.
Jaewoo Choi 0004, Dongin Shin, Youngjin Yu, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Storage3
2013 Dynamic Interval Polling and Pipelined Post I/O Processing for Low-Latency Storage Class Memory
Dongin Shin, Youngjin Yu, Hyeong Seog Kim, Jaewoo Choi 0004, Do Yung Jung, Heon Young Yeom
HotStorage2
2012 Exploiting Peak Device Throughput from Random Access Workload
Youngjin Yu, Dongin Shin, Woong Shin, Nae Young Song, Hyeonsang Eom, Heon Young Yeom
HotStorage1
2011 Enhancing QoS and Energy Efficiency of Realtime Network Application on Smartphone Using Cloud Computing
abstract
This paper proposes a scheme to enhance energy efficiency and QoS of real time network applications on smart phone. The scheme reduces energy consumption and increases the successful interaction rate between the client at smart phone and the busy server of real time network application by deploying a surrogate of the client at smart phone in cloud computing environment. All interactions among the client at smart phone, the application server and the surrogate in the cloud are controlled by tokens. The proposed scheme considers security as well as energy waste in the cloud.
Im Young Jung, Insoon Jo, Youngjin Yu, Hyeonsang Eom, Heon Young Yeom
APSCC3
2011 Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems
abstract
Write-through caching in modern disk drives enables the protection of data in the event of power failures as well as from certain disk errors when the write-back cache does not. Host system can achieve these benefits at the price of significant performance degradation, especially for small disk writes. We present new block-level techniques to address the performance problem of write-through caching disks. Our techniques are strongly motivated by some interesting results when the disk-level caching is turned off. By extending the conventional request merging, request bridging increases the request size and amortizes the inherent delays in the disk drive across more bytes of data. Like sector interleaving, request interleaving rearranges requests to prevent the disk head from missing the target sector position in close proximity, and thus reduces disk latency. We have evaluated our block-level approach using a variety of I/O workloads and shown that it increases disk I/O throughput by up to about 50%. For some real-world workloads, the disk performance is comparable or even superior to that of using the write-back disk cache. In practice, our simple yet effective solutions achieve better tradeoffs between data reliability and disk performance when applied to write-through caching disks.
Dongin Shin, Youngjin Yu, Hyeong Seog Kim, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Storage2
2010 NCQ vs. I/O scheduler: Preventing unexpected misbehaviors
abstract
Native Command Queueing (NCQ) is an optimization technology to maximize throughput by reordering requests inside a disk drive. It has been so successful that NCQ has become the standard in SATA 2 protocol specification, and the great majority of disk vendors have adopted it for their recent disks. However, there is a possibility that the technology may lead to an information gap between the OS and a disk drive. A NCQ-enabled disk tries to optimize throughput without realizing the intention of an OS, whereas the OS does its best under the assumption that the disk will do as it is told without specific knowledge regarding the details of the disk mechanism. Let us call this expectation discord , which may cause serious problems such as request starvations or performance anomaly. In this article, we (1) confirm that expectation discord actually occurs in real systems; (2) propose software-level approaches to solve them; and (3) evaluate our mechanism. Experimental results show that our solution is simple, cheap (no special hardware required), portable, and effective.
Youngjin Yu, Dongin Shin, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Storage1
2007 Shedding Light in the Black-Box : Structural Modeling of Modern Disk Drives
abstract
The performance of computer systems depends on relatively slow disk I/O performance. In order to improve the disk I/O performance, it is required to reduce a mechanical delay induced by the disk I/O operations. Several approaches have been proposed for it. However, because hard-disk storage hides too much information to the outside world, it makes difficult to predict exact internal layout of disk storage. This paper introduces a technique which brings this black box to light empirically. The technique can be called a gray-box approach due to the use of some prior knowledge about disk drives. We propose a new algorithm that extract disk model parameters and build overall multi-dimensional disk structural model for several up-to-data IDE disk drives whose internals are known as a black-box. We validate the model accuracy through seek time analysis. We expect our modeling result can be applied to many researches optimizing disk I/O performance.
Dongin Shin, Youngjin Yu, Heon Young Yeom
MASCOTS2
2007 Multi-Hop Cooperative Sensing and Transmit Power Control based on Interference Information for Cognitive Radio
abstract
In cognitive radio, reliable detection of the primary system is crucial for the secondary system. However, there is an uncertainty in the received power level due to shadowing. The cooperative sensing where the stations in the secondary system share their sensing information can improve the detection accuracy of the primary system. In this paper, a new secondary system for reducing interference to the primary system is proposed. In the proposed system, a secondary transmitter receives sensing information by using multi-hop transmission from all secondary stations which detect the primary system, and makes decision of its transmission based on the received sensing information. Moreover, it controls the transmit power by the received CINR information from the secondary sensing stations. Simulation results assuming spatially independent shadowing and ideal multi-hop transmission in the secondary system reveal that the proposed system is helpful to protect the primary receivers which have no transmission ability.
Youngjin Yu, Hidekazu Murata, Koji Yamamoto 0001, Susumu Yoshida
PIMRC1
2006 Practical Fault-Tolerant Framework for eScience Infrastructure
abstract
Many areas of science currently use computing resources as a important part of their research, and many research groups adopt cluster architecture to use them efficiently and manage them easily. Therefore, faulttolerance becomes a very important property for the computing resources. However, fault-tolerant systems have not yet been widely adopted because they are either hard to deploy, hard to use, hard to manage, hard to maintain, or hard to justify. This paper proposes a practical fault-tolerant system for eScience infrastructures. Our system uses checkpoint/ restart mechanism for fault-tolerance, and provides a easy mechanism to integrate with Grid services widely used in eScience. Additionally, we run rigorous tests using scientific applications to verify that our system can be used in clusters. We also describe improvements made to our system to solve various problems that arose when deploying it on a cluster. The experimental results show that not only does our system conform to various types of running environment well, but that it can also be practically deployed in clusters.
Hyuck Han, Jai Wug Kim, Jongpil Lee, Youngjin Yu, Kiyoung Kim, Heon Young Yeom
e-Science4
2006 SHIELD: A Fault-Tolerant MPI for an Infiniband Cluster
Hyuck Han, Hyungsoo Jung 0001, Jai Wug Kim, Jongpil Lee, Youngjin Yu, Shin Gyu Kim, Heon Young Yeom
HPCC5