Dongin Shin

dblp:58/4165 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
1since 2021 · last 2026
0009-0004-9906-9434ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 78% Distributed systems · 13% Parallel and multicore computing · 6%
Software engineering, system software, and programming languages
2 papers
Operating systems · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD
solid-state drive
0.422014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Storage systems › storage devices
fast storage devices
0.222014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Operating systems › i/o
i/o subsystem
0.212014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Storage systems › networked storage
storage area network
0.212014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Storage systems › i/o scheduling
disk scheduling
0.112011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Storage systems › magnetic storage
disk storage
0.112011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Operating systems › i/o › i/o subsystem
i/o scheduling
0.112010
NCQ vs. I/O scheduler: Preventing unexpected misbehaviors · ACM Trans. Storage 2010
Storage systems › magnetic storage
hard disk drive
0.112010
NCQ vs. I/O scheduler: Preventing unexpected misbehaviors · ACM Trans. Storage 2010
Storage systems
flash and SSD
0.112014
Optimizing the Block I/O Subsystem for Fast Storage Devices · ACM Trans. Comput. Syst. 2014
Storage systems › i/o optimization
i/o path optimization
0.112014
Towards High-Performance SAN with Fast Storage Devices · ACM Trans. Storage 2014
Distributed systems › fault tolerance
checkpointing
0.112005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005
Distributed systems › fault tolerance › checkpointing
coordinated checkpointing
0.112005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005
Distributed systems
fault tolerance
0.112005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005
Parallel and multicore computing › MPI
fault-tolerant MPI
0.112005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005
Parallel and multicore computing
MPI
0.112005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005
Distributed systems › distributed system architecture › distributed operating systems
process migration
0.112005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005
Storage systems
i/o workload
0.012011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Performance modeling and evaluation
workload characterization
0.012011
Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems · ACM Trans. Storage 2011
Interconnection networks and networks-on-chip › cluster interconnect
myrinet
0.012005
Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3) · SC 2005

Methods — techniques the papers use, named apart from their topics

temporal merge · 0.6hardware interface modification · 0.4context switch elimination · 0.4request reordering · 0.2parallelism optimization · 0.2request merging · 0.1request interleaving · 0.1request bridging · 0.1dynamic process management · 0.1coordinated checkpointing · 0.1
YearPublicationVenuePosition
2026 Performance Characterization and Optimization of LLM Inference on Tenstorrent AI Accelerators
Jangho Lim, Dongin Shin, Uichan Kim, Jinhyeok Choi, Sangwon Shin, Sangwoo Park 0005, Gunjae Koo, Taeweon Suh
Euro-Par (2)2
2014 Optimizing the Block I/O Subsystem for Fast Storage Devices
abstract
Fast storage devices are an emerging solution to satisfy data-intensive applications. They provide high transaction rates for DBMS, low response times for Web servers, instant on-demand paging for applications with large memory footprints, and many similar advantages for performance-hungry applications. In spite of the benefits promised by fast hardware, modern operating systems are not yet structured to take advantage of the hardware’s full potential. The software overhead caused by an OS, negligible in the past, adversely impacts application performance, lessening the advantage of using such hardware. Our analysis demonstrates that the overheads from the traditional storage-stack design are significant and cannot easily be overcome without modifying the hardware interface and adding new capabilities to the operating system. In this article, we propose six optimizations that enable an OS to fully exploit the performance characteristics of fast storage devices. With the support of new hardware interfaces, our optimizations minimize per-request latency by streamlining the I/O path and amortize per-request latency by maximizing parallelism inside the device. We demonstrate the impact on application performance through well-known storage benchmarks run against a Linux kernel with a customized SSD. We find that eliminating context switches in the I/O path decreases the software overhead of an I/O request from 20 microseconds to 5 microseconds and a new request merge scheme called Temporal Merge enables the OS to achieve 87% to 100% of peak device performance, regardless of request access patterns or types. Although the performance improvement by these optimizations on a standard SATA-based SSD is marginal (because of its limited interface and relatively high response times), our sensitivity analysis suggests that future SSDs with lower response times will benefit from these changes. The effectiveness of our optimizations encourages discussion between the OS community and storage vendors about future device interfaces for fast storage devices.
Youngjin Yu, Dongin Shin, Woong Shin, Nae Young Song, Jaewoo Choi 0004, Hyeong Seog Kim, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Comput. Syst.2
2014 Towards High-Performance SAN with Fast Storage Devices
abstract
Storage area network (SAN) is one of the most popular solutions for constructing server environments these days. In these kinds of server environments, HDD-based storage usually becomes the bottleneck of the overall system, but it is not enough to merely replace the devices with faster ones in order to exploit their high performance. In other words, proper optimizations are needed to fully utilize their performance gains. In this work, we first adopted a DRAM-based SSD as a fast backend-storage in the existing SAN environment, and found significant performance degradation compared to its own capabilities, especially in the case of small-sized random I/O pattern, even though a high-speed network was used. We have proposed three optimizations to solve this problem: (1) removing software overhead in the SAN I/O path; (2) increasing parallelism in the procedures for handling I/O requests; and (3) adopting the temporal merge mechanism to reduce network overheads. We have implemented them as a prototype and found that our approaches make substantial performance improvements by up to 39% and 280% in terms of both the latency and bandwidth, respectively.
Jaewoo Choi 0004, Dongin Shin, Youngjin Yu, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Storage2
2013 Dynamic Interval Polling and Pipelined Post I/O Processing for Low-Latency Storage Class Memory
Dongin Shin, Youngjin Yu, Hyeong Seog Kim, Jaewoo Choi 0004, Do Yung Jung, Heon Young Yeom
HotStorage1
2012 Exploiting Peak Device Throughput from Random Access Workload
Youngjin Yu, Dongin Shin, Woong Shin, Nae Young Song, Hyeonsang Eom, Heon Young Yeom
HotStorage2
2011 Dual domain method for single image dehazing and enhancing
abstract
In this paper, we propose a novel method for improving the visibility of an image (with fog or haze), as well as the image's details. The proposed method adjusts the global contrast in the spatial domain for dehazing fog or haze spread over images and then enhances the local contrast in the transform domain for reviving the details of images. Compared to the previous methods performing in the spatial domain only, the proposed method improves the visibility and quality of images significantly.
Dongin Shin, Kristofor B. Gibson, Wonha Kim, Truong Q. Nguyen
ICASSP1
2011 Request Bridging and Interleaving: Improving the Performance of Small Synchronous Updates under Seek-Optimizing Disk Subsystems
abstract
Write-through caching in modern disk drives enables the protection of data in the event of power failures as well as from certain disk errors when the write-back cache does not. Host system can achieve these benefits at the price of significant performance degradation, especially for small disk writes. We present new block-level techniques to address the performance problem of write-through caching disks. Our techniques are strongly motivated by some interesting results when the disk-level caching is turned off. By extending the conventional request merging, request bridging increases the request size and amortizes the inherent delays in the disk drive across more bytes of data. Like sector interleaving, request interleaving rearranges requests to prevent the disk head from missing the target sector position in close proximity, and thus reduces disk latency. We have evaluated our block-level approach using a variety of I/O workloads and shown that it increases disk I/O throughput by up to about 50%. For some real-world workloads, the disk performance is comparable or even superior to that of using the write-back disk cache. In practice, our simple yet effective solutions achieve better tradeoffs between data reliability and disk performance when applied to write-through caching disks.
Dongin Shin, Youngjin Yu, Hyeong Seog Kim, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Storage1
2010 NCQ vs. I/O scheduler: Preventing unexpected misbehaviors
abstract
Native Command Queueing (NCQ) is an optimization technology to maximize throughput by reordering requests inside a disk drive. It has been so successful that NCQ has become the standard in SATA 2 protocol specification, and the great majority of disk vendors have adopted it for their recent disks. However, there is a possibility that the technology may lead to an information gap between the OS and a disk drive. A NCQ-enabled disk tries to optimize throughput without realizing the intention of an OS, whereas the OS does its best under the assumption that the disk will do as it is told without specific knowledge regarding the details of the disk mechanism. Let us call this expectation discord , which may cause serious problems such as request starvations or performance anomaly. In this article, we (1) confirm that expectation discord actually occurs in real systems; (2) propose software-level approaches to solve them; and (3) evaluate our mechanism. Experimental results show that our solution is simple, cheap (no special hardware required), portable, and effective.
Youngjin Yu, Dongin Shin, Hyeonsang Eom, Heon Young Yeom
ACM Trans. Storage2
2007 Shedding Light in the Black-Box : Structural Modeling of Modern Disk Drives
abstract
The performance of computer systems depends on relatively slow disk I/O performance. In order to improve the disk I/O performance, it is required to reduce a mechanical delay induced by the disk I/O operations. Several approaches have been proposed for it. However, because hard-disk storage hides too much information to the outside world, it makes difficult to predict exact internal layout of disk storage. This paper introduces a technique which brings this black box to light empirically. The technique can be called a gray-box approach due to the use of some prior knowledge about disk drives. We propose a new algorithm that extract disk model parameters and build overall multi-dimensional disk structural model for several up-to-data IDE disk drives whose internals are known as a black-box. We validate the model accuracy through seek time analysis. We expect our modeling result can be applied to many researches optimizing disk I/O performance.
Dongin Shin, Youngjin Yu, Heon Young Yeom
MASCOTS1
2005 Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3)
abstract
Advances in network technology and computing power have inspired the emergence of high-performance cluster computing systems. While cluster management and hardware highavailability tools are readily available, practical and easily deployable fault-tolerant systems have not been successfully adopted commercially. We present a fault-tolerant system, Multiple fault-tolerant MPI over Myrinet (M3), that differs in notable respects from other proposed fault-tolerant systems in the literature. M3 is built on top of Myrinet since it is regarded as one of the best solutions for highperformance networks and is widely used in cluster computing systems because it can provide a high-speed switching network that is an inevitable ingredient in interconnecting clusters of workstations or PCs. M^3 is a user-transparent checkpointing system for multiple fault-tolerant MPI implementation that is primarily based on the coordinated checkpointing protocol. M3 supports three critical functionalities that are necessary for faulttolerance: a light-weight failure detection mechanism, dynamic process management that includes process migration, and a consistent checkpoint and recovery mechanism. The features of M are that it requires no modifications of application code and that it preserves much of the high performance characteristics of Myrinet. This paper describes the architecture of M3, its detailed design principles and comprehensive implementation issues. We also propose practical solutions for those involved in constructing highly available cluster systems for parallel programming systems. Experimental results substantiate our assertion that M3 can be a good candidate for practically deployable fault-tolerant systems in very-large and high-performance Myrinet clusters and that its protocol can be applied to a wide variety of parallel communication libraries without difficulty.
Hyungsoo Jung 0001, Dongin Shin, Hyuck Han, Jai Wug Kim, Heon Young Yeom, Jongsuk Lee
SC2