Hyuck Han

dblp:17/5450 · DBLP profile ↗
← Back
48ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0003-0936-9181ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 2 · 1 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Efficient Data Processing using On-the-Fly Host-PIM Interactions in a Commodity PIM System
Hyojune Kim, Jeonghyeon Joo, Taehyeong Park 0001, Yongjun Park, Hyuck Han, Sooyong Kang
ICDE5
2026 Lifetime-Aware Zone Allocation for ZNS SSDS
Doeun Kim, Junmo Seong, Hyuck Han, Sooyong Kang
ICFEC4
2026 Practical Monitoring Tool for ZNS SSD Emulator
Junmo Seong, Doeun Kim, Hyuck Han, Sooyong Kang
ICFEC3
2026 Kafka-Thor: A Kafka-based In-Edge Data Streaming Platform for Enhanced V2X Services
abstract
Enhanced Vehicle-to-Everything (V2X) services require ultra-low latency and robust reliability in end-to-end (E2E) data delivery from producers (e.g., vehicles and roadside units) to consumers (e.g., in-car V2X service applications). One of the candidate data delivery architectures that can meet the stringent requirements is restricting the data delivery path within the network edge by letting an edge server relay data between sources and destinations in its coverage. To that end, edge servers need to equip two important functionalities: 1) a reliable and lightweight data transport protocol to collect data from numerous data sources, and 2) an efficient data streaming platform that delivers collected data to consumers with high throughput and low latency. In this work, we design and implement Kafka- Thor, a Kafka-based high-performance data streaming platform, specifically designed for edge servers to collect and deliver timecritical data for enhanced V2X services. The platform reduces E2E data delivery latency by optimizing the data delivery architecture in Kafka using two novel technologies, the singlepoller, multi-worker (SPMW) architecture and service-specific dynamic batching (SS-batching). Experimental results show that Kafka-Thor significantly improves latency and reliability in E2E data delivery, making enhanced V2X services feasible.
Hyungseok Seo, Hagyeong Lee, Hyuck Han, Minsoo Ryu, Sooyong Kang
IEEE Trans. Serv. Comput.3
2024 TCP: A Tensor Contraction Processor for AI Workloads Industrial Product
abstract
We introduce a novel tensor contraction processor (TCP) architecture that offers a paradigm shift from traditional architectures that rely on fixed-size matrix multiplications. TCP aims at exploiting the rich parallelism and data locality inherent in tensor contractions, thereby enhancing both efficiency and performance of AI workloads.TCP is composed of coarse-grained processing elements (PEs) to simplify software development. In order to efficiently process operations with diverse tensor shapes, the PEs are designed to be flexible enough to be utilized as a large-scale single unit or a set of small independent compute units.We aim at maximizing data reuse on both levels of inter and intra compute units. To do that, we propose a circuit switch-based fetch network to flexibly connect compute units to enable inter-compute unit data reuse. We also exploit input broadcast to multiple contraction engines and input buffer based reuse to further exploit reuse behavior in tensor contraction. Our compiler explores the design space of tensor contractions considering tensor shapes and the order of their associated loop operations as well as the underlying accelerator architecture.A TCP chip was designed and fabricated in 5nm technology as the second-generation product of Furiosa AI, offering 256/512/1024 TOPS (BF16/FP8 or INT8/INT4) with 256 MB SRAM and 1.5 TB/s 48 GB HBM3 under 150 W TDP. Commercialization will start in August 2024.We performed an extensive case study of running the LLaMA-2 7B model and evaluated its performance and power efficiency on various configurations of sequence length and batch size. For this model, TCP is 2.7 × and 4.1 × better than H100 and L40s, respectively, in terms of performance per watt.
Hanjoon Kim, Byeongwook Bae, Hyunmin Jeong, Sang Min Lee 0014, Jeseung Yeon, Changjae Park, Boncheol Gu, Changman Lee, Jaeick Bae, SungGyeong Bae, Yojung Cha, Wooyoung Choe, Jonguk Choi, Juho Ha, Hyuck Han, Namoh Hwang, Seokha Hwang, Kiseok Jang, Haechan Je, Hojin Jeon, Jaewoo Jeon, Hyunjun Jeong, Yeonsu Jung, Dongok Kang, Hyewon Kim, Muhwan Kim, Sewon Kim, Suhyung Kim, Yong Kim, Youngsik Kim, Younki Ku, Jeong Ki Lee, Juyun Lee, Seokho Lee, Minwoo Noh, Hyuntaek Oh, Gyunghee Park, Jimin Seo, Jungyoung Seong, June Paik, Nuno P. Lopes, Sungjoo Yoo
ISCA18
2024 Dynamic zone redistribution for key-value stores on zoned namespaces SSDs
Doeun Kim, Kihan Choi, Hyuck Han, Minsoo Ryu, Sooyong Kang
J. Syst. Archit.4
2023 CredsCache: Making OverlayFS scalable for containerized services
Kihan Choi, Hyungseok Seo, Hyuck Han, Minsoo Ryu, Sooyong Kang
Future Gener. Comput. Syst.3
2022 Workload-optimized sensor data store for industrial IoT gateways
Kihan Choi, Hyuck Han, Hyungsoo Jung 0001, Sooyong Kang
Future Gener. Comput. Syst.2
2021 iEdge: An IoT-assisted Edge Computing Framework
abstract
Edge computing has emerged as a viable solution to bridge the gap between distributed Internet of Things (IoT) devices and centralized distant clouds. In particular, small-scale servers are deployed at the edge of network (i.e., edge servers) to `help' cloud servers process data IoT devices constantly generate. However, these edge servers often struggle to deal with emerging applications that require real-time data processing in situ, such as real-time facial recognition. In this paper, we present iEdge as an IoT-assisted edge computing framework that enables the seamless execution of applications across an edge server and nearby IoT devices. The seamless execution in essence has been realized by transforming platform-dependent monolithic applications to cross-platform composite applications and offloading some tasks/functions of these composite applications to IoT devices considering device context. We have evaluated iEdge using a prototype implementation with a real-time facial recognition application. Experimental results show that iEdge effectively harnesses smart IoT devices as a consolidated edge computing execution environment and enables such an application to process more video streams than typical `edge-only' computing.
Hochul Lee, Seyul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
PerCom4
2020 Design and Implementation of SSD-Assisted Backup and Recovery for Database Systems
abstract
As flash-based solid-state drive (SSD) becomes more prevalent because of the rapid fall in price and the significant increase in capacity, customers expect better data services than traditional disk-based systems. However, the order of magnitude performance provided and new characteristics of flash require a rethinking of data services. For example, backup and recovery is an important service in a database system since it protects data against unexpected hardware and software failures. To provide backup and recovery, backup/recovery tools or backup/recovery methods by operating systems can be used. However, the tools perform time-consuming jobs, and the methods may negatively affect run-time performance during normal operation even though high-performance SSDs are used. To handle these issues, we propose an SSD-assisted backup/recovery scheme for database systems. Our scheme is to utilize the characteristics (e.g., out-of-place update) of flash-based SSD for backup/recovery operations. To this end, we exploit the resources (e.g., flash translation layer and DRAM cache with supercapacitors) inside SSD, and we call our SSD with new backup/ recovery functionality BR-SSD. We design and implement the functionality in the Samsung enterprise-class SSD (i.e., SM843Tn) for more realistic systems. Furthermore, we exploit and integrate BR-SSDs into database systems (i.e., MySQL) in replication and redundant array of independent disks (RAID) environments, as well as a database system in a single BR-SSD. The experimental result demonstrates that our scheme provides fast backup and recovery but does not negatively affect the run-time performance during normal operation.
Yongseok Son, Moonsub Kim, Sunggon Kim, Heon Young Yeom, Nam Sung Kim, Hyuck Han
IEEE Trans. Knowl. Data Eng.6
2019 On the Trade-Off Between Performance and Storage Efficiency of Replication-Based Object Storage
abstract
The object storage systems are used to store and manage unstructured data. Most object storage systems provide the replication policy (REP) or erasure code policy (EC) to ensure the reliability and availability of data. In this paper, we study the trade-off between performance and storage efficiency of these policies with respect to different data sizes of user requests. To this end, we present a hybrid policy management system that takes advantage of both policies by automatically changing policy based on data size. We have implemented the hybrid system in OpenStack Swift. Our evaluation results show the throughput of GET request increases up to 36% while improving storage efficiency by up to 53% compared to that using the REP policy.
Hanbeom Jo, Hochul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
CloudCom5
2019 Border-Collie: A Wait-free, Read-optimal Algorithm for Database Logging on Multicore Hardware
abstract
Actions changing the state of databases are all logged with proper ordering being imposed. Database engines obeying this golden rule of logging enforce total ordering on all events, and this poses challenges in addressing the scalability bottlenecks of database logging on multicore hardware. We reexamined the problem of database logging and realized that in any given log history, obtaining an upper bound on the size of a set that preserves the happen-before relation is the essence of the matter. Based on our understanding, we propose Border-Collie, a wait-free and read-optimal algorithm for database logging that finds such an upper bound even with some worker threads often being idle. We show that (1) Border-Collie always finds the largest set of logged events satisfying the condition in a finite number of steps (i.e., wait-free), (2) the number of logged events to be read is also minimal (i.e., read-optimal), and (3) both properties hold even with threads being in intermittent work. Experimental results demonstrated that Border-Collie proves our claims under various workloads; Border-Collie outperforms the state-of-the-art centralized logging techniques (i.e., Eleda and ERMIA) by up to ~2X and exhibits almost the same throughput with much shorter commit latency than the state-of-the-art decentralized logging techniques (i.e., Silo and FOEDUS).
Jong-Bin Kim, Seohui Son, Hyuck Han, Sooyong Kang, Hyungsoo Jung 0001
SIGMOD Conference4
2019 Mobile collaborative computing on the fly
Hochul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
Pervasive Mob. Comput.4
2018 Efficient Key-Value Stores with Ranged Log-Structured Merge Trees
abstract
The log-structured merge (LSM) tree is designed to provide efficient indexing for data that is frequently updated by using the log-structured approach. It defers merge operations for reordering data, propagating the index changes from a memory-resident component through one or more disk components. Thus, LSM-based storage engines can achieve good write performance. However, processing merge operations incurs high write amplification and memory consumption, ultimately having an adverse effect on system performance. In this paper, we propose the Ranged Log-Structured Merge (RLSM) tree to mitigate the problems of the LSM tree. To reduce the write amplification and memory overhead, RLSM simplifies the logical layout of storage and keeps data as an unsorted order. In addition, we prevent read performance from declining by partitioning data on the disk into multiple files with non-overlapping ranges. We implement our schemes on HBase, one of the most popular key-value storage engines, and evaluate our system by using YCSB benchmark. Our experimental results show that RLSM consequently reduces write amplification by a factor of 3, and memory consumption by up to 24%.
Nae Young Song, Heon Young Yeom, Hyuck Han
IEEE CLOUD3
2018 High-Performance Transaction Processing in Journaling File Systems
Yongseok Son, Sunggon Kim, Heon Young Yeom, Hyuck Han
FAST4
2018 Efficient dentry lookup with backward finding mechanism
abstract
As modern computer systems face the challenge of managing large data, filesystems must deal with a large number of files. This leads to amplified concerns of metadata and data operations. Filesystems in Linux manage the metadata of files by constructing in-memory structures such as directory entry (dentry) and inode. However, we found inefficiencies in metadata management mechanisms, especially in the path traversal mechanism of Linux file systems when searching for a dentry in the dentry cache.
Nae Young Song, Hwajung Kim, Hyuck Han, Heon Young Yeom
HPC Asia3
2018 Hierarchical Recursive Resource Sharing for Containerized Applications
Youngjin Kim 0011, Young Choon Lee, Hyuck Han, Sooyong Kang
ICSOC3
2018 SAMD: Fine-Grained Application Sharing for Mobile Collaboration
abstract
The collective use of ever connected and pervasive mobile devices has been increasingly sought for in mobile collaboration, such as multiplayer mobile gaming and distributed processing. The current model of mobile collaboration requires each device to install a particular, `full' mobile app for a respective collaboration. Besides, collaboration functionalities are typically implemented at application level. In this paper, we present Single Application Multiple Device (SAMD) as a platform-level mobile collaboration framework. A mobile app developed using SAMD is capable of fine-grained application sharing. In particular, SAMD enables devices, agreed to participate in collaboration, to get portions of the app on-the-fly and run them without the prior installation. To achieve this, we have developed three solutions as core functionalities of SAMD: 1) Controller packaging, 2) lookahead transfer and 3) code adaptation. We have implemented SAMD on Android as a proof-of-concept prototype. Our experimental results demonstrate SAMD can provide fine-grained sharing of latency-insensitive applications.
Hochul Lee, Byoungjun Seo, Young Choon Lee, Hyuck Han, Sooyong Kang
PerCom5
2017 Platform Support for Mobile Edge Computing
abstract
Computing resources including mobile devices at the edge of a network are increasingly connected and capable of collaboratively processing what's believed to be too complex to them. Collaboration possibilities with today's feature-rich mobile devices go far beyond simple media content sharing, traditional video conferencing and cloud-based software as a services. The realization of these possibilities for mobile edge computing (MEC) requires non-trivial amounts of efforts in enabling multi-device resource sharing. The current practice of mobile collaborative application development remains largely at the application level. In this paper, we present CollaboRoid, a platform-level solution that provides a set of system services for mobile collaboration. CollaboRoid's platform-level design significantly eases the development of mobile collaborative applications promoting MEC. In particular, it abstracts the sharing of not only hardware resources, but also software resources and multimedia contents between multiple heterogeneous mobile devices. We implement CollaboRoid in the application framework layer of the Android stack and evaluate it with several collaboration scenarios on Nexus 5 and 7 devices. Our experimental results show the feasibility of the platform-level collaboration using CollaboRoid in terms of the latency and energy consumption.
Hochul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
CLOUD4
2017 AUTOBAHN: Accelerating Concurrent, Durable File I/O via a Non-volatile Buffer
abstract
As hardware vendors provision more cores and faster storage devices, attaining fast data durability for concurrent file writes is demanding to high-performance storage systems in cluster systems. We approach the challenge by proposing a system that uses a small amount of fast persistent memory for buffering concurrent file writes while preserving data durability. The main issue in designing a durable file buffer is allowing concurrent file writes to store data in a shared and limited space of persistent memory without incurring lock or resource contention. This paper addresses such issue and presents AUTOBAHN, a durable file buffer that expedites file I/O operations.
Sang Youp Rhee, Jae Eun Kim, Sooyong Kang, Hyuck Han, Hyungsoo Jung 0001
CLUSTER5
2017 SSD-Assisted Backup and Recovery for Database Systems
abstract
Backup and recovery is an important feature of database systems since it protects data against unexpected hardware and software failures. Database systems can provide data safety and reliability by creating a backup and restoring the backup from a failure. Database administrators can use backup/recovery tools that are provided with database systems or backup/recovery methods with operating systems. However, the existing tools perform time-consuming jobs and the existing methods may negatively affect run-time performance during normal operation even though high-performance SSDs are used. In this paper, we present an SSD-assisted backup/recovery scheme for database systems. In our scheme, we extend the out-of-place update characteristics of flash-based SSDs for backup/recovery operations. To this end, we exploit the resources (e.g., flash translation layer and DRAM cache with supercapacitors) inside SSDs, and we call our SSD with new backup/recovery features BR-SSD. We design and implement the backup/recovery functionality in the Samsung enterprise-class SSD (i.e., SM843Tn) for more realistic systems. Furthermore, we conduct a case study of BR-SSDs in replicated database systems and modify MySQL with replication to integrate BR-SSDs. The experimental result demonstrates that our scheme provides fast recovery while it does not negatively affect the run-time performance during normal operation.
Yongseok Son, Jaeyoon Choi, Jekyeom Jeon, Cheolgi Min, Sunggon Kim, Heon Young Yeom, Hyuck Han
ICDE7
2017 Scalable Database Logging for Multicores
abstract
Modern databases, guaranteeing atomicity and durability, store transaction logs in a volatile, central log buffer and then flush the log buffer to non-volatile storage by the write-ahead logging principle. Buffering logs in central log store has recently faced a severe multicore scalability problem, and log flushing has been challenged by synchronous I/O delay. We have designed and implemented a fast and scalable logging method, E leda , that can migrate a surge of transaction logs from volatile memory to stable storage without risking durable transaction atomicity. Our efficient implementation of E leda is enabled by a highly concurrent data structure, G rasshopper , that eliminates a multicore scalability problem of centralized logging and enhances system utilization in the presence of synchronous I/O delay. We implemented E leda and plugged it to WiredTiger and Shore-MT by replacing their log managers. Our evaluation showed that E leda -based transaction systems improve performance up to 71 x, thus showing the applicability of E leda.
Hyungsoo Jung 0001, Hyuck Han, Sooyong Kang
Proc. VLDB Endow.2
2017 Optimizing I/O Operations in File Systems for Fast Storage Devices
abstract
Fast non-volatile memory (NVM) technologies (e.g., phase change memory, spin-transfer torque memory, and MRAM) provide high performance to legacy storage systems. These NVM technologies have attractive features, such as low latency and high throughput to satisfy application performance. Accordingly, fast storage devices based on fast NVM lead to a rapid increase in the demand for diverse computer systems and environments (e.g., cloud platforms, web servers, and database systems) where they are expected to be used as primary storage. Despite the promised benefits provided by fast storage devices, modern file systems do not take advantage of the storage's full performance. In this article, we analyze and explore existing I/O strategies in read, write, journal I/ O, and recovery paths between the file system and the storage device. The analysis shows that existing I/O strategies are an obstacle to get maximum performance of fast storage devices. To address this issue, we propose efficient I/O strategies that enable file systems to fully exploit the performance of fast storage devices. Our main idea is to transfer requests from discontiguous host memory buffers in the file systems to discontiguous storage segments in one I/O request to get maximize I/O performance. We implemented our scheme to read, write, journal I/O and recovery operations in the EXT4 file system and the JBD2 module. We demonstrate the implication of our idea in terms of application performance through well-known benchmarks. The experimental results show that our optimized file system achieves better performance than the existing file system, with improvements of up to 1.54 ×, 1.96×, and 2.28× on ordered mode, data journaling mode, and recovery, respectively.
Yongseok Son, Heon Young Yeom, Hyuck Han
IEEE Trans. Computers3
2016 An Empirical Evaluation of Enterprise and SATA-Based Transactional Solid-State Drives
abstract
In most file systems, performance is usually sacrificed in exchange for crash consistency, which ensures that data and metadata are restored consistently in the event of a system crash. To escape this trade-off between performance and crash consistency, recent researchers designed and implemented the transactional functionality inside Solid State Drives (SSDs). However, in order to investigate its benefit in a more realistic and standard fashion, this scheme should be re-evaluated in enterprise storage with standard interface. This paper explores the challenges and implications of a transactional SSD with extensive experiments. To evaluate the potential benefit of transactional SSD, we design and implement the transaction functionality in Samsung enterprise-class and SATA-based SSD (i.e., SM843TN) and name it TxSSD. We then modify the existing file systems (i.e., ext4 and btrfs) on topof TxSSD, making both file systems crash-consistent without redundant writes. We perform performance evaluation of two filesystems by using file I/O and OLTP benchmarks with a database. We also disclose and analyze the overhead of transactional functionality inside SSD. The experimental results show that TxSSD-aware file systems exhibit better performance compared to crash-consistent modes (i.e., data journaling mode of ext4 and cow mode of btrfs) but worse performance compared to weak consistent modes (i.e., ordered mode of ext4 and no datacow mode of btrfs).
Yongseok Son, Hara Kang, Jinyong Ha 0001, Jongsung Lee 0001, Hyuck Han, Hyungsoo Jung 0001, Heon Young Yeom
MASCOTS5
2016 Efficient Memory-Mapped I/O on Fast Storage Device
abstract
In modern operating systems, memory-mapped I/O ( mmio ) is an important access method that maps a file or file-like resource to a region of memory. The mapping allows applications to access data from files through memory semantics (i.e., load/store) and it provides ease of programming. The number of applications that use mmio are increasing because memory semantics can provide better performance than file semantics (i.e., read/write). As more data are located in the main memory, the performance of applications can be enhanced owing to the effect of a large cache. When mmio is used, hot data tend to reside in the main memory and cold data are located in storage devices such as HDD and SSD; data placement in the memory hierarchy depends on the virtual memory subsystem of the operating system. Generally, the performance of storage devices has a direct impact on the performance of mmio . It is widely expected that better storage devices will lead to better performance. However, the expectation is limited when fast storage devices are used since the virtual memory subsystem does not reflect the performance feature of those devices. In this article, we examine the Linux virtual memory subsystem and mmio path to determine the influence of fast storage on the existing Linux kernel. Throughout our investigation, we find that the overhead of the Linux virtual memory subsystem, negligible on the HDD, prevents applications from using the full performance of fast storage devices. To reduce the overheads and fully exploit the fast storage devices, we present several optimization techniques. We modify the Linux kernel to implement our optimization techniques and evaluate our prototyped system with low-latency storage devices. Experimental results show that our optimized mmio has up to 7x better performance than the original mmio . We also compare our system to a system that has enough memory to keep all data in the main memory. The system with insufficient memory and our mmio achieves 92% performance of the resource-rich system. This result implies that our virtual memory subsystem for mmap can effectively extend the main memory with fast storage devices.
Nae Young Song, Yongseok Son, Hyuck Han, Heon Young Yeom
ACM Trans. Storage3
2015 Optimizing file systems for fast storage devices
abstract
Emerging high-performance storage devices have attractive features such as low latency and high throughput. This leads to a rapid increase in the demand for fast storage devices in cloud platforms, social network services, etc. However, there are few block-based file systems that are capable of utilizing superior characteristics of fast storage devices. In this paper, we find that the I/O strategy of modern operating systems prevents file systems from exploiting fast storage devices. To address this problem, we propose several optimization techniques for block-based file systems. Then, we apply our techniques to two well-known file systems and evaluate them with multiple benchmarks. The experimental results show that our optimized file systems achieve 32% on average and up to 54% better performance than existing file systems.
Yongseok Son, Hyuck Han, Heon Young Yeom
SYSTOR2
2015 Resource-efficient workflow scheduling in clouds
Young Choon Lee, Hyuck Han, Albert Y. Zomaya, Mazin Yousif
Knowl. Based Syst.2
2014 Scalable serializable snapshot isolation for multicore systems
abstract
Since 1990's, Snapshot Isolation (SI) has been widely studied, and it was successfully deployed in commercial and open-source database engines. Berenson et al. showed that data consistency can be violated under SI. Recently, a new class of Serializable SI algorithms (SSI) has been proposed to achieve serializable execution while still allowing concurrency between reads and updates.
Hyuck Han, Seongjae Park, Hyungsoo Jung 0001, Alan D. Fekete, Uwe Röhm, Heon Young Yeom
ICDE1
2014 Design and evaluation of mobile offloading system for web-centric devices
Sehoon Park, Qichen Chen, Hyuck Han, Heon Young Yeom
J. Netw. Comput. Appl.3
2014 A Scalable Lock Manager for Multicores
abstract
Modern implementations of DBMS software are intended to take advantage of high core counts that are becoming common in high-end servers. However, we have observed that several database platforms, including MySQL, Shore-MT, and a commercial system, exhibit throughput collapse as load increases into oversaturation (where there are more request threads than cores), even for a workload with little or no logical contention for locks, such as a read-only workload. Our analysis of MySQL identifies latch contention within the lock manager as the bottleneck responsible for this collapse. We design a lock manager with reduced latching, implement it in MySQL, and show that it avoids the collapse and generally improves performance. Our efficient implementation of a lock manager is enabled by a staged allocation and deallocation of locks. Locks are preallocated in bulk, so that the lock manager only has to perform simple list manipulation operations during the acquire and release phases of a transaction. Deallocation of the lock data structures is also performed in bulk, which enables the use of fast implementations of lock acquisition and release as well as concurrent deadlock checking.
Hyungsoo Jung 0001, Hyuck Han, Alan D. Fekete, Gernot Heiser, Heon Young Yeom
ACM Trans. Database Syst.2
2013 Performance of Serializable Snapshot Isolation on Multicore Servers
Hyungsoo Jung 0001, Hyuck Han, Alan D. Fekete, Uwe Röhm, Heon Young Yeom
DASFAA (2)2
2013 A scalable lock manager for multicores
abstract
Modern implementations of DBMS software are intended to take advantage of high core counts that are becoming common in high-end servers. However, we have observed that several database platforms, including MySQL, Shore-MT, and a commercial system, exhibit throughput collapse as load increases, even for a workload with little or no logical contention for locks. Our analysis of MySQL identifies latch contention within the lock manager as the bottleneck responsible for this collapse.
Hyungsoo Jung 0001, Hyuck Han, Alan D. Fekete, Gernot Heiser, Heon Young Yeom
SIGMOD Conference2
2012 Cashing in on the Cache in the Cloud
abstract
Over the past decades, caching has become the key technology used for bridging the performance gap across memory hierarchies via temporal or spatial localities; in particular, the effect is prominent in disk storage systems. Applications that involve heavy I/O activities, which are common in the cloud, probably benefit the most from caching. The use of local volatile memory as cache might be a natural alternative, but many well-known restrictions, such as capacity and the utilization of host machines, hinder its effective use. In addition to technical challenges, providing cache services in clouds encounters a major practical issue (quality of service or service level agreement issue) of pricing. Currently, (public) cloud users are limited to a small set of uniform and coarse-grained service offerings, such as High-Memory and High-CPU in Amazon EC2. In this paper, we present the cache as a service (CaaS) model as an optional service to typical infrastructure service offerings. Specifically, the cloud provider sets aside a large pool of memory that can be dynamically partitioned and allocated to standard infrastructure services as disk cache. We first investigate the feasibility of providing CaaS with the proof-of-concept elastic cache system (using dedicated remote memory servers) built and validated on the actual system, and practical benefits of CaaS for both users and providers (i.e., performance and profit, respectively) are thoroughly studied with a novel pricing scheme. Our CaaS model helps to leverage the cloud economy greatly in that 1) the extra user cost for I/O performance gain is minimal if ever exists, and 2) the provider's profit increases due to improvements in server consolidation resulting from that performance gain. Through extensive experiments with eight resource allocation strategies, we demonstrate that our CaaS model can be a promising cost-efficient solution for both users and providers.
Hyuck Han, Young Choon Lee, Woong Shin, Hyungsoo Jung 0001, Heon Young Yeom, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2011 An efficient skyline framework for matchmaking applications
Hyuck Han, Hyungsoo Jung 0001, Hyeonsang Eom, Heon Young Yeom
J. Netw. Comput. Appl.1
2011 Serializable Snapshot Isolation for Replicated Databases in High-Update Scenarios
Hyungsoo Jung 0001, Hyuck Han, Alan D. Fekete, Uwe Röhm
Proc. VLDB Endow.2
2011 Athanasia: A User-Transparent and Fault-Tolerant System for Parallel Applications
abstract
This article presents Athanasia, a user-transparent and fault-tolerant system, for parallel applications running on large-scale cluster systems. Cluster systems have been regarded as a de facto standard to achieve multitera-flop computing power. These cluster systems, as we know, have an inherent failure factor that can cause computation failure. The reliability issue in parallel computing systems, therefore, has been studied for a relatively long time in the literature, and we have seen many theoretical promises arise from the extensive research. However, despite the rigorous studies, practical and easily deployable fault-tolerant systems have not been successfully adopted commercially. Athanasia is a user-transparent checkpointing system for a fault-tolerant Message Passing Interface (MPI) implementation that is primarily based on the sync-and-stop protocol. Athanasia supports three critical functionalities that are necessary for fault tolerance: a light-weight failure detection mechanism, dynamic process management that includes process migration, and a consistent checkpoint and recovery mechanism. The main features of Athanasia are that it does not require any modifications to the application code and that it preserves many of the high performance characteristics of high-speed networks. Experimental results show that Athanasia can be a good candidate for practically deployable fault-tolerant systems in very-large and high-performance clusters and that its protocol can be applied to a variety of parallel communication libraries easily.
Hyungsoo Jung 0001, Hyuck Han, Heon Young Yeom, Sooyong Kang
IEEE Trans. Parallel Distributed Syst.2
2010 Large Graph Processing Based on Remote Memory System
abstract
This paper focuses on large graph processing based on the remote memory system. Using our remote memory system enables applications to deal with large data sets, especially graph data, which do not fit into the machines main memory. Although recent dramatic increases in DRAM capacity now allow us to build inexpensive computers with very large amounts of main memory, the rise in brand-new Internet services has resulted in rapid increases in data size. This is especially true for on-line social network services that generate various data sets that can be represented as graphs. On the other hand, high-speed networking technologies such as Infini Band, Myrinet and 10G Ethernet now enable us to transfer data with low latency and high throughput. The advanced networking technologies reduce the latency/bandwidth gap between main memory and remote memory. Thus, remote memory based processing could now be helpful in accelerating large-scale graph process when main memory space is insufficient to store application data. In this paper, we present our design and implementation of remote memory system that efficiently processes large graph data. We also evaluate a breadth-first search of various types of graphs using our system and show that our approach is good for large graph data processing.
Kyungho Jeon, Hyuck Han, Shin Gyu Kim, Hyeonsang Eom, Heon Young Yeom
HPCC2
2010 A fast and progressive algorithm for skyline queries with totally- and partially-ordered domains
Hyungsoo Jung 0001, Hyuck Han, Heon Young Yeom, Sooyong Kang
J. Syst. Softw.2
2009 A RESTful Approach to the Management of Cloud Infrastructure
abstract
Recently, REpresentational State Transfer (REST) has been proposed as an alternative architecture for Web services.In the era of Cloud and Web 2.0, many complex Web service-based systems such as e-Business an de-Government applications have adopted REST. Unfortunately, the REST approach has been applied to few cases in management systems, especially for a management system for cloud computing infrastructures.In this paper, we design and implement a RESTful Cloud Management System (CMS).Managed elements can be modeled as resources in REST and operations in existing systems can be evaluated using four methods of REST or a combination of them.We also show how components of existing management systems can be realized as REST-style Web services.
Hyuck Han, Shin Gyu Kim, Hyungsoo Jung 0001, Heon Young Yeom, Changho Yoon, Jong-Won Park
IEEE CLOUD1
2009 Intelligent Management of Remote Facilities through a Ubiquitous Cloud Middleware
abstract
This paper introduces a tele-management system as a part of SmartUM which is a ubiquitous cloud middleware for ubiquitous city (u-city). The cloud computing platform allows users to control remote devices. The users get data from a various kinds of remote sensors and scene images about the place of sensors from remote video cameras and control remote devices seeing the scene images of the remote place. Our cloud computing platform has context-awareness and can intelligently control the remote devices according to the circumstance scenario. We used ontology for the context aware intelligence processing.
Chang-Ho Yun, Hyuck Han, Hae-Sun Jung, Heon Young Yeom
IEEE CLOUD2
2009 A Skyline Approach to the Matchmaking Web Service
abstract
Item matchmaking that finds items for users is an essential service framework in the web service infrastructure. The current way of carrying out the matchmaking procedure is the selection of items based on a user's specifications. We rethink the item matchmaking framework in such a way that a matchmaker can find items that can satisfy a specific computing demand from a user and recommend a collection of better items candidates among the identified items. This endows a user with the right of choice on deciding best-possible items. We approach the problem in the view of skyline query processing that has become one of the major topics in the database community, and present the efficient skyline algorithm that gathers interesting item candidates efficiently. To this end, we adopt (i) lattice-based indexing using a lattice composition technique,and (ii) an optimized dominance-check algorithm. Our extensive experimental results show that our algorithm outperforms the current state-of-the-art algorithm.
Hyuck Han, Hyungsoo Jung 0001, Shin Gyu Kim, Heon Young Yeom
CCGRID1
2008 MRBench: A Benchmark for MapReduce Framework
abstract
MapReduce is Google's programming model for easy development of scalable parallel applications which process huge quantity of data on many clusters. Due to its conveniency and efficiency, MapReduce is used in various applications (e.g., Web search services and online analytical processing). However, there are only few good benchmarks to evaluate MapReduce implementations by realistic testsets. In this paper, we present MRBench that is a benchmark for evaluating MapReduce systems. MRBench focuses on processing business oriented queries and concurrent data modifications. To this end, we build MRBench to deal with large volumes of relational data and execute highly complex queries. By MRBench, users can evaluate the performance of MapReduce systems while varying environmental parameters such as data size and the number of (map/reduce) tasks. Our extensive experimental results show that MRBench is a useful tool to benchmark the capability of answering critical business questions.
Kiyoung Kim, Kyungho Jeon, Hyuck Han, Shin Gyu Kim, Hyungsoo Jung 0001, Heon Young Yeom
ICPADS3
2007 Taste of AOP : Blending concerns in cluster computing software
abstract
Pioneering work on Aspect Oriented Programming (AOP) has not flourished enough to enrich the design of distributed systems with the refined AOP paradigm. The more generous perspective today is that a decade of growing research on AOP has brought the paradigm into many exciting areas. We investigate two case studies that cover time-honored issues, fault tolerant computing and parallel computing, in the cluster computing world using the AOP paradigm. Aspects that we define here are simple, intuitive and reusable. We believe that our implementation is very useful in developing other cluster computing software, and AOP can be a powerful method in modularizing source codes.
Hyuck Han, Hyungsoo Jung 0001, Heon Young Yeom, Dong-Young Lee
CLUSTER1
2006 HVEM Grid: Experiences in Constructing an Electron Microscopy Grid
Hyuck Han, Hyungsoo Jung 0001, Heon Young Yeom, Hee S. Kweon, Jysoo Lee
APWeb1
2006 COEDIG: Collaborative Editor in Grid Computing
Hyunjoon Jung, Hyuck Han, Heon Young Yeom, Hee-Jae Park, Jysoo Lee
APWeb2
2006 Practical Fault-Tolerant Framework for eScience Infrastructure
abstract
Many areas of science currently use computing resources as a important part of their research, and many research groups adopt cluster architecture to use them efficiently and manage them easily. Therefore, faulttolerance becomes a very important property for the computing resources. However, fault-tolerant systems have not yet been widely adopted because they are either hard to deploy, hard to use, hard to manage, hard to maintain, or hard to justify. This paper proposes a practical fault-tolerant system for eScience infrastructures. Our system uses checkpoint/ restart mechanism for fault-tolerance, and provides a easy mechanism to integrate with Grid services widely used in eScience. Additionally, we run rigorous tests using scientific applications to verify that our system can be used in clusters. We also describe improvements made to our system to solve various problems that arose when deploying it on a cluster. The experimental results show that not only does our system conform to various types of running environment well, but that it can also be practically deployed in clusters.
Hyuck Han, Jai Wug Kim, Jongpil Lee, Youngjin Yu, Kiyoung Kim, Heon Young Yeom
e-Science1
2006 SHIELD: A Fault-Tolerant MPI for an Infiniband Cluster
Hyuck Han, Hyungsoo Jung 0001, Jai Wug Kim, Jongpil Lee, Youngjin Yu, Shin Gyu Kim, Heon Young Yeom
HPCC1
2005 Design and Implementation of Multiple Fault-Tolerant MPI over Myrinet (M^3)
abstract
Advances in network technology and computing power have inspired the emergence of high-performance cluster computing systems. While cluster management and hardware highavailability tools are readily available, practical and easily deployable fault-tolerant systems have not been successfully adopted commercially. We present a fault-tolerant system, Multiple fault-tolerant MPI over Myrinet (M3), that differs in notable respects from other proposed fault-tolerant systems in the literature. M3 is built on top of Myrinet since it is regarded as one of the best solutions for highperformance networks and is widely used in cluster computing systems because it can provide a high-speed switching network that is an inevitable ingredient in interconnecting clusters of workstations or PCs. M^3 is a user-transparent checkpointing system for multiple fault-tolerant MPI implementation that is primarily based on the coordinated checkpointing protocol. M3 supports three critical functionalities that are necessary for faulttolerance: a light-weight failure detection mechanism, dynamic process management that includes process migration, and a consistent checkpoint and recovery mechanism. The features of M are that it requires no modifications of application code and that it preserves much of the high performance characteristics of Myrinet. This paper describes the architecture of M3, its detailed design principles and comprehensive implementation issues. We also propose practical solutions for those involved in constructing highly available cluster systems for parallel programming systems. Experimental results substantiate our assertion that M3 can be a good candidate for practically deployable fault-tolerant systems in very-large and high-performance Myrinet clusters and that its protocol can be applied to a wide variety of parallel communication libraries without difficulty.
Hyungsoo Jung 0001, Dongin Shin, Hyuck Han, Jai Wug Kim, Heon Young Yeom, Jongsuk Lee
SC3