Hwanju Kim

dblp:48/3659 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Memory systems · 37% Cloud and datacenter computing · 35% Storage systems · 13%
Software engineering, system software, and programming languages
5 papers
Operating systems · 100%

Topics — the 30 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
virtualization
0.862016
vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Memory systems
cache coherence
0.422016
Virtual Snooping Coherence for Multi-Core Virtualized Systems · IEEE Trans. Parallel Distributed Syst. 2016
Virtual Snooping: Filtering Snoops in Virtualized Multi-cores · MICRO 2010
Storage systems › i/o architecture › i/o subsystem
i/o path
0.312017
Enlightening the I/O Path: A Holistic Approach for Application Performance · FAST 2017
Operating systems › resource management
memory management
0.212016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Cloud and datacenter computing › resource management
datacenter resource management
0.212016
TPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive Services · ASPLOS 2016
Storage systems
file systems
0.212016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Memory systems › cache management › storage caching
page cache
0.212016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Memory systems › cache coherence › cache coherence protocol
snoopy coherence
0.212016
Virtual Snooping Coherence for Multi-Core Virtualized Systems · IEEE Trans. Parallel Distributed Syst. 2016
Memory systems
cache
0.212015
vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015
Memory systems › cache management
cache isolation
0.212015
vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.212015
vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015
Cloud and datacenter computing › resource management
resource isolation
0.212015
vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015
Memory systems › cache › cache organization
write cache
0.212015
Request-Oriented Durable Write Caching for Application Performance · USENIX ATC 2015
Parallel and multicore computing › task scheduling
coordinated scheduling
0.212013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
Electronic design automation › high-level synthesis
scheduling
0.212013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine scheduling
0.212013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
Storage systems › buffer management
buffer cache management
0.112011
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Memory systems › cache management › storage caching
cooperative caching
0.112011
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Embedded and real-time systems
device drivers
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Distributed systems › fault tolerance
fault detection and recovery
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Distributed systems
fault tolerance
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Cloud and datacenter computing › virtualization
virtual machine
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Cloud and datacenter computing › virtualization › virtualization security
virtual machine isolation
0.112010
Virtual Snooping: Filtering Snoops in Virtualized Multi-cores · MICRO 2010
Operating systems › i/o › i/o subsystem
i/o scheduling
0.112017
Enlightening the I/O Path: A Holistic Approach for Application Performance · FAST 2017
Cloud and datacenter computing › datacenter services › online service systems › internet services
interactive services
0.112016
TPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive Services · ASPLOS 2016
Embedded and real-time systems
mobile computing
0.112016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Memory systems
memory management
0.012011
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation
0.012011
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Operating systems › i/o › i/o subsystem
device drivers
0.012010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Operating systems
i/o
0.012010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010

Methods — techniques the papers use, named apart from their topics

prototype implementation · 0.5cost-based region selection · 0.5request-oriented caching · 0.4durability guarantees · 0.4target-driven parallelism · 0.2execution time prediction · 0.2emulation · 0.2dynamic correction · 0.2guest physical address indexing · 0.2architectural support · 0.2demand-based scheduling · 0.2communication-aware scheduling · 0.2progress monitoring · 0.1
YearPublicationVenuePosition
2017 Enlightening the I/O Path: A Holistic Approach for Application Performance
Hwanju Kim, Joonwon Lee, Jinkyu Jeong
FAST2
2016 TPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive Services
abstract
In interactive services such as web search, recommendations, games and finance, reducing the tail latency is crucial to provide fast response to every user. Using web search as a driving example, we systematically characterize interactive workload to identify the opportunities and challenges for reducing tail latency. We find that the workload consists of mainly short requests that do not benefit from parallelism, and a few long requests which significantly impact the tail but exhibit high parallelism speedup. This motivates estimating request execution time, using a predictor, to identify long requests and to parallelize them. Prediction, however, is not perfect; a long request mispredicted as short is likely to contribute to the server tail latency, setting a ceiling on the achievable tail latency. We propose TPC, an approach that combines prediction information judiciously with dynamic correction for inaccurate prediction. Dynamic correction increases parallelism to accelerate a long request that is mispredicted as short. TPC carefully selects the appropriate target latencies based on system load and parallelism efficiency to reduce tail latency.
Myeongjae Jeon, Yuxiong He, Hwanju Kim, Sameh Elnikety, Scott Rixner, Alan L. Cox
ASPLOS3
2016 Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems
abstract
Most embedded systems require contiguous memory space to be reserved for devices, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses. Our scheme, on the other hand, utilizes reserved memory as an eviction-based file cache. It guarantees contiguous memory allocation to devices while providing idle device memory as an additional file cache called eCache for general-purpose usage. Because eCache stores only evicted data from the in-kernel page cache, the memory efficiency is preserved and the allocation time for devices is minimized. Cost-based region selection also minimizes additional read I/O operations by carefully discarding cached data from eCache. The additional indexing cost incurred by adding eCache is minimized by integrating its index structure with the kernel page cache. The prototype is implemented on the Nexus S smartphone and is evaluated using popular Android applications. The evaluation results show that our scheme outperforms previous approaches in terms of the application launch performance. The device memory reallocation time is also limited to a few milliseconds, which is sufficiently small to make our scheme transparent.
Jinkyu Jeong, Hwanju Kim, Joonwon Lee
IEEE Trans. Mob. Comput.2
2016 Virtual Snooping Coherence for Multi-Core Virtualized Systems
abstract
Proliferation of virtualized systems opens a new opportunity to improve the scalability of multi-core architectures. Among the scalability bottlenecks in multi-cores, cache coherence has been one of the most critical problems. Although snoop-based protocols have been dominating commercial multi-core designs, it has been difficult to scale them for more cores, as snooping protocols require high network bandwidth and power consumption for snooping all the caches.In this paper, we propose a novel snoop-based cache coherence protocol, called virtual snooping, for virtualized multi-core architectures. Virtual snooping exploits memory isolation across virtual machines and prevents unnecessary snoop requests from crossing the virtual machine boundaries. Each virtual machine becomes a virtual snoop domain, consisting of a subset of the cores in a system. Although the majority of virtual machine memory is isolated, sharing of cachelines across VMs still occur. To address such data sharing, this paper investigates three factors, data sharing through the hypervisor, virtual machine relocation, and content-based sharing. In this paper, we explore the design space of virtual snooping with experiments on emulated and real virtualized systems including the mechanisms and overheads of the hypervisor. In addition, the paper discusses the scheduling impact on the effectiveness of virtual snooping.
Daehoon Kim 0001, Chang Hyun Park 0001, Hwanju Kim, Jaehyuk Huh 0001
IEEE Trans. Parallel Distributed Syst.3
2015 vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments
abstract
A key role of virtualization is to give an illusion that a consolidated workload runs on a dedicated machine although the underlying resources are actively shared by multiple workloads. Technical advances have enabled a virtual machine (VM) to exercise many shared resources of a machine in a transparent and isolated manner. However, such an illusion of resource dedication has not been supported for the last-level cache (LLC), although the LLC is the largest on-chip shared resource with a significant performance impact. In this paper, we propose vCache--architectural support to provide a transparent and isolated virtual LLC (vLLC) for each VM and interfaces to manage the vLLC. More specifically, this study first proposes architectural support for the guest OS of a VM to index the LLC with its guest physical address instead of a host physical address. This in turn allows that the guest OS transparently view its vLLC and preserve the effectiveness of its page placement policy. Second, this study extends the architectural support for each VM to keep its vLLC strongly isolated from other VMs. Such resource dedication is critical to offer performance isolation and preserve vLLC transparency for each VM in a highly consolidated machine. With little hardware overhead, vCache can facilitate various unchartered vLLC capacity-based services for the public clouds while providing up to 17% higher performance than a traditional virtualized system.
Daehoon Kim 0001, Hwanju Kim, Nam Sung Kim, Jaehyuk Huh 0001
MICRO2
2015 Request-Oriented Durable Write Caching for Application Performance
Hwanju Kim, Sang-Hoon Kim, Joonwon Lee, Jinkyu Jeong
USENIX ATC2
2014 Virtual asymmetric multiprocessor for interactive performance of consolidated desktops
abstract
This paper presents virtual asymmetric multiprocessor, a new scheme of virtual desktop scheduling on multi-core processors for user-interactive performance. The proposed scheme enables virtual CPUs to be dynamically performance-asymmetric based on their hosted workloads. To enhance user experience on consolidated desktops, our scheme provides interactive workloads with fast virtual CPUs, which have more computing power than those hosting background workloads in the same virtual machine. To this end, we devise a hypervisor extension that transparently classifies background tasks from potentially interactive workloads. In addition, we introduce a guest extension that manipulates the scheduling policy of an operating system in favor of our hypervisor-level scheme so that interactive performance can be further improved. Our evaluation shows that the proposed scheme significantly improves interactive performance of application launch, Web browsing, and video playback applications when CPU-intensive workloads highly disturb the interactive workloads.
Hwanju Kim, Jinkyu Jeong, Joonwon Lee
VEE1
2014 Group-based memory oversubscription for virtualized clouds
Hwanju Kim, Joonwon Lee, Jinkyu Jeong
J. Parallel Distributed Comput.2
2013 Demand-based coordinated scheduling for SMP VMs
abstract
As processor architectures have been enhancing their computing capacity by increasing core counts, independent workloads can be consolidated on a single node for the sake of high resource efficiency in data centers. With the prevalence of virtualization technology, each individual workload can be hosted on a virtual machine for strong isolation between co-located workloads. Along with this trend, hosted applications have increasingly been multithreaded to take advantage of improved hardware parallelism. Although the performance of many multithreaded applications highly depends on communication (or synchronization) latency, existing schemes of virtual machine scheduling do not explicitly coordinate virtual CPUs based on their communication behaviors.
Hwanju Kim, Jinkyu Jeong, Joonwon Lee, Seung Ryoul Maeng
ASPLOS1
2013 Rigorous rental memory management for embedded systems
abstract
Memory reservation in embedded systems is a prevalent approach to provide a physically contiguous memory region to its integrated devices, such as a camera device and a video decoder. Inefficiency of the memory reservation becomes a more significant problem in emerging embedded systems, such as smartphones and smart TVs. Many ways of using these systems increase the idle time of their integrated devices, and eventually decrease the utilization of their reserved memory. In this article, we propose a scheme to minimize the memory inefficiency caused by the memory reservation. The memory space reserved for a device can be rented for other purposes when the device is not active. For this scheme to be viable, latencies associated with reallocating the memory space should be minimal. Volatile pages are good candidates for such page reallocation since they can be reclaimed immediately as they are needed by the original device. We also provide two optimization techniques, lazy-migration and adaptive-activation. The former increases the lowered utilization of the rental memory by our volatile page allocations, and the latter saves active pages in the rental memory during the reallocation. We implemented our scheme on a smartphone development board with the Android Linux kernel. Our prototype has shown that the time for the return operation is less than 0.77 seconds in the tested cases. We believe that this time is acceptable to end-users in terms of transparency since the time can be hidden in application initialization time. The rental memory also brings throughput increases ranging from 2% to 200% based on the available memory and the applications' memory intensiveness.
Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
ACM Trans. Embed. Comput. Syst.2
2013 Analysis of virtual machine live-migration as a method for power-capping
Jinkyu Jeong, Sung-hun Kim 0005, Hwanju Kim, Joonwon Lee, Euiseong Seo
J. Supercomput.3
2012 DaaC: device-reserved memory as an eviction-based file cache
abstract
Most embedded systems require contiguous memory space to be reserved for each device, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses.
Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
CASES2
2012 Scheduler support for video-oriented multimedia on client-side virtualization
abstract
Virtualization has recently been adopted for client devices to provide strong isolation between services and efficient manageability. Even though multimedia service is not rare for the devices, the virtual machine hosting this service is not guaranteed to receive proper scheduling support from the underlying hypervisor. The quality of multimedia service is often compromised when several virtual machines compete for computing power. This paper presents a new scheduling scheme for the hypervisor to transparently identify if the workload handles multimedia and to provide proper scheduling supports. An implementation of our scheme has shown that the virtual machine hosting a video-oriented application receives propoer CPU scheduling even when other virtual machines host CPU intensive workloads.
Hwanju Kim, Jinkyu Jeong, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
MMSys1
2011 Transparently bridging semantic gap in CPU management for virtualized environments
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee, Seung Ryoul Maeng
J. Parallel Distributed Comput.1
2011 XHive: Efficient Cooperative Caching for Virtual Machines
abstract
Since a virtual machine independently uses its own caching policy, redundant disk operations exacerbate the I/O virtualization overhead when virtual machines access large amounts of data on shared storage. This paper presents XHive, an efficient cooperative caching system that is implemented at the virtualization layer, for consolidated environments. Our proposed scheme globally manages buffer caches of consolidated virtual machines in order to accommodate a shared working set in machine memory. A singlet, which is a block cached solely by a virtual machine, is preferentially given more chances to be cached in machine memory by XHive, when it is evicted by a guest operating system. For efficient use of limited memory, singlets are cached in memory that is collaboratively donated from idle memory of virtual machines. Our evaluation shows that XHive significantly reduces disk I/O operations for shared working sets, thereby achieving high read performance and scalability. Improved scalability enables a high degree of workload consolidation with respect to virtual machines that have shared working sets.
Hwanju Kim, Heeseung Jo, Joonwon Lee
IEEE Trans. Computers1
2010 Virtual Snooping: Filtering Snoops in Virtualized Multi-cores
abstract
Virtualization has been rapidly expanding its applications in numerous server and desktop environments to improve the utilization and manageability of physical systems. Such proliferation of virtualized systems opens a new opportunity to improve the scalability of future multi-core architectures. Among the scalability bottlenecks in multi-cores, cache coherence has been a critical problem. Although snoop-based protocols have been dominating commercial multi-core designs, it has been difficult to scale them for more cores, as snooping protocols require high network bandwidth and power consumption for snooping all the caches. In this paper, we propose a novel snoop-based cache coherence protocol, called virtual snooping, for virtualized multi-core architectures. Virtual snooping exploits memory isolation across virtual machines and prevents unnecessary snoop requests from crossing the virtual machine boundaries. Each virtual machine becomes a virtual snoop domain, consisting of a subset of the cores in a system. However, in real virtualized systems, virtual machines cannot partition the cores perfectly without any data sharing across the snoop partitions. This paper investigates three factors, which break the memory isolation among virtual machines: data sharing with a hyper visor, virtual machine relocation, and content-based data sharing. In this paper, we explore the design space of virtual snooping with experiments on real virtualized systems and approximate simulations. The results show that virtual snooping can reduce snoops significantly even if virtual machines migrate frequently. We also propose mechanisms to address content-based data sharing by exploiting its read-only property.
Daehoon Kim 0001, Hwanju Kim, Jaehyuk Huh 0001
MICRO2
2010 KAL: kernel-assisted non-invasive memory leak tolerance with a general-purpose memory allocator
abstract
Abstract Memory leaks are a continuing problem in the software developed with programming languages, such as C and C++. A recent approach adopted by some researchers is to tolerate leaks in the software application and to reclaim the leaked memory by use of specially constructed memory allocation routines. However, such routines replace the usual general‐purpose memory allocator and tend to be less efficient in speed and in memory utilization. We propose a new scheme which coexists with the existing memory allocation routines and which reclaims memory leaks. Our scheme identifies and reclaims leaked memory at the kernel level. There are some major advantages to our approach: (1) the application software does not need to be modified; (2) the application does not need to be suspended while leaked memory is reclaimed; (3) a remote host can be used to identify the leaked memory, thus minimizing impact on the application program's performance; and (4) our scheme does not degrade the service availability of the application while detecting and reclaiming memory leaks. We have implemented a prototype that works with the GNU C library and with the Linux kernel. Our prototype has been tested and evaluated with various real‐world applications. Our results show that the computational overhead of our approach is around 2% of that incurred by the conventional memory allocator in terms of throughput and average response time. We also verified that the prototype successfully suppressed address space expansion caused by memory leaks when the applications are run on synthetic workloads. Copyright © 2010 John Wiley & Sons, Ltd.
Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, Joonwon Lee
Softw. Pract. Exp.4
2010 Transparent Fault Tolerance of Device Drivers for Virtual Machines
abstract
In a consolidated server system using virtualization, physical device accesses from guest virtual machines (VMs) need to be coordinated. In this environment, a separate driver VM is usually assigned to this task to enhance reliability and to reuse existing device drivers. This driver VM needs to be highly reliable, since it handles all the I/O requests. This paper describes a mechanism to detect and recover the driver VM from faults to enhance the reliability of the whole system. The proposed mechanism is transparent in that guest VMs cannot recognize the fault and the driver VM can recover and continue its I/O operations. Our mechanism provides a progress monitoring-based fault detection that is isolated from fault contamination with low monitoring overhead. When a fault occurs, the system recovers by switching the faulted driver VM to another one. The recovery is performed without service disconnection or data loss and with negligible delay by fully exploiting the I/O structure of the virtualized system.
Heeseung Jo, Hwanju Kim, Jae-Wan Jang, Joonwon Lee, Seung Ryoul Maeng
IEEE Trans. Computers2
2009 Task-aware virtual machine scheduling for I/O performance
abstract
The use of virtualization is progressively accommodating diverse and unpredictable workloads as being adopted in virtual desktop and cloud computing environments. Since a virtual machine monitor lacks knowledge of each virtual machine, the unpredictableness of workloads makes resource allocation difficult. Particularly, virtual machine scheduling has a critical impact on I/O performance in cases where the virtual machine monitor is agnostic about the internal workloads of virtual machines. This paper presents a task-aware virtual machine scheduling mechanism based on inference techniques using gray-box knowledge. The proposed mechanism infers the I/O-boundness of guest-level tasks and correlates incoming events with I/O-bound tasks. With this information, we introduce partial boosting, which is a priority boosting mechanism with task-level granularity, so that an I/O-bound task is selectively scheduled to handle its incoming events promptly. Our technique focuses on improving the performance of I/O-bound tasks within heterogeneous workloads by lightweight mechanisms with complete CPU fairness among virtual machines. All implementation is confined to the virtualization layer based on the Xen virtual machine monitor and the credit scheduler. We evaluate our prototype in terms of I/O performance and CPU fairness over synthetic mixed workloads and realistic applications.
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee
VEE1
2008 Guest-Aware Priority-Based Virtual Machine Scheduling for Highly Consolidated Server
Hwanju Kim, Myeongjae Jeon, Euiseong Seo, Joonwon Lee
Euro-Par2