EDBT 2026 Demo / reviewers in the wild / expert
Hwanju Kim
dblp:48/3659
· DBLP profile ↗
20ranked-venue papers
6as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Memory systems · 37% Cloud and datacenter computing · 35% Storage systems · 13% | |
| Software engineering, system software, and programming languages
5 papers |
Operating systems · 100% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
virtualization |
0.8 | 6 | 2016 | vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015 Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011 |
Memory systems
cache coherence |
0.4 | 2 | 2016 | Virtual Snooping Coherence for Multi-Core Virtualized Systems · IEEE Trans. Parallel Distributed Syst. 2016 Virtual Snooping: Filtering Snoops in Virtualized Multi-cores · MICRO 2010 |
Storage systems › i/o architecture › i/o subsystem
i/o path |
0.3 | 1 | 2017 | Enlightening the I/O Path: A Holistic Approach for Application Performance · FAST 2017 |
Operating systems › resource management
memory management |
0.2 | 1 | 2016 | Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016 |
Cloud and datacenter computing › resource management
datacenter resource management |
0.2 | 1 | 2016 | TPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive Services · ASPLOS 2016 |
Storage systems
file systems |
0.2 | 1 | 2016 | Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016 |
Memory systems › cache management › storage caching
page cache |
0.2 | 1 | 2016 | Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016 |
Memory systems › cache coherence › cache coherence protocol
snoopy coherence |
0.2 | 1 | 2016 | Virtual Snooping Coherence for Multi-Core Virtualized Systems · IEEE Trans. Parallel Distributed Syst. 2016 |
Memory systems
cache |
0.2 | 1 | 2015 | vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015 |
Memory systems › cache management
cache isolation |
0.2 | 1 | 2015 | vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.2 | 1 | 2015 | vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015 |
Cloud and datacenter computing › resource management
resource isolation |
0.2 | 1 | 2015 | vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments · MICRO 2015 |
Memory systems › cache › cache organization
write cache |
0.2 | 1 | 2015 | Request-Oriented Durable Write Caching for Application Performance · USENIX ATC 2015 |
Parallel and multicore computing › task scheduling
coordinated scheduling |
0.2 | 1 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 |
Electronic design automation › high-level synthesis
scheduling |
0.2 | 1 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 |
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine scheduling |
0.2 | 1 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 |
Storage systems › buffer management
buffer cache management |
0.1 | 1 | 2011 | XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011 |
Memory systems › cache management › storage caching
cooperative caching |
0.1 | 1 | 2011 | XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011 |
Embedded and real-time systems
device drivers |
0.1 | 1 | 2010 | Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Distributed systems › fault tolerance
fault detection and recovery |
0.1 | 1 | 2010 | Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Distributed systems
fault tolerance |
0.1 | 1 | 2010 | Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Cloud and datacenter computing › virtualization
virtual machine |
0.1 | 1 | 2010 | Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Cloud and datacenter computing › virtualization › virtualization security
virtual machine isolation |
0.1 | 1 | 2010 | Virtual Snooping: Filtering Snoops in Virtualized Multi-cores · MICRO 2010 |
Operating systems › i/o › i/o subsystem
i/o scheduling |
0.1 | 1 | 2017 | Enlightening the I/O Path: A Holistic Approach for Application Performance · FAST 2017 |
Cloud and datacenter computing › datacenter services › online service systems › internet services
interactive services |
0.1 | 1 | 2016 | TPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive Services · ASPLOS 2016 |
Embedded and real-time systems
mobile computing |
0.1 | 1 | 2016 | Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016 |
Memory systems
memory management |
0.0 | 1 | 2011 | XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011 |
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation |
0.0 | 1 | 2011 | XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011 |
Operating systems › i/o › i/o subsystem
device drivers |
0.0 | 1 | 2010 | Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Operating systems
i/o |
0.0 | 1 | 2010 | Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Methods — techniques the papers use, named apart from their topics
prototype implementation · 0.5cost-based region selection · 0.5request-oriented caching · 0.4durability guarantees · 0.4target-driven parallelism · 0.2execution time prediction · 0.2emulation · 0.2dynamic correction · 0.2guest physical address indexing · 0.2architectural support · 0.2demand-based scheduling · 0.2communication-aware scheduling · 0.2progress monitoring · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Enlightening the I/O Path: A Holistic Approach for Application Performance
Hwanju Kim, Joonwon Lee, Jinkyu Jeong |
FAST | 2 |
| 2016 | TPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive ServicesabstractIn interactive services such as web search, recommendations, games and finance, reducing the tail latency is crucial to provide fast response to every user. Using web search as a driving example, we systematically characterize interactive workload to identify the opportunities and challenges for reducing tail latency. We find that the workload consists of mainly short requests that do not benefit from parallelism, and a few long requests which significantly impact the tail but exhibit high parallelism speedup. This motivates estimating request execution time, using a predictor, to identify long requests and to parallelize them. Prediction, however, is not perfect; a long request mispredicted as short is likely to contribute to the server tail latency, setting a ceiling on the achievable tail latency. We propose TPC, an approach that combines prediction information judiciously with dynamic correction for inaccurate prediction. Dynamic correction increases parallelism to accelerate a long request that is mispredicted as short. TPC carefully selects the appropriate target latencies based on system load and parallelism efficiency to reduce tail latency. Myeongjae Jeon, Yuxiong He, Hwanju Kim, Sameh Elnikety, Scott Rixner, Alan L. Cox |
ASPLOS | 3 |
| 2016 | Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile SystemsabstractMost embedded systems require contiguous memory space to be reserved for devices, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses. Our scheme, on the other hand, utilizes reserved memory as an eviction-based file cache. It guarantees contiguous memory allocation to devices while providing idle device memory as an additional file cache called eCache for general-purpose usage. Because eCache stores only evicted data from the in-kernel page cache, the memory efficiency is preserved and the allocation time for devices is minimized. Cost-based region selection also minimizes additional read I/O operations by carefully discarding cached data from eCache. The additional indexing cost incurred by adding eCache is minimized by integrating its index structure with the kernel page cache. The prototype is implemented on the Nexus S smartphone and is evaluated using popular Android applications. The evaluation results show that our scheme outperforms previous approaches in terms of the application launch performance. The device memory reallocation time is also limited to a few milliseconds, which is sufficiently small to make our scheme transparent. Jinkyu Jeong, Hwanju Kim, Joonwon Lee |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | Virtual Snooping Coherence for Multi-Core Virtualized SystemsabstractProliferation of virtualized systems opens a new opportunity to improve the scalability of multi-core architectures. Among the scalability bottlenecks in multi-cores, cache coherence has been one of the most critical problems. Although snoop-based protocols have been dominating commercial multi-core designs, it has been difficult to scale them for more cores, as snooping protocols require high network bandwidth and power consumption for snooping all the caches.In this paper, we propose a novel snoop-based cache coherence protocol, called virtual snooping, for virtualized multi-core architectures. Virtual snooping exploits memory isolation across virtual machines and prevents unnecessary snoop requests from crossing the virtual machine boundaries. Each virtual machine becomes a virtual snoop domain, consisting of a subset of the cores in a system. Although the majority of virtual machine memory is isolated, sharing of cachelines across VMs still occur. To address such data sharing, this paper investigates three factors, data sharing through the hypervisor, virtual machine relocation, and content-based sharing. In this paper, we explore the design space of virtual snooping with experiments on emulated and real virtualized systems including the mechanisms and overheads of the hypervisor. In addition, the paper discusses the scheduling impact on the effectiveness of virtual snooping. Daehoon Kim 0001, Chang Hyun Park 0001, Hwanju Kim, Jaehyuk Huh 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | vCache: architectural support for transparent and isolated virtual LLCs in virtualized environmentsabstractA key role of virtualization is to give an illusion that a consolidated workload runs on a dedicated machine although the underlying resources are actively shared by multiple workloads. Technical advances have enabled a virtual machine (VM) to exercise many shared resources of a machine in a transparent and isolated manner. However, such an illusion of resource dedication has not been supported for the last-level cache (LLC), although the LLC is the largest on-chip shared resource with a significant performance impact. In this paper, we propose vCache--architectural support to provide a transparent and isolated virtual LLC (vLLC) for each VM and interfaces to manage the vLLC. More specifically, this study first proposes architectural support for the guest OS of a VM to index the LLC with its guest physical address instead of a host physical address. This in turn allows that the guest OS transparently view its vLLC and preserve the effectiveness of its page placement policy. Second, this study extends the architectural support for each VM to keep its vLLC strongly isolated from other VMs. Such resource dedication is critical to offer performance isolation and preserve vLLC transparency for each VM in a highly consolidated machine. With little hardware overhead, vCache can facilitate various unchartered vLLC capacity-based services for the public clouds while providing up to 17% higher performance than a traditional virtualized system. Daehoon Kim 0001, Hwanju Kim, Nam Sung Kim, Jaehyuk Huh 0001 |
MICRO | 2 |
| 2015 | Request-Oriented Durable Write Caching for Application Performance
Hwanju Kim, Sang-Hoon Kim, Joonwon Lee, Jinkyu Jeong |
USENIX ATC | 2 |
| 2014 | Virtual asymmetric multiprocessor for interactive performance of consolidated desktopsabstractThis paper presents virtual asymmetric multiprocessor, a new scheme of virtual desktop scheduling on multi-core processors for user-interactive performance. The proposed scheme enables virtual CPUs to be dynamically performance-asymmetric based on their hosted workloads. To enhance user experience on consolidated desktops, our scheme provides interactive workloads with fast virtual CPUs, which have more computing power than those hosting background workloads in the same virtual machine. To this end, we devise a hypervisor extension that transparently classifies background tasks from potentially interactive workloads. In addition, we introduce a guest extension that manipulates the scheduling policy of an operating system in favor of our hypervisor-level scheme so that interactive performance can be further improved. Our evaluation shows that the proposed scheme significantly improves interactive performance of application launch, Web browsing, and video playback applications when CPU-intensive workloads highly disturb the interactive workloads. Hwanju Kim, Jinkyu Jeong, Joonwon Lee |
VEE | 1 |
| 2014 | Group-based memory oversubscription for virtualized clouds
Hwanju Kim, Joonwon Lee, Jinkyu Jeong |
J. Parallel Distributed Comput. | 2 |
| 2013 | Demand-based coordinated scheduling for SMP VMsabstractAs processor architectures have been enhancing their computing capacity by increasing core counts, independent workloads can be consolidated on a single node for the sake of high resource efficiency in data centers. With the prevalence of virtualization technology, each individual workload can be hosted on a virtual machine for strong isolation between co-located workloads. Along with this trend, hosted applications have increasingly been multithreaded to take advantage of improved hardware parallelism. Although the performance of many multithreaded applications highly depends on communication (or synchronization) latency, existing schemes of virtual machine scheduling do not explicitly coordinate virtual CPUs based on their communication behaviors. Hwanju Kim, Jinkyu Jeong, Joonwon Lee, Seung Ryoul Maeng |
ASPLOS | 1 |
| 2013 | Rigorous rental memory management for embedded systemsabstractMemory reservation in embedded systems is a prevalent approach to provide a physically contiguous memory region to its integrated devices, such as a camera device and a video decoder. Inefficiency of the memory reservation becomes a more significant problem in emerging embedded systems, such as smartphones and smart TVs. Many ways of using these systems increase the idle time of their integrated devices, and eventually decrease the utilization of their reserved memory. In this article, we propose a scheme to minimize the memory inefficiency caused by the memory reservation. The memory space reserved for a device can be rented for other purposes when the device is not active. For this scheme to be viable, latencies associated with reallocating the memory space should be minimal. Volatile pages are good candidates for such page reallocation since they can be reclaimed immediately as they are needed by the original device. We also provide two optimization techniques, lazy-migration and adaptive-activation. The former increases the lowered utilization of the rental memory by our volatile page allocations, and the latter saves active pages in the rental memory during the reallocation. We implemented our scheme on a smartphone development board with the Android Linux kernel. Our prototype has shown that the time for the return operation is less than 0.77 seconds in the tested cases. We believe that this time is acceptable to end-users in terms of transparency since the time can be hidden in application initialization time. The rental memory also brings throughput increases ranging from 2% to 200% based on the available memory and the applications' memory intensiveness. Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Analysis of virtual machine live-migration as a method for power-capping
Jinkyu Jeong, Sung-hun Kim 0005, Hwanju Kim, Joonwon Lee, Euiseong Seo |
J. Supercomput. | 3 |
| 2012 | DaaC: device-reserved memory as an eviction-based file cacheabstractMost embedded systems require contiguous memory space to be reserved for each device, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses. Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng |
CASES | 2 |
| 2012 | Scheduler support for video-oriented multimedia on client-side virtualizationabstractVirtualization has recently been adopted for client devices to provide strong isolation between services and efficient manageability. Even though multimedia service is not rare for the devices, the virtual machine hosting this service is not guaranteed to receive proper scheduling support from the underlying hypervisor. The quality of multimedia service is often compromised when several virtual machines compete for computing power. This paper presents a new scheduling scheme for the hypervisor to transparently identify if the workload handles multimedia and to provide proper scheduling supports. An implementation of our scheme has shown that the virtual machine hosting a video-oriented application receives propoer CPU scheduling even when other virtual machines host CPU intensive workloads. Hwanju Kim, Jinkyu Jeong, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng |
MMSys | 1 |
| 2011 | Transparently bridging semantic gap in CPU management for virtualized environments
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee, Seung Ryoul Maeng |
J. Parallel Distributed Comput. | 1 |
| 2011 | XHive: Efficient Cooperative Caching for Virtual MachinesabstractSince a virtual machine independently uses its own caching policy, redundant disk operations exacerbate the I/O virtualization overhead when virtual machines access large amounts of data on shared storage. This paper presents XHive, an efficient cooperative caching system that is implemented at the virtualization layer, for consolidated environments. Our proposed scheme globally manages buffer caches of consolidated virtual machines in order to accommodate a shared working set in machine memory. A singlet, which is a block cached solely by a virtual machine, is preferentially given more chances to be cached in machine memory by XHive, when it is evicted by a guest operating system. For efficient use of limited memory, singlets are cached in memory that is collaboratively donated from idle memory of virtual machines. Our evaluation shows that XHive significantly reduces disk I/O operations for shared working sets, thereby achieving high read performance and scalability. Improved scalability enables a high degree of workload consolidation with respect to virtual machines that have shared working sets. Hwanju Kim, Heeseung Jo, Joonwon Lee |
IEEE Trans. Computers | 1 |
| 2010 | Virtual Snooping: Filtering Snoops in Virtualized Multi-coresabstractVirtualization has been rapidly expanding its applications in numerous server and desktop environments to improve the utilization and manageability of physical systems. Such proliferation of virtualized systems opens a new opportunity to improve the scalability of future multi-core architectures. Among the scalability bottlenecks in multi-cores, cache coherence has been a critical problem. Although snoop-based protocols have been dominating commercial multi-core designs, it has been difficult to scale them for more cores, as snooping protocols require high network bandwidth and power consumption for snooping all the caches. In this paper, we propose a novel snoop-based cache coherence protocol, called virtual snooping, for virtualized multi-core architectures. Virtual snooping exploits memory isolation across virtual machines and prevents unnecessary snoop requests from crossing the virtual machine boundaries. Each virtual machine becomes a virtual snoop domain, consisting of a subset of the cores in a system. However, in real virtualized systems, virtual machines cannot partition the cores perfectly without any data sharing across the snoop partitions. This paper investigates three factors, which break the memory isolation among virtual machines: data sharing with a hyper visor, virtual machine relocation, and content-based data sharing. In this paper, we explore the design space of virtual snooping with experiments on real virtualized systems and approximate simulations. The results show that virtual snooping can reduce snoops significantly even if virtual machines migrate frequently. We also propose mechanisms to address content-based data sharing by exploiting its read-only property. Daehoon Kim 0001, Hwanju Kim, Jaehyuk Huh 0001 |
MICRO | 2 |
| 2010 | KAL: kernel-assisted non-invasive memory leak tolerance with a general-purpose memory allocatorabstractAbstract Memory leaks are a continuing problem in the software developed with programming languages, such as C and C++. A recent approach adopted by some researchers is to tolerate leaks in the software application and to reclaim the leaked memory by use of specially constructed memory allocation routines. However, such routines replace the usual general‐purpose memory allocator and tend to be less efficient in speed and in memory utilization. We propose a new scheme which coexists with the existing memory allocation routines and which reclaims memory leaks. Our scheme identifies and reclaims leaked memory at the kernel level. There are some major advantages to our approach: (1) the application software does not need to be modified; (2) the application does not need to be suspended while leaked memory is reclaimed; (3) a remote host can be used to identify the leaked memory, thus minimizing impact on the application program's performance; and (4) our scheme does not degrade the service availability of the application while detecting and reclaiming memory leaks. We have implemented a prototype that works with the GNU C library and with the Linux kernel. Our prototype has been tested and evaluated with various real‐world applications. Our results show that the computational overhead of our approach is around 2% of that incurred by the conventional memory allocator in terms of throughput and average response time. We also verified that the prototype successfully suppressed address space expansion caused by memory leaks when the applications are run on synthetic workloads. Copyright © 2010 John Wiley & Sons, Ltd. Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, Joonwon Lee |
Softw. Pract. Exp. | 4 |
| 2010 | Transparent Fault Tolerance of Device Drivers for Virtual MachinesabstractIn a consolidated server system using virtualization, physical device accesses from guest virtual machines (VMs) need to be coordinated. In this environment, a separate driver VM is usually assigned to this task to enhance reliability and to reuse existing device drivers. This driver VM needs to be highly reliable, since it handles all the I/O requests. This paper describes a mechanism to detect and recover the driver VM from faults to enhance the reliability of the whole system. The proposed mechanism is transparent in that guest VMs cannot recognize the fault and the driver VM can recover and continue its I/O operations. Our mechanism provides a progress monitoring-based fault detection that is isolated from fault contamination with low monitoring overhead. When a fault occurs, the system recovers by switching the faulted driver VM to another one. The recovery is performed without service disconnection or data loss and with negligible delay by fully exploiting the I/O structure of the virtualized system. Heeseung Jo, Hwanju Kim, Jae-Wan Jang, Joonwon Lee, Seung Ryoul Maeng |
IEEE Trans. Computers | 2 |
| 2009 | Task-aware virtual machine scheduling for I/O performanceabstractThe use of virtualization is progressively accommodating diverse and unpredictable workloads as being adopted in virtual desktop and cloud computing environments. Since a virtual machine monitor lacks knowledge of each virtual machine, the unpredictableness of workloads makes resource allocation difficult. Particularly, virtual machine scheduling has a critical impact on I/O performance in cases where the virtual machine monitor is agnostic about the internal workloads of virtual machines. This paper presents a task-aware virtual machine scheduling mechanism based on inference techniques using gray-box knowledge. The proposed mechanism infers the I/O-boundness of guest-level tasks and correlates incoming events with I/O-bound tasks. With this information, we introduce partial boosting, which is a priority boosting mechanism with task-level granularity, so that an I/O-bound task is selectively scheduled to handle its incoming events promptly. Our technique focuses on improving the performance of I/O-bound tasks within heterogeneous workloads by lightweight mechanisms with complete CPU fairness among virtual machines. All implementation is confined to the virtualization layer based on the Xen virtual machine monitor and the credit scheduler. We evaluate our prototype in terms of I/O performance and CPU fairness over synthetic mixed workloads and realistic applications. Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee |
VEE | 1 |
| 2008 | Guest-Aware Priority-Based Virtual Machine Scheduling for Highly Consolidated Server
Hwanju Kim, Myeongjae Jeon, Euiseong Seo, Joonwon Lee |
Euro-Par | 2 |