Jinkyu Jeong

dblp:57/1613 · DBLP profile ↗
← Back
40ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-4905-9244ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 MemSOS: OS-Guided Selective Memory Mirroring
abstract
Memory errors pose an escalating threat to datacenter reliability as DRAM technology scales to smaller nodes and servers process ever-larger datasets. While Error Correction Code (ECC) provides a first-line defense, uncorrectable errors still cause catastrophic server failures with significant economic impact. Memory mirroring offers complementary protection against these errors, but existing memory mirroring solutions require reserving specific memory regions exclusively for mirroring, which incurs significant capacity overhead and thus limits wide adoption. A recent proposal suggests leveraging free memory space for mirroring, but it leaves a critical research question unanswered: which data to mirror when free memory is limited. Thus, we propose MemSOS, a selective memory mirroring system that dynamically chooses which pages to mirror based on their impact on system reliability. Specifically, MemSOS selects pages to mirror based on their criticality and recency. Criticality is evaluated by examining the page type, while recency serves as a proxy for the likelihood of future access. Our evaluation demonstrates that MemSOS reduces system Failures In Time (FIT) by up to 19,000× compared to a state-of-the-art partial mirroring scheme, while maintaining less than 3 % performance overhead. In many cases, MemSOS achieves reliability levels comparable to full mirroring, underscoring its effectiveness in maximizing system availability under limited free memory space.
Junghoon Kim 0008, Jongheon Jeong, Seokwon Moon, Seong Hoon Seo, Yeonhong Park, Jinkyu Jeong, Nam Sung Kim, Jae W. Lee
HPCA6
2026 Beyond Static Policies: Dynamic-PRAC for Balanced and Efficient Rowhammer Mitigation
Minwoo Ahn, Jisung Park 0001, Jinkyu Jeong
ISLPED4
2025 CPC: Coordinated Page Cache for Serverless Computing
abstract
Virtual machine-based serverless computing suffers from high memory overhead due to data duplication. This duplication occurs not only across function (or VM) images but also within in-memory caches shared between VMs. Such redundancy limits the scalability and efficiency of serverless computing infrastructure. While existing memory deduplication techniques can help, they are ill-suited to serverless computing due to their high CPU overhead and delayed deduplication, which often comes too late to provide meaningful memory savings given the short lifetime of a function. However, serverless platforms inherently know which functions (or VMs) will be executed in advance, presenting a unique opportunity for proactive memory optimization. In this paper, we propose CPC, a proactive deduplication scheme that enables efficient page sharing across microVMs, lightweight variants of traditional VMs, in serverless environments. Upon function submission, CPC deduplicates identical files across function images using existing techniques. During function execution, CPC leverages this deduplication information to further deduplicate redundant cache data across VMs and the hypervisor. Comprehensive experiments using serverless function benchmarks show that CPC reduces memory consumption by up to 32.0% compared to AWS Lambda when hosting heterogeneous functions on a shared worker node. Furthermore, with Azure Functions traces, CPC achieves a 41.2% reduction in 99th percentile tail latency by reducing unnecessary instance evictions through more efficient memory utilization.
Keun Soo Lim, Yunjay Hong, Jongheon Jeong, Sam Son, Yeonhong Park, Jae W. Lee, Jinkyu Jeong
PACT8
2025 A Scalable and Overflow-Tolerant Mechanism for Minimum Virtual Time Tracking
abstract
In computer systems, fair-share scheduling of resources is essential, and virtual time-based algorithms are widely adopted for their work-conserving nature. These algorithms rely on tracking the minimum virtual time, as they always schedule the entity with the smallest value to ensure fairness. Maintaining this minimum efficiently is crucial for scalable performance, particularly in multi-core systems where contention can be high. Mindicator, a scalable and low-overhead data structure, is wellsuited for tracking minimum values and is a natural candidate for monitoring minimum virtual time. However, its use in virtual time management is limited because virtual time grows monotonically and can exceed the 32-bit integer range supported by Mindicator. This leads to incorrect minimum tracking when values overflow, potentially causing fairness violations and even malfunctioning behavior. To overcome these limitations, this paper proposes TMindicator (Twin-Mindicator), a scalable approach to virtual time tracking that tolerates integer overflow and supports values with arbitrary bit widths. T-Mindicator uses two Mindicator instances, each managing a 32-bit value, while independently tracking the number of even and odd overflow events. By concatenating the overflow counters with the 32-bit minimum values from each instance, T-Mindicator effectively extends support to 64-bit and larger virtual time representations. Our evaluation demonstrates that T-Mindicator preserves fairness among competing entities and ensures stable workload execution without anomalies when integrated into a state-of-the-art fair I/O scheduler.
Gyusun Lee, Seungwoo Jin, Jiwon Woo, Jinkyu Jeong
ICCD4
2024 Identifying On-/Off-CPU Bottlenecks Together with Blocked Samples
Minwoo Ahn, Jeongmin Han, Youngjin Kwon, Jinkyu Jeong
OSDI4
2024 An Adaptive Zone-Grouping Scheme Enabling General-Purpose File Systems on ZNS SSDs
abstract
Zoned namespace solid state drives (ZNS SSDs) provide superior capacity and life span compared to conventional namespace (CNS) SSDs at the same cost. Especially, a small-zone ZNS SSD can achieve a low write amplification factor (WAF) and performance isolation among I/O streams by providing the host with finer-grained control of the SSD's flash chips than the large-zone counterpart. To maximize these advantages, multiple zones in a small-zone SSD should be grouped and operated simultaneously. A ZNS-aware file system encapsulates the ZNS interface and transforms a ZNS SSD into a general-purpose storage to the host layer. However, the current ZNS-aware file systems were designed for the large-zone ZNS SSDs, and thus lack the zone group management feature. This paper proposes a dynamic zone group management scheme for ZNS-aware file systems that dynamically adjusts the number of zone groups in accordance with the workload characteristics and the size of each zone group to achieve high performance at a low WAF. The proposed scheme was implemented in the F2FS. Our evaluation with FIO, Filebench and db_bench showed that, across all cases, the proposed scheme improved the performance by up to 71.2% while reducing the WAF by 29.8% in comparison to the static zone grouping scheme.
Jungyun Choi, Yuhun Jun, Jinkyu Jeong, Euiseong Seo
SYSTOR4
2024 A Secure, Fast, and Resource-Efficient Serverless Platform with Function REWIND
Jaehyun Song, Bumsuk Kim, Minwoo Kwak, Byoungyoung Lee, Euiseong Seo, Jinkyu Jeong
USENIX ATC6
2023 NVMe-Driven Lazy Cache Coherence for Immutable Data with NVMe over Fabrics
abstract
In this work, we explore opportunities to design shared storage systems that leverage the distance connectivity of NVMe over Fabrics (NVMe-oF). NVMe-oF enables the use of NVMe storage devices in a shared storage environment, where multiple servers can access the same storage device via RDMA. Leveraging the distance connectivity of NVMe-oF, we develop a shared file system called EXT4-oF by extending the EXT4 file system. EXT4-oF uses RDMA to enable a local file system to function as a shared file system without requiring remote daemon processes. EXT4-oF employs a novel NVMe-driven lazy cache coherence to maintain cache coherence of file system metadata across multiple compute nodes upon creating new files, all achieved without the need for any daemon processes. To ensure cache coherence, NVMe-driven lazy cache coherence mechanism requires compute nodes to perform a re-read of NVMe-oF to avoid false negative file open errors. Through our experiments, we demonstrate that EXT4-oF improves the performance of MinIO by minimizing network traffic between compute and storage nodes and eliminating the need for TCP/IP communication during remote reads.
Hyeongjun Jeon, Daegyu Han, Duck-Ho Bae, Youngjin Yu, Kyeungpyo Kim, Sung-Soon Park 0001, Jinkyu Jeong, Beomseok Nam
CLOUD8
2022 Efficient hybrid polling for ultra-low latency storage devices
Gyusun Lee, Seokha Shin, Jinkyu Jeong
J. Syst. Archit.3
2021 D2FQ: Device-Direct Fair Queueing for NVMe SSDs
Jiwon Woo, Minwoo Ahn, Gyusun Lee, Jinkyu Jeong
FAST4
2021 Z-Journal: Scalable Per-Core Journaling
Cassiano Campes, Joo Young Hwang, Jinkyu Jeong, Euiseong Seo
USENIX ATC4
2021 ASAP: Fast Mobile Application Switch via Adaptive Prepaging
Sam Son, Seung Yul Lee, Yunho Jin, Jonghyun Bae, Jinkyu Jeong, Tae Jun Ham, Jae W. Lee, Hongil Yoon
USENIX ATC5
2021 SCOZ: A system-wide causal profiler for multicore systems
abstract
Abstract The increased complexity of hardware and software makes it difficult to analyze programs with conventional profilers. The causal profiling technique is introduced to solve the problem of conventional profilers. The causal profiling technique finds the bottleneck of the program and shows the effect of optimizing it. COZ, the newest causal profiler, exploits a technique called virtual speedup to perform causal profiling without actually optimizing program codes. However, it can only profile multithreaded applications, and cannot profile multiprogram applications and operating system (OS) kernel codes, thereby limiting the use of causal profiling. This article introduces SCOZ, a system‐wide causal profiler that addresses these limitations. The proposed profiler changes the target of virtual speedup from threads to CPU cores, thereby expanding the profiling coverage to diverse applications as well as OS kernel codes. To verify our profiler, we profiled multithreaded and OS kernel‐intensive applications. For multithread applications, our profiler shows identical results to what COZ provides. For the OS kernel‐intensive applications, our profiler identifies identical bottlenecks that previous OS scalability studies have pinpointed. Finally, we verified the profiling capability of the proposed profiler by profiling and optimizing multiprocess applications in the NAS parallel benchmark suite.
Minwoo Ahn, Taekeun Nam, Jinkyu Jeong
Softw. Pract. Exp.4
2020 A Case for Hardware-Based Demand Paging
abstract
The virtual memory system is pervasive in today's computer systems, and demand paging is the key enabling mechanism for it. At a page miss, the CPU raises an exception, and the page fault handler is responsible for fetching the requested page from the disk. The OS typically performs a context switch to run other threads as traditional disk access is slow. However, with the widespread adoption of high-performance storage devices, such as low-latency solid-state drives (SSDs), the traditional OS-based demand paging is no longer effective because a considerable portion of the demand paging latency is now spent inside the OS kernel. Thus, this paper makes a case for hardware-based demand paging that mostly eliminates OS involvement in page miss handling to provide a near-disk-access-time latency for demand paging. To this end, two architectural extensions are proposed: LBA-augmented page table that moves I/O stack operations to the control plane and Storage Management Unit that enables CPU to directly issue I/O commands without OS intervention in most cases. OS support is also proposed to detach tasks for memory resource management from the critical path. The evaluation results using both a cycle-level simulator and a real x86 machine with an ultra-low latency SSD show that the proposed scheme reduces the demand paging latency by 37.0%, and hence improves the performance of FIO read random benchmark by up to 57.1% and a NoSQL server by up to 27.3% with real-world workloads. As a side effect of eliminating OS intervention, the IPC of the user-level code is also increased by up to 7.0%.
Gyusun Lee, Wenjing Jin 0001, Wonsuk Song, Jeonghun Gong, Jonghyun Bae, Tae Jun Ham, Jae W. Lee, Jinkyu Jeong
ISCA8
2019 Asynchronous I/O Stack: A Low-latency Kernel I/O Stack for Ultra-Low Latency SSDs
Gyusun Lee, Seokha Shin, Wonsuk Song, Tae Jun Ham, Jae W. Lee, Jinkyu Jeong
USENIX ATC6
2017 Enlightening the I/O Path: A Holistic Approach for Application Performance
Hwanju Kim, Joonwon Lee, Jinkyu Jeong
FAST4
2017 Request-aware Cooperative I/O Scheduling for Scale-out Database Applications
Hyungil Jo, Sung-hun Kim 0005, Jinkyu Jeong, Joonwon Lee
HotStorage4
2017 Enhancing network I/o performance for a virtualized Hadoop cluster
abstract
Summary A MapReduce programming model is proposed to process big data using Hadoop, one of the major cloud computing frameworks. With the increasing adoption of cloud computing, running a Hadoop framework on a virtualized cluster is a compelling approach to reducing costs and increasing efficiency. In this paper, we measure the performance of a virtualized network and analyze the impact of network performance on Hadoop workloads running on a virtualized cluster. Then, we propose a virtualized network I/O architecture as a novel optimization for a virtualized Hadoop cluster for a public/private cloud provider. The proposed network architecture combines traditional network configurations and achieves better performance for Hadoop workloads. We also show a better way to utilize the rack awareness feature of the Hadoop framework in the proposed computing environment. The evaluation demonstrates that the proposed network architecture and mechanisms improve performance by up to 4.1 times compared with a bridge network architecture. This novel architecture can even virtually match the performance of the expensive, hardware‐based single root I/O virtualization network architecture.
Jinkyu Jeong, Dong Hoon Choi, Heeseung Jo
Concurr. Comput. Pract. Exp.1
2017 Application-Aware Swapping for Mobile Systems
abstract
There has been a constant demand for memory in modern mobile systems to provide users with better experience. Swapping is one of the cost-effective software solutions to provide extra usable memory by reclaiming inactive pages and improving memory utilization. However, swapping has not been actively adopted to mobile systems since it incurs a significant amount of I/O, which in fact impairs system performance as well as user experience. In this paper, we propose a novel scheme to properly harness the swapping to mobile systems. We identify that a vast amount of I/O for swapping comes from the conflict of the traditional page-level approach of the swapping and the process-level memory management scheme tailored to mobile systems. Moreover, we find out that the current victim page selection policy is not effective due to the process-level policy. To address these problems, we revise the victim selection policy to resolve the conflict and to selectively perform swapping according to the efficacy of swapping. Evaluation using a running prototype with realistic workloads indicates that the propose scheme effectively reduces the paging traffic, thereby improving user experience as well as energy consumption.
Sang-Hoon Kim, Jinkyu Jeong, Jin-Soo Kim 0001
ACM Trans. Embed. Comput. Syst.2
2016 Efficient footprint caching for Tagless DRAM Caches
abstract
Efficient cache tag management is a primary design objective for large, in-package DRAM caches. Recently, Tagless DRAM Caches (TDCs) have been proposed to completely eliminate tagging structures from both on-die SRAM and in-package DRAM, which are a major scalability bottleneck for future multi-gigabyte DRAM caches. However, TDC imposes a constraint on DRAM cache block size to be the same as OS page size (e.g., 4KB) as it takes a unified approach to address translation and cache tag management. Caching at a page granularity, or page-based caching, incurs significant off-package DRAM bandwidth waste by over-fetching blocks within a page that are not actually used. Footprint caching is an effective solution to this problem, which fetches only those blocks that will likely be touched during the page's lifetime in the DRAM cache, referred to as the page's footprint. In this paper we demonstrate TDC opens up unique opportunities to realize efficient footprint caching with higher prediction accuracy and a lower hardware cost than the original footprint caching scheme. Since there are no cache tags in TDC, the footprints of cached pages are tracked at TLB, instead of cache tag array, to incur much lower on-die storage overhead than the original design. Besides, when a cached page is evicted, its footprint will be stored in the corresponding page table entry, instead of an auxiliary on-die structure (i.e., Footprint History Table), to prevent footprint thrashing among different pages, thus yielding higher accuracy in footprint prediction. The resulting design, called Footprint-augmented Tagless DRAM Cache (F-TDC), significantly improves the bandwidth efficiency of TDC, and hence its performance and energy efficiency. Our evaluation with 3D Through-Silicon-Via-based in-package DRAM demonstrates an average reduction of off-package bandwidth by 32.0%, which, in turn, improves IPC and EDP by 17.7% and 25.4%, respectively, over the state-of-the-art TDC with no footprint caching.
Hakbeom Jang, Yongjun Lee, Youngsok Kim, Jangwoo Kim, Jinkyu Jeong, Jae W. Lee
HPCA6
2016 SmartLMK: A Memory Reclamation Scheme for Improving User-Perceived App Launch Time
abstract
As the mobile computing environment evolves, users demand high-quality apps and better user experience. Consequently, memory demand in mobile devices has soared. Device manufacturers have fulfilled the demand by equipping devices with more RAM. However, such a hardware approach is only a temporary solution and does not scale well in the resource-constrained mobile environment. Meanwhile, mobile systems adopt a new app life cycle and a memory reclamation scheme tailored for the life cycle. When a user leaves an app, the app is not terminated but cached in memory as long as there is enough free memory. If the free memory gets low, a victim app is terminated and the associated memory to the app is reclaimed. This process-level approach has worked well in the mobile environment. However, user experience can be impaired severely because the victim selection policy does not consider the user experience. In this article, we propose a novel memory reclamation scheme called SmartLMK . SmartLMK minimizes the impact of the process-level reclamation on user experience. The worthiness to keep an app in memory is modeled by means of user-perceived app launch time and app usage statistics. The memory footprint and impending memory demand are estimated from the history of the memory usage. Using these values and memory models, SmartLMK picks up the least valuable apps and terminates them at once. Our evaluation on a real Android-based smartphone shows that SmartLMK efficiently distinguishes the valuable apps among cached apps and keeps those valuable apps in memory. As a result, the user-perceived app launch time can be improved by up to 13.2%.
Sang-Hoon Kim, Jinkyu Jeong, Jin-Soo Kim 0001, Seung Ryoul Maeng
ACM Trans. Embed. Comput. Syst.2
2016 Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems
abstract
Most embedded systems require contiguous memory space to be reserved for devices, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses. Our scheme, on the other hand, utilizes reserved memory as an eviction-based file cache. It guarantees contiguous memory allocation to devices while providing idle device memory as an additional file cache called eCache for general-purpose usage. Because eCache stores only evicted data from the in-kernel page cache, the memory efficiency is preserved and the allocation time for devices is minimized. Cost-based region selection also minimizes additional read I/O operations by carefully discarding cached data from eCache. The additional indexing cost incurred by adding eCache is minimized by integrating its index structure with the kernel page cache. The prototype is implemented on the Nexus S smartphone and is evaluated using popular Android applications. The evaluation results show that our scheme outperforms previous approaches in terms of the application launch performance. The device memory reallocation time is also limited to a few milliseconds, which is sufficiently small to make our scheme transparent.
Jinkyu Jeong, Hwanju Kim, Joonwon Lee
IEEE Trans. Mob. Comput.1
2015 Managing gpu buffers for caching more apps in mobile systems
abstract
Modern mobile systems cache apps actively to quickly respond to a user’s call to launch apps. Since the amount of usable memory is critical to the number of cacheable apps, it is important to maximize memory utilization. Meanwhile, modern mobile apps make use of graphics processingunits (GPUs) to accelerate their graphic operations and to provide better user experience. In resource-constrained mobile systems, GPU cannot afford its private memory but shares the main memory with CPU. It leads to a considerable amount of main memory to be allocated for GPU buffers which are used for processingGPU operations. These GPU buffers are, however, not managed effectively so that inactive GPU buffers occupy a large fraction of the memory and decrease memory utilization. This paper proposes a scheme to manage GPU buffers to increase the memory utilization in mobile systems. Our scheme identifies inactive GPU buffers by exploitingthe state of an app from a user’s perspective, and reduces their memory footprint by compressingthem. Our sophisticated design approach prevents GPU-specific issues from causing an unpleasant overhead. Our evaluation on a runningprototype with realistic workloads shows that the proposed scheme can secure up to 215.9 MB of extra memory from 1.5 GB of main memory and increase the average number of cached apps by up to 31.3%.
Sejun Kwon, Sang-Hoon Kim, Jin-Soo Kim 0001, Jinkyu Jeong
EMSOFT4
2015 A fully associative, tagless DRAM cache
abstract
This paper introduces a tagless cache architecture for large in-package DRAM caches. The conventional die-stacked DRAM cache has both a TLB and a cache tag array, which are responsible for virtual-to-physical and physical-to-cache address translation, respectively. We propose to align the granularity of caching with OS page size and take a unified approach to address translation and cache tag management. To this end, we introduce cache-map TLB (cTLB), which stores virtual-to-cache, instead of virtual-to-physical, address mappings. At a TLB miss, the TLB miss handler allocates the requested block into the cache if it is not cached yet, and updates both the page table and cTLB with the virtual-to-cache address mapping. Assuming the availability of large in-package DRAM caches, this ensures that an access to the memory region within the TLB reach always hits in the cache with low hit latency since a TLB access immediately returns the exact location of the requested block in the cache, hence saving a tag-checking operation. The remaining cache space is used as victim cache for memory pages that are recently evicted from cTLB. By completely eliminating data structures for cache tag management, from either on-die SRAM or in-package DRAM, the proposed DRAM cache achieves best scalability and hit latency, while maintaining high hit rate of a fully associative cache. Our evaluation with 3D Through-Silicon Via (TSV)-based in-package DRAM demonstrates that the proposed cache improves the IPC and energy efficiency by 30.9% and 39.5%, respectively, compared to the baseline with no DRAM cache. These numbers translate to 4.3% and 23.8% improvements over an impractical SRAM-tag cache requiring megabytes of on-die SRAM storage, due to low hit latency and zero energy waste for cache tags.
Yongjun Lee, Hakbeom Jang, Hyunggyun Yang, Jangwoo Kim, Jinkyu Jeong, Jae W. Lee
ISCA6
2015 Controlling physical memory fragmentation in mobile systems
abstract
Since the adoption of hardware-accelerated features (e.g., hardware codec) improves the performance and quality of mobile devices, it revives the need for contiguous memory allocation. However, physical memory in mobile systems is highly fragmented due to the frequent spawn and exit of processes and the lack of proactive anti-fragmentation scheme. As a result, the memory allocation for large and contiguous I/O buffers suffer from the highly fragmented memory, thereby incurring high CPU usage and power consumption. This paper presents a proactive anti-fragmentation approach that groups pages with the same lifetime, and stores them contiguously in fixed-size contiguous regions. When a process is killed to secure free memory, a set of contiguous regions are freed and subsequent contiguous memory allocations can be easily satisfied without incurring additional overhead. Our prototype implementation on a Nexus 10 tablet with the Android kernel shows that the proposed scheme greatly alleviates fragmentation, thereby reducing the I/O buffer allocation time, associated CPU usage, and energy consumption.
Sang-Hoon Kim, Sejun Kwon, Jin-Soo Kim 0001, Jinkyu Jeong
ISMM4
2015 Request-Oriented Durable Write Caching for Application Performance
Hwanju Kim, Sang-Hoon Kim, Joonwon Lee, Jinkyu Jeong
USENIX ATC5
2014 Virtual asymmetric multiprocessor for interactive performance of consolidated desktops
abstract
This paper presents virtual asymmetric multiprocessor, a new scheme of virtual desktop scheduling on multi-core processors for user-interactive performance. The proposed scheme enables virtual CPUs to be dynamically performance-asymmetric based on their hosted workloads. To enhance user experience on consolidated desktops, our scheme provides interactive workloads with fast virtual CPUs, which have more computing power than those hosting background workloads in the same virtual machine. To this end, we devise a hypervisor extension that transparently classifies background tasks from potentially interactive workloads. In addition, we introduce a guest extension that manipulates the scheduling policy of an operating system in favor of our hypervisor-level scheme so that interactive performance can be further improved. Our evaluation shows that the proposed scheme significantly improves interactive performance of application launch, Web browsing, and video playback applications when CPU-intensive workloads highly disturb the interactive workloads.
Hwanju Kim, Jinkyu Jeong, Joonwon Lee
VEE3
2014 Group-based memory oversubscription for virtualized clouds
Hwanju Kim, Joonwon Lee, Jinkyu Jeong
J. Parallel Distributed Comput.4
2013 Demand-based coordinated scheduling for SMP VMs
abstract
As processor architectures have been enhancing their computing capacity by increasing core counts, independent workloads can be consolidated on a single node for the sake of high resource efficiency in data centers. With the prevalence of virtualization technology, each individual workload can be hosted on a virtual machine for strong isolation between co-located workloads. Along with this trend, hosted applications have increasingly been multithreaded to take advantage of improved hardware parallelism. Although the performance of many multithreaded applications highly depends on communication (or synchronization) latency, existing schemes of virtual machine scheduling do not explicitly coordinate virtual CPUs based on their communication behaviors.
Hwanju Kim, Jinkyu Jeong, Joonwon Lee, Seung Ryoul Maeng
ASPLOS3
2013 Rigorous rental memory management for embedded systems
abstract
Memory reservation in embedded systems is a prevalent approach to provide a physically contiguous memory region to its integrated devices, such as a camera device and a video decoder. Inefficiency of the memory reservation becomes a more significant problem in emerging embedded systems, such as smartphones and smart TVs. Many ways of using these systems increase the idle time of their integrated devices, and eventually decrease the utilization of their reserved memory. In this article, we propose a scheme to minimize the memory inefficiency caused by the memory reservation. The memory space reserved for a device can be rented for other purposes when the device is not active. For this scheme to be viable, latencies associated with reallocating the memory space should be minimal. Volatile pages are good candidates for such page reallocation since they can be reclaimed immediately as they are needed by the original device. We also provide two optimization techniques, lazy-migration and adaptive-activation. The former increases the lowered utilization of the rental memory by our volatile page allocations, and the latter saves active pages in the rental memory during the reallocation. We implemented our scheme on a smartphone development board with the Android Linux kernel. Our prototype has shown that the time for the return operation is less than 0.77 seconds in the tested cases. We believe that this time is acceptable to end-users in terms of transparency since the time can be hidden in application initialization time. The rental memory also brings throughput increases ranging from 2% to 200% based on the available memory and the applications' memory intensiveness.
Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
ACM Trans. Embed. Comput. Syst.1
2013 Analysis of virtual machine live-migration as a method for power-capping
Jinkyu Jeong, Sung-hun Kim 0005, Hwanju Kim, Joonwon Lee, Euiseong Seo
J. Supercomput.1
2012 DaaC: device-reserved memory as an eviction-based file cache
abstract
Most embedded systems require contiguous memory space to be reserved for each device, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses.
Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
CASES1
2012 Scheduler support for video-oriented multimedia on client-side virtualization
abstract
Virtualization has recently been adopted for client devices to provide strong isolation between services and efficient manageability. Even though multimedia service is not rare for the devices, the virtual machine hosting this service is not guaranteed to receive proper scheduling support from the underlying hypervisor. The quality of multimedia service is often compromised when several virtual machines compete for computing power. This paper presents a new scheduling scheme for the hypervisor to transparently identify if the workload handles multimedia and to provide proper scheduling supports. An implementation of our scheme has shown that the virtual machine hosting a video-oriented application receives propoer CPU scheduling even when other virtual machines host CPU intensive workloads.
Hwanju Kim, Jinkyu Jeong, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
MMSys2
2012 TwoB: a two-tier web browser architecture optimized for mobile network
abstract
The connection establishment phase including DNS lookups and TCP handshakes takes significantly long time during web browsing through mobile network. In this paper, we propose a novel web browser architecture that aims at improving mobile web browsing performance. Our approach delegates the connection establishment phase and HTTP header field delivery to a dedicated proxy server located at the joint point between WAN and mobile network to reduce both the number and size of packets on mobile network. Our evaluation showed that the proposed scheme reduces the number of mobile network packets by up to 52% and, consequently, shortens the average page loading time by up to 37%.
Junguk Cho, Jinkyu Jeong, Euiseong Seo
MoMM2
2011 Transparently bridging semantic gap in CPU management for virtualized environments
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee, Seung Ryoul Maeng
J. Parallel Distributed Comput.3
2010 KAL: kernel-assisted non-invasive memory leak tolerance with a general-purpose memory allocator
abstract
Abstract Memory leaks are a continuing problem in the software developed with programming languages, such as C and C++. A recent approach adopted by some researchers is to tolerate leaks in the software application and to reclaim the leaked memory by use of specially constructed memory allocation routines. However, such routines replace the usual general‐purpose memory allocator and tend to be less efficient in speed and in memory utilization. We propose a new scheme which coexists with the existing memory allocation routines and which reclaims memory leaks. Our scheme identifies and reclaims leaked memory at the kernel level. There are some major advantages to our approach: (1) the application software does not need to be modified; (2) the application does not need to be suspended while leaked memory is reclaimed; (3) a remote host can be used to identify the leaked memory, thus minimizing impact on the application program's performance; and (4) our scheme does not degrade the service availability of the application while detecting and reclaiming memory leaks. We have implemented a prototype that works with the GNU C library and with the Linux kernel. Our prototype has been tested and evaluated with various real‐world applications. Our results show that the computational overhead of our approach is around 2% of that incurred by the conventional memory allocator in terms of throughput and average response time. We also verified that the prototype successfully suppressed address space expansion caused by memory leaks when the applications are run on synthetic workloads. Copyright © 2010 John Wiley & Sons, Ltd.
Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, Joonwon Lee
Softw. Pract. Exp.1
2010 Power Consumption Prediction and Power-Aware Packing in Consolidated Environments
abstract
Consolidation of workloads has emerged as a key mechanism to dampen the rapidly growing energy expenditure within enterprise-scale data centers. To gainfully utilize consolidation-based techniques, we must be able to characterize the power consumption of groups of colocated applications. Such characterization is crucial for effective prediction and enforcement of appropriate limits on power consumption-power budgets-within the data center. We identify two kinds of power budgets: 1) an average budget to capture an upper bound on long-term energy consumption within that level and 2) a sustained budget to capture any restrictions on sustained draw of current above a certain threshold. Using a simple measurement infrastructure, we derive power profiles-statistical descriptions of the power consumption of applications. Based on insights gained from detailed profiling of several applications-both individual and consolidated-we develop models for predicting average and sustained power consumption of consolidated applications. We conduct an experimental evaluation of our techniques on a Xen-based server that consolidates applications drawn from a diverse pool. For a variety of consolidation scenarios, we are able to predict average power consumption within five percent error margin and sustained power within 10 percent error margin. Using prediction techniques allows us to ensure safe yet efficient system operation-in a representative case, we are able to improve the number of applications consolidated on a server from two to three (compared to existing baseline techniques) by choosing the appropriate power state that satisfies the power budgets associated with the server.
Jeonghwan Choi, Sriram Govindan, Jinkyu Jeong, Bhuvan Urgaonkar, Anand Sivasubramaniam
IEEE Trans. Computers3
2009 Task-aware virtual machine scheduling for I/O performance
abstract
The use of virtualization is progressively accommodating diverse and unpredictable workloads as being adopted in virtual desktop and cloud computing environments. Since a virtual machine monitor lacks knowledge of each virtual machine, the unpredictableness of workloads makes resource allocation difficult. Particularly, virtual machine scheduling has a critical impact on I/O performance in cases where the virtual machine monitor is agnostic about the internal workloads of virtual machines. This paper presents a task-aware virtual machine scheduling mechanism based on inference techniques using gray-box knowledge. The proposed mechanism infers the I/O-boundness of guest-level tasks and correlates incoming events with I/O-bound tasks. With this information, we introduce partial boosting, which is a priority boosting mechanism with task-level granularity, so that an I/O-bound task is selectively scheduled to handle its incoming events promptly. Our technique focuses on improving the performance of I/O-bound tasks within heterogeneous workloads by lightweight mechanisms with complete CPU fairness among virtual machines. All implementation is confined to the virtualization layer based on the Xen virtual machine monitor and the credit scheduler. We evaluate our prototype in terms of I/O performance and CPU fairness over synthetic mixed workloads and realistic applications.
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee
VEE3
2009 Catching two rabbits: adaptive real-time support for embedded Linux
abstract
Abstract The trend of digital convergence makes multitasking common in many digital electronic products. Some applications in those systems have inherent real‐time properties, while many others have few or no timeliness requirements. Therefore the embedded Linux kernels, which are widely used in those devices, provide real‐time features in many forms. However, providing real‐time scheduling usually induces throughput degradation in heavy multitasking due to the increased context switches. Usually the throughput degradation becomes a critical problem, since the performance of the embedded processors is generally limited for cost, design and energy efficiency reasons. This paper proposes schemes to lessen the throughput degradation, which is from real‐time scheduling, by suppressing unnecessary context switches and applying real‐time scheduling mechanisms only when it is necessary. Also the suggested schemes enable the complete priority inheritance protocol to prevent the well‐known priority inversion problem. We evaluated the effectiveness of our approach with open‐source benchmarks. By using the suggested schemes, the throughput is improved while the scheduling latency is kept same or better in comparison with the existing approaches. Copyright © 2008 John Wiley & Sons, Ltd.
Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee
Softw. Pract. Exp.2
2008 Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors
abstract
Multicore processors deliver a higher throughput at lower power consumption than unicore processors. In the near future, they will thus be widely used in mobile real-time systems. There have been many research on energy-efficient scheduling of real-time tasks using DVS. These approaches must be modified for multicore processors, however, since normally all the cores in a chip must run at the same performance level. Thus, blindly adopting existing DVS algorithms that do not consider the restriction will result in a waste of energy. This article suggests Dynamic Repartitioning algorithm based on existing partitioning approaches of multiprocessor systems. The algorithm dynamically balances the task loads of multiple cores to optimize power consumption during execution. We also suggest Dynamic Core Scaling algorithm, which adjusts the number of active cores to reduce leakage power consumption under low load conditions. Simulation results show that Dynamic Repartitioning can produce energy savings of about 8 percent even with the best energy-efficient partitioning algorithm. The results also show that Dynamic Core Scaling can reduce energy consumption by about 26 percent under low load conditions.
Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Joonwon Lee
IEEE Trans. Parallel Distributed Syst.2