Joonwon Lee

dblp:29/3723 · DBLP profile ↗
← Back
57ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 4Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Energy-efficient computing · 20% Memory systems · 18% Storage systems · 18%
Human-computer interaction and pervasive computing
4 papers
Collaborative and social computing · 47% Health and well-being technologies · 28% Learning and educational technologies · 13%
Software engineering, system software, and programming languages
6 papers
Operating systems · 100%

Topics — the 30 heaviest of 55, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
virtualization
0.432013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Storage systems › i/o architecture › i/o subsystem
i/o path
0.312017
Enlightening the I/O Path: A Holistic Approach for Application Performance · FAST 2017
Operating systems › resource management
memory management
0.322016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
PABC: Power-Aware Buffer Cache Management for Low Power Consumption · IEEE Trans. Computers 2007
Collaborative and social computing › computer-mediated communication
online social interaction
0.212016
Understanding Mass Interactions in Online Sports Viewing: Chatting Motives and Usage Patterns · ACM Trans. Comput. Hum. Interact. 2016
Collaborative and social computing › social media
social media use
0.212016
Understanding Mass Interactions in Online Sports Viewing: Chatting Motives and Usage Patterns · ACM Trans. Comput. Hum. Interact. 2016
Storage systems
file systems
0.212016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Memory systems › cache management › storage caching
page cache
0.212016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Learning and educational technologies › informal learning
family co-learning
0.212015
FamiLync: facilitating participatory parental mediation of adolescents' smartphone use · UbiComp 2015
Collaborative and social computing › family communication
parental mediation
0.212015
FamiLync: facilitating participatory parental mediation of adolescents' smartphone use · UbiComp 2015
Memory systems › cache › cache organization
write cache
0.212015
Request-Oriented Durable Write Caching for Application Performance · USENIX ATC 2015
Storage systems › buffer management
buffer cache management
0.222011
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
PABC: Power-Aware Buffer Cache Management for Low Power Consumption · IEEE Trans. Computers 2007
Health and well-being technologies › digital well-being
smartphone overuse
0.212014
Hooked on smartphones: an exploratory study on smartphone overuse among college students · CHI 2014
Parallel and multicore computing › task scheduling
coordinated scheduling
0.212013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
Electronic design automation › high-level synthesis
scheduling
0.212013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine scheduling
0.212013
Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013
Energy-efficient computing
thermal management
0.222008
A CFD-Based Tool for Studying Temperature in Rack-Mounted Servers · IEEE Trans. Computers 2008
Modeling and Managing Thermal Profiles of Rack-mounted Servers with ThermoStat · HPCA 2007
Energy-efficient computing
thermal modeling
0.222008
A CFD-Based Tool for Studying Temperature in Rack-Mounted Servers · IEEE Trans. Computers 2008
Modeling and Managing Thermal Profiles of Rack-mounted Servers with ThermoStat · HPCA 2007
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.122008
Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2008
Optimal intratask dynamic voltage-scaling technique and its practical extensions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Embedded and real-time systems
real-time scheduling
0.122008
Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2008
Optimal intratask dynamic voltage-scaling technique and its practical extensions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Memory systems › cache management › storage caching
cooperative caching
0.112011
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Memory systems
memory management
0.122011
PABC: Power-Aware Buffer Cache Management for Low Power Consumption · IEEE Trans. Computers 2007
XHive: Efficient Cooperative Caching for Virtual Machines · IEEE Trans. Computers 2011
Embedded and real-time systems
device drivers
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Distributed systems › fault tolerance
fault detection and recovery
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Distributed systems
fault tolerance
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Cloud and datacenter computing › virtualization
virtual machine
0.112010
Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010
Operating systems › i/o › i/o subsystem
i/o scheduling
0.112017
Enlightening the I/O Path: A Holistic Approach for Application Performance · FAST 2017
Embedded and real-time systems › energy-efficient embedded systems
energy-efficient scheduling
0.112008
Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2008
Embedded and real-time systems › real-time scheduling
multicore scheduling
0.112008
Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2008
Embedded and real-time systems
mobile computing
0.112016
Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems · IEEE Trans. Mob. Comput. 2016
Energy-efficient computing › thermal management
dynamic thermal management
0.112007
Modeling and Managing Thermal Profiles of Rack-mounted Servers with ThermoStat · HPCA 2007

Methods — techniques the papers use, named apart from their topics

survey · 0.6prototype implementation · 0.5cost-based region selection · 0.5request-oriented caching · 0.4durability guarantees · 0.4user study · 0.4logged data analysis · 0.4interviews · 0.4demand-based scheduling · 0.3communication-aware scheduling · 0.3self-report · 0.2chat-log analysis · 0.2social cognitive theory · 0.2progress monitoring · 0.2mobile service design · 0.2computational fluid dynamics · 0.2memory donation · 0.1buffer cache management · 0.1
YearPublicationVenuePosition
2019 Optimizing the Ceph Distributed File System for High Performance Computing
abstract
With increasing demand for running big data analytics and machine learning workloads with diverse data types, high performance computing (HPC) systems consequently need to support diverse types of storage services. Ceph is one possible candidate for such HPC environments, as Ceph provides interfaces for object, block, and file storage. Ceph, however, is not designed for HPC environments, thus it needs to be optimized for HPC workloads. In this paper, we find and analyze problems that arise when running HPC workloads on Ceph, and propose a novel optimization technique called F2FS-split, based on the F2FS file system and several other optimizations. We measure the performance of Ceph in HPC environments, and show that F2FS-split outperforms both F2FS and XFS by 39% and 59%, respectively, in a write dominant workload. We also observe that modifying the Ceph RADOS object size can improve read speed further.
Kisik Jeong, Carl Duffy, Joonwon Lee
PDP4
2018 vStream: Virtual Stream Management for Multi-streamed SSDs
Hwanjin Yong, Kisik Jeong, Joonwon Lee
HotStorage3
2017 Enlightening the I/O Path: A Holistic Approach for Application Performance
Hwanju Kim, Joonwon Lee, Jinkyu Jeong
FAST3
2017 Request-aware Cooperative I/O Scheduling for Scale-out Database Applications
Hyungil Jo, Sung-hun Kim 0005, Jinkyu Jeong, Joonwon Lee
HotStorage5
2016 Transparently Exploiting Device-Reserved Memory for Application Performance in Mobile Systems
abstract
Most embedded systems require contiguous memory space to be reserved for devices, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses. Our scheme, on the other hand, utilizes reserved memory as an eviction-based file cache. It guarantees contiguous memory allocation to devices while providing idle device memory as an additional file cache called eCache for general-purpose usage. Because eCache stores only evicted data from the in-kernel page cache, the memory efficiency is preserved and the allocation time for devices is minimized. Cost-based region selection also minimizes additional read I/O operations by carefully discarding cached data from eCache. The additional indexing cost incurred by adding eCache is minimized by integrating its index structure with the kernel page cache. The prototype is implemented on the Nexus S smartphone and is evaluated using popular Android applications. The evaluation results show that our scheme outperforms previous approaches in terms of the application launch performance. The device memory reallocation time is also limited to a few milliseconds, which is sufficiently small to make our scheme transparent.
Jinkyu Jeong, Hwanju Kim, Joonwon Lee
IEEE Trans. Mob. Comput.3
2016 Understanding Mass Interactions in Online Sports Viewing: Chatting Motives and Usage Patterns
abstract
This article aims to deepen understanding of these mass interactions in online sports viewing through studying Naver Sports, the largest online sports viewing service in Korea. We examined the diverse aspects of mass interactions, including interactive experiences, usage motives, and relationships between usage patterns and motives, through analysis of almost 6 million chats from Naver Sports and from self-reporting survey data from 1,123 users. First, we found that online sports viewing provides unique interactive experiences when compared to other settings such as offline sports viewing and social TV viewing with friends. Second, we found the key motives inspiring online sports viewing include the following: sharing feelings/thoughts, wanting to be entertained, sharing information, and wanting to feel membership in a group. Third, these motives were significantly related to specific usage patterns. Finally, we explored how the study’s key findings can offer practical design implications to enhance online sports viewing services, and to show system designers how to support particular usage patterns to better accommodate specific user motives.
Minsam Ko, Seung-Woo Choi, Joonwon Lee, Uichin Lee, Aviv Segev
ACM Trans. Comput. Hum. Interact.3
2015 NUGU: A Group-based Intervention App for Improving Self-Regulation of Limiting Smartphone Use
abstract
Our preliminary study reveals that individuals use various management strategies for limiting smartphone use, ranging from keeping smartphones out of reach to removing apps. However, we also found that users often had difficulties in maintaining their chosen management strategies due to lack of self-regulation. In this paper, we present NUGU, a group-based intervention app for improving self-regulation of limiting smartphone use through leveraging social support: groups of people limit their use together by sharing their limiting information. NUGU is designed based on social cognitive theory, and it has been developed iteratively through two pilot tests. Our three-week user study (n = 62) demonstrated that compared with its non-social counterpart, the NUGU users' usage amount significantly decreased and their perceived level of managing disturbances improved. Furthermore, our exit interview confirmed that NUGU's design elements are effective for achieving limiting goals.
Minsam Ko, Subin Yang, Joonwon Lee, Christian Heizmann, Jinyoung Jeong, Uichin Lee, Daehee Shin, Koji Yatani, Junehwa Song, Kyong-Mee Chung
CSCW3
2015 FamiLync: facilitating participatory parental mediation of adolescents' smartphone use
abstract
We consider participatory parental mediation in which children engage with their parents in activities that encourage both parents and children to participate in co-learning of digital media use. To this end, we developed FamiLync, a mobile service that treats use-limiting as a family activity and provides the family with a virtual public space to foster social awareness and improve self-regulation. A three-week user study conducted with twelve families in Korea (17 parents and 18 teenagers) showed that FamiLync improves mutual understanding of usage behavior, thereby providing common grounds for parental mediation. Further, parents actively participated in use-limiting with their children, which significantly increased the children's desire to participate. As a consequence, parental mediation methods and parent-child interaction in relation to smartphone usage changed appreciably, and the participants smartphone usage amount significantly decreased.
Minsam Ko, Seung-Woo Choi, Subin Yang, Joonwon Lee, Uichin Lee
UbiComp4
2015 Request-Oriented Durable Write Caching for Application Performance
Hwanju Kim, Sang-Hoon Kim, Joonwon Lee, Jinkyu Jeong
USENIX ATC4
2014 Hooked on smartphones: an exploratory study on smartphone overuse among college students
abstract
The negative aspects of smartphone overuse on young adults, such as sleep deprivation and attention deficits, are being increasingly recognized recently. This emerging issue motivated us to analyze the usage patterns related to smartphone overuse. We investigate smartphone usage for 95 college students using surveys, logged data, and interviews. We first divide the participants into risk and non-risk groups based on self-reported rating scale for smartphone overuse. We then analyze the usage data to identify between-group usage differences, which ranged from the overall usage patterns to app-specific usage patterns. Compared with the non-risk group, our results show that the risk group has longer usage time per day and different diurnal usage patterns. Also, the risk group users are more susceptible to push notifications, and tend to consume more online content. We characterize the overall relationship between usage features and smartphone overuse using analytic modeling and provide detailed illustrations of problematic usage behaviors based on interview data.
Uichin Lee, Joonwon Lee, Minsam Ko, Changhun Lee, Yuhwan Kim, Subin Yang, Koji Yatani, Gahgene Gweon, Kyong-Mee Chung, Junehwa Song
CHI2
2014 Virtual asymmetric multiprocessor for interactive performance of consolidated desktops
abstract
This paper presents virtual asymmetric multiprocessor, a new scheme of virtual desktop scheduling on multi-core processors for user-interactive performance. The proposed scheme enables virtual CPUs to be dynamically performance-asymmetric based on their hosted workloads. To enhance user experience on consolidated desktops, our scheme provides interactive workloads with fast virtual CPUs, which have more computing power than those hosting background workloads in the same virtual machine. To this end, we devise a hypervisor extension that transparently classifies background tasks from potentially interactive workloads. In addition, we introduce a guest extension that manipulates the scheduling policy of an operating system in favor of our hypervisor-level scheme so that interactive performance can be further improved. Our evaluation shows that the proposed scheme significantly improves interactive performance of application launch, Web browsing, and video playback applications when CPU-intensive workloads highly disturb the interactive workloads.
Hwanju Kim, Jinkyu Jeong, Joonwon Lee
VEE4
2014 Group-based memory oversubscription for virtualized clouds
Hwanju Kim, Joonwon Lee, Jinkyu Jeong
J. Parallel Distributed Comput.3
2013 Demand-based coordinated scheduling for SMP VMs
abstract
As processor architectures have been enhancing their computing capacity by increasing core counts, independent workloads can be consolidated on a single node for the sake of high resource efficiency in data centers. With the prevalence of virtualization technology, each individual workload can be hosted on a virtual machine for strong isolation between co-located workloads. Along with this trend, hosted applications have increasingly been multithreaded to take advantage of improved hardware parallelism. Although the performance of many multithreaded applications highly depends on communication (or synchronization) latency, existing schemes of virtual machine scheduling do not explicitly coordinate virtual CPUs based on their communication behaviors.
Hwanju Kim, Jinkyu Jeong, Joonwon Lee, Seung Ryoul Maeng
ASPLOS4
2013 Rigorous rental memory management for embedded systems
abstract
Memory reservation in embedded systems is a prevalent approach to provide a physically contiguous memory region to its integrated devices, such as a camera device and a video decoder. Inefficiency of the memory reservation becomes a more significant problem in emerging embedded systems, such as smartphones and smart TVs. Many ways of using these systems increase the idle time of their integrated devices, and eventually decrease the utilization of their reserved memory. In this article, we propose a scheme to minimize the memory inefficiency caused by the memory reservation. The memory space reserved for a device can be rented for other purposes when the device is not active. For this scheme to be viable, latencies associated with reallocating the memory space should be minimal. Volatile pages are good candidates for such page reallocation since they can be reclaimed immediately as they are needed by the original device. We also provide two optimization techniques, lazy-migration and adaptive-activation. The former increases the lowered utilization of the rental memory by our volatile page allocations, and the latter saves active pages in the rental memory during the reallocation. We implemented our scheme on a smartphone development board with the Android Linux kernel. Our prototype has shown that the time for the return operation is less than 0.77 seconds in the tested cases. We believe that this time is acceptable to end-users in terms of transparency since the time can be hidden in application initialization time. The rental memory also brings throughput increases ranging from 2% to 200% based on the available memory and the applications' memory intensiveness.
Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
ACM Trans. Embed. Comput. Syst.4
2013 Analysis of virtual machine live-migration as a method for power-capping
Jinkyu Jeong, Sung-hun Kim 0005, Hwanju Kim, Joonwon Lee, Euiseong Seo
J. Supercomput.4
2012 DaaC: device-reserved memory as an eviction-based file cache
abstract
Most embedded systems require contiguous memory space to be reserved for each device, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses.
Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
CASES4
2012 Scheduler support for video-oriented multimedia on client-side virtualization
abstract
Virtualization has recently been adopted for client devices to provide strong isolation between services and efficient manageability. Even though multimedia service is not rare for the devices, the virtual machine hosting this service is not guaranteed to receive proper scheduling support from the underlying hypervisor. The quality of multimedia service is often compromised when several virtual machines compete for computing power. This paper presents a new scheduling scheme for the hypervisor to transparently identify if the workload handles multimedia and to provide proper scheduling supports. An implementation of our scheme has shown that the virtual machine hosting a video-oriented application receives propoer CPU scheduling even when other virtual machines host CPU intensive workloads.
Hwanju Kim, Jinkyu Jeong, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng
MMSys4
2012 Workload Characterization and Performance Implications of Large-Scale Blog Servers
abstract
With the ever-increasing popularity of Social Network Services (SNSs), an understanding of the characteristics of these services and their effects on the behavior of their host servers is critical. However, there has been a lack of research on the workload characterization of servers running SNS applications such as blog services. To fill this void, we empirically characterized real-world Web server logs collected from one of the largest South Korean blog hosting sites for 12 consecutive days. The logs consist of more than 96 million HTTP requests and 4.7TB of network traffic. Our analysis reveals the following: (i) The transfer size of nonmultimedia files and blog articles can be modeled using a truncated Pareto distribution and a log-normal distribution, respectively; (ii) user access for blog articles does not show temporal locality, but is strongly biased towards those posted with image or audio files. We additionally discuss the potential performance improvement through clustering of small files on a blog page into contiguous disk blocks, which benefits from the observed file access patterns. Trace-driven simulations show that, on average, the suggested approach achieves 60.6% better system throughput and reduces the processing time for file access by 30.8% compared to the best performance of the Ext4 filesystem.
Myeongjae Jeon, Youngjae Kim 0001, Jeaho Hwang, Joonwon Lee, Euiseong Seo
ACM Trans. Web4
2011 Transparently bridging semantic gap in CPU management for virtualized environments
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee, Seung Ryoul Maeng
J. Parallel Distributed Comput.5
2011 Replicated abstract data types: Building blocks for collaborative applications
Hyun-Gul Roh, Myeongjae Jeon, Jin-Soo Kim 0001, Joonwon Lee
J. Parallel Distributed Comput.4
2011 A comprehensive study of energy efficiency and performance of flash-based SSD
Seon-Yeong Park, Youngjae Kim 0001, Bhuvan Urgaonkar, Joonwon Lee, Euiseong Seo
J. Syst. Archit.4
2011 XHive: Efficient Cooperative Caching for Virtual Machines
abstract
Since a virtual machine independently uses its own caching policy, redundant disk operations exacerbate the I/O virtualization overhead when virtual machines access large amounts of data on shared storage. This paper presents XHive, an efficient cooperative caching system that is implemented at the virtualization layer, for consolidated environments. Our proposed scheme globally manages buffer caches of consolidated virtual machines in order to accommodate a shared working set in machine memory. A singlet, which is a block cached solely by a virtual machine, is preferentially given more chances to be cached in machine memory by XHive, when it is evicted by a guest operating system. For efficient use of limited memory, singlets are cached in memory that is collaboratively donated from idle memory of virtual machines. Our evaluation shows that XHive significantly reduces disk I/O operations for shared working sets, thereby achieving high read performance and scalability. Improved scalability enables a high degree of workload consolidation with respect to virtual machines that have shared working sets.
Hwanju Kim, Heeseung Jo, Joonwon Lee
IEEE Trans. Computers3
2010 Log' version vector: Logging version vectors concisely in dynamic replication
Hyun-Gul Roh, Myeongjae Jeon, Euiseong Seo, Jin-Soo Kim 0001, Joonwon Lee
Inf. Process. Lett.5
2010 KAL: kernel-assisted non-invasive memory leak tolerance with a general-purpose memory allocator
abstract
Abstract Memory leaks are a continuing problem in the software developed with programming languages, such as C and C++. A recent approach adopted by some researchers is to tolerate leaks in the software application and to reclaim the leaked memory by use of specially constructed memory allocation routines. However, such routines replace the usual general‐purpose memory allocator and tend to be less efficient in speed and in memory utilization. We propose a new scheme which coexists with the existing memory allocation routines and which reclaims memory leaks. Our scheme identifies and reclaims leaked memory at the kernel level. There are some major advantages to our approach: (1) the application software does not need to be modified; (2) the application does not need to be suspended while leaked memory is reclaimed; (3) a remote host can be used to identify the leaked memory, thus minimizing impact on the application program's performance; and (4) our scheme does not degrade the service availability of the application while detecting and reclaiming memory leaks. We have implemented a prototype that works with the GNU C library and with the Linux kernel. Our prototype has been tested and evaluated with various real‐world applications. Our results show that the computational overhead of our approach is around 2% of that incurred by the conventional memory allocator in terms of throughput and average response time. We also verified that the prototype successfully suppressed address space expansion caused by memory leaks when the applications are run on synthetic workloads. Copyright © 2010 John Wiley & Sons, Ltd.
Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, Joonwon Lee
Softw. Pract. Exp.6
2010 Transparent Fault Tolerance of Device Drivers for Virtual Machines
abstract
In a consolidated server system using virtualization, physical device accesses from guest virtual machines (VMs) need to be coordinated. In this environment, a separate driver VM is usually assigned to this task to enhance reliability and to reuse existing device drivers. This driver VM needs to be highly reliable, since it handles all the I/O requests. This paper describes a mechanism to detect and recover the driver VM from faults to enhance the reliability of the whole system. The proposed mechanism is transparent in that guest VMs cannot recognize the fault and the driver VM can recover and continue its I/O operations. Our mechanism provides a progress monitoring-based fault detection that is isolated from fault contamination with low monitoring overhead. When a fault occurs, the system recovers by switching the faulted driver VM to another one. The recovery is performed without service disconnection or data loss and with negligible delay by fully exploiting the I/O structure of the virtualized system.
Heeseung Jo, Hwanju Kim, Jae-Wan Jang, Joonwon Lee, Seung Ryoul Maeng
IEEE Trans. Computers4
2010 Superblock FTL: A superblock-based flash translation layer with a hybrid address translation scheme
abstract
In NAND flash-based storage systems, an intermediate software layer called a Flash Translation Layer (FTL) is usually employed to hide the erase-before-write characteristics of NAND flash memory. We propose a novel superblock-based FTL scheme, which combines a set of adjacent logical blocks into a superblock. In the proposed Superblock FTL, superblocks are mapped at coarse granularity, while pages inside the superblock are mapped freely at fine granularity to any location in several physical blocks. To reduce extra storage and flash memory operations, the fine-grain mapping information is stored in the spare area of NAND flash memory. This hybrid address translation scheme has the flexibility provided by fine-grain address translation, while reducing the memory overhead to the level of coarse-grain address translation. Our experimental results show that the proposed FTL scheme significantly outperforms previous block-mapped FTL schemes with roughly the same memory overhead.
Da Woon Jung 0001, Jeong-Uk Kang, Heeseung Jo, Jin-Soo Kim 0001, Joonwon Lee
ACM Trans. Embed. Comput. Syst.5
2010 Dynamic alteration schemes of real-time schedules for I/O device energy efficiency
abstract
Many I/O devices provide multiple power states known as the dynamic power management (DPM) feature. However, activating from sleep state requires significant transition time and this obstructs utilizing DPM in nonpreemptive real-time systems. This article suggests nonpreemptive real-time task scheduling schemes maximizing the effectiveness of the I/O device DPM support. First, we introduce a runtime schedulability check algorithm for nonpreemptive real-time systems that can check whether a modification from a valid schedule is still valid. By using this, we suggest three heuristic algorithms. The first algorithm reorders the execution sequence of tasks according to the similarity of their required device sets. The second one gathers dispersed short idle periods into one long idle period to extend sleeping state of I/O devices and the last one inserts an idle period between two consecutively scheduled tasks to prepare the required devices of a task right before the starting time of the task. The suggested schemes were evaluated for both the real-world task sets and the hypothetical task sets with simulation and the results showed that the suggested algorithms produced better energy efficiency than the existing comparative algorithms.
Euiseong Seo, Seon-Yeong Park, Joonwon Lee
ACM Trans. Embed. Comput. Syst.4
2009 Task-aware virtual machine scheduling for I/O performance
abstract
The use of virtualization is progressively accommodating diverse and unpredictable workloads as being adopted in virtual desktop and cloud computing environments. Since a virtual machine monitor lacks knowledge of each virtual machine, the unpredictableness of workloads makes resource allocation difficult. Particularly, virtual machine scheduling has a critical impact on I/O performance in cases where the virtual machine monitor is agnostic about the internal workloads of virtual machines. This paper presents a task-aware virtual machine scheduling mechanism based on inference techniques using gray-box knowledge. The proposed mechanism infers the I/O-boundness of guest-level tasks and correlates incoming events with I/O-bound tasks. With this information, we introduce partial boosting, which is a priority boosting mechanism with task-level granularity, so that an I/O-bound task is selectively scheduled to handle its incoming events promptly. Our technique focuses on improving the performance of I/O-bound tasks within heterogeneous workloads by lightweight mechanisms with complete CPU fairness among virtual machines. All implementation is confined to the virtualization layer based on the Xen virtual machine monitor and the credit scheduler. We evaluate our prototype in terms of I/O performance and CPU fairness over synthetic mixed workloads and realistic applications.
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee
VEE5
2009 Catching two rabbits: adaptive real-time support for embedded Linux
abstract
Abstract The trend of digital convergence makes multitasking common in many digital electronic products. Some applications in those systems have inherent real‐time properties, while many others have few or no timeliness requirements. Therefore the embedded Linux kernels, which are widely used in those devices, provide real‐time features in many forms. However, providing real‐time scheduling usually induces throughput degradation in heavy multitasking due to the increased context switches. Usually the throughput degradation becomes a critical problem, since the performance of the embedded processors is generally limited for cost, design and energy efficiency reasons. This paper proposes schemes to lessen the throughput degradation, which is from real‐time scheduling, by suppressing unnecessary context switches and applying real‐time scheduling mechanisms only when it is necessary. Also the suggested schemes enable the complete priority inheritance protocol to prevent the well‐known priority inversion problem. We evaluated the effectiveness of our approach with open‐source benchmarks. By using the suggested schemes, the throughput is improved while the scheduling latency is kept same or better in comparison with the existing approaches. Copyright © 2008 John Wiley & Sons, Ltd.
Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee
Softw. Pract. Exp.5
2008 Context-aware address translation for high performance SMP cluster system
abstract
User-level communication allows an application process to access the network interface directly. Bypassing the kernel requires that a user process accesses the network interface using its own virtual address which should be translated to a physical address. A small caching structure which is similar to the hardware TLB on the host processor has been used to cache the mappings between virtual and physical addresses on the network interface memory. In this study, we propose a new TLB architecture for the network interface. The proposed architecture splits an original caching structure into as many partitions as the number of processors on the SMP system and assigns a separate partition to each application process. In addition, the architecture becomes aware of user contexts and switches the content of caching structure in accordance with context switching. According to our experiments, our scheme achieves significant reduction in application execution time compared to the previous approach.
Moon-Sang Lee, Joonwon Lee, Seung Ryoul Maeng
CLUSTER2
2008 Guest-Aware Priority-Based Virtual Machine Scheduling for Highly Consolidated Server
Hwanju Kim, Myeongjae Jeon, Euiseong Seo, Joonwon Lee
Euro-Par5
2008 TSB: A DVS algorithm with quick response for general purpose operating systems
Euiseong Seo, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee
J. Syst. Archit.4
2008 A CFD-Based Tool for Studying Temperature in Rack-Mounted Servers
abstract
Temperature-aware computing is becoming more important in design of computer systems as power densities are increasing and the implications of high operating temperatures result in higher failure rates of components and increased demand for cooling capability. Computer architects and system software designers need to understand the thermal consequences of their proposals, and develop techniques to lower operating temperatures to reduce both transient and permanent component failures. Recognizing the need for thermal modeling tools to support those researches, there has been work on modeling temperatures of processors at the micro-architectural level which can be easily understood and employed by computer architects for processor designs. However, there is a dearth of such tools in the academic/research community for undertaking architectural/systems studies beyond a processor - a server box, rack or even a machine room. In this paper we presents a detailed 3-dimensional computational fluid dynamics based thermal modeling tool, called ThermoStat, for rack-mounted server systems. We conduct several experiments with this tool to show how different load conditions affect the thermal profile, and also illustrate how this tool can help design dynamic thermal management techniques. We propose reactive and proactive thermal management for rack mounted server and isothermal workload distribution for rack.
Jeonghwan Choi, Youngjae Kim 0001, Anand Sivasubramaniam, Jelena Srebric, Qian Wang 0029, Joonwon Lee
IEEE Trans. Computers6
2008 ScaleFFS: A scalable log-structured flash file system for mobile multimedia systems
abstract
NAND flash memory has become one of the most popular storage media for mobile multimedia systems. A key issue in designing storage systems for mobile multimedia systems is handling large-capacity storage media and numerous large files with limited resources such as memory. However, existing flash file systems, including JFFS2 and YAFFS in particular, exhibit many limitations in addressing the storage capacity of mobile multimedia systems. In this article, we design and implement a scalable flash file system, called ScaleFFS, for mobile multimedia systems. ScaleFFS is designed to require only a small fixed amount of memory space and to provide fast mount time, even if the file system size grows to more than tens of gigabytes. The measurement results show that ScaleFFS can be instantly mounted regardless of the file system size, while achieving the same write bandwidth and up to 22% higher read bandwidth compared to JFFS2.
Da Woon Jung 0001, Jaegeuk Kim, Jin-Soo Kim 0001, Joonwon Lee
ACM Trans. Multim. Comput. Commun. Appl.4
2008 Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors
abstract
Multicore processors deliver a higher throughput at lower power consumption than unicore processors. In the near future, they will thus be widely used in mobile real-time systems. There have been many research on energy-efficient scheduling of real-time tasks using DVS. These approaches must be modified for multicore processors, however, since normally all the cores in a chip must run at the same performance level. Thus, blindly adopting existing DVS algorithms that do not consider the restriction will result in a waste of energy. This article suggests Dynamic Repartitioning algorithm based on existing partitioning approaches of multiprocessor systems. The algorithm dynamically balances the task loads of multiple cores to optimize power consumption during execution. We also suggest Dynamic Core Scaling algorithm, which adjusts the number of active cores to reduce leakage power consumption under low load conditions. Simulation results show that Dynamic Repartitioning can produce energy savings of about 8 percent even with the best energy-efficient partitioning algorithm. The results also show that Dynamic Core Scaling can reduce energy consumption by about 26 percent under low load conditions.
Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Joonwon Lee
IEEE Trans. Parallel Distributed Syst.4
2007 Domain Level Page Sharing in Xen Virtual Machine Systems
Myeongjae Jeon, Euiseong Seo, Joonwon Lee
APPT4
2007 A group-based wear-leveling algorithm for large-capacity flash memory storage systems
abstract
Although NAND flash memory has become one of the most popular storage media for portable devices, it has a serious problem with respect to lifetime. Each block of NAND flash memory has a limited number of program/erase cycles, usually 10,000–100,000, and data in a block become unreliable after the limit. For this reason, distributing erase operations evenly across the whole flash memory media is an important concern in designing flash memory storage systems. In this paper, we propose a memory-efficient group-based wear-leveling algorithm. Our group-based algorithm achieves a small memory footprint by grouping several logically sequential blocks and managing only the summary information for each group. We also propose an effective group summary structure and a method to reduce unnecessary wearleveling operations in order to enhance the wear-leveling performance. The evaluation results show that our group-based algorithm consumes only 8.75 % of memory space compared to the previous scheme that manages per-block information, while showing roughly the same wear-leveling performance.
Da Woon Jung 0001, Yoon-Hee Chae, Heeseung Jo, Jin-Soo Kim 0001, Joonwon Lee
CASES5
2007 Modeling and Managing Thermal Profiles of Rack-mounted Servers with ThermoStat
abstract
High power densities and the implications of high operating temperatures on the failure rates of components are key driving factors of temperature-aware computing. Computer architects and system software designers need to understand the thermal consequences of their proposals, and develop techniques to lower operating temperatures to reduce both transient and permanent component failures. Tools for understanding temperature ramifications of designs have been mainly restricted to industry for studying packaging and cooling mechanisms, with little access to such toolsets for academic researchers. Developing such tools is an arduous task since it usually requires cross-cutting areas of expertise spanning architecture, systems software, thermodynamics, and cooling systems. Recognizing the need for such tools, there has been work on modeling temperatures of processors at the micro-architectural level which can be easily understood and employed by computer architects for processor designs. However, there is a dearth of such tools in the academic/research community for undertaking architectural/systems studies beyond a processor - a server box, rack or even a machine room. This paper presents a detailed 3-dimensional computational fluid dynamics based thermal modeling tool, called ThermoStat, for rack-mounted server systems. Using this tool, we model a 20 (each with dual Xeon processors) node rack-mounted server system, and validate it with over 30 temperature sensor measurements at different points in the servers/rack. We conduct several experiments with this tool to show how different load conditions affect the thermal profile, and also illustrate how this tool can help design dynamic thermal management techniques
Jeonghwan Choi, Youngjae Kim 0001, Anand Sivasubramaniam, Jelena Srebric, Qian Wang 0029, Joonwon Lee
HPCA6
2007 A multi-channel architecture for high-performance NAND flash-based storage system
Jeong-Uk Kang, Jin-Soo Kim 0001, Chanik Park, Hyoungjun Park, Joonwon Lee
J. Syst. Archit.5
2007 PABC: Power-Aware Buffer Cache Management for Low Power Consumption
abstract
Power consumed by memory systems becomes a serious issue as the size of the memory installed increases. With various low power modes that can be applied to each memory unit, the operating system can reduce the number of active memory units by collocating active pages onto a few memory units. This paper presents a memory management scheme based on this observation, which differs from other approaches in that all of the memory space is considered, while previous methods deal only with pages mapped to user address spaces. The buffer cache usually takes more than half of the total memory and the pages access patterns are different from those in user address spaces. Based on an analysis of buffer cache behavior and its interaction with the user space, our scheme achieves up to 63 percent more power reduction. Migrating a page to a different memory unit increases memory latencies, but it is shown to reduce the power consumed by an additional 4.4 percent
Min Lee, Euiseong Seo, Joonwon Lee, Jin-Soo Kim 0001
IEEE Trans. Computers3
2006 CFLRU: a replacement algorithm for flash memory
abstract
In most operating systems which are customized for disk-based storage system, the replacement algorithm concerns only the number of memory hits. However, flash memory has different read and write cost in the aspects of time and energy so the replacement algorithm with flash memory should consider not only the hit count but also the replacement cost caused by selecting dirty victims. The replacement cost of dirty page is higher than that of clean page with regard to both access time and energy consumption. In this paper, we propose the Clean-First LRU (CFLRU) replacement algorithm that exploits the characteristics of flash memory. CFLRU splits the LRU list into the working region and the clean-first region and adopts a policy that evicts clean pages preferentially in the clean-first region until the number of page hits in the working region is preserved in a suitable level. Using the trace-driven simulation, the proposed algorithm reduces the average replacement cost by 28.4% in swap system and by 26.2% in buffer cache, compared with LRU algorithm. We also implement the CFLRU algorithm in the Linux kernel and present some optimization issues.
Seon-Yeong Park, Da Woon Jung 0001, Jeong-Uk Kang, Jin-Soo Kim 0001, Joonwon Lee
CASES5
2006 A superblock-based flash translation layer for NAND flash memory
abstract
In NAND flash-based storage systems, an intermediate software layer called a flash translation layer (FTL) is usually employed to hide the erase-before-write characteristics of NAND flash memory. This paper proposes a novel superblockbased FTL scheme, which combines a set of adjacent logical blocks into a superblock. In the proposed FTL scheme, superblocks are mapped at coarse granularity, while pages inside the superblock are mapped freely at fine granularity to any location in several physical blocks. To reduce extra storage and flash memory operations, the fine-grain mapping information is stored in the spare area of NAND flash memory. This hybrid mapping technique has the flexibility provided by fine-grain address translation, while reducing the memory overhead to the level of coarse-grain address translation. Our experimental results show that the proposed FTL scheme decreases the garbage collection overhead up to 40 % compared to previous FTL schemes.
Jeong-Uk Kang, Heeseung Jo, Jin-Soo Kim 0001, Joonwon Lee
EMSOFT4
2006 Dynamic Repartitioning of Real-Time Schedule on a Multicore Processor for Energy Efficiency
Euiseong Seo, Yongbon Koo, Joonwon Lee
EUC3
2006 Runtime feasibility check for non-preemptive real-time periodic tasks
Joonwon Lee, Jin-Soo Kim 0001
Inf. Process. Lett.2
2006 Optimal intratask dynamic voltage-scaling technique and its practical extensions
abstract
This paper presents a set of comprehensive techniques for the intratask voltage-scheduling problem to reduce energy consumption in hard real-time tasks of embedded systems. Based on the execution profile of the task, a voltage-scheduling technique that optimally determines the operating voltages to individual basic blocks in the task is proposed. The obtained voltage schedule guarantees minimum average energy consumption. The proposed technique is then extended to solve practical issues regarding transition overheads, which are totally or partially ignored in the existing approaches. Finally, a technique involving a novel extension of our optimal scheduler is proposed to solve the scheduling problem in a discretely variable voltage environment. We also present a novel voltage set-up technique to determine each voltage level for customizable systems-on-chips (SoCs) with discretely variable voltages. In summary, it is confirmed from experiments that the proposed optimal scheduling technique reduces energy consumption by 20.2% over that of one of the state-of-the-art schedulers (Shin and Kim, 2001) and, further, the extended technique in a discrete-voltage environment reduces energy consumption by 45.3% on average.
Jaewon Seo, Joonwon Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 Design and implementation of an OGSI-compliant Grid broker service
abstract
Grid computing promises the ability to share geographically and organizationally distributed resources to increase effective computational power and resource utilization. However, for the Grid computing to be successful, it is very important to provide middleware services that assist Grid users to easily interact with Grid environments. In this paper, we have designed and implemented a new general-purpose OGSI-compliant Grid resource broker service to hide the underlying complexity of the Grid resources from Grid users and to meet not only Grid user requirements but also resource owner policies. It focuses on discovering and scheduling dynamic resources scattered across multiple organizations. Furthermore, it can be integrated with various scheduling services. We also present experimental results and demonstrate the effectiveness of our Grid broker service.
Youn-Seok Kim, Jung-Lok Yu, Jae-Gyoon Hahm, Jin-Soo Kim 0001, Joonwon Lee
CCGRID5
2001 Adaptive Prefetching Technique for Shared Virtual Memory
abstract
Though shared virtual memory (SVM) systems promise low cost solutions for high performance computing, they suffer from long memory latencies. These latencies are usually caused by repetitive invalidation on shared data. Since shared data are accessed through synchronizations and the patterns by which threads synchronize are repetitive, a prefetching scheme based on such repetitiveness would reduce memory latencies. Based on this observation, we propose a prefetching technique which predicts future access behavior by analyzing access history per synchronization variable. Our technique was evaluated on an 8-node SVM system using the SPLASH-2 benchmark. The results show that our technique could achieve 34%-45% reduction in memory access latencies.
Sang-Kwon Lee, Hee-Chul Yun, Joonwon Lee, Seung Ryoul Maeng
CCGRID3
2001 An Efficient Lock Protocol for Home-Based Lazy Release Consistency
abstract
Home-based lazy release consistency (HLRC) shows poor performance on lock based applications because of two reasons: a whole page is fetched on a page fault while actual modification is much smaller; and a home is at the fixed location while the access pattern is migratory. We present an efficient lock protocol for HLRC. In this protocol, the pages that are expected to be used by the acquirer are selectively updated using diffs. The diff accumulation problem is minimized by limiting the size of diffs to be sent for each page. Our protocol reduces the number of page faults inside critical sections because pages can be updated by applying locally stored diffs. This reduction yields the reduction of average lock waiting time and the reduction of message amount. The experiment with five applications shows that our protocol archives 2%-40% speedup against base HLRC for four applications.
Hee-Chul Yun, Sang-Kwon Lee, Joonwon Lee, Seung Ryoul Maeng
CCGRID3
2000 A scheduling policy for preserving cache locality in a multiprogrammed system
Inbum Jung, Jongwoong Hyun, Joonwon Lee
J. Syst. Archit.3
2000 Broadcast directory: A scalable cache coherent architecture for mesh-connected multiprocessors
Yunseok Rhee, Joonwon Lee
J. Syst. Archit.2
1999 A Scalable Cache Coherent Scheme Exploiting Wormhole Routing Networks
abstract
Large scale shared memory multiprocessors favor a directory-based cache coherence scheme for its scalability. The directory space needed to record the information for sharers has a complexity of /spl Theta/(N/sup 2/) when a full-mapped vector is used for an N-node system. Though this overhead can be reduced by limiting the directory size assuming that the sharing degree is small, it will experience significant inefficiency when a data is widely shared. In this paper, we propose a new directory scheme and a cache coherence scheme based on it for a mesh interconnection. Deterministic and wormhole routing enables a pointer to represent a set of nodes. Also a message traversing on the mesh performs a broadcast mission to a set of nodes without extra traffic, which can be utilized for the cache coherence problem. Only a slight change on a generic router is needed to implement our scheme. This scheme is also applicable to any k-ary n-cube networks including a mesh.
Yunseok Rhee, Joonwon Lee
HPCA2
1997 A Virtual-Physical On-Chip Cache for Shared Memory Multiprocessors
Joonwon Lee
Euro-Par2
1997 A partitioned on-chip virtual cache for fast processors
Joonwon Lee, Seungkyu Park
J. Syst. Archit.2
1996 Prefetching scheme for image processing on shared memory multiprocessors
abstract
We propose a prefetching scheme for image processing which fully exploits regular memory references, and hides or reduces latencies from memory to cache in shared memory multiprocessors. Since the image processing applications usually require quite huge memory references, it is obvious that the memory latencies dominate the overall execution time. Though software-controlled prefetching can be also an alternative to alleviate such latency impact, it usually suffers from large overheads due to the prefetch instructions themselves which would offset the benefit obtained from prefetching. This paper proposes a new software-controlled prefetching scheme that prefetches multiple blocks using a prefetch instruction. It is designed in consideration of very regular memory references which is intrinsic to most image processing applications. Our scheme reduces almost all the memory stall time for loads and significantly the overall execution time. Simulation applications have shown a reduction in the execution time by 10-21% compared to the original methods without prefetching.
Yunseok Rhee, Joonwon Lee
ICIP (2)2
1996 Cache-Based Synchronization in Shared Memory Multiprocessors
Umakishore Ramachandran, Joonwon Lee
J. Parallel Distributed Comput.2
1991 Architectural Primitives for a Scalable Shared Memory Multiprocessor
abstract
Article Free Access Share on Architectural primitives for a scalable shared memory multiprocessor Authors: Joonwon Lee College of Computing, Georgia Institute of Technology, Atlanta, Georgia College of Computing, Georgia Institute of Technology, Atlanta, GeorgiaView Profile , Umakishore Ramachandran College of Computing, Georgia Institute of Technology, Atlanta, Georgia College of Computing, Georgia Institute of Technology, Atlanta, GeorgiaView Profile Authors Info & Claims SPAA '91: Proceedings of the third annual ACM symposium on Parallel algorithms and architecturesJune 1991 Pages 103–114https://doi.org/10.1145/113379.113389Published:01 June 1991Publication History 4citation145DownloadsMetricsTotal Citations4Total Downloads145Last 12 Months4Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Joonwon Lee, Umakishore Ramachandran
SPAA1
1990 Synchronization with Multiprocessor Caches
abstract
Introducing private caches in bus-based shared memory multiprocessors leads to the cache consistency problem since there may be multiple copies of shared data. However, the ability to snoop on the bus coupled with the fast broadcast capability allows the design of special hardware support for synchronization. We present a new lock-based cache scheme which incorporates synchronization into the cache coherency mechanism. With this scheme high-level synchronization primitives as well as low-level ones can be implemented without excessive overhead. Cost functions for well-known synchronization methods are derived for invalidation schemes, write update schemes, and our lock-based scheme. To accurately predict the performance implications of the new scheme, a new simulation model is developed embodying a widely accepted paradigm of parallel programming. It is shown that our lock-based protocol outperforms existing cache protocols.
Joonwon Lee, Umakishore Ramachandran
ISCA1