Chuck Yoo

dblp:12/4989 · DBLP profile ↗
← Back
61ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0002-1115-1862ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 7 since 2021Computer networks · 16 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prediction-based GPU sharing for distributed training
abstract
• Formulate the inconsistent JCT problem using gSLA for the first time. • Design a new JCT increase prediction model and job scheduler for GPU sharing. • Achieve up to 47.3× better gSLA satisfaction and 50× lower gSLA excess ratio. • Improve JCT and GPU efficiency by ∼ 60% and ∼ 44% over existing methods. • Demonstrate TensorShare’s effectiveness in improving gSLA and JCT for unseen jobs. GPU sharing aims to enhance the efficiency of GPU utilization by running distributed deep learning training jobs concurrently. However, GPU sharing poses a significant challenge: the increase in job completion time (JCT) caused by interference between jobs is inconsistent, complicating job scheduling. Our experiments reveal that the degree of JCT increase varies by as much as ∼ 3.7 × . While previous studies have analyzed this JCT inconsistency problem, none of them have been able to minimize the inconsistency. We propose TensorShare, a proactive GPU sharing technique that leverages a deep learning model to predict the extent of JCT increase. This study defines a new metric, called GPU SLA, which represents the upper threshold of JCT increase. TensorShare then introduces a novel scheduler that proactively identifies which jobs meet GPU SLA while minimizing the JCT increase. Our evaluation shows that TensorShare improves GPU SLA satisfaction rates by 26.1 × –47.3 × and reduces the JCT increase by 37%–60%. Furthermore, we evaluate TensorShare with large language models that are not included in training TensorShare’s prediction model, achieving ∼ 7 × and ∼ 10.3 × improvements in GPU SLA satisfaction and JCT inconsistency, respectively.
Changyong Shin, Younghun Go, Yeonho Yoo, Jae-Hyun Hwang, Gyeongsik Yang, Chuck Yoo
Future Gener. Comput. Syst.7
2025 LLM-based Interactive Coding Education via Predictive Query Management and Student-Centered Fine-Tuning: Design and Implementation with 1500-Student Class Data
abstract
Large-scale university courses face significant challenges. Teaching assistants are overwhelmed by the large number of student questions, limiting their ability to provide detailed and individualized support. As a result, students-especially those who are struggling-receive less tailored assistance, further widening gaps in academic performance. To address these challenges, we propose student-centered AI learning assistant (SCALA), a large language model (LLM)-based interactive tutoring system that incorporates student needs and learning expectations. SCALA consists of two main components, i.e., predictive query management and student-centered fine tuning. The first component anticipates common student questions via LLM agent debate. Each agent interacts using a combination of lecture content and student interactions in chat logs, collaboratively predicting what the students will likely ask. This fosters learning in students by generating and presenting relevant queries that guide their learning. On the other hand, the second component is fine-tuned on a 14k-question Python-tutoring dataset, curated based on in-depth student interviews to reflect real learning expectations. Our real-world experiments with 1500-student large-scale Python classes demonstrate that SCALA delivers more helpful and accurate responses compared to closed-form models (e.g., GPT-4o), while significantly reducing latency.
Geonjae Youn, Jonghoon Lee 0001, Joongheon Kim, Chuck Yoo
CIKM4
2025 Parameter-Efficient 12-Lead ECG Reconstruction from a Single Lead
Yeonho Yoo, Jinkyu Kim 0001, Dosun Lim, Gyeongsik Yang, Chuck Yoo
MICCAI (2)6
2024 Harmonia: Accurate Federated Learning with All-Inclusive Dataset
abstract
Federated learning (FL) is an appealing model training technique that utilizes heterogeneous datasets and user devices, ensuring user data privacy. Existing FL research proposed device selection schemes to balance the computing speeds of devices. However, we observe that these schemes compromise prediction accuracy by ~57. 7 %. To solve this problem, we present Harmonia that enhances prediction accuracy, while also balancing the diverse computing speeds of devices. Our evaluation shows that Harmonia improves prediction accuracy by ~ 1.7 x over existing schemes.
Wonmi Choi, Juyoung Ahn, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang
CLOUD4
2024 Predictive Placement of Geo-Distributed Blockchain Nodes for Performance Guarantee
abstract
Blockchain-as-a-service (BaaS) in cloud datacenters is gaining widespread attention due to its high performance and privacy. However, existing BaaS solutions lack a method for deciding the proper placement of blockchain nodes across virtual machines in worldwide datacenters to achieve desired performance. Our motivating experiments show that transaction processing performance (TPS) varies ~31.6% depending on the placements. To provide an automatic placement solution for BaaS, we propose Cyan that predicts the TPS for blockchain node placements. Our evaluations on Google Cloud Platform demonstrate that Cyan improves the TPS guarantee ~2.39x compared to existing techniques.
Yeonho Yoo, Chuck Yoo, Gyeongsik Yang
CLOUD3
2024 Intelligent Packet Processing for Performant Containers in IoT
abstract
This article explores the computing and communication overhead of network processing in Internet of Things (IoT) devices, focusing on containers, a major building block for the edge computing. Our experiments reveal that containers on IoT devices suffer$\sim 2.6\times $higher CPU usage for SoftIRQ processing, ~59% less network throughput, and$2\times $higher per-packet latency on average than native processes. While several existing studies enhance networking performance, they often sacrifice interoperability by requiring special hardware or modifying networking semantics or APIs. Thus, we design and implement a kernel networking accelerator, called SCON, that maintains interoperability, crucial for IoT devices. SCON addresses major bottlenecks in container networking through system-level profiling. We evaluate SCON with three types of IoT devices. On the Raspberry Pi 4, SCON reduces the latencies of major IoT application protocols (e.g., HTTP and MQTT) by$\sim 10\times $, achieving a similar level of latency to the native process. Further analysis shows that SCON reduces CPU usage for SoftIRQ processing by ~26%. We also report similar improvements on the other two IoT devices. Our conclusion is that SCON is unique in significantly reducing the computing and communication overhead of container networking in IoT devices while maintaining interoperability. Furthermore, it works consistently across different types of devices, whether wired or wireless, and regardless of heavy or sporadic traffic.
Wonmi Choi, Yeonho Yoo, Kyungwoon Lee, Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Gyeongsik Yang, Chuck Yoo
IEEE Internet Things J.8
2023 Selective Preemption of Distributed Deep Learning Training
abstract
As more distributed deep learning (DDL) jobs run in public clouds, their effective scheduling becomes a major challenge. Current studies prioritize the execution of jobs with less remaining time, which is known to be the best in reducing average job completion time (JCT). However, we observe that this approach does not work when the preemption for pausing and loading jobs weighs in; sometimes, the preemption overheads of DDL jobs take up to hundreds of seconds. This results in very ineffective scheduling, so in some cases, the first-in-first-out policy performs much better. This paper proposes a new scheduling framework called Xion that takes into account the preemption overheads and only preempts DDL jobs when it is beneficial. Our evaluation results demonstrate that Xion effectively reduces the average JCT by 19% and improves the waiting time by 1.64×.
Younghun Go, Changyong Shin, Jeunghwan Lee, Yeonho Yoo, Gyeongsik Yang, Chuck Yoo
CLOUD6
2023 SegaNet: An Advanced IoT Cloud Gateway for Performant and Priority-Oriented Message Delivery
abstract
With the tremendous growth of IoT, the role of IoT cloud gateways in facilitating communication between IoT devices and the cloud has become more important than ever before. Most previous studies have focused on developing interoperability between IoT and cloud to accommodate various radio protocols. However, they have often neglected the performance aspect of the IoT cloud gateway, leaving users with limited options: either purchasing multiple gateways or connecting only a small number of IoT devices. Through our comprehensive measurements and analysis, we identified five key issues in IoT cloud gateways related to high latency, CPU bottlenecks, inefficient network stacks on ARM, substantial encryption overhead, and the lack of priority support. To address these issues, we propose a new IoT cloud gateway - SegaNet. We carefully design with 1) multiple agents management, 2) efficient TLS encryption, and 3) priority-oriented message delivery. Our prototype evaluation shows up to 16.7 × lower latency and 4.5 × lower CPU consumption than gateways of the existing IoT-cloud ecosystem.
Yeonho Yoo, Zhixiong Niu, Chuck Yoo, Peng Cheng 0005, Yongqiang Xiong
APNet3
2023 Control Channel Isolation in SDN Virtualization: A Machine Learning Approach
abstract
Performance isolation is an essential property that network virtualization must provide for clouds. This study addresses the performance isolation of the control plane in virtualized software-defined networking (SDN), which we call control channel isolation. First, we report that the control channel isolation is seriously broken in the existing network hypervisor in that the end-to-end control latency grows by up to 15 x as the number of virtual switches increases. This jeopardizes the key network operations, such as routing, in datacenters. To address this issue, we take a machine learning approach that learns from the past control traffic as time-series data. We propose a new network hypervisor, Meteor, that designs an LSTM autoencoder to predict the control traffic per virtual switch. Our evaluation results show that Meteor improves the processing latency per control message by up to 12.7x. Furthermore, Meteor reduces the end-to-end control latency by up to 73.7%, which makes it comparable to the non-virtualized SDN.
Yeonho Yoo, Gyeongsik Yang, Changyong Shin, Jeunghwan Lee, Chuck Yoo
CCGrid5
2023 Autothrottle: Satisfying Network Performance Requirements for Containers
abstract
This article investigates how to satisfy network performance requirements that are crucial in achieving the service level objectives (SLOs) in clouds. Traditional techniques for network performance management have a limited ability to satisfy the network SLOs. Our in-depth analysis reveals that the fundamental reason comes from decoupling of the CPU scheduler and the network traffic controller as the current CPU scheduler is not aware of such network requirements but only provides a fair-share amount of CPU to all containers. Thus, the container cannot perform the amount of network processing as needed to satisfy its SLO when the CPU allocation is insufficient. In this article, we propose Autothrottle that dynamically adjusts the CPU allocation for the containers to satisfy their network SLOs. The key element of Autothrottle is a throttle algorithm that autonomously determines the amount of CPU for each container needed to satisfy the requirement. We implement Autothrottle in the Linux kernel and evaluate it with massive real-world workloads such as Apache Kafka. Our evaluation results show that Autothrottle successfully satisfies the given network SLO only with a 2% gap while the existing scheme achieves 20% less than the SLO. We further observe that Autothrottle also reduces the CPU overhead in network processing by 19%, improving the network throughput by 27% compared to the existing scheme.
Kyungwoon Lee, Kwanhoon Lee, Hyunchan Park, Jae-Hyun Hwang, Chuck Yoo
IEEE Trans. Cloud Comput.5
2023 Accurate and Efficient Monitoring for Virtualized SDN in Clouds
abstract
This article presents V-Sight, a network monitoring framework for programmable virtual networks in clouds. Network virtualization based on software-defined networking (SDN-NV) in clouds makes it possible to realize programmable virtual networks; consequently, this technology offers many benefits to cloud services for tenants. However, to the best of our knowledge, network monitoring, which is a prerequisite for managing and optimizing virtual networks, has not been investigated in the context of SDN-NV systems. As the first framework for network monitoring in SDN-NV, we identify three challenges: non-isolated and inaccurate statistics, high monitoring delay, and excessive control channel consumption for gathering statistics. To address these challenges, V-Sight introduces three key mechanisms: 1) statistics virtualization for isolated statistics, 2) transmission disaggregation for reduced transmission delay, and 3) pCollector aggregation for efficient control channel consumption. The evaluation results reveal that V-Sight successfully provides accurate and isolated statistics while reducing the monitoring delay and control channel consumption in orders of magnitude. We also show that V-Sight can achieve a data plane throughput close to that of non-virtualized SDN.
Gyeongsik Yang, Yeonho Yoo, Minkoo Kang 0001, Heesang Jin, Chuck Yoo
IEEE Trans. Cloud Comput.5
2023 TeaVisor: Network Hypervisor for Bandwidth Isolation in SDN-NV
abstract
We introduce TeaVisor that provides bandwidth isolation guarantee for network virtualization (NV) based on software-defined networking (SDN). SDN-based NV (SDN-NV) offers many benefits to clouds, such as topology and address virtualization while allowing flexible resource provisioning, control, and monitoring on virtual networks. In SDN-NV, however, routing is done by tenants independently; thus, existing studies have difficulties in bandwidth isolation guarantee due to the overloaded link problem. Bandwidth isolation guarantee is essential for providing stable and reliable throughput on network services in SDN-NV. Without bandwidth isolation guarantee, tenants suffer degraded service qualities and significant loss in revenue. To address this problem, we design and implement TeaVisor in three components: path virtualization, bandwidth reservation, and path establishment. Through extensive experiments, TeaVisor shows that bandwidth isolation is guaranteed with near-zero errors, which is three orders of magnitude better than existing studies. In addition, TeaVisor guarantees the minimum and maximum bandwidth at the same time. We also present an overhead analysis of TeaVisor in control traffic and memory consumption.
Yeonho Yoo, Gyeongsik Yang, Jeunghwan Lee, Changyong Shin, Hoseok Kim, Chuck Yoo
IEEE Trans. Cloud Comput.6
2023 Network SLO-aware container scheduling in Kubernetes
Eunsook Kim, Kyungwoon Lee, Chuck Yoo
J. Supercomput.3
2023 Machine Learning-Based Prediction Models for Control Traffic in SDN Systems
abstract
This article presentsElixir, an automated prediction model formulation framework for control traffic using machine learning. Control traffic is vital in software-defined networking (SDN) systems because it determines the reliability and scalability of the entire system. Various studies have sought to design control traffic prediction models for the proper provisioning and planning of SDN systems. However, previously proposed models are based on descriptive modeling, well-suited for only specific SDN system instances. Furthermore, these models exhibit poor accuracy (errors of up to 85%) because of the heterogeneity of SDN systems. Because descriptive modeling requires a significant amount of human contemplation, it is impossible to formulate adequate prediction models for countless SDN system instances.Elixiraddresses this problem by applying machine learning.Elixirstarts the model formulation through self-generated datasets. Then,Elixirsearches prediction models to fit the accuracy for respective SDN systems. Also,Elixirpicks robust models that exhibit reasonable accuracy even in a network topology that differs from the topology used for model training. We evaluate theElixirframework on nine heterogeneous SDN systems. As a key outcome,Elixirsignificantly reduces prediction errors, achieving up to 10.6× improvement compared to the previous model for control traffic throughput of OpenDayLight controller.
Yeonho Yoo, Gyeongsik Yang, Changyong Shin, Chuck Yoo
IEEE Trans. Serv. Comput.5
2022 Xonar: Profiling-based Job Orderer for Distributed Deep Learning
abstract
Deep learning models have a wide spectrum of GPU execution time and memory size. When running distributed training jobs, however, their GPU execution time and memory size have not been taken into account, which leads to the high variance of job completion time (JCT). Moreover, the jobs often run into the GPU out-of-memory (OoM) problem so that the unlucky job has to restart all over. To address the problems, we propose Xonar to profile the deep learning jobs and order them in the queue. The experiments show that Xonar with TensorFlow v1.6 reduces the tail JCT by 44% with the OoM problem eliminated.
Changyong Shin, Gyeongsik Yang, Yeonho Yoo, Jeunghwan Lee, Chuck Yoo
CLOUD5
2021 Bandwidth Isolation Guarantee for SDN Virtual Networks
abstract
We introduce TeaVisor, which provides bandwidth isolation guarantee for software-defined networking (SDN)-based network virtualization (NV). SDN-NV provides topology and address virtualization while allowing flexible resource provisioning, control, and monitoring of virtual networks. However, to the best of our knowledge, the bandwidth isolation guarantee, which is essential for providing stable and reliable throughput on network services, is missing in SDN-NV. Without bandwidth isolation guarantee, tenants suffer degraded service quality and significant revenue loss. In fact, we find that the existing studies on bandwidth isolation guarantees are insufficient for SDN-NV. With SDN-NV, routing is performed by tenants, and existing studies have not addressed the overloaded link problem. To solve this problem, TeaVisor designs three components: path virtualization, bandwidth reservation, and path establishment, which utilize multipath routing. With these, TeaVisor achieves the bandwidth isolation guarantee while preserving the routing of the tenants. In addition, TeaVisor guarantees the minimum and maximum amounts of bandwidth simultaneously. We fully implement TeaVisor, and the comprehensive evaluation results show that near-zero error rates on achieving the bandwidth isolation guarantee. We also present an overhead analysis of control traffic and memory consumption.
Gyeongsik Yang, Yeonho Yoo, Minkoo Kang 0001, Heesang Jin, Chuck Yoo
INFOCOM5
2021 A Case for SDN-based Network Virtualization
abstract
Network virtualization (NV) becomes an essential technology in cloud computing that isolates network flows for tenants. However, because existing NV technologies like overlay do not enable tenants to directly program (i.e., provision, control, and monitor) network resources, software-defined networking (SDN)-based NV (SDN-NV) has been proposed. Despite its great benefits, SDN-NV has been believed to bring considerable overheads due to the network hypervisor (NH). However, to date, there is no definite performance evaluation that proves the overheads of SDN-NV. To this end, this paper comprehensively investigates the performance and overheads of SDN-NV. Our experiment results reveal that SDN-NV provides the data plane performance comparable to or even better (up to 10.5× better TCP throughput) than the existing NV technologies. Also, the results on NH show that its overheads remain mostly constant, even when the number of switches, virtual networks, or network flows increases. In short, our evaluation indicates that the overhead of SDN-NV should not deter its practical use in datacenters.
Gyeongsik Yang, Changyong Shin, Yeonho Yoo, Chuck Yoo
MASCOTS4
2020 TensorExpress: In-Network Communication Scheduling for Distributed Deep Learning
abstract
TensorExpress provides in-network communication scheduling for distributed deep learning (DDL). In cloud-based DDL, parameter communication over a network is a key bottleneck. Previous studies proposed tensor packet reordering approaches to reduce network blocking time. However, network contention still exists in DDL. TensorExpress mitigates network contention and reduces overall training time. It schedules tensor packets in-network using P4, a switch programming language. TensorExpress improves latency and network blocking time up to 2.5 and 2.44 times, respectively.
Minkoo Kang 0001, Gyeongsik Yang, Yeonho Yoo, Chuck Yoo
CLOUD4
2020 Adaptive Control Channel Traffic Shaping for Virtualized SDN in Clouds
abstract
As the number of tenants in clouds grows, virtualized SDN faces a challenge of control channel fairness as the control channels interfere with each other. This paper proposes an adaptive traffic shaping scheme for control channels called “Sincon.” Through experiments, Sincon achieves traffic shaping for control channels and improves the variances of control channel throughputs and forwarding setup times up to 3.8 and 2.86 times, respectively.
Yeonho Yoo, Gyeongsik Yang, Minkoo Kang 0001, Chuck Yoo
CLOUD4
2020 Network Monitoring for SDN Virtual Networks
abstract
This paper proposes V-Sight, a network monitoring framework for software-defined networking (SDN)-based virtual networks. Network virtualization with SDN (SDN-NV) makes it possible to realize programmable virtual networks; so, the technology can be beneficial to cloud services for tenants. However, to the best of our knowledge, although network monitoring is a vital prerequisite for managing and optimizing virtual networks, it has not been investigated in the context of SDN-NV. Thus, virtual networks suffer from non-isolated statistics between virtual networks, high monitoring delays, and excessive control channel consumption for gathering statistics, which critically hinders the benefits of SDN-NV. To solve these problems, V-Sight presents three key mechanisms: 1) statistics virtualization for isolated statistics, 2) transmission disaggregation for reduced transmission delay, and 3) pCollector aggregation for efficient control channel consumption. V-Sight is implemented on top of OpenVirteX, and the evaluation results demonstrate that V-Sight successfully reduces monitoring delay and control channel consumption up to 454 times.
Gyeongsik Yang, Heesang Jin, Minkoo Kang 0001, Gi Jun Moon, Chuck Yoo
INFOCOM5
2019 FAVE: Bandwidth-Aware Failover in Virtualized SDN for Clouds
abstract
Network virtualization based on SDN has gained attention in cloud networking. However, existing studies have not provided any failover technique in the event of physical link failure. We propose FAVE, which provides seamless failover and bandwidth-aware protection. FAVE carefully allocates backup routes to handle both failure and interference between tenants. Evaluation shows that FAVE is effective. To our knowledge, FAVE is the first attempt to address failover in virtualized SDN environments.
Heesang Jin, Gyeongsik Yang, Bong-yeol Yu, Chuck Yoo
CLOUD4
2019 Multimedia file forensics system exploiting file similarity search
Min-Ja Kim, Chuck Yoo, Young Woong Ko
Multim. Tools Appl.2
2018 FlowVirt: Flow Rule Virtualization for Dynamic Scalability of Programmable Network Virtualization
abstract
We propose a new concept called "flow rule virtualization" (FlowVirt) for programmable network virtualization (P-NV). In P-NV, network hypervisor is a key component in that it plays a role in creating and managing virtual networks. This paper first reports a critical limitation of network hypervisor - scalability problem, which results in the high consumption of the switch memory, control channel, and CPU cycles: 3.9, 4.7, and 1.7 times higher than host-based network virtualization, respectively. This scalability problem arises because all the flow rules from the virtual network controllers are directly installed into switches. To resolve the scalability problem, FlowVirt introduces a flow rule abstraction: virtual and physical flow rules. By separating virtual and physical flow rules, the abstraction virtualizes flow rules so that FlowVirt can merge virtual flow rules to a smaller number of physical flow rules to be installed in switches. The evaluation results show the enhanced scalability of FlowVirt. The number of flow rules to be installed in switches decreases by up to 10 times compared to the previous P-NV. The control channel bandwidth and CPU cycles are also reduced by up to 14 and 3 times, respectively.
Gyeongsik Yang, Bong-yeol Yu, Wontae Jeong, Chuck Yoo
IEEE CLOUD4
2018 Kafe: Can OS Kernels Forward Packets Fast Enough for Software Routers?
abstract
It is widely believed that software routers based on commodity operating systems cannot deliver high-speed packet processing, and a number of alternative approaches (including user-space network stacks) have been proposed. This paper revisits the inefficiency of kernel-level packet processing inside modern OS-based software routers and explores whether a redesign of kernel network stacks can improve the incompetence. We present a case contrary to the belief through a redesign: Kafe-a kernel-based advanced forwarding engine that can process packets as fast as user-space network stacks. The Kafe neither adds any new API nor depends on proprietary hardware features, but the Kafe outperforms Linux by seven times and RouteBricks by three times. The current implementation of the Kafe can forward 64-byte IPv4 packets at 28.2 Gbps using eight cores running at 2.6 GHz. Our evaluation results show that the Kafe achieves similar packet forwarding performance to Intel DPDK while consuming much less CPU and memory resources.
Cheol-Ho Hong, Kyungwoon Lee, Jae-Hyun Hwang, Hyunchan Park, Chuck Yoo
IEEE/ACM Trans. Netw.5
2017 KVS: high-efficiency kernel-level virtual switch
abstract
In clouds, virtual switch (vSwitch) is in charge of packet forwarding between virtual machines (VMs). However, kernel-based vSwitches show throughput degradation for intensive packet processing; this becomes a bottleneck for the network performance of clouds. DPDK-based vSwitch (DPDK vSwitch) [1] has been developed to resolve the performance problem. Although it exhibits high throughput, DPDK vSwitch has two weak points. First, it consumes excessive memory. DPDK vSwitch uses huge page to reduce the number of memory operations, and this design causes high memory consumption even when the traffic is low. According to [2], memory determines the available number of VMs per single physical server. Thus, saving the memory decreases the capital expenditure of clouds. Second, security is another concern of the DPDK vSwitch, because its data plane is exposed to user space with the shared memory [3]. Therefore, the isolation of packets across VMs cannot be guaranteed. To overcome the excessive memory use and security concern, we propose a new kernel-level vSwitch (KVS) based on Linux. KVS do not use huge page nor bypass kernel stack. Instead, KVS applies the following key ideas to enhance the throughput.
Heungsik Choi, Gyeongsik Yang, Kyungwoon Lee, Chuck Yoo
SoCC4
2017 AKC: advanced KSM for cloud computing
abstract
Kernel samepage merging (KSM) in Linux kernel archive is a memory deduplication scheme that finds duplicate pages and shares the page in order to alleviate memory bottleneck in cloud. However, because the KSM has to scan all pages in memory to find duplicate pages, KSM consumes high CPU cycles and so causes virtual machines (VMs) performance degradation [1]. This degradation of VMs performance is an obstacle in cloud to service real-time applications (i.e. Netflix) [3]. A previous work, CMD [1] proposed page grouping scheme to reduce page comparisons, but it requires special monitoring hardware, XLH [2] enhanced page sharing with the information of guest VM I/O operation. However, the CPU overhead of XLH is still very high - similar to the default KSM. to make KSM more useful, we need an optimization scheme that consume less CPU cycles. Therefore, we first profile the CPU cycle consumption of KSM and the results show that page comparison (28.77%) and page checksum (26.14%) take most of cycles. Based on the results, we propose advanced KSM for cloud computing (AKC) that consumes less CPU cycles than the default KSM. to reduce the number of page comparisons, we apply checksum based RB-tree structure. In addition, AKC decreases page checksum overhead with hardware-accelerated crc32 hash function.
Sioh Lee, Bongkyu Kim, Young-Pil Kim, Chuck Yoo
SoCC4
2017 Efficient big link allocation scheme in virtualized software-defined networking
abstract
We propose an efficient resource allocation scheme for big links in virtualized software-defined networking. Network virtualization based on software-defined networking provides big link concept to facilitate simple network management - big link maps a set of switches and links into a single virtual link. However, this paper reports an issue of the big link in that there is a severe performance degradation in virtualized SDN environments. We find the cause: the existing network hypervisors do not consider the network traffic when allocating physical resources to a big link. To address this issue, we present big link allocation scheme (BAS) that considers network traffic when allocating and reallocating resources to a big link. A prototype implementation is done with OpenVirteX, and experiments demonstrate that the big link with BAS achieves four times greater throughput than that of the big link without BAS. Moreover, by including a timer in OpenVirteX, the BAS decreases unnecessary resource reallocations, which reduces overhead.
Wontae Jeong, Gyeongsik Yang, Seong-Mun Kim, Chuck Yoo
CNSM4
2017 BFD-based link latency measurement in software defined networking
abstract
5G networks offer various network services based on software defined networking and network function virtualization. However, certain services are sensitive to link latency which is why it is consistently observed to provide high quality services. Previous studies have proposed two approaches to this task: measuring the latency by probe packets and link-layer discovery protocol (LLDP) packets. However, they have several limitations like flow rule preconfiguration, influence of the control plane traffic, and necessity of calibration. In this paper, Bidirectional forwarding detection (BFD) based approach is proposed. The approach measures latency at the data plane with simply implemented echo mode in Open vSwitch. We evaluates and compare the proposed approach to LLDP-based one in terms of single link latency and path latency, and error rate. In addition, we verify that the control plane throughput affects link latency according to the increased number of switches. As a result, the proposed approach can resolve the limitations and provides accuracy link latency.
Seong-Mun Kim, Gyeongsik Yang, Chuck Yoo, Sung-Gi Min
CNSM3
2016 Eliminating bandwidth estimation from adaptive video streaming in wireless networks
Jae-Hyun Hwang, Chuck Yoo
Signal Process. Image Commun.3
2016 VADI: GPU Virtualization for an Automotive Platform
abstract
Modern vehicles are evolving with more electronic components than ever before (In this paper, “vehicle” means “automotive vehicle.” It is also equal to “car.”) One notable example is graphical processing unit (GPU), which is a key component to implement a digital cluster. To implement the digital cluster that displays all the meters (such as speed and fuel gauge) together with infotainment services (such as navigator and browser), the GPU needs to be virtualized; however, GPU virtualization for the digital cluster has not been addressed yet. This paper presents a Virtualized Automotive DIsplay (VADI) system to virtualize a GPU and its attached display device. VADI manages two execution domains: one for the automotive control software and the other for the in-vehicle infotainment (IVI) software. Through GPU virtualization, VADI provides GPU rendering to both execution domains, and it simultaneously displays their images on a digital cluster. In addition, VADI isolates GPU from the IVI software in order to protect it from potential failures of the IVI software. We implement VADI with Vivante GC2000 GPU and perform experiments to ensure requirements of International Standard Organization (ISO) safety standards. The results show that VADI guarantees 30 frames per second (fps), which is the minimum frame rate for digital cluster mandated by ISO safety standards even with the failure of the IVI software. It also achieves 60 fps in a synthetic workload.
Chi-Young Lee, Se-Won Kim, Chuck Yoo
IEEE Trans. Ind. Informatics3
2016 Synchronization support for parallel applications in virtualized clouds
Cheol-Ho Hong, Young-Pil Kim, Hyunchan Park, Chuck Yoo
J. Supercomput.4
2016 Storage SLA Guarantee with Novel SSD I/O Scheduler in Virtualized Data Centers
abstract
Service level agreements (SLAs) for storage performance in virtualized systems are difficult to guarantee, because different consolidated virtual machines have their own performance requirements. Moreover, hard disk drives (HDDs) in virtualized systems are being replaced by solid-state drives (SSDs). SSDs have higher throughput and lower latency than HDDs; however, they pose new challenges in terms of SLAs. In this paper, we determine that existing I/O schedulers working with SSDs fail to guarantee SLAs among virtualmachines, and do not effectively utilize the high performance of SSDs. To address this issue, we propose the opportunistic I/O scheduler (OIOS), a novel I/O scheduler for SSDs. OIOS guarantees SLAs and fully utilizes the high performance of SSDs. To support realistic SLAs, OIOS provides diverse SLA support functions, including reservations, limitations, and proportional sharing. In addition, OIOS accepts SLAs that are specified in four measurement types: bandwidth, I/Os per second (IOPS), latency, and utilization. Experimental results show that OIOS increases the aggregated bandwidth of VMs by 80 percent compared to mClock, while achieving a similar level of fairness. In addition, we evaluate the proposed scheduler with realistic benchmarks, such as Filebench and the Yahoo CloudServing Benchmark. OIOS successfully guarantees the requirements of diverse SLAs with different metrics.
Hyunchan Park, See-hwan Yoo, Cheol-Ho Hong, Chuck Yoo
IEEE Trans. Parallel Distributed Syst.4
2015 SSD-Tailor: Automated Customization System for Solid-State Drives
abstract
Enterprise servers require customized solid-state drives (SSDs) to satisfy their specialized I/O performance and reliability requirements. For effective use of SSDs for enterprise purposes, SSDs must be designed considering requirements such as those related to performance, lifetime, and cost constraints. However, SSDs have numerous hardware and software design options, such as flash memory types and block allocation methods, which have not been well analyzed yet, but on which the SSD performance depends. Furthermore, there is no methodology for determining the optimal design for a particular I/O workload. This paper proposes SSD-Tailor, a customization tool for SSDs. SSD-Tailor determines a near-optimal set of design options for a given workload. SSD designers can use SSD-Tailor to customize SSDs in the early design stage to meet the customer requirements. We evaluate SSD-Tailor with nine I/O workload traces collected from real-world enterprise servers. We observe that SSD-Tailor finds near-optimal SSD designs for these workloads by exploring only about 1% of the entire set of design candidates. We also show that the near-optimal designs increase the average I/O operations per second by up to 17% and decrease the average response time by up to 163% as compared to an SSD with a general design.
Hyunchan Park, Hanchan Jo, Cheol-Ho Hong, Young-Pil Kim, See-hwan Yoo, Chuck Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2014 Performance Prediction and Evaluation of Parallel Applications in KVM, Xen, and VMware
Cheol-Ho Hong, Young-Pil Kim, Hyunchan Park, Chuck Yoo
Euro-Par5
2014 O1FS: Flash file system with O(1) crash recovery time
Hyunchan Park, Sam H. Noh, Chuck Yoo
J. Syst. Softw.3
2014 Real-Time Scheduling for Xen-ARM Virtual Machines
abstract
This paper investigates the feasibility of real-time scheduling with mobile hypervisor, Xen-ARM. Particularly for mobile virtual machines, real-time support is in high demand. However, it is difficult to guarantee real-time scheduling with virtual machines because inter-VM and intra-VM schedulability have to be determined in multi-OS environments. To address the schedulability, first, this paper presents a definition of a real-time virtual machine. Second, this paper analyzes intra-VM schedulability, taking quantization overhead into account. Quantization overhead comes from tick-based scheduling of Xen-ARM, which requires integer presentation of scheduling period and execution slice. Third, to minimize quantization overhead, this paper provides a new algorithm, called SH-quantization that provides accurate and efficient parameterization of a real-time virtual machine. Fourth, this paper presents an inter-VM schedulability test for incorporating multiple real-time virtual machines. To evaluate the approach, we implement the SH-quantization algorithm in Xen-ARM and paravirtualize a real-time OS, called xeno-μC/OS-II. We ran extensive experiments with various configurations of real-time tasks on a real hardware platform in order to characterize the scheduling behavior of real-time virtual machine with quantization. The results show that quantization overhead consumes additional CPU bandwidth up to 90% and the proposed algorithm guarantees intra/inter-VM schedulability with minimal CPU bandwidth.
See-hwan Yoo, Chuck Yoo
IEEE Trans. Mob. Comput.2
2013 Cache replacement strategies for scalable video streaming in CCN
abstract
To deal with the structural limitations of the current Internet, a new concept of a future Internet has been suggested, among which a content centric network(CCN) with in-router storage and the use of content as an address has taken center stage. We describe the design and implementation of new cache-replacement algorithms for layered video content(H.264/SVC) streaming in a CCN. The cache-replacement algorithms(LGD-size, and Lmix)proposed in this paper impose value on each object based on its size, recent reference trend, and the reference frequency stored in the cache. In addition, the proposed algorithms operate by reflecting the layered features of SVC video streaming. The proposed algorithms are applied to a CCN node, and when a client repeatedly requests HTTP adaptive video-streaming encoded with H.264/SVC, the existing algorithms(FIFO, LFU, LRU) and cache hit rates of each CCN node are measured. A comparison of the cache hit rate at each CCN node and the quality of the video streaming, show that the efficiency of layered video streaming in a CCN can be improved by modifying the cache-replacement algorithm used.
Kyubo Lim, Chuck Yoo
APCC3
2013 Virtualizing ARM VFP (Vector Floating-Point) with Xen-ARM
See-hwan Yoo, Sung-bae Yoo, Chuck Yoo
J. Syst. Archit.3
2012 Data Deduplication Using Dynamic Chunking Algorithm
Young Chan Moon, Ho Min Jung, Chuck Yoo, Young Woong Ko
ICCCI (2)3
2011 PARFAIT: A new scheduler framework supporting heterogeneous Xen-ARM schedulers
abstract
In recent consumer electronics devices, virtualization is widely adopted for diverse reasons such as enhanced reliability, stronger security and better user customizability. Recent CE virtual machines require not only diverse functionalities of GPOS, but also real-time performance of RTOS. However, the current hypervisor-based virtual machine monitors cannot sufficiently support both RTOS and GPOS at the same time. The reason is that the current hypervisor scheduler is biased to only one kind of guest OS and it cannot support heterogeneous schedulers. Therefore, this paper proposes a new scheduler architecture, PARFAIT for both RT and GP guest OSs. Our new scheduler enables to guarantee CPU bandwidth for RTOS and provides fairness among GPOSs using hierarchical structure.
Jae-Woo Jeong, See-hwan Yoo, Chuck Yoo
CCNC3
2009 A Step to Support Real-Time in Virtual Machine
abstract
Real-time is one of the unique requirements in embedded systems. In this paper, we perform a feasibility study on how to support real-time in an embedded virtual machine system. Firstly, we argue that the I/O model of the current virtual machine monitor like Xen is not suitable to support real-time applications because it lacks in predictability and it does not guarantee a deterministic I/O processing. We provide an alternative I/O model for virtualized embedded systems. Devices are categorized into four groups: dedicated, active, running, dynamic. Dedicated devices make a virtual machine simple because they do not need to be virtualized for isolation. However, dedication does not mean the performance isolation. Our experimental results with dedicated device show that traditional dedication cannot guarantee the timely responsiveness in heavy interrupt cases. Specifically, responsiveness of real-time OS degrades as interrupt load increases. Therefore, a proper interrupt control mechanism is required at virtual machine monitor level in order to support timely responsiveness. In addition, our result supports that (1) short and prioritized interrupt processing helps responsiveness in a virtual machine system; (2) smaller time quantum results in better responsiveness also.
See-hwan Yoo, Miri Park, Chuck Yoo
CCNC3
2009 TCP Feno: Enhancement for higher accuracy of loss differentiation over small buffer heterogeneous networks
abstract
It is well known that TCP shows performance degradation over wired/wireless networks since TCP regards packet loss as network congestion. TCP Veno has successfully addressed this fundamental problem by proposing an end-to-end loss differentiation algorithm, which distinguishes the cause of packet loss by the number of packets in the router buffer. Unfortunately, Veno's algorithm shows very low accuracy on small buffer routers that are emerging recently as a new challenge for Internet routers. In the small buffer networks, routers can overflow easily in a short time although Veno diagnoses wireless loss, and this leads to the failure in loss differentiation. Furthermore, when congestion loss is misdiagnosed as wireless loss, Veno shows poor Reno-friendliness since it can increase sending rate by setting ssthresh to a larger value than Reno. In this paper, we propose a more accurate loss differentiation algorithm for small buffer heterogeneous networks. Our algorithm accurately distinguishes wireless and wired packet loss by newly defining congestive rate. Through extensive network simulations, we confirm that our new TCP Feno achieves not only higher accuracy, but also better Reno-friendliness while not losing performance efficiency.
Jae-Hyun Hwang, See-hwan Yoo, Chuck Yoo
LCN3
2008 Minimum DVS gateway deployment in DVS-based overlay streaming
Sang-Seon Byun, Chuck Yoo
Comput. Commun.2
2008 DR-TCP: Downloadable and reconfigurable TCP
Jae-Hyun Hwang, Jin-Hee Choi, Se-Won Kim, Chuck Yoo
J. Syst. Softw.4
2008 Towards building large scale live media streaming framework for a U-city
Eun-Seok Ryu, Chuck Yoo
Multim. Tools Appl.2
2007 Efficient MD Coding Core Selection to Reduce the Bandwidth Consumption
abstract
Multiple distribution trees and multiple description (MD) coding are highly robust since they provide redundancy both in network paths and data. However, MD coded streaming includes a redundant information, which results in additional bandwidth consumptions in entire distribution trees. In this paper, we deploy core nodes in distribution tree, and give a role of MD coding to each core node, instead of a source node then we show how amount of bandwidth consumption can be reduced. Since the problem of finding an optimal set of core nodes is proved to be NP-hard, an intuitional heuristic-based algorithm is proposed. The simulation results show that our heuristic algorithm reduces the bandwidth consumptions by about 25% in the hierarchical topology compared to the MD coding in source node only.
Sunoh Choi, Sang-Seon Byun, Chuck Yoo
LCN3
2007 Self-prevention of socket buffer overflow
Jin-Hee Choi, Young-Pil Kim, Chuck Yoo
Comput. Networks3
2007 Proxy location for minimizing delivery delay in HRM networks
Sang-Seon Byun, Chuck Yoo
Comput. Commun.2
2007 Impact of protocol overheads on network throughput over high-speed interconnects: measurement, analysis, and improvement
Hyun-Wook Jin, Chuck Yoo
J. Supercomput.2
2006 Reducing Delivery Delay in HRM Tree
Sang-Seon Byun, Chuck Yoo
ICCSA (2)2
2005 Analytic end-to-end estimation for the one-way delay and its variation
abstract
Delay estimation is a difficult problem in computer networks. RTT (round trip time) is often used as an approximation of the delay, but because it is a sum of the forward and reverse delays, the actual one-way delay cannot he estimated accurately from RTT. Accurate one-way delay estimation becomes crucial because it serves a very important role in network and application design. This paper proposes a new scheme to estimate one-way delay and its variation. The scheme calibrates estimated one-way delay so in a brief duration as to be used in many protocols that adjust their behavior depending on the network condition. We analytically derive one-way delay, forward and reverse delay respectively, and show that our one-way delay estimation is much more accurate than RTT estimation by simulation.
Jin-Hee Choi, Chuck Yoo
CCNC2
2005 One-way delay estimation and its application
Jin-Hee Choi, Chuck Yoo
Comput. Commun.2
2005 Exploiting NIC architectural support for enhancing IP-based protocols on high-performance networks
Hyun-Wook Jin, Pavan Balaji, Chuck Yoo, Dhabaleswar K. Panda 0001
J. Parallel Distributed Comput.3
2005 A Video Streaming System for Mobile Phones: Practice and Experience
Hojung Cha, Jongmin Lee 0001, Jongho Nang, Sungyong Park, Jin-Hwan Jeong, Chuck Yoo
Wirel. Networks6
2004 An approach to interactive media system for mobile devices
abstract
The interactive system which interacts human with computer has been recognized as one direction of computer development for a long time. For example, in cinema, a person gets information he wants or plays the media data while moving by using a mobile device. As the development of this system, we designed and implemented the system interacts with users in a small terminal. Our study has three categories. The first category is the development of new interactive media markup language (IML) for the writing interactive media data. The second category is the IML translator which translates IML into the best form to be played on mobile device. And the third category is the IM player, which plays the transferred media data and interacts with user. IML was designed for controlling vector graphics and general media objects in detail and supporting synchronization. Also, it was designed to be operated in small mobile device as well as desktop PC or set-top box which has high CPU performance. The player, implemented finally, is operated on PDA (HP iPAQ) and plays the multimedia data consist of vector graphics (OpenGL), H.264 and AAC etc. according to the choice of user. This system can be used in the ways of interactive cinema and interactive game, and can substitute new interactive web services for existing web services.
Eun-Seok Ryu, Chuck Yoo
ACM Multimedia2
2003 TCP-aware Source Routing in Mobile Ad Hoc Networks
abstract
Temporary link failures and route changes occur frequently in mobile ad hoc networks. Since TCP assumes that all packet looses are due to network congestion, TCP does not show satisfactory performance in ad hoc networks. In this paper, we propose a simple and new mechanism called TSR, TCP-aware source routing, which can improve TCP performance in mobile ad hoc networks. By reducing the number of invalid routes, TSR minimizes TCP's consecutive timeouts. In our simulation study, TSR achieves up to 50% TCP performance improvement without requiring any modification of TCP stack in end systems.
Jin-Hee Choi, Chuck Yoo
ISCC2
2003 Firmware-Level Latency Analysis on a Gigabit Network
Hyun-Wook Jin, Chuck Yoo
J. Supercomput.2
2002 Stepwise Optimizations of UDP/IP on a Gigabit Network (Research Note)
Hyun-Wook Jin, Chuck Yoo, Sung-Kyun Park
Euro-Par2
2001 Distributed Test using Logical Clock
Hee Yong Youn, Soonuk Seol, Chuck Yoo
FORTE4
2001 Comments on 'The Model Checker SPIN'
abstract
The paper by G.J. Holzmann (see ibid., vol.23, no.5, p.279-95, 1997) describes how to apply SPIN to the verification of a synchronization algorithm (L.M. Ruane, 1990) in process scheduling of an operating system. We report an error in the verification model presented by G.J. Holzmann and present a revised model with verification result. Our result explains the reason why SPIN found the race condition in the synchronization algorithm. We also show that the suggested fix by G.J. Holzmann is incorrect.
Ki-Seok Bang, Chuck Yoo
IEEE Trans. Software Eng.3
1999 Latency analysis of UDP and BPI on Myrinet
abstract
High-speed networks such as ATM, Myrinet, and Gigabit Ethernet are available today, and many researchers make efforts to enhance the performance of end-to-end communication on these high-speed networks. One of the efforts is to develop new light-weight communication primitives for high-speed network. However the latency of the new primitives has not been characterized thoroughly, partly because existing measurement methodologies do not take into account the features of high-speed networks. Therefore, there are only incomplete comparisons of the new primitives and traditional protocols, and they cannot really prove the usefulness of new primitives. In order to address this issue, this paper suggests a new measurement methodology and uses the methodology to perform a detailed latency analysis of UDP and a light-weight primitive, called BPI, on Myrinet. Our results clearly show the difference of per-byte overhead between BPI and UDP. A surprising result is that BPI is found to be slower than UDP for 4KB or larger data size.
Hyun-Wook Jin, Chuck Yoo
IPCCC2