EDBT 2026 Demo / reviewers in the wild / expert
Gyeongsik Yang
dblp:181/1435
· DBLP profile ↗
21ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-4560-2972ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prediction-based GPU sharing for distributed trainingabstract• Formulate the inconsistent JCT problem using gSLA for the first time. • Design a new JCT increase prediction model and job scheduler for GPU sharing. • Achieve up to 47.3× better gSLA satisfaction and 50× lower gSLA excess ratio. • Improve JCT and GPU efficiency by ∼ 60% and ∼ 44% over existing methods. • Demonstrate TensorShare’s effectiveness in improving gSLA and JCT for unseen jobs. GPU sharing aims to enhance the efficiency of GPU utilization by running distributed deep learning training jobs concurrently. However, GPU sharing poses a significant challenge: the increase in job completion time (JCT) caused by interference between jobs is inconsistent, complicating job scheduling. Our experiments reveal that the degree of JCT increase varies by as much as ∼ 3.7 × . While previous studies have analyzed this JCT inconsistency problem, none of them have been able to minimize the inconsistency. We propose TensorShare, a proactive GPU sharing technique that leverages a deep learning model to predict the extent of JCT increase. This study defines a new metric, called GPU SLA, which represents the upper threshold of JCT increase. TensorShare then introduces a novel scheduler that proactively identifies which jobs meet GPU SLA while minimizing the JCT increase. Our evaluation shows that TensorShare improves GPU SLA satisfaction rates by 26.1 × –47.3 × and reduces the JCT increase by 37%–60%. Furthermore, we evaluate TensorShare with large language models that are not included in training TensorShare’s prediction model, achieving ∼ 7 × and ∼ 10.3 × improvements in GPU SLA satisfaction and JCT inconsistency, respectively. Changyong Shin, Younghun Go, Yeonho Yoo, Jae-Hyun Hwang, Gyeongsik Yang, Chuck Yoo |
Future Gener. Comput. Syst. | 6 |
| 2025 | Parameter-Efficient 12-Lead ECG Reconstruction from a Single Lead
Yeonho Yoo, Jinkyu Kim 0001, Dosun Lim, Gyeongsik Yang, Chuck Yoo |
MICCAI (2) | 5 |
| 2024 | Harmonia: Accurate Federated Learning with All-Inclusive DatasetabstractFederated learning (FL) is an appealing model training technique that utilizes heterogeneous datasets and user devices, ensuring user data privacy. Existing FL research proposed device selection schemes to balance the computing speeds of devices. However, we observe that these schemes compromise prediction accuracy by ~57. 7 %. To solve this problem, we present Harmonia that enhances prediction accuracy, while also balancing the diverse computing speeds of devices. Our evaluation shows that Harmonia improves prediction accuracy by ~ 1.7 x over existing schemes. Wonmi Choi, Juyoung Ahn, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang |
CLOUD | 5 |
| 2024 | Predictive Placement of Geo-Distributed Blockchain Nodes for Performance GuaranteeabstractBlockchain-as-a-service (BaaS) in cloud datacenters is gaining widespread attention due to its high performance and privacy. However, existing BaaS solutions lack a method for deciding the proper placement of blockchain nodes across virtual machines in worldwide datacenters to achieve desired performance. Our motivating experiments show that transaction processing performance (TPS) varies ~31.6% depending on the placements. To provide an automatic placement solution for BaaS, we propose Cyan that predicts the TPS for blockchain node placements. Our evaluations on Google Cloud Platform demonstrate that Cyan improves the TPS guarantee ~2.39x compared to existing techniques. Yeonho Yoo, Chuck Yoo, Gyeongsik Yang |
CLOUD | 4 |
| 2024 | Intelligent Packet Processing for Performant Containers in IoTabstractThis article explores the computing and communication overhead of network processing in Internet of Things (IoT) devices, focusing on containers, a major building block for the edge computing. Our experiments reveal that containers on IoT devices suffer$\sim 2.6\times $higher CPU usage for SoftIRQ processing, ~59% less network throughput, and$2\times $higher per-packet latency on average than native processes. While several existing studies enhance networking performance, they often sacrifice interoperability by requiring special hardware or modifying networking semantics or APIs. Thus, we design and implement a kernel networking accelerator, called SCON, that maintains interoperability, crucial for IoT devices. SCON addresses major bottlenecks in container networking through system-level profiling. We evaluate SCON with three types of IoT devices. On the Raspberry Pi 4, SCON reduces the latencies of major IoT application protocols (e.g., HTTP and MQTT) by$\sim 10\times $, achieving a similar level of latency to the native process. Further analysis shows that SCON reduces CPU usage for SoftIRQ processing by ~26%. We also report similar improvements on the other two IoT devices. Our conclusion is that SCON is unique in significantly reducing the computing and communication overhead of container networking in IoT devices while maintaining interoperability. Furthermore, it works consistently across different types of devices, whether wired or wireless, and regardless of heavy or sporadic traffic. Wonmi Choi, Yeonho Yoo, Kyungwoon Lee, Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Gyeongsik Yang, Chuck Yoo |
IEEE Internet Things J. | 7 |
| 2023 | Selective Preemption of Distributed Deep Learning TrainingabstractAs more distributed deep learning (DDL) jobs run in public clouds, their effective scheduling becomes a major challenge. Current studies prioritize the execution of jobs with less remaining time, which is known to be the best in reducing average job completion time (JCT). However, we observe that this approach does not work when the preemption for pausing and loading jobs weighs in; sometimes, the preemption overheads of DDL jobs take up to hundreds of seconds. This results in very ineffective scheduling, so in some cases, the first-in-first-out policy performs much better. This paper proposes a new scheduling framework called Xion that takes into account the preemption overheads and only preempts DDL jobs when it is beneficial. Our evaluation results demonstrate that Xion effectively reduces the average JCT by 19% and improves the waiting time by 1.64×. Younghun Go, Changyong Shin, Jeunghwan Lee, Yeonho Yoo, Gyeongsik Yang, Chuck Yoo |
CLOUD | 5 |
| 2023 | Control Channel Isolation in SDN Virtualization: A Machine Learning ApproachabstractPerformance isolation is an essential property that network virtualization must provide for clouds. This study addresses the performance isolation of the control plane in virtualized software-defined networking (SDN), which we call control channel isolation. First, we report that the control channel isolation is seriously broken in the existing network hypervisor in that the end-to-end control latency grows by up to 15 x as the number of virtual switches increases. This jeopardizes the key network operations, such as routing, in datacenters. To address this issue, we take a machine learning approach that learns from the past control traffic as time-series data. We propose a new network hypervisor, Meteor, that designs an LSTM autoencoder to predict the control traffic per virtual switch. Our evaluation results show that Meteor improves the processing latency per control message by up to 12.7x. Furthermore, Meteor reduces the end-to-end control latency by up to 73.7%, which makes it comparable to the non-virtualized SDN. Yeonho Yoo, Gyeongsik Yang, Changyong Shin, Jeunghwan Lee, Chuck Yoo |
CCGrid | 2 |
| 2023 | Accurate and Efficient Monitoring for Virtualized SDN in CloudsabstractThis article presents V-Sight, a network monitoring framework for programmable virtual networks in clouds. Network virtualization based on software-defined networking (SDN-NV) in clouds makes it possible to realize programmable virtual networks; consequently, this technology offers many benefits to cloud services for tenants. However, to the best of our knowledge, network monitoring, which is a prerequisite for managing and optimizing virtual networks, has not been investigated in the context of SDN-NV systems. As the first framework for network monitoring in SDN-NV, we identify three challenges: non-isolated and inaccurate statistics, high monitoring delay, and excessive control channel consumption for gathering statistics. To address these challenges, V-Sight introduces three key mechanisms: 1) statistics virtualization for isolated statistics, 2) transmission disaggregation for reduced transmission delay, and 3) pCollector aggregation for efficient control channel consumption. The evaluation results reveal that V-Sight successfully provides accurate and isolated statistics while reducing the monitoring delay and control channel consumption in orders of magnitude. We also show that V-Sight can achieve a data plane throughput close to that of non-virtualized SDN. Gyeongsik Yang, Yeonho Yoo, Minkoo Kang 0001, Heesang Jin, Chuck Yoo |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | TeaVisor: Network Hypervisor for Bandwidth Isolation in SDN-NVabstractWe introduce TeaVisor that provides bandwidth isolation guarantee for network virtualization (NV) based on software-defined networking (SDN). SDN-based NV (SDN-NV) offers many benefits to clouds, such as topology and address virtualization while allowing flexible resource provisioning, control, and monitoring on virtual networks. In SDN-NV, however, routing is done by tenants independently; thus, existing studies have difficulties in bandwidth isolation guarantee due to the overloaded link problem. Bandwidth isolation guarantee is essential for providing stable and reliable throughput on network services in SDN-NV. Without bandwidth isolation guarantee, tenants suffer degraded service qualities and significant loss in revenue. To address this problem, we design and implement TeaVisor in three components: path virtualization, bandwidth reservation, and path establishment. Through extensive experiments, TeaVisor shows that bandwidth isolation is guaranteed with near-zero errors, which is three orders of magnitude better than existing studies. In addition, TeaVisor guarantees the minimum and maximum bandwidth at the same time. We also present an overhead analysis of TeaVisor in control traffic and memory consumption. Yeonho Yoo, Gyeongsik Yang, Jeunghwan Lee, Changyong Shin, Hoseok Kim, Chuck Yoo |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Machine Learning-Based Prediction Models for Control Traffic in SDN SystemsabstractThis article presentsElixir, an automated prediction model formulation framework for control traffic using machine learning. Control traffic is vital in software-defined networking (SDN) systems because it determines the reliability and scalability of the entire system. Various studies have sought to design control traffic prediction models for the proper provisioning and planning of SDN systems. However, previously proposed models are based on descriptive modeling, well-suited for only specific SDN system instances. Furthermore, these models exhibit poor accuracy (errors of up to 85%) because of the heterogeneity of SDN systems. Because descriptive modeling requires a significant amount of human contemplation, it is impossible to formulate adequate prediction models for countless SDN system instances.Elixiraddresses this problem by applying machine learning.Elixirstarts the model formulation through self-generated datasets. Then,Elixirsearches prediction models to fit the accuracy for respective SDN systems. Also,Elixirpicks robust models that exhibit reasonable accuracy even in a network topology that differs from the topology used for model training. We evaluate theElixirframework on nine heterogeneous SDN systems. As a key outcome,Elixirsignificantly reduces prediction errors, achieving up to 10.6× improvement compared to the previous model for control traffic throughput of OpenDayLight controller. Yeonho Yoo, Gyeongsik Yang, Changyong Shin, Chuck Yoo |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Xonar: Profiling-based Job Orderer for Distributed Deep LearningabstractDeep learning models have a wide spectrum of GPU execution time and memory size. When running distributed training jobs, however, their GPU execution time and memory size have not been taken into account, which leads to the high variance of job completion time (JCT). Moreover, the jobs often run into the GPU out-of-memory (OoM) problem so that the unlucky job has to restart all over. To address the problems, we propose Xonar to profile the deep learning jobs and order them in the queue. The experiments show that Xonar with TensorFlow v1.6 reduces the tail JCT by 44% with the OoM problem eliminated. Changyong Shin, Gyeongsik Yang, Yeonho Yoo, Jeunghwan Lee, Chuck Yoo |
CLOUD | 2 |
| 2021 | Bandwidth Isolation Guarantee for SDN Virtual NetworksabstractWe introduce TeaVisor, which provides bandwidth isolation guarantee for software-defined networking (SDN)-based network virtualization (NV). SDN-NV provides topology and address virtualization while allowing flexible resource provisioning, control, and monitoring of virtual networks. However, to the best of our knowledge, the bandwidth isolation guarantee, which is essential for providing stable and reliable throughput on network services, is missing in SDN-NV. Without bandwidth isolation guarantee, tenants suffer degraded service quality and significant revenue loss. In fact, we find that the existing studies on bandwidth isolation guarantees are insufficient for SDN-NV. With SDN-NV, routing is performed by tenants, and existing studies have not addressed the overloaded link problem. To solve this problem, TeaVisor designs three components: path virtualization, bandwidth reservation, and path establishment, which utilize multipath routing. With these, TeaVisor achieves the bandwidth isolation guarantee while preserving the routing of the tenants. In addition, TeaVisor guarantees the minimum and maximum amounts of bandwidth simultaneously. We fully implement TeaVisor, and the comprehensive evaluation results show that near-zero error rates on achieving the bandwidth isolation guarantee. We also present an overhead analysis of control traffic and memory consumption. Gyeongsik Yang, Yeonho Yoo, Minkoo Kang 0001, Heesang Jin, Chuck Yoo |
INFOCOM | 1 |
| 2021 | A Case for SDN-based Network VirtualizationabstractNetwork virtualization (NV) becomes an essential technology in cloud computing that isolates network flows for tenants. However, because existing NV technologies like overlay do not enable tenants to directly program (i.e., provision, control, and monitor) network resources, software-defined networking (SDN)-based NV (SDN-NV) has been proposed. Despite its great benefits, SDN-NV has been believed to bring considerable overheads due to the network hypervisor (NH). However, to date, there is no definite performance evaluation that proves the overheads of SDN-NV. To this end, this paper comprehensively investigates the performance and overheads of SDN-NV. Our experiment results reveal that SDN-NV provides the data plane performance comparable to or even better (up to 10.5× better TCP throughput) than the existing NV technologies. Also, the results on NH show that its overheads remain mostly constant, even when the number of switches, virtual networks, or network flows increases. In short, our evaluation indicates that the overhead of SDN-NV should not deter its practical use in datacenters. Gyeongsik Yang, Changyong Shin, Yeonho Yoo, Chuck Yoo |
MASCOTS | 1 |
| 2020 | TensorExpress: In-Network Communication Scheduling for Distributed Deep LearningabstractTensorExpress provides in-network communication scheduling for distributed deep learning (DDL). In cloud-based DDL, parameter communication over a network is a key bottleneck. Previous studies proposed tensor packet reordering approaches to reduce network blocking time. However, network contention still exists in DDL. TensorExpress mitigates network contention and reduces overall training time. It schedules tensor packets in-network using P4, a switch programming language. TensorExpress improves latency and network blocking time up to 2.5 and 2.44 times, respectively. Minkoo Kang 0001, Gyeongsik Yang, Yeonho Yoo, Chuck Yoo |
CLOUD | 2 |
| 2020 | Adaptive Control Channel Traffic Shaping for Virtualized SDN in CloudsabstractAs the number of tenants in clouds grows, virtualized SDN faces a challenge of control channel fairness as the control channels interfere with each other. This paper proposes an adaptive traffic shaping scheme for control channels called “Sincon.” Through experiments, Sincon achieves traffic shaping for control channels and improves the variances of control channel throughputs and forwarding setup times up to 3.8 and 2.86 times, respectively. Yeonho Yoo, Gyeongsik Yang, Minkoo Kang 0001, Chuck Yoo |
CLOUD | 2 |
| 2020 | Network Monitoring for SDN Virtual NetworksabstractThis paper proposes V-Sight, a network monitoring framework for software-defined networking (SDN)-based virtual networks. Network virtualization with SDN (SDN-NV) makes it possible to realize programmable virtual networks; so, the technology can be beneficial to cloud services for tenants. However, to the best of our knowledge, although network monitoring is a vital prerequisite for managing and optimizing virtual networks, it has not been investigated in the context of SDN-NV. Thus, virtual networks suffer from non-isolated statistics between virtual networks, high monitoring delays, and excessive control channel consumption for gathering statistics, which critically hinders the benefits of SDN-NV. To solve these problems, V-Sight presents three key mechanisms: 1) statistics virtualization for isolated statistics, 2) transmission disaggregation for reduced transmission delay, and 3) pCollector aggregation for efficient control channel consumption. V-Sight is implemented on top of OpenVirteX, and the evaluation results demonstrate that V-Sight successfully reduces monitoring delay and control channel consumption up to 454 times. Gyeongsik Yang, Heesang Jin, Minkoo Kang 0001, Gi Jun Moon, Chuck Yoo |
INFOCOM | 1 |
| 2019 | FAVE: Bandwidth-Aware Failover in Virtualized SDN for CloudsabstractNetwork virtualization based on SDN has gained attention in cloud networking. However, existing studies have not provided any failover technique in the event of physical link failure. We propose FAVE, which provides seamless failover and bandwidth-aware protection. FAVE carefully allocates backup routes to handle both failure and interference between tenants. Evaluation shows that FAVE is effective. To our knowledge, FAVE is the first attempt to address failover in virtualized SDN environments. Heesang Jin, Gyeongsik Yang, Bong-yeol Yu, Chuck Yoo |
CLOUD | 2 |
| 2018 | FlowVirt: Flow Rule Virtualization for Dynamic Scalability of Programmable Network VirtualizationabstractWe propose a new concept called "flow rule virtualization" (FlowVirt) for programmable network virtualization (P-NV). In P-NV, network hypervisor is a key component in that it plays a role in creating and managing virtual networks. This paper first reports a critical limitation of network hypervisor - scalability problem, which results in the high consumption of the switch memory, control channel, and CPU cycles: 3.9, 4.7, and 1.7 times higher than host-based network virtualization, respectively. This scalability problem arises because all the flow rules from the virtual network controllers are directly installed into switches. To resolve the scalability problem, FlowVirt introduces a flow rule abstraction: virtual and physical flow rules. By separating virtual and physical flow rules, the abstraction virtualizes flow rules so that FlowVirt can merge virtual flow rules to a smaller number of physical flow rules to be installed in switches. The evaluation results show the enhanced scalability of FlowVirt. The number of flow rules to be installed in switches decreases by up to 10 times compared to the previous P-NV. The control channel bandwidth and CPU cycles are also reduced by up to 14 and 3 times, respectively. Gyeongsik Yang, Bong-yeol Yu, Wontae Jeong, Chuck Yoo |
IEEE CLOUD | 1 |
| 2017 | KVS: high-efficiency kernel-level virtual switchabstractIn clouds, virtual switch (vSwitch) is in charge of packet forwarding between virtual machines (VMs). However, kernel-based vSwitches show throughput degradation for intensive packet processing; this becomes a bottleneck for the network performance of clouds. DPDK-based vSwitch (DPDK vSwitch) [1] has been developed to resolve the performance problem. Although it exhibits high throughput, DPDK vSwitch has two weak points. First, it consumes excessive memory. DPDK vSwitch uses huge page to reduce the number of memory operations, and this design causes high memory consumption even when the traffic is low. According to [2], memory determines the available number of VMs per single physical server. Thus, saving the memory decreases the capital expenditure of clouds. Second, security is another concern of the DPDK vSwitch, because its data plane is exposed to user space with the shared memory [3]. Therefore, the isolation of packets across VMs cannot be guaranteed. To overcome the excessive memory use and security concern, we propose a new kernel-level vSwitch (KVS) based on Linux. KVS do not use huge page nor bypass kernel stack. Instead, KVS applies the following key ideas to enhance the throughput. Heungsik Choi, Gyeongsik Yang, Kyungwoon Lee, Chuck Yoo |
SoCC | 2 |
| 2017 | Efficient big link allocation scheme in virtualized software-defined networkingabstractWe propose an efficient resource allocation scheme for big links in virtualized software-defined networking. Network virtualization based on software-defined networking provides big link concept to facilitate simple network management - big link maps a set of switches and links into a single virtual link. However, this paper reports an issue of the big link in that there is a severe performance degradation in virtualized SDN environments. We find the cause: the existing network hypervisors do not consider the network traffic when allocating physical resources to a big link. To address this issue, we present big link allocation scheme (BAS) that considers network traffic when allocating and reallocating resources to a big link. A prototype implementation is done with OpenVirteX, and experiments demonstrate that the big link with BAS achieves four times greater throughput than that of the big link without BAS. Moreover, by including a timer in OpenVirteX, the BAS decreases unnecessary resource reallocations, which reduces overhead. Wontae Jeong, Gyeongsik Yang, Seong-Mun Kim, Chuck Yoo |
CNSM | 2 |
| 2017 | BFD-based link latency measurement in software defined networkingabstract5G networks offer various network services based on software defined networking and network function virtualization. However, certain services are sensitive to link latency which is why it is consistently observed to provide high quality services. Previous studies have proposed two approaches to this task: measuring the latency by probe packets and link-layer discovery protocol (LLDP) packets. However, they have several limitations like flow rule preconfiguration, influence of the control plane traffic, and necessity of calibration. In this paper, Bidirectional forwarding detection (BFD) based approach is proposed. The approach measures latency at the data plane with simply implemented echo mode in Open vSwitch. We evaluates and compare the proposed approach to LLDP-based one in terms of single link latency and path latency, and error rate. In addition, we verify that the control plane throughput affects link latency according to the increased number of switches. As a result, the proposed approach can resolve the limitations and provides accuracy link latency. Seong-Mun Kim, Gyeongsik Yang, Chuck Yoo, Sung-Gi Min |
CNSM | 2 |