Kyungyong Lee 0001

dblp:99/7801-1 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-0312-4386ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Orchestrating WASM-Based MCP Tool Runtimes for AI Agents Across Edge-Cloud Continuum
Moohyun Song, Hayoung Kim, Kyoohyun Lee, Jae Gi Son, Kyungyong Lee 0001
CCGrid5
2026 Pygration : Workload-Aware Live Migratable Cloud Instance Detector for Python Runtime
abstract
Supporting live migration in the cloud can be beneficial to dynamically build a reliable and cost-optimal environment, especially when using spot instances. When a spot instance interruption event occurs, users can apply live migration using the Checkpoint/Restore In Userspace (CRIU) to a more reliable instance. In the process of migration, ensuring the compatibility of the CPU features between the source and target hosts is crucial for flawless execution after migration. However, the standard approach implemented by CRIU is workload-agnostic and overly conservative, comparing the full CPU feature sets of the hosts, which results in unnecessarily restricting the pool of viable migration targets. To mitigate this limitation, a workload-aware analysis can be employed to identify the precise set of CPU features that an application utilizes at runtime. However, it can be very challenging for Python workloads due to their layers of abstraction between bytecode and invoked native libraries, obscuring the true hardware dependencies. To overcome the challenge, this paper presents Pygration, a novel workload aware migratable cloud instance detection system for Python runtimes. Pygration implements a hybrid analysis pipeline that combines Python bytecode tracking to build a precise call graph with a native code execution path tracking heuristic to identify the minimal set of required CPU features. A comprehensive evaluation on 522 AWS instance types shows that Pygration achieves perfect precision while improving recall by over 5× compared to the CRIU baseline. In a practical spot instance scenario, this increased recall translates to a 16% improvement in median cost savings while enhancing reliability.
Soohyuk Lee, Junho Lim, Kyungyong Lee 0001
IEEE Trans. Cloud Comput.3
2025 Multi-Node Spot Instances Availability Score Collection System
abstract
Spot instances let users access unused cloud resources at significantly reduced costs. While cloud vendors offer availability information, existing tools like Spotlake only provide single-node availability data, which falls short for modern distributed applications. This paper highlighted the limitations of single-node availability data and introduced a multi-node availability dataset collection system. We analyzed the collected data and enhanced Spotlake to share these multi-node datasets publicly for broader use.
Sungkyu Cheon, Kyunghwan Kim, Moohyun Song, Kyungyong Lee 0001
HPDC5
2024 CNN Training Latency Prediction Using Hardware Metrics on Cloud GPUs
abstract
Convolutional neural network (CNN) models are becoming larger and more sophisticated over time, and training requires a significant amount of time and computing resources. To meet computing demand, graphics processing units (GPUs) are widely used for training. Due to the excessive cost and overhead of maintaining a GPU cluster, users may prefer to use GPUs provided by a public cloud vendor rather than creating their own servers. The initial cloud service is offered in the Infrastructure-as-a-Service (IaaS) model. As the cloud evolves, the abstraction level of the public cloud service becomes higher, and serverless computing is considered the next-generation cloud service. In the new way of offering cloud services, vendors are required to provide an efficient environment for diverse workloads. To meet the new requirement of the evolving cloud service, this paper proposes heuristics to predict the training latency on various GPU devices without using model information to help cloud-service vendors prepare an efficient training environment with minimal exposure to users’ model architectures. Unlike previous work that relies on internal model details for latency prediction, the proposed system uses only the hardware metrics that are extracted during training. Using the information, we first propose an algorithm to detect an epoch period that we aim to predict. The detected epoch period becomes the target latency to predict, for which we apply a stacked regressor to achieve superior prediction accuracy. Detailed experiments revealed that the average prediction accuracy of the proposed training latency prediction model is 11.17%, which is similar to the state-of-the-art approach that references the internal architecture of the model. Unlike previous work, the proposed work does not reference the model’s internal architecture, which proves the applicability of this proposed work in the next-generation cloud service.
Yoonseo Hur, Kyungyong Lee 0001
CCGrid2
2024 Workload-Aware Live Migratable Cloud Instance Detector
abstract
Cloud computing provides a variety of distinct computing resources on demand. Supporting live migration in the cloud can be beneficial to dynamically build a reliable and cost-optimal environment, especially when using spot instances. Users can apply the process of live migration technology using the Checkpoint/Restore In Userspace (CRIU) to achieve the goal. Due to the nature of live migration, ensuring the compatibility of the central processing unit (CPU) features between the source and target hosts is crucial for flawsless execution after migration. To detect migratable instances precisely while lowering false-negative detection on the cloud-scale, we propose a workload-aware migratable instance detector. Unlike the implementation of the CRIU compatibility checking algorithm, which audits the source and target host CPU features, the proposed system thoroughly investigates instructions used in a migrating process to consider CPU features that are actually in use. With a thorough evaluation under various workloads, we demonstrate that the proposed system improves the recall of migratable instance detection over 5× compared to the default CRIU implementation with 100% detection accuracy. To demonstrate its practicability, we apply it to the spot-instance environment, revealing that it can improve the median cost savings by 16% and the interruption ratio by 15% for quarter cases.
Junho Lim, Kyunghwan Kim, Kyungyong Lee 0001
CCGrid3
2024 Making Cloud Spot Instance Interruption Events Visible
abstract
Public cloud computing providers offer a surplus of computing resources at a lower price with a service of a spot instance. Despite the possible great cost savings from using spot instances, sudden resource interruption can occur as resource demand changes. To help users estimate cost savings and the possibility of interruption when using spot instances, vendors provide diverse datasets. However, the effectiveness of using the datasets has not yet been quantitatively evaluated, and many users still rely on the guess when choosing spot instances. To help users lower the chance of interruption of the spot instance for reliable usage, in this paper, we thoroughly analyze various datasets of the spot instance and present the feasibility for value prediction. Then, to measure how the public datasets reflect real-world spot instance interruption events, we conduct real-world experiments for spot instances of AWS, Azure, and Google Cloud. Combining the dataset analysis, modeling, and the real-world spot instance interruption experiment, we present a significant improvement in reducing the possibility of interruption events.
Kyunghwan Kim, Kyungyong Lee 0001
WWW2
2023 Dense or Sparse : Elastic SPMM Implementation for Optimal Big-Data Processing
abstract
Many real-world graph datasets can be represented using a sparse matrix format, and they are widely used for various big-data applications. The multiplication of two sparse matrices (SPMM) is a major kernel for various machine learning algorithms when using a sparsely expressed dataset. Apache Spark, a general-purpose big-data processing engine, includes the SPMM operation in its linear algebra package. The default Spark SPMM implementation, however, always converts a right sparse matrix to a dense format before performing multiplication, which can result in significant performance overhead for diverse SPMM scenarios. To address a limitation of the current Spark implementation, we describe an SPMM implementation that keeps the right matrix in a Compressed Sparse Column (CSC) format and propose an SPMM task latency prediction model based on a Deep Neural Network (DNN) architecture. Using the SPMM latency prediction model, we implement an elastic SPMM implementation recommendation service, which we name DoS (DenseorSparse). The proposed DoS recommends an optimal SPMM implementation method of either transforming a right matrix to a dense format or keeping it as a sparse format during the multiplication. Through evaluation of the proposed system using a real-world graph reveals that the proposed service can improve the SPMM latency of default Spark implementation by 2.2 times while shortening the overall execution time.
Unho Choi, Kyungyong Lee 0001
IEEE Trans. Big Data2
2022 PROFET: PROFiling-based CNN Training Latency ProphET for GPU Cloud Instances
abstract
Training a Convolutional Neural Network (CNN) model typically requires significant computing power, and cloud computing resources are widely used as a training environment. However, it is difficult for CNN algorithm developers to keep up with system updates and apply them to their training environment due to quickly evolving cloud services. Thus, it is important for cloud computing service vendors to design and deliver an optimal training environment for various training tasks to lessen system operation management overhead of algorithm developers. To achieve the goal, we propose PROFET, which can predict the training latency of arbitrary CNN implementation on various Graphical Processing Unit (GPU) devices to develop a cost-effective and time-efficient training cloud environment. Different from the previous training latency prediction work, PROFET does not rely on the implementation details of the CNN architecture, and it is suitable for use in a public cloud environment. Thorough evaluations reveal the superior prediction accuracy of PROFET compared to the state-of-the-art related work, and the demonstration service presents the practicality of the proposed system.
Sungjae Lee 0004, Yoonseo Hur, Kyungyong Lee 0001
IEEE Big Data4
2022 MPEC: Distributed Matrix Multiplication Performance Modeling on a Scale-Out Cloud Environment for Data Mining Jobs
abstract
Many data mining workloads are being analyzed in large-scale distributed cloud computing environments which provide nearly infinite resources with diverse hardware configurations. To maintain cost-efficiency in such environments, understanding the characteristics and estimating the overheads of a distributed matrix multiplication task that is a core computation kernel in many machine learning algorithms are essential. This article aims to propose a Matrix Multiplication Performance Estimator on Cloud (MPEC) algorithm. The proposed algorithm predicts the latency incurred when executing distributed matrix multiplication tasks of various input sizes and shapes with diverse instance types and a different number of worker nodes on cloud computing environments. To achieve this goal, we first analyze the characteristics of distributed matrix multiplication tasks. With characteristics generated from qualitative analysis, we propose to apply an ensemble of non-linear regression algorithm to predict the execution time of arbitrary matrix multiplication tasks. Thorough experimental results reveal that the proposed algorithm demonstrates higher accuracy than a state-of-the-art machine learning task performance estimation engine, Ernest, by decreasing the Mean Absolute Percentage Error (MAPE) in half.
Jeongchul Kim, Myungjun Son, Kyungyong Lee 0001
IEEE Trans. Cloud Comput.3
2021 Accelerator-Aware Kubernetes Scheduler for DNN Tasks on Edge Computing Environment
Jungae Park, Un-Sook Choi, Jaewon Moon, Kyungyong Lee 0001
SEC5
2020 Evaluating Concurrent Executions of Multiple Function-as-a-Service Runtimes with MicroVM
abstract
Serverless computing and public Function-as-a-Service (FaaS) systems are gaining significant attention because they help easily build a highly available system. With recent advances in micro virtual machines (microVM), the internal architecture of FaaS systems substantially changes. This paper focuses on a thorough investigation of the recent improvement in public FaaS systems concerning numerous concurrent executions. The adoption of microVM has changed the nature of FaaS, especially for runtime reservations. As a result, the performance degradation has decreased significantly compared to the previous generation FaaS, as shown in the experiment.
Jungae Park, Hyunjune Kim, Kyungyong Lee 0001
CLOUD3
2019 FunctionBench: A Suite of Workloads for Serverless Cloud Function Service
abstract
Serverless computing is attracting considerable attention recently, but many published papers use micro-benchmarks for evaluation that might result in impracticality. To address this, we present FunctionBench, a suite of practical function workloads for public services. It contains realistic data-oriented applications that utilize various resources during execution. The source codes customized for various cloud service providers are publicly available. We are positive that it suggests opportunities for new function applications with lessen experiment setup overheads.
Jeongchul Kim, Kyungyong Lee 0001
CLOUD2
2019 RnR: Extraction of Visual Attributes from Large-Scale Fashion Dataset
abstract
Supervised visual perception models require a large number of annotated inputs so that a model can be trained and validated. However, manual annotation of large-scale datasets can be prohibitive in terms of time and expense. We present an automatic fashion image tagging algorithm that uses detailed item descriptions and carefully selected category information produced by a service provider. Compared to previous work, which used a relatively simple frequency-based heuristic, a set of image attributes extracted by our proposed algorithm could be used to create a model (ResNet-18 architecture) with exhibited 8% higher accuracy than a model created with attributes suggested by DeepFashion. This in-progress work identifies opportunities that carefully chosen visually recognizable attributes can be used to help build a more general and descriptive model.
Sungjae Lee 0004, Yeonji Lee, Kyungyong Lee 0001
IEEE BigData4
2019 Practical Cloud Workloads for Serverless FaaS
abstract
Serverless computing is gaining popularity with the Function-asa-Service (FaaS) execution model. Without incurring overheads involved in provisioning cloud instances and with high availability and scalability, serverless computing allows developers to focus on implementation of core application logic using other well-developed cloud services. By abstracting the complex resource management task, serverless computing opens new opportunities for the cloud service adoption even to non-cloud experts [2]. With the popularity, many research results have been published using the FaaS execution model. They include investigation of serverless computing opportunities [1], proposing new serverless applications, function run-time optimization, and public service comparison. Without a common test benchmark suite, authors in the previous work had evaluated proposed systems using fairly simple FaaS applications, such as micro-benchmarks that emphasize specific resources exclusively, e.g., CPU, disk I/O, and network. However, such simple workloads do not represent realistic FaaS system applications, and the evaluations might not compare proposed systems appropriately.
Jeongchul Kim, Kyungyong Lee 0001
SoCC2
2018 Distributed Matrix Multiplication Performance Estimator for Machine Learning Jobs in Cloud Computing
abstract
Matrix multiplication is an important kernel task in many machine learning algorithms. As the size of input datasets increases, multiple workloads are analyzed in large-scale distributed cloud computing environments. Therefore, understanding the characteristics of a distributed matrix multiplication task is essential for running machine learning jobs in the cloud. Herein, we propose Matrix multiplication Performance Estimator for Cloud computing, a method to predict the latency of matrix multiplication of various sizes and shapes in a distributed cloud computing environment. We first characterize the overhead of a distributed matrix multiplication task and propose features to model the latency of a task with different input types. Using the proposed features, a latency prediction model is developed by applying a data mining algorithm and a parameter optimization step iteratively. In experiments with 236 distinct types of matrix multiplications on diverse cloud instances running Apache Spark, we confirm that the proposed method can model the latency of various types of matrix multiplication tasks effectively and capture the non-linear interactions among the proposed features. A comparison with the state-of-the-art cloud computing performance predictor, Ernest, reveals that the proposed method provides 63% lower Root Mean Square Error (RMSE) for a distributed matrix multiplication latency prediction task and confirms the uniqueness of the distributed matrix multiplication workload.
Myungjun Son, Kyungyong Lee 0001
IEEE CLOUD2
2017 DeepSpotCloud: Leveraging Cross-Region GPU Spot Instances for Deep Learning
abstract
Cloud computing resources that are equipped with GPU devices are widely used for applications that require extensive parallelism, such as deep learning. When the demand of cloud computing instance is low, the surplus of resources is provided at a lower price in the form of spot instance by AWS EC2. This paper proposes DeepSpotCloud that utilizes GPU-equipped spot instances to run deep learning tasks in a cost efficient and fault-tolerant way. Thorough analysis about spot instance price history logs reveals that GPU spot instances show more dynamic price change pattern than other general types of cloud computing resources. To deal with the price dynamicity of the GPU spot instance, DeepSpotCloud utilizes instances in different regions across continents as a single resource pool. This paper also proposes a task migration heuristic by utilizing a checkpointing mechanism of existing deep learning analysis platform to conduct fast task migration when a running spot instance is interrupted. Extensive experiments using real AWS services prove that the proposed task migration method is effective even in a WAN environment with limited network bandwidth. Comprehensive simulations by replaying AWS EC2 price history logs reveal that DeepSpotCloud can achieve 13% more cost gain than a state-of-the-art interrupt-driven scheduling policy. The prototype of DeepSpotCloud is implemented using various cloud computing services provided by AWS to serve real deep learning tasks.
Kyungyong Lee 0001, Myungjun Son
CLOUD1
2013 Enabling decentralized microblogging through P2PVPNs
abstract
In the past few years, many peer-to-peer microblogging solutions have been proposed and/or implemented utilizing various technologies such as DHTs, multicast trees, and/or gossip protocols. These previous works address the issue of privacy and performance in a variety of ways including the use of session keys for message encryption or direct connections for low latency communication. We propose a decentralized microblogging service which takes advantage of available peer-to-peer virtual private networking (P2PVPN) technologies which provide privacy and low-latency communication in the common case of P2P messaging among social peers. Leveraging the private IP connectivity of P2PVPNs, our design utilizes both IP multicasting and random walks to ensure that peers are able to publish messages with varying degree of scope (i.e. friends, friends of friends, and/or the public). We study the implications of our data dissemination mechanism for a decentralized microblogging service through simulation-based analysis based on synthetic social graphs. Overall, our experimental results show that peers can effectively follow each other's updates with acceptable overhead. Through the use of our pseudo-random-walk algorithm, we estimate that, in a 900K social graph, with a TTL of 100, a user can retreive updates from anyone in the social graph 55% of the time, but by increasing the TTL to 400 that hit rate increases to 95%.
Pierre St. Juste, Heungsik Eom, Kyungyong Lee 0001, Renato J. O. Figueiredo
CCNC3
2013 MatchTree: Flexible, scalable, and fault-tolerant wide-area resource discovery with distributed matchmaking and aggregation
Kyungyong Lee 0001, Tae Woong Choi, P. Oscar Boykin, Renato J. O. Figueiredo
Future Gener. Comput. Syst.1
2012 MapReduce on opportunistic resources leveraging resource availability
abstract
MapReduce is a popular large-scale parallel data processing framework. In the context of MapReduce processing on volunteer computing environments, it is important to devise scheduling and data placement policies that account for characteristics of opportunistic resources. This paper investigates availability characteristics of opportunistic resources with analyses based on log traces from the SETI@Home project. Based on the analysis, the paper devises heuristics to leverage the uptime of each available session to detect possibly long-lasting resources. Our proposed session uptime-based resource availability prediction approach shows a two-fold reduction in the number of service disturbance compared to an availability-rate based model. The paper paper investigates a heuristic that differentiates stable nodes from unstable nodes while increasing the chance of leveraging existing data blocks.
Kyungyong Lee 0001, Renato J. O. Figueiredo
CloudCom1
2012 PonD: dynamic creation of HTC pool on demand using a decentralized resource discovery system
abstract
High Throughput Computing (HTC) platforms aggregate heterogeneous resources to provide vast amounts of computing power over a long period of time. Typical HTC systems, such as Condor and BOINC, rely on central managers for resource discovery and scheduling. While this approach simplifies deployment, it requires careful system configuration and management to ensure high availability and scalability. In this paper, we present a novel approach that integrates a self-organizing P2P overlay for scalable and timely discovery of resources with unmodified client/server job scheduling middleware in order to create HTC virtual resource Pools on Demand (PonD). This approach decouples resource discovery and scheduling from job execution/monitoring - a job submission dynamically generates an HTC platform based upon resources discovered through match-making from a large "sea" of resources in the P2P overlay and forms a "PonD" capable of leveraging unmodified HTC middleware for job execution and monitoring. We show that job scheduling time of our approach scales with O(log N), where N is the number of resources in a pool, through first-order analytical models and large-scale simulation results. To verify the practicality of PonD, we have implemented a prototype using Condor (called C-PonD), a structured P2P overlay, and a PonD creation module. Experimental results with the prototype in two WAN environments (PlanetLab and the FutureGrid cloud computing testbed) demonstrates the utility of C-PonD as a HTC approach without relying on a central repository for maintaining all resource information. Though the prototype is based on Condor, the decoupled nature of the system components - decentralized resource discovery, PonD creation, job execution/monitoring - is generally applicable to other grid computing middleware systems.
Kyungyong Lee 0001, David Wolinsky, Renato J. O. Figueiredo
HPDC1
2010 SocialDNS: A decentralized naming service for collaborative P2P VPNs
abstract
The ability to define domain names for resources in a collaborative virtual organization is usually reserved to network administrators through centralized domain name servers. We propose SocialDNS, a decentralized, naming service that gives individual collaborators the power to choose the domain nam
Pierre St. Juste, David Wolinsky, Kyungyong Lee 0001, P. Oscar Boykin, Renato J. O. Figueiredo
CollaborateCom3
2010 On the design of autonomic, decentralized VPNs
abstract
Decentralized and P2P (peer-to-peer) VPNs (virtual private networks) have recently become quite popular for connecting users in small to medium collaborative environments, such as academia, businesses, and homes. In the realm of VPNs, there exist centralized, decentralized, and P2P solutions. Centra
David Wolinsky, Kyungyong Lee 0001, P. Oscar Boykin, Renato J. O. Figueiredo
CollaborateCom2