VLDB 2026 Research / reviewers in the wild / expert
In Kee Kim
dblp:96/69
· DBLP profile ↗
25ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0003-1330-7784ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Coral: Covariance-Guided Resource Adaptive Learning for Efficient Edge Inference
Ahmad N. L. Nabhaan, Zaki Sukma, Rakandhiya D. Rachmanto, Muhammad Husni Santriaji, Byungjin Cho, Arief Setyanto, In Kee Kim |
ICFEC | 7 |
| 2026 | $\varDelta $-NeRF: Incremental Refinement of Neural Radiance Fields Through Residual Control and Knowledge Transfer
Kriti Ghosh, Devjyoti Chakraborty, Lakshmish Ramaswamy, Suchendra M. Bhandarkar, In Kee Kim, Nancy O'Hare, Deepak Mishra 0005 |
ICPR (9) | 5 |
| 2024 | An Empirical Evaluation of the Impact of Solar Correction in NeRFs for Satellite Imagery
Devjyoti Chakraborty, Kriti Ghosh, Zaki Sukma, In Kee Kim, Lakshmish Ramaswamy, Suchendra M. Bhandarkar, Deepak Mishra 0005 |
ICPR (18) | 4 |
| 2023 | CNT: Semi-Automatic Translation from CWL to Nextflow for Genomic WorkflowsabstractWith the rise of advanced workflow languages for scientific computations, Nextflow has gained increased attention from the bioinformatics community. Nextflow offers native support for advanced parallelism, which can greatly enhance resource utilization and throughput. Still, a significant portion of bioinformatics workflows are developed with the Common Workflow Language (CWL). Transitioning from CWL to Nextflow poses a significant challenge due to the differences in programming models, scripting language compatibilities, and the prerequisite for in-depth knowledge in both languages. To address this challenge, we present CNT, a novel, semi-automated translator converting CWL workflows into Nextflow ones. At its core, CNT uses an automated translation mechanism that converts the CommandLineTool, the most basic unit of CWL, into Nextflow's Process class. This component integrates tool-level conversion, graph dependency analysis, and correctness checks to provide highly automated translation coverage, significantly reducing the development time while satisfying language-specific requirements like building a proper dataflow model when creating workflows. Furthermore, CNT incorporates a module for aiding manual translation. Specifically, it can identify three common JavaScript patterns in CWL workflows, offering further guidance for developers during the translation phase. We evaluated CNT with production-grade workflows and found that it can cover up to 81% of the original workflows, substantially reducing development time. Additionally, transitioning from a cwltool-based system to Nextflow with CNT can result in a 72% speedup and 85% increased CPU utilization. Martin L. Putra, In Kee Kim, Haryadi S. Gunawi, Robert L. Grossman |
BIBE | 2 |
| 2023 | DynaES: Dynamic Energy Scheduling for Energy Harvesting Environmental SensorsabstractLow-cost sensors and IoT technologies have facilitated the deployment of environmental sensors to collect and analyze various factors, such as soil properties. Due to the lack of electric power networks in many deployment locations, these sensors rely on energy harvesting (EH) systems that use natural energy sources such as solar power. Specifically, solar-powered EH systems benefit from timely obtaining weather information for efficient future energy scheduling. However, obtaining weather information for EH environmental sensors is always challenging, as they are commonly deployed in harsh environments without network access. To address this problem, we present DynaES, a novel energy scheduling method for EH sensors without relying on online weather forecasts. DynaES comprises two components: a DC power gain estimator that predicts future power gain by individually estimating changes in environmental parameters and ensembling them, and a dynamic energy scheduler that distributes energy to sensors based on priority and adjusts sensing intervals and frequency. We evaluate DynaES via simulation-based studies on real-world datasets and compare its performance against state-of-the-art baselines. Evaluation results show that DynaES accurately predicts future energy gain with low estimation errors and enables $1.8 \times$ – $4 \times$ more frequent sensing operations with shorter sensing intervals while achieving longer operation hours without complete battery drains. Jianwei Hao, Emmanuel Oni, In Kee Kim, Lakshmish Ramaswamy |
IPCCC | 3 |
| 2023 | Reaching for the Sky: Maximizing Deep Learning Inference Throughput on Edge Devices with AI Multi-TenancyabstractThe wide adoption of smart devices and Internet-of-Things (IoT) sensors has led to massive growth in data generation at the edge of the Internet over the past decade. Intelligent real-time analysis of such a high volume of data, particularly leveraging highly accurate deep learning (DL) models, often requires the data to be processed as close to the data sources (or at the edge of the Internet) to minimize the network and processing latency. The advent of specialized, low-cost, and power-efficient edge devices has greatly facilitated DL inference tasks at the edge. However, limited research has been done to improve the inference throughput (e.g., number of inferences per second) by exploiting various system techniques. This study investigates system techniques, such as batched inferencing, AI multi-tenancy, and cluster of AI accelerators, which can significantly enhance the overall inference throughput on edge devices with DL models for image classification tasks. In particular, AI multi-tenancy enables collective utilization of edge devices’ system resources (CPU, GPU) and AI accelerators (e.g., Edge Tensor Processing Units; EdgeTPUs). The evaluation results show that batched inferencing results in more than 2.4× throughput improvement on devices equipped with high-performance GPUs like Jetson Xavier NX. Moreover, with multi-tenancy approaches, e.g., concurrent model executions (CME) and dynamic model placements (DMP), the DL inference throughput on edge devices (with GPUs) and EdgeTPU can be further improved by up to 3× and 10×, respectively. Furthermore, we present a detailed analysis of hardware and software factors that change the DL inference throughput on edge devices and EdgeTPUs, thereby shedding light on areas that could be further improved to achieve high-performance DL inference at the edge. Jianwei Hao, Piyush Subedi, Lakshmish Ramaswamy, In Kee Kim |
ACM Trans. Internet Techn. | 4 |
| 2022 | Wasserstein Adversarial Transformer for Cloud Workload PredictionabstractPredictive VM (Virtual Machine) auto-scaling is a promising technique to optimize cloud applications’ operating costs and performance. Understanding the job arrival rate is crucial for accurately predicting future changes in cloud workloads and proactively provisioning and de-provisioning VMs for hosting the applications. However, developing a model that accurately predicts cloud workload changes is extremely challenging due to the dynamic nature of cloud workloads. Long- Short-Term-Memory (LSTM) models have been developed for cloud workload prediction. Unfortunately, the state-of-the-art LSTM model leverages recurrences to predict, which naturally adds complexity and increases the inference overhead as input sequences grow longer. To develop a cloud workload prediction model with high accuracy and low inference overhead, this work presents a novel time-series forecasting model called WGAN-gp Transformer, inspired by the Transformer network and improved Wasserstein-GANs. The proposed method adopts a Transformer network as a generator and a multi-layer perceptron as a critic. The extensive evaluations with real-world workload traces show WGAN- gp Transformer achieves 5× faster inference time with up to 5.1% higher prediction accuracy against the state-of-the-art. We also apply WGAN-gp Transformer to auto-scaling mechanisms on Google cloud platforms, and the WGAN-gp Transformer-based auto-scaling mechanism outperforms the LSTM-based mechanism by significantly reducing VM over-provisioning and under-provisioning rates. Shivani Arbat, Vinodh Kumaran Jayakumar, Wei Wang 0054, In Kee Kim |
AAAI | 5 |
| 2022 | CloudBruno: A Low-Overhead Online Workload Prediction Framework for Cloud ComputingabstractAccurate prediction of future incoming workloads to cloud applications, such as future user request count, is critical to proactive auto-scaling, and in general, critical to the cost-effectiveness of cloud deployments. However, designing a generic predictive framework that can accurately predict for any types of workloads is difficult, especially when the workload is dynamic and can change to a pattern that has not been observed in the training data sets. However, existing workload prediction solutions typically rely on complex machine learning models, which require comprehensive training data, making it difficult for them to handle dynamic workloads. Moreover, the training of existing workload prediction solutions are also expensive in terms of both time and computing resources. This paper presents a generic and low-cost online workload prediction framework, called Cloud Bruno, which combines the more accurate LSTM models with less expensive but fast SVM models to achieve high accuracy and low training overhead. When compared to existing predictors, CloudBruno had at least 8.8 % lower error than existing deep learning-based predictors for a highly-dynamic workload that does not have comprehensive training data (i.e, has changes unknown to training data). For workloads with comprehensive training data, Cloud Bruno's error was at most 2.5 % higher than optimized deep learning-based predictors. More importantly, Cloud Bruno can effectively execute on a free cloud CPU, allowing it to be used as an online workload predictor without additional cost. Vinodh Kumaran Jayakumar, Shivani Arbat, In Kee Kim, Wei Wang 0054 |
IC2E | 3 |
| 2022 | Privacy invasion via smart-home hub in personal area networks
Omid Setayeshfar, Karthika Subramani, Xingzi Yuan, Raunak Dey, Dezhi Hong, In Kee Kim, Kyu Hyung Lee |
Pervasive Mob. Comput. | 6 |
| 2022 | Guaranteeing Performance SLAs of Cloud Applications Under Resource StormsabstractIn modern data centers, enterprise cloud instances run not only foreground applications like web and databases, but also different background services (e.g., backup, virus/compliance scan, batch) to manage the cloud instances securely and improve the overall resource utilization. These background services often incur resource storms that suddenly consume a lot of shared resources on cloud instances. The resource storms significantly degrade the performance of foreground applications by interfering in the preemption of the shared resources, resulting in frequent SLA violations. However, stock OS schedulers are not designed to handle these situations, and prior works are insufficient to address such resource storms under highly dynamic cloud workloads. This article presents Orchestra, a cloud-specific framework for controlling multiple applications in the user space, aiming at meeting corresponding SLAs. Orchestra takes an online approach with lightweight monitoring and performance models for both applications on the fly. It optimizes the resource allocations to meet corresponding SLAs. We evaluate the performance of Orchestra on a production cloud with a diverse range of SLAs. Orchestra guarantees the foreground application's performance SLAs at all times. At the same time, Orchestra maintains the background's performance by minimizing its performance penalty with proper allocation of the shared resources. In Kee Kim, Jinho Hwang, Wei Wang 0054, Marty Humphrey |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | Forecasting Cloud Application Workloads With CloudInsight for Predictive Resource ManagementabstractPredictive cloud resource management has been widely adopted to overcome the limitations of reactive cloud autoscaling. The predictive resource management is highly relying on workload predictors, which estimate short-/long-term fluctuations of cloud application workloads. These predictors tend to be pre-optimized for specific workload patterns. However, such predictors are still insufficient to handle real-world cloud workloads whose patterns may be unknown a priori, may dynamically change over time and may be irregular. As a result, these predictors often cause over-/under-provisioning of cloud resources. To address this problem, we have created CloudInsight, a novel cloud workload prediction framework, leveraging the combined power of multiple workload predictors. CloudInsight creates an ensemble model using multiple predictors to make accurate predictions for real workloads. The weights of the predictors in CloudInsight are determined at runtime with their accuracy for the current workload using multi-class regression. The ensemble model is periodically optimized to handle sudden changes in the workload. We evaluated CloudInsight with various real workload traces. The results show that CloudInsight has 13–27 percent higher accuracy than state-of-the-art predictors. Moreover, the results from trace-based simulations with a cloud resource manager show that CloudInsight has 15–20 percent less under-/over-provisioning periods, resulting in high cost-efficiency and low SLA violations. In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | An Empirical Analysis of VM Startup Times in Public IaaS CloudsabstractVM startup time is an essential factor in designing elastic cloud applications. VM autoscaling can reduce the under-/over-provisioning period of VMs with a precise estimation of VM startup time, and in turn, it can guarantee the application's SLOs with improved cost-efficiency. Unfortunately, VM startup time has been little studied, and previous measurement results did not consider various configurations of VMs. This work performs a thorough analysis of VM startup times in two IaaS clouds (AWS, GCP). Specifically, we collected 300,000 data points from each provider by applying diverse VM configurations. i.e., different VM types, image sizes, location, purchase models. With extensive analysis, we found several important factors that can change VM startup time significantly. Moreover, by comparing with a previous study, we confirm that AWS made significant improvements for reducing VM startup times and currently has much quicker VM startup times than in the past. Jianwei Hao, Wei Wang 0054, In Kee Kim |
CLOUD | 4 |
| 2021 | AI Multi-Tenancy on Edge: Concurrent Deep Learning Model Executions and Dynamic Model Placements on Edge DevicesabstractMany real-world applications are widely adopting the edge computing paradigm due to its low latency and better privacy protection. With notable success in AI and deep learning (DL), edge devices and AI accelerators play a crucial role in deploying DL inference services at the edge of the Internet. While prior works quantified various edge devices’ efficiency, most studies focused on the performance of edge devices with single DL tasks. Therefore, there is an urgent need to investigate AI multi-tenancy on edge devices, required by many advanced DL applications for edge computing. This work investigates two techniques – concurrent model executions and dynamic model placements – for AI multi-tenancy on edge devices. With image classification as an example scenario, we empirically evaluate AI multi-tenancy on various edge devices, AI accelerators, and DL frameworks to identify its benefits and limitations. Our results show that multi-tenancy significantly improves DL inference throughput by up to 3.3 × − 3.8 × on Jetson TX2. These AI multi-tenancy techniques also open up new opportunities for flexible deployment of multiple DL services on edge devices and AI accelerators. Piyush Subedi, Jianwei Hao, In Kee Kim, Lakshmish Ramaswamy |
CLOUD | 3 |
| 2021 | Performance Testing for Cloud Computing with Dependent Data BootstrappingabstractTo effectively utilize cloud computing, cloud practice and research require accurate knowledge of the performance of cloud applications. However, due to the random performance fluctuations, obtaining accurate performance results in the cloud is extremely difficult. To handle this random fluctuation, prior research on cloud performance testing relied on a non-parametric statistic tool called bootstrapping to design their stop criteria. However, in this paper, we show that the basic bootstrapping employed by prior work overlooks the internal dependency within cloud performance test data, which leads to inaccurate performance results.We then present Metior, a novel automated cloud performance testing methodology, which is designed based on statistical tools of block bootstrapping, the law of large numbers, and autocorrelation. These statistical tools allow Metior to properly consider the internal dependency within cloud performance test data. They also provide better coverage of cloud performance fluctuation and reduce the testing cost. Experimental evaluation on two public clouds showed that 98% of Metior’s tests could provide performance results with less than 3% error. Metior also significantly outperformed existing cloud performance testing methodologies in terms of accuracy and cost – with up to 14% increase in the accurate test count and up to 3.1 times reduction in testing cost. Sen He 0002, Palden Lama, In Kee Kim, Wei Wang 0054 |
ASE | 5 |
| 2021 | ChatterHub: Privacy Invasion via Smart Home HubabstractSmart-home devices promise to make users’ lives more convenient. However, at the same time, such devices increase the possibility of breaching users’ privacy as they are tightly connected to the users’ daily lives and activities. To address privacy invasion through smart-home devices, we present ChatterHub. This novel approach accurately identifies smart-home devices’ activities with minimal monitoring of encrypted traffic in the home network. ChatterHub targets devices that can only connect to the Internet through a centralized smart-home hub (e.g., Samsung SmartThings) using Zigbee or Z-wave. Specifically, ChatterHub passively eavesdrops on encrypted network traffic from the hub and leverages machine learning techniques to classify events and states of smart-home devices. Using ChatterHub, an adversary can identify smart-home devices’ specific activities without prior knowledge of the target smart home (e.g., list of deployed devices, types of communication protocols). We evaluated the accuracy and efficiency of ChatterHub in three real-world smart-home environments, and the evaluation results show that an attacker can successfully disclose smart-home devices’ behaviors with over 88% F1 score. We further demonstrate that ChatterHub successfully recognizes privacy-sensitive activities, including open and close of a smart door lock and turn on and off of smart LED. Additionally, to mitigate the threats posed by ChatterHub, we introduce two approaches, packet padding and random sequence injection. These mitigation approaches can effectively prevent threats from ChatterHub with only 9.2MB of additional network traffic per day. Omid Setayeshfar, Karthika Subramani, Xingzi Yuan, Raunak Dey, Dezhi Hong, Kyu Hyung Lee, In Kee Kim |
SMARTCOMP | 7 |
| 2020 | A Self-Optimized Generic Workload Prediction Framework for Cloud ComputingabstractThe accurate prediction of the future workload, such as the job arrival rate and the user request rate, is critical to the efficiency of resource management and elasticity in the cloud. However, designing a generic workload predictor that works properly for various types of workload is very challenging due to the large variety of workload patterns and the dynamic changes within a workload. Because of these challenges, existing workload predictors are usually hand-tuned for specific (types of) workloads for maximum accuracy. This necessity to individually tune the predictors also makes it very difficult to reproduce the results from prior research, as the predictor designs have a strong dependency on the workloads.In this paper, we present a novel generic workload prediction framework, LoadDynamics, that can provide high accuracy predictions for any workloads. LoadDynamics employs Long-Short-Term-Memory models and can automatically optimize its internal parameters for an individual workload to achieve high prediction accuracy. We evaluated LoadDynamics with a mixture of workload traces representing public cloud applications, scientific applications, data center jobs and web applications. The evaluation results show that LoadDynamics have only 18% prediction error on average, which is at least 6.7% lower than state-of-the-art workload prediction techniques. The error of LoadDynamics was also only 1% higher than the best predictor found by exhaustive search for each workload. When applied in the Google Cloud, LoadDynamics-enabled auto-scaling policy also outperformed the state-of-the-art predictors by reducing the job turnaround time by at least 24.6% and reducing virtual machine over-provisioning by at least 4.8%. Vinodh Kumaran Jayakumar, In Kee Kim, Wei Wang 0054 |
IPDPS | 3 |
| 2018 | CloudInsight: Utilizing a Council of Experts to Predict Future Cloud Application WorkloadsabstractMany predictive approaches have been proposed to overcome the limitations of reactive autoscaling on clouds. These approaches leverage workload predictors that are usually targeted for a particular workload pattern and can fail to handle real-world cloud workloads whose patterns may be unknown a priori, may dynamically change over time, or may be irregular. The result is that resources are frequently under-and overprovisioned. To address this problem, we create a novel cloud workload prediction framework called CloudInsight, leveraging the combined power of multiple workload predictors that collectively provide a "council of experts". The weights of the predictors in this ensemble model are determined in real-time based on their accuracy for current workload using multi-class regression. Under real workload traces, CloudInsight has 13% - 27% better accuracy than state-of-the-art predictors. It also has low overhead for predicting future workload changes (<; 100 ms) and creating a new ensemble workload predictor (<; 1.1 sec.). In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
IEEE CLOUD | 1 |
| 2018 | Orchestra: Guaranteeing Performance SLAs for Cloud Applications by Avoiding Resource StormsabstractThis paper presents Orchestra, a cloud-specific framework for managing both foreground applications (e.g., Web, DBMS) and background services (e.g., backup, security check, batch jobs) in the user space. Orchestra is designed to address "resource storms" caused by sudden executions of the background services on the cloud instances. The resource storms significantly degrade the performance of foreground applications by interfering in the preemption of the shared resources, resulting in frequent SLA violations and poor user experience. Orchestra takes an online approach using lightweight monitoring and creates performance models for multiple cloud applications on the fly. It then optimizes the allocations of shared resources to meet SLAs. We evaluate the performance of Orchestra on a production cloud (Amazon EC2) with a diverse range of SLA requirements. The experiment results show that Orchestra successfully guarantees the foreground application's performance to meet its SLA targets at all times. Moreover, Orchestra maintains the background's performance by minimizing its performance penalty with proper allocation of the shared resources. In Kee Kim, Jinho Hwang, Wei Wang 0054, Marty Humphrey |
ISPDC | 1 |
| 2017 | iCSI: A Cloud Garbage VM Collector for Addressing Inactive VMs with Machine LearningabstractAccording to a recent study, 30% of VMs in private cloud data centers are "comatose", in part because there is generally no strong incentive for their human owners to delete them at an appropriate time. These inactive VMs are still scheduled and executed on physical cloud resources, taking valuable access away from productive VMs. In an extreme, cloud infrastructure may deny legitimate requests for new VMs because capacity limits have been hit. It is not sufficient for cloud infrastructure to identify such inactive VMs by monitoring resource utilization (e.g., CPU utilization) - e.g., management processes (e.g. virus-scan, software update) on inactive VMs often consume high CPU and memory resources, and active VMs with lightweight jobs (e.g. text editing) show almost zero resource utilization. To properly detect and address such inactive VMs, we present iCSI: a cloud garbage VM collector to improve resource utilization and cost efficiency of enterprise data centers. iCSI includes three main components, a lightweight data collector, a VM identification model and a recommendation engine. The data collector periodically gathers primitive information from VMs. The identification model infers the purpose of a VM from the data collection and extracts the most relevant features associated with the purpose. The recommendation engine offers proper actions to end users i.e., suspending or resizing VMs. In this prototype phase, iCSI is deployed into multiple data centers in IBM and manages more than 750 production VMs. iCSI achieves 20% better accuracy (90%) in identifying active/inactive VMs compared with state-of-the-art methods. With recommendations to end users, our estimation results show that iCSI can improve internal cost efficiency with 23% and resource utilization more than 45%. In Kee Kim, Sai Zeng, Christopher C. Young, Jinho Hwang, Marty Humphrey |
IC2E | 1 |
| 2016 | Empirical Evaluation of Workload Forecasting Techniques for Predictive Cloud Resource ScalingabstractMany predictive resource scaling approaches have been proposed to overcome the limitations of the conventional reactive approaches most often used in clouds today. In general, due to the complexity of clouds, these reactive approaches were often forced to make significant limiting assumptions in either the operating conditions/requirements or expected workload patterns. As such, it is extremely difficult for cloud users to know which - if any - existing workload predictor will work best for their particular cloud activity, especially when considering highly-variable workload patterns, non-trivial billing models, variety of resources to add/subtract, etc. To solve this problem, we conduct comprehensive evaluations for a variety of workload predictors under real-world cloud configurations. The workload predictors cover four classes of 21 predictors: naive, regression, temporal, and non-temporal methods. We simulate a cloud application under four realistic workload patterns, two different cloud billing models, and three different styles of predictive scaling. Our evaluation confirms that no workload predictor is universally best for all workload patterns, and shows that Predictive Scaling-out + Predictive Scaling-in has the best cost efficiency and the lowest job deadline miss rate in cloud resource management, on average providing 30% better cost efficiency and 80% less job deadline miss rate compared to other styles of predictive scaling. In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
CLOUD | 1 |
| 2015 | PICS: A Public IaaS Cloud SimulatorabstractPublic clouds become essential for many organizations to run their applications because they provide huge financial benefits and great flexibility. However, it is very challenging to accurately evaluate the performance and cost of applications without actual deployment on the clouds. Existing cloud simulators are generally designed from the perspective of cloud service providers, thus they can be under-developed for answering questions for the perspective of cloud users. To solve this prediction and evaluation problem, we created a Public Cloud IaaS Simulator (PICS). PICS enables the cloud user to evaluate the cost and performance of public IaaS clouds along with such dimensions like VM and storage service, resource scaling, job scheduling, and diverse workload patterns. We extensively validated PICS by comparing its results with the data acquired from real public IaaS cloud using real cloud-applications. We show that PICS provides highly accurate simulation results (less than 5% of average errors) under a variety of use cases. Moreover, we evaluated PICS' sensitivity with imprecise simulation parameters. The results show that PICS still provides very reliable simulation results with imprecise simulation parameters and performance uncertainty. In Kee Kim, Wei Wang 0054, Marty Humphrey |
CLOUD | 1 |
| 2015 | Toward Optimal Resource Provisioning for Cloud MapReduce and Hybrid Cloud ApplicationsabstractGiven the variety of resources available in public clouds and locally (hybrid clouds), it can be very difficult to determine the best number and type of resources to allocate (and where) for a given activity. In order to solve this problem we first define the requested computation in terms of an Integer Linear Programming (ILP) problem and then use an efficient ILP solver to make a provisioning decision in a few milliseconds. Our approach is based on the two most important metrics for the user: cost and job execution time. Thus, based on the user's preferences we can favor solutions that optimize speed or cost or a certain combination of both (e.g. Cheapest solution that meets a certain deadline). We evaluate our approach with two classes of cloud applications: MapReduce applications, and Monte Carlo simulations. A significant advantage in our approach is that our solution has been proved optimal by the ILP solver, the set of the scheduling decisions based on our model are plotted on a time vs. Cost graph that forms a Pareto efficient frontier. This way, we can avoid the pitfalls of a naïve strategy that can lead to a great increase in cost (91%) or job running time (21%) compared to optimal. Arkaitz Ruiz-Alvarez, In Kee Kim, Marty Humphrey |
CLOUD | 2 |
| 2015 | WDCloud: An end to end system for large-scale watershed delineation on cloudabstractWatershed delineation is a process to compute the drainage area for a point on the land surface, which is a critical step in hydrologic and water resources analysis. However, existing watershed delineation tools are still insufficient to support hydrologists and watershed researchers due to the lack of essential capabilities such as fully leveraging scalable and high performance computing infrastructure (public cloud), and providing predictable performance for the delineation tasks. To solve these problems, this paper reports on WDCloud, which is a system for large-scale watershed delineation on public cloud. For the design and implementation of WDCloud, we employ three main approaches: 1) an automated catchment search mechanism for a public data set, 2) three performance improvement strategies (Data-reuse, parallel-union, and MapReduce), and 3) local linear regression-based execution time estimator for watershed delineation. Moreover, WDCloud extensively utilizes several compute and storage capabilities from Amazon Web Services in order to maximize the performance, scalability, and elasticity of watershed delineation system. Our evaluations on WDCloud focus on two main aspects of WDCloud; the performance improvement for watershed delineation via three strategies and the estimation accuracy for watershed delineation time by local linear regression. The evaluation results show that WDCloud can achieve 18x-111x of speed-ups for delineating any scale of watersheds in the contiguous United States as compared to commodity laptop environments, and accurately predict execution time for watershed delineation with 85.6% of prediction accuracy, which is 23%-13% higher than other state-of-the-art approaches. In Kee Kim, Jacob Steele, Anthony M. Castronova, Jonathan L. Goodall, Marty Humphrey |
IEEE BigData | 1 |
| 2013 | CloudDRN: A Lightweight, End-to-End System for Sharing Distributed Research Data in the CloudabstractThe cloud has proven itself as a scalable platform for Web-based applications. However, scientists and medical researchers are still searching for a simple cloud-based architecture that enables secure collaboration and sharing of distributed datasets. To date, attempts at using the cloud for this purpose generally view the cloud as simply a pool of servers upon which to run their legacy software. This approach fails to leverage the unique platform capabilities of the cloud. In this paper, we describe our Cloud Distributed Research Network (CloudDRN). We leverage the cloud for availability, reliability, scalability, and improved security as compared to legacy distributed systems while still supporting site autonomy. Our philosophy is to adapt commercial software tooling that was originally designed for business use-cases, thereby benefiting from the large built-in user community. We describe our general architecture and show an example of our system created to share distributed clinical research data. We evaluate our system in Amazon Web Services (AWS) and in Microsoft Windows Azure and find that while each cloud achieves similar financial cost, representative queries are 3.5x slower on average in Windows Azure. Marty Humphrey, Jacob Steele, In Kee Kim, Michael G. Kahn, Jessica Bondy, Michael Ames |
e-Science | 3 |
| 2006 | Resource Demand Prediction-Based Grid Resource Transaction Network Model in Grid Computing Environment
In Kee Kim, Jong Sik Lee |
ICCSA (5) | 1 |