Byung-Chul Tak

dblp:10/6711 · also Byungchul Tak · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-8204-6816ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ForumSeeker: Fusion Retrieval of Online Technical Forums for Effective Troubleshooting
Youyang Kim, Yaoping Ruan, Young-Kyoon Suh, Liqiang Wang 0001, Byung-Chul Tak
FASE5
2026 Measuring Intrinsic Dimension of Multiobjective Landscapes and Dimensionality-Reduced Neuroevolution of Deep Reinforcement Learning
Oladayo S. Ajani, Jeonggeun Kim, Jong Taek Lee, Byung-Chul Tak
GECCO4
2026 Falconf: Configuration Error Diagnosis via Log Sequence Learning and Automated Misconfiguration Injections
Youyang Kim, Sahil Suneja, Yunja Choi, Young-Woo Kwon 0001, Byung-Chul Tak
ICDCS5
2026 Batcher: Learning to Construct Cost-Efficient Batches of Small Queries in Big Data Processing Platforms
Yeonsu Park 0001, Taesung Lee, Byung-Chul Tak, Wook-Shin Han
ICDE3
2026 A two-stage evolutionary framework with structural-parametric decoupling for sparse large-scale multi-objective optimization
Oladayo S. Ajani, Jeonggeun Kim, Jong Taek Lee, Byung-Chul Tak
Expert Syst. Appl.4
2026 Performance analysis of microVMs and containers for edge computing: A focus on file and network I/O
Kyungwoon Lee, Yunha Choi, Byung-Chul Tak
Future Gener. Comput. Syst.3
2026 TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
Taesung Lee, Jaehyun Ha, Byung-Chul Tak, Wook-Shin Han
Proc. VLDB Endow.3
2025 PROBA: Enhancing Serverless Edge Computing via Adaptive Task Scheduling and Probabilistic Resource Sharing
abstract
Serverless edge computing improves performance by processing data closer to its source, reducing operational costs, and increasing server utilization. Despite these benefits, edge servers face scalability challenges and queuing delays due to limited resources. Horizontal offloading can alleviate excessive workloads by efficiently distributing tasks across edge servers. However, it introduces higher waiting time, cold starts latency, and missed task deadlines at the receiving edge servers. To address these challenges, we introduce PROBA, which utilizes a Double Dueling Deep Q-learning algorithm and In-node scheduling that optimize the task offloading between the edge servers and improve task scheduling within the edge server. The approach uses probabilistic resource sharing, where edge servers share their real-time availability to a central cloud system. The cloud analyzes these performance metrics based on user-specified rewards to determine optimal scheduling decisions, which the edge servers execute to maintain balanced and responsive work-load distribution. We evaluated PROBA in a serverless edge computing simulator that focuses on horizontal offloading and in-node scheduling. In the evaluation using real-world trace data from Alibaba, our PROBA technique decreased the average wait time from 3.1 s to below 0.69$s$. PROBA also gave 1.37 % better task completion time than competitors.
Byung-Chul Tak, Young-Woo Kwon 0001
CLOUD2
2025 A Resource Provisioning Framework with Adaptive Task Distribution for Edge Devices
abstract
Edge devices such as wearables, drones, and CCTV systems have been widely deployed to collect real-world data, playing a crucial role in enhancing and securing urban life. However, these devices often struggle with significant performance challenges due to their limited computational and storage capacities when processing data locally. Offloading computation and data to the public cloud is straightforward but introduces high costs and latency. Alternatively, relying on an edge server to support a diverse array of heterogeneous edge devices can standardize operations. Still, it may lead to underutilization of high-performance devices such as Jetson Xavier if all tasks are centralized on the server. To address these concerns, we introduce ERPF, an edge resource provisioner that virtually extends edge devices' computation and storage capabilities, enabling them to handle complex tasks beyond their capacity. ERPF supports dynamic volume provisioning, GPU provisioning, and online execution context migration. Also, we propose a novel technique (ATS) that schedules AI workloads on distributed edge devices and edge servers with adjustment of task partition sizes based on the computational and network performance of the edge devices. ATS is seamlessly integrated into the ERPF prototype, which is implemented on a Kubernetes cluster using the Rook-Ceph storage orchestrator. Experimental results show that ERPF efficiently scales resources for edge devices through strategic offloading, while ATS delivers a substantial performance improvement of up to 23 × compared to baseline methods.
Youngwoo Jang, Soonbeom Kwon, Illyoung Choi, Dukyun Nam, Byung-Chul Tak, Young-Kyoon Suh
NOMS5
2024 POSTER: Seccomp profiling with Dynamic Analysis via ChatGPT-assisted Test Code Generation
abstract
The effectiveness of Seccomp kernel feature depends on how tightly and accurately the necessary system calls are specified in the seccomp policy. Static code analysis may miss out or over-approximate required system calls. With dynamic analysis, it is difficult to cover all possible execution paths. In this work, we aim to advance the state-of-the-art dynamic analysis approach by enabling it to increase the coverage of the target application's functionalities. Our approach takes as input the application's online documentation and leverages ChatGPT to generate a large number of test codes for functionalities in the documentation. This automated process eliminates the barrier to manually writing a large number of test codes for conducting dynamic analysis. Through our preliminary evaluation, we confirmed that ChatGPT can be used effectively to automatically generate a large number of test codes. Also, we observed early evidence that the seccomp policy generated from running the test codes could be more sound than the ones generated by static analysis.
Somin Song, Ashish Kundu, Byung-Chul Tak
AsiaCCS3
2024 K-RAF: A Kubernetes-based Resource Augmentation Framework for Edge Devices
abstract
Internet of Things (IoT) (or edge) devices are typically resource-constrained in terms of CPU, memory, and storage. Thus, it is viable for the devices to request resource provisioning to an edge server in the presence of growing data and heavy computation, as the edge server provides better accessibility than cloud servers. Consequently, the edge devices often perform computation and storage provisioning to the edge servers in large-scale data operations. However, the conventional methods for provisioning edge devices take into little consideration the characteristics of resources that jobs executed at the devices rely on. In particular, fully migrating computation jobs from the device to the server may waste valuable resources of the server without considering the computation and I/O characteristics of the jobs, thereby making the devices' resources idle. To overcome these limitations, we propose a novel Kubernetes-based resource augmentation framework, termed K-RAF, for provisioning edge devices with limited capabilities and accelerating the devices' job processing. Our experiment demonstrates that utilizing GPU acceleration, on average, K-RAF can run tasks 306 times faster than local computation on an edge device. Also, we show that utilizing the task distribution between an edge device and K-RAF can offer an average speedup of about 40% compared to K-RAF alone.
Youngwoo Jang, Jiseob Byun, Soonbeom Kwon, Illyoung Choi, Dukyun Nam, Byung-Chul Tak, Gap-Joo Na, Young-Kyoon Suh
HPDC6
2024 Distributed Page Table: Harnessing Physical Memory as an Unbounded Hashed Page Table
abstract
Virtual memory systems rely on the page table, a crucial component that maps virtual addresses to physical addresses (i.e., address translation). While the Radix Page Table (RPT) has traditionally been used for this task, its limitations have become more apparent with the rise of memory-intensive applications. Recently, Hashed Page Tables (HPTs) have been explored as an alternative page table structure to offer faster address translation. However, the HPT introduces its own set of challenges particularly in resizing the page table and allocating contiguous physical memory space for storing the table. To tackle the fundamental problem of the existing HPT designs, this paper introduces Distributed Page Table (DPT), a novel approach that utilizes the physical memory as a huge hashed page table. DPT distributes Page Table Entries (PTEs) across the entire physical memory space, significantly reducing the hash collisions while avoiding the table resizing overheads. When distributing the PTEs across the physical memory, they can be mapped to memory locations already allocated to data pages. This new type of collision, referred to as address collision, may reduce the effectiveness of the DPT. This paper showcases that the DPT can effectively resolve the address collision with three simple yet efficient techniques: Strided Open Addressing (SOA), Collision-Aware Virtual Address Allocation (CVA) and Collided Page Displacement (CPD). Our experimental results demonstrate that DPT achieves average performance improvements of 12.6%, 11.6%, and 8.7% compared to traditional RPT, the latest large-coverage TLB design, and state-of-the-art HPTs, respectively.
Osang Kwon, Junhyeok Park 0001, Sungbin Jang, Byung-Chul Tak, Seokin Hong
MICRO5
2024 ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework
abstract
The schemalessness, one of the major advantages of JSON representation format, comes with high penalties in querying and operations by denying various critical functions such as query optimizations, indexing, or data verification. There have been continuous efforts to develop an accurate JSON schema discovery algorithm from a bag of JSON documents. Unfortunately, existing schema discovery techniques, being top-down algorithms, face challenges from the lack of visibility into children nodes of JSON tree. With absence of the information about lower-level JSON elements, top-down algorithms need to employ assumptions and heuristics to decide the schema type of nodes. However, such static decisions are often violated in datasets which causes top-down algorithms to perform poorly. To overcome this, we propose an algorithm, called ReCG, that processes JSON documents in a bottom-up manner. It builds up schemas from leaf elements upward in the JSON document tree and, thus, can make more informed decisions of the schema node types. In addition, we adopt MDL (Minimum Description Length) principles systematically while building up the schemas to choose among candidate schemas the most concise yet accurate one with well-balanced generality. Evaluations show that our technique improves the recall and precision of found schemas by as high as 47%, resulting in 46% better F1 score while also performing 2.11× faster on average against the state-of-the-art.
Joohyung Yun, Byung-Chul Tak, Wook-Shin Han
Proc. VLDB Endow.2
2023 On the Value of Sequence-Based System Call Filtering for Container Security
abstract
One critical attack that exploits kernel vulnerabilities through system call invocations is considered a serious threat to container security since it results in the privilege escalation followed by the infamous container escape. The seccomp kernel feature provides the first line of defense against it. Further, secure container runtimes such as gVisor also make use of it to strengthen security. However, it is known to be brittle since it operates at the granularity of the individual system call. Inadvertent filtering of necessary system calls may inhibit the correct execution while overly generous rules allow the attacks. We believe that, by looking at the sequence of system calls, we can achieve more accurate and effective blocking of attacks in containers. To this end, we built a software tool, Nimos, that performs a combination of static and dynamic analyses of exploit codes in an automated way and investigated the existence of such commonly occurring system call sequences. Then, we analyzed the expected defensive power from applying the sequence-based filtering mechanisms using a large set of collected kernel vulnerabilities to assess the feasibility. We found that there exist a significant number and forms of commonly appearing system call sequences that can be used as a clear signature of the class of attacks. We characterize these common system call sequences that exist among the exploit codes and evaluate the expected effectiveness of a sequence-based system call filtering mechanism for containers.
Somin Song, Sahil Suneja, Michael V. Le, Byung-Chul Tak
CLOUD4
2023 MicroVM on Edge: Is It Ready for Prime Time?
abstract
Container virtualization is recognized as indispensable for realizing the vision of edge computing due to its advantages. However, OS-level virtualization suffers from a relatively low degree of security. Recently, microVM technology has emerged in response to this deficiency to provide stronger isolation and security while delivering performance comparable to the containers. In this work, we aim to gain a better understanding of microVM's suitability for edge computing in comparison with containers. We conduct extensive experiments on diverse workloads to test how microVMs compare against containers in several aspects. Through rigorous measurements and analysis, we extract several important findings. Despite having a more complex architecture than containers, microVMs perform comparably to the containers in terms of I/o performance. MicroVMs can even outperform containers in certain I/O workload types by 69%. Network I/O performance of microVMs can be 3x better than containers. We provide our findings and insights on the performance characteristics of microVMs on edge.
Kyungwoon Lee, Byung-Chul Tak
MASCOTS2
2023 A Close Look at Shared Resource Consumption in NoSQL Databases for Accurate Accounting
abstract
A NoSQL database plays an essential role in all parts of today’s multi-tenant large-scale IT services. Accurate information about per-tenant resource accounting is invaluable for optimal service management. However, there is a lack of study on the characteristics of resource consumption by multi-tenants in NoSQL database services, particularly for those resources consumed as part of activities asynchronously triggered by requests accumulated over time from multiple tenants. According to our investigation, this shared resource usage takes up a significant portion of total resource consumption in NoSQL databases.We assert that an accurate understanding of the shared resource consumption pattern is required to design correct and effective techniques for resource management. To this end, we conduct a detailed investigation of the shared resource consumption in popular NoSQL databases. Our focuses are to find out what portion of the overall resource consumption occurs in a shared manner, what type of operations cause them, and what effect the workload intensity has on the amount of shared resource consumption. We have developed a set of techniques for monitoring, recording, and analyzing the resource consumption data toward answering these questions. Our investigation revealed that the shared resource consumption can be as large as 30% of the total resource consumption, implying that resource accounting based only on directly causal events, may lead to underestimation of the true amount. We believe our study highlights the importance of shared resource accounting and provides crucial insights for building accurate resource accounting techniques.
Jaeryun Lee, Byung-Chul Tak, Euiseong Seo
NOMS2
2023 ECM: An Energy-efficient HVAC Control Framework for Stable Construction Environment
abstract
A cargo containment system (CCS) of liquefied natural gas (LNG) is an essential component of an LNG carrier (LNGC). During the manufacturing process of the LNGC CCS, it is critical that the heating, ventilation, and air conditioning (HVAC) facility stabilizes the environmental states inside the CCS at all times to prevent devastating rust and dew from forming inside the LNGC CCS. One critical problem is that it consumes enormous power, resulting in high expenses. To alleviate this problem, we propose our design of a novel data-driven framework, termed ECM, that uses a combination of machine learning and deep reinforcement learning (DRL) models to robustly and automatically control the HVAC system. Based on selected features, we develop the best indoor-environment forecasting model from several candidate models and build an HVAC control agent by training the DRL model with the reward function that uses the predicted temperature and humidity through the forecasting model. To validate our proposed framework, we have assessed the performance of our models on the real-world sensor data obtained from one of the major world-class shipyards. As a result, we show that our DRL-based model trained in the proposed framework stably controls the temperature inside the CCS within only 1.$5^{\mathrm{o}}$C variance in the set range from 2$3^{\mathrm{o}}$C to 2$5^{\mathrm{o}}$C while on average consuming power up to about 34% less than the compared existing methods. We expect our framework will bring an annual savings of about ${\$}$ 14 million or more once deployed in the actual field.
Jin-Sung Ok, Youngeun Chae, Harin Seo, Soon-Do Kwon, Byung-Chul Tak, Young-Kyoon Suh
SECON5
2023 QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark
abstract
Spark big data processing platform is heavily used in today's IT services for various critical applications such as machine learning tasks for service recommendations or massive volumes of raw sales data analysis. Spark is designed to deliver high performance by enabling a high degree of parallelism while processing various heavy-weight queries that require homogeneous operations on large data. However, it has been observed that workloads made of small and short-running queries coming from various sources are becoming dominant in practice. Unfortunately, the current Spark architecture is unfit to process workloads made of a large number of small queries optimally due to excessive I/Os with small computations. We present a technique, called QaaD, that addresses this problem fundamentally by applying i) transparent conversion of workloads made of small queries into one with large queries and ii) dynamic partition size adjustment for runtime overhead minimization. For this, we introduce a new abstraction, microRDD, to support our design of query merging, the embedding of queries as part of data, and an opportunistic sharing of common input data among queries. Comprehensive evaluation using real-world data shows that QaaD is able to deliver 10.6x to 36.6x speed-up against standard Spark executions for small query workloads.
Yeonsu Park 0001, Byung-Chul Tak, Wook-Shin Han
Proc. ACM Manag. Data2
2022 SecQuant: Quantifying Container System Call Exposure
Sunwoo Jang, Somin Song, Byung-Chul Tak, Sahil Suneja, Michael V. Le, Chuan Yue, Dan Williams 0001
ESORICS (2)3
2022 NoSQL Database Performance Diagnosis through System Call-level Introspection
abstract
Since its emergence, NoSQL databases have firmly established themselves as an indispensable software component of modern cloud-native applications. However, it also becomes increasingly challenging to perform critical management tasks such as troubleshooting unexpected performance problems. This is due to the ever-increasing diversity and specialization of NoSQL databases that make it difficult to observe the internal activities. To address these challenges, we have designed and built a technique for introspecting NoSQL databases. Our technique traces system call sequences of key operations under controlled workloads and filters scaling patterns from constant components. Novel algorithms are developed to uncover repeating patterns of system calls from massive amounts of traces and filter out background noise with high efficiency. The evaluation shows that our technique can greatly enhance the visibility into the NoSQL databases enabling us to diagnose performance problems or gain insights into internal activities.
Changho Seo, Yunchang Chae, Jaeryun Lee, Euiseong Seo, Byung-Chul Tak
NOMS5
2022 Privacy-Aware Collaborative Task Offloading in Fog Computing
abstract
Numerous new applications have been proliferated with the mature of 5G, which generates a large number of latency-sensitive and computationally intensive mobile data requests. The real-time requirement of these mobile data has been accommodated well by fog computing in the past few years, mainly through offloading tasks to fog nodes in the vicinity. On the other hand, the user-privacy hidden in the Internet-of-Things (IoT) data has not been sufficiently considered in the presence of insecure fog nodes. It is risky to offload an entire mission-critical task to just one fog node or several fog nodes owned by the same service provider (SP), especially when the SP is marked with low-security credit and tends to collect data information of users for malicious use. To address this issue, we classify IoT user tasks based on their security requirements, divide them into different numbers of smaller fragments, and, finally, offload the segments of a task to multiple fog nodes owned by the same or various SPs according to their security requirements. The selected fog nodes will collaboratively serve the divided fragments to avoid the possible damage caused by the leak of sensitive data due to compromised fog nodes of malicious SPs. For this, we propose an integer linear programming (ILP) model and a dynamic programming algorithm to maximize the number of successfully served IoT data tasks with satisfactory security requirements while minimizing the end-to-end transmission delay. The numerical results show that the proposed ILP model and algorithm can significantly increase the successful provisioning ratio for tasks with high-security requirements.
Mian Muaz Razaq, Byung-Chul Tak, Limei Peng, Mohsen Guizani
IEEE Trans. Comput. Soc. Syst.2
2022 BlackEye: automatic IP blacklisting using machine learning from security logs
Dooyong Jeon, Byung-Chul Tak
Wirel. Networks2
2021 Lognroll: discovering accurate log templates by iterative filtering
abstract
Modern IT systems rely heavily on log analytics for critical operational tasks. Since the volume of logs produced from numerous distributed components is overwhelming, it requires us to employ automated processing. The first step of automated log processing is to convert streams of log lines into the sequence of log format IDs, called log templates. A log template serves as a base string with unfilled parts from which logs are generated during runtime by substitution of contextual information. The problem of log template discovery from the volume of collected logs poses a great challenge due to the semi-structured nature of the logs and the computational overheads. Our investigation reveals that existing techniques show various limitations. We approach the log template discovery problem as search-based learning by applying the ILP (Inductive Logic Programming) framework. The algorithm core consists of narrowing down the logs into smaller sets by analyzing value compositions on selected log column positions. Our evaluation shows that it produces accurate log templates from diverse application logs with small computational costs compared to existing methods. With the quality metric we defined, we obtained about 21%-51% improvements of log template quality.
Byung-Chul Tak, Wook-Shin Han
Middleware1
2019 LADRA: Log-based abnormal task detection and root-cause analysis in big data processing with Spark
Siyang Lu, Wei Xiang 0007, BingBing Rao, Byung-Chul Tak, Long Wang 0003, Liqiang Wang 0001
Future Gener. Comput. Syst.4
2017 Log-based Abnormal Task Detection and Root Cause Analysis for Spark
abstract
Application delays caused by abnormal tasks arecommon problems in big data computing frameworks. Anabnormal task in Spark, which may run slowly withouterror or warning logs, not only reduces its resident node'sperformance, but also affects other nodes' efficiency.Spark log files report neither root causes of abnormal tasks,nor where and when abnormal scenarios happen. AlthoughSpark provides a “speculation” mechanism to detect stragglertasks, it can only detect tailed stragglers in each stage. Sincethe root causes of abnormal happening are complicated, thereare no effective ways to detect root causes.This paper proposes an approach to detect abnormality andanalyzes root causes using Spark log files. Unlike commononline monitoring or analysis tools, our approach is a pureoff-line method that can analyze abnormality accurately. Ourapproach consists of four steps. First, a parser preprocessesraw log files to generate structured log data. Second, ineach stage of Spark application, we choose features relatedto execution time and data locality of each task, as well asmemory usage and garbage collection of each node. Third,based on the selected features, we detect where and whenabnormalities happen. Finally, we analyze the problems usingweighted factors to decide the probability of root causes. In thispaper, we consider four potential root causes of abnormalities,which include CPU, memory, network, and disk. The proposedmethod has been tested on real-world Spark benchmarks.To simulate various scenario of root causes, we conductedinterference injections related to CPU, memory, network,and Disk. Our experimental results show that the proposedapproach is accurate on detecting abnormal tasks as well asfinding the root causes
Siyang Lu, BingBing Rao, Wei Xiang 0007, Byung-Chul Tak, Long Wang 0003, Liqiang Wang 0001
ICWS4
2017 Understanding Security Implications of Using Containers in the Cloud
Byung-Chul Tak, Canturk Isci, Sastry S. Duri, Nilton Bila, Shripad Nadgowda, James Doran
USENIX ATC1
2017 Failure Diagnosis for Distributed Systems Using Targeted Fault Injection
abstract
This paper introduces a novel approach to automating failure diagnostics in distributed systems by combining fault injection and data analytics. We use fault injection to populate the database of failures for a target distributed system. When a failure is reported from production environment, the database is queried to find “matched” failures generated by fault injections. Relying on the assumption that similar faults generate similar failures, we use information from the matched failures as hints to locate the actual root cause of the reported failures. In order to implement this approach, we introduce techniques for (i) reconstructing end-to-end execution flows of distributed software components, (ii) computing the similarity of the reconstructed flows, and (iii) performing precise fault injection at pre-specified executing points in distributed systems. We have evaluated our approach using an OpenStack cloud platform, a popular cloud infrastructure management system. Our experimental results showed that this approach is effective in determining the root causes, e.g., fault types and affected components, for 71-100 percent of tested failures. Furthermore, it can provide fault locations close to actual ones and can easily be used to find and fix actual root causes. We have also validated this technique by localizing real bugs that occurred in OpenStack.
Cuong Pham 0003, Long Wang 0003, Byung-Chul Tak, Salman Baset, Chunqiang Tang, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Parallel Distributed Syst.3
2017 Resource Accounting of Shared IT Resources in Multi-Tenant Clouds
abstract
In today's IT platforms, the capability to accurately account overall resource usage among applications is crucial for variety of management actions (e.g., capacity planning, dynamic resource reallocation and/or load balancing). However, in the environments where small number of shared services cater to a large number of distinct entities' requests, resource accounting becomes significantly challenging. First, the overall resource consumption at the shared service is the aggregate of the resource consumption for multiple remote entities whose identities are not visible to the shared service. Second, even if such information becomes available, common monitoring tools (e.g., top, iostat) are unable to deliver accurate break-down of resource consumption since sharing occurs at sub-instance level (i.e., service instances are not exclusive). We study inherent challenges of performing resource accounting of shared resource. We compare two nonintrusive approaches having different balance between local monitoring and collective inference - (i) LR that uses easily-available tools which provide aggregate measurement and applying well-known linear regression as inference, and (ii) Rameter that puts more emphasis on gathering fine-grained per-thread information from within the hypervisor and applying light inference on the data. Evaluation shows that Rameter offers less than 1% error in accounting whereas LR's error fluctuates between 5-150%.
Byung-Chul Tak, Youngjin Kwon, Bhuvan Urgaonkar
IEEE Trans. Serv. Comput.1
2016 Auto-tuning Performance of MPI Parallel Programs Using Resource Management in Container-Based Virtual Cloud
abstract
Load imbalance problem is one of the major obstacles to achieving optimal performance of High Performance Computing applications. The approach of trying to distribute the problem pieces to each node with the hope of balancing execution time has limits since the performance depends not only on data size but also on many other dynamic factors. This paper describes an approach that uses adaptive resource management enabled by the container-based virtualization to solve the load imbalance problem of MPI programs running in the cloud. Our techniques dynamically adjust CPU resource allocation to MPI processes running as container instances according to the current program execution state and system resource status. The resource allocation among MPI processes is adjusted in two ways: the intra-host level, which dynamically adjusts resources within a host, and the inter-host level, which migrates containers together with MPI processes from one host to another host. We have implemented and evaluated our approach on Amazon EC2 platform using real-world scientific benchmarks and applications, which demonstrates that the performance can be improved up to 31% (with an average of 15%) when compared with the baseline.
Hongyi Ma, Liqiang Wang 0001, Byung-Chul Tak, Long Wang 0003, Chunqiang Tang
CLOUD3
2016 LOGAN: Problem Diagnosis in the Cloud Using Log-Based Reference Models
abstract
Problem diagnosis is one crucial aspect in the cloud operation that is becoming increasingly challenging. On the one hand, the volume of logs generated in today's cloud is overwhelmingly large. On the other hand, cloud architecture becomes more distributed and complex, which makes it more difficult to troubleshoot failures. In order to address these challenges, we have developed a tool, called LOGAN, that enables operators to quickly identify the log entries that potentially lead to the root cause of a problem. It constructs behavioral reference models from logs that represent the normal patterns. When problem occurs, our tool enables operators to inspect the divergence of current logs from the reference model and highlight logs likely to contain the hints to the root cause. To support these capabilities we have designed and developed several mechanisms. First, we developed log correlation algorithms using various IDs embedded in logs to help identify and isolate log entries that belong to the failed request. Second, we provide efficient log comparison to help understand the differences between different executions. Finally we designed mechanisms to highlight critical log entries that are likely to contain information pertaining to the root cause of the problem. We have implemented the proposed approach in a popular cloud management system, OpenStack, and through case studies, we demonstrate this tool can help operators perform problem diagnosis quickly and effectively.
Byung-Chul Tak, Shu Tao, Yaoping Ruan
IC2E1
2015 SafeSky: A Secure Cloud Storage Middleware for End-User Applications
abstract
As the popularity of cloud storage services grows rapidly, it is desirable and even essential for both legacy and new end-user applications to have the cloud storage capability to improve their functionality, usability, and accessibility. However, incorporating the cloud storage capability into applications must be done in a secure manner to ensure the confidentiality, integrity, and availability of users' data in the cloud. Unfortunately, it is non-trivial for ordinary application developers to either enhance legacy applications or build new applications to properly have the secure cloud storage capability, due to the development efforts involved as well as the security knowledge and skills required. In this paper, we propose SafeSky, a middleware that can immediately enable an application to use the cloud storage services securely and efficiently, without any code modification or recompilation. A SafeSky-enabled application does not need to save a user's data to the local disk, but instead securely saves them to different cloud storage services to significantly enhance the data security. We have implemented SafeSky as a shared library on Linux. SafeSky supports applications written in different languages, supports various popular cloud storage services, and supports common user authentication methods used by those services. Our evaluation and analysis of SafeSky with real-world applications demonstrate that SafeSky is a feasible and practical approach for equipping end-user applications with the secure cloud storage capability.
Rui Zhao 0005, Chuan Yue, Byung-Chul Tak, Chunqiang Tang
SRDS3
2014 AppCloak: Rapid Migration of Legacy Applications into Cloud
abstract
Although cloud has been adopted by many organizations as their main infrastructure for IT delivery, there are still a large number of legacy applications running in non-cloud hosting environments. Thus, it is crucial to have migration techniques for such legacy applications so that they can benefit from many advantages of cloud such as elasticity, low upfront investment, and fast time-to-market. However, migrating large number of legacy applications into cloud in a timely manner is a daunting task. Common techniques such as redeveloping (i.e., modernizing) them or reinstalling from the scratch entails high costs. To mitigate these problems, we have developed a rapid migration technique, called AppCloak, that allows users to literally copy an already-installed application to cloud and run it without any modifications. The technique is based on intercepting a selected set of system calls and replacing the parameters and return values to hide any differences of environments to the application. We demonstrate that our technique works in Amazon EC2 and quantify the performance overhead.
Byung-Chul Tak, Chunqiang Tang
IEEE CLOUD1
2013 CAP3: A Cloud Auto-Provisioning Framework for Parallel Processing Using On-Demand and Spot Instances
abstract
Cloud computing has drawn increasing attention from the scientific computing community due to its ease of use, elasticity, and relatively low cost. Because a high-performance computing (HPC) application is usually resource demanding, without careful planning, it can incur a high monetary expense even in Cloud. We design a tool called CAP3 (Cloud Auto-Provisioning framework for Parallel Processing) to help a user minimize the expense of running an HPC application in Cloud, while meeting the user-specified job deadline. Given an HPC application, CAP3 automatically profiles the application, builds a model to predict its performance, and infers a proper cluster size that can finish the job within its deadline while minimizing the total cost. To further reduce the cost, CAP3 intelligently chooses the Cloud's reliable on-demand instances or low-cost spot instances, depending on whether the remaining time is tight in meeting the application's deadline. Experiments on Amazon EC2 show that the execution strategy given by CAP3 is cost-effective, by choosing a proper cluster size and a proper instance type (on-demand or spot).
Liqiang Wang 0001, Byung-Chul Tak, Long Wang 0003, Chunqiang Tang
IEEE CLOUD3
2013 PseudoApp: Performance prediction for application migration to cloud
Byung-Chul Tak, Chunqiang Tang, Long Wang 0003
IM1
2013 Cloudy with a Chance of Cost Savings
abstract
Cloud-based hosting is claimed to possess many advantages over traditional in-house (on-premise) hosting such as better scalability, ease of management, and cost savings. It is not difficult to understand how cloud-based hosting can be used to address some of the existing limitations and extend the capabilities of many types of applications. However, one of the most important questions is whether cloud-based hosting will be economically feasible for my application if migrated into the cloud. It is not straightforward to answer this question because it is not clear how my application will benefit from the claimed advantages, and, in turn, be able to convert them into tangible cost savings. Within cloud-based hosting offerings, there is a wide range of hosting options one can choose from, each impacting the cost in a different way. Answering these questions requires an in-depth understanding of the cost implications of all the possible choices specific to my circumstances. In this study, we identify a diverse set of key factors affecting the costs of deployment choices. Using benchmarks representing two different applications (TPC-W and TPC-E) we investigate the evolution of costs for different deployment choices. We consider important application characteristics such as workload intensity, growth rate, traffic size, storage, and software license to understand their impact on the overall costs. We also discuss the impact of workload variance and cloud elasticity, and certain cost factors that are subjective in nature.
Byung-Chul Tak, Bhuvan Urgaonkar, Anand Sivasubramaniam
IEEE Trans. Parallel Distributed Syst.1
2011 Reducing the Delay and Power Consumption of Web Browsing on Smartphones in 3G Networks
abstract
Smart phone is becoming a key element in providing greater user access to the mobile Internet. Many complex applications, which are used to be only on PCs, have been developed and run on smart phones. These applications extend the functionalities of smart phones and make them more convenient for users to be connected. However, they also greatly increase the power consumption of smart phones and many users are frustrated with the long delay of web browsing when using smart phones. In this paper, we have discovered that the key reason of the long delay and high power consumption in web browsing is not due to the bandwidth limitation most of time in 3G networks. The local computation limitation at the smart phone is the real bottleneck for opening most web pages. To address this issue, we propose an architecture, called Virtual-Machine based Proxy (VMP), to shift the computing from smart phones to the VMP. To illustrate the feasibility of deploying the proposed VMP system in 3G networks, we have built a prototype using Xen virtual machines and Android Phones with T-Mobile UMTS network. Experimental results show that compared to normal smart phone browser, our VMP approach reduces the delay by more than 80% and reduces the power consumption during web browsing by more than 45%.
Bo Zhao 0009, Byung-Chul Tak, Guohong Cao
ICDCS2
2011 A dynamic energy management scheme for multi-tier data centers
abstract
Multi-tier data centers have become a norm for hosting modern Internet applications because they provide a flexible, modular, scalable and high performance environment. However, these benefits come at a price of the economic dent incurred in powering and cooling these large hosting centers. Thus, energy efficiency has become a critical consideration in designing Internet data centers. In this paper, we propose a multifaceted approach, Hybrid, consisting of dynamic provisioning, frequency scaling and dynamic power management (DPM) schemes to reduce the energy consumption of multi-tier data centers, while meeting the Service Level Agreements (SLAs). We formulate a mathematical model of the energy and performance/SLA optimization problem followed by a queueing theory based approach to develop two heuristics for solving the optimization problem. The first heuristic dynamically provisions the optimal number of servers required in each tier. The second heuristic proactively decides the CPU speed and the duration of sleep states of a server to achieve further energy savings. We evaluate our heuristics using a simulator that was validated with real measurements on a prototype three-tier data center consisting of 25 servers with two multi-tier application benchmarks. Our experimental results indicate that the proposed scheme, Hybrid, can reduce the energy consumption by 50% relative to static provisioning without CPU frequency scaling and DPM. We demonstrate that Hybrid satisfies the SLAs for dynamically varying workloads. In addition, the proposed multifaceted approach is more energy efficient than the other methods such as dynamic provisioning with exploiting deep sleep states.
Seung-Hwan Lim, Bikash Sharma, Byung-Chul Tak, Chita R. Das
ISPASS3
2009 vPath: Precise Discovery of Request Processing Paths from Black-Box Observations of Thread and Network Activities
Byung-Chul Tak, Chunqiang Tang, Sriram Govindan, Bhuvan Urgaonkar, Rong Chang 0001
USENIX ATC1
2005 Rotational Lease: Providing High Availability in a Shared Storage File System
Byung-Chul Tak, Yon Dohn Chung, Sunja Kim, Myoung-Ho Kim
HPCC1