EDBT 2026 Demo / reviewers in the wild / expert
Kwangsung Oh
dblp:149/7532
· DBLP profile ↗
10ranked-venue papers
6as first author
3since 2021 · last 2023
0000-0003-3281-7325ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 39% Cloud and datacenter computing · 25% Parallel and multicore computing · 25% | |
| Computer networks
2 papers |
Edge and fog computing · 62% Network optimization and economics · 38% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › distributed storage
geo-distributed storage |
0.7 | 2 | 2020 | Wiera: Policy-Driven Multi-Tiered Geo-Distributed Cloud Storage System · IEEE Trans. Parallel Distributed Syst. 2020 Wiera: Towards Flexible Multi-Tiered Geo-Distributed Cloud Storage Instances · HPDC 2016 |
Storage systems › storage hierarchy
tiered storage |
0.7 | 2 | 2020 | Wiera: Policy-Driven Multi-Tiered Geo-Distributed Cloud Storage System · IEEE Trans. Parallel Distributed Syst. 2020 Wiera: Towards Flexible Multi-Tiered Geo-Distributed Cloud Storage Instances · HPDC 2016 |
Cloud and datacenter computing › geo-distributed cloud
geo-distributed data analytics |
0.6 | 1 | 2022 | Network Cost-Aware Geo-Distributed Data Analytics System · IEEE Trans. Parallel Distributed Syst. 2022 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.6 | 1 | 2022 | Network Cost-Aware Geo-Distributed Data Analytics System · IEEE Trans. Parallel Distributed Syst. 2022 |
Parallel and multicore computing
task scheduling |
0.6 | 1 | 2022 | Network Cost-Aware Geo-Distributed Data Analytics System · IEEE Trans. Parallel Distributed Syst. 2022 |
Storage systems
data placement |
0.4 | 1 | 2020 | Wiera: Policy-Driven Multi-Tiered Geo-Distributed Cloud Storage System · IEEE Trans. Parallel Distributed Syst. 2020 |
Cloud and datacenter computing › cloud deployment
multi-cloud |
0.3 | 2 | 2022 | Network Cost-Aware Geo-Distributed Data Analytics System · IEEE Trans. Parallel Distributed Syst. 2022 Wiera: Policy-Driven Multi-Tiered Geo-Distributed Cloud Storage System · IEEE Trans. Parallel Distributed Syst. 2020 |
Edge and fog computing
edge cloud |
0.3 | 1 | 2017 | Nebula: Distributed Edge Cloud for Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2017 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.3 | 1 | 2017 | Nebula: Distributed Edge Cloud for Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2017 |
Interconnection networks and networks-on-chip
network bandwidth |
0.2 | 1 | 2022 | Network Cost-Aware Geo-Distributed Data Analytics System · IEEE Trans. Parallel Distributed Syst. 2022 |
Distributed systems
fault tolerance |
0.2 | 2 | 2017 | Nebula: Distributed Edge Cloud for Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2017 Wiera: Towards Flexible Multi-Tiered Geo-Distributed Cloud Storage Instances · HPDC 2016 |
Distributed systems › distributed system architecture
geo-distributed systems |
0.1 | 1 | 2020 | Wiera: Policy-Driven Multi-Tiered Geo-Distributed Cloud Storage System · IEEE Trans. Parallel Distributed Syst. 2020 |
Distributed systems › replication › replication and fault tolerance
replication and recovery |
0.1 | 1 | 2017 | Nebula: Distributed Edge Cloud for Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2017 |
Methods — techniques the papers use, named apart from their topics
mapreduce models · 1.1cost-aware scheduling · 1.1mapreduce · 0.6policy specification · 0.4optimal data placement · 0.4dynamic policy specification · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Smartpick: Workload Prediction for Serverless-enabled Scalable Data Analytics SystemsabstractMany data analytic systems have adopted a newly emerging compute resource, serverless (SL), to handle data analytics queries in a timely and cost-efficient manner, i.e., serverless data analytics. While these systems can start processing queries quickly thanks to the agility and scalability of SL, they may encounter performance-and cost-bottlenecks based on workloads due to SL's worse performance and more expensive cost than traditional compute resources, e.g., virtual machine (VM). In this paper, we introduce Smartpick, a SL-enabled scalable data analytics system that exploits SL and VM together to realize composite benefits, i.e., agility from SL and better performance with reduced cost from VM. Smartpick uses a machine learning prediction scheme, decision-tree based Random Forest with Bayesian Optimizer, to determine SL and VM configurations, i.e., how many SL and VM instances for queries, that meet cost-performance goals. Smartpick offers a knob for applications to allow them to explore a richer cost-performance tradeoff space opened by exploiting SL and VM together. To maximize the benefits of SL, Smartpick supports a simple but strong mechanism, called relay-instances. Smartpick also supports event-driven prediction model retraining to deal with workload dynamics. A Smartpick prototype was implemented on Spark and deployed on live testbeds, Amazon AWS and Google Cloud Platform. Evaluation results indicate 97.05% and 83.49% prediction accuracies respectively with up to 50% cost reduction as opposed to the baselines. The results also confirm that Smartpick allows data analytics applications to navigate the richer cost-performance tradeoff space efficiently and to handle workload dynamics effectively and automatically. Anshuman Das Mohapatra, Kwangsung Oh |
Middleware | 2 |
| 2022 | Network Cost-Aware Geo-Distributed Data Analytics SystemabstractMany geo-distributed data analytics (GDA) systems have focused on the network performance-bottleneck: inter-data center network bandwidth to improve performance. Unfortunately, these systems may encounter acost-bottleneck(${\$}$) because they have not considered data transfer cost (${\$}$), one of the most expensive and heterogeneous resources in a multi-cloud environment. In this article, we presentKimchi, a network cost-aware GDA system to meet the cost-performance tradeoff by exploiting data transfer cost heterogeneity to avoid the cost-bottleneck. Kimchi determines cost-aware task placement decisions for scheduling tasks given inputs including data transfer cost, network bandwidth, input data size and locations, and desired cost-performance tradeoff preference. In addition, Kimchi is also mindful of data transfer cost in the presence of dynamics. Kimchi has been applied to two common GDA MapReduce models: synchronous barrier and asynchronous push-based shuffle. A Kimchi prototype has been implemented on Spark, and experiments show that it reduces cost by 5%$\scriptstyle \sim$24% without impacting performance and reduces query execution time by 45%$\scriptstyle \sim$70% without impacting cost compared to other baseline approaches centralized, vanilla Spark, and bandwidth-aware (e.g., Iridium). More importantly, Kimchi allows applications to explore a much richer cost-performance tradeoff space in a multi-cloud environment. Kwangsung Oh, Minmin Zhang, Abhishek Chandra, Jon B. Weissman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Cocoa: Towards a Scalable Compute Cost-aware Data Analytics SystemabstractRecently, many data analytics systems have focused on adopting a newly emerging compute resource, serverless, which offers scalability and agility to deal with peak workloads in a timely and cost-efficient manner, i.e., serverless data analytics (SDA). Unfortunately, these systems may encounter a cost bottleneck ($) because they have ignored the per unit time cost ($) of serverless, which is more expensive by up to 5.8 times for the same compute capacity than a traditional compute resource such as a virtual machine (VM). In addition, SDA may also encounter a performance bottleneck due to serverless' worse performance than VM. In this paper, we first study and report when serverless is beneficial for data analytics. Then, we present a scalable compute cost-aware data analytics system, Cocoa, that exploits serverless and VM together to achieve composite benefits. A Cocoa prototype was implemented on Spark. Evaluation results show a richer cost-performance tradeoff space opened by exploiting heterogeneous compute resources together, and identify substantial opportunities for future serverless-enabled systems research. Kwangsung Oh, Myoungkyu Song |
IC2E | 1 |
| 2020 | A Network Cost-aware Geo-distributed Data Analytics SystemabstractMany geo-distributed data analytics (GDA) systems have focused on the network performance-bottleneck: interdata center network bandwidth to improve performance. Unfortunately, these systems may encounter a cost-bottleneck ($) because they have not considered data transfer cost ($), one of the most expensive and heterogeneous resources in a multi-cloud environment. In this paper, we present Kimchi, a network cost-aware GDA system to meet the cost-performance tradeoff by exploiting data transfer cost heterogeneity to avoid the cost-bottleneck. Kimchi determines cost-aware task placement decisions for scheduling tasks given inputs including data transfer cost, network bandwidth, input data size and locations, and desired cost-performance tradeoff preference. In addition, Kim- chi is also mindful of data transfer cost in the presence of dynamics. A Kimchi prototype has been implemented on Spark and experiments show that it reduces cost by 14% ~ 24% without impacting performance and reduces query execution time by 45% ~ 70% without impacting cost compared to other baseline approaches centralized, vanilla Spark, and bandwidth-aware (e.g. Iridium). More importantly, Kimchi allows applications to explore a much richer cost-performance tradeoff space in a multi-cloud environment. Kwangsung Oh, Abhishek Chandra, Jon B. Weissman |
CCGRID | 1 |
| 2020 | Code Inspection Support for Recurring Changes with Deep Learning in Evolving SoftwareabstractDevelopers often make recurring changes, similar but different changes across multiple locations. They inspect such code changes per source file (i.e., a diff patch) during code reviews; however, diff patches represent low-level code modification without summarizing recurring changes, leading to tedious and error-prone code inspection. To address this problem, we propose a novel code review approach, Recurring Code Changes Inspection with Deep Learning (RIDL) that leverages change patterns of an edit script by learning code clones, identical or nearly similar code fragments. To train a classifier, RIDL learns 13,940 clones with four different clone types (e.g., Type-1, Type-2, Type-3, and Type-4 clones) from a clone database mined from 25,000 subject programs. Our approach then leverages the classifier to (1) interactively summarize recurring changes and (2) detect change mistakes, potential anomalies in a given codebase. In the evaluation, after 2 hours of training, RIDL analyzes code changes in four open source projects. It summarizes recurring changes with 95.1% accuracy and detects change anomalies with 93.1% accuracy. Our results show that RIDL should help developers effectively inspect recurring changes during code reviews. Krishna Teja Ayinala, Kwok Sun Cheng, Kwangsung Oh, Teukseob Song, Myoungkyu Song |
COMPSAC | 3 |
| 2020 | Wiera: Policy-Driven Multi-Tiered Geo-Distributed Cloud Storage SystemabstractMulti-tiered geo-distributed cloud storage systems must tame complexity at many levels: uniform APIs for storage access, supporting flexible storage policies that meet a wide array of application metrics, determining an optimal data placement, handling uncertain network dynamics and access dynamism, and operating across many levels of heterogeneity both within and across data-centers (DCs). In this paper, we present an integrated solution called Wiera. Wiera enables the specification of data management policies both within a local DC and across DCs. Such policies enable the user to optimize for cost, performance, reliability, durability, and consistency, and to express their tradeoffs. In addition, Wiera determines an optimal data placement for the user to meet their desired tradeoffs easily in such an environment. A key aspect of Wiera is first-class support for dynamism due to network, workload, and access patterns changes. As far as we know, Wiera is the first geo-distributed cloud storage system which handles dynamism actively at run-time. Wiera allowsunmodified applicationsto reap the benefits of flexible data/storage policies by externalizing the policy specification. We show how Wiera enables a rich specification of dynamic policies using a concise notation and describe the design and implementation of the system. We have implemented a Wiera prototype on multiple cloud environments, AWS and Azure, that illustrates potential benefits from managing dynamics and in using multiple cloud storage tiers both within and across DCs. Kwangsung Oh, Nan Qin, Abhishek Chandra, Jon B. Weissman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | TripS: automated multi-tiered data placement in a geo-distributed cloud environmentabstractExploiting the cloud storage hierarchy both within and across data-centers of different cloud providers empowers Internet applications to choose data centers (DCs) and storage services based on storage needs. However, using multiple storage services across multiple data centers brings a complex data placement problem that depends on a large number of factors including, e.g., desired goals, storage and network characteristics, and pricing policies. In addition, dynamics e.g., changing user locations and access patterns, make it impossible to determine the best data placement statically. In this paper, we present TripS, a lightweight system that considers both data center locations and storage tiers to determine the data placement for geo-distributed storage systems. Such systems make use of TripS by providing inputs including SLA, consistency model, fault tolerance, latency information, and cost information. With given inputs, TripS models and solves the data placement problem using mixed integer linear programming (MILP) to determine data placement. In addition, to adapt quickly to dynamics, we introduce the notion of Target Locale List (TLL), a pro-active approach to avoid expensive re-evaluation of the optimal placement. The TripS prototype is running on Wiera, a policy driven geo-distributed storage system, to show how a storage system can easily utilize TripS for data placement. We evaluate TripS/Wiera on multiple data centers of AWS and Azure. The results show that TripS/Wiera can reduce cost 14.96% ∼ 98.1% based on workloads in comparison with other works' approaches and can handle both short- and long-term dynamics to avoid SLA violations. Kwangsung Oh, Abhishek Chandra, Jon B. Weissman |
SYSTOR | 1 |
| 2017 | Nebula: Distributed Edge Cloud for Data Intensive ComputingabstractCentralized cloud infrastructures have become the popular platforms for data-intensive computing today. However, they suffer from inefficient data mobility due to the centralization of cloud resources, and hence, are highly unsuited for geo-distributed data-intensive applications where the data may be spread at multiple geographical locations. In this paper, we present Nebula: a dispersed edge cloud infrastructure that explores the use of voluntary resources for both computation and data storage. We describe the lightweight Nebula architecture that enables distributed data-intensive computing through a number of optimization techniques including location-aware data and computation placement, replication, and recovery. We evaluate Nebula performance on an emulated volunteer platform that spans over 50 PlanetLab nodes distributed across Europe, and show how a common data-intensive computing framework, MapReduce, can be easily deployed and run on Nebula. We show Nebula MapReduce is robust to a wide array of failures and substantially outperforms other wide-area versions based on emulated existing systems. Albert Jonathan, Mathew Ryden, Kwangsung Oh, Abhishek Chandra, Jon B. Weissman |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Wiera: Towards Flexible Multi-Tiered Geo-Distributed Cloud Storage InstancesabstractGeo-distributed cloud storage systems must tame complexity at many levels: uniform APIs for storage access, supporting flexible storage policies that meet a wide array of application metrics, handling uncertain network dynamics and access dynamism, and operating across many levels of heterogeneity both within and across data-centers. In this paper, we present an integrated solution called Wiera. Wiera extends our earlier cloud storage system, Tiera, that is targeted to multi-tiered policy-based single cloud storage, to the wide-area and multiple data-centers (even across different providers). Wiera enables the specification of global data management policies built on top of local Tiera policies. Such policies enable the user to optimize for cost, performance, reliability, durability, and consistency, both within and across data-centers, and to express their tradeoffs. A key aspect of Wiera is first-class support for dynamism due to network, workload, and access patterns changes. Wiera policies can adapt to changes in user workload, poorly performing data tiers, failures, and changes in user metrics (e.g., cost). Wiera allows unmodified applications to reap the benefits of flexible data/storage policies by externalizing the policy specification. As far as we know, Wiera is the first geo-distributed cloud storage system which handles dynamism actively at run-time. We show how Wiera enables a rich specification of dynamic policies using a concise notation and describe the design and implementation of the system. We have implemented a Wiera prototype on multiple cloud environments, AWS and Azure, that illustrates potential benefits from managing dynamics and in using multiple cloud storage tiers both within and across data-centers. Kwangsung Oh, Abhishek Chandra, Jon B. Weissman |
HPDC | 1 |
| 2014 | Nebula: Distributed Edge Cloud for Data Intensive ComputingabstractCentralized cloud infrastructures have become the de-facto platform for data-intensive computing today. However, they suffer from inefficient data mobility due to the centralization of cloud resources, and hence, are highly unsuited for dispersed-data-intensive applications, where the data may be spread at multiple geographical locations. In this paper, we present Nebula: a dispersed cloud infrastructure that uses voluntary edge resources for both computation and data storage. We describe the lightweight Nebula architecture that enables distributed data-intensive computing through a number of optimizations including location-aware data and computation placement, replication, and recovery. We evaluate Nebula's performance on an emulated volunteer platform that spans over 50 PlanetLab nodes distributed across Europe, and show how a common data-intensive computing framework, MapReduce, can be easily deployed and run on Nebula. We show Nebula MapReduce is robust to a wide array of failures and substantially outperforms other wide-area versions based on a BOINC like model. Mathew Ryden, Kwangsung Oh, Abhishek Chandra, Jon B. Weissman |
IC2E | 2 |