EDBT 2026 Demo / reviewers in the wild / expert
Byung-Gon Chun
dblp:34/3515
· DBLP profile ↗
60ranked-venue papers
13as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 5 first-author · 7 since 2021Computer networks · 15 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Garen: Reliable Cluster Management with Atomic State ReconciliationabstractModern cluster managers orchestrate large-scale services and resources through a set of controllers, each managing a specific part of the cluster by iteratively reconciling the cluster states into the desired states. However, controllers are prone to various state inconsistencies stemming from asynchrony, concurrency, and failures, posing significant challenges in reliable cluster operation. Our analysis of 51 consistency bugs in Kubernetes controllers reveals that many issues remain unresolved, and even proposed fixes are often rejected for backward incompatibility or performance loss. Ahnjae Shin, Jaewoo Maeng, Myeongjae Jeon, Byung-Gon Chun |
EuroSys | 5 |
| 2024 | Blaze: Holistic Caching for Iterative Data ProcessingabstractModern data processing workloads, such as machine learning and graph processing, involve iterative computations to converge generated models into higher accuracy. An effective caching mechanism is vital to expedite iterative computations since the intermediate data that needs to be stored in memory grows larger over iterations, often exceeding the memory capacity. However, existing systems handle intermediate data through separate operational layers (e.g., caching, eviction, and recovery), with each layer working independently in a greedy or cost-agnostic manner. These layers typically rely on user annotations and past access patterns, failing to make globally optimal decisions for the workload. Won Wook Song, Jeongyoon Eo, Taegeon Um, Myeongjae Jeon, Byung-Gon Chun |
EuroSys | 5 |
| 2024 | Maestro: The Analysis-Simulation Integrated Framework for Mixed RealityabstractMixed reality devices with near-eye displays unlock new possibilities for innovation and user experiences. Mixed reality applications require a new unified framework that enables seamless analysis of the real world and simulation of realistic virtual content. Designing such a framework faces various challenges, including huge programming efforts of analysis and simulation pipelines, and inconsistencies between real-world and virtual content caused by end-to-end processes across pipelines. Jingyu Lee, Minjae Kim 0001, Byung-Gon Chun, Youngki Lee 0001 |
MobiSys | 4 |
| 2024 | Poster: Maestro: The Analysis-Simulation Integrated Framework for Mixed RealityabstractThe recent development of DNN and hardware has created new opportunities for mixed-reality applications. These applications demand the ability to analyze the real world and simulate realistic virtual content. However, designing mixed-reality applications faces diverse challenges due to the absence of a unified framework, such as huge programming effort and inconsistencies between the real scene and virtual content induced by end-to-end latency. Jingyu Lee, Minjae Kim 0001, Byung-Gon Chun, Youngki Lee 0001 |
MobiSys | 4 |
| 2023 | FlowKV: A Semantic-Aware Store for Large-Scale State Management of Stream Processing EnginesabstractWe propose FlowKV, a persistent store tailored for large-scale state management of streaming applications. Unlike existing KV stores, FlowKV leverages information from stream processing engines by taking a principled approach toward exploiting information about how and when the applications access data. FlowKV categorizes data access patterns of window operations according to how window boundaries are set and how tuples inside a window are aggregated, and deploys customized in-memory and on-disk data structures optimized for each pattern. In addition, FlowKV takes window metadata as explicit arguments of read and write methods to predict the moment when a window is read, and then loads the tuples of windows in batches from storage ahead of time. Using the NEXMark benchmark as workload, our experiments show that Apache Flink on FlowKV outperforms Flink on RocksDB or Faster with up to 4.12× throughput gain. Gyewon Lee, Jaewoo Maeng, Jinsol Park, Jangho Seo, Haeyoon Cho 0001, Youngseok Yang, Taegeon Um, Jongsung Lee 0001, Jae W. Lee, Byung-Gon Chun |
EuroSys | 10 |
| 2023 | BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language ModelsabstractPipeline parallelism is a key technique for training large language models within GPU clusters. However, it often leads to a memory imbalance problem, where certain GPUs face high memory pressure while others underutilize their capacity. This imbalance results in suboptimal training performance, even when the overall GPU memory capacity is sufficient for more efficient setups. To address this inefficiency, we propose BPipe, a novel approach for achieving memory balance in pipeline parallelism. BPipe employs an activation balancing method to transfer intermediate activations between GPUs during training, enabling all GPUs to utilize comparable amounts of memory. With balanced memory utilization, BPipe enhances the training efficiency of large language models like GPT-3 by eliminating redundant recomputations or increasing the micro-batch size. Our evaluation conducted on 48 A100 GPUs across six nodes interconnected with HDR InfiniBand shows that BPipe accelerates the training of GPT-3 96B and GPT-3 134B models by 1.25x-2.17x compared to Megatron-LM, a state-of-the-art framework for training large language models. Taebum Kim, Gyeong-In Yu, Byung-Gon Chun |
ICML | 4 |
| 2023 | Sponge: Fast Reactive Scaling for Stream Processing with Serverless Frameworks
Won Wook Song, Taegeon Um, Sameh Elnikety, Myeongjae Jeon, Byung-Gon Chun |
USENIX ATC | 5 |
| 2022 | OpenNetLab: Open Platform for RL-based Congestion Control for Real-Time CommunicationsabstractWith the growing importance of real-time communications (RTC), designing congestion control (CC) algorithms for RTC that achieve high network performance and QoE is gaining attention. Recently, data-driven, reinforcement learning (RL)-based CC algorithms for RTC have shown great potential, outperforming traditional rule-based counterparts. However, there are no open platforms tailored for training, evaluation, and validation of the algorithms that can facilitate this emerging research area. Jeongyoon Eo, Zhixiong Niu, Wenxue Cheng, Francis Y. Yan, Jorina Kardhashi, Scott Inglis, Michael Revow, Byung-Gon Chun, Peng Cheng 0005, Yongqiang Xiong |
APNet | 9 |
| 2022 | SUMNAS: Supernet with Unbiased Meta-Features for Neural Architecture Search
Hyeonmin Ha, Jihoon Kim 0002, Semin Park, Byung-Gon Chun |
ICLR | 4 |
| 2022 | Band: coordinated multi-DNN inference on heterogeneous mobile processorsabstractThe rapid development of deep learning algorithms, as well as innovative hardware advancements, encourages multi-DNN workloads such as augmented reality applications. However, existing mobile inference frameworks like TensorFlow Lite and MNN fail to efficiently utilize heterogeneous processors available on mobile platforms, because they focus on running a single DNN on a specific processor. As mobile processors are too resource-limited to deliver reasonable performance for such workloads by their own, it is challenging to serve multi-DNN workloads with existing frameworks. Joo Seong Jeong, Jingyu Lee, Changmin Jeon, Changjin Jeong, Youngki Lee 0001, Byung-Gon Chun |
MobiSys | 7 |
| 2022 | Orca: A Distributed Serving System for Transformer-Based Generative Models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun |
OSDI | 5 |
| 2022 | Hippo: Sharing Computations in Hyper-Parameter OptimizationabstractHyper-parameter optimization is crucial for pushing the accuracy of a deep learning model to its limits. However, a hyper-parameter optimization job, referred to as a study, involves numerous trials of training a model using different training knobs, and therefore is very computation-heavy, typically taking hours and days to finish. We observe that trials issued from hyper-parameter optimization algorithms often share common hyper-parameter sequence prefixes. Based on this observation, we propose Hippo, a hyper-parameter optimization system that reuses computation across trials to reduce the overall amount of computation significantly. Instead of treating each trial independently as in existing hyper-parameter optimization systems, Hippo breaks down the hyper-parameter sequences into stages and merges common stages to form a tree of stages (a stage tree). Hippo maintains an internal data structure, search plan, to manage the current status and history of a study, and employs a critical path based scheduler to minimize the overall study completion time. Hippo applies to not only single studies but multi-study scenarios as well. Evaluations show that Hippo's stage-based execution strategy outperforms trial-based methods for several models and hyper-parameter optimization algorithms, reducing end-to-end training time by up to 2.76X (3.53x) and GPU-hours by up to 4.81X (6.77x), for single (multiple) studies. Ahnjae Shin, Joo Seong Jeong, Do Yoon Kim, Soyoung Jung, Byung-Gon Chun |
Proc. VLDB Endow. | 5 |
| 2021 | SCOPA: Soft Code-Switching and Pairwise Alignment for Zero-Shot Cross-lingual TransferabstractThe recent advent of cross-lingual embeddings, such as multilingual BERT (mBERT), provides a strong baseline for zero-shot cross-lingual transfer. There also exists increasing research attention to reduce the alignment discrepancy of cross-lingual embeddings between source and target languages, via generating code-switched sentences by substituting randomly selected words in the source languages with their counterparts of the target languages. Although these approaches improve the performance, naively code-switched sentences can have inherent limitations. In this paper, we propose SCOPA, a novel technique to improve the performance of zero-shot cross-lingual transfer. Instead of using the embeddings of code-switched sentences directly, SCOPA mixes them softly with the embeddings of original sentences. In addition, SCOPA utilizes an additional pairwise alignment objective, which aligns the vector differences of word pairs instead of word-level embeddings, in order to transfer contextualized information between different languages while preserving language-specific information. Experiments on the PAWS-X and MLDoc dataset show the effectiveness of SCOPA. Dohyeon Lee, Jaeseong Lee 0002, Gyewon Lee, Byung-Gon Chun, Seung-won Hwang |
CIKM | 4 |
| 2021 | Harmony: A Scheduling Framework Optimized for Multiple Distributed Machine Learning JobsabstractWe introduce Harmony, a new scheduling framework that executes multiple Parameter-Server ML training jobs together to improve cluster resource utilization. Harmony coordinates a fine-grained execution of co-located jobs with complementary resource usages to avoid contention and to efficiently share resources between the jobs. To resolve the memory pressure due to the increased number of simultaneous jobs, Harmony uses a data spill/reload mechanism optimized for multiple jobs with the iterative execution pattern. Our evaluation shows that Harmony improves cluster resource utilization by up to 1.65×, resulting in a reduction of the mean ML training job time by about 53%, and makespan, the total time to process all given jobs, by about 38%, compared to the traditional approaches that allocate dedicated resources to each job. Woo-Yeon Lee, Yunseong Lee, Won Wook Song, Youngseok Yang, Byung-Gon Chun |
ICDCS | 6 |
| 2021 | Pluto: High-Performance IoT-Aware Stream ProcessingabstractNowadays, large numbers of small IoT stream queries are created from diverse IoT applications and executed on cloud backend servers. However, existing distributed stream processing systems such as Storm and Flink do not efficiently handle the large numbers of IoT stream queries because of their tightly-coupled query/code submission layer and inefficient query execution layer. In this paper, we propose Pluto, a new IoT-aware stream processing system. As a first step for IoT stream processing, this paper focuses on optimizing the execution of many IoT stream queries on a node. Pluto optimizes the end-to-end query processing with a three-phase execution, harnessing IoT-query characteristics. First, Pluto minimizes bottlenecks in the IoT query submission by decoupling the code registration from the query submission process with new APIs, which eliminates duplicate code registration and enables code sharing across queries. Second, in the execution phase, Pluto shares system resources as much as possible and minimizes resource bottlenecks in a machine by exploiting commonalities among IoT stream queries and information exposed in the API. Our evaluations show that Pluto improves the throughput by an order of magnitude compared to other stream processing systems on a 24-core machine, keeping P99 latency less than one second. Taegeon Um, Gyewon Lee, Byung-Gon Chun |
ICDCS | 3 |
| 2021 | Terra: Imperative-Symbolic Co-Execution of Imperative Deep Learning ProgramsabstractImperative programming allows users to implement their deep neural networks (DNNs) easily and has become an essential part of recent deep learning (DL) frameworks. Recently, several systems have been proposed to combine the usability of imperative programming with the optimized performance of symbolic graph execution. Such systems convert imperative Python DL programs to optimized symbolic graphs and execute them. However, they cannot fully support the usability of imperative programming. For example, if an imperative DL program contains a Python feature with no corresponding symbolic representation (e.g., third-party library calls or unsupported dynamic control flows) they fail to execute the program. To overcome this limitation, we propose Terra, an imperative-symbolic co-execution system that can handle any imperative DL programs while achieving the optimized performance of symbolic graph execution. To achieve this, Terra builds a symbolic graph by decoupling DL operations from Python features. Then, Terra conducts the imperative execution to support all Python features, while delegating the decoupled operations to the symbolic execution. We evaluated Terra’s performance improvement and coverage with ten imperative DL programs for several DNN architectures. The results show that Terra can speed up the execution of all ten imperative DL programs, whereas AutoGraph, one of the state-of-the-art systems, fails to execute five of them. Taebum Kim, Eunji Jeong, Geon-Woo Kim, Yunmo Koo, Sehoon Kim 0001, Gyeong-In Yu, Byung-Gon Chun |
NeurIPS | 7 |
| 2021 | Finding Consensus Bugs in Ethereum via Multi-transaction Differential Fuzzing
Youngseok Yang, Taesoo Kim, Byung-Gon Chun |
OSDI | 3 |
| 2021 | Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network Training
Gyewon Lee, Irene Lee, Hyeonmin Ha, Kyung-Geun Lee, Hwarim Hyun, Ahnjae Shin, Byung-Gon Chun |
USENIX ATC | 7 |
| 2021 | WindTunnel: Towards Differentiable ML Pipelines Beyond a Single ModeleabstractWhile deep neural networks (DNNs) have shown to be successful in several domains like computer vision, non-DNN models such as linear models and gradient boosting trees are still considered state-of-the-art over tabular data. When using these models, data scientists often author machine learning (ML) pipelines: DAG of ML operators comprising data transforms and ML models, whereby each operator is sequentially trained one-at-a-time. Conversely, when training DNNs, layers composing the neural networks are simultaneously trained using backpropagation. In this paper, we argue that the training scheme of ML pipelines is sub-optimal because it tries to optimize a single operator at a time thus losing the chance of global optimization. We therefore propose WindTunnel: a system that translates a trained ML pipeline into a pipeline of neural network modules and jointly optimizes the modules using backpropagation. We also suggest translation methodologies for several non-differentiable operators such as gradient boosting trees and categorical feature encoders. Our experiments show that fine-tuning of the translated WindTunnel pipelines is a promising technique able to increase the final accuracy. Gyeong-In Yu, Saeed Amizadeh, Sehoon Kim 0001, Artidoro Pagnoni, Ce Zhang 0001, Byung-Gon Chun, Markus Weimer, Matteo Interlandi |
Proc. VLDB Endow. | 6 |
| 2020 | Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningabstractDeep learning (DL) frameworks take advantage of GPUs to improve the speed of DL inference and training. Ideally, DL frameworks should be able to fully utilize the computation power of GPUs such that the running time depends on the amount of computation assigned to GPUs. Yet, we observe that in scheduling GPU tasks, existing DL frameworks suffer from inefficiencies such as large scheduling overhead and unnecessary serial execution. To this end, we propose Nimble, a DL execution engine that runs GPU tasks in parallel with minimal scheduling overhead. Nimble introduces a novel technique called ahead-of-time (AoT) scheduling. Here, the scheduling procedure finishes before executing the GPU kernel, thereby removing most of the scheduling overhead during run time. Furthermore, Nimble automatically parallelizes the execution of GPU tasks by exploiting multiple GPU streams in a single GPU. Evaluation on a variety of neural networks shows that compared to PyTorch, Nimble speeds up inference and training by up to 22.34× and 3.61×, respectively. Moreover, Nimble outperforms state-of-the-art inference systems, TensorRT and TVM, by up to 2.81× and 1.70×, respectively. Woosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon Chun |
NeurIPS | 4 |
| 2020 | Apache Nemo: A Framework for Optimizing Distributed Data ProcessingabstractOptimizing scheduling and communication of distributed data processing for resource and data characteristics is crucial for achieving high performance. Existing approaches to such optimizations largely fall into two categories. First, distributed runtimes provide low-level policy interfaces to apply the optimizations, but do not ensure the maintenance of correct application semantics and thus often require significant effort to use. Second, policy interfaces that extend a high-level application programming model ensure correctness, but do not provide sufficient fine control. We describe Apache Nemo, an optimization framework for distributed dataflow processing that provides fine control for high performance and also ensures correctness for ease of use. We combine several techniques to achieve this, including an intermediate representation of dataflow, compiler optimization passes, and runtime extensions. Our evaluation results show that Nemo enables composable and reusable optimizations that bring performance improvements on par with existing specialized runtimes tailored for a specific deployment scenario. Apache Nemo is open-sourced at https://nemo.apache.org as an Apache incubator project. Won Wook Song, Youngseok Yang, Jeongyoon Eo, Jangho Seo, Sanha Lee, Gyewon Lee, Taegeon Um, Haeyoon Cho 0001, Byung-Gon Chun |
ACM Trans. Comput. Syst. | 10 |
| 2019 | Parallax: Sparsity-aware Data Parallel Training of Deep Neural NetworksabstractThe employment of high-performance servers and GPU accelerators for training deep neural network models have greatly accelerated recent advances in deep learning (DL). DL frameworks, such as TensorFlow, MXNet, and Caffe2, have emerged to assist DL researchers to train their models in a distributed manner. Although current DL frameworks scale well for image classification models, there remain opportunities for scalable distributed training on natural language processing (NLP) models. We found that current frameworks show relatively low scalability on training NLP models due to the lack of consideration to the difference in sparsity of model parameters. In this paper, we propose Parallax, a framework that optimizes data parallel training by utilizing the sparsity of model parameters. Parallax introduces a hybrid approach that combines Parameter Server and AllReduce architectures to optimize the amount of data transfer according to the sparsity. Experiments show that Parallax built atop Tensor-Flow achieves scalable training throughput on both dense and sparse models while requiring little effort from its users. Parallax achieves up to 2.8x, 6.02x speedup for NLP models than TensorFlow and Horovod with 48 GPUs, respectively. The training speed for the image classification models is equal to Horovod and 1.53x faster than TensorFlow. Soojeong Kim, Gyeong-In Yu, Hojin Park, Sungwoo Cho, Eunji Jeong, Hyeonmin Ha, Sanha Lee, Joo Seong Jeong, Byung-Gon Chun |
EuroSys | 9 |
| 2019 | Automating System Configuration of Distributed Machine LearningabstractThe performance of distributed machine learning systems is dependent on their system configuration. However, configuring the system for optimal performance is challenging and time consuming even for experts due to the diverse runtime factors such as workloads or the system environment. We present cost-based optimization to automatically find a good system configuration for parameter server (PS) machine learning (ML) frameworks. We design and implement Cruise that applies the optimization technique to tune distributed PS ML execution automatically. Evaluation results on three ML applications verify that Cruise automates the system configuration of the applications to achieve good performance with minor reconfiguration costs. Woo-Yeon Lee, Markus Weimer, Byung-Gon Chun, Yunseong Lee, Joo Seong Jeong, Gyeong-In Yu, Hojin Park, Beomyeol Jeon, Won Wook Song, Gunhee Kim |
ICDCS | 4 |
| 2019 | JANUS: Fast and Flexible Deep Learning via Symbolic Graph Execution of Imperative Programs
Eunji Jeong, Sungwoo Cho, Gyeong-In Yu, Joo Seong Jeong, Dongjin Shin, Byung-Gon Chun |
NSDI | 6 |
| 2019 | Apache Nemo: A Framework for Building Distributed Dataflow Optimization Policies
Youngseok Yang, Jeongyoon Eo, Geon-Woo Kim, Sanha Lee, Jangho Seo, Won Wook Song, Byung-Gon Chun |
USENIX ATC | 8 |
| 2018 | Improving the expressiveness of deep learning frameworks with recursionabstractRecursive neural networks have widely been used by researchers to handle applications with recursively or hierarchically structured data. However, embedded control flow deep learning frameworks such as TensorFlow, Theano, Caffe2, and MXNet fail to efficiently represent and execute such neural networks, due to lack of support for recursion. In this paper, we add recursion to the programming model of existing frameworks by complementing their design with recursive execution of dataflow graphs as well as additional APIs for recursive definitions. Unlike iterative implementations, which can only understand the topological index of each node in recursive data structures, our recursive implementation is able to exploit the recursive relationships between nodes for efficient execution based on parallel computation. We present an implementation on TensorFlow and evaluation results with various recursive neural network models, showing that our recursive implementation not only conveys the recursive nature of recursive neural networks better than other implementations, but also uses given resources more effectively to reduce training and inference time. Eunji Jeong, Joo Seong Jeong, Soojeong Kim, Gyeong-In Yu, Byung-Gon Chun |
EuroSys | 5 |
| 2018 | PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
Yunseong Lee, Alberto Scolari, Byung-Gon Chun, Marco D. Santambrogio, Markus Weimer, Matteo Interlandi |
OSDI | 3 |
| 2017 | Breaking Ad-hoc Runtime Integrity Protection Mechanisms in Android Financial AppsabstractTo protect customers' sensitive information, many mobile financial applications include steps to probe the runtime environment and abort their execution if the environment is deemed to have been tampered with. This paper investigates the security of such self-defense mechanisms used in 76 popular financial Android apps in the Republic of Korea. Our investigation found that existing tools fail to analyze these Android apps effectively because of their highly obfuscated code and complex, non-traditional control flows. We overcome this challenge by extracting a call graph with a self-defense mechanism, from a detailed runtime trace record of a target app's execution. To generate the call graph, we identify the causality between the system APIs (Android APIs and system calls) used to check device rooting and app integrity, and those used to stop an app's execution. Our analysis of 76 apps shows that we can pinpoint methods to bypass a self-defense mechanism using a causality graph in most cases. We successfully bypassed self-defense mechanisms in 67 out of 73 apps that check device rooting and 39 out of 44 apps that check app integrity. While analyzing the self-defense mechanisms, we found that many apps rely on third-party security libraries for their self-defense mechanisms. Thus we present in-depth studies of the top five security libraries. Our results demonstrate the necessity of a platform-level solution for integrity checks. Hyeonmin Ha, Seoyoon Choi, Jaeyeon Jung, Byung-Gon Chun |
AsiaCCS | 5 |
| 2017 | Pado: A Data Processing Engine for Harnessing Transient Resources in DatacentersabstractDatacenters are under-utilized, primarily due to unused resources on over-provisioned nodes of latency-critical jobs. Such idle resources can be used to run batch data analytic jobs to increase datacenter utilization, but these transient resources must be evicted whenever latency-critical jobs require them again. Resource evictions often lead to cascading recomputations, which is usually handled by checkpointing intermediate results on stable storages of eviction-free reserved resources. However, checkpointing has major shortcomings in its substantial overhead of transferring data back and forth. In this work, we step away from such approaches and focus on observing the job structure and the relationships between computations of the job. We carefully mark the computations that are most likely to cause a large number of recomputations upon evictions, to run them reliably using reserved resources. This lets us retain corresponding intermediate results effortlessly without any additional checkpointing. We design Pado, a general data processing engine, which carries out our idea with several optimizations that minimize the number of additional reserved nodes. Evaluation results show that Pado outperforms Spark 2.0.0 by up to 5.1×, and checkpoint-enabled Spark by up to 3.8×. Youngseok Yang, Geon-Woo Kim, Won Wook Song, Yunseong Lee, Zhengping Qian, Byung-Gon Chun |
EuroSys | 8 |
| 2017 | Apache REEF: Retainable Evaluator Execution FrameworkabstractResource Managers like YARN and Mesos have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault tolerance, task scheduling and coordination) and reimplement common mechanisms (e.g., caching, bulk-data transfers). This article presents REEF, a development framework that provides a control plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource reuse for data caching and state management abstractions that greatly ease the development of elastic data processing pipelines on cloud platforms that support a Resource Manager service. We illustrate the power of REEF by showing applications built atop: a distributed shell application, a machine-learning framework, a distributed in-memory caching system, and a port of the CORFU system. REEF is currently an Apache top-level project that has attracted contributors from several institutions and it is being used to develop several commercial offerings such as the Azure Stream Analytics service. Byung-Gon Chun, Tyson Condie, Yingda Chen, Carlo Curino, Chris Douglas, Matteo Interlandi, Beomyeol Jeon, Joo Seong Jeong, Gyewon Lee, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Mariia Mykhailova, Shravan M. Narayanamurthy, Joseph Noor, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Taegeon Um, Julia Wang, Markus Weimer, Youngseok Yang |
ACM Trans. Comput. Syst. | 1 |
| 2016 | Collaborative analytics for data silosabstractAs a great deal of data has been accumulated in various disciplines, the need for the integrative analysis of separate but relevant data sources is becoming more important. Combining data sources can provide global insight that is otherwise difficult to obtain from individual sources. Because of privacy, regulations, and other issues, many large-scale data repositories remain closed off from the outside, raising what has been termed the data silo issue. The huge volume of today's big data often leads to computational challenges, adding another layer of complexity to the solution. In this paper, we propose a novel method called collaborative analytics by ensemble learning (CABEL), which attempts to resolve the main hurdles regarding the silo issue: accuracy, privacy, and computational efficiency. CABEL represents the data stored in each silo as a compact aggregate of samples called the silo signature. The compact representation provides computational efficiency and privacy preservation but makes it challenging to produce accurate analytics. To resolve this challenge, we formulate the problem of attribute domain sampling and reconstruction, and propose a solution called the Chebyshev subset. To model collaborative efforts to analyze semantically linked but structurally disconnected databases, CABEL utilizes a new ensemble learning technique termed the weighted bagging of base classifiers. We demonstrate the effectiveness of CABEL by testing with a nationwide health-insurance data set containing approximately 4,182,000,000 records collected from the entire population of an Organisation for Economic Co-operation and Development (OECD) country in 2012. In our binary classification tests, CABEL achieved median recall, precision, and F-measure values of 89%, 64%, and 76%, respectively, although only 0.001–0.00001% of the original data was used for model construction, while maintaining data privacy and computational efficiency. Jinkyu Kim 0001, Heonseok Ha, Byung-Gon Chun, Sungroh Yoon, Sang Kyun Cha |
ICDE | 3 |
| 2015 | Elastic Memory: Bring Elasticity Back to In-Memory Big Data Analytics
Joo Seong Jeong, Woo-Yeon Lee, Yunseong Lee, Youngseok Yang, Byung-Gon Chun |
HotOS | 6 |
| 2015 | Making Sense of Performance in Data Analytics Frameworks
Kay Ousterhout, Ryan Rasti, Sylvia Ratnasamy, Scott Shenker, Byung-Gon Chun |
NSDI | 5 |
| 2015 | REEF: Retainable Evaluator Execution FrameworkabstractResource Managers like Apache YARN have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low-level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault-tolerance, task scheduling and coordination) and re-implement common mechanisms (e.g., caching, bulk-data transfers). This paper presents REEF, a development framework that provides a control-plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource re-use for data caching, and state management abstractions that greatly ease the development of elastic data processing work-flows on cloud platforms that support a Resource Manager service. REEF is being used to develop several commercial offerings such as the Azure Stream Analytics service. Furthermore, we demonstrate REEF development of a distributed shell application, a machine learning algorithm, and a port of the CORFU [4] system. REEF is also currently an Apache Incubator project that has attracted contributors from several instititutions. Markus Weimer, Yingda Chen, Byung-Gon Chun, Tyson Condie, Carlo Curino, Chris Douglas, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Shravan M. Narayanamurthy, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Julia Wang |
SIGMOD Conference | 3 |
| 2015 | Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile DevicesabstractWe present Mantis, a framework for predicting the computational resource consumption (CRC) of Android applications on given inputs accurately, and efficiently. A key insight underlying Mantis is that program codes often contain features that correlate with performance and these features can be automatically computed efficiently. Mantis synergistically combines techniques from program analysis and machine learning. It constructs concise CRC models by choosing from many program execution features only a handful that are most correlated with the program's CRC metric yet can be evaluated efficiently from the program's input. We apply program slicing to reduce evaluation time of a feature and automatically generate executable code snippets for efficiently evaluating features. Our evaluation shows that Mantis predicts four CRC metrics of seven Android apps with estimation error in the range of 0-11.1 percent by executing predictor code spending at most 1.3 percent of their execution time on Galaxy Nexus. Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
IEEE Trans. Mob. Comput. | 6 |
| 2014 | Collecting, organizing, and sharing pins in pinterest: interest-driven or social-driven?abstractPinterest, a popular social curating service where people collect, organize, and share content (pins in Pinterest), has gained great attention in recent years. Despite the increasing interest in Pinterest, little research has paid attention to how people collect, manage, and share pins in Pinterest. In this paper, to shed insight on such issues, we study the following questions. How do people collect and manage pins by their tastes in Pinterest? What factors do mainly drive people to share their pins in Pinterest? How do the characteristics of users (e.g., gender, popularity, country) or properties of pins (e.g., category, topic) play roles in propagating pins in Pinterest? To answer these questions, we have conducted a measurement study on patterns of pin curating and sharing in Pinterest. By keeping track of all the newly posted and shared pins in each category (e.g., animal, kids, women's fashion) from June 5 to July 18, 2013, we built 350 K pin propagation trees for 3 M users. With the dataset, we investigate: (1) how users collect and curate pins, (2) how users share their pins and why, and (3) how users are related by shared pins of interest. Our key finding is that pin propagation in Pinterest is mostly driven by pin's properties like its topic, not by user's characteristics like her number of followers. We further show that users in the same community in the interest graph (i.e., representing the relations among users) of Pinterest share pins (i) in the same category with 94% probability and (ii) of the same URL where pins come from with 89% probability. Finally, we explore the implications of our findings for predicting how pins are shared in Pinterest. Jinyoung Han, Daejin Choi, Byung-Gon Chun, Ted Taekyoung Kwon, Hyunchul Kim, Yanghee Choi |
SIGMETRICS | 3 |
| 2014 | TaintDroid: An Information-Flow Tracking System for Realtime Privacy Monitoring on SmartphonesabstractToday’s smartphone operating systems frequently fail to provide users with visibility into how third-party applications collect and share their private data. We address these shortcomings with TaintDroid, an efficient, system-wide dynamic taint tracking and analysis system capable of simultaneously tracking multiple sources of sensitive data. TaintDroid enables realtime analysis by leveraging Android’s virtualized execution environment. TaintDroid incurs only 32% performance overhead on a CPU-bound microbenchmark and imposes negligible overhead on interactive third-party applications. Using TaintDroid to monitor the behavior of 30 popular third-party Android applications, in our 2010 study we found 20 applications potentially misused users’ private information; so did a similar fraction of the tested applications in our 2012 study. Monitoring the flow of privacy-sensitive data with TaintDroid provides valuable input for smartphone users and security service firms seeking to identify misbehaving applications. William Enck, Peter Gilbert, Seungyeop Han, Vasant Tendulkar, Byung-Gon Chun, Landon P. Cox, Jaeyeon Jung, Patrick D. McDaniel, Anmol N. Sheth |
ACM Trans. Comput. Syst. | 5 |
| 2013 | Mantis: Automatic Performance Prediction for Smartphone Applications
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
USENIX ATC | 6 |
| 2013 | REEF: Retainable Evaluator Execution FrameworkabstractIn this demo proposal, we describe REEF, a framework that makes it easy to implement scalable, fault-tolerant runtime environments for a range of computational models. We will demonstrate diverse workloads, including extract-transform-load MapReduce jobs, iterative machine learning algorithms, and ad-hoc declarative query processing. At its core, REEF builds atop YARN (Apache Hadoop 2's resource manager) to provide retainable hardware resources with lifetimes that are decoupled from those of computational tasks. This allows us to build persistent (cross-job) caches and cluster-wide services, but, more importantly, supports high-performance iterative graph processing and machine learning algorithms. Unlike existing systems, REEF aims for composability of jobs across computational models, providing significant performance and usability gains, even with legacy code. REEF includes a library of interoperable data management primitives optimized for communication and data movement (which are distinct from storage locality). The library also allows REEF applications to access external services, such as user-facing relational databases. We were careful to decouple lower levels of REEF from the data models and semantics of systems built atop it. The result was two new standalone systems: Tang, a configuration manager and dependency injector, and Wake, a state-of-the-art event-driven programming and data movement framework. Both are language independent, allowing REEF to bridge the JVM and .NET. Byung-Gon Chun, Tyson Condie, Carlo Curino, Raghu Ramakrishnan 0001, Russell Sears, Markus Weimer |
Proc. VLDB Endow. | 1 |
| 2012 | Mobius: unified messaging and data serving for mobile appsabstractMobile application development is challenging for several reasons: intermittent and limited network connectivity, tight power constraints, server-side scalability concerns, and a number of fault-tolerance issues. Developers handcraft complex solutions that include client-side caching, conflict resolution, disconnection tolerance, and backend database sharding. To simplify mobile app development, we present Mobius, a system that addresses the messaging and data management challenges of mobile application development. Mobius introduces MUD (Messaging Unified with Data). MUD presents the programming abstraction of a logical table of data that spans devices and clouds. Applications using Mobius can asynchronously read from/write to MUD tables, and also receive notifications when tables change via continuous queries on the tables. The system combines dynamic client-side caching (with intelligent policies chosen on the server-side, based on usage patterns across multiple applications), notification services, flexible query processing, and a scalable and highly available cloud storage system. We present an initial prototype to demonstrate the feasibility of our design. Even in our initial prototype, remote read and write latency overhead is less than 52% when compared to a hand-tuned solution. Our dynamic caching reduces the number of messages by a factor of 4 to 8.5 when compared to fixed strategies, thus reducing latency, bandwidth, power, and server load costs, while also reducing data staleness. Byung-Gon Chun, Carlo Curino, Russell Sears, Alexander Shraer, Samuel Madden 0001, Raghu Ramakrishnan 0001 |
MobiSys | 1 |
| 2012 | MegaPipe: A New Programming Interface for Scalable Network I/O
Sangjin Han, Scott Marshall, Byung-Gon Chun, Sylvia Ratnasamy |
OSDI | 3 |
| 2011 | CloneCloud: elastic execution between mobile device and cloudabstractMobile applications are becoming increasingly ubiquitous and provide ever richer functionality on mobile devices. At the same time, such devices often enjoy strong connectivity with more powerful machines ranging from laptops and desktops to commercial clouds. This paper presents the design and implementation of CloneCloud, a system that automatically transforms mobile applications to benefit from the cloud. The system is a flexible application partitioner and execution runtime that enables unmodified mobile applications running in an application-level virtual machine to seamlessly off-load part of their execution from mobile devices onto device clones operating in a computational cloud. CloneCloud uses a combination of static analysis and dynamic profiling to partition applications automatically at a fine granularity while optimizing execution time and energy use for a target computation and communication environment. At runtime, the application partitioning is effected by migrating a thread from the mobile device at a chosen point to the clone in the cloud, executing there for the remainder of the partition, and re-integrating the migrated thread back to the mobile device. Our evaluation shows that CloneCloud can adapt application partitioning to different environments, and can help some applications achieve as much as a 20x execution speed-up and a 20-fold decrease of energy spent on the mobile device. Byung-Gon Chun, Sunghwan Ihm, Petros Maniatis, Mayur Naik, Ashwin Patti |
EuroSys | 1 |
| 2011 | Making Programs Forget: Enforcing Lifetime for Sensitive Data
Jayanthkumar Kannan, Byung-Gon Chun |
HotOS | 2 |
| 2010 | Predicting Execution Time of Computer Programs Using Sparse Polynomial RegressionabstractPredicting the execution time of computer programs is an important but challenging problem in the community of computer systems. Existing methods require experts to perform detailed analysis of program code in order to construct predictors or select important features. We recently developed a new system to automatically extract a large number of features from program execution on sample inputs, on which prediction models can be constructed without expert knowledge. In this paper we study the construction of predictive models for this problem. We propose the SPORE (Sparse POlynomial REgression) methodology to build accurate prediction models of program performance using feature data collected from program execution on sample inputs. Our two SPORE algorithms are able to build relationships between responses (e.g., the execution time of a computer program) and features, and select a few from hundreds of the retrieved features to construct an explicitly sparse and non-linear model to predict the response variable. The compact and explicitly polynomial form of the estimated model could reveal important insights into the computer program (e.g., features and their non-linear combinations that dominate the execution time), enabling a better understanding of the program’s behavior. Our evaluation on three widely used computer programs shows that SPORE methods can give accurate prediction with relative error less than 7% by using a moderate number of training data samples. In addition, we compare SPORE algorithms to state-of-the-art sparse regression algorithms, and show that SPORE methods, motivated by real applications, outperform the other methods in terms of both interpretability and prediction accuracy. Ling Huang 0001, Jinzhu Jia, Bin Yu 0001, Byung-Gon Chun, Petros Maniatis, Mayur Naik |
NIPS | 4 |
| 2010 | TaintDroid: An Information-Flow Tracking System for Realtime Privacy Monitoring on Smartphones
William Enck, Peter Gilbert, Byung-Gon Chun, Landon P. Cox, Jaeyeon Jung, Patrick D. McDaniel, Anmol Sheth |
OSDI | 3 |
| 2009 | Macroscope: end-point approach to networked application dependency discoveryabstractEnterprise and data center networks consist of a large number of complex networked applications and services that depend upon each other. For this reason, they are difficult to manage and diagnose. In this paper we propose Macroscope, a new approach to extracting the dependencies of networked applications automatically by combining application process information with network level packet traces. We evaluate Macroscope on traces collected at 52 laptops within a large enterprise and show that Macroscope is accurate in finding the dependencies of networked applications. We also show that Macroscope requires less human involvement and is significantly more accurate than state of the art approaches that use only packet traces. Using our rich profiles of the application-service dependencies, we explore and uncover some interesting characteristics about this relationship. Finally, we discuss several usage scenarios that can benefit from Macroscope. Lucian Popa 0002, Byung-Gon Chun, Ion Stoica, Jaideep Chandrashekar, Nina Taft |
CoNEXT | 2 |
| 2009 | Tiered Fault Tolerance for Long-Term Integrity
Byung-Gon Chun, Petros Maniatis, Scott Shenker, John Kubiatowicz |
FAST | 1 |
| 2009 | Minuet: Rethinking Concurrency Control in Storage Area Networks
Andrey Ermolinskiy, Daekyeong Moon, Byung-Gon Chun, Scott Shenker |
FAST | 3 |
| 2009 | Augmented Smartphone Applications Through Clone Cloud Execution
Byung-Gon Chun, Petros Maniatis |
HotOS | 1 |
| 2009 | RouteBricks: exploiting parallelism to scale software routersabstractWe revisit the problem of scaling software routers, motivated by recent advances in server technology that enable high-speed parallel processing--a feature router workloads appear ideally suited to exploit. We propose a software router architecture that parallelizes router functionality both across multiple servers and across multiple cores within a single server. By carefully exploiting parallelism at every opportunity, we demonstrate a 35Gbps parallel router prototype; this router capacity can be linearly scaled through the use of additional servers. Our prototype router is fully programmable using the familiar Click/Linux environment and is built entirely from off-the-shelf, general-purpose server hardware. Mihai Dobrescu, Norbert Egi, Katerina J. Argyraki, Byung-Gon Chun, Kevin R. Fall, Gianluca Iannaccone, Allan Knies, Maziar Manesh, Sylvia Ratnasamy |
SOSP | 4 |
| 2008 | NetComplex: A Complexity Metric for Networked System Designs
Byung-Gon Chun, Sylvia Ratnasamy, Eddie Kohler |
NSDI | 1 |
| 2008 | Diverse Replication for Single-Machine Byzantine-Fault Tolerance
Byung-Gon Chun, Petros Maniatis, Scott Shenker |
USENIX ATC | 1 |
| 2007 | Antiquity: exploiting a secure log for wide-area distributed storageabstractAntiquity is a wide-area distributed storage system designed to provide a simple storage service for applications like file systems and back-up. The design assumes that all servers eventually fail and attempts to maintain data despite those failures. Antiquity uses a secure log to maintain data integrity, replicates each log on multiple servers for durability, and uses dynamic Byzantine fault-tolerant quorum protocols to ensure consistency among replicas. We present Antiquity's design and an experimental evaluation with global and local testbeds. Antiquity has been running for over two months on 400+ PlanetLab servers storing nearly 20,000 logs totaling more than 84 GB of data. Despite constant server churn, all logs remain durable. Hakim Weatherspoon, Patrick R. Eaton, Byung-Gon Chun, John Kubiatowicz |
EuroSys | 3 |
| 2007 | Resolving inter-domain policy disputesabstractThe Border Gateway Protocol (BGP) allows each autonomous system (AS) to select routes to destinations based on semantically rich and locally determined policies. This autonomously exercised policy freedom can cause instability, where unresolvable policy-based disputes in the network result in interdomain route oscillations. Several recent works have established that such instabilities can only be eliminated by enforcing a globally accepted preference ordering on routes (such as shortest path). To resolve this conflict between policy autonomy and system stability, we propose a distributed mechanism that enforces a preference ordering only when disputes resulting in oscillations exist. This preserves policy freedom when possible, and imposes stability when required. Cheng Tien Ee, Vijay Ramachandran, Byung-Gon Chun, Kaushik Lakshminarayanan, Scott Shenker |
SIGCOMM | 3 |
| 2007 | A data-oriented (and beyond) network architectureabstractThe Internet has evolved greatly from its original incarnation. For instance, the vast majority of current Internet usage is data retrieval and service access, whereas the architecture was designed around host-to-host applications such as telnet and ftp. Moreover, the original Internet was a purely transparent carrier of packets, but now the various network stakeholders use middleboxes to improve security and accelerate applications. To adapt to these changes, we propose the Data-Oriented Network Architecture (DONA), which involves a clean-slate redesign of Internet naming and name resolution. Teemu Koponen, Mohit Chawla, Byung-Gon Chun, Andrey Ermolinskiy, Kye Hyun Kim, Scott Shenker, Ion Stoica |
SIGCOMM | 3 |
| 2007 | Attested append-only memory: making adversaries stick to their wordabstractResearchers have made great strides in improving the fault tolerance of both centralized and replicated systems against arbitrary (Byzantine) faults. However, there are hard limits to how much can be done with entirely untrusted components; for example, replicated state machines cannot tolerate more than a third of their replica population being Byzantine. In this paper, we investigate how minimal trusted abstractions can push through these hard limits in practical ways. We propose Attested Append-Only Memory (A2M), a trusted system facility that is small, easy to implement and easy to verify formally. A2M provides the programming abstraction of a trusted log, which leads to protocol designs immune to equivocation -- the ability of a faulty host to lie in different ways to different clients or servers -- which is a common source of Byzantine headaches. Using A2M, we improve upon the state of the art in Byzantine-fault tolerant replicated state machines, producing A2M-enabled protocols (variants of Castro and Liskov's PBFT) that remain correct (linearizable) and keep making progress (live) even when half the replicas are faulty, in contrast to the previous upper bound. We also present an A2M-enabled single-server shared storage protocol that guarantees linearizability despite server faults. We implement A2M and our protocols, evaluate them experimentally through micro- and macro-benchmarks, and argue that the improved fault tolerance is cost-effective for a broad range of uses, opening up new avenues for practical, more reliable services. Byung-Gon Chun, Petros Maniatis, Scott Shenker, John Kubiatowicz |
SOSP | 1 |
| 2006 | Efficient Replica Maintenance for Distributed Storage Systems
Byung-Gon Chun, Frank Dabek, Andreas Haeberlen, Emil Sit, Hakim Weatherspoon, M. Frans Kaashoek, John Kubiatowicz, Robert Morris 0005 |
NSDI | 1 |
| 2004 | Characterizing Selfishly Constructed Overlay Routing NetworksabstractWe analyze the characteristics of overlay routing networks generated by selfish nodes playing competitive network construction games. We explore several networking scenarios - some simplistic, others more realistic - and analyze the resulting Nash equilibrium graphs with respect to topology, performance, and resilience. We find a fundamental tradeoff between performance and resilience, and show that limiting the degree of nodes is of great importance in controlling this balance. Further, by varying the cost function, the game produces widely different topologies; one parameter in particular - the relative cost between maintaining an overlay link and increasing the path length to other nodes - can generate topologies with node-degree distributions whose tails vary from exponential to power-law. We conclude that competitive games can create overlay routing networks satisfying very diverse goals. Byung-Gon Chun, Rodrigo Fonseca, Ion Stoica, John Kubiatowicz |
INFOCOM | 1 |
| 2004 | Selfish caching in distributed systems: a game-theoretic analysisabstractWe analyze replication of resources by server nodes that act selfishly, using a game-theoretic approach. We refer to this as the selfish caching problem. In our model, nodes incur either cost for replicating resources or cost for access to a remote replica. We show the existence of pure strategy Nash equilibria and investigate the price of anarchy, which is the relative cost of the lack of coordination. The price of anarchy can be high due to undersupply problems, but with certain network topologies it has better bounds. With a payment scheme the game can always implement the social optimum in the best case by giving servers incentive to replicate. Byung-Gon Chun, Kamalika Chaudhuri, Hoeteck Wee, Marco Barreno, Christos H. Papadimitriou, John Kubiatowicz |
PODC | 1 |
| 1997 | Auxiliary Timeout and Selective Packet Discard Schemes to Improve TCP Performance in PCN EnvironmentabstractWe consider how to improve the performance of the TCP connections over personal communication networks by introducing the auxiliary timeout and selective packet discard schemes. Because TCP is a reliable transport protocol tuned to reliable wired networks, we need a data link protocol to supplement the TCP to overcome unreliable error environment in wireless networks. In this case, however, the data link protocol can interact with the TCP and thus influence the TCP performance. According to computer simulations, when the selective repeat ARQ is used as the data link protocol in the fading channel, the throughput decreases as the fading frequency increases; and when the fading frequency is low, TCP retransmission occurs spuriously and the throughput decreases due to delay variation and the misunderstanding of congestion. In order to solve these problems we introduce the auxiliary timeout and selective packet discard schemes. The auxiliary timeout keeps the rule of the congestion control in traditional networks and avoids spurious retransmissions and misbehaved congestion controls by acknowledging the delayed ACK. The selective packet discard is designed to avoid retransmitting data packets that the mobile station has acknowledged or will acknowledge soon at the base station, thus saving the wireless bandwidth. Therefore the proposed schemes can help to improve the TCP performance in wireless networks by increasing throughput and reducing spurious retransmissions. Byung-Gon Chun, Byeong Gi Lee |
ICC (1) | 1 |