Won Wook Song

dblp:198/6829 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-8530-2184ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Blaze: Holistic Caching for Iterative Data Processing
abstract
Modern data processing workloads, such as machine learning and graph processing, involve iterative computations to converge generated models into higher accuracy. An effective caching mechanism is vital to expedite iterative computations since the intermediate data that needs to be stored in memory grows larger over iterations, often exceeding the memory capacity. However, existing systems handle intermediate data through separate operational layers (e.g., caching, eviction, and recovery), with each layer working independently in a greedy or cost-agnostic manner. These layers typically rely on user annotations and past access patterns, failing to make globally optimal decisions for the workload.
Won Wook Song, Jeongyoon Eo, Taegeon Um, Myeongjae Jeon, Byung-Gon Chun
EuroSys1
2023 Sponge: Fast Reactive Scaling for Stream Processing with Serverless Frameworks
Won Wook Song, Taegeon Um, Sameh Elnikety, Myeongjae Jeon, Byung-Gon Chun
USENIX ATC1
2021 Harmony: A Scheduling Framework Optimized for Multiple Distributed Machine Learning Jobs
abstract
We introduce Harmony, a new scheduling framework that executes multiple Parameter-Server ML training jobs together to improve cluster resource utilization. Harmony coordinates a fine-grained execution of co-located jobs with complementary resource usages to avoid contention and to efficiently share resources between the jobs. To resolve the memory pressure due to the increased number of simultaneous jobs, Harmony uses a data spill/reload mechanism optimized for multiple jobs with the iterative execution pattern. Our evaluation shows that Harmony improves cluster resource utilization by up to 1.65×, resulting in a reduction of the mean ML training job time by about 53%, and makespan, the total time to process all given jobs, by about 38%, compared to the traditional approaches that allocate dedicated resources to each job.
Woo-Yeon Lee, Yunseong Lee, Won Wook Song, Youngseok Yang, Byung-Gon Chun
ICDCS3
2020 Apache Nemo: A Framework for Optimizing Distributed Data Processing
abstract
Optimizing scheduling and communication of distributed data processing for resource and data characteristics is crucial for achieving high performance. Existing approaches to such optimizations largely fall into two categories. First, distributed runtimes provide low-level policy interfaces to apply the optimizations, but do not ensure the maintenance of correct application semantics and thus often require significant effort to use. Second, policy interfaces that extend a high-level application programming model ensure correctness, but do not provide sufficient fine control. We describe Apache Nemo, an optimization framework for distributed dataflow processing that provides fine control for high performance and also ensures correctness for ease of use. We combine several techniques to achieve this, including an intermediate representation of dataflow, compiler optimization passes, and runtime extensions. Our evaluation results show that Nemo enables composable and reusable optimizations that bring performance improvements on par with existing specialized runtimes tailored for a specific deployment scenario. Apache Nemo is open-sourced at https://nemo.apache.org as an Apache incubator project.
Won Wook Song, Youngseok Yang, Jeongyoon Eo, Jangho Seo, Sanha Lee, Gyewon Lee, Taegeon Um, Haeyoon Cho 0001, Byung-Gon Chun
ACM Trans. Comput. Syst.1
2019 Automating System Configuration of Distributed Machine Learning
abstract
The performance of distributed machine learning systems is dependent on their system configuration. However, configuring the system for optimal performance is challenging and time consuming even for experts due to the diverse runtime factors such as workloads or the system environment. We present cost-based optimization to automatically find a good system configuration for parameter server (PS) machine learning (ML) frameworks. We design and implement Cruise that applies the optimization technique to tune distributed PS ML execution automatically. Evaluation results on three ML applications verify that Cruise automates the system configuration of the applications to achieve good performance with minor reconfiguration costs.
Woo-Yeon Lee, Markus Weimer, Byung-Gon Chun, Yunseong Lee, Joo Seong Jeong, Gyeong-In Yu, Hojin Park, Beomyeol Jeon, Won Wook Song, Gunhee Kim
ICDCS11
2019 Apache Nemo: A Framework for Building Distributed Dataflow Optimization Policies
Youngseok Yang, Jeongyoon Eo, Geon-Woo Kim, Sanha Lee, Jangho Seo, Won Wook Song, Byung-Gon Chun
USENIX ATC7
2017 Pado: A Data Processing Engine for Harnessing Transient Resources in Datacenters
abstract
Datacenters are under-utilized, primarily due to unused resources on over-provisioned nodes of latency-critical jobs. Such idle resources can be used to run batch data analytic jobs to increase datacenter utilization, but these transient resources must be evicted whenever latency-critical jobs require them again. Resource evictions often lead to cascading recomputations, which is usually handled by checkpointing intermediate results on stable storages of eviction-free reserved resources. However, checkpointing has major shortcomings in its substantial overhead of transferring data back and forth. In this work, we step away from such approaches and focus on observing the job structure and the relationships between computations of the job. We carefully mark the computations that are most likely to cause a large number of recomputations upon evictions, to run them reliably using reserved resources. This lets us retain corresponding intermediate results effortlessly without any additional checkpointing. We design Pado, a general data processing engine, which carries out our idea with several optimizations that minimize the number of additional reserved nodes. Evaluation results show that Pado outperforms Spark 2.0.0 by up to 5.1×, and checkpoint-enabled Spark by up to 3.8×.
Youngseok Yang, Geon-Woo Kim, Won Wook Song, Yunseong Lee, Zhengping Qian, Byung-Gon Chun
EuroSys3