EDBT 2026 Demo / reviewers in the wild / expert
Woo-Yeon Lee
dblp:163/1527
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ORBITFLOW: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
Heelim Hong, Taegeon Um, Jongseop Lee, Seoyeong Choy, Woo-Yeon Lee, Myeongjae Jeon |
Proc. VLDB Endow. | 6 |
| 2024 | Metis: Fast Automatic Distributed Training on Heterogeneous GPUs
Taegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee, Goeun Kim, Dongseob Kim, Youngtaek Kim, Mohd Muzzammil, Myeongjae Jeon |
USENIX ATC | 4 |
| 2023 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineabstractWhen training a deep learning (DL) model, input data are pre-processed on CPUs and transformed into tensors, which are then fed into GPUs for gradient computations of model training. Expensive GPUs must be fully utilized during training to accelerate the training speed. However, intensive CPU operations for input data preprocessing (input pipeline) often lead to CPU bottlenecks; correspondingly, various DL training jobs suffer from GPU under-utilization. We propose FastFlow, a DL training system that automatically mitigates the CPU bottleneck by offloading (scaling out) input pipelines to remote CPUs. FastFlow carefully decides various offloading decisions based on performance metrics specific to applications and allocated resources, while leveraging both local and remote CPUs to prevent the inefficient use of remote resources and minimize the training time. FastFlow's smart offloading policy and mechanisms are seamlessly integrated with TensorFlow for users to enjoy the smart offloading features without modifying the main logic. Our evaluations on our private DL cloud with diverse workloads on various resource environments show that FastFlow improves the training throughput by 1 ~ 4.34X compared to TensorFlow without offloading, by 1 ~ 4.52X compared to TensorFlow with manual CPU offloading (tf.data.service), and by 0.63 ~ 2.06X compared to GPU offloading (DALI). Taegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun, Goeun Kim, Woo-Yeon Lee |
Proc. VLDB Endow. | 6 |
| 2022 | PokéMem: Taming Wild Memory Consumers in Apache SparkabstractApache Spark is a widely used in-memory processing system due to its high performance. For fast data processing, Spark manages in-memory data such as cached or shuffling (aggregate and sorting) data in its own managed memory pools. However, despite its sophisticated memory management scheme, we found that Spark still suffers from out-of-memory (OOM) exceptions and high garbage collection (GC) overheads when wild memory consumers, who are not tracked by Spark and execute external codes, use a large amount of memory. To resolve the problems, we propose PokéMem, which is an enhanced Spark that incorporates wild memory consumers into the managed ones to prevent them from taking up memory spaces excessively in stealth. Our main idea is to open the black-box of unmanaged memory regions in external codes by providing customized data collections. PokéMem enables fine-grained controls of created objects within running tasks, by spilling and reloading the objects of custom data collections based on the memory pressure and access patterns. To further reduce memory pressures, PokéMem exploits pre-built memory estimation models to predict the external code's memory usage and proactively acquires memory before the execution of external code, and also performs JVM heap-usage monitoring to avoid critical memory pressures. With the help of these techniques, our evaluations show that PokéMem outperforms vanilla Spark with at most 3× faster execution with 3.9× smaller GC overheads, and successfully runs workloads without OOM exception that vanilla Spark has failed to run. Minhyeok Kweun, Goeun Kim, Byungsoo Oh, Seongho Jung, Taegeon Um, Woo-Yeon Lee |
IPDPS | 6 |
| 2021 | Harmony: A Scheduling Framework Optimized for Multiple Distributed Machine Learning JobsabstractWe introduce Harmony, a new scheduling framework that executes multiple Parameter-Server ML training jobs together to improve cluster resource utilization. Harmony coordinates a fine-grained execution of co-located jobs with complementary resource usages to avoid contention and to efficiently share resources between the jobs. To resolve the memory pressure due to the increased number of simultaneous jobs, Harmony uses a data spill/reload mechanism optimized for multiple jobs with the iterative execution pattern. Our evaluation shows that Harmony improves cluster resource utilization by up to 1.65×, resulting in a reduction of the mean ML training job time by about 53%, and makespan, the total time to process all given jobs, by about 38%, compared to the traditional approaches that allocate dedicated resources to each job. Woo-Yeon Lee, Yunseong Lee, Won Wook Song, Youngseok Yang, Byung-Gon Chun |
ICDCS | 1 |
| 2020 | Lineage Checkpoint Approach for Long-lineage Problem in Apache SparkabstractIn distributed data processing frameworks, data lineage is widely used to achieve fault-tolerance efficiently. However, as data analytics apps become more complex, long lineage incurs the performance and reliability issues. Existing data checkpoint solutions can alleviate the problem, however they bring new types of heavy inevitable overheads or lack fault-tolerance. To overcome the limitations, we propose the solution that checkpoints lineage graph instead of data itself, preserving the lineage information and reconciling the full lineage graph when needed. In evaluations, we show that our lineage checkpoint solution outperforms the data checkpoint solutions in terms of performance and fault-tolerance. Minhyeok Kweun, Woo-Yeon Lee, Goeun Kim, Jisoo Hwang, Yoonkyong Lee |
IEEE BigData | 2 |
| 2019 | Automating System Configuration of Distributed Machine LearningabstractThe performance of distributed machine learning systems is dependent on their system configuration. However, configuring the system for optimal performance is challenging and time consuming even for experts due to the diverse runtime factors such as workloads or the system environment. We present cost-based optimization to automatically find a good system configuration for parameter server (PS) machine learning (ML) frameworks. We design and implement Cruise that applies the optimization technique to tune distributed PS ML execution automatically. Evaluation results on three ML applications verify that Cruise automates the system configuration of the applications to achieve good performance with minor reconfiguration costs. Woo-Yeon Lee, Markus Weimer, Byung-Gon Chun, Yunseong Lee, Joo Seong Jeong, Gyeong-In Yu, Hojin Park, Beomyeol Jeon, Won Wook Song, Gunhee Kim |
ICDCS | 1 |
| 2015 | Elastic Memory: Bring Elasticity Back to In-Memory Big Data Analytics
Joo Seong Jeong, Woo-Yeon Lee, Yunseong Lee, Youngseok Yang, Byung-Gon Chun |
HotOS | 2 |