EDBT 2026 Demo / reviewers in the wild / expert
Gyewon Lee
dblp:205/6847
· DBLP profile ↗
6ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0002-6543-8877ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | FlowKV: A Semantic-Aware Store for Large-Scale State Management of Stream Processing EnginesabstractWe propose FlowKV, a persistent store tailored for large-scale state management of streaming applications. Unlike existing KV stores, FlowKV leverages information from stream processing engines by taking a principled approach toward exploiting information about how and when the applications access data. FlowKV categorizes data access patterns of window operations according to how window boundaries are set and how tuples inside a window are aggregated, and deploys customized in-memory and on-disk data structures optimized for each pattern. In addition, FlowKV takes window metadata as explicit arguments of read and write methods to predict the moment when a window is read, and then loads the tuples of windows in batches from storage ahead of time. Using the NEXMark benchmark as workload, our experiments show that Apache Flink on FlowKV outperforms Flink on RocksDB or Faster with up to 4.12× throughput gain. Gyewon Lee, Jaewoo Maeng, Jinsol Park, Jangho Seo, Haeyoon Cho 0001, Youngseok Yang, Taegeon Um, Jongsung Lee 0001, Jae W. Lee, Byung-Gon Chun |
EuroSys | 1 |
| 2021 | SCOPA: Soft Code-Switching and Pairwise Alignment for Zero-Shot Cross-lingual TransferabstractThe recent advent of cross-lingual embeddings, such as multilingual BERT (mBERT), provides a strong baseline for zero-shot cross-lingual transfer. There also exists increasing research attention to reduce the alignment discrepancy of cross-lingual embeddings between source and target languages, via generating code-switched sentences by substituting randomly selected words in the source languages with their counterparts of the target languages. Although these approaches improve the performance, naively code-switched sentences can have inherent limitations. In this paper, we propose SCOPA, a novel technique to improve the performance of zero-shot cross-lingual transfer. Instead of using the embeddings of code-switched sentences directly, SCOPA mixes them softly with the embeddings of original sentences. In addition, SCOPA utilizes an additional pairwise alignment objective, which aligns the vector differences of word pairs instead of word-level embeddings, in order to transfer contextualized information between different languages while preserving language-specific information. Experiments on the PAWS-X and MLDoc dataset show the effectiveness of SCOPA. Dohyeon Lee, Jaeseong Lee 0002, Gyewon Lee, Byung-Gon Chun, Seung-won Hwang |
CIKM | 3 |
| 2021 | Pluto: High-Performance IoT-Aware Stream ProcessingabstractNowadays, large numbers of small IoT stream queries are created from diverse IoT applications and executed on cloud backend servers. However, existing distributed stream processing systems such as Storm and Flink do not efficiently handle the large numbers of IoT stream queries because of their tightly-coupled query/code submission layer and inefficient query execution layer. In this paper, we propose Pluto, a new IoT-aware stream processing system. As a first step for IoT stream processing, this paper focuses on optimizing the execution of many IoT stream queries on a node. Pluto optimizes the end-to-end query processing with a three-phase execution, harnessing IoT-query characteristics. First, Pluto minimizes bottlenecks in the IoT query submission by decoupling the code registration from the query submission process with new APIs, which eliminates duplicate code registration and enables code sharing across queries. Second, in the execution phase, Pluto shares system resources as much as possible and minimizes resource bottlenecks in a machine by exploiting commonalities among IoT stream queries and information exposed in the API. Our evaluations show that Pluto improves the throughput by an order of magnitude compared to other stream processing systems on a 24-core machine, keeping P99 latency less than one second. Taegeon Um, Gyewon Lee, Byung-Gon Chun |
ICDCS | 2 |
| 2021 | Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network Training
Gyewon Lee, Irene Lee, Hyeonmin Ha, Kyung-Geun Lee, Hwarim Hyun, Ahnjae Shin, Byung-Gon Chun |
USENIX ATC | 1 |
| 2020 | Apache Nemo: A Framework for Optimizing Distributed Data ProcessingabstractOptimizing scheduling and communication of distributed data processing for resource and data characteristics is crucial for achieving high performance. Existing approaches to such optimizations largely fall into two categories. First, distributed runtimes provide low-level policy interfaces to apply the optimizations, but do not ensure the maintenance of correct application semantics and thus often require significant effort to use. Second, policy interfaces that extend a high-level application programming model ensure correctness, but do not provide sufficient fine control. We describe Apache Nemo, an optimization framework for distributed dataflow processing that provides fine control for high performance and also ensures correctness for ease of use. We combine several techniques to achieve this, including an intermediate representation of dataflow, compiler optimization passes, and runtime extensions. Our evaluation results show that Nemo enables composable and reusable optimizations that bring performance improvements on par with existing specialized runtimes tailored for a specific deployment scenario. Apache Nemo is open-sourced at https://nemo.apache.org as an Apache incubator project. Won Wook Song, Youngseok Yang, Jeongyoon Eo, Jangho Seo, Sanha Lee, Gyewon Lee, Taegeon Um, Haeyoon Cho 0001, Byung-Gon Chun |
ACM Trans. Comput. Syst. | 7 |
| 2017 | Apache REEF: Retainable Evaluator Execution FrameworkabstractResource Managers like YARN and Mesos have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault tolerance, task scheduling and coordination) and reimplement common mechanisms (e.g., caching, bulk-data transfers). This article presents REEF, a development framework that provides a control plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource reuse for data caching and state management abstractions that greatly ease the development of elastic data processing pipelines on cloud platforms that support a Resource Manager service. We illustrate the power of REEF by showing applications built atop: a distributed shell application, a machine-learning framework, a distributed in-memory caching system, and a port of the CORFU system. REEF is currently an Apache top-level project that has attracted contributors from several institutions and it is being used to develop several commercial offerings such as the Azure Stream Analytics service. Byung-Gon Chun, Tyson Condie, Yingda Chen, Carlo Curino, Chris Douglas, Matteo Interlandi, Beomyeol Jeon, Joo Seong Jeong, Gyewon Lee, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Mariia Mykhailova, Shravan M. Narayanamurthy, Joseph Noor, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Taegeon Um, Julia Wang, Markus Weimer, Youngseok Yang |
ACM Trans. Comput. Syst. | 11 |