VLDB 2026 Research / reviewers in the wild / expert
Dasol Kim
dblp:280/7399
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-4389-5070ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Panacea: An Automatic Data Migration Framework for Constructing Internet-Scale Open Data LakesabstractABSTRACT Background With the recent growth in data‐driven research and development, the need to build integrated data lakes targeting Internet‐scale open data is rapidly increasing. In this study, we investigate the limitations of open data management and Internet‐scale migration to construct an integrated data lake, deriving related problems. First, open data lakes (ODLs) face problems of preprocessing complexity, scalability limitation, and platform dependency owing to their data management method and the characteristics of open data. Second, migrating data from distributed sources to a single data lake incurs problems such as migration incompleteness, scalability limitation, and resource wastage. Aims In this study, we propose Panacea, a novel automation framework designed to solve these problems in the construction and migration of ODLs. Methods Panacea addresses the first three problems caused by the characteristics of ODLs through automation expansion. Specifically, it resolves preprocessing complexity and scalability limitation by supporting the automation of catalog collection and preprocessing tasks, and alleviates platform dependency through universality across representative platforms by automating detailed catalog processing logic. Panacea also addresses the latter three problems of migration. Specifically, it supports the construction of domain‐based Internet‐scale data lakes by addressing migration incompleteness and scalability limitation through original data migration and automation expansion. Additionally, it tackles resource wastage by providing metadata management functions. Results We address the aforementioned problems through automation logic and management modules that consider the characteristics of open data lakes. Comparative experiments with the legacy IMP‐CKAN demonstrate that Panacea performs migrations up to 216% faster in large‐scale experiments and up to 409% faster in large‐volume experiments. Moreover, its automation performance is up to 313% better than that of IMP‐CKAN. Conclusion These experimental results indicate that Panacea is an excellent automation framework that enhances the utilization of open data and supports Internet‐scale migration for collecting research data. Furthermore, we demonstrate how to integrate Panacea with the open‐source framework Demeter, which can further enhance data usability. This framework will significantly assist many researchers facing data scarcity challenges. Dasol Kim, Hee-Sun Won, Myeong-Seon Gil, Yang-Sae Moon |
Softw. Pract. Exp. | 1 |
| 2024 | Demeter: An automatic framework for data migration in open data lakesabstractAbstract An open data lake stores various forms and types of open data, and there is an increasing demand to manage raw data in tables rather than files for efficient data exploration and analysis. In this paper, we investigate the data management of open data lakes and recognize the limitations of table migration and related problems. First, open data lakes have problems of preprocessing complexity, scale limitation, and platform dependency due to the traditional data management method and open data characteristics. Second, existing studies for table migration have problems of lack of scalability, migration incompleteness, and scale limitation. In this work, we present a novel automation framework, called Demeter, which solves three problems inherent in open data lakes by expanding automation. Specifically, it supports automating catalog collection and preprocessing tasks to solve preprocessing complexity and scale limitation. It also supports platform universality for representative data platforms through the automation of catalog analysis and detailed processing logic. Demeter then solves three problems in table migration by adopting Airbyte, an open‐source ELT platform, and by enhancing automation capability with the Airbyte manager. We verify that Demeter resolves all the problems above through extensive experiments and proves its scalability and universality. In addition, significantly outperforms CKAN by Demeter up to 508.5% in automation performance, up to 207.28% in processing time, and up to 917.17% in migration performance. These results indicate that Demeter is an excellent automation framework that increases the utilization of large‐scale open data and supports reliable Internet‐scale migration. Dasol Kim, Jiwoo Han, Siwoon Son, Myeong-Seon Gil, Yang-Sae Moon, Hee-Sun Won |
Softw. Pract. Exp. | 1 |