Myeong-Seon Gil

dblp:11/8263 · also Myeong-Sun Gil · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-4611-644XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HBAC: Hierarchical-Based Access Control Model for Storage Management in Data Lake Environments
abstract
ABSTRACT Background Traditional storage systems typically use simple access control models that manage permissions at the user or group level. These models, however, are not well suited for data lake environments, where large‐scale and diverse datasets must be accessed concurrently by users across multiple hierarchical levels, and they also have many limitations in terms of maintenance. Aims In this paper, we propose a novel access control model, HBAC (Hierarchical‐Based Access Control), designed to provide hierarchical access control in Ceph‐based distributed storage environments. Methods HBAC introduces a hierarchical structure into permission management, ensuring that higher‐level users are granted broader permissions than lower‐level users. Furthermore, by providing hierarchical group‐based management capabilities, HBAC overcomes the limitations of conventional ACL (Access Control List) models, which only support simple mappings between users and objects. This design also enables fine‐grained and flexible access control, particularly in collaborative, large‐scale data environments. We present a formal permission granting algorithm with specific example scenarios to ensure the stable implementation of HBAC on top of Ceph. This model leverages Ceph's metadata storage to centrally manage permissions without requiring additional software, while ensuring efficient data handling through its distributed architecture. Results and Conclusions Experimental evaluation confirms that Ceph‐based HBAC effectively reflects user hierarchies and reliably provides complex access control. Furthermore, an efficiency comparison on policy representation shows that HBAC achieves approximately a 15‐fold reduction in policy complexity compared to AWS IAM for equivalent permission configurations, and up to 21 times faster permission processing speeds compared to Ceph's default access control mechanisms, demonstrating its superior performance.
Yisac Hong, Myeong-Seon Gil, Yang-Sae Moon
Softw. Pract. Exp.2
2025 Panacea: An Automatic Data Migration Framework for Constructing Internet-Scale Open Data Lakes
abstract
ABSTRACT Background With the recent growth in data‐driven research and development, the need to build integrated data lakes targeting Internet‐scale open data is rapidly increasing. In this study, we investigate the limitations of open data management and Internet‐scale migration to construct an integrated data lake, deriving related problems. First, open data lakes (ODLs) face problems of preprocessing complexity, scalability limitation, and platform dependency owing to their data management method and the characteristics of open data. Second, migrating data from distributed sources to a single data lake incurs problems such as migration incompleteness, scalability limitation, and resource wastage. Aims In this study, we propose Panacea, a novel automation framework designed to solve these problems in the construction and migration of ODLs. Methods Panacea addresses the first three problems caused by the characteristics of ODLs through automation expansion. Specifically, it resolves preprocessing complexity and scalability limitation by supporting the automation of catalog collection and preprocessing tasks, and alleviates platform dependency through universality across representative platforms by automating detailed catalog processing logic. Panacea also addresses the latter three problems of migration. Specifically, it supports the construction of domain‐based Internet‐scale data lakes by addressing migration incompleteness and scalability limitation through original data migration and automation expansion. Additionally, it tackles resource wastage by providing metadata management functions. Results We address the aforementioned problems through automation logic and management modules that consider the characteristics of open data lakes. Comparative experiments with the legacy IMP‐CKAN demonstrate that Panacea performs migrations up to 216% faster in large‐scale experiments and up to 409% faster in large‐volume experiments. Moreover, its automation performance is up to 313% better than that of IMP‐CKAN. Conclusion These experimental results indicate that Panacea is an excellent automation framework that enhances the utilization of open data and supports Internet‐scale migration for collecting research data. Furthermore, we demonstrate how to integrate Panacea with the open‐source framework Demeter, which can further enhance data usability. This framework will significantly assist many researchers facing data scarcity challenges.
Dasol Kim, Hee-Sun Won, Myeong-Seon Gil, Yang-Sae Moon
Softw. Pract. Exp.3
2024 Demeter: An automatic framework for data migration in open data lakes
abstract
Abstract An open data lake stores various forms and types of open data, and there is an increasing demand to manage raw data in tables rather than files for efficient data exploration and analysis. In this paper, we investigate the data management of open data lakes and recognize the limitations of table migration and related problems. First, open data lakes have problems of preprocessing complexity, scale limitation, and platform dependency due to the traditional data management method and open data characteristics. Second, existing studies for table migration have problems of lack of scalability, migration incompleteness, and scale limitation. In this work, we present a novel automation framework, called Demeter, which solves three problems inherent in open data lakes by expanding automation. Specifically, it supports automating catalog collection and preprocessing tasks to solve preprocessing complexity and scale limitation. It also supports platform universality for representative data platforms through the automation of catalog analysis and detailed processing logic. Demeter then solves three problems in table migration by adopting Airbyte, an open‐source ELT platform, and by enhancing automation capability with the Airbyte manager. We verify that Demeter resolves all the problems above through extensive experiments and proves its scalability and universality. In addition, significantly outperforms CKAN by Demeter up to 508.5% in automation performance, up to 207.28% in processing time, and up to 917.17% in migration performance. These results indicate that Demeter is an excellent automation framework that increases the utilization of large‐scale open data and supports reliable Internet‐scale migration.
Dasol Kim, Jiwoo Han, Siwoon Son, Myeong-Seon Gil, Yang-Sae Moon, Hee-Sun Won
Softw. Pract. Exp.4
2021 An Advanced Open Data Platform for Integrated Support of Data Management, Distribution, and Analysis
abstract
With the growing applications of big data and artificial intelligence, the quality of the service is an outcome of the quality of the data. Nevertheless, there is still a significant lack of data that has practical application value. To solve these problems, we propose SODAS (Smart Open Data As a Service) as a novel open data platform for efficient data sharing and utilization. We first analyze the major problems in the legacy CKAN and then draw up their solutions through core strategies. We next define four components and nine function blocks of SODAS for each core strategy. As a result, SODAS drives Open Data Portal, Open Data Reference Model, DataMap Publisher, and ADE Provisioning (Analytics and Development Environment Provisioning) by connecting the defined function blocks. We confirm that each function works correctly through the SODAS Web portal and apply SODAS to actual data distribution sites to prove its efficiency and practical use. SODAS is the first open data platform that provides secure interoperability between heterogeneous platforms based on international standards and enables domain-free data management with flexible metadata.
Hee-Sun Won, Minh Chau Nguyen, Myeong-Seon Gil, Yang-Sae Moon
IEEE BigData3
2019 Prefetching-based metadata management in Advanced Multitenant Hadoop
Minh Chau Nguyen, Hee-Sun Won, Siwoon Son, Myeong-Seon Gil, Yang-Sae Moon
J. Supercomput.4
2017 Efficient Two-Step Protocol and Its Discriminative Feature Selections in Secure Similar Document Detection
abstract
Recently, the risk of information disclosure is increasing significantly. Accordingly, privacy-preserving data mining (PPDM) is being actively studied to obtain accurate mining results while preserving the data privacy. We here focus on secure similar document detection (SSDD), which identifies similar documents of two parties when each party does not disclose its own sensitive documents to the another party. In this paper, we propose an efficient two-step protocol that exploits a feature selection as a lower-dimensional transformation, and we present discriminative feature selections to maximize the performance of the protocol. The proposed protocol consists of two steps: thefilteringstep and thepostprocessingstep. For the feature selection, we first consider the simplest one, random projection (RP), and propose its two-step solution,SSDD-RP. We then present two discriminative feature selections and their solutions:SSDD-LFwhich selects a few dimensions locally frequent in the current querying vector andSSDD-GFwhich selects ones globally frequent in the set of all document vectors. We finally propose a hybrid one,SSDD-HF, which takes advantage of bothSSDD-LFandSSDD-GF. We empirically show that the proposed two-step protocol significantly outperforms the previous one-step protocol by three or four orders of magnitude.
Sang-Pil Kim, Myeong-Seon Gil, Hajin Kim, Mi-Jung Choi, Yang-Sae Moon, Hee-Sun Won
Secur. Commun. Networks2
2017 Moving metadata from ad hoc files to database tables for robust, highly available, and scalable HDFS
Hee-Sun Won, Minh Chau Nguyen, Myeong-Seon Gil, Yang-Sae Moon, Kyu-Young Whang
J. Supercomput.3