Yang-Sae Moon

dblp:13/4420 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-2396-0405ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 32 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 1 since 2021Security and privacy · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorSystems, architecture and hardware · 4 · 1 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 HBAC: Hierarchical-Based Access Control Model for Storage Management in Data Lake Environments
abstract
ABSTRACT Background Traditional storage systems typically use simple access control models that manage permissions at the user or group level. These models, however, are not well suited for data lake environments, where large‐scale and diverse datasets must be accessed concurrently by users across multiple hierarchical levels, and they also have many limitations in terms of maintenance. Aims In this paper, we propose a novel access control model, HBAC (Hierarchical‐Based Access Control), designed to provide hierarchical access control in Ceph‐based distributed storage environments. Methods HBAC introduces a hierarchical structure into permission management, ensuring that higher‐level users are granted broader permissions than lower‐level users. Furthermore, by providing hierarchical group‐based management capabilities, HBAC overcomes the limitations of conventional ACL (Access Control List) models, which only support simple mappings between users and objects. This design also enables fine‐grained and flexible access control, particularly in collaborative, large‐scale data environments. We present a formal permission granting algorithm with specific example scenarios to ensure the stable implementation of HBAC on top of Ceph. This model leverages Ceph's metadata storage to centrally manage permissions without requiring additional software, while ensuring efficient data handling through its distributed architecture. Results and Conclusions Experimental evaluation confirms that Ceph‐based HBAC effectively reflects user hierarchies and reliably provides complex access control. Furthermore, an efficiency comparison on policy representation shows that HBAC achieves approximately a 15‐fold reduction in policy complexity compared to AWS IAM for equivalent permission configurations, and up to 21 times faster permission processing speeds compared to Ceph's default access control mechanisms, demonstrating its superior performance.
Yisac Hong, Myeong-Seon Gil, Yang-Sae Moon
Softw. Pract. Exp.4
2025 Panacea: An Automatic Data Migration Framework for Constructing Internet-Scale Open Data Lakes
abstract
ABSTRACT Background With the recent growth in data‐driven research and development, the need to build integrated data lakes targeting Internet‐scale open data is rapidly increasing. In this study, we investigate the limitations of open data management and Internet‐scale migration to construct an integrated data lake, deriving related problems. First, open data lakes (ODLs) face problems of preprocessing complexity, scalability limitation, and platform dependency owing to their data management method and the characteristics of open data. Second, migrating data from distributed sources to a single data lake incurs problems such as migration incompleteness, scalability limitation, and resource wastage. Aims In this study, we propose Panacea, a novel automation framework designed to solve these problems in the construction and migration of ODLs. Methods Panacea addresses the first three problems caused by the characteristics of ODLs through automation expansion. Specifically, it resolves preprocessing complexity and scalability limitation by supporting the automation of catalog collection and preprocessing tasks, and alleviates platform dependency through universality across representative platforms by automating detailed catalog processing logic. Panacea also addresses the latter three problems of migration. Specifically, it supports the construction of domain‐based Internet‐scale data lakes by addressing migration incompleteness and scalability limitation through original data migration and automation expansion. Additionally, it tackles resource wastage by providing metadata management functions. Results We address the aforementioned problems through automation logic and management modules that consider the characteristics of open data lakes. Comparative experiments with the legacy IMP‐CKAN demonstrate that Panacea performs migrations up to 216% faster in large‐scale experiments and up to 409% faster in large‐volume experiments. Moreover, its automation performance is up to 313% better than that of IMP‐CKAN. Conclusion These experimental results indicate that Panacea is an excellent automation framework that enhances the utilization of open data and supports Internet‐scale migration for collecting research data. Furthermore, we demonstrate how to integrate Panacea with the open‐source framework Demeter, which can further enhance data usability. This framework will significantly assist many researchers facing data scarcity challenges.
Dasol Kim, Hee-Sun Won, Myeong-Seon Gil, Yang-Sae Moon
Softw. Pract. Exp.4
2024 Themis: A GPU-accelerated Relational Query Execution Engine
abstract
GPU-accelerated relational query execution engines have parallelized the execution of a pipeline, a sequence of operators. For the parallelization, the engines evenly partition the tuples in a table that will be scanned by the pipeline's first operator (a scan), and each thread executes the pipeline for the tuples in a partition. However, this approach leads to load imbalances since an operator returns a varying number of output tuples per input tuple, particularly under non-uniform data distributions such as skewed join key values. The load imbalances are classified into intra- and inter-warp load imbalances (intra-WLIs and inter-WLIs) since 1) threads are grouped into warps and 2) every thread in a warp evaluates the same operator for an input tuple concurrently following a single-instruction-multiple-thread manner. In contrast, threads in different warps can evaluate different operators concurrently. Although load balancing techniques have been proposed, however, they fail to solve the load imbalances on various workloads. In this paper, we propose a query execution engine, Themis, named after the deity of fairness, which symbolizes balanced workloads within our context. Themis minimizes intra-WLIs and inter-WLIs across various workloads. First, Themis minimizes intra-WLIs by redistributing tuples between the threads in a warp and making the threads evaluate an operator only when all of them hold inputs. Second, Themis mitigates the inter-WLIs by redistributing the tuples of warps with heavy workloads to idle warps. To check whether a warp's workload is heavy, we propose a method to approximate the sizes of warps' workloads. Based on these approximations, Themis adaptively adjusts the threshold for determining a warp's workload as heavy. In a recent benchmark JCC-H, which introduces skewed join key distributions to TPC-H, Themis significantly alleviates the inter-WLIs and intra-WLIs, outperforming the runner-up by up to 379x.
Kijae Hong, Kyoungmin Kim 0002, Young-Koo Lee, Yang-Sae Moon, Sourav S. Bhowmick, Wook-Shin Han
Proc. VLDB Endow.4
2024 Demeter: An automatic framework for data migration in open data lakes
abstract
Abstract An open data lake stores various forms and types of open data, and there is an increasing demand to manage raw data in tables rather than files for efficient data exploration and analysis. In this paper, we investigate the data management of open data lakes and recognize the limitations of table migration and related problems. First, open data lakes have problems of preprocessing complexity, scale limitation, and platform dependency due to the traditional data management method and open data characteristics. Second, existing studies for table migration have problems of lack of scalability, migration incompleteness, and scale limitation. In this work, we present a novel automation framework, called Demeter, which solves three problems inherent in open data lakes by expanding automation. Specifically, it supports automating catalog collection and preprocessing tasks to solve preprocessing complexity and scale limitation. It also supports platform universality for representative data platforms through the automation of catalog analysis and detailed processing logic. Demeter then solves three problems in table migration by adopting Airbyte, an open‐source ELT platform, and by enhancing automation capability with the Airbyte manager. We verify that Demeter resolves all the problems above through extensive experiments and proves its scalability and universality. In addition, significantly outperforms CKAN by Demeter up to 508.5% in automation performance, up to 207.28% in processing time, and up to 917.17% in migration performance. These results indicate that Demeter is an excellent automation framework that increases the utilization of large‐scale open data and supports reliable Internet‐scale migration.
Dasol Kim, Jiwoo Han, Siwoon Son, Myeong-Seon Gil, Yang-Sae Moon, Hee-Sun Won
Softw. Pract. Exp.5
2022 Regular Path Query Evaluation Sharing a Reduced Transitive Closure Based on Graph Reduction
abstract
Regular path queries (RPQs) find pairs of vertices of paths satisfying given regular expressions on an edge-labeled, directed multigraph. When evaluating an RPQ, the evaluation of a Kleene closure is very expensive. Furthermore, when multiple RPQs include a Kleene closure as a common sub-query, repeated evaluations of the common sub-query cause serious performance degradation. In this paper, we present a novel concept of RPQ-based graph reduction, which significantly simplifies the original graph through edge-level and vertex-level reductions. Interestingly, RPQ-based graph reduction can replace the evaluation of the Kleene closure on the large original graph to that of the transitive closure to the small reduced graph. We then propose a reduced transitive closure (RTC) as a lightweight structure for efficiently sharing the result of a Kleene closure. We also present an RPQ evaluation algorithm, RTCSharing, which treats each clause in the disjunctive normal form of the given RPQ as a batch unit. If the batch units include a Kleene closure as a common sub-query, we share the lightweight RTC instead of the heavyweight result of the Kleene closure. RPQ-based graph reduction further enables us to formally represent the result of an RPQ including a Kleene closure as a relational algebra expression including the RTC. Through the formal expression, we optimize the evaluation of the batch unit by eliminating useless and redundant operations of the previous method. Experiments show that RTCSharing improves the performance significantly by up to 73.86 times compared with existing methods in terms of query response time.
Inju Na, Yang-Sae Moon, Ilyeop Yi, Kyu-Young Whang, Soon J. Hyun
ICDE2
2021 An Advanced Open Data Platform for Integrated Support of Data Management, Distribution, and Analysis
abstract
With the growing applications of big data and artificial intelligence, the quality of the service is an outcome of the quality of the data. Nevertheless, there is still a significant lack of data that has practical application value. To solve these problems, we propose SODAS (Smart Open Data As a Service) as a novel open data platform for efficient data sharing and utilization. We first analyze the major problems in the legacy CKAN and then draw up their solutions through core strategies. We next define four components and nine function blocks of SODAS for each core strategy. As a result, SODAS drives Open Data Portal, Open Data Reference Model, DataMap Publisher, and ADE Provisioning (Analytics and Development Environment Provisioning) by connecting the defined function blocks. We confirm that each function works correctly through the SODAS Web portal and apply SODAS to actual data distribution sites to prove its efficiency and practical use. SODAS is the first open data platform that provides secure interoperability between heterogeneous platforms based on international standards and enables domain-free data management with flexible metadata.
Hee-Sun Won, Minh Chau Nguyen, Myeong-Seon Gil, Yang-Sae Moon
IEEE BigData4
2021 Efficient Deep Learning Models for DGA Domain Detection
abstract
In recent years, cyberattacks using command and control (C&C) servers have significantly increased. To hide their C&C servers, attackers often use a domain generation algorithm (DGA), which automatically generates domain names for the C&C servers. Accordingly, extensive research on DGA domain detection has been conducted. However, existing methods cannot accurately detect continuously generated DGA domains and can easily be evaded by an attacker. Recently, long short-term memory- (LSTM-) based deep learning models have been introduced to detect DGA domains in real time using only domain names without feature extraction or additional information. In this paper, we propose an efficient DGA domain detection method based on bidirectional LSTM (BiLSTM), which learns bidirectional information as opposed to unidirectional information learned by LSTM. We further maximize the detection performance with a convolutional neural network (CNN) + BiLSTM ensemble model using Attention mechanism, which allows the model to learn both local and global information in a domain sequence. Experimental results show that existing CNN and LSTM models achieved F1-scores of 0.9384 and 0.9597, respectively, while the proposed BiLSTM and ensemble models achieved higher F1-scores of 0.9618 and 0.9666, respectively. In addition, the ensemble model achieved the best performance for most DGA domain classes, enabling more accurate DGA domain detection than existing models.
Juhong Namgung, Siwoon Son, Yang-Sae Moon
Secur. Commun. Networks3
2021 Stochastic distributed data stream partitioning using task locality: design, implementation, and optimization
Siwoon Son, Hyeonseung Im, Yang-Sae Moon
J. Supercomput.3
2020 Experimental Comparison of Machine Learning Models in Malware Packing Detection
abstract
Recently , malware is widely distributed by combining recent technologies such as packing, encoding and obfuscation to bypass anti-virus software. These kinds of technologies allow malware to survive longer, infect various computers and devices for longer periods of time, create a number of mutated malware, and make experts spend longer to analyze malware. Packers disrupt the reverse engineering process, making it difficult for security researchers to analyze new or unknown malware. Thus, we need to analyze as many malware as possible by first detecting the packed malware and analyzing not-packed malware, and then unpack the packed malware. Previously, the packing detection methods were based on mainly signature and entropy detection. However, these methods have increased the undetected rate with the appearance of custom packers. Due to these problems, there have been many research efforts on machine learning-based malware packing detection and classification. In this paper, we present an extensive experimental comparison of these machine learning-based algorithms. In particular, we extract a total of 13 important features and considers eight machine learning algorithms to detect the packing of malware. Experimental results show that we can also detect well malware packed by custom packers which did not studied in previous studies.
Jong-Wouk Kim, Juhong Namgung, Yang-Sae Moon, Mi-Jung Choi
APNOMS3
2019 All-in-One Framework for Detection, Unpacking, and Verification for Malware Analysis
abstract
Packing is the most common analysis avoidance technique for hiding malware. Also, packing can make it harder for the security researcher to identify the behaviour of malware and increase the analysis time. In order to analyze the packed malware, we need to perform unpacking first to release the packing. In this paper, we focus on unpacking and its related technologies to analyze the packed malware. Through extensive analysis on previous unpacking studies, we pay attention to four important drawbacks: no phase integration, no detection combination, no real-restoration, and no unpacking verification. To resolve these four drawbacks, in this paper, we present an all-in-one structure of the unpacking system that performs packing detection, unpacking (i.e., restoration), and verification phases in an integrated framework. For this, we first greatly increase the packing detection accuracy in the detection phase by combining four existing and new packing detection techniques. We then improve the unpacking phase by using the state-of-the-art static and dynamic unpacking techniques. We also present a verification algorithm evaluating the accuracy of unpacking results. Experimental results show that the proposed all-in-one unpacking system performs all of the three phases well in an integrated framework. In particular, the proposed hybrid detection method is superior to the existing methods, and the system performs unpacking very well up to 100% of restoration accuracy for most of the files except for a few packers.
Mi-Jung Choi, Jiwon Bang, Jongwook Kim, Hajin Kim, Yang-Sae Moon
Secur. Commun. Networks5
2019 Prefetching-based metadata management in Advanced Multitenant Hadoop
Minh Chau Nguyen, Hee-Sun Won, Siwoon Son, Myeong-Seon Gil, Yang-Sae Moon
J. Supercomput.5
2019 Performance improvement of Apache Storm using InfiniBand RDMA
Seokwoo Yang, Siwoon Son, Mi-Jung Choi, Yang-Sae Moon
J. Supercomput.4
2018 Design and implementation of a load shedding engine for solving starvation problems in Apache Kafka
abstract
Real-time data stream processing technologies such as Apache Storm and Apache Spark are being actively studied to deal with large-capacity data streams that generated rapidly in real time. Because it is difficult to use most real-time processing techniques alone, it is common to use it with a messaging system that supports input and output of data streams. Apache Kafka is a representative distributed messaging system, specialized in delivering large amounts of real-time log data. However, if the production rate of data in Kafka is faster than the consumption rate, data starvation problem may arise. In order to solve the starvation problem, a load shedding technique is needed to limit the incoming data and maintain system performance when the system is under load. Thus, in this paper confirmed the starvation problem that can occur in Kafka, and we designed and implemented a load shedding engine to solve this problem and proposed a solution to the starvation problem in Kafka based on the performance experiment.
Jiwon Bang, Siwoon Son, Hajin Kim, Yang-Sae Moon, Mi-Jung Choi
NOMS4
2018 A time-series matching approach for symmetric-invariant boundary image matching
Hajin Kim, Mi-Jung Choi, Yang-Sae Moon
Multim. Tools Appl.4
2017 Feasibility study for simulating community based content caching on CCN network using ndnSIM simulator
abstract
In this paper we have done a feasibility study to develop a simulation model for Content Centric Network. Our goal is to form communities based on popular clustering algorithm on the simulated CCN network. Furthermore we aim to do a guided content delivery on CCN based on the community preferences. Though community formation and a guided delivery of content is not a new field. However an experimental simulation of these concepts on CCN network is novel. Basically we want to utilize CCN network and its advantages to implement a next generation CDN.
Suman Pandey, Yang-Sae Moon, Mi-Jung Choi
APNOMS2
2017 Boundary image matching supporting partial denoising using time-series matching techniques
Yang-Sae Moon, Jae-Gil Lee 0001
Multim. Tools Appl.2
2017 Efficient Two-Step Protocol and Its Discriminative Feature Selections in Secure Similar Document Detection
abstract
Recently, the risk of information disclosure is increasing significantly. Accordingly, privacy-preserving data mining (PPDM) is being actively studied to obtain accurate mining results while preserving the data privacy. We here focus on secure similar document detection (SSDD), which identifies similar documents of two parties when each party does not disclose its own sensitive documents to the another party. In this paper, we propose an efficient two-step protocol that exploits a feature selection as a lower-dimensional transformation, and we present discriminative feature selections to maximize the performance of the protocol. The proposed protocol consists of two steps: thefilteringstep and thepostprocessingstep. For the feature selection, we first consider the simplest one, random projection (RP), and propose its two-step solution,SSDD-RP. We then present two discriminative feature selections and their solutions:SSDD-LFwhich selects a few dimensions locally frequent in the current querying vector andSSDD-GFwhich selects ones globally frequent in the set of all document vectors. We finally propose a hybrid one,SSDD-HF, which takes advantage of bothSSDD-LFandSSDD-GF. We empirically show that the proposed two-step protocol significantly outperforms the previous one-step protocol by three or four orders of magnitude.
Sang-Pil Kim, Myeong-Seon Gil, Hajin Kim, Mi-Jung Choi, Yang-Sae Moon, Hee-Sun Won
Secur. Commun. Networks5
2017 Moving metadata from ad hoc files to database tables for robust, highly available, and scalable HDFS
Hee-Sun Won, Minh Chau Nguyen, Myeong-Seon Gil, Yang-Sae Moon, Kyu-Young Whang
J. Supercomput.4
2016 Secure principal component analysis in multiple distributed nodes
abstract
Abstract Privacy preservation becomes an important issue in recent big data analysis, and many secure multiparty computations have been proposed for the purpose of privacy preservation in the environment of distributed nodes. As a secure multiparty computations of principal component analysis (PCA), in this paper, we propose S‐PCA, which compute PCA securely among the distributed nodes. PCA is widely used in many applications including time‐series analysis, text mining, and image compression. In general, we compute PCA after concentrating all data in a single server, but this approach discloses data privacy of each node. In contrast, the proposed S‐PCA computes PCA without disclosing the sensitive data of individual nodes. In S‐PCA, the nodes share non‐sensitive mean vectors first and compute covariance matrices and PCA securely using the shared mean vectors. In this paper, we formally prove the correctness and secureness of S‐PCA and apply it to an application of secure similar document detection. Experimental results show that the performance of S‐PCA is slightly worse than that of PCA due to guarantee of secureness, but it significantly improves the performance of secure similar document detection by up to two orders of magnitudes. Copyright © 2016 John Wiley & Sons, Ltd.
Hee-Sun Won, Sang-Pil Kim, Mi-Jung Choi, Yang-Sae Moon
Secur. Commun. Networks5
2015 Envelope-based boundary image matching for smart devices under arbitrary rotations
Woong-Kee Loh, Sang-Pil Kim, Sun-Kyong Hong, Yang-Sae Moon
Multim. Syst.4
2015 Triangular inequality-based rotation-invariant boundary image matching for smart devices
Yang-Sae Moon, Woong-Kee Loh
Multim. Syst.1
2014 Safe MBR-transformation in similar sequence matching
Yang-Sae Moon, Byung Suk Lee 0001
Inf. Sci.1
2014 Interactive noise-controlled boundary image matching using the time-series moving average transform
Yang-Sae Moon, Mi-Jung Choi
Multim. Tools Appl.2
2012 Efficient Distributed Parallel Top-Down Computation of ROLAP Data Cube Using MapReduce
Suan Lee, Yang-Sae Moon, Wookey Lee
DaWaK3
2012 Horizontal Reduction: Instance-Level Dimensionality Reduction for Similarity Search in Large Document Databases
abstract
Dimensionality reduction is essential in text mining since the dimensionality of text documents could easily reach several tens of thousands. Most recent efforts on dimensionality reduction, however, are not adequate to large document databases due to lack of scalability. We hence propose a new type of simple but effective dimensionality reduction, called horizontal (dimensionality) reduction, for large document databases. Horizontal reduction converts each text document to a few bitmap vectors and provides tight lower bounds of inter-document distances using those bitmap vectors. Bitmap representation is very simple and extremely fast, and its instance-based nature makes it suitable for large and dynamic document databases. Using the proposed horizontal reduction, we develop an efficient k-nearest neighbor (k-NN) search algorithm for text mining such as classification and clustering, and we formally prove its correctness. The proposed algorithm decreases I/O and CPU overheads simultaneously since horizontal reduction (1) reduces the number of accesses to documents significantly by exploiting the bitmap-based lower bounds in filtering dissimilar documents at an early stage, and accordingly, (2) decreases the number of CPU-intensive computations for obtaining a real distance between high-dimensional document vectors. Extensive experimental results show that horizontal reduction improves the performance of the reduction (preprocessing) process by one to two orders of magnitude compared with existing reduction techniques, and our k-NN search algorithm significantly outperforms the existing ones by one to three orders of magnitude.
Min Soo Kim 0001, Kyu-Young Whang, Yang-Sae Moon
ICDE3
2011 Photo Cube: An Automatic Management and Search for Photos Using Mobile Smartphones
abstract
Recently new mobile devices such as cellular phones, smart phones, and digital cameras are popularly used to take photos. By the virtue of these convenient instruments, we can take many photos easily, but we suffer from the difficulty of managing and searching photos due to their large volume. This paper develops a mobile application software, called Photo Cube, which automatically extracts various metadata for photos (e.g., date/time, place/address, weather, personal event, etc.) by taking advantage of sensors and networking functions embedded in mobile smart phones like Android phones or iPhones. The metadata can be used to manage and to search photos. Using this Photo Cube, users will be able to classify, store, manage, and search a large number of photos easily, without specifying any information but just clicking the shutter in a camera. The Photo Cube system was implemented on smart phones using Google's Android.
Suan Lee, Ji-Seop Won, Yang-Sae Moon
DASC4
2011 An Envelope-Based Approach to Rotation-Invariant Boundary Image Matching
Sang-Pil Kim, Yang-Sae Moon, Sun-Kyong Hong
DaWaK2
2011 A new approach for processing ranked subsequence matching based on ranked union
abstract
Ranked subsequence matching finds top-k subsequences most similar to a given query sequence from data sequences. Recently, Han et al. [12] proposed a solution (referred to here as HLMJ) to this problem by using the concept of the minimum distance matching window pair (MDMWP) and a global priority queue. By using the concept of MDMWP, HLMJ can prune many unnecessary accesses to data subsequences using a lower bound distance. However, we notice that HLMJ may incur serious performance overhead for important types of queries. In this paper, we propose a novel systematic framework to solve this problem by viewing ranked subsequence matching as ranked union. Specifically, we propose a notion of the matching subsequence equivalence class (MSEQ) and a novel lower bound called the MSEQ-distance. To completely eliminate the performance problem of HLMJ, we also propose a cost-aware density-based scheduling technique, where we consider both the density and cost of the priority queue. Extensive experimental results with many real datasets show that the proposed algorithm outperforms HLMJ and the adapted PSM [22], a state-of-the-art index-based merge algorithm supporting non-monotonic distance functions, by up to two to three orders of magnitude, respectively.
Wook-Shin Han, Jinsoo Lee, Yang-Sae Moon, Seung-won Hwang, Hwanjo Yu
SIGMOD Conference3
2010 Publishing Time-Series Data under Preservation of Privacy and Distance Orders
Yang-Sae Moon, Hea-Suk Kim, Sang-Pil Kim, Elisa Bertino
DEXA (2)1
2010 Scaling-invariant boundary image matching using time-series matching techniques
Yang-Sae Moon, Min Soo Kim 0001, Kyu-Young Whang
Data Knowl. Eng.1
2010 Distortion-free predictive streaming time-series matching
Woong-Kee Loh, Yang-Sae Moon, Jaideep Srivastava
Inf. Sci.2
2009 Assessing the trustworthiness of location data based on provenance
abstract
Trustworthiness of location information about particular individuals is of particular interest in the areas of forensic science and epidemic control. In many cases, location information is not precise and may include fraudulent information. With the growth of mobile computing and positioning systems, e.g., GPS and cell phones, it has become possible to trace the location of moving objects. Such Systems provide us an opportunity to find out the true locations of individuals. In this paper, we present a model to compute trustworthiness of the location information of an individual based on different evidences from different sources. We also introduce a collusion attack that may bias the computation. Based on the analysis of the attack, we present the algorithm to detect and reduce the effect of collusion attacks. Our experimental results show the efficiency and effectiveness of our approach.
Chenyun Dai, Hyo-Sang Lim, Elisa Bertino, Yang-Sae Moon
GIS4
2008 Noise Control Boundary Image Matching Using Time-Series Moving Average Transform
Yang-Sae Moon
DEXA2
2008 SAMSTAR: An Automatic Tool for Generating Star Schemas from an Entity-Relationship Diagram
Il-Yeol Song, Ritu Khare, Suan Lee, Sang-Pil Kim, Yang-Sae Moon
ER7
2008 Similar sequence matching supporting variable-length and variable-tolerance continuous queries on time-series data stream
Hyo-Sang Lim, Kyu-Young Whang, Yang-Sae Moon
Inf. Sci.3
2007 An MBR-Safe Transform for High-Dimensional MBRs in Similar Sequence Matching
Yang-Sae Moon
DASFAA1
2007 Ranked Subsequence Matching in Time-Series Databases
Wook-Shin Han, Jinsoo Lee, Yang-Sae Moon
VLDB3
2007 Efficient moving average transform-based subsequence matching algorithms in time-series databases
Yang-Sae Moon
Inf. Sci.1
2006 An Efficient Algorithm for Computing Range-Groupby Queries
Young-Koo Lee, Woong-Kee Loh, Yang-Sae Moon, Kyu-Young Whang, Il-Yeol Song
DASFAA3
2006 A Single Index Approach for Time-Series Subsequence Matching That Supports Moving Average Transform of Arbitrary Order
Yang-Sae Moon
PAKDD1
2005 Parallel Consistency Maintenance of Materialized Views Using Referential Integrity Constraints in Data Warehouses
Yang-Sae Moon, Sooho Ok, Wookey Lee
DaWaK3
2005 A Formal Framework for Prefetching Based on the Type-Level Access Pattern in Object-Relational DBMSs
abstract
Prefetching is an effective method for minimizing the number of fetches between the client and the server in a database management system. In this paper, we formally define the notion of prefetching. We also formally propose new notions of the type-level access locality and type-level access pattern. The type-level access locality is a phenomenon that repetitive patterns exist in the attributes referenced. The type-level access pattern is a pattern of attributes that are referenced in accessing the objects. We then develop an efficient capturing and prefetching policy based on this formal framework. Existing prefetching methods are based on object-level or page-level access patterns, which consist of object-ids or page-ids of the objects accessed. However, the drawback of these methods is that they work only when exactly the same objects or pages are accessed repeatedly. In contrast, even though the same objects are not accessed repeatedly, our technique effectively prefetches objects if the same attributes are referenced repeatedly, i.e., if there is type-level access locality. Many navigational applications in object-relational database management systems (ORDBMSs) have type-level access locality. Therefore, our technique can be employed in ORDBMSs to effectively reduce the number of fetches, thereby significantly enhancing the performance. We also address issues in implementing the proposed algorithm. We have conducted extensive experiments in a prototype ORDBMS to show effectiveness of our algorithm. Experimental results using the 007 benchmark, a real GIS application, and an XML application show that our technique reduces the number of fetches by orders of magnitude and improves the elapsed time by several factors over on-demand fetching and context-based prefetching, which is a state-of-the-art prefetching method. These results indicate that our approach provides a new paradigm in prefetching that improves performance of navigational applications significantly and is a practical method that can be implemented in commercial ORDBMSs.
Wook-Shin Han, Kyu-Young Whang, Yang-Sae Moon
IEEE Trans. Knowl. Data Eng.3
2004 A top-down approach for density-based clustering using multidimensional indexes
Jae-Joon Hwang, Kyu-Young Whang, Yang-Sae Moon, Byung Suk Lee 0001
J. Syst. Softw.3
2003 PrefetchGuide: capturing navigational access patterns for prefetching in client/server object-oriented/object-relational DBMSs
Wook-Shin Han, Yang-Sae Moon, Kyu-Young Whang
Inf. Sci.2
2003 An aggregation algorithm using a multidimensional file in multidimensional OLAP
Young-Koo Lee, Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
Inf. Sci.3
2003 Dynamic Buffer Allocation in Video-on-Demand Systems
abstract
In video-on-demand (VOD) systems, as the size of the buer allocated to user requests increases, initial latency and mem-ory requirements increase. Hence, the buer size must be minimized. The existing static buer allocation scheme, however, determines the buer size based on the assumption that the system is in the fully loaded state. Thus, when the system is in a partially loaded state, the scheme allocates a buer larger than necessary to a user request. This paper proposes a dynamic buer allocation scheme that allocates to user requests buers of the minimum size in a partially loaded state as well as in the fully loaded state. The inherent diÆculty in determining the buer size in the dynamic buer allocation scheme is that the size of the buer currently be-ing allocated is dependent on the number of and the sizes of the buers to be allocated in the next service period. We solve this problem by the predict-and-enforce strategy, where we predict the number and the sizes of future buers based on inertia assumptions and enforce these assumptions at runtime. Any violation of these assumptions is resolved by deferring service to the violating new user request until the assumptions are satised. Since the size of the current buer is dependent on the sizes of the future buers, the size is represented by a recurrence equation. We provide a solution to this equation, which can be computed at the system initialization time for runtime eÆciency. We have performed extensive analysis and simulation. The results show that the dynamic buer allocation scheme reduces ini-tial latency (averaged over the number of user requests in service from one to the maximum capacity) to
Kyu-Young Whang, Yang-Sae Moon, Wook-Shin Han, Il-Yeol Song
IEEE Trans. Knowl. Data Eng.3
2002 General match: a subsequence matching method in time-series databases based on generalized windows
abstract
We generalize the method of constructing windows in subsequence matching. By this generalization, we can explain earlier subsequence matching methods as special cases of a common framework. Based on the generalization, we propose a new subsequence matching method, General Match. The earlier work by Faloutsos et al. (called FRM for convenience) causes a lot of false alarms due to lack of point-filtering effect. Dual Match, recently proposed as a dual approach of FRM, improves performance significantly over FRM by exploiting point filtering effect. However, it has the problem of having a smaller allowable window size---half that of FRM---given the minimum query length. A smaller window increases false alarms due to window size effect. General Match offers advantages of both methods: it can reduce window size effect by using large windows like FRM and, at the same time, can exploit point-filtering effect like Dual Match. General Match divides data sequences into generalized sliding windows (J-sliding windows) and the query sequence into generalized disjoint windows (J-disjoint windows). We formally prove that General Match is correct, i.e., it incurs no false dismissal. We then propose a method of estimating the optimal value of the sliding factor J that minimizes the number of page accesses. Experimental results for real stock data show that, for low selectivities (10-6∼10-4), General Match improves average performance by 117% over Dual Match and by 998% over FRM; for high selectivities (10-3∼10-1), by 45% over Dual Match and by 64% over FRM. The proposed generalization provides an excellent theoretical basis for understanding the underlying mechanisms of subsequence matching.
Yang-Sae Moon, Kyu-Young Whang, Wook-Shin Han
SIGMOD Conference1
2002 A One-Pass Aggregation Algorithm with the Optimal Buffer Size in Multidimensional OLAP
Young-Koo Lee, Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
VLDB3
2001 Prefetching Based on Type-Level Access Pattern in Object-Relational DBMSs
abstract
Prefetching is an effective method for minimizing the number of round-trips between the client and the server in database management systems. We propose new notions of the type-level access locality and the type-level access pattern. We also formally define the notions of capturing and prefetching to help understand the underlying mechanisms. We then develop an efficient prefetching policy based on these notions and the framework. The type-level access locality is a phenomenon that repetitive patterns exist in the attributes referenced. The type-level access pattern is a pattern of attributes that are referenced in accessing the objects. Existing prefetching methods are based on object-level or page-level access patterns, which consist of object-ids or page-ids of the objects accessed. However the drawback of these methods is that they work only when exactly the same objects or pages are accessed repeatedly. In contrast even though the same objects are not accessed repeatedly our technique effectively prefetches objects if the same attributes are referenced repeatedly, i.e., if there is type-level access locality. Many navigational applications in object-relational database management systems (ORDBMSs) have type-level access locality. Therefore, our technique can be employed in ORDBMSs to effectively reduce the number of round trips, thereby significantly enhancing the performance.
Wook-Shin Han, Yang-Sae Moon, Kyu-Young Whang, Il-Yeol Song
ICDE2
2001 Duality-Based Subsequence Matching in Time-Series Databases
abstract
The authors propose a subsequence matching method, Dual Match, which exploits duality in constructing windows and significantly improves performance. Dual Match divides data sequences into disjoint windows and the query sequence into sliding windows, and thus, is a dual approach of the one by C. Faloutsos et al. (1994), which divides data sequences into sliding windows and the query sequence into disjoint windows. We formally prove that our dual approach is correct, i.e., it incurs no false dismissal. We also prove that, given the minimum query length, there is a maximum bound of the window size to guarantee correctness of Dual Match and discuss the effect of the window size on performance. FRM causes a lot of false alarms by storing minimum bounding rectangles rather than individual points representing windows to avoid excessive storage space required for the index. Dual Match solves this problem by directly storing points, but without incurring excessive storage overhead. Experimental results show that, in most cases, Dual Match provides large improvement in both false alarms and performance over FRM, given the same amount of storage space. In particular, for low selectivities (less than 10/sup -4/), Dual Match significantly improves performance up to 430-fold. On the other hand, for high selectivities(more than 10/sup -2/), it shows a very minor degradation (less than 29%). For selectivities in between (10/sup -4//spl sim/10/sup -2/), Dual Match shows performance slightly better than that of FRM. Dual Match is also 4.10/spl sim/25.6 times faster than FRM in building indexes of approximately the same size. Overall, these results indicate that our approach provides a new paradigm in subsequence matching that improves performance significantly in large database applications.
Yang-Sae Moon, Kyu-Young Whang, Woong-Kee Loh
ICDE1
2001 Dynamic Buffer Allocation in Video-on-Demand Systems
abstract
In video-on-demand (VOD) systems, as the size of the buffer allocated to user requests increases, initial latency and memory requirements increase. Hence, the buffer size must be minimized. The existing static buffer allocation scheme, however, determines the buffer size based on the assumption that the system is in the fully loaded state. Thus, when the system is in a partially loaded state, the scheme allocates a buffer larger than necessary to a user request. This paper proposes a dynamic buffer allocation scheme that allocates to user requests buffers of the minimum size in a partially loaded state as well as in the fully loaded state. The inherent difficulty in determining the buffer size in the dynamic buffer allocation scheme is that the size of the buffer currently being allocated is dependent on the number of and the sizes of the buffers to be allocated in the next service period. We solve this problem by the predict-and-enforce strategy, where we predict the number and the sizes of future buffers based on inertia assumptions and enforce these assumptions at runtime. Any violation of these assumptions is resolved by deferring service to the violating new user request until the assumptions are satisfied. Since the size of the current buffer is dependent on the sizes of the future buffers, the size is represented by a recurrence equation. We provide a solution to this equation, which can be computed at the system initialization time for runtime efficiency. We have performed extensive analysis and simulation. The results show that the dynamic buffer allocation scheme reduces initial latency (averaged over the number of user requests in service from one to the maximum capacity) to 1 ÷ 29.4 ≁ 1 ÷ 11.0 of that for the static one and, by reducing the memory requirement, increases the number of concurrent user requests to 2.36 ∼ 3.25 times that of the static one when averaged over the amount of system memory available. These results demonstrate that the dynamic buffer allocation scheme significantly improves the performance and capacity of VOD systems.
Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
SIGMOD Conference3
2001 Efficient time-series subsequence matching using duality in constructing windows
Yang-Sae Moon, Kyu-Young Whang, Woong-Kee Loh
Inf. Syst.1
2001 DyBASe: A buffer allocation scheme for reducing average initial latency in video-on-demand systems
Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
Inf. Sci.3