Il-Yeol Song

dblp:s/IlYeolSong · DBLP profile ↗
← Back
103ranked-venue papers in the field
16as first author
7since 2021 · last 2025
0000-0001-7706-959XORCID · reported

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 57 (8 first)Information Retrieval & Web Search · 18 (2 first)Business Process & Enterprise Data · 15 (6 first)Data Mining & Knowledge Discovery · 7Knowledge Engineering, Semantic Web & Information Systems · 5Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 Four decades of data & knowledge engineering: A bibliometric analysis and topic evolution study (1985-2024)
Tatsawan Timakum, Soobin Lee, Min Song 0001, Il-Yeol Song
Data Knowl. Eng.5
2024 DOLAP: A 25 Year Journey Through Research Trends and Performance (Invited Paper)
Tatsawan Timakum, Soobin Lee, Haotian Hu, Il-Yeol Song, Min Song 0001
DOLAP4
2023 About JASIST special issue on "Data Science in the iField"
Yin Zhang 0007, Il-Yeol Song, Theresa Dirndorfer Anderson, Dan Wu 0003
J. Assoc. Inf. Sci. Technol.2
2023 Data science curriculum in the iField
abstract
Many disciplines, including the broad Field of Information (iField), have been offering Data Science (DS) programs. There have been significant efforts exploring an individual discipline's identity and unique contributions to the broader DS education landscape. To advance DS education in the iField, the iSchool Data Science Curriculum Committee (iDSCC) was formed and charged with building and recommending a DS education framework for iSchools. This paper reports on the research process and findings of a series of studies to address important questions: What is the iField identity in the multidisciplinary DS education landscape? What is the status of DS education in iField schools? What knowledge and skills should be included in the core curriculum for iField DS education? What are the jobs available for DS graduates from the iField? What are the differences between graduate-level and undergraduate-level DS education? Answers to these questions will not only distinguish an iField approach to DS education but also define critical components of DS curriculum. The results will inform individual DS programs in the iField to develop curriculum to support undergraduate and graduate DS education in their local context.
Yin Zhang 0007, Dan Wu 0003, Loni Hagen, Il-Yeol Song, Javed Mostafa, Sam Gyun Oh, Theresa Dirndorfer Anderson, Chirag Shah 0001, Bradley Wade Bishop, Frank Hopfgartner, Kai Eckert 0001, Lisa Federer, Jeffrey S. Saltz
J. Assoc. Inf. Sci. Technol.4
2022 Trends in Design, Optimization, Languages, and Analytical Processing of Big Data (DOLAP 2020)
Katja Hose, Oscar Romero 0001, Il-Yeol Song
Inf. Syst.3
2022 Special issue on DOLAP 2021: Design, Optimization, Languages and Analytical Processing of Big Data
Kostas Stefanidis, Patrick Marcel, Il-Yeol Song
Inf. Syst.3
2021 Preface
Il-Yeol Song
Data Knowl. Eng.1
2020 Analyzing the Research Landscape of DaWaK Papers from 1999 to 2019
Tatsawan Timakum, Soobin Lee, Il-Yeol Song, Min Song 0001
DaWaK3
2020 Guest Editorial - DaWaK 2019 Special Issue - Evolving Big Data Analytics Towards Data Science
Carlos Ordonez 0001, Il-Yeol Song
Data Knowl. Eng.2
2020 An Alternative View on Data Processing Pipelines from the DOLAP 2019 Perspective
Oscar Romero 0001, Robert Wrembel, Il-Yeol Song
Inf. Syst.3
2019 Special issue on DOLAP 2017: Design, Optimization, Languages and Analytical Processing of Big Data
Patrick Marcel, Il-Yeol Song
Inf. Syst.2
2019 DOLAP data warehouse research over two decades: Trends and challenges
Robert Wrembel, Alberto Abelló, Il-Yeol Song
Inf. Syst.3
2018 Guest editorial - Special issue on conceptual modeling - 35th International Conference on Conceptual Modeling (ER2016)
Isabelle Comyn-Wattiau, Il-Yeol Song, Katsumi Tanaka, Motoshi Saeki, Shuichiro Yamamoto
Data Knowl. Eng.2
2018 The landscape of smart aging: Topics, applications, and agenda
Il-Yeol Song, Min Song 0001, Tatsawan Timakum, Su-Ryeon Ryu, Hanju Lee
Data Knowl. Eng.1
2017 Big data technologies and Management: What conceptual modeling can do
Veda C. Storey, Il-Yeol Song
Data Knowl. Eng.2
2017 A natural language interface to a graph-based bibliographic information retrieval system
Yongjun Zhu 0001, Erjia Yan, Il-Yeol Song
Data Knowl. Eng.3
2017 Big Data Management: New Frontiers, New Paradigms
Alfredo Cuzzocrea, Alkis Simitsis, Il-Yeol Song
Inf. Syst.3
2017 Special issue on DOLAP 2015: Evolving data warehousing and OLAP cubes to big data analytics
Carlos Ordonez 0001, Carlos Garcia-Alvarado, Il-Yeol Song
Inf. Syst.3
2017 The use of a graph-based system to improve bibliographic information retrieval: System design, implementation, and evaluation
abstract
In this article, we propose a graph‐based interactive bibliographic information retrieval system—GIBIR. GIBIR provides an effective way to retrieve bibliographic information. The system represents bibliographic information as networks and provides a form‐based query interface. Users can develop their queries interactively by referencing the system‐generated graph queries. Complex queries such as “papers on information retrieval, which were cited by John's papers that had been presented in SIGIR” can be effectively answered by the system. We evaluate the proposed system by developing another relational database‐based bibliographic information retrieval system with the same interface and functions. Experiment results show that the proposed system executes the same queries much faster than the relational database‐based system, and on average, our system reduced the execution time by 72% (for 3‐node query), 89% (for 4‐node query), and 99% (for 5‐node query).
Yongjun Zhu 0001, Erjia Yan, Il-Yeol Song
J. Assoc. Inf. Sci. Technol.3
2015 DOLAP 2015 Workshop Summary
abstract
The ACM DOLAP workshop presents research that bridges data warehousing, On-Line Analytical Processing (OLAP), and other large-scale data processing platforms. The program has four interesting sessions on data warehouse design, database modeling, query processing, and text processing, as well as an invited paper on Big Data Database Design.
Carlos Garcia-Alvarado, Carlos Ordonez 0001, Il-Yeol Song
CIKM3
2015 Methodologies for Semi-automated Conceptual Data Modeling from Requirements
Il-Yeol Song, Yongjun Zhu 0001, Hyithaek Ceong, Ornsiri Thonggoom
ER1
2015 Advances in data warehousing and OLAP in the big Data Era
Ladjel Bellatreche, Alfredo Cuzzocrea, Il-Yeol Song
Inf. Syst.3
2014 Big Graph Analytics: The State of the Art and Future Research Agenda
abstract
Analytics over big graphs is becoming a first-class challenge in database research, with fast-growing interest from both the academia and the industrial community. This problem arises in several application scenarios, ranging from social networks to large-scale network systems, from knowledge discovery to cybersecurity, and so forth. Following this major trend, this paper explores actual state-of-the-art results in the area of analytics over big graphs and discusses open research issues and actual trends in such area.
Alfredo Cuzzocrea, Il-Yeol Song
DOLAP2
2014 Editorial
Matteo Golfarelli, Il-Yeol Song
Inf. Syst.2
2013 DOLAP 2013 workshop summary
abstract
The ACM DOLAP workshop presents research on data warehousing and On-Line Analytical Processing (OLAP). The DOLAP 2013 program has three interesting sessions on Design and Exploitation of Social Data Warehouses, ETL and modeling and new trends, as well as a keynote talk on OLAP query processing and a panel on OLAP and DataWarehousing Technology in Big Data era.
Ladjel Bellatreche, Alfredo Cuzzocrea, Il-Yeol Song
CIKM3
2013 Data warehousing and OLAP over big data: current challenges and future research directions
abstract
In this paper, we highlight open problems and actual research trends in the field of Data Warehousing and OLAP over Big Data, an emerging term in Data Warehousing and OLAP research. We also derive several novel research directions arising in this field, and put emphasis on possible contributions to be achieved by future research efforts.
Alfredo Cuzzocrea, Ladjel Bellatreche, Il-Yeol Song
DOLAP3
2013 ODYS: an approach to building a massively-parallel search engine using a DB-IR tightly-integrated parallel DBMS for higher-level functionality
abstract
Recently, parallel search engines have been implemented based on scalable distributed file systems such as Google File System. However, we claim that building a massively-parallel search engine using a parallel DBMS can be an attractive alternative since it supports a higher-level (i.e., SQL-level) interface than that of a distributed file system for easy and less error-prone application development while providing scalability. Regarding higher-level functionality, we can draw a parallel with the traditional O/S file system vs. DBMS. In this paper, we propose a new approach of building a massively-parallel search engine using a DB-IR tightly-integrated parallel DBMS. To estimate the performance, we propose a hybrid (i.e., analytic and experimental) performance model for the parallel search engine. We argue that the model can accurately estimate the performance of a massively-parallel (e.g., 300-node) search engine using the experimental results obtained from a small-scale (e.g., 5-node) one. We show that the estimation error between the model and the actual experiment is less than 2.13% by observing that the bulk of the query processing time is spent at the slave (vs. at the master and network) and by estimating the time spent at the slave based on actual measurement. Using our model, we demonstrate a commercial-level scalability and performance of our architecture. Our proposed system ODYS is capable of handling 1 billion queries per day (81 queries/sec) for 30 billion Web pages by using only 43,472 nodes with an average query response time of 194 ms. By using twice as many (86,944) nodes, ODYS can provide an average query response time of 148 ms. These results show that building a massively-parallel search engine using a parallel DBMS is a viable approach with advantages of supporting the high-level (i.e., DBMS-level), SQL-like programming interface.
Kyu-Young Whang, Tae-Seob Yun, Yeon-Mi Yeo, Il-Yeol Song, Hyukyoon Kwon, In-Joong Kim
SIGMOD Conference4
2013 CAM: A Conceptual Modeling Framework based on the Analysis of Entity Classes and Association Types
abstract
The problem of identifying relevant classes (entities) and associations (relationships) is a fundamental problem for conceptual modeling. In a previous work the authors introduced a conceptual modeling methodology named OMP (Ontology-based Modeling Patterns), which is based on the analysis of class categories representing entity types that are organized in the form of ontology. Since then the authors have explored a way to improve the methodology. As a result, in this paper the authors introduce a new conceptual modeling framework, entitled CAM (Class/Association-analysis-based Modeling), which is based on the analysis and classification of association types as well as entity types. The main objective of CAM is to serve as a tool to facilitate teaching the fundamentals of conceptual modeling to students in a systematic way, by providing extensible and adaptable entity/association classificatory systems that can be directly used in the problem-solving process. In this paper the authors present the CAM framework and illustrate its application.
Sofia J. Athenikos, Il-Yeol Song
J. Database Manag.2
2012 Learning to discover complex mappings from web forms to ontologies
abstract
In order to realize the Semantic Web, various structures on the Web including Web forms need to be annotated with and mapped to domain ontologies. We present a machine learning-based automatic approach for discovering complex mappings from Web forms to ontologies. A complex mapping associates a set of semantically related elements on a form to a set of semantically related elements in an ontology. Existing schema mapping solutions mainly rely on integrity constraints to infer complex schema mappings. However, it is difficult to extract rich integrity constraints from forms. We show how machine learning techniques can be used to automatically discover complex mappings between Web forms and ontologies. The challenge is how to capture and learn the complicated knowledge encoded in existing complex mappings. We develop an initial solution that takes a naive Bayesian approach. We evaluated the performance of the solution on various domains. Our experimental results show that the solution returns the expected mappings as the top-1 results usually among several hundreds candidate mappings for more than 80% of the test cases. Furthermore, the expected mappings are always returned as the top-k results with k<4. The experiments have demonstrated that the approach is effective and has the potential to save significant human efforts.
Xiaohua Hu 0001, Il-Yeol Song
CIKM3
2012 DOLAP 2012 workshop summary
abstract
The ACM DOLAP workshop presents research on data warehousing and On-Line Analytical Processing (OLAP). The DOLAP 2012 program is organized in four interesting sessions on data warehouse design and maintainability, OLAP querying and trends, warehousing of complex data, performance optimization and benchmarking.
Matteo Golfarelli, Il-Yeol Song
CIKM2
2012 Improving the maintainability of data warehouse designs: modeling relationships between sources and user concepts
abstract
In data warehouse (DW) development, a series of mappings must be specified between user concepts and data source elements, in order to identify which sources must undergo an integration process. Until now, these mappings are either assumed to be implied by name matching or identified according to the designer's experience. Then, the result is implemented as Extraction/Transformation/Loading (ETL) processes. Since ETL processes relate elements at the logical level, designers cannot adequately analyze how a change in requirements or in the data sources affects the analysis capabilities. Furthermore, this approach makes it difficult to perform incremental changes in DW design, requiring in some cases to perform the whole analysis again. In this paper we present a set of semantic mappings that relate user concepts specified by requirements to those obtained from data sources. In turn, this allows us to accurately identify how any potential change affects the different structures and ETL processes. As a DW evolves over time, our approach easily allows us to incorporate new concepts, as well as any change introduced at requirements or data sources into the DW repository with no need to redesign the whole DW. In order to show the application of our proposal, we show a real case study focusing on the Digital library of the University of Alicante.
Alejandro Maté, Juan Trujillo 0001, Elisa de Gregorio, Il-Yeol Song
DOLAP4
2012 Special Issue of DOLAP 2010 Information Systems
Carlos Ordonez 0001, Il-Yeol Song
Inf. Syst.2
2011 DOLAP 2011: overview of the 14th international workshop on data warehousing and olap
abstract
The ACM 14th International Workshop on Data Warehousing and OLAP (DOLAP 2011), held in Glasgow, Scotland, UK on October 28, 2011, in conjunction with the ACM 20th International Conference on Information and Knowledge Management (CIKM 2011), presents research on data warehousing and On-Line Analytical Processing (OLAP). The DOLAP 2011 program has three interesting sessions on data warehouse modeling and maintenance, ETL and performance, and OLAP visualization and extensions, and a panel discussing analytics in data warehouses.
Alfredo Cuzzocrea, Karen C. Davis, Il-Yeol Song
CIKM3
2011 Analytics over large-scale multidimensional data: the big data revolution!
abstract
In this paper, we provide an overview of state-of-the-art research issues and achievements in the field of analytics over big data, and we extend the discussion to analytics over big multidimensional data as well, by highlighting open problems and actual research trends. Our analytical contribution is finally completed by several novel research directions arising in this field, which plays a leading role in next-generation Data Warehousing and OLAP research.
Alfredo Cuzzocrea, Il-Yeol Song, Karen C. Davis
DOLAP2
2011 Automatically Mapping and Integrating Multiple Data Entry Forms into a Database
Ritu Khare, Il-Yeol Song, Xiaohua Hu 0001
ER3
2011 Semi-automatic Conceptual Data Modeling Using Entity and Relationship Instance Repositories
Ornsiri Thonggoom, Il-Yeol Song
ER2
2010 DOLAP 2010 workshop summary
abstract
The ACM DOLAP workshop presents research on data warehousing and On-Line Analytical Processing (OLAP). The program has three interesting sessions on modeling, query processing and new trends, as well as a keynote talk on OLAP query processing and a panel comparing relational and non-relational technology for data warehousing.
Carlos Ordonez 0001, Il-Yeol Song
CIKM2
2010 Relational versus non-relational database systems for data warehousing
abstract
Relational database systems have been the dominating technology to manage and analyze large data warehouses. Moreover, the ER model, the standard in database design has a close relationship with the relational model. Recently, there has been a surge of alternative technologies for large scale analytic processing, most of which are not based on the relational model. Out of these proposals, distributed file systems together with MapReduce have become strong competitors to relational database systems to analyze large data sets, exploiting parallel processing. Moreover, there is progress on using MapReduce to evaluate relational queries. With that motivation in mind, this panel will compare pros and cons of each technology for data warehousing and will identify research issues, considering practical aspects like ease of use, programming flexibility and cost; as well as technical aspects like data modeling, storage, hardware, scalability, query processing, fault tolerance and data mining.
Carlos Ordonez 0001, Il-Yeol Song, Carlos Garcia-Alvarado
DOLAP2
2010 Page-differential logging: an efficient and DBMS-independent approach for storing data into flash memory
abstract
Flash memory is widely used as the secondary storage in lightweight computing devices due to its outstanding advantages over magnetic disks. Flash memory has many access characteristics different from those of magnetic disks, and how to take advantage of them is becoming an important research issue. There are two existing approaches to storing data into flash memory: page-based and log-based. The former has good performance for read operations, but poor performance for write operations. In contrast, the latter has good performance for write operations when updates are light, but poor performance for read operations. In this paper, we propose a new method of storing data, called page-differential logging, for flash-based storage systems that solves the drawbacks of the two methods. The primary characteristics of our method are: (1) writing only the difference (which we define as the page-differential) between the original page in flash memory and the up-to-date page in memory; (2) computing and writing the page-differential only once at the time the page needs to be reflected into flash memory. The former contrasts with existing page-based methods that write the whole page including both changed and unchanged parts of data or from log-based ones that keep track of the history of all the changes in a page. Our method allows existing disk-based DBMSs to be reused as flash-based DBMSs just by modifying the flash memory driver, i.e., it is DBMS-independent. Experimental results show that the proposed method is superior in I/O performance, except for some special cases, to existing ones. Specifically, it improves the performance of various mixes of read-only and update operations by 0.5 (the special case when all transactions are read-only on updated pages) ~3.4 times over the page-based method and by 1.6 ~ 3.1 times over the log-based one for synthetic data of approximately 1 Gbytes. The TPC-C benchmark also shows improvement of the I/O time over existing methods by 1.2 ~ 6.1 times. This result indicates the effectiveness of our method under (semi) real workloads. We note that the performance advantage of our method can be further enhanced up to two folds by obviating the need to write to the spare area of the page a second time.
Yi-Reun Kim, Kyu-Young Whang, Il-Yeol Song
SIGMOD Conference3
2010 Data warehousing and OLAP (DOLAP'08)
Alberto Abelló, Il-Yeol Song
Data Knowl. Eng.2
2010 Maintaining Mappings between Conceptual Models and Relational Schemas
abstract
This paper describes a round-trip engineering approach for incrementally maintaining mappings between conceptual models and relational schemas. When either schema or conceptual model evolves to accommodate new information needs, the existing mapping must be maintained accordingly to continuously provide valid services. In this paper, the authors examine the mappings specifying “consistent” relationships between models. First, they define the consistency of a conceptual-relational mapping through “semantically compatible” instances. Next, the authors analyze the knowledge encoded in the standard database design process and develop round-trip algorithms for incrementally maintaining the consistency of conceptual-relational mappings under evolution. Finally, they conduct a set of comprehensive experiments. The results show that the proposed solution is efficient and provides significant benefits in comparison to the mapping reconstructing approach.
Xiaohua Hu 0001, Il-Yeol Song
J. Database Manag.3
2009 The partitioned-layer index: Answering monotone top-k queries using the convex skyline and partitioning-merging technique
Jun-Seok Heo, Kyu-Young Whang, Min-Soo Kim 0002, Yi-Reun Kim, Il-Yeol Song
Inf. Sci.5
2008 Round-Trip Engineering for Maintaining Conceptual-Relational Mappings
Xiaohua Hu 0001, Il-Yeol Song
CAiSE3
2008 Discovering Semantically Similar Associations (SeSA) for Complex Mappings between Conceptual Models
Il-Yeol Song
ER2
2008 A Multi-level Methodology for Developing UML Sequence Diagrams
Il-Yeol Song, Ritu Khare, Margaret Hilsbos
ER1
2008 SAMSTAR: An Automatic Tool for Generating Star Schemas from an Entity-Relationship Diagram
Il-Yeol Song, Ritu Khare, Suan Lee, Sang-Pil Kim, Yang-Sae Moon
ER1
2008 The thematic and citation landscape of Data and Knowledge Engineering
Chaomei Chen, Il-Yeol Song, Xiaojun Yuan 0001, Jian Zhang 0006
Data Knowl. Eng.2
2007 SAMSTAR: a semi-automated lexical method for generating star schemas from an entity-relationship diagram
abstract
The star schema is widely accepted as the de facto data model for data warehouse design. A popular approach for developing a star schema is to develop it from an entity-relationship diagram with some heuristics. Most of the existing approaches analyze the semantics of an ERD to generate a star schema. In this paper, we present the SAMSTAR method, which semi-automatically generates star schemas from an ERD by analyzing its semantics as well as structure. The novel features of SAMSTAR are (1) the use of the notion of Connection Topology Value (CTV) in identifying the candidates of facts and dimensions and (2) the use of Annotated Dimensional Design Patterns (A_DDP) as well as WordNet to extend the list of dimensions. We illustrate our method by applying it to the examples from existing literature. We prove that the outputs of our method are a superset of those of the existing methods. The SAMSTAR method simplifies the work of experienced designers and gives a smooth head-start to novices.
Il-Yeol Song, Ritu Khare, Bing Dai
DOLAP1
2007 Integration of association rules and ontologies for semantic query expansion
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen
Data Knowl. Eng.2
2007 The dynamic predicate: integrating access control with query processing in XML databases
Jae-Gil Lee 0001, Kyu-Young Whang, Wook-Shin Han, Il-Yeol Song
VLDB J.4
2006 An Efficient Algorithm for Computing Range-Groupby Queries
Young-Koo Lee, Woong-Kee Loh, Yang-Sae Moon, Kyu-Young Whang, Il-Yeol Song
DASFAA5
2006 Automatic Extraction for Creating a Lexical Repository of Abbreviations in the Biomedical Literature
Min Song 0001, Il-Yeol Song, Ki Jung Lee
DaWaK2
2006 A Coherent Biomedical Literature Clustering and Summarization Approach Through Ontology-Enriched Graphical Representations
Illhoi Yoo, Xiaohua Hu 0001, Il-Yeol Song
DaWaK3
2006 Integration of semantic-based bipartite graph representation and mutual refinement strategy for biomedical literature clustering
abstract
We introduce a novel document clustering approach that overcomes those problems by combining a semantic-based bipartite graph representation and a mutual refinement strategy. The primary contributions of this paper are the following. First, we introduce a new representation of documents using a bipartite graph between documents and co-occurrence concepts in the documents. Second, we show how to enhance clustering quality by applying the mutual refinement strategy to the initial clustering results. Third, through the experiments on MEDLINE documents, we show that our integrated method significantly enhances cluster quality and clustering reliability compared to existing clustering methods. Our approach improves on the average 29.5 cluster quality and 26.3 clustering reliability, in terms of misclassification index, over Bisecting K-means with the best parameters.
Illhoi Yoo, Xiaohua Hu 0001, Il-Yeol Song
KDD3
2006 Context-sensitive semantic smoothing for the language modeling approach to genomic IR
abstract
Semantic smoothing, which incorporates synonym and sense information into the language models, is effective and potentially significant to improve retrieval performance. The implemented semantic smoothing models, such as the translation model which statistically maps document terms to query terms, and a number of works that have followed have shown good experimental results. However, these models are unable to incorporate contextual information. Thus, the resulting translation might be mixed and fairly general. To overcome this limitation, we propose a novel context-sensitive semantic smoothing method that decomposes a document or a query into a set of weighted context-sensitive topic signatures and then translate those topic signatures into query terms. In detail, we solve this problem through (1) choosing concept pairs as topic signatures and adopting an ontology-based approach to extract concept pairs; (2) estimating the translation model for each topic signature using the EM algorithm; and (3) expanding document and query models based on topic signature translations. The new smoothing method is evaluated on TREC 2004/05 Genomics Track collections and significant improvements are obtained. The MAP (mean average precision) achieves a 33.6 % maximal gain over the simple language model, as well as a 7.8 % gain over the language model with context-insensitive semantic smoothing.
Xiaohua Zhou, Xiaohua Hu 0001, Xiaodan Zhang 0001, Xia Lin, Il-Yeol Song
SIGIR5
2006 Continuous query processing in data streams using duality of data and queries
abstract
Recent data stream systems such as TelegraphCQ have employed the well-known property of duality between data and queries. In these systems, query processing methods are classified into two dual categories -- data-initiative and query-initiative -- depending on whether query processing is initiated by selecting a data element or a query. Although the duality property has been widely recognized, previous data stream systems do not fully take advantages of this property since they use the two dual methods independently: data-initiative methods only for continuous queries and query-initiative methods only for ad-hoc queries. We contend that continuous query processing can be better optimized by adopting an approach that integrates the two dual methods. Our primary contribution is based on the observation that spatial join is a powerful tool for achieving this objective. In this paper, we first present a new viewpoint of transforming the continuous query processing problem to a multi-dimensional spatial join problem. We then present a continuous query processing algorithm based on spatial join, which we name Spatial Join CQ. This algorithm processes continuous queries by finding the pairs of overlapping regions from a set of data elements and a set of queries, both defined as regions in the multi-dimensional space. The algorithm achieves the advantages of the two dual methods simultaneously. Experimental results show that the proposed algorithm outperforms earlier algorithms by up to 36 times for simple selection continuous queries and by up to 7 times for sliding window join queries.
Hyo-Sang Lim, Jae-Gil Lee 0001, Min-Jae Lee 0002, Kyu-Young Whang, Il-Yeol Song
SIGMOD Conference5
2006 A UML profile for multidimensional modeling in data warehouses
Sergio Luján-Mora, Juan Trujillo 0001, Il-Yeol Song
Data Knowl. Eng.3
2006 Twenty Second International Conference on Conceptual Modeling (ER 2003)
Il-Yeol Song, Stephen W. Liddle, Tok Wang Ling
Data Knowl. Eng.1
2006 Data Warehouse Design to Support Customer Relationship Management Analysis
abstract
CRM is a strategy that integrates concepts of knowledge management, data mining, and data warehousing in order to support an organization’s decision-making process to retain long-term and profitable relationships with its customers. This research is part of a long-term study to examine systematically CRM factors that affect design decisions for CRM data warehouses in order to build a taxonomy of CRM analyses and to determine the impact of those analyses on CRM data warehousing design decisions. This article presents the design implications that CRM poses to data warehousing and then proposes a robust multidimensional starter model that supports CRM analyses. Additional research contributions include the introduction of two new measures, percent success ratio and CRM suitability ratio by which CRM models can be evaluated, the identification of and classification of CRM queries, and a preliminary heuristic for designing data warehouses to support CRM analyses.
Colleen Cunningham, Il-Yeol Song, Peter P. Chen
J. Database Manag.2
2006 Transform-Space View: Performing Spatial Join in the Transform Space Using Original-Space Indexes
abstract
Spatial joins find all pairs of objects that satisfy a given spatial relationship. In spatial joins using indexes, original-space indexes such as the R-tree are widely used. An original-space index is the one that indexes objects as represented in the original space. Since original-space indexes deal with extents of objects, it is relatively complex to optimize join algorithms using these indexes. On the other hand, transform-space indexes, which transform objects in the original space into points in the transform space and index them, deal only with points but no extents. Thus, optimization of join algorithms using these indexes can be relatively simple. However, the disadvantage of these join algorithms is that they cannot be applied to original-space indexes such as the R-tree. In this paper, we present a novel mechanism for achieving the best of these two types of algorithms. Specifically, we propose the new notion of the transform-space view and present the transform-space view join algorithm. The transform-space view is a virtual transform-space index based on an original-space index. It allows us to "interpret" or "view" an existing original-space index as a transform-space index with no space and negligible time overhead and without actually modifying the structure of the original-space index or changing object representation. The transform-space view join algorithm joins two original-space indexes in the transform space through the notion of the transform-space view. Through analysis and experiments, we verify the excellence of the transform-space view join algorithm. The transform-space view join algorithm always outperforms existing ones for all the data sets tested in terms of all three measures used: the one-pass buffer size (the minimum buffer size required for guaranteeing one disk access per page), the number of disk accesses for a given buffer size, and the wall clock time. Thus, it constitutes a lower-bound algorithm. We believe that the proposed transform-space view can be applied to developing various new spatial query processing algorithms in the transform space.
Min-Jae Lee 0002, Kyu-Young Whang, Wook-Shin Han, Il-Yeol Song
IEEE Trans. Knowl. Data Eng.4
2005 Mining undiscovered public knowledge from complementary and non-interactive biomedical literature through semantic pruning
abstract
Two complementary and non-interactive literature sets of articles, when they are considered together, can reveal useful information of scientific interest not apparent in either of the two document sets. Swanson called the existence of such knowledge, undiscovered public knowledge (UDPK). This paper proposes a semantic-based mining model for UDPK. Our method replaces manual ad-hoc pruning with using semantic knowledge from the biomedical ontologies. Using the semantic types and semantic relationships of the biomedical concepts, our prototype system can identify the relevant concepts collected from Medline and generate the novel hypothesis between these concepts. The system successfully replicates Swanson's two famous discoveries: Raynaud disease/fish oils and migraine/magnesium. Compared with previous approaches, our methods generate much fewer but more relevant novel hypotheses, and require much less human intervention in the discovery procedure.
Xiaohua Hu 0001, Illhoi Yoo, Min Song 0001, Yan-Qing Zhang 0001, Il-Yeol Song
CIKM5
2005 XML-OLAP: A Multidimensional Analysis Framework for XML Warehouses
Byung-Kwon Park, Hyoil Han, Il-Yeol Song
DaWaK3
2005 Semantic Query Expansion Combining Association Rules with Ontologies and Information Retrieval Techniques
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen
DaWaK2
2005 Dimensional modeling: identifying, classifying & applying patterns
abstract
Software design is a complex activity. A successful designer requires knowledge and training in specific design techniques combined with practical experience. Designing a dimensional model embodies this challenge. This paper presents Dimensional Design Patterns (DDPs) and their applications to the design of dimensional models. We describe a metamodel of the DDPs and show their integration into Kimball's dimensional modeling design process so they can be applied to design problems using a known practice. By providing a metamodel and a method for DDP use, we combine theory and a practical design technique with the goal of increasing the efficiency and effectiveness of the software designer. The initial experimental results regarding the classroom use of DDPs revealed a significant increase in the efficiency of students to design a dimensional model, but more testing is necessary in order to evaluate the effectiveness measure.
Mary Elizabeth Jones, Il-Yeol Song
DOLAP2
2005 A Taxonomy of Inaccurate Summaries and Their Management in OLAP Systems
John Horner, Il-Yeol Song
ER2
2005 An Automatic Unsupervised Querying Algorithm for Efficient Information Extraction in Biomedical Domain
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen
PAKDD2
2005 PIES: A Web Information Extraction System Using Ontology and Tag Patterns
Byung-Kwon Park, Hyoil Han, Il-Yeol Song
WAIM3
2004 Managing and Mining Clinical Outcomes
Hyoil Han, Il-Yeol Song, Xiaohua Hu 0001, Ann Prestrud, Murray F. Brennan, Ari D. Brooks
DASFAA2
2004 Data warehouse design to support customer relationship management analyses
abstract
CRM is a strategy that integrates the concepts of Knowledge Management, Data Mining, and Data Warehousing in order to support the organization's decision-making process to retain long-term and profitable relationships with its customers. In this paper, we first present the design implications that CRM poses to data warehousing, and then propose a robust multidimensional starter model that supports CRM analyses. We then present sample CRM queries, test our starter model using those queries and define two measures (% success ratio and CRM suitability ratio) by which CRM models can be evaluated. We finally introduce a preliminary heuristic for designing data warehouses to support CRM analyses. Our study shows that our starter model can be used to analyze various profitability analyses such as customer profitability analysis, market profitability analysis, product profitability analysis, and channel profitability analysis.
Colleen Cunningham, Il-Yeol Song, Peter P. Chen
DOLAP2
2004 An analysis of additivity in OLAP systems
abstract
Accurate summary data is of paramount concern in data warehouse systems; however, there have been few attempts to completely characterize the ability to summarize measures. The sum operator is the typical aggregate operator for summarizing the large amount of data in these systems. We look to uncover and characterize potentially inaccurate summaries resulting from aggregating measures using the sum operator. We discuss the effect of classification hierarchies, and non-, semi-, and fully- additive measures on summary data, and develop a taxonomy of the additive nature of measures. Additionally, averaging and rounding rules can add complexity to seemingly simple aggregations. To deal with these problems, we describe the importance of storing metadata that can be used to restrict potentially inaccurate aggregate queries. These summary constraints could be integrated into data warehouses, just as integrity constraints and are integrated into OLTP systems. We conclude by suggesting methods for identifying and dealing with non- and semi- additive attributes.
John Horner, Il-Yeol Song, Peter P. Chen
DOLAP2
2004 Use of Tabular Analysis Method to Construct UML Sequence Diagrams
Margaret Hilsbos, Il-Yeol Song
ER2
2004 Entity-Relationship Modeling Re-revisited
Don Goelman, Il-Yeol Song
ER2
2004 Ontology-Based Scalable and Portable Information Extraction System to Extract Biological Knowledge from Huge Collection of Biomedical Web Documents
abstract
Automated discovery and extraction of biological knowledge from biomedical web documents has become essential because of the enormous amount of biomedical literature published each year. In this paper we present an ontology-based scalable and portable information extraction system to automatically extract biological knowledge from huge collection of online biomedical web documents. Our method integrates ontology-based semantic tagging, information extraction and data mining together, automatically learns the patterns based on a few user seed tuples, and then extract new tuples from the biomedical web documents based on the discovered patterns. A novel system SPIE (Scalable and Portable Information Extraction) is implemented and tested on the PuBMed to find the chromatin protein-protein interaction and the experimental results indicate our approach is very effective in extracting biological knowledge from huge collection of biomedical web documents.
Xiaohua Hu 0001, Tsau Young Lin, Il-Yeol Song, Xia Lin, Illhoi Yoo, Mark Lechner, Min Song 0001
Web Intelligence3
2004 Applying UML and XML for designing and interchanging information for data warehouses and OLAP applications
abstract
Multidimensional (MD) modeling is the basis for data warehouses (DW), multidimensional databases (MDB) and on-line analytical processing (OLAP) applications. In this paper, we present how the unified modeling language (UML) can be successfully used to represent both structural and dynamic properties of these systems at the conceptual level. The structure of the system is specified by means of a UML class diagram that considers the main properties of MD modeling with minimal use of constraints and extensions of the UML. If the system to be modeled is too complex, thereby leading us to a considerable number of classes and relationships, we describe how to use the package grouping mechanism provided by the UML to simplify the final model. Furthermore, we provide a UML-compliant class notation (called cube class) to represent OLAP users’ initial requirements. We also describe how we can use the UML state and interaction diagrams to model the behavior of a data warehouse system. To facilitate the interchange of conceptual MD models, we provide a Document Type Definition (DTD) which allows us to represent the same MD modeling properties that can be considered by using our approach. From this DTD, we can directly generate valid eXtensible Markup Language (XML) documents that represent MD models at the conceptual level. We believe that our innovative approach provides a theoretical foundation for simplifying the conceptual design of MD systems and the examples included in this paper clearly illustrate the use of our approach.
Juan Trujillo 0001, Sergio Luján-Mora, Il-Yeol Song
J. Database Manag.3
2003 An analysis of structural validity in entity-relationship modeling
James Dullea, Il-Yeol Song, Ioanna Lamprou
Data Knowl. Eng.2
2003 An aggregation algorithm using a multidimensional file in multidimensional OLAP
Young-Koo Lee, Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
Inf. Sci.4
2003 Dynamic Buffer Allocation in Video-on-Demand Systems
abstract
In video-on-demand (VOD) systems, as the size of the buer allocated to user requests increases, initial latency and mem-ory requirements increase. Hence, the buer size must be minimized. The existing static buer allocation scheme, however, determines the buer size based on the assumption that the system is in the fully loaded state. Thus, when the system is in a partially loaded state, the scheme allocates a buer larger than necessary to a user request. This paper proposes a dynamic buer allocation scheme that allocates to user requests buers of the minimum size in a partially loaded state as well as in the fully loaded state. The inherent diÆculty in determining the buer size in the dynamic buer allocation scheme is that the size of the buer currently be-ing allocated is dependent on the number of and the sizes of the buers to be allocated in the next service period. We solve this problem by the predict-and-enforce strategy, where we predict the number and the sizes of future buers based on inertia assumptions and enforce these assumptions at runtime. Any violation of these assumptions is resolved by deferring service to the violating new user request until the assumptions are satised. Since the size of the current buer is dependent on the sizes of the future buers, the size is represented by a recurrence equation. We provide a solution to this equation, which can be computed at the system initialization time for runtime eÆciency. We have performed extensive analysis and simulation. The results show that the dynamic buer allocation scheme reduces ini-tial latency (averaged over the number of user requests in service from one to the maximum capacity) to
Kyu-Young Whang, Yang-Sae Moon, Wook-Shin Han, Il-Yeol Song
IEEE Trans. Knowl. Data Eng.5
2002 Multidimensional Modeling with UML Package Diagrams
Sergio Luján-Mora, Juan Trujillo 0001, Il-Yeol Song
ER3
2002 A One-Pass Aggregation Algorithm with the Optimal Buffer Size in Multidimensional OLAP
Young-Koo Lee, Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
VLDB4
2001 Developing Sequence Diagrams in UML
Il-Yeol Song
ER1
2001 Prefetching Based on Type-Level Access Pattern in Object-Relational DBMSs
abstract
Prefetching is an effective method for minimizing the number of round-trips between the client and the server in database management systems. We propose new notions of the type-level access locality and the type-level access pattern. We also formally define the notions of capturing and prefetching to help understand the underlying mechanisms. We then develop an efficient prefetching policy based on these notions and the framework. The type-level access locality is a phenomenon that repetitive patterns exist in the attributes referenced. The type-level access pattern is a pattern of attributes that are referenced in accessing the objects. Existing prefetching methods are based on object-level or page-level access patterns, which consist of object-ids or page-ids of the objects accessed. However the drawback of these methods is that they work only when exactly the same objects or pages are accessed repeatedly. In contrast even though the same objects are not accessed repeatedly our technique effectively prefetches objects if the same attributes are referenced repeatedly, i.e., if there is type-level access locality. Many navigational applications in object-relational database management systems (ORDBMSs) have type-level access locality. Therefore, our technique can be employed in ORDBMSs to effectively reduce the number of round trips, thereby significantly enhancing the performance.
Wook-Shin Han, Yang-Sae Moon, Kyu-Young Whang, Il-Yeol Song
ICDE4
2001 Dynamic Buffer Allocation in Video-on-Demand Systems
abstract
In video-on-demand (VOD) systems, as the size of the buffer allocated to user requests increases, initial latency and memory requirements increase. Hence, the buffer size must be minimized. The existing static buffer allocation scheme, however, determines the buffer size based on the assumption that the system is in the fully loaded state. Thus, when the system is in a partially loaded state, the scheme allocates a buffer larger than necessary to a user request. This paper proposes a dynamic buffer allocation scheme that allocates to user requests buffers of the minimum size in a partially loaded state as well as in the fully loaded state. The inherent difficulty in determining the buffer size in the dynamic buffer allocation scheme is that the size of the buffer currently being allocated is dependent on the number of and the sizes of the buffers to be allocated in the next service period. We solve this problem by the predict-and-enforce strategy, where we predict the number and the sizes of future buffers based on inertia assumptions and enforce these assumptions at runtime. Any violation of these assumptions is resolved by deferring service to the violating new user request until the assumptions are satisfied. Since the size of the current buffer is dependent on the sizes of the future buffers, the size is represented by a recurrence equation. We provide a solution to this equation, which can be computed at the system initialization time for runtime efficiency. We have performed extensive analysis and simulation. The results show that the dynamic buffer allocation scheme reduces initial latency (averaged over the number of user requests in service from one to the maximum capacity) to 1 ÷ 29.4 ≁ 1 ÷ 11.0 of that for the static one and, by reducing the memory requirement, increases the number of concurrent user requests to 2.36 ∼ 3.25 times that of the static one when averaged over the amount of system memory available. These results demonstrate that the dynamic buffer allocation scheme significantly improves the performance and capacity of VOD systems.
Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
SIGMOD Conference4
2001 DyBASe: A buffer allocation scheme for reducing average initial latency in video-on-demand systems
Kyu-Young Whang, Yang-Sae Moon, Il-Yeol Song
Inf. Sci.4
2001 A cost-based buffer replacement algorithm for object-oriented database systems
Chong-Mok Park, Kyu-Young Whang, Jeong-Joon Lee, Il-Yeol Song
Inf. Sci.4
2000 Binary Equivalents of Ternary Relationships in Entity-Relationship Modeling: A Logical Decomposition Approach
abstract
Little work has been completed which addresses the logical composition and use of ternary relationships in entity-relationship modeling. Many modeling notations and most CASE tools do not allow for ternary relationships. Alternative methods and substitutes for ternary relationship structures do not necessarily reflect the original logic, semantics or constraints of a given situation. Furthermore, it has been shown that ternary relationships can be constrained by additional implicit binary constraints which do not occur in the logic of binary relationships. This paper develops an analytical perspective of ternary relationships. We investigate the logical relationships implicit to the ternary structure and then identify potential simplification through decomposition into binary equivalents. These alternative binary equivalents allow retention of the implicit logical structure, and consequently also retain the semantics of the original structure. The analysis investigates equivalency of lossless decompositions, preservation of functional dependencies and finally the ability to preserve update constraints (insertions and deletions). We identify which ternary relationships have true, fully equivalent, binary equivalents and those which do not. We provide an exhaustive analysis of cardinality combinations found in ternary relationships which practitioners can use to guide the way in which they deal with ternary relationships in conceptual modeling.
Trevor H. Jones, Il-Yeol Song
J. Database Manag.2
1999 A Taxonomy of Recursive Relationships and Their Structural Validity in ER Modeling
James Dullea, Il-Yeol Song
ER2
1998 An Analysis of the Structural Validity of Ternary Relatinships in Entity Relationship Modeling
abstract
____________________________________________________________________________________________________
James Dullea, Il-Yeol Song
CIKM2
1998 A Framework for Object-Oriented On-line Analytical Processing
abstract
Article A framework for object-oriented on-line analytic processing Share on Authors: Jan W. Buzydlowski School of Information Science and Technology, Drexel University, Philadelphia, PA School of Information Science and Technology, Drexel University, Philadelphia, PAView Profile , Il-Yeol Song School of Information Science and Technology, Drexel University, Philadelphia, PA School of Information Science and Technology, Drexel University, Philadelphia, PAView Profile , Lewis Hassell School of Information Science and Technology, Drexel University, Philadelphia, PA School of Information Science and Technology, Drexel University, Philadelphia, PAView Profile Authors Info & Claims DOLAP '98: Proceedings of the 1st ACM international workshop on Data warehousing and OLAPNovember 1998 Pages 10–15https://doi.org/10.1145/294260.294264Online:01 November 1998Publication History 19citation790DownloadsMetricsTotal Citations19Total Downloads790Last 12 Months10Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jan W. Buzydlowski, Il-Yeol Song, Lewis Hassell
DOLAP2
1997 An Analysis of Cardinality Constraints in Redundant Relationships
James Dullea, Il-Yeol Song
CIKM2
1997 Intensional Query Processing Using Data Mining Approaches
abstract
Article Free Access Share on Intensional query processing using data mining approaches Authors: S. C. Yoon Dept. of Computer Science, Widener University, Chester, PA Dept. of Computer Science, Widener University, Chester, PAView Profile , I. Y. Song College of Information Studies, Drexel University, Philadelphia, PA College of Information Studies, Drexel University, Philadelphia, PAView Profile , E. K. Park Dept. of Software Architecture, University of Missouri, Kansas City, MO Dept. of Software Architecture, University of Missouri, Kansas City, MOView Profile Authors Info & Claims CIKM '97: Proceedings of the sixth international conference on Information and knowledge managementJanuary 1997 Pages 201–208https://doi.org/10.1145/266714.266896Published:01 January 1997Publication History 2citation383DownloadsMetricsTotal Citations2Total Downloads383Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Suk-Chung Yoon, Il-Yeol Song, E. K. Park
CIKM2
1997 A Region Splitting Strategy for Physical Database Design of Multidimensional File Organizations
Jong-Hak Lee, Young-Koo Lee, Kyu-Young Whang, Il-Yeol Song
VLDB4
1997 A Physical Database Design Method for Multidimensional File Organizations
Jong-Hak Lee, Young-Koo Lee, Kyu-Young Whang, Il-Yeol Song
Inf. Sci.4
1996 Analysis of Binary/Ternary Cardinality Combinations in Entity-Relationship Modeling
Trevor H. Jones, Il-Yeol Song
Data Knowl. Eng.2
1995 Semantic Query Processing in Object-Oriented Databases Using Deductive Approach
abstract
Article Semantic query processing in object-oriented databases using deductive approach Share on Authors: S. C. Yoon Dept. of Computer Science, Widener University, Chester, PA Dept. of Computer Science, Widener University, Chester, PAView Profile , I. Y. Song College of Information Studies, Drexel University, Philadelphia, PA College of Information Studies, Drexel University, Philadelphia, PAView Profile , E. K. Park Dept. of Computer Science, U.S. Naval Academy, Annapolis, MD Dept. of Computer Science, U.S. Naval Academy, Annapolis, MDView Profile Authors Info & Claims CIKM '95: Proceedings of the fourth international conference on Information and knowledge managementDecember 1995 Pages 150–157https://doi.org/10.1145/221270.221365Online:02 December 1995Publication History 12citation445DownloadsMetricsTotal Citations12Total Downloads445Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Suk-Chung Yoon, Il-Yeol Song, E. K. Park
CIKM2
1995 Ternary Relationship Decomposition Strategies Based on Binary Imposition Rules
abstract
We review a set of rules identifying which combinations of ternary and binary relationships can be combined simultaneously in semantically related situations. We investigate the effect of these rules on decomposing ternary relationships to simpler, multiple binary relationships. We also discuss the relevance of these decomposition strategies to ER modeling. We show that if at least one 1:1 or 1:M binary constraint can be identified within the construct of the ternary itself, then any ternary relationship can be decomposed to a binary format. From this methodology we construct a heuristic-the Constrained Ternary Decomposition (CTD) rule.>
Il-Yeol Song, Trevor H. Jones
ICDE1
1994 Intelligent Query Answering in Deductive and Object-Oriented Databases
abstract
In the near future, we believe that we will need much more sophisticated answer-finding schemes in an object-oriented database in order to satisfy the needs of truly intelligent information system. In this paper, we introduce a method to apply the intensional query processing techniques of deductive databases to object-oriented databases. So, we can generate intensional answers to represent answer-set abstractly for a given query in object-oriented databases.
Suk-Chung Yoon, Il-Yeol Song, E. K. Park
CIKM2
1993 Binary Relationship Imposition Rules on Ternary Relationships
Il-Yeol Song, Trevor H. Jones
CIKM1
1993 A Knowledge Based System Converting ER Model into an Object-Oriented Database Schema
Il-Yeol Song, Heather M. Godsey
DASFAA1
1993 GemCode: An Expert System Generating Mnemonic Codes for Data Elements and Data Items
Il-Yeol Song, Heather M. Godsey, Judith Newton, Bruce Bargmeyer
DEXA1
1993 Analysis of Binary Relationships within Ternary Relationships in ER Modeling
Il-Yeol Song, Trevor H. Jones
ER1
1992 Object-Oriented Database Design Methodologies: A Survey
Il-Yeol Song, E. K. Park
CIKM1
1991 Schema Conversion Rules Between EER and NIAM Models
Il-Yeol Song, Edward A. Forbes
ER1
1990 Intensional Query Processing: A Three-Step Approach
Il-Yeol Song, Petra Geutner
DEXA1