Lipyeow Lim

dblp:55/2304 · DBLP profile ↗
← Back
36ranked-venue papers
13as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 30 · 13 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 2 first-authorArtificial intelligence and machine learning · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
14 papers
Query processing and optimization · 24% Data integration and cleaning · 23% Information retrieval · 20%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
link discovery
0.222009
Linkage Query Writer · Proc. VLDB Endow. 2009
A declarative framework for semantic link discovery over relational data · WWW 2009
Query processing and optimization
selectivity estimation
0.132005
CXHist : An On-line Classification-Based Histogram for XML String Selectivity Estimation · VLDB 2005
SASH: A Self-Adaptive Histogram Set for Dynamically Changing Workloads · VLDB 2003
XPathLearner: An On-line Self-Tuning Markov Histogram for XML Path Selectivity Estimation · VLDB 2002
Information retrieval › web search
data freshness
0.112010
Optimizing content freshness of relations extracted from the web using keyword search · SIGMOD Conference 2010
Information retrieval › indexing
index compression
0.112009
Efficient Index Compression in DB2 LUW · Proc. VLDB Endow. 2009
Data integration and cleaning › link discovery
semantic link discovery
0.112009
A declarative framework for semantic link discovery over relational data · WWW 2009
Query processing and optimization › OLAP
OLAP query optimization
0.112008
Optimizing Hierarchical Access in OLAP Environment · ICDE 2008
Query processing and optimization
query rewriting
0.112008
Optimizing Hierarchical Access in OLAP Environment · ICDE 2008
Knowledge graphs
ontology
0.112007
Semantic Data Management: Towards Querying Data with their Meaning · ICDE 2007
Database theory
ontology-mediated queries
0.112007
Semantic Data Management: Towards Querying Data with their Meaning · ICDE 2007
Database system architecture and tuning › database design
physical database design
0.112007
Schema advisor for hybrid relational-XML DBMS · SIGMOD Conference 2007
Data models and query languages › schema management
schema evolution
0.112007
Preserving XML queries during schema evolution · WWW 2007
Knowledge graphs › semantic web
semantic data management
0.112007
Semantic Data Management: Towards Querying Data with their Meaning · ICDE 2007
Data models and query languages › XML data management
XML data model
0.112007
Schema advisor for hybrid relational-XML DBMS · SIGMOD Conference 2007
Query processing and optimization
XML query processing
0.112007
Preserving XML queries during schema evolution · WWW 2007
Data models and query languages › schema management › schema evolution
XML schema evolution
0.112007
Preserving XML queries during schema evolution · WWW 2007
Query processing and optimization › cardinality estimation
histogram
0.012003
SASH: A Self-Adaptive Histogram Set for Dynamically Changing Workloads · VLDB 2003
Information retrieval › search engines › web crawling
incremental crawling
0.012003
Dynamic maintenance of web indexes using landmarks · WWW 2003
Information retrieval › indexing
search engine indexing
0.012003
Dynamic maintenance of web indexes using landmarks · WWW 2003
Information retrieval › search engines
web crawling
0.012003
Dynamic maintenance of web indexes using landmarks · WWW 2003
Storage systems
indexing and storage engines
0.012003
Dynamic maintenance of web indexes using landmarks · WWW 2003
Information retrieval
keyword search
0.012010
Optimizing content freshness of relations extracted from the web using keyword search · SIGMOD Conference 2010
Information retrieval
search interfaces
0.012010
Optimizing content freshness of relations extracted from the web using keyword search · SIGMOD Conference 2010
Data models and query languages
relational data
0.012009
A declarative framework for semantic link discovery over relational data · WWW 2009
Data models and query languages › SQL
SQL query generation
0.012009
Linkage Query Writer · Proc. VLDB Endow. 2009
Data models and query languages
XML query languages
0.022005
CXHist : An On-line Classification-Based Histogram for XML String Selectivity Estimation · VLDB 2005
XPathLearner: An On-line Self-Tuning Markov Histogram for XML Path Selectivity Estimation · VLDB 2002
Data models and query languages
fuzzy query
0.012007
Supporting ranking and clustering as generalized order-by and group-by · SIGMOD Conference 2007
Database system architecture and tuning › self-managing database systems
workload adaptation
0.012003
SASH: A Self-Adaptive Histogram Set for Dynamically Changing Workloads · VLDB 2003

Methods — techniques the papers use, named apart from their topics

web communication optimization · 0.1query selectivity estimation · 0.1query translation · 0.1declarative specification · 0.1query rewriting · 0.1hierarchy overlap detection · 0.1schema recommendation · 0.1ontology-based semantic matching · 0.1XQuery · 0.1SQL/XML · 0.1landmark-diff algorithm · 0.0
YearPublicationVenuePosition
2019 An Experimental Survey of Evaluation Strategies for Constellation Queries
abstract
Given a set of query points within an image coordinate system, constellation queries identify the matching points in a database of known points within a standard coordinate system. Constellation queries are an integral part of orientation determination systems used in spacecrafts to orient and navigate themselves. The query points are bright spots in an image captured by a camera on the spacecraft and the database contains known celestial objects in a celestial coordinate system. This paper studies six existing constellation query processing strategies (Angle, Interior Angle, Spherical Triangle, Planar Triangle, Pyramid, Composite Pyramid) using a unified algorithmic framework and presents experimental evaluation of the six strategies. We find that the Pyramid strategy in its simplified form has the best accuracy to runtime ratio given simulated images with false positives, false negatives, and Gaussian noise.
Glenn Galvizo, Lipyeow Lim
SSDBM2
2018 What's That Plant? WTPlant is a Deep Learning System to Identify Plants in Natural Images
Jonas Krause, Gavin Sugita, Kyungim Baek, Lipyeow Lim
BMVC4
2018 WTPlant (What's That Plant?): A Deep Learning System for Identifying Plants in Natural Images
abstract
Despite the availability of dozens of plant identification mobile applications, identifying plants from a natural image remains a challenging problem - most of the existing applications do not address the complexity of natural images, the large number of plant species, and the multi-scale nature of natural images. In this technical demonstration, we present the WTPlant system for identifying plants in natural images. WTPlant is based on deep learning approaches. Specifically, it uses stacked Convolutional Neural Networks for image segmentation, a novel preprocessing stage for multi-scale analyses, and deep convolutional networks to extract the most discriminative features. WTPlant employs different classification architectures for plants and flowers, thus enabling plant identification throughout all the seasons. The user interface also shows, in an interactive way, the most representative areas in the image that are used to predict each plant species. The first version of WTPlant is trained to classify 100 different plant species present in the campus of the University of Hawai'i at Manoa. First experiments support the hypothesis that an initial segmentation process helps guide the extraction of representative samples and, consequently, enables Convolutional Neural Networks to better recognize objects of different scales in natural images. Future versions aim to extend the recognizable species to cover the land-based flora of the Hawaiian Islands.
Jonas Krause, Gavin Sugita, Kyungim Baek, Lipyeow Lim
ICMR4
2017 Cloud-based query evaluation for energy-efficient mobile sensing
Tianli Mo, Lipyeow Lim, Sougata Sen, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001
Pervasive Mob. Comput.2
2015 Probabilistic Models for One-Day Ahead Solar Irradiance Forecasting in Renewable Energy Applications
abstract
Solar irradiance forecasting is an important problem in renewable energy management where any dips in solar energy generation must be made up for by reserves in order to ensure an uninterrupted energy supply. In this paper, we study several data mining methods for short term solar irradiance forecasting at a given location. In particular, we apply linear regression, probabilistic models, and naive Bayes classifier to forecast solar irradiance one day ahead, i.e., we forecast what tomorrow's solar irradiance will be like at sundown today. We evaluate the forecasting performance of our adaptations of the three models using land-based weather data from several weather stations on the island of Oahu in Hawai'i.
Carlos V. Paradis, Lipyeow Lim, Duane Stevens, Dora Nakafuji
ICMLA2
2014 Cost-Optimal Execution of Boolean Query Trees with Shared Streams
abstract
The processing of queries expressed as trees of boolean operators applied to predicates on sensor data streams has several applications in mobile computing. Sensor data must be retrieved from the sensors, which incurs a cost, e.g., an energy expense that depletes the battery of a mobile query processing device. The objective is to determine the order in which predicates should be evaluated so as to shortcut part of the query evaluation and minimize the expected cost. This problem has been studied assuming that each data stream occurs at a single predicate. In this work we remove this assumption since it does not necessarily hold in practice. Our main results are an optimal algorithm for single-level trees and a proof of NP-completeness for DNF trees. For DNF trees, however, we show that there is an optimal predicate evaluation order that corresponds to a depth-first traversal. This result provides inspiration for a class of heuristics. We show that one of these heuristics largely outperforms other sensible heuristics, including a heuristic proposed in previous work.
Henri Casanova, Lipyeow Lim, Yves Robert, Frédéric Vivien, Dounia Zaidouni
IPDPS2
2014 Cloud-Based Query Evaluation for Energy-Efficient Mobile Sensing
abstract
In this paper, we reduce the energy overheads of continuous mobile sensing for context-aware applications that are interested in collective context or events. We propose a cloud-based query management and optimization framework, called CloQue, which can support concurrent queries, executing over thousands of individual smartphones. CloQue exploits correlation across context of different users to reduce energy overheads via two key innovations: i) Dynamically reordering the order of predicate processing to preferentially select predicates with not just lower sensing cost and higher selectivity, but that maximally reduce the uncertainty about other context predicates, and ii) intelligently propagating the query evaluation results to dynamically update the uncertainty of other correlated, but yet-to-be evaluated, context predicates. An evaluation, using real cell phone traces from a real world dataset shows significant energy savings (between 30 to 50% compared with traditional short-circuit systems) with little loss in accuracy (5% at most).
Tianli Mo, Sougata Sen, Lipyeow Lim, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001
MDM (1)3
2013 Elastic data partitioning for cloud-based SQL processing systems
abstract
One of the key advantages of cloud computing is the elasticity in which computing resources such as virtual machines can be increased or decreased. Current state-of-the-art shared-nothing parallel SQL processing systems, on the other hand, are often designed and optimized for a fixed number of database nodes. To take advantage of the elasticity afforded by cloud computing, cloud-based SQL processing systems need the ability to repartition the data easily when the number of database nodes is scaled up or down. In this paper, we investigate the problem of supporting elastic partitioning of data in cloud-based parallel SQL processing systems. We propose several algorithms and associated data organization techniques that minimizes the re-partitioning of tuples and the movement of data between nodes. Our experimental evaluation demonstrates the effectiveness of the proposed methods.
Lipyeow Lim
IEEE BigData1
2013 Semantic queries by example
abstract
With the ever increasing quantities of electronic data, there is a growing need to make sense out of the data. Many advanced database applications are beginning to support this need by integrating domain knowledge encoded as ontologies into queries over relational data. However, it is extremely difficult to express queries against graph structured ontology in the relational SQL query language or its extensions. Moreover, semantic queries are usually not precise, especially when data and its related ontology are complicated. Users often only have a vague notion of their information needs and are not able to specify queries precisely. In this paper, we address these challenges by introducing a novel method to support semantic queries in relational databases with ease. Instead of casting ontology into relational form and creating new language constructs to express such queries, we ask the user to provide a small number of examples that satisfy the query she has in mind. Using those examples as seeds, the system infers the exact query automatically, and the user is therefore shielded from the complexity of interfacing with the ontology. Our approach consists of three steps. In the first step, the user provides several examples that satisfy the query. In the second step, we use machine learning techniques to mine the semantics of the query from the given examples and related ontologies. Finally, we apply the query semantics on the data to generate the full query result. We also implement an optional active learning mechanism to find the query semantics accurately and quickly. Our experiments validate the effectiveness of our approach.
Lipyeow Lim, Haixun Wang, Min Wang 0001
EDBT1
2013 Query-aware compression of join results
abstract
Client-server database query processing has become an important paradigm in many data processing applications today. In cloud-based data services, for example, queries over structured data are sent to cloud-based servers for processing and the results relayed back to the client devices. Network bandwidth between client devices and cloud-based servers is often a limited resource and the use of data compression to reduce the amount of query result data transmitted would not only conserve bandwidth but also help with battery lifetime in the case of mobile client devices. For query result compression, current data compression methods do not exploit redundancy information that can be inferred from the query structure itself for greater compression. In this paper we propose a novel query-aware compression method for compressing query results sent from database servers to client applications. Our method is based on two key ideas. We exploit redundancy information obtained from the query plan and possibly from the database schema to achieve better compression than standard non-query aware compressors. We use a collection of memory-limited dictionaries to encode attribute values in a lightweight and efficient manner. Each dictionary in the collection of dictionaries are also dynamically resized to adapt to changing temporal access characteristics. We evaluated our method empirically using the TPC-H benchmark show that this technique is effective especially when used in conjunction with standard compressors. Our results show that compression ratios of up to twice that of gzip are possible.
Christopher M. Mullins, Lipyeow Lim, Christian A. Lang
EDBT2
2013 Energy-Efficient Collaborative Query Processing Framework for Mobile Sensing Services
abstract
Many emerging context-aware mobile applications involve the execution of continuous queries over sensor data streams generated by a variety of on-board sensors on multiple personal mobile devices (aka smartphones). To reduce the energy-overheads of such large-scale, continuous mobile sensing and query processing, this paper introduces CQP, a collaborative query processing framework that exploits the overlap (in both the sensor sources and the query predicates) across multiple smartphones. The framework automatically identifies the shareable parts of multiple executing queries, and then reduces the overheads of repetitive execution and data transmissions, by having a set of `leader' mobile nodes execute and disseminate these shareable partial results. To further reduce energy, CQP utilizes lower-energy short-range wireless links (such as Bluetooth) to disseminate such results directly among proximate smartphones. We describe algorithms to support our server-assisted distributed query sharing and optimization strategy. Simulation experiments indicate that this approach can result in 60% reduction in the energy overhead of continuous query processing; when `leader' selection is dynamically rotated to equitably share the burden, we observe an increase of up to 65% in operational lifetime.
Jin Yang 0001, Tianli Mo, Lipyeow Lim, Kai-Uwe Sattler, Archan Misra
MDM (1)3
2013 Adaptive data acquisition strategies for energy-efficient, smartphone-based, continuous processing of sensor streams
Lipyeow Lim, Archan Misra, Tianli Mo
Distributed Parallel Databases1
2011 Optimizing Sensor Data Acquisition for Energy-Efficient Smartphone-Based Continuous Event Processing
abstract
Many pervasive applications, such as activity recognition or remote wellness monitoring, utilize a personal mobile device (aka smart phone) to perform continuous processing of data streams acquired from locally-connected, wearable, sensors. To ensure the continuous operation of such applications on a battery-limited mobile device, it is essential to dramatically reduce the energy overhead associated with the process of sensor data acquisition and processing. To achieve this goal, this paper introduces a technique of 'acquisition-cost' aware continuous query processing, as part of the Acquisition Cost-Aware Query Adaptation (ACQUA) framework. ACQUA replaces the current paradigm, where the data is typically streamed (pushed) from the sensors to the smart phone, with a pull-based asynchronous model, where the phone retrieves appropriate blocks of sensor data from individual sensors, only when the stream elements are judged to be relevant to the query being processed. We describe algorithms that dynamically optimize the sequence (for complex stream queries with conjunctive and disjunctive predicates) in which such sensor data streams are retrieved by the phone, based on a combination of the communication cost and selectivity properties of individual sensor streams. Simulation experiments indicate that this approach can result in 70% reduction in the energy overhead of continuous query processing, without affecting the fidelity of the processing logic.
Archan Misra, Lipyeow Lim
Mobile Data Management (1)2
2010 Statistics-based parallelization of XPath queries in shared memory systems
abstract
The wide availability of commodity multi-core systems presents an opportunity to address the latency issues that have plaqued XML query processing. However, simply executing multiple XML queries over multiple cores merely addresses the throughput issue: intra-query parallelization is needed to exploit multiple processing cores for better latency. Toward this effort, this paper investigates the parallelization of individual XPath queries over shared-address space multi-core processors. Much previous work on parallelizing XPath in a distributed setting failed to exploit the shared memory parallelism of multi-core systems. We propose a novel, end-to-end parallelization framework that determines the optimal way of parallelizing an XML query. This decision is based on a statistics-based approach that relies both on the query specifics and the data statistics. At each stage of the parallelization process, we evaluate three alternative approaches, namely, data-, query-, and hybrid-partitioning. For a given XPath query, our parallelization algorithm uses XML statistics to estimate the relative efficiencies of these different alternatives and find an optimal parallel XPath processing plan. Our experiments using well-known XML documents validate our parallel cost model and optimization framework, and demonstrate that it is possible to accelerate XPath processing using commodity multi-core systems.
Rajesh Bordawekar, Lipyeow Lim, Anastasios Kementsietsidis, Bryant Wei-Lun Kok
EDBT2
2010 Optimizing content freshness of relations extracted from the web using keyword search
abstract
An increasing number of applications operate on data obtained from the Web. These applications typically maintain local copies of the web data to avoid network latency in data accesses. As the data on the Web evolves, it is critical that the local copy be kept up-to-date. Data freshness is one of the most important data quality issues, and has been extensively studied for various applications including web crawling. However, web crawling is focused on obtaining as many raw web pages as possible. Our applications, on the other hand, are interested in specific content from specific data sources. Knowing the content or the semantics of the data enables us to differentiate data items based on their importance and volatility, which are key factors that impact the design of the data synchronization strategy. In this work, we formulate the concept of content freshness, and present a novel approach that maintains content freshness with least amount of web communication. Specifically, we assume data is accessible through a general keyword search interface, and we form keyword queries based on their selectivity, as well their contribution to content freshness of the local copy. Experiments show the effectiveness of our approach compared with several naive methods for keeping data fresh.
Mohan Yang, Haixun Wang, Lipyeow Lim, Min Wang 0001
SIGMOD Conference3
2009 Profile-based Retrieval of Records in Medical Databases
Anastasios Kementsietsidis, Lipyeow Lim, Min Wang 0001
AMIA2
2009 A framework for semantic link discovery over relational data
abstract
Discovering links between different data items in a single data source or across different data sources is a challenging problem faced by many information systems today. In particular, the recent Linking Open Data (LOD) community project has highlighted the paramount importance of establishing semantic links among web data sources. Currently, LOD sources provide billions of RDF triples, but only millions of links between data sources. Many of these data sources are published using tools that operate over relational data stored in a standard RDBMS. In this paper, we present a framework for discovery of semantic links from relational data. Our framework is based on declarative specification of linkage requirements by a user. We illustrate the use of our framework using several link discovery algorithms on a real world scenario. Our framework allows data publishers to easily find and publish high-quality links to other data sources, and therefore could significantly enhance the value of the data in the next generation of web.
Oktie Hassanzadeh, Anastasios Kementsietsidis, Lipyeow Lim, Renée J. Miller, Min Wang 0001
CIKM3
2009 Semantic queries in databases: problems and challenges
abstract
Supporting semantic queries in relational databases is essential to many advanced applications. Recently, with the increasing use of ontology in various applications, the need for querying relational data together with its related ontology has become more urgent. In this paper, we identify and discuss the problem of querying relational data with its ontologies. Two fundamental challenges make the problem interesting. First, it is extremely difficult to express queries against graph structured ontology in the relational query language SQL, and second, in many cases where data and its related ontology are complicated, queries are usually not precise, that is, users often have only a vague notion, rather than a clear understanding and definition, of what they query for. We outline a query-by-example approach that enables us to support semantic queries in relational databases with ease. Instead of endeavoring to incorporate ontology into relational form and create new language constructs to express such queries, we ask the user to provide a small number of examples that satisfy the query she has in mind. Using these examples as seeds, the system infers the exact query automatically, and the user is therefore shielded from the complexity of interfacing with the ontology.
Lipyeow Lim, Haixun Wang, Min Wang 0001
CIKM1
2009 Parallelization of XPath queries using multi-core processors: challenges and experiences
abstract
In this study, we present experiences of parallelizing XPath queries using the Xalan XPath engine on shared-address space multi-core systems. For our evaluation, we consider a scenario where an XPath processor uses multiple threads to concurrently navigate and execute individual XPath queries on a shared XML document. Given the constraints of the XML execution and data models, we propose three strategies for parallelizing individual XPath queries: Data partitioning, Query partitioning, and Hybrid (query and data) partitioning. We experimentally evaluated these strategies on an x86 Linux multi-core system using a set of XPath queries, invoked on a variety of XML documents using the Xalan XPath APIs. Experimental results demonstrate that the proposed parallelization strategies work very effectively in practice; for a majority of XPath queries under evaluation, the execution performance scaled linearly as the number of threads was increased. Results also revealed the pros and cons of the different parallelization strategies for different XPath query patterns.
Rajesh Bordawekar, Lipyeow Lim, Oded Shmueli
EDBT2
2009 A declarative framework for semantic link discovery over relational data
abstract
In this paper, we present a framework for online discovery of semantic links from relational data. Our framework is based on declarative specification of the linkage requirements by the user, that allows matching data items in many real-world scenarios. These requirements are translated to queries that can run over the relational data source, potentially using the semantic knowledge to enhance the accuracy of link discovery. Our framework lets data publishers to easily find and publish high-quality links to other data sources, and therefore could significantly enhance the value of the data in the next generation of web.
Oktie Hassanzadeh, Lipyeow Lim, Anastasios Kementsietsidis, Min Wang 0001
WWW2
2009 Efficient Index Compression in DB2 LUW
abstract
In database systems, the cost of data storage and retrieval are important components of the total cost and response time of the system. A popular mechanism to reduce the storage footprint is by compressing the data residing in tables and indexes. Compressing indexes efficiently, while maintaining response time requirements, is known to be challenging. This is especially true when designing for a workload spectrum covering both data warehousing and transaction processing environments. DB2 Linux, UNIX, Windows (LUW) recently introduced index compression for use in both environments. This uses techniques that are able to compress index data efficiently while incurring virtually no performance penalty for query processing. On the contrary, for certain operations, the performance is actually better. In this paper, we detail the design of index compression in DB2 LUW and discuss the challenges that were encountered in meeting the design goals. We also demonstrate its effectiveness by showing performance results on typical customer scenarios.
Bishwaranjan Bhattacharjee, Lipyeow Lim, Timothy Malkemus, George A. Mihaila, Kenneth A. Ross, Sherman Lau, Cathy McCarthur, Zoltan Toth, Reza Sherkat
Proc. VLDB Endow.2
2009 Linkage Query Writer
abstract
We present Linkage Query Writer (LinQuer), a system for generating SQL queries for semantic link discovery over relational data. The LinQuer framework consists of (a) LinQL, a language for specification of linkage requirements; (b) a web interface and an API for translating LinQL queries to standard SQL queries; (c) an interface that assists users in writing LinQL queries. We discuss the challenges involved in the design and implementation of a declarative and easy to use framework for discovering links between different data items in a single data source or across different data sources. We demonstrate different steps of the linkage requirements specification and discovery process in several real world scenarios and show how the LinQuer system can be used to create high-quality linked data sources.
Oktie Hassanzadeh, Reynold Xin, Renée J. Miller, Anastasios Kementsietsidis, Lipyeow Lim, Min Wang 0001
Proc. VLDB Endow.5
2008 Supporting Ontology-based Keyword Search over Medical Databases
Anastasios Kementsietsidis, Lipyeow Lim, Min Wang 0001
AMIA2
2008 Modeling and Querying E-Commerce Data in Hybrid Relational-XML DBMSs
Lipyeow Lim, Haixun Wang, Min Wang 0001
ER1
2008 Optimizing Hierarchical Access in OLAP Environment
abstract
In online analytic processing (OLAP) deployments, different users, lines of businesses and business units often create adhoc aggregation hierarchies tailor-made for specific reporting or analytical applications. As a result, a large number of these application specific hierarchies accumulate over time. System administrators typically are not able to optimize all these hierarchical accesses by hand due to the large number of hierarchies. However, many optimization opportunities exist due to the significant amount of overlap between some hierarhies. In this paper, we sketch a novel method for optimizing OLAP aggregation queries using precomputed aggregates on other overlapping hierarchies. Our method detects common sub-structures among hierarchies and provides a rewriting algorithm to exploit any precomputations on these shared sub-structures.
Lipyeow Lim, Bishwaranjan Bhattacharjee
ICDE1
2007 Semantic Data Management: Towards Querying Data with their Meaning
abstract
Relational database management systems are constantly being extended and augmented to accommodate data in different domains. Recently, with the increasing use of ontology in various applications, the need to support ontology, especially the related inferencing operation, in DBMS has become more concrete and urgent. However, manipulating knowledge along with relational data in DBMSs is not a trivial undertaking due to the mismatch in data models. In this paper, we introduce a framework for managing relational data and hierarchical domain knowledge together. Our framework persists taxonomies contained in ontologies by leveraging XML support in hybrid relational-XML DBMSs (e.g., IBM's DB2 v9) and rewrites ontology-based semantic matching queries using the industry-standard query languages, SQL/XML and XQuery. Compared with previous approaches, our approach does not materialize transitive closures of ontological relationships to support inferencing. Consequently, our method has wide applicability and good performance.
Lipyeow Lim, Haixun Wang, Min Wang 0001
ICDE1
2007 Supporting ranking and clustering as generalized order-by and group-by
abstract
The Boolean semantics of SQL queries cannot adequately capture the "fuzzy" preferences and "soft" criteria required in non-traditional data retrieval applications. One way to solve this problem is to add a flavor of "information retrieval" into database queries by allowing fuzzy query conditions and flexibly supporting grouping and ranking of the query results within the DBMS engine. While ranking is already supported by all major commercial DBMSs natively, support of flexibly grouping is still very limited (i.e., group-by).
Chengkai Li 0001, Min Wang 0001, Lipyeow Lim, Haixun Wang, Kevin Chen-Chuan Chang
SIGMOD Conference3
2007 Schema advisor for hybrid relational-XML DBMS
abstract
In response to the widespread use of the XML format for document representation and message exchange, major database vendors support XML in terms of persistence, querying and indexing. Specifically, the recently released IBM DB2 9 (for Linux, Unix and Windows) is a hybrid data server with optimized management of both XML and relational data. With the new option of storing and querying XML in a relational DBMS, data architects face the the decision of what portion of their data to persist as XML and what portion as relational data. This problem has not been addressed yet and represents a serious need in the industry. Hence, this paper describes ReXSA, a schema advisor tool that is being prototyped for IBM DB2 9. ReXSA proposes candidate database schemas given an information model of the enterprise data. It has the advantage of considering qualitative properties of the information model such as reuse, evolution and performance profiles for deciding how to persist the data. Finally, we show the viability and practicality of ReXSA by applying it to custom and real usecases.
Mirella M. Moro, Lipyeow Lim, Yuan-Chi Chang
SIGMOD Conference2
2007 Unifying Data and Domain Knowledge Using Virtual Views
Lipyeow Lim, Haixun Wang, Min Wang 0001
VLDB1
2007 Preserving XML queries during schema evolution
abstract
In XML databases, new schema versions may be released as frequently as once every two weeks. This poster describes a taxonomy of changes for XML schema evolution. It examines the impact of those changes on schema validation and query evaluation. Based on that study, it proposes guidelines for XML schema evolution and for writing queries in such a way that they continue to operate as expected across evolving schemas.
Mirella M. Moro, Susan Malaika, Lipyeow Lim
WWW3
2005 CXHist : An On-line Classification-Based Histogram for XML String Selectivity Estimation
Lipyeow Lim, Min Wang 0001, Jeffrey Scott Vitter
VLDB1
2003 SASH: A Self-Adaptive Histogram Set for Dynamically Changing Workloads
Lipyeow Lim, Min Wang 0001, Jeffrey Scott Vitter
VLDB1
2003 Dynamic maintenance of web indexes using landmarks
abstract
Recent work on incremental crawling has enabled the indexed document collection of a search engine to be more synchronized with the changing World Wide Web. However, this synchronized collection is not immediately searchable, because the keyword index is rebuilt from scratch less frequently than the collection can be refreshed. An inverted index is usually used to index documents crawled from the web. Complete index rebuild at high frequency is expensive. Previous work on incremental inverted index updates have been restricted to adding and removing documents. Updating the inverted index for previously indexed documents that have changed has not been addressed.In this paper, we propose an efficient method to update the inverted index for previously indexed documents whose contents have changed. Our method uses the idea of landmarks together with the diff algorithm to significantly reduce the number of postings in the inverted index that need to be updated. Our experiments verify that our landmark-diff method results in significant savings in the number of update operations on the inverted index.
Lipyeow Lim, Min Wang 0001, Sriram Padmanabhan, Jeffrey Scott Vitter, Ramesh C. Agarwal
WWW1
2002 XPathLearner: An On-line Self-Tuning Markov Histogram for XML Path Selectivity Estimation
Lipyeow Lim, Min Wang 0001, Sriram Padmanabhan, Jeffrey Scott Vitter, Ronald Parr
VLDB1
2001 Wavelet-Based Cost Estimation for Spatial Queries
Min Wang 0001, Jeffrey Scott Vitter, Lipyeow Lim, Sriram Padmanabhan
SSTD3
2001 Characterizing Web Document Change
Lipyeow Lim, Min Wang 0001, Sriram Padmanabhan, Jeffrey Scott Vitter, Ramesh C. Agarwal
WAIM1