EDBT 2026 Demo / reviewers in the wild / expert
Maik Thiele
dblp:48/2949
· DBLP profile ↗
41ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-1665-977XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 35 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Survey of Active Learning Hyperparameters: Insights From a Large-Scale Experimental GridabstractAnnotating data is a time-consuming and costly task, but it is inherently required for supervised machine learning. Active Learning (AL) is an established method that minimizes human labeling effort by iteratively selecting the most informative unlabeled samples for expert annotation, thereby improving the overall classification performance. Even though AL has been known for decades [1], AL is still rarely used in real-world applications. As indicated in the two community web surveys among the NLP community about AL [2], [3], two main reasons continue to hold practitioners back from using AL: first, the complexity of setting AL up, and second, a lack of trust in its effectiveness. We hypothesize that both reasons share the same culprit: the large hyperparameter space of AL. This mostly unexplored hyperparameter space often leads to misleading and irreproducible glsAL experiment results. In this study, we first compiled a large hyperparameter grid of over 4.6 million hyperparameter combinations, second, recorded the performance of all combinations in the so-far biggest conducted AL study, and third, analyzed the impact of each hyperparameter in the experiment results. Rather than merely reporting correlations, we explicitly focus on distilling these results into practitioner-oriented rulesof-thumb for designing AL experiments under realistic resource constraints. In the end, we give recommendations about the influence of each hyperparameter, demonstrate the surprising influence of the concrete AL strategy implementation, and outline an experimental study design for reproducible AL experiments with minimal computational effort, thus contributing to more reproducible and trustworthy AL research in the future. Julius Gonsior, Tim Rieß, Anja Reusch, Claudio Hartmann, Maik Thiele, Wolfgang Lehner |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Active Learning with Aggregated Uncertainties from Image Augmentations
Tamás Janusko, Colin Simon, Kevin Kirsten, Serhiy Bolkun, Eric Weinzierl, Julius Gonsior, Maik Thiele |
EANN | 7 |
| 2023 | Comparing and Improving Active Learning Uncertainty Measures for Transformer Models
Julius Gonsior, Christian Falkenberg, Silvio Magino, Anja Reusch, Claudio Hartmann, Maik Thiele, Wolfgang Lehner |
ADBIS | 6 |
| 2022 | ImitAL: Learned Active Learning Strategy on Synthetic Data
Julius Gonsior, Maik Thiele, Wolfgang Lehner |
DS | 2 |
| 2022 | ALWars: Combat-Based Evaluation of Active Learning Strategies
Julius Gonsior, Jakob Krude, Janik Schönfelder, Maik Thiele, Wolfgang Lehner |
ECIR (2) | 4 |
| 2021 | An ALBERT-based Similarity Measure for Mathematical Answer RetrievalabstractMathematical Language Processing (MLP) deals with the automated processing and analysis of mathematical documents and relies heavily on good representations of mathematical symbols and texts. The aim of this work is to explore the modeling capabilities of state-of-the-art unsupervised deep learning methods to create such representations. Therefore, we pre-trained different instances of an ALBERT model on Mathematics StackExchange data and fine-tuned it on the task of Mathematical Answer Retrieval. Our evaluation shows that ALBERT outperforms all previous systems and is on par with current state-of-the-art systems for math retrieval indicating strong capabilities of modeling mathematical posts. This implies that our approach can also be beneficial to various other tasks in MLP such as automatic proof checking or summarization of scientific texts. Anja Reusch, Maik Thiele, Wolfgang Lehner |
SIGIR | 2 |
| 2020 | Learning from Textual Data in Database SystemsabstractRelational database systems hold massive amounts of text, valuable for many machine learning (ML) tasks. Since ML techniques depend on numerical input representations, pre-trained word embeddings are increasingly utilized to convert text values into meaningful numbers. However, a naïve one-to-one mapping of each word in a database to a word embedding vector misses incorporating rich context information given by the database schema. Thus, we propose a novel relational retrofitting framework Retro to learn numerical representations of text values in databases, capturing the rich information encoded by pre-trained word embedding models as well as context information provided by tabular and foreign key relations in the database. We defined relation retrofitting as an optimization problem, present an efficient algorithm solving it, and investigate the influence of various hyperparameters. Further, we develop simple feed-forward and complex graph convolutional neural network architectures to operate on those representations. Our evaluation shows that the proposed embeddings and models are ready-to-use for many ML tasks, such as text classification, imputation, and link prediction, and even outperform state-of-the-art techniques. Michael Günther 0002, Philipp Oehme, Maik Thiele, Wolfgang Lehner |
CIKM | 3 |
| 2020 | WeakAL: Combining Active Learning and Weak Supervision
Julius Gonsior, Maik Thiele, Wolfgang Lehner |
DS | 2 |
| 2020 | Retro: Relation Retrofitting For In-Database Machine Learning on Textual Data
Michael Günther 0002, Maik Thiele, Wolfgang Lehner |
EDBT | 2 |
| 2020 | RetroLive: Analysis of Relational Retrofitted Word Embeddings
Michael Günther 0002, Maik Thiele, Erik Nikulski, Wolfgang Lehner |
EDBT | 2 |
| 2020 | A cost-based storage format selector for materialized results in big data frameworks
Rana Faisal Munir, Alberto Abelló, Oscar Romero 0001, Maik Thiele, Wolfgang Lehner |
Distributed Parallel Databases | 4 |
| 2019 | XLIndy: Interactive Recognition and Information Extraction in SpreadsheetsabstractOver the years, spreadsheets have established their presence in many domains, including business, government, and science. However, challenges arise due to spreadsheets being partially-structured and carrying implicit (visual and textual) information. This translates into a bottleneck, when it comes to automatic analysis and extraction of information. Therefore, we present XLIndy, a Microsoft Excel add-in with a machine learning back-end, written in Python. It showcases our novel methods for layout inference and table recognition in spreadsheets. For a selected task and method, users can visually inspect the results, change configurations, and compare different runs. This enables iterative fine-tuning. Additionally, users can manually revise the predicted layout and tables, and subsequently save them as annotations. The latter is used to measure performance and (re-)train classifiers. Finally, data in the recognized tables can be extracted for further processing. XLIndy supports several standard formats, such as CSV and JSON. Elvis Koci, Dana Kuban, Nico Luettig, Dominik Olwig, Maik Thiele, Julius Gonsior, Wolfgang Lehner, Oscar Romero 0001 |
DocEng | 5 |
| 2019 | A Genetic-Based Search for Adaptive Table Recognition in SpreadsheetsabstractSpreadsheets are very successful content generation tools, used in almost every enterprise to create a wealth of information. However, this information is often intermingled with various formatting, layout, and textual metadata, making it hard to identify and interpret the tabular payload. Previous works proposed to solve this problem by mainly using heuristics. Although fast to implement, these approaches fail to capture the high variability of user-generated spreadsheet tables. Therefore, in this paper, we propose a supervised approach that is able to adapt to arbitrary spreadsheet datasets. We use a graph model to represent the contents of a sheet, which carries layout and spatial features. Subsequently, we apply genetic-based approaches for graph partitioning, to recognize the parts of the graph corresponding to tables in the sheet. The search for tables is guided by an objective function, which is tuned to match the specific characteristics of a given dataset. We present the feasibility of this approach with an experimental evaluation, on a large, real-world spreadsheet corpus. Elvis Koci, Maik Thiele, Oscar Romero 0001, Wolfgang Lehner |
ICDAR | 2 |
| 2019 | DECO: A Dataset of Annotated Spreadsheets for Layout and Table RecognitionabstractThis paper presents DECO (Dresden Enron COrpus), a dataset of spreadsheet files, annotated on the basis of layout and contents. It comprises of 1,165 files, extracted from the Enron corpus. Three different annotators (judges) assigned layout roles (e.g., Header, Data, and Notes) to non-empty cells and marked the borders of tables. Files that do not contain tables were flagged using categories such as Template, Form, and Report. Subsequently, a thorough analysis is performed to uncover the characteristics of the overall dataset and specific annotations. The results are discussed in this paper, providing several takeaways for future works. Furthermore, this work describes in detail the annotation methodology, going through the individual steps. The dataset, methodology, and tools are made publicly available, so that they can be adopted for further studies. DECO is available at: https://wwwdb.inf.tu-dresden.de/research-projects/deexcelarator/, Elvis Koci, Maik Thiele, Josephine Rehak, Oscar Romero 0001, Wolfgang Lehner |
ICDAR | 2 |
| 2018 | ATUN-HL: Auto Tuning of Hybrid Layouts Using Workload and Data Characteristics
Rana Faisal Munir, Alberto Abelló, Oscar Romero 0001, Maik Thiele, Wolfgang Lehner |
ADBIS | 4 |
| 2018 | Table Recognition in Spreadsheets via a Graph RepresentationabstractSpreadsheet software are very popular data management tools. Their ease of use and abundant functionalities equip novices and professionals alike with the means to generate, transform, analyze, and visualize data. As a result, spreadsheets are a great resource of factual and structured information. This accentuates the need to automatically understand and extract their contents. In this paper, we present a novel approach for recognizing tables in spreadsheets. Having inferred the layout role of the individual cells, we build layout regions. We encode the spatial interrelations between these regions using a graph representation. Based on this, we propose Remove and Conquer (RAC), an algorithm for table recognition that implements a list of carefully curated rules. An extensive experimental evaluation shows that our approach is viable. We achieve significant accuracy in a dataset of real spreadsheets from various domains. Elvis Koci, Maik Thiele, Wolfgang Lehner, Oscar Romero 0001 |
DAS | 2 |
| 2018 | Modeling Customers and Products with Word Embeddings from Receipt DataabstractFor many tasks in market research it is important to model customers and products as comparable instances. Usually, the integration of customers and products into one model is not trivial. In this paper, we will detail an approach for a combined vector space of customers and products based on word embeddings learned from receipt data. To highlight the strengths of this approach we propose four different applications: recommender systems, customer and product segmentation and purchase prediction. Experimental results on a real-world dataset with 200M order receipts for 2M customers show that our word embedding approach is promising and helps to improve the quality in these applications scenarios. Lucas Woltmann, Maik Thiele, Wolfgang Lehner |
IDEAS | 2 |
| 2018 | Intermediate Results Materialization Selection and Format for Data-Intensive FlowsabstractData-intensive flows deploy a variety of complex data transformations to build information pipelines from data sources to different end users. As data are processed, these workflows generate large intermediate results, typically pipelined from one operator to the following ones. Materializing intermediate results, shared among multiple flows, brings benefits not only in terms of performance but also in resource usage and consistency. Similar ideas have been proposed in the context of data warehouses, which are studied under the materialized view selection problem. With the rise of Big Data systems, new challenges emerge due to new quality metrics captured by service level agreements which must be taken into account. Moreover, the way such results are stored must be reconsidered, as different data layouts can be used to reduce the I/O cost. In this paper, we propose a novel approach for automatic selection of multi-objective materialization of intermediate results in data-intensive flows, which can tackle multiple and conflicting quality objectives. In addition, our approach chooses the optimal storage data format for selected materialized intermediate results based on subsequent access patterns. The experimental results show that our approach provides 40% better average speedup with respect to the current state-of-the-art, as well as an improvement on disk access time of 18% as compared to fixed format solutions. Rana Faisal Munir, Sergi Nadal, Oscar Romero 0001, Alberto Abelló, Petar Jovanovic 0001, Maik Thiele, Wolfgang Lehner |
Fundam. Informaticae | 6 |
| 2017 | Context Similarity for Retrieval-Based ImputationabstractCompleteness as one of the four major dimensions of data quality is a pervasive issue in modern databases. Although data imputation has been studied extensively in the literature, most of the research is focused on inference-based approach. We propose to harness Web tables as an external data source to effectively and efficiently retrieve missing data while taking into account the inherent uncertainty and lack of veracity that they contain. Ahmad Ahmadov, Maik Thiele, Wolfgang Lehner, Robert Wrembel |
ASONAM | 2 |
| 2017 | Table Identification and Reconstruction in Spreadsheets
Elvis Koci, Maik Thiele, Oscar Romero 0001, Wolfgang Lehner |
CAiSE | 2 |
| 2017 | Frequent patterns in ETL workflows: An empirical approach
Vasileios Theodorou, Alberto Abelló, Maik Thiele, Wolfgang Lehner |
Data Knowl. Eng. | 3 |
| 2016 | Cell Classification for Layout Recognition in Spreadsheets
Elvis Koci, Maik Thiele, Oscar Romero 0001, Wolfgang Lehner |
IC3K | 2 |
| 2016 | DebEAQ - debugging empty-answer queries on large data graphsabstractThe large volume of freely available graph data sets impedes the users in analyzing them. For this purpose, they usually pose plenty of pattern matching queries and study their answers. Without deep knowledge about the data graph, users can create ‘failing’ queries, which deliver empty answers. Analyzing the causes of these empty answers is a time-consuming and complicated task especially for graph queries. To help users in debugging these ‘failing’ queries, there are two common approaches: one is focusing on discovering missing subgraphs of a data graph, the other one tries to rewrite the queries such that they deliver some results. In this demonstration, we will combine both approaches and give the users an opportunity to discover why empty results were delivered by the requested queries. Therefore, we propose DebEAQ, a debugging tool for pattern matching queries, which allows to compare both approaches and also provides functionality to debug queries manually. Elena Vasilyeva, Thomas Heinze 0001, Maik Thiele, Wolfgang Lehner |
ICDE | 3 |
| 2016 | ResilientStore: A Heuristic-Based Data Format Selector for Intermediate Results
Rana Faisal Munir, Oscar Romero 0001, Alberto Abelló, Besim Bilalli, Maik Thiele, Wolfgang Lehner |
MEDI | 5 |
| 2016 | Quality measures for ETL processes: from goals to implementationabstractSummary Extraction transformation loading (ETL) processes play an increasingly important role for the support of modern business operations. These business processes are centred around artifacts with high variability and diverse lifecycles, which correspond to key business entities. The apparent complexity of these activities has been examined through the prism of business process management, mainly focusing on functional requirements and performance optimization. However, the quality dimension has not yet been thoroughly investigated, and there is a need for a more human‐centric approach to bring them closer to business‐users requirements. In this paper, we take a first step towards this direction by defining a sound model for ETL process quality characteristics and quantitative measures for each characteristic, based on existing literature. Our model shows dependencies among quality characteristics and can provide the basis for subsequent analysis using goal modeling techniques. We showcase the use of goal modeling for ETL process design through a use case, where we employ the use of a goal model that includes quantitative components (i.e., indicators) for evaluation and analysis of alternative design decisions. Copyright © 2015 John Wiley & Sons, Ltd. Vasileios Theodorou, Alberto Abelló, Wolfgang Lehner, Maik Thiele |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Answering "Why Empty?" and "Why So Many?" queries in graph databases
Elena Vasilyeva, Maik Thiele, Christof Bornhövd, Wolfgang Lehner |
J. Comput. Syst. Sci. | 2 |
| 2015 | POIESIS: a Tool for Quality-aware ETL Process RedesignabstractWe present a tool, called POIESIS, for automatic ETL process enhancement. ETL processes are essential data-centric activities in modern business intelligence environments and they need to be examined through a viewpoint that concerns their quality characteristics (e.g., data quality, performance, manageability) in the era of Big Data.\nPOIESIS responds to this need by providing a user-centered environment for quality-aware analysis and redesign of ETL flows. It generates thousands of alternative flows by adding flow patterns to the initial flow, in varying positions and combinations, thus creating alternative design options in a multidimensional space of different quality attributes.\nThrough the demonstration of POIESIS we introduce the tool's capabilities and highlight its efficiency, usability and modifiability, thanks to its polymorphic design. © 2015, Copyright is with the authors. Vasileios Theodorou, Alberto Abelló, Maik Thiele, Wolfgang Lehner |
EDBT | 3 |
| 2015 | From Web Tables to Concepts: A Semantic Normalization Approach
Katrin Braunschweig, Maik Thiele, Wolfgang Lehner |
ER | 2 |
| 2015 | Top-k entity augmentation using consistent set coveringabstractEntity augmentation is a query type in which, given a set of entities and a large corpus of possible data sources, the values of a missing attribute are to be retrieved. State of the art methods return a single result that, to cover all queried entities, is fused from a potentially large set of data sources. We argue that queries on large corpora of heterogeneous sources using information retrieval and automatic schema matching methods can not easily return a single result that the user can trust, especially if the result is composed from a large number of sources that user has to verify manually. We therefore propose to process these queries in a Top-k fashion, in which the system produces multiple minimal consistent solutions from which the user can choose to resolve the uncertainty of the data sources and methods used. In this paper, we introduce and formalize the problem of consistent, multi-solution set covering, and present algorithms based on a greedy and a genetic optimization approach. We then apply these algorithms to Web table-based entity augmentation. The publication further includes a Web table corpus with 100M tables, and a Web table retrieval and matching system in which these algorithms are implemented. Our experiments show that the consistency and minimality of the augmentation results can be improved using our set covering approach, without loss of precision or coverage and while producing multiple alternative query results. Julian Eberius, Maik Thiele, Katrin Braunschweig, Wolfgang Lehner |
SSDBM | 2 |
| 2015 | DrillBeyond: processing multi-result open world SQL queriesabstractIn a traditional relational database management system, queries can only be defined over attributes defined in the schema, but are guaranteed to give single, definitive answer structured exactly as specified in the query. In contrast, an information retrieval system allows the user to pose queries without knowledge of a schema, but the result will be a top-k list of possible answers, with no guarantees about the structure or content of the retrieved documents. Julian Eberius, Maik Thiele, Katrin Braunschweig, Wolfgang Lehner |
SSDBM | 2 |
| 2015 | Relaxation of subgraph queries delivering empty resultsabstractGraph databases with the property graph model are used in multiple domains including social networks, biology, and data integration. They provide schema-flexible storage for data of a different degree of a structure and support complex, expressive queries such as subgraph isomorphism queries. The exibility and expressiveness of graph databases make it difficult for the users to express queries correctly and can lead to unexpected query results, e.g. empty results. Therefore, we propose a relaxation approach for subgraph isomorphism queries that is able to automatically rewrite a graph query, such that the rewritten query is similar to the original query and returns a non-empty result set. In detail, we present relaxation operations applicable to a query, cardinality estimation heuristics, and strategies for prioritizing graph query elements to be relaxed. To determine the similarity between the original query and its relaxed variants, we propose a novel cardinality-based graph edit distance. The feasibility of our approach is shown by using real-world queries from the DBpedia query log. Elena Vasilyeva, Maik Thiele, Adrian Mocan, Wolfgang Lehner |
SSDBM | 2 |
| 2015 | Considering User Intention in Differential Graph QueriesabstractEmpty answers are a major problem by processing pattern matching queries in graph databases. Especially, there can be multiple reasons why a query failed. To support users in such situations, differential queries can be used that deliver missing parts of a graph query. Multiple heuristics are proposed for differential queries, which reduce the search space. Although they are successful in increasing the performance, they can discard query subgraphs relevant to a user. To address this issue, the authors extend the concept of differential queries and introduce top-k differential queries that calculate the ranking based on users' preferences and significantly support the users' understanding of query database management systems. A user assigns relevance weights to elements of a graph query that steer the search and are used for the ranking. In this paper the authors propose different strategies for selection of relevance weights and their propagation. As a result, the search is modelled along the most relevant paths. The authors evaluate their solution and both strategies on the DBpedia data graph. Elena Vasilyeva, Maik Thiele, Christof Bornhövd, Wolfgang Lehner |
J. Database Manag. | 2 |
| 2014 | Top-k Differential Queries in Graph Databases
Elena Vasilyeva, Maik Thiele, Christof Bornhövd, Wolfgang Lehner |
ADBIS | 2 |
| 2014 | A Framework for User-Centered Declarative ETLabstractAs business requirements evolve with increasing information density and velocity, there is a growing need for efficiency and automation of Extract-Transform-Load (ETL) processes. Current approaches for the modeling and optimization of ETL processes provide platform-independent optimization solutions for the (semi-)automated transition among different abstraction levels, focusing on cost and performance. However, the suggested representations are not abstract enough to communicate business requirements and the role of the process quality in a user-centered perspective has not yet been adequately examined. In this paper, we introduce a novel methodology for the end-to-end design of ETL processes that takes under consideration both functional and non-functional requirements. Based on existing work, we raise the level of abstraction for the conceptual representation of ETL operations and we show how process quality characteristics can generate specific patterns on the process design. Vasileios Theodorou, Alberto Abelló, Maik Thiele, Wolfgang Lehner |
DOLAP | 3 |
| 2013 | DeExcelerator: a framework for extracting relational data from partially structured documentsabstractOf the structured data published on the web, for instance as datasets on Open Data Platforms such as data.gov, but also in the form of HTML tables on the general web, only a small part is in a relational form. Instead the data is intermingled with formatting, layout and textual metadata, i.e., it is contained in partially structured documents. This makes transformation into a true relational form necessary, which is a precondition for most forms of data analysis and data integration. Studying data.gov as an example source for partially structured documents, we present a classification of typical normalization problems. We then present the DeExcelerator, which is a framework for extracting relations from partially structured documents such as spreadsheets and HTML tables. Julian Eberius, Christopher Werner, Maik Thiele, Katrin Braunschweig, Lars Dannecker, Wolfgang Lehner |
CIKM | 3 |
| 2012 | DrillBeyond: Enabling Business Analysts to Explore the Web of Open DataabstractFollowing the Open Data trend, governments and public agencies have started making their data available on the Web and established platforms such as data.gov or data.un.org. These Open Data platforms provide a huge amount of data for various topics such as demographics, transport, finance or health in various data formats. One typical usage scenario for this kind of data is their integration into a database or data warehouse in order to apply data analytics. However, in today's business intelligence tools there is an evident lack of support for so-called situational or ad-hoc data integration. In this demonstration we will therefore present DrillBeyond , a novel database and information retrieval engine which allows users to query a local database as well as the Web of Open Data in a seamless and integrated way with standard SQL. The audience will be able to pose queries to our DrillBeyond system which will be answered partly from local data in the database and partly from datasets that originate from the Web of Data. We will show how such queries are divided into known and unknown parts and how missing attributes are mapped to open datasets. We will demonstrate the integration of the open datasets back into the DBMS in order to apply its analytical features. Julian Eberius, Maik Thiele, Katrin Braunschweig, Wolfgang Lehner |
Proc. VLDB Endow. | 2 |
| 2009 | Cardinality estimation in ETL processesabstractThe cardinality estimation in ETL processes is particularly difficult. Aside from the well-known SQL operators, which are also used in ETL processes, there are a variety of operators without exact counterparts in the relational world. In addition to those, we find operators that support very specific data integration aspects. For such operators, there are no well-examined statistic approaches for cardinality estimations. Therefore, we propose a black-box approach and estimate the cardinality using a set of statistic models for each operator. We discuss different model granularities and develop an adaptive cardinality estimation framework for ETL processes. We map the abstract model operators to specific statistic learning approaches (regression, decision trees, support vector machines, etc.) and evaluate our cardinality estimations in an extensive experimental study. Maik Thiele, Tim Kiefer, Wolfgang Lehner |
DOLAP | 1 |
| 2009 | Partition-based workload scheduling in living data warehouse environments
Maik Thiele, Ulrike Fischer, Wolfgang Lehner |
Inf. Syst. | 1 |
| 2007 | Partition-based workload scheduling in living data warehouse environmentsabstractThe demand for so-called living or real-time data warehouses is increasing in many application areas such as manufacturing, event monitoring and telecommunications. In these fields users usually expect short response times for their queries and high freshness for the requested data. However, meeting these fundamental requirements is challenging due to the high loads and the continuous flow of write-only updates and read-only queries, which may be in conflict with each other. Therefore, we present the concept of Workload Balancing by Election (WINE), which allows users to express their individual demands on the Quality of Service and the Quality of Data respectively. WINE applies this information to balance and prioritize over both types of transactions -- queries and update -- according to the varying user needs. A simulation study shows that our proposed algorithm outperforms competitor baseline algorithms over the entire spectrum of workloads and user requirements. Maik Thiele, Ulrike Fischer, Wolfgang Lehner |
DOLAP | 1 |
| 2006 | Shrinked Data Marts Enabled for Negative CachingabstractData marts storing pre-aggregated data, prepared for further roll-ups, play an essential role in data warehouse environments and lead to significant performance gains in the query evaluation. However, in order to ensure the completeness of query results on the data mart without to access the underlying data warehouse, null values need to be stored explicitly; this process is denoted as negative caching. Such null values typically occur in multidimensional data sets, which are naturally very sparse. To our knowledge, there is no work on shrinking the null tuples in a multi-dimensional data set within ROLAP. For these tuples, we propose a lossless compression technique, leading to a dramatic reduction in size of the data mart. Queries depending on null value information can be answered with 100% precision by partially inflating the shrunken data mart. We complement our analytical approach with an experimental evaluation using real and synthetic data sets, and demonstrate our results Maik Thiele, Wolfgang Lehner |
IDEAS | 1 |
| 2006 | Optimistic Coarse-Grained Cache Semantics for Data MartsabstractData marts and caching are two closely related concepts in the domain of multi-dimensional data. Both store pre-computed data to provide fast response times for complex OLAP queries, and for both it must be guaranteed that every query can be completely processed. However, they differ extremely in their update behaviour which we utilise to build a specific data mart extended by cache semantics. In this paper, we introduce a novel cache exploitation concept for data marts - coarse-grained caching - in which the containedness check for a multi-dimensional query is done through the comparison of the expected and the actual cardinalities. Therefore, we subdivide the multi-dimensional data into coarse partitions, the so called cubletets, which allow to specify the completeness criteria for incoming queries. We show that during query processing, the completeness check is done with no additional costs Maik Thiele, Jens Albrecht, Wolfgang Lehner |
SSDBM | 1 |