EDBT 2026 Demo / reviewers in the wild / expert
Stefan Manegold
dblp:m/StefanManegold
· DBLP profile ↗
55ranked-venue papers in the field
9as first author
5since 2021 · last 2025
0000-0001-7938-4358ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 53 (9 first)Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | G-ALP: Rethinking Light-weight Encodings for GPUsabstractThis paper introduces G-ALP, a GPU-optimized version of ALP, which is a recent and state-of-the-art compression scheme for floating-point.This GPU-optimization is based on two core ideas.First, all parts of the decoding process must be fully data-parallelized.In this paper, we fully data-parallelize exception patching, which typically applies to only 1% of the data.While patching has negligible performance cost on CPUs, it can become the main bottleneck on GPUs if it is not data-parallel.Second, the decoding API must minimize its register footprint, a highly scarce resource on GPUs, and hence deliver just one value-at-a-time.Our unique aim is to integrate G-ALP decoding into GPU kernels that consume data from global memory, rather than let decompression be a separate kernel.We consider these two ideas general guidelines for future GPU-optimized lightweight encodings, and a significant evolution of our new FastLanes file format, making it GPU-friendly.We extensively test G-ALP in a series of microbenchmarks and evaluate its performance on an NVIDIA V100 GPU and an NVIDIA RTX4070 Super Ti GPU, demonstrating superior performance compared to NVIDIA nvCOMP and ndzip in both decoding and filtering queries. Sven Hepkema, Azim Afroozeh, Charlotte Felius, Peter Boncz, Stefan Manegold |
DaMoN | 5 |
| 2021 | Progressive Join Algorithms Considering User Preference
Mengsu Ding, Shimin Chen, Nantia Makrynioti, Stefan Manegold |
CIDR | 4 |
| 2021 | Progressive Mergesort: Merging Batches of Appends into Progressive IndexesabstractInteractive exploratory data analysis consists of workloads that are composed of filter-aggregate queries with highly selective filters [1]. Hence, their performance is dependent on how much data they can skip during their scans, with indexes being the most efficient technique for aggressive data-skipping. Progressive Indexes are the state-of-the-art on automatic index creation for interactive exploratory data analysis. These indexes are partially constructed during query execution, eventually refining to a full index. However, progressive indexes have been designed for static databases, while in exploratory data analysis updates - usually batch-appends of newly acquired data - are frequent. In this paper, we propose Progressive Mergesort, a novel merging technique to make Progressive Indexes cope with updates. Progressive Mergesort differs from other merging techniques for partial indexes as it incorporates the index budget strategy design from Progressive Indexing. It follows the same three principles as Progressive Indexes: (1) fast query execution, (2) high robustness,(3) guaranteed convergence. Our experimental evaluation demonstrates that Progressive Mergesort is capable of achieving a 2x speedup when merging updates and up to 3 orders of magnitude lower variance than the state of the art. Pedro Holanda, Stefan Manegold |
EDBT | 2 |
| 2021 | Multidimensional Adaptive & Progressive IndexesabstractExploratory data analysis is the primary technique used by data scientists to extract knowledge from new data sets. This type of workload is composed of trial-and-error hypothesis-driven queries with a human in the loop. To keep up with the data scientist's productivity, the system must be capable of answering queries in interactive times. Given that these queries are highly selective multidimensional queries, multidimensional indexes are necessary to ensure low latency. However, creating the appropriate indexes is not a given due to the highly exploratory and interactive nature of such human-in-the-loop scenarios.In this paper, we identify four main objectives that are desirable for exploratory data analysis workloads: (1) low overhead over the initial queries, (2) low query variance (i.e., high robustness), (3) predictable index convergence, and (4) low total workload time. Given that not all of them can be achieved at the same time, we present three novel incremental multidimensional indexing techniques that represent three sample points on a Pareto front for this multi-objective optimization problem. (a) The Adaptive KD-Tree is designed to achieve the lowest total workload time at the expense of a higher indexing penalty for the initial queries, lack of robustness, and unpredictable convergence. (b) The Progressive KD-Tree has predictable convergence and a user-defined indexing cost for the initial queries. However, total workload time can be higher than with Adaptive KD-Trees, and per-query time still varies. (c) The Greedy Progressive KD-Tree aims at full robustness at the expense of only improving the per-query cost after full index convergence.Our extensive experimental evaluation using both synthetic and real-life data sets and workloads shows that (a) the Adaptive KD-Tree reduces total workload time by up to a factor 2 compared to the state-of-the-art, (b) the Progressive KD-Tree achieves predictable convergence with up to one order of magnitude lower initial query cost, and (c) the Greedy Progressive KDTree exhibits the lowest query variance up to three orders of magnitude lower than the state-of-the-art. Matheus Nerone, Pedro Holanda, Eduardo C. de Almeida, Stefan Manegold |
ICDE | 4 |
| 2021 | Front Matter
Stefan Manegold |
Proc. VLDB Endow. | 1 |
| 2019 | SQALPEL: A database performance platform
Martin L. Kersten, Stefan Manegold, Ying Zhang 0027, Panos Kuoutsourakis |
CIDR | 2 |
| 2019 | devUDF: Increasing UDF development efficiency through IDE Integration. It works like a PyCharm!abstractUser-defined functions (UDFs) facilitate the execution of analytics pipelines inside the database. They provide many advantages over traditional methods, such as close-to-data execution and automatic parallelization. However, the standard workflow for developing and debugging UDFs does not allow developers to use their regular toolchains and Integrated Development Environments (IDEs). As a result, writing functional UDFs is challenging. In this demo, we present the devUDF, a plugin to the PyCharm IDE that allows developers to develop and debug their MonetDB/Python UDFs directly from within the IDE. Mark Raasveldt, Pedro Holanda, Stefan Manegold |
EDBT | 3 |
| 2019 | Progressive Indexes: Indexing for Interactive Data AnalysisabstractInteractive exploration of large volumes of data is increasingly common, as data scientists attempt to extract interesting information from large opaque data sets. This scenario presents a difficult challenge for traditional database systems, as (1) nothing is known about the query workload in advance, (2) the query workload is constantly changing, and (3) the system must provide interactive responses to the issued queries. This environment is challenging for index creation, as traditional database indexes require upfront creation, hence a priori workload knowledge, to be efficient. In this paper, we introduce Progressive Indexing , a novel performance-driven indexing technique that focuses on automatic index creation while providing interactive response times to incoming queries. Its design allows queries to have a limited budget to spend on index creation. The indexing budget is automatically tuned to each query before query processing. This allows for systems to provide interactive answers to queries during index creation while being robust against various workload patterns and data distributions. Pedro Holanda, Stefan Manegold, Hannes Mühleisen, Mark Raasveldt |
Proc. VLDB Endow. | 2 |
| 2018 | From Big Data to Big Information and Big Knowledge: the Case of Earth Observation DataabstractSome particularly important rich sources of open and free big geospatial data are the Earth observation (EO) programs of various countries such as the Landsat program of the US and the Copernicus programme of the European Union. EO data is a paradigmatic case of big data and the same is true for the big information and big knowledge extracted from it. EO data (satellite images and in-situ data), and the information and knowledge extracted from it, can be utilized in many applications with financial and environmental impact in areas such as emergency management, climate change, agriculture and security. Konstantina Bereta, Manolis Koubarakis, Stefan Manegold, George Stamoulis 0001, Begüm Demir |
CIKM | 3 |
| 2018 | Cracking KD-Tree: The First Multidimensional Adaptive Indexing (Position Paper)abstractWorkload-aware physical data access structures are crucial to achieve short response time with (exploratory) data analysis tasks as commonly required for Big Data and Data Science applications. Recently proposed techniques such as automatic index advisers (for a priori known static workloads) and query-driven adaptive incremental indexing (for a priori unknown dynamic workloads) form the state-of-the-art to build single-dimensional indexes for single-attribute query predicates. However, similar techniques for more demanding multi-attribute query predicates, which are vital for any data analysis task, have not been proposed, yet. In this paper, we present our on-going work on a new set of workload-adaptive indexing techniques that focus on creating multidimensional indexes. We present our proof-of-concept, the Cracking KD-Tree, an adaptive indexing approach that generates a KD-Tree based on multidimensional range query predicates. It works by incrementally creating partial multidimensional indexes as a by-product of query processing. The indexes are produced only on those parts of the data that are accessed, and their creation cost is effectively distributed across a stream of queries. Experimental results show that the Cracking KD-Tree is three times faster than creating a full KD-Tree, one order of magnitude faster than executing full scans and two orders of magnitude faster than using uni-dimensional full or adaptive indexes on multiple columns. Pedro Holanda, Matheus Nerone, Eduardo C. de Almeida, Stefan Manegold |
DATA | 4 |
| 2018 | Deep Integration of Machine Learning Into Column StoresabstractWe leverage vectorized User-Defined Functions (UDFs) to efficiently integrate unchanged machine learning pipelines into an analytical data management system.The entire pipelines including data, models, parameters and evaluation outcomes are stored and executed inside the database system.Experiments using our MonetDB/Python UDFs show greatly improved performance due to reduced data movement and parallel processing opportunities.In addition, this integration enables meta-analysis of models using relational queries. Mark Raasveldt, Pedro Holanda, Hannes Mühleisen, Stefan Manegold |
EDBT | 4 |
| 2018 | GeoTriples: Transforming geospatial data into RDF graphs using R2RML and RML mappings
Kostis Kyzirakos, Dimitrianos Savva, Ioannis Vlachopoulos, Alexandros Vasileiou 0003, Nikolaos Karalis, Manolis Koubarakis, Stefan Manegold |
J. Web Semant. | 7 |
| 2015 | Capturing the Laws of (Data) Nature
Hannes Mühleisen, Martin L. Kersten, Stefan Manegold |
CIDR | 3 |
| 2015 | The DBMS - your big data sommelierabstractWhen addressing the problem of “big” data volume, preparation costs are one of the key challenges: the high costs for loading, aggregating and indexing data leads to a long data-to-insight time. In addition to being a nuisance to the end-user, this latency prevents real-time analytics on “big” data. Fortunately, data often comes in semantic chunks such as files that contain data items that share some characteristics such as acquisition time or location. A data management system that exploits this trait can significantly lower the data preparation costs and the associated data-to-insight time by only investing in the preparation of the relevant chunks. In this paper, we develop such a system as an extension of an existing relational DBMS (MonetDB). To this end, we develop a query processing paradigm and data storage model that are partial-loading aware. The result is a system that can make a 1.2 TB dataset (consisting of 4000 chunks) ready for querying in less than 3 minutes on a single server-class machine while maintaining good query processing performance. Yagiz Kargin, Martin L. Kersten, Stefan Manegold, Holger Pirk |
ICDE | 3 |
| 2015 | Holistic Indexing in Main-memory Column-storesabstractGreat database systems performance relies heavily on index tuning, i.e., creating and utilizing the best indices depending on the workload. However, the complexity of the index tuning process has dramatically increased in recent years due to ad-hoc workloads and shortage of time and system resources to invest in tuning. Eleni Petraki, Stratos Idreos, Stefan Manegold |
SIGMOD Conference | 3 |
| 2015 | Front Matter
Stefan Manegold |
Proc. VLDB Endow. | 1 |
| 2014 | Database cracking: fancy scan, not poor man's sort!abstractDatabase Cracking is an appealing approach to adaptive indexing: on every range-selection query, the data is partitioned using the supplied predicates as pivots. The core of database cracking is, thus, pivoted partitioning. While pivoted partitioning, like scanning, requires a single pass through the data it tends to have much higher costs due to lower CPU efficiency. In this paper, we conduct an in-depth study of the reasons for the low CPU efficiency of pivoted partitioning. Based on the findings, we develop an optimized version with significantly higher (single-threaded) CPU efficiency. We also develop a number of multi-threaded implementations that are effectively bound by memory bandwidth. Combining all of these optimizations we achieve an implementation that has costs close to or better than an ordinary scan on a variety of systems ranging from low-end (cheaper than $300) desktop machines to high-end (above $60,000) servers. Holger Pirk, Eleni Petraki, Stratos Idreos, Stefan Manegold, Martin L. Kersten |
DaMoN | 4 |
| 2014 | Waste not... Efficient co-processing of relational dataabstractThe variety of memory devices in modern computer systems holds opportunities as well as challenges for data management systems. In particular, the exploitation of Graphics Processing Units (GPUs) and their fast memory has been studied quite intensively. However, current approaches treat GPUs as systems in their own right and fail to provide a generic strategy for efficient CPU/GPU cooperation. We propose such a strategy for relational query processing: calculating an approximate result based on lossily compressed, GPU-resident data and refine the result using residuals, i.e., the lost data, on the CPU.We developed the required algorithms, implemented the strategy in an existing DBMS and found up to 8 times performance improvement, even for datasets larger than the available GPU memory. Holger Pirk, Stefan Manegold, Martin L. Kersten |
ICDE | 2 |
| 2014 | Transactional support for adaptive indexing
Goetz Graefe, Felix Halim, Stratos Idreos, Harumi A. Kuno, Stefan Manegold, Bernhard Seeger |
VLDB J. | 5 |
| 2013 | Real-time wildfire monitoring using scientific database and linked data technologiesabstractWe present a real-time wildfire monitoring service that exploits satellite images and linked geospatial data to detect hotspots and monitor the evolution of fire fronts. The service makes heavy use of scientific database technologies (array databases, SciQL, data vaults) and linked data technologies (ontologies, linked geospatial data, stSPARQL) and is implemented on top of MonetDB and Strabon. The service is now operational at the National Observatory of Athens and has been used during the previous summer by emergency managers monitoring wildfires in Greece. Manolis Koubarakis, Charalambos Kontoes, Stefan Manegold |
EDBT | 3 |
| 2013 | Enhanced stream processing in a DBMS kernelabstractContinuous query processing has emerged as a promising query processing paradigm with numerous applications. A recent development is the need to handle both streaming queries and typical one-time queries in the same application. For example, data warehousing can greatly benefit from the integration of stream semantics, i.e., online analysis of incoming data and combination with existing data. This is especially useful to provide low latency in data-intensive analysis in big data warehouses that are augmented with new data on a daily basis. Erietta Liarou, Stratos Idreos, Stefan Manegold, Martin L. Kersten |
EDBT | 3 |
| 2013 | CPU and cache efficient management of memory-resident databasesabstractMemory-Resident Database Management Systems (MRDBMS) have to be optimized for two resources: CPU cycles and memory bandwidth. To optimize for bandwidth in mixed OLTP/OLAP scenarios, the hybrid or Partially Decomposed Storage Model (PDSM) has been proposed. However, in current implementations, bandwidth savings achieved by partial decomposition come at increased CPU costs. To achieve the aspired bandwidth savings without sacrificing CPU efficiency, we combine partially decomposed storage with Just-in-Time (JiT) compilation of queries, thus eliminating CPU inefficient function calls. Since existing cost based optimization components are not designed for JiT-compiled query execution, we also develop a novel approach to cost modeling and subsequent storage layout optimization. Our evaluation shows that the JiT-based processor maintains the bandwidth savings of previously presented hybrid query processors but outperforms them by two orders of magnitude due to increased CPU efficiency. Holger Pirk, Florian Funke 0001, Martin Grund, Thomas Neumann 0001, Ulf Leser, Stefan Manegold, Alfons Kemper, Martin L. Kersten |
ICDE | 6 |
| 2013 | SciQL: array data processing inside an RDBMSabstractScientific discoveries increasingly rely on the ability to efficiently grind massive amounts of experimental data using database technologies. To bridge the gap between the needs of the Data-Intensive Research fields and the current DBMS technologies, we have introduced SciQL (pronounced as 'cycle'). SciQL is the first SQL-based declarative query language for scientific applications with both tables and arrays as first class citizens. It provides a seamless symbiosis of array-, set- and sequence- interpretations. A key innovation is the extension of value-based grouping of SQL:2003 with structural grouping, i.e., group array elements based on their positions. This leads to a generalisation of window-based query processing with wide applicability in science domains. Ying Zhang 0027, Martin L. Kersten, Stefan Manegold |
SIGMOD Conference | 3 |
| 2013 | Data vaults: a database welcome to scientific file repositoriesabstractEfficient management and exploration of high-volume scientific file repositories have become pivotal for advancement in science. We propose to demonstrate the Data Vault, an extension of the database system architecture that transparently opens scientific file repositories for efficient in-database processing and exploration. Milena Ivanova, Yagiz Kargin, Martin L. Kersten, Stefan Manegold, Ying Zhang 0027, Mihai Datcu, Daniela Espinoza-Molina |
SSDBM | 4 |
| 2013 | Front Matter
Sihem Amer-Yahia, Stefan Manegold |
Proc. VLDB Endow. | 2 |
| 2013 | Hardware-Oblivious Parallelism for In-Memory Column-StoresabstractThe multi-core architectures of today's computer systems make parallelism a necessity for performance critical applications. Writing such applications in a generic, hardware-oblivious manner is a challenging problem: Current database systems thus rely on labor-intensive and error-prone manual tuning to exploit the full potential of modern parallel hardware architectures like multi-core CPUs and graphics cards. We propose an alternative design for a parallel database engine, based on a single set of hardware-oblivious operators, which are compiled down to the actual hardware at runtime. This design reduces the development overhead for parallel database engines, while achieving competitive performance to hand-tuned systems. We provide a proof-of-concept for this design by integrating operators written using the parallel programming framework OpenCL into the open-source database MonetDB. Following this approach, we achieve efficient, yet highly portable parallel code without the need for optimization by hand. We evaluated our implementation against MonetDB using TPC-H derived queries and observed a performance that rivals that of MonetDB's query execution on the CPU and surpasses it on the GPU. In addition, we show that the same set of operators runs nearly unchanged on a GPU, demonstrating the feasibility of our approach. Max Heimel, Michael Saecker, Holger Pirk, Stefan Manegold, Volker Markl |
Proc. VLDB Endow. | 4 |
| 2013 | Lazy ETL in Action: ETL Technology Dates Scientific DataabstractBoth scientific data and business data have analytical needs. Analysis takes place after a scientific data warehouse is eagerly filled with all data from external data sources (repositories). This is similar to the initial loading stage of Extract, Transform, and Load (ETL) processes that drive business intelligence. ETL can also help scientific data analysis. However, the initial loading is a time and resource consuming operation. It might not be entirely necessary, e.g. if the user is interested in only a subset of the data. We propose to demonstrate Lazy ETL, a technique to lower costs for initial loading. With it, ETL is integrated into the query processing of the scientific data warehouse. For a query, only the required data items are extracted, transformed, and loaded transparently on-the-fly. The demo is built around concrete implementations of Lazy ETL for seismic data analysis. The seismic data warehouse is ready for query processing, without waiting for long initial loading. The audience fires analytical queries to observe the internal mechanisms and modifications that realize each of the steps; lazy extraction, transformation, and loading. Yagiz Karæz, Milena Ivanova, Ying Zhang 0027, Stefan Manegold, Martin L. Kersten |
Proc. VLDB Endow. | 4 |
| 2012 | X-device query processing by bitwise distributionabstractThe diversity of hardware components within a single system calls for strategies for efficient cross-device data processing. For example, existing approaches to CPU/GPU co-processing distribute individual relational operators to the "most appropriate" device. While pleasantly simple, this strategy has a number of problems: it may leave the "inappropriate" devices idle while overloading the "appropriate" device and putting a high pressure on the PCI bus. To address these issues we distribute data among the devices by partially decomposing relations at the granularity of individual bits. Each of the resulting bit-partitions is stored and processed on one of the available devices. Using this strategy, we implemented a processor for spatial range queries that makes efficient use of all available devices. The performance gains achieved indicate that bitwise distribution makes a good cross-device processing strategy. Holger Pirk, Thibault Sellam, Stefan Manegold, Martin L. Kersten |
DaMoN | 3 |
| 2012 | Adaptive indexing in modern database kernelsabstractPhysical design represents one of the hardest problems for database management systems. Without proper tuning, systems cannot achieve good performance. Offline indexing creates indexes a priori assuming good workload knowledge and idle time. More recently, online indexing monitors the workload trends and creates or drops indexes online. Adaptive indexing takes another step towards completely automating the tuning process of a database system, by enabling incremental and partial online indexing. The main idea is that physical design changes continuously, adaptively, partially, incrementally and on demand while processing queries as part of the execution operators. As such it brings a plethora of opportunities for rethinking and improving every single corner of database system design. Stratos Idreos, Stefan Manegold, Goetz Graefe |
EDBT | 2 |
| 2012 | Data Vaults: A Symbiosis between Database Technology and Scientific File Repositories
Milena Ivanova, Martin L. Kersten, Stefan Manegold |
SSDBM | 3 |
| 2012 | Concurrency Control for Adaptive IndexingabstractAdaptive indexing initializes and optimizes indexes incrementally, as a side effect of query processing. The goal is to achieve the benefits of indexes while hiding or minimizing the costs of index creation. However, index-optimizing side effects seem to turn read-only queries into update transactions that might, for example, create lock contention. This paper studies concurrency control in the context of adaptive indexing. We show that the design and implementation of adaptive indexing rigorously separates index structures from index contents ; this relaxes the constraints and requirements during adaptive indexing compared to those of traditional index updates. Our design adapts to the fact that an adaptive index is refined continuously, and exploits any concurrency opportunities in a dynamic way. A detailed experimental analysis demonstrates that (a) adaptive indexing maintains its adaptive properties even when running concurrent queries, (b) adaptive indexing can exploit the opportunity for parallelism due to concurrent queries, (c) the number of concurrency conflicts and any concurrency administration overheads follow an adaptive behavior, decreasing as the workload evolves and adapting to the workload needs. Goetz Graefe, Felix Halim, Stratos Idreos, Harumi A. Kuno, Stefan Manegold |
Proc. VLDB Endow. | 5 |
| 2012 | TELEIOS: A Database-Powered Virtual Earth ObservatoryabstractTELEIOS is a recent European project that addresses the need for scalable access to petabytes of Earth Observation data and the discovery and exploitation of knowledge that is hidden in them. TELEIOS builds on scientific database technologies (array databases, SciQL, data vaults) and Semantic Web technologies (stRDF and stSPARQL) implemented on top of a state of the art column store database system (MonetDB). We demonstrate a first prototype of the TELEIOS Virtual Earth Observatory (VEO) architecture, using a forest fire monitoring application as example. Manolis Koubarakis, Kostis Kyzirakos, Manos Karpathiotakis, Charalampos Nikolaou, Stavros Vassos, George Garbis, Michael Sioutis, Konstantina Bereta, Dimitrios Michail 0001, Charalambos Kontoes, Ioannis Papoutsis, Themos Herekakis, Stefan Manegold, Martin L. Kersten, Milena Ivanova, Holger Pirk, Ying Zhang 0027, Mihai Datcu, Gottfried Schwarz, Corneliu Octavian Dumitru, Daniela Espinoza-Molina, Katrin Molch, Ugo Di Giammatteo, Manuela Sagona, Sergio Perelli, Thorsten Reitz, Eva Klien, Robert Gregor |
Proc. VLDB Endow. | 13 |
| 2012 | MonetDB/DataCell: Online Analytics in a Streaming Column-StoreabstractIn DataCell, we design streaming functionalities in a modern relational database kernel which targets big data analytics. This includes exploitation of both its storage/execution engine and its optimizer infrastructure. We investigate the opportunities and challenges that arise with such a direction and we show that it carries significant advantages for modern applications in need for online analytics such as web logs, network monitoring and scientific data management. The major challenge then becomes the efficient support for specialized stream features, e.g., multi-query processing and incremental window-based processing as well as exploiting standard DBMS functionalities in a streaming environment such as indexing. This demo presents DataCell, an extension of the MonetDB open-source column-store for online analytics. The demo gives users the opportunity to experience the features of DataCell such as processing both stream and persistent data and performing window based processing. The demo provides a visual interface to monitor the critical system components, e.g., how query plans transform from typical DBMS query plans to online query plans, how data flows through the query plans as the streams evolve, how DataCell maintains intermediate results in columnar form to avoid repeated evaluation of the same stream portions, etc. The demo also provides the ability to interactively set the test scenarios and various DataCell knobs. Erietta Liarou, Stratos Idreos, Stefan Manegold, Martin L. Kersten |
Proc. VLDB Endow. | 3 |
| 2011 | Merging What's Cracked, Cracking What's Merged: Adaptive Indexing in Main-Memory Column-StoresabstractAdaptive indexing is characterized by the partial creation and refinement of the index as side effects of query execution. Dynamic or shifting workloads may benefit from preliminary index structures focused on the columns and specific key ranges actually queried --- without incurring the cost of full index construction. The costs and benefits of adaptive indexing techniques should therefore be compared in terms of initialization costs, the overhead imposed upon queries, and the rate at which the index converges to a state that is fully-refined for a particular workload component. Based on an examination of database cracking and adaptive merging, which are two techniques for adaptive indexing, we seek a hybrid technique that has a low initialization cost and also converges rapidly. We find the strengths and weaknesses of database cracking and adaptive merging complementary. One has a relatively high initialization cost but converges rapidly. The other has a low initialization cost but converges relatively slowly. We analyze the sources of their respective strengths and explore the space of hybrid techniques. We have designed and implemented a family of hybrid algorithms in the context of a column-store database system. Our experiments compare their behavior against database cracking and adaptive merging, as well as against both traditional full index lookup and scan of unordered data. We show that the new hybrids significantly improve over past methods while at least two of the hybrids come very close to the "ideal performance" in terms of both overhead per query and convergence to a final state. Stratos Idreos, Stefan Manegold, Harumi A. Kuno, Goetz Graefe |
Proc. VLDB Endow. | 2 |
| 2011 | The Researcher's Guide to the Data Deluge: Querying a Scientific Database in Just a Few Seconds
Martin L. Kersten, Stratos Idreos, Stefan Manegold, Erietta Liarou |
Proc. VLDB Endow. | 3 |
| 2010 | ROX: The robustness of a run-time XQuery optimizer against correlated dataabstractWe demonstrate ROX, a run-time optimizer of XQueries, that focuses on finding the best execution order of XPath steps and relational joins in an XQuery. The problem of join ordering has been extensively researched, but the proposed techniques are still unsatisfying. These either rely on a cost model which might result in inaccurate estimations, or explore only a restrictive number of plans from the search space. ROX is developed to tackle these problems. ROX does not need any cost model, and defers query optimization to run-time intertwining optimization and execution steps. In every optimization step, sampling techniques are used to estimate the cardinality of unexecuted steps and joins to make a decision which sequence of operators to process next. Consequently, each execution step will provide updated and accurate knowledge about intermediate results, which will be used during the next optimization round. This demonstration will focus on: (i) illustrating the steps that ROX follows and the decisions it makes to choose a good join order, (ii) showing ROX's robustness in the face of data with different degree of correlation, (iii) comparing the performance of the plan chosen by ROX to different plans picked from the search space, (iv) proving that the run-time overhead needed by ROX is restricted to a small fraction of the execution time. Riham Abdel Kader, Peter Boncz, Stefan Manegold, Maurice van Keulen |
ICDE | 3 |
| 2009 | Performance evaluation in database research: principles and experienceabstractA significant part of today's database research focuses on improving performance of a specific system. Quantitative experiments are the best way to validate such results. However, performing experiments is not always easy. Besides the complexity of the system under test, designing an experiment, choosing the right environment and parameter values, analyzing the data which is gathered, and reporting it to a third party in an expressive and intelligible way is hard.In this tutorial, we present a general road-map to the above steps, including tips and tricks on how to organize and present code that performs experiments, so that an outsider can repeat them.The tutorial is primarily aimed at MS and PhD students seeking to improve their experiment practices, but more senior attendants may also find it interesting. Stefan Manegold, Ioana Manolescu |
EDBT | 1 |
| 2009 | Self-organizing tuple reconstruction in column-storesabstractColumn-stores gained popularity as a promising physical design alternative. Each attribute of a relation is physically stored as a separate column allowing queries to load only the required attributes. The overhead incurred is on-the-fly tuple reconstruction for multi-attribute queries. Each tuple reconstruction is a join of two columns based on tuple IDs, making it a significant cost component. The ultimate physical design is to have multiple presorted copies of each base table such that tuples are already appropriately organized in multiple different orders across the various columns. This requires the ability to predict the workload, idle time to prepare, and infrequent updates. Stratos Idreos, Martin L. Kersten, Stefan Manegold |
SIGMOD Conference | 3 |
| 2009 | ROX: run-time optimization of XQueriesabstractOptimization of complex XQueries combining many XPath steps and joins is currently hindered by the absence of good cardinality estimation and cost models for XQuery. Additionally, the state-of-the-art of even relational query optimization still struggles to cope with cost model estimation errors that increase with plan size, as well as with the effect of correlated joins and selections. Riham Abdel Kader, Peter Boncz, Stefan Manegold, Maurice van Keulen |
SIGMOD Conference | 3 |
| 2009 | Database Architecture Evolution: Mammals Flourished long before Dinosaurs became ExtinctabstractThe holy grail for database architecture research is to find a solution that is Scalable & Speedy , to run on anything from small ARM processors up to globally distributed compute clusters, Stable & Secure , to service a broad user community, Small & Simple , to be comprehensible to a small team of programmers, Self-managing , to let it run out-of-the-box without hassle. In this paper, we provide a trip report on this quest, covering both past experiences, ongoing research on hardware-conscious algorithms, and novel ways towards self-management specifically focused on column store solutions. Peter Boncz, Stefan Manegold, Martin L. Kersten |
Proc. VLDB Endow. | 2 |
| 2008 | An empirical evaluation of XQuery processors
Stefan Manegold |
Inf. Syst. | 1 |
| 2008 | Column-store support for RDF data management: not all swans are whiteabstractThis paper reports on the results of an independent evaluation of the techniques presented in the VLDB 2007 paper "Scalable Semantic Web Data Management Using Vertical Partitioning", authored by D. Abadi, A. Marcus, S. R. Madden, and K. Hollenbach [1]. We revisit the proposed benchmark and examine both the data and query space coverage. The benchmark is extended to cover a larger portion of the query space in a canonical way. Repeatability of the experiments is assessed using the code base obtained from the authors. Inspired by the proposed vertically-partitioned storage solution for RDF data and the performance figures using a column-store, we conduct a complementary analysis of state-of-the-art RDF storage solutions. To this end, we employ MonetDB/SQL, a fully-functional open source column-store, and a well-known -- for its performance -- commercial row-store DBMS. We implement two relational RDF storage solutions -- triple-store and vertically-partitioned -- in both systems. This allows us to expand the scope of [1] with the performance characterization along both dimensions -- triple-store vs. vertically-partitioned and row-store vs. column-store -- individually, before analyzing their combined effects. A detailed report of the experimental test-bed, as well as an in-depth analysis of the parameters involved, clarify the scope of the solution originally presented and position the results in a broader context by covering more systems. Lefteris Sidirourgos, Romulo Goncalves, Martin L. Kersten, Niels Nes, Stefan Manegold |
Proc. VLDB Endow. | 5 |
| 2007 | Database Cracking
Stratos Idreos, Martin L. Kersten, Stefan Manegold |
CIDR | 3 |
| 2007 | Updating a cracked databaseabstractA cracked database is a datastore continuously reorganized based on operations being executed. For each query, the data of interest is physically reclustered to speed-up future access to the same, overlapping or even disjoint data. This way, a cracking DBMS self-organizes and adapts itself to the workload. Stratos Idreos, Martin L. Kersten, Stefan Manegold |
SIGMOD Conference | 3 |
| 2007 | Performance Evaluation and Experimental Assessment - Conscience or Curse of Database Research?
Ioana Manolescu, Stefan Manegold |
VLDB | 2 |
| 2006 | MonetDB/XQuery-Consistent and Efficient Updates on the Pre/Post Plane
Peter Boncz, Jan Flokstra, Torsten Grust, Maurice van Keulen, Stefan Manegold, K. Sjoerd Mullender, Jan Rittinger, Jens Teubner |
EDBT | 5 |
| 2006 | MonetDB/XQuery: a fast XQuery processor powered by a relational engineabstractRelational XQuery systems try to re-use mature relational data management infrastructures to create fast and scalable XML database technology. This paper describes the main features, key contributions, and lessons learned while implementing such a system. Its architecture consists of (i) a range-based encoding of XML documents into relational tables, (ii) a compilation technique that translates XQuery into a basic relational algebra, (iii) a restricted (order) property-aware peephole relational query optimization strategy, and (iv) a mapping from XML update statements into relational updates. Thus, this system implements all essential XML database functionalities (rather than a single feature) such that we can learn from the full consequences of our architectural decisions. While implementing this system, we had to extend the state-of-the-art with a number of new technical contributions, such as loop-lifted staircase join and efficient relational query evaluation strategies for XQuery theta-joins with existential semantics. These contributions as well as the architectural lessons learned are also deemed valuable for other relational back-end engines. The performance and scalability of the resulting system is evaluated on the XMark benchmark up to data sizes of 11GB. The performance section also provides an extensive benchmark comparison of all major XMark results published previously, which confirm that the goal of purely relational XQuery processing, namely speed and scalability, was met. Peter Boncz, Torsten Grust, Maurice van Keulen, Stefan Manegold, Jan Rittinger, Jens Teubner |
SIGMOD Conference | 4 |
| 2005 | Cracking the Database Store
Martin L. Kersten, Stefan Manegold |
CIDR | 2 |
| 2005 | Pathfinder: XQuery - The Relational Way
Peter Boncz, Torsten Grust, Maurice van Keulen, Stefan Manegold, Jan Rittinger, Jens Teubner |
VLDB | 4 |
| 2004 | Cache-Conscious Radix-Decluster Projections
Stefan Manegold, Peter Boncz, Niels Nes |
VLDB | 1 |
| 2002 | Generic Database Cost Models for Hierarchical Memory Systems
Stefan Manegold, Peter Boncz, Martin L. Kersten |
VLDB | 1 |
| 2002 | Optimizing Main-Memory Join on Modern HardwareabstractIn the past decade, the exponential growth in commodity CPU's speed has far outpaced advances in memory latency. A second trend is that CPU performance advances are not only brought by increased clock rates, but also by increasing parallelism inside the CPU. Current database systems have not yet adapted to these trends and show poor utilization of both CPU and memory resources on current hardware. In this paper, we show how these resources can be optimized for large joins and translate these insights into guidelines for future database architectures, encompassing data structures, algorithms, cost modeling and implementation. In particular, we discuss how vertically fragmented data structures optimize cache performance on sequential data access. On the algorithmic side, we refine the partitioned hash-join with a new partitioning algorithm called "radix-cluster", which is specifically designed to optimize memory access. The performance of this algorithm is quantified using a detailed analytical model that incorporates memory access costs in terms of a limited number of parameters, such as cache sizes and miss penalties. We also present a calibration tool that extracts such parameters automatically from any computer hardware. The accuracy of our models is proven by exhaustive experiments conducted with the Monet database system on three different hardware platforms. Finally, we investigate the effect of implementation techniques that optimize CPU resource usage. Our experiments show that large joins can be accelerated almost an order of magnitude on modern RISC hardware when both memory and CPU resources are optimized. Stefan Manegold, Peter Boncz, Martin L. Kersten |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2000 | What Happens During a Join? Dissecting CPU and Memory Optimization Effects
Stefan Manegold, Peter Boncz, Martin L. Kersten |
VLDB | 1 |
| 2000 | Optimizing database architecture for the new bottleneck: memory access
Stefan Manegold, Peter Boncz, Martin L. Kersten |
VLDB J. | 1 |
| 1999 | Database Architecture Optimized for the New Bottleneck: Memory Access
Peter Boncz, Stefan Manegold, Martin L. Kersten |
VLDB | 2 |