EDBT 2026 Demo / reviewers in the wild / expert
Knut Stolze
dblp:s/KnutStolze
· DBLP profile ↗
8ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Prefiltering for Massive-Scale Equi- and Geospatial Joins
Jason Arnold, Knut Stolze, Ellis Saupe |
EDBT | 2 |
| 2026 | A Declarative, Recursive SQL Framework for Composable Machine Learning Ensembles
Jason Arnold, Knut Stolze, Dan Zollers, Poojan Khanpara |
EDBT | 2 |
| 2023 | SPA: Economical and Workload-Driven Indexing for Data Analytics in the CloudabstractSelective queries are not uncommon in large-scale data analytics, for example, when drilling down into a specific customer in a dashboard. Traditionally, selective queries are accelerated by creating secondary indexes. However, because of their large size, expensive maintenance, and difficulty to tune and automate, indexes are typically not used in modern cloud data warehouses or data lakes. Instead, such systems rely mostly on full table scans and lightweight optimizations like min/max filtering, whose effectiveness depends heavily on the data layout and value distributions.We propose SPA as the vision for automatically optimizing selective queries for immutable copy-on-write data formats. SPA adaptively indexes subsets of the data in an incremental and workload-driven manner. It makes fine-grained decisions and continuously monitors their benefit, dynamically allocating an optimization budget in a way that bounds the additional cost of indexing. Furthermore, it guarantees a performance improvement in the cases where indexes—potentially partial ones—prove to be beneficial. When indexes lose their benefit due to a shifting workload, they are gradually deconstructed in favor of optimizations that accommodate recent trends. As SPA does not require information about updates performed on the data, it can also be employed as an accelerator for systems that do not control the data, e.g., in cloud data lakes. Peter Boncz, Yannis Chronis, Jan Finis, Stefan Halfpap, Viktor Leis, Thomas Neumann 0001, Anisoara Nica, Caetano Sauer, Knut Stolze, Marcin Zukowski |
ICDE | 9 |
| 2020 | Replication at the Speed of Change - a Fast, Scalable Replication Solution for Near Real-Time HTAP ProcessingabstractThe IBM Db2 Analytics Accelerator (IDAA) is a state-of-the art hybrid database system that seamlessly extends the strong transactional capabilities of Db2 for z/OS with the very fast column-store processing in Db2 Warehouse. The Accelerator maintains a copy of the data from Db2 for z/OS in its backend database. Data can be synchronized at a single point in time with a granularity of a table, one or more of its partitions, or incrementally as rows changed using replication technology. IBM Change Data Capture (CDC) has been employed as replication technology since IDAA version 3. Version 7.5.0 introduces a superior replication technology as a replacement for IDAA's use of CDC - Integrated Synchronization. In this paper, we present how Integrated Synchronization improves the performance by orders of magnitudes, paving the way for near real-time Hybrid Transactional and Analytical (HTAP) processing. Dennis Butterstein, Knut Stolze, Felix Beier, Jia Zhong |
Proc. VLDB Endow. | 3 |
| 2016 | Extending Database Accelerators for Data Transformations and Predictive AnalyticsabstractThe IBM DB2 Analytics Accelerator (IDAA) integrates the strong OLTP capabilities of DB2 for z/OS with very fast processing of OLAP workloads using Netezza technology. The accelerator is attached to DB2 as analytical process- ing resource { completely transparent for user applications. But all data modi_cations must be carried out by DB2 and are replicated to the accelerator internally. However, this behavior is not optimized for ELT processing and predic- tive analytics or data mining workloads where multi-staged data transformations are involved. We present our work for extending IDAA with accelerator-only tables, which enable direct data transformations without any necessary interven- tions by DB2. Further, we present a framework for executing arbitrary in-database analytics operations on the accelerator while ensuring data governance aspects like privilege man- agement on DB2 and allowing to ingest data from any other source directly to the accelerator to enrich analytics e. g., with social media data. The evolutionary framework design maintains compatibility with existing infrastructure and ap- plications, a must-have for the majority of customers, while allowing complex analytics beyond read-only reporting. Felix Beier, Knut Stolze |
EDBT | 2 |
| 2014 | Joins on Encoded and Partitioned DataabstractCompression has historically been used to reduce the cost of storage, I/Os from that storage, and buffer pool utilization, at the expense of the CPU required to decompress data every time it is queried. However, significant additional CPU efficiencies can be achieved by deferring decompression as late in query processing as possible and performing query processing operations directly on the still-compressed data. In this paper, we investigate the benefits and challenges of performing joins on compressed (or encoded) data. We demonstrate the benefit of independently optimizing the compression scheme of each join column, even though join predicates relating values from multiple columns may require translation of the encoding of one join column into the encoding of the other. We also show the benefit of compressing "payload" data other than the join columns "on the fly," to minimize the size of hash tables used in the join. By partitioning the domain of each column and defining separate dictionaries for each partition, we can achieve even better overall compression as well as increased flexibility in dealing with new values introduced by updates. Instead of decompressing both join columns participating in a join to resolve their different compression schemes, our system performs a light-weight mapping of only qualifying rows from one of the join columns to the encoding space of the other at run time. Consequently, join predicates can be applied directly on the compressed data. We call this procedure encoding translation. Two alternatives of encoding translation are developed and compared in the paper. We provide a comprehensive evaluation of these alternatives using product implementations of each on the TPC-H data set, and demonstrate that performing joins on encoded and partitioned data achieves both superior performance and excellent compression. Jae-Gil Lee 0001, Gopi K. Attaluri, Ron Barber, Naresh Chainani, Oliver Draese, Frederick Ho, Stratos Idreos, Min-Soo Kim 0002, Sam Lightstone, Guy M. Lohman, Konstantinos Morfonios, Keshava Murthy, Ippokratis Pandis, Lin Qiao 0001, Vijayshankar Raman, Vincent KulandaiSamy, Richard Sidle, Knut Stolze |
Proc. VLDB Endow. | 18 |
| 2012 | WYSIWYE: An Algebra for Expressing Spatial and Textual Rules for Information Extraction
Vijil Chenthamarakshan, Ramakrishna Varadarajan, Prasad Deshpande, Raghu Krishnapuram, Knut Stolze |
WAIM | 5 |
| 2011 | Online reorganization in read optimized MMDBSabstractQuery performance is a critical factor in modern business intelligence and data warehouse systems. An increasing number of companies uses detailed analyses for conducting daily business and supporting management decisions. Thus, several techniques have been developed for achieving near realtime response times - techniques which try to alleviate I/O bottlenecks while increasing the throughputs of available processing units, i.e. by keeping relevant data in compressed main-memory data structures and exploiting the read-only characteristics of analytical workloads. Felix Beier, Knut Stolze, Kai-Uwe Sattler |
SIGMOD Conference | 2 |