Maximilian Rieger

dblp:354/9687 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0007-8887-2960ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 72% Information retrieval · 28%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › SQL query processing
nested query processing
0.912025
Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins · Proc. ACM Manag. Data 2025
Query processing and optimization › cost estimation
query execution time prediction
0.912025
T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision Trees · Proc. ACM Manag. Data 2025
Information retrieval › evaluation
query performance prediction
0.912025
T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision Trees · Proc. ACM Manag. Data 2025
Storage systems › data management › database storage
columnar storage
0.912025
Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins · Proc. ACM Manag. Data 2025
Query processing and optimization
join processing
0.312025
Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins · Proc. ACM Manag. Data 2025
Query processing and optimization › query planning
query plan representation
0.312025
T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision Trees · Proc. ACM Manag. Data 2025

Methods — techniques the papers use, named apart from their topics

on-the-fly key generation · 1.7join-based nesting reconstruction · 1.7decision tree · 0.9compiled native code · 0.9
YearPublicationVenuePosition
2025 Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins
abstract
Parquet is the most commonly used file format to store data in a columnar, binary structure. The format also supports storing nested data in this flattened columnar layout. However, many query engines either do not support nested data or process it with substantially worse performance than relational data. In this work, we close this gap and present a new way to leverage relational query engines for nested data that is stored in this flat columnar file format. Specifically, we demonstrate how to process nested Parquet files much more efficiently. Our approach does not store a copy of the data in an internal format but reads directly from the Parquet file. During query computation, the required flat columns are scanned independently and the nesting is reconstructed using joins with on-the-fly generated join keys. Our approach can be easily integrated into existing query engines to support querying nested Parquet files. Furthermore, we achieve orders of magnitude faster analytical query performance than existing solutions, which makes it a valuable addition.
Alice Rey, Maximilian Rieger, Thomas Neumann 0001
Proc. ACM Manag. Data2
2025 T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision Trees
abstract
Query performance prediction is used for scheduling, resource scaling, tenant placement, and various other use-cases. Here, the main goal is to estimate the execution time of a query without running it. To be effective, predictors need to be both accurate and fast. In contrast, neural networks that were used in recent work deliver very accurate predictions but suffer from high latency. In this work, we propose the Tuple Time Tree (T3), a new model that is both accurate and fast. It is orders of magnitude faster than comparable methods and has competitive accuracy to state-of-the-art approaches. Additionally, T3 works for new database instances without re-training because it generalizes across database instances. We achieve T3's speed by relying on a low-latency decision tree model that is compiled to native machine code. We maintain high accuracy with two novel techniques: pipeline-based query plan representation and tuple-centric prediction targets. In our pipeline-based query plan representation, T3 decomposes query plans into pipelines. Then, T3 predicts the execution time of each pipeline individually, instead of the whole query in one step. With tuple-centric prediction targets, T3 predicts the expected time it takes to push a single tuple through a pipeline. It then multiplies this predicted value by the input cardinality of the pipeline to estimate its execution time. As a result, T3 achieves state-of-the-art accuracy with a low-latency decision tree model.
Maximilian Rieger, Thomas Neumann 0001
Proc. ACM Manag. Data1