VLDB 2026 Research / reviewers in the wild / expert
Mihail Stoian
dblp:255/9132
· DBLP profile ↗
8ranked-venue papers in the field
4as first author
7since 2021 · last 2026
0000-0002-8843-3374ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 8 (4 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Waiting to Decompress: The Economics of LLM-Based Compression
Andreas Kipf, Ping-Lin Kuo, Skander Krid, Moritz Rengert, Luca Heller, Andreas Zimmerer, Mihail Stoian, Varun Pandey, Alexander van Renen |
CIDR | 8 |
| 2026 | Redbench: Workload Synthesis From Cloud Traces
Johannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid, Martin Stemmer, Andreas Kipf, Carsten Binnig, Muhammad El-Hindi |
Proc. VLDB Endow. | 3 |
| 2025 | Virtual: Compressing Data Lake Files
Mihail Stoian, Alexander van Renen, Jan Kobiolka, Ping-Lin Kuo, Andreas Zimmerer, Josif Grabocka, Andreas Kipf |
EDBT | 1 |
| 2025 | Parachute: Single-Pass Bi-Directional Information PassingabstractSideways information passing is a well-known technique for mitigating the impact of large build sides in a database query plan. As currently implemented in production systems, sideways information passing enables only a uni-directional information flow, as opposed to instance-optimal algorithms, such as Yannakakis'. On the other hand, the latter require an additional pass over the input, which hinders adoption in production systems. In this paper, we make a step towards enabling single-pass bidirectional information passing during query execution. We achieve this by statically analyzing between which tables the information flow is blocked and by leveraging precomputed join-induced fingerprint columns on FK-tables. On the JOB benchmark, Parachute improves DuckDB v1.2's end-to-end execution time without and with semi-join filtering by 1.54x and 1.24x, respectively, when allowed to use 15% extra space. Mihail Stoian, Andreas Zimmerer, Skander Krid, Amadou Ngom, Jialin Ding 0001, Tim Kraska, Andreas Kipf |
Proc. VLDB Endow. | 1 |
| 2024 | DPconv: Super-Polynomially Faster Join OrderingabstractWe revisit the join ordering problem in query optimization. The standard exact algorithm, DPccp, has a worst-case running time of O(3 n ). This is prohibitively expensive for large queries, which are not that uncommon anymore. We develop a new algorithmic framework based on subset convolution. DPconv achieves a super-polynomial speedup over DPccp, breaking the O(3 n ) time-barrier for the first time. We show that the framework instantiation for the C max cost function is up to 30x faster than DPccp for large clique queries. Mihail Stoian, Andreas Kipf |
Proc. ACM Manag. Data | 1 |
| 2024 | DataLoom: Simplifying Data Loading with LLMsabstractSchema discovery and data loading is a crucial step in any data analysis pipeline. While this used to be a rare task, in the highly dynamic field of machine learning and modern business intelligence on top of data lakes, today it has become a frequent, but often underestimated, activity. Existing tools often focus on single files, presume prior knowledge of the data on the user's side or a significant amount of manual labor. In this paper, we improve the process of mapping a "chaotic" set of files to an initial database schema that can then be iteratively refined and loaded. The idea is to take the previously tedious parts of this process and automate them through the use of Large Language Models (LLMs) while leaving already well-understood problems such as constraint discovery to existing algorithms. We thus carefully orchestrate the use of LLMs for the "soft" problems and traditional algorithms for the "hard" problems. This creates a more seamless schema discovery and data loading experience that minimizes the time to first insight for users. We show this vision on modern schema discovery and data loading in our web-based prototype called DataLoom that serves as our demonstration. Alexander van Renen, Mihail Stoian, Andreas Kipf |
Proc. VLDB Endow. | 2 |
| 2022 | Concurrent Link-Cut TreesabstractShare on Concurrent Link-Cut Trees Author: Mihail M. Stoian Technische Universität München, Munich, Germany Technische Universität München, Munich, GermanyView Profile Authors Info & Claims SIGMOD '22: Proceedings of the 2022 International Conference on Management of DataJune 2022 Pages 2503–2505https://doi.org/10.1145/3514221.3520247Online:11 June 2022Publication History 0citation47DownloadsMetricsTotal Citations0Total Downloads47Last 12 Months47Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Mihail Stoian |
SIGMOD Conference | 1 |
| 2020 | Benchmarking Learned IndexesabstractRecent advancements in learned index structures propose replacing existing index structures, like B-Trees, with approximate learned models. In this work, we present a unified benchmark that compares well-tuned implementations of three learned index structures against several state-of-the-art "traditional" baselines. Using four real-world datasets, we demonstrate that learned index structures can indeed outperform non-learned indexes in read-only in-memory workloads over a dense array. We investigate the impact of caching, pipelining, dataset size, and key size. We study the performance profile of learned index structures, and build an explanation for why learned models achieve such good performance. Finally, we investigate other important properties of learned index structures, such as their performance in multi-threaded systems and their build times. Ryan Marcus, Andreas Kipf, Alexander van Renen, Mihail Stoian, Sanchit Misra, Alfons Kemper, Thomas Neumann 0001, Tim Kraska |
Proc. VLDB Endow. | 4 |