Till Döhmen

dblp:169/5218 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0001-7483-0099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Distilled documentation for dialect-specific SQL generation
Till Döhmen, Adithya Krishnan, Hamilton Ulmer, Peter Boncz, Sebastian Schelter
VLDB J.1
2024 MotherDuck: DuckDB in the cloud and in the client
R. J. Atwal, Peter Boncz, Ryan Boyd, Antony Courtney, Till Döhmen, Florian Gerlinghoff, Jeff Huang 0007, Joseph Hwang, Raphael Hyde, Elena Felder, Jacob Lacouture, Yves Le Maout, Boaz Leskes, Alex Monahan, Dan Perkins, Tino Tereshko, Jordan Tigani, Nick Ursa, Stephanie Wang, Yannick Welsch
CIDR5
2024 SchemaPile: A Large Collection of Relational Database Schemas
abstract
Access to fine-grained schema information is crucial for understanding how relational databases are designed and used in practice, and for building systems that help users interact with them. Furthermore, such information is required as training data to leverage the potential of large language models (LLMs) for improving data preparation, data integration and natural language querying. Existing single-table corpora such as GitTables provide insights into how tables are structured in-the-wild, but lack detailed schema information about how tables relate to each other, as well as metadata like data types or integrity constraints. On the other hand, existing multi-table (or database schema) datasets are rather small and attribute-poor, leaving it unclear to what extent they actually represent typical real-world database schemas. In order to address these challenges, we present SchemaPile, a corpus of 221,171 database schemas, extracted from SQL files on GitHub. It contains 1.7 million tables with 10 million column definitions, 700 thousand foreign key relationships, seven million integrity constraints, and data content for more than 340 thousand tables. We conduct an in-depth analysis on the millions of schema metadata properties in our corpus, as well as its highly diverse language and topic distribution. In addition, we showcase the potential of \corpus to improve a variety of data management applications, e.g., fine-tuning LLMs for schema-only foreign key detection, improving CSV header detection and evaluating multi-dialect SQL parsers. We publish the code and data for recreating SchemaPile and a permissively licensed subset SchemaPile-Perm.
Till Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian Schelter
Proc. ACM Manag. Data1
2023 Interpreting Black-box Machine Learning Models for High Dimensional Datasets
abstract
Many datasets are of increasingly high dimension- ality, where a large number of features could be irrelevant to the learning task. The inclusion of such features would not only introduce unwanted noise but also increase computational complexity. Deep neural networks (DNNs) outperform machine learning (ML) algorithms in a variety of applications due to their effectiveness in modelling complex problems and handling high-dimensional datasets. However, due to non-linearity and higher-order feature interactions, DNN models are unavoidably opaque, making them black-box methods. In contrast, an interpretable model can identify statistically significant features and explain the way they affect the model’s outcome. In this paper, we propose a novel method to improve the interpretability of blackbox models in the case of high-dimensional datasets. First, a black-box model is trained on full feature space that learns useful embeddings on which the classification is performed. To decompose the inner principles of the black-box and to identify top-k important features (global explainability), probing and perturbing techniques are applied. An interpretable surrogate model is then trained on top-k feature space to approximate the black-box. Finally, decision rules and counterfactuals are derived from the surrogate to provide local decisions. Our approach outperforms tabular learners, e.g., TabNet and XGboost, and SHAP-based interpretability techniques, when tested on a number of datasets having dimensionality between 54 and 20,5311.1GitHub: https://github.com/rezacsedu/DeepExplainHidim
Md. Rezaul Karim 0001, Md Shajalal, Alexander Graß, Till Döhmen, Sisay Adugna Chala, Alexander Boden, Christian Beecks, Stefan Decker
DSAA4
2020 DeepCOVIDExplainer: Explainable COVID-19 Diagnosis from Chest X-ray Images
abstract
In this paper1, we proposed an explainable deep neural networks (DNN)-based method for automatic detection of COVID-19 symptoms from chest radiography (CXR) images, which we call ‘DeepCOVIDExplainer’. We used 15,959 CXR images of 15,854 patients, covering normal, pneumonia, and COVID-19 cases. CXR images are first comprehensively preprocessed and augmented before classifying with a neural ensemble method, followed by highlighting class-discriminating regions using gradient-guided class activation maps (Grad-CAM ++) and layer-wise relevance propagation (LRP). Further, we provide human-interpretable explanations for the diagnosis. Evaluation results show that our approach can identify COVID-19 cases with a positive predictive value (PPV) of 91.6%, 92.45%, and 96.12%, respectively for normal, pneumonia, and COVID-19 cases, respectively, outperforming recent approaches.1Read longer version of this paper: https://arxiv.org/pdf/2004.04582.pdf
Md. Rezaul Karim 0001, Till Döhmen, Michael Cochez, Oya Beyan, Dietrich Rebholz-Schuhmann, Stefan Decker
BIBM2
2017 Multi-Hypothesis CSV Parsing
abstract
Comma Separated Value (CSV) files are commonly used to represent data. CSV is a very simple format, yet we show that it gives rise to a surprisingly large amount of ambiguities in its parsing and interpretation. We summarize the state-of-the-art in CSV parsers, which typically make a linear series of parsing and interpretation decisions, such that any wrong decision at an earlier stage can negatively affect all downstream decisions. Since computation time is much less scarce than human time, we propose to turn CSV parsing into a ranking problem. Our quality-oriented multi-hypothesis CSV parsing approach generates several concurrent hypotheses about dialect, table structure, etc. and ranks these hypotheses based on quality features of the resulting table. This approach makes it possible to create an advanced CSV parser that makes many different decisions, yet keeps the overall parser code a simple plug-in infrastructure. The complex interactions between these decisions are taken care of by searching the hypothesis space rather than by having to program these many interactions in code. We show that our approach leads to better parsing results than the state of the art and facilitates the parsing of large corpora of heterogeneous CSV files.
Till Döhmen, Hannes Mühleisen, Peter Boncz
SSDBM1
2016 Towards a Benchmark for the Maintainability Evolution of Industrial Software Systems
abstract
The maintainability of software is an important cost factor for organizations across all industries, as maintenance makes up approximately 40% to 70% of the total development costs of a software system. Organizations are often stuck in the situation where software maintenance costs dominate IT budgets, leaving no room for enhancement and innovation. Building a benchmark for maintainability evolution is helpful in this context because it can help organizations decide on software improvement or replacement strategies. The prototype benchmark we study in this paper shows that software volume and maintainability levels are strong determinants of future maintainability evolution rates. We further describe the data collection and cleaning procedures that were applied to approximately 1,750 industrial software systems, and we provide an exploratory analysis of the resulting benchmark dataset.
Till Döhmen, Magiel Bruntink, Davide Ceolin, Joost Visser 0001
IWSM-Mensura1