EDBT 2026 Demo / reviewers in the wild / expert
Saeed Fathollahzadeh
dblp:164/3312
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-3723-6191ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 56% Machine learning and data management · 44% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 3 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
programming by example |
0.7 | 1 | 2023 | GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example · Proc. ACM Manag. Data 2023 |
Data integration and cleaning › data preprocessing
data cleaning |
0.3 | 1 | 2025 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines · Proc. VLDB Endow. 2025 |
Data integration and cleaning
data mapping |
0.2 | 1 | 2023 | GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example · Proc. ACM Manag. Data 2023 |
Methods — techniques the papers use, named apart from their topics
program synthesis by example · 1.3multi-threaded code generation · 1.3large language model · 0.9data catalog · 0.9automatic validation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML PipelinesabstractData-centric machine learning (ML) pipelines extend traditional ML pipelines—of feature transformations, hyper-parameter tuning, and model training—by additional pre-processing steps for data cleaning, data augmentation, and feature engineering to create high-quality data with good coverage. Finding effective data-centric ML pipelines is still a labor- and compute-intensive process though. While AutoML tools use effective search strategies, they struggle to scale with large datasets. Large language models (LLMs) show promise for code generation but face challenges in generating data-centric ML pipelines due to private datasets not seen during training, complex pre-processing requirements, and the need for mitigating hallucinations. These demands exceed typical code generation as it requires actions tailored to the characteristics and requirements of a particular dataset. This paper introduces CatDB, a comprehensive, LLM-based system for generating effective, error-free, and efficient data-centric ML pipelines. CatDB leverages data catalog information and refined metadata to dynamically create dataset-specific rules (instructions) to guide the LLM. Moreover, CatDB includes a robust mechanism for automatic validation and error handling of the generated pipeline. Our experimental results show that CatDB reliably generates effective ML pipelines across diverse datasets, achieving accuracy comparable to or better than existing LLM-based systems, standalone AutoML tools, and combined workflows of data cleaning and AutoML tools, while delivering up to orders of magnitude faster performance on large datasets. Saeed Fathollahzadeh, Essam Mansour 0001, Matthias Boehm 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by ExampleabstractData Scientists deal with a wide variety of file data formats and data representations. Probably the most difficult to handle are custom data formats that liberally define their own particular flat or nested structure with multiple custom delimiters, multi-line records, or undocumented semantics of attribute sequences, co-appearances, and repetitions. As a prerequisite for exploratory ML model training, data scientists need to map these data representations into regular frames or matrices. Unfortunately, existing tools and frameworks provide only limited support for aiding this process, which causes redundant manual efforts and unnecessary data quality issues. In this paper, we initiate work on automatic matrix and frame reader generation by example. A user provides a sample of raw text data and its mapped matrix or frame representation. Our GIO framework then first identifies the mapping rules from raw to structured data, and subsequently generates source code of an efficient, multi-threaded reader for reading full raw datasets of this format. In order to facilitate manual improvements, both the mapping rules, and generated reader can be modified as needed. Our experiments show that GIO is able to correctly identify the mapping rules for basic text formats like CSV, LibSVM, MatrixMarket; custom text formats from publishing, automotive, and health care; as well as various nested formats such as JSON and XML. Additionally, the automatically generated readers yield competitive performance compared to hand-coded readers and tuned libraries like RapidJSON. Saeed Fathollahzadeh, Matthias Boehm 0001 |
Proc. ACM Manag. Data | 1 |