EDBT 2026 Demo / reviewers in the wild / expert
Gerardo Vitagliano
dblp:249/4023
· DBLP profile ↗
11ranked-venue papers in the field
3as first author
11since 2021 · last 2026
0000-0001-7782-2596ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 10 (3 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao, Andreas Zimmerer, Olga Ovcharenko, Tianji Cong, Matthew Russo, Gerardo Vitagliano, Michael Cochez, Fatma Özcan 0001, Gautam Gupta, Thibaud Hottelier, H. V. Jagadish, Kris Kissel, Sebastian Schelter, Andreas Kipf, Immanuel Trummer |
Proc. VLDB Endow. | 6 |
| 2026 | Abacus: A Cost-Based Optimizer for Semantic Operator Systems
Matthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano, Michael J. Cafarella, Tim Kraska, Samuel Madden 0001 |
Proc. VLDB Endow. | 4 |
| 2025 | Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing
Chunwei Liu, Matthew Russo, Michael J. Cafarella, Lei Cao 0004, Peter Baile Chen, Zui Chen, Michael J. Franklin, Tim Kraska, Samuel Madden 0001, Rana Shahout, Gerardo Vitagliano |
CIDR | 11 |
| 2024 | TASHEEH: Repairing Row-Structure in Raw CSV Files
Mazhar Hameed 0001, Gerardo Vitagliano, Fabian Panse, Felix Naumann |
EDBT | 2 |
| 2023 | MORPHER: Structural Transformation of Ill-formed RowsabstractOpen data portals contain a plethora of data files, with comma-separated value (CSV) files being particularly popular with users and businesses due to their flexible standard. However, this flexibility comes with much responsibility for data consumers, as many files contain various structural problems, e.g., a different number of cells across data rows, multiple value formats within the same column, different variants of quoted fields due to user specifications, etc. We refer to rows that contain such structural inconsistencies as ill-formed. Consequently, ingesting them into a host system, such as a database or an analytics platform, often requires prior data preparation steps. Mazhar Hameed 0001, Gerardo Vitagliano, Felix Naumann |
CIKM | 2 |
| 2023 | Pollock: A Data Loading BenchmarkabstractAny system at play in a data-driven project has a fundamental requirement: the ability to load data. The de-facto standard format to distribute and consume raw data is csv. Yet, the plain text and flexible nature of this format make such files often difficult to parse and correctly load their content, requiring cumbersome data preparation steps. We propose a benchmark to assess the robustness of systems in loading data from non-standard csv formats and with structural inconsistencies. First, we formalize a model to describe the issues that affect real-world files and use it to derive a systematic "pollution" process to generate dialects for any given grammar. Our benchmark leverages the pollution framework for the csv format. To guide pollution, we have surveyed thousands of real-world, publicly available csv files, recording the problems we encountered. We demonstrate the applicability of our benchmark by testing and scoring 16 different systems: popular csv parsing frameworks, relational database tools, spreadsheet systems, and a data visualization tool. Gerardo Vitagliano, Mazhar Hameed 0001, Lan Jiang 0001, Lucas Reisener, Eugene Wu 0002, Felix Naumann |
Proc. VLDB Endow. | 1 |
| 2022 | Aggregation Detection in CSV Files
Lan Jiang 0001, Gerardo Vitagliano, Mazhar Hameed 0001, Felix Naumann |
EDBT | 2 |
| 2022 | SURAGH: Syntactic Pattern Matching to Identify Ill-Formed Records
Mazhar Hameed 0001, Gerardo Vitagliano, Lan Jiang 0001, Felix Naumann |
EDBT | 2 |
| 2022 | Mondrian: Spreadsheet Layout DetectionabstractSpreadsheet datasets are valuable sources of data, but often ill-suited for machine consumption. Their unstructured nature allows users to arrange data and metadata freely in a human-readable format, often in canvas-like layouts. To extract their content, data practitioners need to resort to manual inspection and run cumbersome preparation pipelines. The Mondrian system assists users in identifying and handling multiregion layout templates: spreadsheet layouts composed of independent regions that appear repeatedly across different files. Mondrian comprises an automated approach to detect multiple regions within a single file and an algorithm that leverages mapping region layouts to graphs to compute layout similarity and identify templates. Users interact with Mondrian through a web-based visual interface, that serves as a practical toolkit to handle collections of multiregion spreadsheets and enables their automated preparation. Gerardo Vitagliano, Lucas Reisener, Lan Jiang 0001, Mazhar Hameed 0001, Felix Naumann |
SIGMOD Conference | 1 |
| 2021 | Structure Detection in Verbose CSV Files
Lan Jiang 0001, Gerardo Vitagliano, Felix Naumann |
EDBT | 2 |
| 2021 | Detecting Layout Templates in Complex Multiregion FilesabstractSpreadsheets are among the most commonly used file formats for data management, distribution, and analysis. Their widespread employment makes it easy to gather large collections of data, but their flexible canvas-based structure makes automated analysis difficult without heavy preparation. One of the common problems that practitioners face is the presence of multiple, independent regions in a single spreadsheet, possibly separated by repeated empty cells. We define such files as "multiregion" files. In collections of various spreadsheets, we can observe that some share the same layout. We present the Mondrian approach to automatically identify layout templates across multiple files and systematically extract the corresponding regions. Our approach is composed of three phases: first, each file is rendered as an image and inspected for elements that could form regions; then, using a clustering algorithm, the identified elements are grouped to form regions; finally, every file layout is represented as a graph and compared with others to find layout templates. We compare our method to state-of-the-art table recognition algorithms on two corpora of real-world enterprise spreadsheets. Our approach shows the best performances in detecting reliable region boundaries within each file and can correctly identify recurring layouts across files. Gerardo Vitagliano, Lan Jiang 0001, Felix Naumann |
Proc. VLDB Endow. | 1 |