VLDB 2026 Research / reviewers in the wild / expert
Mahdi Esmailoghli
dblp:238/4380
· DBLP profile ↗
7ranked-venue papers in the field
5as first author
6since 2021 · last 2026
0009-0009-2148-6402ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 7 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Every Data Lake Has a Past: Analytical Exploration of Wikipedia History as a Temporal Data Lake
Mahdi Esmailoghli, Steven Purtzel, Roee Shraga, Renée J. Miller, Matthias Weidlich 0001 |
DOLAP | 1 |
| 2026 | Data Discovery in Data Lakes: Operations, Indexes, SystemsabstractData discovery has gained significant traction in the database community resulting in various discovery operations, index schemes, and discovery systems. This tutorial explores the architecture and components of data discovery systems, focusing on indexing structures and scalable algorithms for typical operations, such as join and union discovery. While giving insights into individual algorithms, we point out open challenges for holistic systems, data discovery evaluation, and discovery in federated setups. Ziawasch Abedjan, Mahdi Esmailoghli, Sainyam Galhotra |
ICDE | 2 |
| 2025 | BLEND: A Unified Data Discovery SystemabstractMost research on data discovery has so far focused on improving individual discovery operators such as join, correlation, or union discovery. However, in practice, a combination of these techniques and their corresponding indexes may be necessary to support arbitrary discovery tasks. We propose BLEND, a comprehensive data discovery system that supports existing operators and enables their flexible pipelining. BLEND is based on a set of lower-level operators that serve as fundamental building blocks for more complex and sophisticated user tasks. To reduce the execution runtime of discovery pipelines, we propose a unified index structure and a rule- and cost-based optimizer that rewrites SQL statements into low-level operators when possible. We show the superior flexibility and efficiency of our system compared to ad-hoc discovery pipelines and stand-alone solutions. Mahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch Abedjan |
ICDE | 1 |
| 2025 | Data Disovery in Data Lakes: Operations, Indexes, Systems
Ziawasch Abedjan, Mahdi Esmailoghli, Sainyam Galhorta |
Proc. VLDB Endow. | 2 |
| 2022 | MATE: Multi-Attribute Table ExtractionabstractA core operation in data discovery is to find joinable tables for a given table. Real-world tables include both unary and n-ary join keys. However, existing table discovery systems are optimized for unary joins and are ineffective and slow in the existence of n-ary keys. In this paper, we introduce Mate, a table discovery system that leverages a novel hash-based index that enables n-ary join discovery through a space-efficient super key. We design a filtering layer that uses a novel hash, Xash. This hash function encodes the syntactic features of all column values and aggregates them into a super key, which allows the system to efficiently prune tables with non-joinable rows. Our join discovery system is able to prune up to 1000 x more false positives and leads to over 60 x faster table discovery in comparison to state-of-the-art. Mahdi Esmailoghli, Jorge-Arnulfo Quiané-Ruiz, Ziawasch Abedjan |
Proc. VLDB Endow. | 1 |
| 2021 | COCOA: COrrelation COefficient-Aware Data Augmentation
Mahdi Esmailoghli, Jorge-Arnulfo Quiané-Ruiz, Ziawasch Abedjan |
EDBT | 1 |
| 2020 | CAFE: Constraint-Aware Feature Extraction from Large Databases
Mahdi Esmailoghli, Ziawasch Abedjan |
CIDR | 1 |