VLDB 2026 Research / reviewers in the wild / expert
Marc Maynou
dblp:347/1146
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0004-3908-4392ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 67% Distributed and cloud data management · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed and cloud data management
data lake |
1.0 | 1 | 2026 | Freyja: Efficient Join Discovery in Data Lakes · IEEE Trans. Knowl. Data Eng. 2026 |
Data integration and cleaning
data profiling |
1.0 | 1 | 2026 | Freyja: Efficient Join Discovery in Data Lakes · IEEE Trans. Knowl. Data Eng. 2026 |
Data integration and cleaning › table discovery
joinable table discovery |
1.0 | 1 | 2026 | Freyja: Efficient Join Discovery in Data Lakes · IEEE Trans. Knowl. Data Eng. 2026 |
Methods — techniques the papers use, named apart from their topics
predictive model · 1.0multiset jaccard · 1.0cardinality proportion · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Freyja: Efficient Join Discovery in Data LakesabstractWe study the problem of efficiently computing rankings of joinable attributes in data lakes. Traditional set-overlap measures produce numerous false positives in this scenario, while modern, more accurate Table Representation Learning (TRL) techniques incur prohibitive computational costs. In contrast to the state-of-the-art, we adopt a novel notion of join quality tailored to data lakes relying on a metric that combines multiset Jaccard and cardinality proportion. The proposed metric merges the best of both worlds by leveraging syntactic measures while achieving accuracy scores comparable to those of TRL approaches. Generating rankings of joinable pairs is highly scalable at both preparation and query time, since we train a general-purpose predictive model. Predictions are based on data profiles, succinct and efficiently computed representations of dataset characteristics. Our experiments show that our system, Freyja, matches and improves upon, the results obtained by the state-of-the-art while reducing execution costs by orders of magnitude. Marc Maynou, Sergi Nadal, Raquel Panadero, Javier Flores 0002, Oscar Romero 0001, Anna Queralt |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Supporting Data Discovery Tasks at Scale with FREYJA
Marc Maynou, Sergi Nadal |
EDBT | 1 |
| 2024 | Discovery of Semantic Non-Syntactic Joins
Marc Maynou, Sergi Nadal |
DOLAP | 1 |
| 2023 | Generating valid test data through data cloningabstractOne of the most difficult, time-consuming and error-prone tasks during software testing is that of manually generating the data required to properly run the test. This is even harder when we need to generate data of a certain size and such that it satisfies a set of conditions, or business rules, specified over an ontology. To solve this problem, some proposals exist to automatically generate database sample data. However, they are only able to generate data satisfying primary or foreign key constraints but not more complex business rules in the ontology. We propose here a more general solution for generating test data which is able to deal with expressive business rules. Our approach, which is entirely based on the chase algorithm, first generates a small sample of valid test data (by means of an automated reasoner), then clones this sample data, and finally, relates the cloned data with the original data. All the steps are performed iteratively until a valid database of a certain size is obtained. We theoretically prove the correctness of our approach, and experimentally show its practical applicability. Xavier Oriol, Ernest Teniente, Marc Maynou, Sergi Nadal |
Future Gener. Comput. Syst. | 3 |