EDBT 2026 Demo / reviewers in the wild / expert
Alberto Berenguer
dblp:311/1756
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-2867-8329ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Content-Based Catalogs for Enhanced Discovery Services in Data Spaces
Adriana Morejón, Alberto Berenguer, Lucia de Espona, David Tomás 0001, Jose-Norberto Mazón |
DOLAP | 2 |
| 2024 | Evaluating the Impact of Content Deletion on Tabular Data Similarity and Retrieval Using Contextual Word Embeddings
Alberto Berenguer, David Tomás 0001, Jose-Norberto Mazón |
ECIR (2) | 1 |
| 2024 | Word embeddings for retrieving tabular data from research publicationsabstractAbstract Scientists face challenges when finding datasets related to their research problems due to the limitations of current dataset search engines. Existing tools for searching research datasets rely on publication content or metadata, do not considering the data contained in the publication in the form of tables. Moreover, scientists require more elaborate inputs and functionalities to retrieve different parts of an article, such as data presented in tables, based on their search purposes. Therefore, this paper proposes a novel approach to retrieve relevant tabular datasets from publications. The input of our system is a research problem stated as an abstract from a scientific paper, and the output is a set of relevant tables from publications that are related to the research problem. This approach aims to provide a better solution for scientists to find useful datasets that support them in addressing their research problems. To validate this approach, experiments were conducted using word embedding from different language models to calculate the semantic similarity between abstracts and tables. The results showed that contextual models significantly outperformed non-contextual models, especially when pre-trained with scientific data. Furthermore, the importance of context was found to be crucial for improving the results. Alberto Berenguer, Jose-Norberto Mazón, David Tomás 0001 |
Mach. Learn. | 1 |
| 2023 | Tabular Open Government Data Search for Data Spaces based on Word Embeddings
Alberto Berenguer, David Tomás 0001, Jose-Norberto Mazón |
DOLAP | 1 |
| 2021 | Towards a tabular open data search engine for public sector informationabstractPublic Sector Information (PSI) scenarios require tools that support retrieval of tabular open data beyond keyword-based search on metadata. This paper presents a novel interface for searching tabular open data, as well as a search engine that retrieves tabular data by considering table contents apart from metadata. Our search engine uses word embeddings to calculate the semantic similarity between tabular open data, providing a ranking of candidate tabular datasets to be integrated with an input query table according to the different intentions of the open data reuser (e.g. column or row extension as well as data completion). An initial set of experiments have been conducted, showing promising results in this task. Alberto Berenguer, Jose-Norberto Mazón, David Tomás 0001 |
IEEE BigData | 1 |