Martin Pekár Christensen

dblp:380/1888 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-3168-6810ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 An End-To-End Re-Evaluation of Table Entity-Linkers
abstract
Abstract Knowledge graph (KG) entity linkers link entity mentions from a source data representation to their corre sponding entities in a target KG. A knowledge graph (KG) is a popular graph model expressing semantic information about entities, concepts, and relationships. Consequently, entity linking from tables to KGs has become increasingly important, allowing semantic table augmentation and other advanced data integration tasks. However, existing evaluations of entity linkers are incomplete, as they only evaluate for specific applications and only consider aggregated output quality metrics: they do not consider the performance and effectiveness of the individual entity linking components along with their scalability. To address this gap, we provide an in-depth analysis and taxonomy of state of-the-art entity linkers by thoroughly evaluating the quality and scalability of the individual entity linking components. Hence, we evaluate entity linkers on four existing entity linking benchmarks using DBpedia and Wikidata and identify major bottlenecks to be overcome to make these entity linkers fully applicable in real-world use-cases. We identified candidate generation as the most crucial entity linking step, which commonly is overlooked. Furthermore, we show that most entity linkers are irreproducible, either because they are not open-source, or because they use irreproducible public endpoints and datasets. Acknowledgements This research was partially funded by the Danish Council for Independent Research (DFF), grant agreement no. DFF 804800051B, the Poul Due Jensen Fond, and DataGEMS, funded by the European Union’s Horizon Europe Research and Innovation programme, grant agreement no. 101188416.
Martin Pekár Christensen, Matteo Lissandrini, Katja Hose
ICDE1
2026 Jazero: A Semantic Table Search System
abstract
Abstract Finding relevant tables is challenging and commonly performed using keyword, join, or union search. However, these are not always able to retrieve all the relevant tables since they usually require some form of exact matches or structural similarity. Semantic table search has recently been proposed as a novel technique that exploits knowledge graphs and various forms of semantic similarity for the retrieval of semantically relevant data lake tables. Specifically, semantic table search computes the relevance of a table based on the contained entity mentions without requiring strict structural similarities. This enables the discovery of a much larger set of relevant tables compared to existing approaches. In this paper, we demonstrate Jazero, the first semantic search system for data lakes featuring various semantic similarity metrics. Jazero uses the query-by-example paradigm, where a user provides an example table containing data of interest, enabling entity-centric discovery of data lake tables w.r.t. the query entity tuples. Jazero further enables scalable semantic table search with search space pre-filtering using the popular hierarchical navigable small world index overdifferent vector representations for knowledge graph entities, leading to an average runtime improvement of 87.4%. This corresponds to an average runtime of 3.3s. Jazero allows users to experience this novel discovery paradigm with various entity representations and similarity functions and to perform data discovery in multiple data lake instances. Jazerohas proven to retrieve near-disjoint results to keyword search, whilst retaining the same NDCG ranking performance. Acknowledgements This research was partially funded by the Danish Council for Independent Research (DFF), grant agreement no. DFF 804800051B, the Poul Due Jensen Fond, and DataGEMS, funded by the European Union’s Horizon Europe Research and Innovation programme, grant agreement no. 101188416.
Martin Pekár Christensen, Matteo Lissandrini, Katja Hose
ICDE1
2025 Fantastic Tables and Where to Find Them: Table Search in Semantic Data Lakes
abstract
In data lakes, one of the core challenges remains finding relevant tables. We introduce the notion of semantic data lakes, i.e., repositories where datasets are linked to concepts and entities described in a knowledge graph (KG). We formalize the problem of semantic table search, i.e., retrieving tables containing information semantically related to a given set of entities, and provide the first formal definition of semantic relatedness of a dataset to tuples of entities. Our solution offers the first general framework to compute the semantic relevance of the contents of a table w.r.t. entity tuples, as well as efficient algorithms (exploiting semantic signals, such as entity types and embeddings) to scale the semantic search to repositories with hundreds of thousands of distinct tables. Our extensive experiments on both real-world and synthetic benchmarks show that our approach is able to retrieve more relevant tables (up to 5.4 times higher recall) in comparison to existing methods while ensuring fast response times (up to 17 times faster with LSH).
Martin Pekár Christensen, Aristotelis Leventidis, Matteo Lissandrini, Laura Di Rocco, Renée J. Miller, Katja Hose
EDBT1
2024 A Large Scale Test Corpus for Semantic Table Search
abstract
Table search aims to answer a query with a ranked list of tables. Unfortunately, current test corpora have focused mostly on needle-in-the-haystack tasks, where only a few tables are expected to exactly match the query intent. Instead, table search tasks often arise in response to the need for retrieving new datasets or augmenting existing ones, e.g., for data augmentation within data science or machine learning pipelines. Existing table repositories and benchmarks are limited in their ability to test retrieval methods for table search tasks. Thus, to close this gap, we introduce a novel dataset for query-by-example Semantic Table Search. This novel dataset consists of two snapshots of the large-scale Wikipedia tables collection from 2013 and 2019 with two important additions: (1) a page and topic aware ground truth relevance judgment and (2) a large-scale DBpedia entity linking annotation. Moreover, we generate a novel set of entity-centric queries that allows testing existing methods under a novel search scenario: semantic exploratory search. The resulting resource consists of 9,296 novel queries, 610,553 query-table relevance annotations, and 238,038 entity-linked tables from the 2013 snapshot. Similarly, on the 2019 snapshot, the resource consists of 2,560 queries, 958,214 relevance annotations, and 457,714 total tables. This makes our resource the largest annotated table-search corpus to date (97 times more queries and 956 times more annotated tables than any existing benchmark). We perform a user study among domain experts and prove that these annotators agree with the automatically generated relevance annotations. As a result, we can re-evaluate some basic assumptions behind existing table search approaches identifying their shortcomings along with promising novel research directions.
Aristotelis Leventidis, Martin Pekár Christensen, Matteo Lissandrini, Laura Di Rocco, Katja Hose, Renée J. Miller
SIGIR2