Basil Ell

dblp:25/8883 · DBLP profile ↗
← Back
12ranked-venue papers in the field
3as first author
8since 2021 · last 2025
0000-0002-8863-3157ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 10 (3 first)Information Retrieval & Web Search · 2
YearPublicationVenuePosition
2025 Finding Good Neighbors: Examining the Importance of Neighborhood Selection for Link Prediction
abstract
Link Prediction (LP) approaches based on Language Models (LMs) operate over the labels and descriptions of entities and relations in a Knowledge Graph (KG). Recent approaches have shown that incorporating a local graph neighborhood can improve the LP capabilities of LMs. These approaches usually sample a context from the neighborhood around a query triple randomly, thereby incorporating noise that might hinder the model in making correct predictions.
Moritz Blum, Moritz Plenz, Basil Ell, Philipp Cimiano
K-CAP3
2024 Numerical Literals in Link Prediction: A Critical Examination of Models and Datasets
Moritz Blum, Basil Ell, Hannes Ill, Philipp Cimiano
ISWC (1)2
2024 ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries
abstract
Dataset search, or more specifically, ad hoc dataset retrieval which is a trending specialized IR task, has received increasing attention in both academia and industry. While methods and systems continue evolving, existing test collections for this task exhibit shortcomings, particularly suffering from lexical bias in pooling and limited to keyword-style queries for evaluation. To address these limitations, in this paper, we construct ACORDAR 2.0, a new test collection for this task which is also the largest to date. To reduce lexical bias in pooling, we adapt dense retrieval models to large structured data, using them to find an extended set of semantically relevant datasets to be annotated. To diversify query forms, we employ a large language model to rewrite keyword queries into high-quality question-style queries. We use the test collection to evaluate popular sparse and dense retrieval models to establish a baseline for future studies. The test collection and source code are publicly available.
Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang 0001, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
SIGIR7
2023 LexExMachinaQA: A framework for the automatic induction ofontology lexica for Question Answering over Linked Data
Mohammad Fazleh Elahi, Basil Ell, Philipp Cimiano
LDK2
2023 Human-Machine Collaborative Annotation: A Case Study with GPT-3
Ole Magnus Holter, Basil Ell
LDK2
2022 ACORDAR: A Test Collection for Ad Hoc Content-Based (RDF) Dataset Retrieval
abstract
Ad hoc dataset retrieval is a trending topic in IR research. Methods and systems are evolving from metadata-based to content-based ones which exploit the data itself for improving retrieval accuracy but thus far lack a specialized test collection. In this paper, we build and release the first test collection for ad hoc content-based dataset retrieval, where content-oriented dataset queries and content-based relevance judgments are annotated by human experts who are assisted with a dashboard designed specifically for comprehensively and conveniently browsing both the metadata and data of a dataset. We conduct extensive experiments on the test collection to analyze its difficulty and provide insights into the underlying task.
Tengteng Lin, Qiaosheng Chen, Gong Cheng 0001, Ahmet Soylu, Basil Ell, Ruoqi Zhao, Xiaxia Wang 0001, Yu Gu 0016, Evgeny Kharlamov
SIGIR5
2021 Bridging the Gap Between Ontology and Lexicon via Class-Specific Association Rules Mined from a Loosely-Parallel Text-Data Corpus
abstract
There is a well-known lexical gap between content expressed in the form of natural language (NL) texts and content stored in an RDF knowledge base (KB). For tasks such as Information Extraction (IE), this gap needs to be bridged from NL to KB, so that facts extracted from text can be represented in RDF and can then be added to an RDF KB. For tasks such as Natural Language Generation, this gap needs to be bridged from KB to NL, so that facts stored in an RDF KB can be verbalized and read by humans. In this paper we propose LexExMachina, a new methodology that induces correspondences between lexical elements and KB elements by mining class-specific association rules. As an example of such an association rule, consider the rule that predicts that if the text about a person contains the token "Greek", then this person has the relation nationality to the entity Greece. Another rule predicts that if the text about a settlement contains the token "Greek", then this settlement has the relation country to the entity Greece. Such a rule can help in question answering, as it maps an adjective to the relevant KB terms, and it can help in information extraction from text. We propose and empirically investigate a set of 20 types of class-specific association rules together with different interestingness measures to rank them. We apply our method on a loosely-parallel text-data corpus that consists of data from DBpedia and texts from Wikipedia, and evaluate and provide empirical evidence for the utility of the rules for Question Answering.
Basil Ell, Mohammad Fazleh Elahi, Philipp Cimiano
LDK1
2021 Towards Scope Detection in Textual Requirements
abstract
Requirements are an integral part of industry operation and projects. Not only do requirements dictate industrial operations, but they are used in legally binding contracts between supplier and purchaser. Some companies even have requirements as their core business. Most requirements are found in textual documents, this brings a couple of challenges such as ambiguity, scalability, maintenance, and finding relevant and related requirements. Having the requirements in a machine-readable format would be a solution to these challenges, however, existing requirements need to be transformed into machine-readable requirements using NLP technology. Using state-of-the-art NLP methods based on end-to-end neural modelling on such documents is not trivial because the language is technical and domain-specific and training data is not available. In this paper, we focus on one step in that direction, namely scope detection of textual requirements using weak supervision and a simple classifier based on BERT general domain word embeddings and show that using openly available data, it is possible to get promising results on domain-specific requirements documents.
Ole Magnus Holter, Basil Ell
LDK2
2015 Learning a Cross-Lingual Semantic Representation of Relations Expressed in Text
Achim Rettinger, Artem Schumilin, Steffen Thoma, Basil Ell
ESWC4
2014 SPARQL Query Verbalization for Explaining Semantic Search Engine Queries
Basil Ell, Andreas Harth, Elena Simperl
ESWC1
2011 Labels in the Web of Data
Basil Ell, Denny Vrandecic, Elena Simperl
ISWC (1)1
2010 Semantic MediaWiki in Operation: Experiences with Building a Semantic Portal
Daniel M. Herzig, Basil Ell
ISWC (2)2