Ricardo Usbeck

dblp:65/9656 · DBLP profile ↗
← Back
22ranked-venue papers in the field
4as first author
9since 2021 · last 2025
0000-0002-0191-7211ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 13 (3 first)Information Retrieval & Web Search · 6 (1 first)Database Systems & Data Management · 1Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 ReportGRI: Automating GRI Alignment and Report Assessment
abstract
Organisations disclose their sustainability performance in corporate sustainability reports (CSRs). CSRs vary widely in structure and depth depending on the reporting framework. Such disparity, together with report complexity and volume, poses significant challenges to transparency, comparability and standardisation. To address this problem, we introduce ReportGRI, an automated system for Global Reporting Initiative (GRI) indexing and qualitative assessment of CSRs. The interactive framework leverages information retrieval techniques and zero-shot prompting to enable GRI disclosure-based report indexing and report coverage assessment by visualising well-covered topics and reporting gaps. The tool facilitates scalable and explainable benchmarking of Environmental, Social and Governance (ESG) reporting quality, enhancing report interpretation, transparency, and corporate accountability. The system is open-sourced on GitHub with an introduction video.
Aida Usmanova, Rana Abdullah, Debayan Banerjee, Markus Leippold, Ricardo Usbeck
CIKM5
2025 Analyzing the Influence of Knowledge Graph Information on Relation Extraction
Cedric Möller, Ricardo Usbeck
ESWC (1)2
2025 LLM Agents for Georelating - A New Task for Locating Events
abstract
Accurately identifying disaster-affected areas is crucial for data-driven disaster resilience. In response, we introduce Georelating, a task that infers affected areas from textual reports containing complex locative expressions, moving beyond traditional geoparsing approaches that rely on explicit point locations. Georelating instead combines resolving unnamed regions and reasoning about spatial relations to represent event-affected areas within standardized Discrete Global Grid Systems (DGGSs).
Kai Moltzen, Junbo Huang, Ricardo Usbeck
SIGSPATIAL/GIS3
2025 DBLP QuAD 2.0: Scholarly Natural Questions from SPARQL
abstract
We present DBLP-QuAD 2.0, designed to evaluate Scholarly Knowledge Graph Question Answering (KGQA) over DBLP. Recent updates in the underlying DBLP KG, including new entities and relationships such as venues, research streams, and citation links, have necessitated a corresponding update to existing KG QA benchmarking resources. While the DBLP-QuAD dataset focused on author and publication-centered queries, DBLP-QuAD 2.0 broadens the coverage to reflect the enriched structure of the updated KG. Specifically, the questions in our dataset are formulated from SPARQL query logs that cover a wide range of entities involving authors, publications, venues, research streams, and citation relationships. DBLP-QuAD 2.0 thus provides a more comprehensive benchmark for evaluating KGQA systems with a baseline.
Tilahun Abedissa Taffa, Patrick Neises, Stefan Ollinger, Patrick Westphal, Marcel R. Ackermann, Debayan Banerjee, Ricardo Usbeck
K-CAP7
2025 ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs
Mikhail Salnikov, Andrey Sakhovskiy, Irina Nikishina, Aida Usmanova, Angelie Kraft, Cedric Möller, Debayan Banerjee, Junbo Huang, Longquan Jiang 0001, Rana Abdullah, Xi Yan 0001, Elena Tutubalina, Ricardo Usbeck, Alexander Panchenko
NLDB (1)13
2024 DISCIE-Discriminative Closed Information Extraction
Cedric Möller, Ricardo Usbeck
ISWC (2)2
2023 GETT-QA: Graph Embedding Based T2T Transformer for Knowledge Graph Question Answering
Debayan Banerjee, Pranav Ajit Nair, Ricardo Usbeck, Chris Biemann
ESWC3
2022 Modern Baselines for SPARQL Semantic Parsing
abstract
In this work, we focus on the task of generating SPARQL queries from natural language questions, which can then be executed on Knowledge Graphs (KGs). We assume that gold entity and relations have been provided, and the remaining task is to arrange them in the right order along with SPARQL vocabulary, and input tokens to produce the correct SPARQL query. Pre-trained Language Models (PLMs) have not been explored in depth on this task so far, so we experiment with BART, T5 and PGNs (Pointer Generator Networks) with BERT embeddings, looking for new baselines in the PLM era for this task, on DBpedia and Wikidata KGs. We show that T5 requires special input tokenisation, but produces state of the art performance on LC-QuAD 1.0 and LC-QuAD 2.0 datasets, and outperforms task-specific models from previous works. Moreover, the methods enable semantic parsing for questions where a part of the input needs to be copied to the output query, thus enabling a new paradigm in KG semantic parsing.
Debayan Banerjee, Pranav Ajit Nair, Jivat Neet Kaur, Ricardo Usbeck, Chris Biemann
SIGIR4
2022 Knowledge Graph Question Answering Datasets and Their Generalizability: Are They Enough for Future Research?
abstract
Existing approaches on Question Answering over Knowledge Graphs (KGQA) have weak generalizability. That is often due to the standard i.i.d. assumption on the underlying dataset. Recently, three levels of generalization for KGQA were defined, namely i.i.d., compositional, zero-shot. We analyze 25 well-known KGQA datasets for 5 different Knowledge Graphs (KGs). We show that according to this definition many existing and online available KGQA datasets are either not suited to train a generalizable KGQA system or that the datasets are based on discontinued and out-dated KGs. Generating new datasets is a costly process and, thus, is not an alternative to smaller research groups and companies. In this work, we propose a mitigation method for re-splitting available KGQA datasets to enable their applicability to evaluate generalization, without any cost and manual effort. We test our hypothesis on three KGQA datasets, i.e., LC-QuAD, LC-QuAD 2.0 and QALD-9). Experiments on re-splitted KGQA datasets demonstrate its effectiveness towards generalizability. The code and a unified way to access 18 available datasets is online at https://github.com/semantic-systems/KGQA-datasets as well as https://github.com/semantic-systems/KGQA-datasets-generalization.
Longquan Jiang 0001, Ricardo Usbeck
SIGIR2
2018 Why Reinvent the Wheel: Let's Build Question Answering Systems Together
abstract
Modern question answering (QA) systems need to flexibly integrate a number of components specialised to fulfil specific tasks in a QA pipeline. Key QA tasks include Named Entity Recognition and Disambiguation, Relation Extraction, and Query Building. Since a number of different software components exist that implement different strategies for each of these tasks, it is a major challenge to select and combine the most suitable components into a QA system, given the characteristics of a question. We study this optimisation problem and train classifiers, which take features of a question as input and have the goal of optimising the selection of QA components based on those features. We then devise a greedy algorithm to identify the pipelines that include the suitable components and can effectively answer the given question. We implement this model within Frankenstein, a QA framework able to select QA components and compose QA pipelines. We evaluate the effectiveness of the pipelines generated by Frankenstein using the QALD and LC-QuAD benchmarks. These results not only suggest that Frankenstein precisely solves the QA optimisation problem but also enables the automatic composition of optimised QA pipelines, which outperform the static Baseline QA pipeline. Thanks to this flexible and fully automated pipeline generation process, new QA components can be easily included in Frankenstein, thus improving the performance of the generated pipelines.
Kuldeep Singh 0001, Arun Sethupat Radhakrishna, Andreas Both 0001, Saeedeh Shekarpour, Ioanna Lytra, Ricardo Usbeck, Akhilesh Vyas, Akmal Khikmatullaev, Dharmen Punjani, Christoph Lange 0002, Maria-Esther Vidal, Jens Lehmann 0001, Sören Auer
WWW6
2017 Holistic and scalable ranking of RDF data
abstract
The volume and number of data sources published using Semantic Web standards such as RDF grows continuously. The largest of these data sources now contain billions of facts and are updated periodically. A large number of applications driven by such data sources requires the ranking of entities and facts contained in such knowledge graphs. Hence, there is a need for time-efficient approaches that can compute ranks for entities and facts simultaneously. In this paper, we present the first holistic ranking approach for RDF data. Our approach, dubbed HARE, allows the simultaneous computation of ranks for RDF triples, resources, properties and literals. To this end, HARE relies on the representation of RDF graphs as bi-partite graphs. It then employs a time-efficient extension of the random walk paradigm to bi-partite graphs. We show that by virtue of this extension, the worst-case complexity of HARE is O(n5) while that of PageRank is O(n6). In addition, we evaluate the practical efficiency of our approach by comparing it with PageRank on 6 real and 6 synthetic datasets with sizes up to 108triples. Our results show that HARE is up to 2 orders of magnitude faster than PageRank. We also present a brief evaluation of HARE's ranking accuracy by comparing it with that of PageRank applied directly to RDF graphs. Our evaluation on 19 classes of DBpedia demonstrates that there is no statistical difference between HARE and PageRank. We hence conclude that our approach goes beyond the state of the art by allowing the ranking of all RDF entities and of RDF triples without being worse w.r.t. the ranking quality it achieves on resources. HARE is open-source and is available at http://github.com/dice-group/hare.
Axel-Cyrille Ngonga Ngomo, Michael Hoffmann 0007, Ricardo Usbeck, Kunal Jha
IEEE BigData3
2017 MAG: A Multilingual, Knowledge-base Agnostic and Deterministic Entity Linking Approach
abstract
Entity linking has recently been the subject of a significant body of research. Currently, the best performing approaches rely on trained mono-lingual models. Porting these approaches to other languages is consequently a difficult endeavor as it requires corresponding training data and retraining of the models. We address this drawback by presenting a novel multilingual, knowledge-base agnostic and deterministic approach to entity linking, dubbed MAG. MAG is based on a combination of context-based retrieval on structured knowledge bases and graph algorithms. We evaluate MAG on 23 data sets and in 7 languages. Our results show that the best approach trained on English datasets (PBOH) achieves a micro F-measure that is up to 4 times worse on datasets in other languages. MAG on the other hand achieves state-of-the-art performance on English datasets and reaches a micro F-measure that is up to 0.6 higher than that of PBOH on non-English languages.
Diego Moussallem, Ricardo Usbeck, Michael Röder, Axel-Cyrille Ngonga Ngomo
K-CAP2
2017 GENESIS: a generic RDF data access interface
abstract
The availability of billions of facts represented in RDF on the Web provides novel opportunities for data discovery and access. In particular, keyword search and question answering approaches enable even lay people to access this data. However, the interpretation of the results of these systems, as well as the navigation through these results, remains challenging. In this paper, we present Genesis, a generic RDF data access interface. Genesis can be deployed on top of any knowledge base and search engine with minimal effort and allows for the representation of RDF data in a layperson-friendly way. This is facilitated by the modular architecture for reusable components underlying our framework. Currently, these include a generic search back-end, together with corresponding interactive user interface components based on a service for similar and related entities as well as verbalization services to bridge between RDF and natural language.
Timofey Ermilov, Diego Moussallem, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo
WI3
2016 CubeQA - Question Answering on RDF Data Cubes
Konrad Höffner, Jens Lehmann 0001, Ricardo Usbeck
ISWC (1)3
2015 HAWK - Hybrid Question Answering Using Linked Data
Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Lorenz Bühmann, Christina Unger
ESWC1
2015 ASSESS - Automatic Self-Assessment Using Linked Data
Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo
ISWC (2)2
2015 GERBIL: General Entity Annotator Benchmarking Framework
abstract
We present GERBIL, an evaluation framework for semantic entity annotation. The rationale behind our framework is to provide developers, end users and researchers with easy-to-use interfaces that allow for the agile, fine-grained and uniform evaluation of annotation tools on multiple datasets. By these means, we aim to ensure that both tool developers and end users can derive meaningful insights pertaining to the extension, integration and use of annotation applications. In particular, GERBIL provides comparable results to tool developers so as to allow them to easily discover the strengths and weaknesses of their implementations with respect to the state of the art. With the permanent experiment URIs provided by our framework, we ensure the reproducibility and archiving of evaluation results. Moreover, the framework generates data in machine-processable format, allowing for the efficient querying and post-processing of evaluation results. Finally, the tool diagnostics provided by GERBIL allows deriving insights pertaining to the areas in which tools should be further refined, thus allowing developers to create an informed agenda for extensions and end users to detect the right tools for their purposes. GERBIL aims to become a focal point for the state of the art, driving the research agenda of the community by presenting comparable objective evaluation results.
Ricardo Usbeck, Michael Röder, Axel-Cyrille Ngonga Ngomo, Ciro Baron, Andreas Both 0001, Martin Brümmer, Diego Ceccarelli, Marco Cornolti, Didier Cherix, Bernd Eickmann, Paolo Ferragina, Christiane Lemke, Andrea Moro 0001, Roberto Navigli, Francesco Piccinno, Giuseppe Rizzo 0002, Harald Sack, René Speck, Raphaël Troncy, Jörg Waitelonis, Lars Wesemann
WWW1
2015 DeFacto - Temporal and multilingual Deep Fact Validation
Daniel Gerber, Diego Esteves, Jens Lehmann 0001, Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, René Speck
J. Web Semant.5
2014 Combining Linked Data and Statistical Information Retrieval - Next Generation Information Systems
Ricardo Usbeck
ESWC1
2014 Web-Scale Extension of RDF Knowledge Bases from Templated Websites
Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Muhammad Saleem 0002, Andreas Both 0001, Valter Crescenzi, Paolo Merialdo, Disheng Qiu
ISWC (1)2
2014 AGDISTIS - Graph-Based Disambiguation of Named Entities Using Linked Data
Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Michael Röder, Daniel Gerber, Sandro A. Coelho, Sören Auer, Andreas Both 0001
ISWC (1)1
2013 Real-Time RDF Extraction from Unstructured Data Streams
Daniel Gerber, Sebastian Hellmann 0001, Lorenz Bühmann, Tommaso Soru, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo
ISWC (1)5