Andreas Both 0001

dblp:07/4780 · DBLP profile ↗
← Back
32ranked-venue papers in the field
4as first author
16since 2021 · last 2026
0000-0002-9177-5463ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 18 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 10 (1 first)Data Mining & Knowledge Discovery · 3 (1 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 DynBench Generator: A Web-Based Platform for On-Demand KGQA Benchmark Creation
Aleksandr Gashkov, Maria Eltsova, Andreas Both 0001
ICWE3
2026 Toward Reliable LLM-Integrated Web Architectures for Teacher-Aligned Automatic Student Grading
Jonas Gwozdz, Andreas Both 0001
ICWE2
2026 Toward Trustworthy, Teacher-Aligned Adaptive Learning on the Web
Jonas Gwozdz, Andreas Both 0001
ICWE2
2025 SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
Aleksandr Gashkov, Aleksandr Perevalov, Maria Eltsova, Andreas Both 0001
ICWE4
2025 Post-hoc LLM-Supported Debugging of Distributed Processes
Dennis Schiese, Andreas Both 0001
ICWE2
2024 AuthApp - Portable, Reusable Solid App for GDPR-Compliant Access Granting
Andreas Both 0001, Thorsten Kastner, Dustin Yeboah, Christoph Braun 0002, Daniel Schraudner, Sebastian Schmid 0001, Tobias Käfer, Andreas Harth
ICWE1
2024 Language Models as SPARQL Query Filtering for Improving the Quality of Multilingual Question Answering over Knowledge Graphs
Aleksandr Perevalov, Aleksandr Gashkov, Maria Eltsova, Andreas Both 0001
ICWE4
2024 Understanding SPARQL Queries: Are We Already There? Multilingual Natural Language Generation Based on SPARQL Queries and Large Language Models
Aleksandr Perevalov, Aleksandr Gashkov, Maria Eltsova, Andreas Both 0001
ISWC (2)4
2023 Lingua Franca - Entity-Aware Machine Translation Approach for Question Answering over Knowledge Graphs
abstract
This research paper proposes an approach called Lingua Franca that improves machine translation quality by utilizing information from a knowledge graph to translate named entities accurately. The accurate entity translation is crucial when applied to entity-oriented search including Knowledge Graph Question Answering systems. In a nutshell, the approach preserves recognized named entities with an entity-replacement technique during the translation process. It replaces the entities back with their labels found in a knowledge graph for the target language to ensure that questions are translated correctly before answering them using a Knowledge Graph Question Answering system. The paper also introduces an open-source modular framework that enables researchers to design their own named entity-aware machine translation pipelines. The presented experimental results demonstrate the effectiveness of the Lingua Franca approach in comparison to baseline Machine Translation models. The approach shows a statistically significant improvement in the quality provided by several Knowledge Graph Question Answering systems using Lingua Franca on different datasets.
Nikit Srivastava, Aleksandr Perevalov, Denis Kuchelev, Diego Moussallem, Axel-Cyrille Ngonga Ngomo, Andreas Both 0001
K-CAP6
2023 Information extraction pipelines for knowledge graphs
abstract
In the last decade, a large number of knowledge graph (KG) completion approaches were proposed. Albeit effective, these efforts are disjoint, and their collective strengths and weaknesses in effective KG completion have not been studied in the literature. We extend Plumber, a framework that brings together the research community's disjoint efforts on KG completion. We include more components into the architecture of Plumber to comprise 40 reusable components for various KG completion subtasks, such as coreference resolution, entity linking, and relation extraction. Using these components, Plumber dynamically generates suitable knowledge extraction pipelines and offers overall 432 distinct pipelines. We study the optimization problem of choosing optimal pipelines based on input sentences. To do so, we train a transformer-based classification model that extracts contextual embeddings from the input and finds an appropriate pipeline. We study the efficacy of Plumber for extracting the KG triples using standard datasets over three KGs: DBpedia, Wikidata, and Open Research Knowledge Graph. Our results demonstrate the effectiveness of Plumber in dynamically generating KG completion pipelines, outperforming all baselines agnostic of the underlying KG. Furthermore, we provide an analysis of collective failure cases, study the similarities and synergies among integrated components and discuss their limitations.
Mohamad Yaser Jaradeh, Kuldeep Singh 0001, Markus Stocker, Andreas Both 0001, Sören Auer
Knowl. Inf. Syst.4
2022 Improving Question Answering Quality Through Language Feature-Based SPARQL Query Candidate Validation
Aleksandr Gashkov, Aleksandr Perevalov, Maria Eltsova, Andreas Both 0001
ESWC4
2022 Towards Bridging the Gap Between Knowledge Graphs and Chatbots
Annemarie Wittig, Aleksandr Perevalov, Andreas Both 0001
ICWE3
2022 Quality Assurance of a German COVID-19 Question Answering Systems using Component-based Microbenchmarking
abstract
Question Answering (QA) has become an often used method to retrieve data as part of chatbots and other natural-language user interfaces. In particular, QA systems of official institutions have high expectations regarding the answers computed by the system, as the provided information might be critical. In this demonstration, we use the official COVID-19 QA system that was developed together with the German Federal government to provide German citizens access to data regarding incident values, number of deaths, etc. To ensure high quality, a component-based approach was used that enables exchanging data between QA components using RDF and validating the functionality of the QA system using SPARQL. Here, we will demonstrate how our solution enables developers of QA systems to use a descriptive approach to validate the quality of their implementation before the system's deployment and also within a live environment.
Andreas Both 0001, Paul Heinze, Aleksandr Perevalov, Johannes Richard Bartsch, Rostislav Iudin, Johannes Rudolf Herkner, Tim Schrader, Jonas Wunsch, René Gürth, Ann Kristin Falkenhain
WSDM1
2022 Can Machine Translation be a Reasonable Alternative for Multilingual Question Answering Systems over Knowledge Graphs?
abstract
Providing access to information is the main and most important purpose of the Web. However, despite available easy-to-use tools (e.g., search engines, chatbots, question answering) the accessibility is typically limited by the capability of using the English language. This excludes a huge amount of people. In this work, we discuss Knowledge Graph Question Answering (KGQA) systems that aim at providing natural language access to data stored in Knowledge Graphs (KG). While several KGQA systems have been proposed, only very few have dealt with a language other than English. In this work, we follow our research agenda of enabling speakers of any language to access the knowledge stored in KGs. Because of the lack of native support for many languages, we use machine translation (MT) tools to evaluate KGQA systems regarding questions in languages that are unsupported by a KGQA system. In total, our evaluation is based on 8 different languages (including some that never were evaluated before). For the intensive evaluation, we extend the QALD-9 dataset for KGQA with Wikidata queries and high-quality translations. The extension was done in a crowdsourcing manner by native speakers of the different languages. By using multiple KGQA systems for the evaluation, we were enabled to investigate and answer the main research question: “Can MT be an alternative for multilingual KGQA systems?”. The evaluation results demonstrated that the monolingual KGQA systems can be effectively ported to the new languages with MT tools.
Aleksandr Perevalov, Andreas Both 0001, Dennis Diefenbach, Axel-Cyrille Ngonga Ngomo
WWW2
2021 Better Call the Plumber: Orchestrating Dynamic Information Extraction Pipelines
Mohamad Yaser Jaradeh, Kuldeep Singh 0001, Markus Stocker, Andreas Both 0001, Sören Auer
ICWE4
2021 Improving Answer Type Classification Quality Through Combined Question Answering Datasets
Aleksandr Perevalov, Andreas Both 0001
KSEM2
2020 QAnswer KG: Designing a Portable Question Answering System over RDF Data
abstract
While RDF was designed to make data easily readable by machines, it does not make data easily usable by end-users. Question Answering (QA) over Knowledge Graphs (KGs) is seen as the technology which is able to bridge this gap. It aims to build systems which are capable of extracting the answer to a user’s natural language question from an RDF dataset. In recent years, many approaches were proposed which tackle the problem of QA over KGs. Despite such efforts, it is hard and cumbersome to create a Question Answering system on top of a new RDF dataset. The main open challenge remains portability, i.e., the possibility to apply a QA algorithm easily on new and previously untested RDF datasets. In this publication, we address the problem of portability by presenting an architecture for a portable QA system. We present a novel approach called QAnswer KG, which allows the construction of on-demand QA systems over new RDF datasets. Hence, our approach addresses non-expert users in QA domain. In this paper, we provide the details of QA system generation process. We show that it is possible to build a QA system over any RDF dataset while requiring minimal investments in terms of training. We run experiments using 3 different datasets. To the best of our knowledge, we are the first to design a process for non-expert users. We enable such users to efficiently create an on-demand, scalable, multilingual, QA system on top of any RDF dataset.
Dennis Diefenbach, José M. Giménez-García, Andreas Both 0001, Pierre Maret
ESWC3
2018 Frankenstein: A Platform Enabling Reuse of Question Answering Components
Kuldeep Singh 0001, Andreas Both 0001, Arun Sethupat Radhakrishna, Saeedeh Shekarpour
ESWC2
2018 Why Reinvent the Wheel: Let's Build Question Answering Systems Together
abstract
Modern question answering (QA) systems need to flexibly integrate a number of components specialised to fulfil specific tasks in a QA pipeline. Key QA tasks include Named Entity Recognition and Disambiguation, Relation Extraction, and Query Building. Since a number of different software components exist that implement different strategies for each of these tasks, it is a major challenge to select and combine the most suitable components into a QA system, given the characteristics of a question. We study this optimisation problem and train classifiers, which take features of a question as input and have the goal of optimising the selection of QA components based on those features. We then devise a greedy algorithm to identify the pipelines that include the suitable components and can effectively answer the given question. We implement this model within Frankenstein, a QA framework able to select QA components and compose QA pipelines. We evaluate the effectiveness of the pipelines generated by Frankenstein using the QALD and LC-QuAD benchmarks. These results not only suggest that Frankenstein precisely solves the QA optimisation problem but also enables the automatic composition of optimised QA pipelines, which outperform the static Baseline QA pipeline. Thanks to this flexible and fully automated pipeline generation process, new QA components can be easily included in Frankenstein, thus improving the performance of the generated pipelines.
Kuldeep Singh 0001, Arun Sethupat Radhakrishna, Andreas Both 0001, Saeedeh Shekarpour, Ioanna Lytra, Ricardo Usbeck, Akhilesh Vyas, Akmal Khikmatullaev, Dharmen Punjani, Christoph Lange 0002, Maria-Esther Vidal, Jens Lehmann 0001, Sören Auer
WWW3
2017 Rapid Engineering of QA Systems Using the Light-Weight Qanary Architecture
Andreas Both 0001, Kuldeep Singh 0001, Dennis Diefenbach, Ioanna Lytra
ICWE1
2017 The Qanary Ecosystem: Getting New Insights by Composing Question Answering Pipelines
Dennis Diefenbach, Kuldeep Singh 0001, Andreas Both 0001, Didier Cherix, Christoph Lange 0002, Sören Auer
ICWE3
2016 Qanary - A Methodology for Vocabulary-Driven Open Question Answering Systems
Andreas Both 0001, Dennis Diefenbach, Kuldeep Singh 0001, Saeedeh Shekarpour, Didier Cherix, Christoph Lange 0002
ESWC1
2016 Detecting Similar Linked Datasets Using Topic Modelling
Michael Röder, Axel-Cyrille Ngonga Ngomo, Ivan Ermilov, Andreas Both 0001
ESWC4
2015 Exploring the Space of Topic Coherence Measures
abstract
Quantifying the coherence of a set of statements is a long standing problem with many potential applications that has attracted researchers from different sciences. The special case of measuring coherence of topics has been recently studied to remedy the problem that topic models give no guaranty on the interpretablity of their output. Several benchmark datasets were produced that record human judgements of the interpretability of topics. We are the first to propose a framework that allows to construct existing word based coherence measures as well as new ones by combining elementary components. We conduct a systematic search of the space of coherence measures using all publicly available topic relevance data for the evaluation. Our results show that new combinations of components outperform existing measures with respect to correlation to human ratings. nFinally, we outline how our results can be transferred to further applications in the context of text mining, information retrieval and the world wide web.
Michael Röder, Andreas Both 0001, Alexander Hinneburg
WSDM2
2015 GERBIL: General Entity Annotator Benchmarking Framework
abstract
We present GERBIL, an evaluation framework for semantic entity annotation. The rationale behind our framework is to provide developers, end users and researchers with easy-to-use interfaces that allow for the agile, fine-grained and uniform evaluation of annotation tools on multiple datasets. By these means, we aim to ensure that both tool developers and end users can derive meaningful insights pertaining to the extension, integration and use of annotation applications. In particular, GERBIL provides comparable results to tool developers so as to allow them to easily discover the strengths and weaknesses of their implementations with respect to the state of the art. With the permanent experiment URIs provided by our framework, we ensure the reproducibility and archiving of evaluation results. Moreover, the framework generates data in machine-processable format, allowing for the efficient querying and post-processing of evaluation results. Finally, the tool diagnostics provided by GERBIL allows deriving insights pertaining to the areas in which tools should be further refined, thus allowing developers to create an informed agenda for extensions and end users to detect the right tools for their purposes. GERBIL aims to become a focal point for the state of the art, driving the research agenda of the community by presenting comparable objective evaluation results.
Ricardo Usbeck, Michael Röder, Axel-Cyrille Ngonga Ngomo, Ciro Baron, Andreas Both 0001, Martin Brümmer, Diego Ceccarelli, Marco Cornolti, Didier Cherix, Bernd Eickmann, Paolo Ferragina, Christiane Lemke, Andrea Moro 0001, Roberto Navigli, Francesco Piccinno, Giuseppe Rizzo 0002, Harald Sack, René Speck, Raphaël Troncy, Jörg Waitelonis, Lars Wesemann
WWW5
2014 Ensuring Web Interface Quality through Usability-Based Split Testing
Maximilian Speicher, Andreas Both 0001, Martin Gaedke
ICWE2
2014 WaPPU: Usability-Based A/B Testing
Maximilian Speicher, Andreas Both 0001, Martin Gaedke
ICWE2
2014 StreamMyRelevance!
Maximilian Speicher, Sebastian Nuck, Andreas Both 0001, Martin Gaedke
ICWE3
2014 Web-Scale Extension of RDF Knowledge Bases from Templated Websites
Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Muhammad Saleem 0002, Andreas Both 0001, Valter Crescenzi, Paolo Merialdo, Disheng Qiu
ISWC (1)5
2014 AGDISTIS - Graph-Based Disambiguation of Named Entities Using Linked Data
Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Michael Röder, Daniel Gerber, Sandro A. Coelho, Sören Auer, Andreas Both 0001
ISWC (1)7
2013 TellMyRelevance!: predicting the relevance of web search results from cursor interactions
abstract
It is crucial for the success of a search-driven web application to answer users' queries in the best possible way. A common approach is to use click models for guessing the relevance of search results. However, these models are imprecise and waive valuable information one can gain from non-click user interactions. We introduce TellMyRelevance!---a novel automatic end-to-end pipeline for tracking cursor interactions at the client, analyzing these and learning according relevance models. Yet, the models depend on the layout of the search results page involved, which makes them difficult to evaluate and compare. Thus, we use a Random Mouse Cursor as an extension to our pipeline for generating layout-dependent baselines. Based on these, we can perform evaluations of real-world relevance models. A large-scale interaction log analysis showed that we can learn relevance models whose predictions compare favorably to predictions of an existing state-of-the-art click model.
Maximilian Speicher, Andreas Both 0001, Martin Gaedke
CIKM2
2012 Fast sampling word correlations of high dimensional text data (abstract only)
abstract
Finding correlated words in large document collections is an important ingredient for text analytics. The naïve approach computes the correlations of each word against all other words and filters for highly correlated word pairs. Clearly, this quadratic method cannot be applied to real world scenarios with millions of documents and words. Our main contribution is to transform the task of finding highly correlated word pairs into a word clustering problem that is efficiently solved by locality sensitive hashing (LSH). A key insight of our new method is to note that the empirical Pearson correlation between two words is the cosine of the angle between the centered versions of their word vectors. The angle can be approximated by an LSH scheme. Although centered word vectors are not sparse, the computation of the LSH hash functions can exploit the inherent sparsity of the word data. This leads to an efficient way to detect collisions between centered word vectors having a small angle and therefore provides a fast algorithm to sample highly correlated word pairs. Our new method based on LSH improves run time complexity of the enhanced naïve algorithm. This algorithm reduces the dimensionality of the word vectors using random projection and approximates correlations by computing cosine similarity on the reduced and centered word vectors. However, this method still has quadratic run time. Our new method replaces the filtering for high correlations in the naïve algorithm with finding hash collisions, which can be done by sorting the hash values of the word vectors. We evaluate the scalability of our new algorithm to large text collections.
Frank Rosner, Alexander Hinneburg, Martin Gleditzsch, Matthias Priebe, Andreas Both 0001
SIGMOD Conference5