VLDB 2026 Research / reviewers in the wild / expert
Dennis Diefenbach
dblp:177/7226
· DBLP profile ↗
14ranked-venue papers
7as first author
5since 2021 · last 2024
0000-0002-0046-2219ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Generate and Update Large HDT RDF Knowledge Graphs on Commodity Hardware
Antoine Willerval, Dennis Diefenbach, Angela Bonifati |
ESWC (2) | 2 |
| 2023 | Fine-tuning Strategies for Domain Specific Question Answering under Low Annotation Budget ConstraintsabstractThe progress introduced by pre-trained language models and their fine-tuning has resulted in significant improvements in most downstream NLP tasks. The unsupervised training of a language model combined with further target task finetuning has become the standard QA fine-tuning procedure. In this work, we demonstrate that this strategy is sub-optimal for fine-tuning QA models, especially under a low QA annotation budget, which is a usual setting in practice due to the extractive QA labeling cost. We draw our conclusions by conducting an exhaustive analysis of the performance of the alternatives of the sequential fine-tuning strategy on different QA datasets. Based on the experiments performed, we observed that the best strategy to fine-tune the QA model in low-budget settings is taking a pre-trained language model (PLM) and then fine-tuning PLM with a dataset composed of the target dataset and SQuAD dataset. With zero extra annotation effort, the best strategy outperforms the standard strategy by 2.28% to 6.48%. Our experiments provide one of the first investigations on how to best fine-tune a QA system under a low budget and are therefore of the utmost practical interest to the QA practitioners. Kunpeng Guo, Dennis Diefenbach, Antoine Gourru, Christophe Gravier |
ICTAI | 2 |
| 2023 | Wikidata as a seed for Web ExtractionabstractWikidata has grown to a knowledge graph with an impressive size. To date, it contains more than 17 billion triples collecting information about people, places, films, stars, publications, proteins, and many more. On the other side, most of the information on the Web is not published in highly structured data repositories like Wikidata, but rather as unstructured and semi-structured content, more concretely in HTML pages containing text and tables. Finding, monitoring, and organizing this data in a knowledge graph is requiring considerable work from human editors. The volume and complexity of the data make this task difficult and time-consuming. In this work, we present a framework that is able to identify and extract new facts that are published under multiple Web domains so that they can be proposed for validation by Wikidata editors. The framework is relying on question-answering technologies. We take inspiration from ideas that are used to extract facts from textual collections and adapt them to extract facts from Web pages. For achieving this, we demonstrate that language models can be adapted to extract facts not only from textual collections but also from Web pages. By exploiting the information already contained in Wikidata the proposed framework can be trained without the need for any additional learning signals and can extract new facts for a wide range of properties and domains. Following this path, Wikidata can be used as a seed to extract facts on the Web. Our experiments show that we can achieve a mean performance of 84.07 at F1-score. Moreover, our estimations show that we can potentially extract millions of facts that can be proposed for human validation. The goal is to help editors in their daily tasks and contribute to the completion of the Wikidata knowledge graph. Kunpeng Guo, Dennis Diefenbach, Antoine Gourru, Christophe Gravier |
WWW | 2 |
| 2022 | Can Machine Translation be a Reasonable Alternative for Multilingual Question Answering Systems over Knowledge Graphs?abstractProviding access to information is the main and most important purpose of the Web. However, despite available easy-to-use tools (e.g., search engines, chatbots, question answering) the accessibility is typically limited by the capability of using the English language. This excludes a huge amount of people. In this work, we discuss Knowledge Graph Question Answering (KGQA) systems that aim at providing natural language access to data stored in Knowledge Graphs (KG). While several KGQA systems have been proposed, only very few have dealt with a language other than English. In this work, we follow our research agenda of enabling speakers of any language to access the knowledge stored in KGs. Because of the lack of native support for many languages, we use machine translation (MT) tools to evaluate KGQA systems regarding questions in languages that are unsupported by a KGQA system. In total, our evaluation is based on 8 different languages (including some that never were evaluated before). For the intensive evaluation, we extend the QALD-9 dataset for KGQA with Wikidata queries and high-quality translations. The extension was done in a crowdsourcing manner by native speakers of the different languages. By using multiple KGQA systems for the evaluation, we were enabled to investigate and answer the main research question: “Can MT be an alternative for multilingual KGQA systems?”. The evaluation results demonstrated that the monolingual KGQA systems can be effectively ported to the new languages with MT tools. Aleksandr Perevalov, Andreas Both 0001, Dennis Diefenbach, Axel-Cyrille Ngonga Ngomo |
WWW | 3 |
| 2021 | Wikibase as an Infrastructure for Knowledge Graphs: The EU Knowledge Graph
Dennis Diefenbach, Max De Wilde, Samantha Alipio |
ISWC | 1 |
| 2020 | IntKB: A Verifiable Interactive Framework for Knowledge Base CompletionabstractKnowledge bases (KBs) are essential for many downstream NLP tasks, yet their prime shortcoming is that they are often incomplete.State-of-the-art frameworks for KB completion often lack sufficient accuracy to work fully automated without human supervision.As a remedy, we propose IntKB: a novel interactive framework for KB completion from text based on a question answering pipeline.Our framework is tailored to the specific needs of a human-in-the-loop paradigm: (i) We generate facts that are aligned with text snippets and are thus immediately verifiable by humans.(ii) Our system is designed such that it continuously learns during the KB completion task and, therefore, significantly improves its performance upon initial zero-and few-shot relations over time.(iii) We only trigger human interactions when there is enough information for a correct prediction.Therefore, we train our system with negative examples and a fold-option if there is no answer.Our framework yields a favorable performance: it achieves a hit@1 ratio of 29.7% for initially unseen relations, upon which it gradually improves to 46.2%. Bernhard Kratzwald, Kunpeng Guo, Stefan Feuerriegel, Dennis Diefenbach |
COLING | 4 |
| 2020 | QAnswer KG: Designing a Portable Question Answering System over RDF DataabstractWhile RDF was designed to make data easily readable by machines, it does not make data easily usable by end-users. Question Answering (QA) over Knowledge Graphs (KGs) is seen as the technology which is able to bridge this gap. It aims to build systems which are capable of extracting the answer to a user’s natural language question from an RDF dataset. In recent years, many approaches were proposed which tackle the problem of QA over KGs. Despite such efforts, it is hard and cumbersome to create a Question Answering system on top of a new RDF dataset. The main open challenge remains portability, i.e., the possibility to apply a QA algorithm easily on new and previously untested RDF datasets. In this publication, we address the problem of portability by presenting an architecture for a portable QA system. We present a novel approach called QAnswer KG, which allows the construction of on-demand QA systems over new RDF datasets. Hence, our approach addresses non-expert users in QA domain. In this paper, we provide the details of QA system generation process. We show that it is possible to build a QA system over any RDF dataset while requiring minimal investments in terms of training. We run experiments using 3 different datasets. To the best of our knowledge, we are the first to design a process for non-expert users. We enable such users to efficiently create an on-demand, scalable, multilingual, QA system on top of any RDF dataset. Dennis Diefenbach, José M. Giménez-García, Andreas Both 0001, Pierre Maret |
ESWC | 1 |
| 2020 | HDTCat: Let's Make HDT Generation Scale
Dennis Diefenbach, José M. Giménez-García |
ISWC (2) | 1 |
| 2019 | QAnswer: A Question Answering prototype bridging the gap between a considerable part of the LOD cloud and end-usersabstractWe present QAnswer, a Question Answering system which queries at the same time 3 core datasets of the Semantic Web, that are relevant for end-users. These datasets are Wikidata with Lexemes, LinkedGeodata and Musicbrainz. Additionally, it is possible to query these datasets in English, German, French, Italian, Spanish, Pourtuguese, Arabic and Chinese. Moreover, QAnswer includes a fallback option to the search engine Qwant when the answer to a question cannot be found in the datasets mentioned above. These features make QAnswer as the first prototype of a Question Answering System over a considerable part of the LOD cloud. Dennis Diefenbach, Pedro Henrique Migliatti, Omar Qawasmeh, Vincent Lully, Pierre Maret |
WWW | 1 |
| 2018 | PageRank and Generic Entity Summarization for RDF Knowledge Bases
Dennis Diefenbach, Andreas Thalhammer 0001 |
ESWC | 1 |
| 2018 | Core techniques of question answering systems over knowledge bases: a survey
Dennis Diefenbach, Vanessa López, Kamal Deep Singh, Pierre Maret |
Knowl. Inf. Syst. | 1 |
| 2017 | Rapid Engineering of QA Systems Using the Light-Weight Qanary Architecture
Andreas Both 0001, Kuldeep Singh 0001, Dennis Diefenbach, Ioanna Lytra |
ICWE | 3 |
| 2017 | The Qanary Ecosystem: Getting New Insights by Composing Question Answering Pipelines
Dennis Diefenbach, Kuldeep Singh 0001, Andreas Both 0001, Didier Cherix, Christoph Lange 0002, Sören Auer |
ICWE | 1 |
| 2016 | Qanary - A Methodology for Vocabulary-Driven Open Question Answering Systems
Andreas Both 0001, Dennis Diefenbach, Kuldeep Singh 0001, Saeedeh Shekarpour, Didier Cherix, Christoph Lange 0002 |
ESWC | 2 |