Patrick Westphal

dblp:143/9337 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-3855-4485ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 DBLP QuAD 2.0: Scholarly Natural Questions from SPARQL
abstract
We present DBLP-QuAD 2.0, designed to evaluate Scholarly Knowledge Graph Question Answering (KGQA) over DBLP. Recent updates in the underlying DBLP KG, including new entities and relationships such as venues, research streams, and citation links, have necessitated a corresponding update to existing KG QA benchmarking resources. While the DBLP-QuAD dataset focused on author and publication-centered queries, DBLP-QuAD 2.0 broadens the coverage to reflect the enriched structure of the updated KG. Specifically, the questions in our dataset are formulated from SPARQL query logs that cover a wide range of entities involving authors, publications, venues, research streams, and citation relationships. DBLP-QuAD 2.0 thus provides a more comprehensive benchmark for evaluating KGQA systems with a baseline.
Tilahun Abedissa Taffa, Patrick Neises, Stefan Ollinger, Patrick Westphal, Marcel R. Ackermann, Debayan Banerjee, Ricardo Usbeck
K-CAP4
2024 Bridging the Gap: Generating a Comprehensive Biomedical Knowledge Graph Question Answering Dataset
abstract
Despite the plethora of resources such as large-scale corpora and manually curated Knowledge Graphs (KGs), the ability to perform reasoning with natural language inputs over biomedical graphs remains challenging due to insufficient training data. We propose a novel method for automatically constructing a Biomedical Knowledge Graph Question Answering (BioKGQA) dataset sourced from PrimeKG, the largest precision medicine-oriented KG. In total, we create 85,368 question-answer pairs along with their respective SPARQL queries. Our approach generates a diverse array of contextually relevant questions covering a wide spectrum of biomedical concepts and levels of complexity. We evaluate our method based on automatic metrics alongside manual annotations. We establish novel standards tailored for KGQA systems to highlight the linguistic correctness and semantical faithfulness of the generated questions based on extracted KG facts. The compiled dataset – PrimeKGQA – serves as a valuable benchmarking resource for advancing knowledge-driven biomedical research and evaluating KGQA systems.
Xi Yan 0001, Patrick Westphal, Jan Seliger, Ricardo Usbeck
ECAI2
2022 Spatial concept learning and inference on geospatial polygon data
Patrick Westphal, Tobias Grubenmann, Diego Collarana, Simon Bin, Lorenz Bühmann, Jens Lehmann 0001
Knowl. Based Syst.1
2021 A Simulated Annealing Meta-heuristic for Concept Learning in Description Logics
Patrick Westphal, Sahar Vahdati, Jens Lehmann 0001
ILP1
2017 Implementing scalable structured machine learning for big data in the SAKE project
abstract
Exploration and analysis of large amounts of machine generated data requires innovative approaches. We propose a combination of Semantic Web and Machine Learning to facilitate the analysis. First, data is collected and converted to RDF according to a schema in the Web Ontology Language OWL. Several components can continue working with the data, to interlink, label, augment, or classify. The size of the data poses new challenges to existing solutions, which we solve in this contribution by transitioning from in-memory to database.
Simon Bin, Patrick Westphal, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo
IEEE BigData2
2017 Distributed Semantic Analytics Using the SANSA Stack
Jens Lehmann 0001, Gezim Sejdiu, Lorenz Bühmann, Patrick Westphal, Claus Stadler, Ivan Ermilov, Simon Bin, Nilesh Chakraborty, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Hajira Jabeen
ISWC (2)4
2016 DL-Learner - A framework for inductive learning on the Semantic Web
Lorenz Bühmann, Jens Lehmann 0001, Patrick Westphal
J. Web Semant.3
2014 Test-driven evaluation of linked data quality
abstract
Linked Open Data (LOD) comprises an unprecedented volume of structured data on the Web. However, these datasets are of varying quality ranging from extensively curated datasets to crowdsourced or extracted data of often relatively low quality. We present a methodology for test-driven quality assessment of Linked Data, which is inspired by test-driven software development. We argue that vocabularies, ontologies and knowledge bases should be accompanied by a number of test cases, which help to ensure a basic level of quality. We present a methodology for assessing the quality of linked data resources, based on a formalization of bad smells and data quality problems. Our formalization employs SPARQL query templates, which are instantiated into concrete quality test case queries. Based on an extensive survey, we compile a comprehensive library of data quality test case patterns. We perform automatic test case instantiation based on schema constraints or semi-automatically enriched schemata and allow the user to generate specific test case instantiations that are applicable to a schema or dataset. We provide an extensive evaluation of five LOD datasets, manual test case instantiation for five schemas and automatic test case instantiations for all available schemata registered with Linked Open Vocabularies (LOV). One of the main advantages of our approach is that domain specific semantics can be encoded in the data quality test cases, thus being able to discover data quality problems beyond conventional quality heuristics.
Dimitris Kontokostas, Patrick Westphal, Sören Auer, Sebastian Hellmann 0001, Jens Lehmann 0001, Roland Cornelissen, Amrapali Zaveri
WWW2