Axel-Cyrille Ngonga Ngomo

dblp:65/4336 · also Axel Ngonga · DBLP profile ↗
← Back
139ranked-venue papers in the field
12as first author
57since 2021 · last 2026
0000-0001-7112-3516ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 97 (10 first)Information Retrieval & Web Search · 25 (1 first)Data Mining & Knowledge Discovery · 6Other / Interdisciplinary · 5Database Systems & Data Management · 4Big Data, Cloud & Distributed Data Systems · 2 (1 first)
YearPublicationVenuePosition
2026 Document-Level Relation Extraction Using Reinforcement Learning with Knowledge Graph Feedback
Manzoor Ali, Hamada M. Zahera, Muhammad Saleem 0002, Yasir Mahmood 0002, Hashim Khan, René Speck, Axel-Cyrille Ngonga Ngomo
ESWC (1)7
2026 Conel: Contrastive Neural Link Discovery Leveraging Literal Similarities
Alexander Becker 0001, Axel-Cyrille Ngonga Ngomo, Mohamed Ahmed Sherif
ESWC (1)2
2026 TIM: Tiered Iterative Knowledge Graph Matching
Alexander Becker 0001, Axel-Cyrille Ngonga Ngomo, Mohamed Ahmed Sherif
ESWC (1)2
2026 No Need to Be a Know-It-All: Fact Checking with Shallow Knowledge
Umair Qudus, Neha Pokharel, Michael Röder, Axel-Cyrille Ngonga Ngomo
ESWC (1)4
2026 Semantics-Aware Caching for Concept Learning
Louis Mozart Kamdem Teyou, Caglar Demir, Axel-Cyrille Ngonga Ngomo
ESWC (1)3
2026 Evaluating Noisy Optimization in Finetuning LMs for Neural Ranking
Daniel Vollmers, Arnab Sharma, Axel-Cyrille Ngonga Ngomo
NLDB3
2026 A Simplex Approach to Synthetic Knowledge Graph Generation
abstract
The growing scale of knowledge graphs demands scalable systems for their subsequent processing. However, accurate benchmarking requires large knowledge graphs. While data-driven synthetic generators based on versioned datasets are promising to generate large realistic graphs, current approaches generate the graph at a triple level without considering higher-order structures. This work introduces SimplexKG, a simplex-based synthetic knowledge graph generator. Our approach analyzes d-dimensional simplices within input knowledge graphs and leverages the identified simplicial networks to generate a synthetic graph of arbitrary size. We explore whether leveraging higher-dimensional structures enhances the realism of synthetic graphs by evaluating the structure and the utility of the generated graphs. Our approach consistently outperforms 2 baseline generators and 6 variants of the state-of-the-art generator LEMMING in structural fidelity and triple store benchmarking scenarios across 3 datasets. Specifically, compared to the second-best approach, our graphs achieve a structuredness value up to 26.62 % closer to the target graph, while reducing the query throughput error by up to 6.59 % across storage solutions.
Ana Alexandra Morim da Silva, Atul Bhopalsing Pundir, Michael Röder, Axel-Cyrille Ngonga Ngomo
WWW4
2026 ELEVATE-ID: Extending Large Language Models for End-to-End Entity Linking Evaluation in Indonesian
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their effectiveness in low-resource languages remains underexplored, particularly in complex tasks such as end-to-end Entity Linking (EL), which requires both mention detection and disambiguation against a knowledge base (KB). In earlier work, we introduced IndEL — the first end-to-end EL benchmark dataset for the Indonesian language — covering both a general domain (news) and a specific domain (religious text from the Indonesian translation of the Quran), and evaluated four traditional end-to-end EL systems on this dataset. In this study, we propose ELEVATE-ID, a comprehensive evaluation framework for assessing LLM performance on end-to-end EL in Indonesian. The framework evaluates LLMs under both zero-shot and fine-tuned conditions, using multilingual and Indonesian monolingual models, with Wikidata as the target KB. Our experiments include performance benchmarking, generalization analysis across domains, and systematic error analysis. Results show that GPT-4 and GPT-3.5 achieve the highest accuracy in zero-shot and fine-tuned settings, respectively. However, even fine-tuned GPT-3.5 underperforms compared to DBpedia Spotlight — the weakest of the traditional model baselines — in the general domain. Interestingly, GPT-3.5 outperforms Babelfy in the specific domain. Generalization analysis indicates that fine-tuned GPT-3.5 adapts more effectively to cross-domain and mixed-domain scenarios. Error analysis uncovers persistent challenges that hinder LLM performance: difficulties with non-complete mentions, acronym disambiguation, and full-name recognition in formal contexts. These issues point to limitations in mention boundary detection and contextual grounding. Indonesian-pretrained LLMs, Komodo and Merak, reveal core weaknesses: template leakage and entity hallucination, respectively—underscoring architectural and training limitations in low-resource end-to-end EL. 1
Ria Hari Gusmita, Asep Fajar Firmansyah, Hamada M. Zahera, Axel-Cyrille Ngonga Ngomo
Data Knowl. Eng.4
2025 ANTS: Abstractive Entity Summarization in Knowledge Graphs
Asep Fajar Firmansyah, Hamada M. Zahera, Mohamed Ahmed Sherif, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
ESWC (1)5
2025 Robustness Evaluation of Knowledge Graph Embedding Models Under Non-targeted Attacks
Sourabh Kapoor, Arnab Sharma, Michael Röder, Caglar Demir, Axel-Cyrille Ngonga Ngomo
ESWC (1)5
2025 Evaluating Approximate Nearest Neighbour Search Systems on Knowledge Graph Embeddings
Gaurav Pandit, Michael Röder, Axel-Cyrille Ngonga Ngomo
ESWC (1)3
2025 NL2LS: LLM-based Automatic Linking of Knowledge Graphs
abstract
Integrated knowledge graphs form the foundation of numerous data-driven applications, including search engines, conversational agents, and e-commerce solutions. Declarative link discovery frameworks utilize link specifications to define the conditions necessary for establishing a link between knowledge graphs’ resources. Despite domain expertise, defining such link specifications remains challenging due to their intricate syntax, threshold tuning, and the need to precisely express complex linking logic. To address this challenge, we propose NL2LS, a novel language-driven approach that leverages large language models to automatically translate natural language (NL) into link specifications (LSs), enabling domain experts and practitioners to express correct and complex linking rules more effectively. NL2LS employs three distinct training paradigms to handle the complexity of link specifications: zero-shot learning, one-shot learning and supervised fine-tuning. We evaluated NL2LS using different large language model architectures in comparison with a rule-based baseline model on different multi-lingual datasets. Our evaluation using BLEU, METEOR, ChrF++, and TER metrics demonstrates that NL2LS effectively translates natural language into link specifications, lowering the technical barrier and assisting users in specifying link rules more intuitively.
Reda Ihtassine, Asep Fajar Firmansyah, Nikit Srivastava, Manzoor Ali, Axel-Cyrille Ngonga Ngomo, Mohamed Ahmed Sherif
K-CAP5
2025 Parameter Averaging in Link Prediction
Rupesh Sapkota, Caglar Demir, Arnab Sharma, Axel-Cyrille Ngonga Ngomo
K-CAP4
2025 Neural Reasoning for Robust Instance Retrieval in SHOIQ
abstract
Concept learning exploits background knowledge in the form of description logic axioms to learn explainable classification models from knowledge bases. Despite recent breakthroughs in neuro-symbolic concept learning, most approaches still cannot be deployed on real-world knowledge bases. This is due to their use of description logic reasoners, which are not robust against inconsistencies nor erroneous data. We address this challenge by presenting a novel neural reasoner dubbed Ebr. Our reasoner relies on embeddings to approximate the results of a symbolic reasoner. We show that Ebr solely requires retrieving instances for atomic concepts and existential restrictions to retrieve or approximate the set of instances of any concept in the description logic \(\mathcal {SHOIQ}\). In our experiments, we compare Ebr with state-of-the-art reasoners. Our results suggest that Ebr is robust against missing and erroneous data in contrast to existing reasoners.
Louis Mozart Kamdem Teyou, Luke Friedrichs, N'Dah Jean Kouagou, Caglar Demir, Yasir Mahmood 0002, Stefan Heindorf, Axel-Cyrille Ngonga Ngomo
K-CAP7
2025 Evaluation of Entity and Relation Linking for Question Answering over Knowledge Graphs
abstract
Entity and relation linking critically impact the accuracy of knowledge graph question answering (KGQA), often limiting the performance of downstream tasks like query generation. While recent advances in large language models (LLMs) offer promising solutions, their role in linking remains underexplored. This work studies how different linking strategies – including both traditional and LLM-based approaches– affect the quality of generated KG queries. We design multiple linking pipelines and use their output to guide structured query construction. Our study not only evaluates linking accuracy, but also the end-to-end impact on query generation. Our experiments show that LLM-based linkers significantly outperform non-LLM methods, particularly in recall. Moreover, we find that high recall—even at the cost of precision—can lead to better overall performance, as LLMs are resilient to input noise. These findings highlight the importance of recall-oriented linking in modern KGQA pipelines.
Daniel Vollmers, René Speck, Hamada M. Zahera, Axel-Cyrille Ngonga Ngomo
K-CAP4
2025 Explainable Benchmarking through the Lense of Concept Learning
abstract
Evaluating competing systems in a comparable way, i.e., benchmarking them, is an undeniable pillar of the scientific method. However, system performance is often summarized via a small number of metrics. The analysis of the evaluation details and the derivation of insights for further development or use remains a tedious manual task with often biased results. Thus, this paper argues for a new type of benchmarking, which is dubbed explainable benchmarking. The aim of explainable benchmarking approaches is to automatically generate explanations for the performance of systems in a benchmark. We provide a first instantiation of this paradigm for knowledge-graph-based question answering systems. We compute explanations by using a novel concept learning approach developed for large knowledge graphs called PruneCEL. Our evaluation shows that PruneCEL outperforms state-of-the-art concept learners on the task of explainable benchmarking by up to 0.55 points F1 measure. A task-driven user study with 41 participants shows that in 80% of the cases, the majority of participants can accurately predict the behavior of a system based on our explanations. Our code and data are available at https://github.com/dice-group/PruneCEL/tree/K-cap2025.
Quannian Zhang, Michael Röder, Nikit Srivastava, N'Dah Jean Kouagou, Axel-Cyrille Ngonga Ngomo
K-CAP5
2025 Tree-Based OWL Class Expression Learner over Large Graphs
Caglar Demir, Moshood Yekini, Michael Röder, Yasir Mahmood 0002, Axel-Cyrille Ngonga Ngomo
ECML/PKDD (3)5
2025 Glide: Knowledge Graph Linking Using Distance-Aware Embeddings
Alexander Becker 0001, Axel-Cyrille Ngonga Ngomo, Mohamed Ahmed Sherif
ISWC (1)2
2025 Efficient Updates for Worst-Case Optimal Join Triple Stores
Alexander Bigerl, Nikolaos Karalis, Liss Heidrich, Axel-Cyrille Ngonga Ngomo
ISWC (1)4
2025 Link Prediction Under Non-targeted Attacks: Do Soft Labels Always Help?
Adel Memariani, Michael Röder, Arnab Sharma, Caglar Demir, Axel-Cyrille Ngonga Ngomo
ISWC (1)5
2025 Benchmarking Knowledge Editing Using Logical Rules
Tatiana Moteu Ngoli, N'Dah Jean Kouagou, Hamada M. Zahera, Axel-Cyrille Ngonga Ngomo
ISWC (2)4
2024 ExPrompt: Augmenting Prompts Using Examples as Modern Baseline for Stance Classification
Umair Qudus, Michael Röder, Daniel Vollmers, Axel-Cyrille Ngonga Ngomo
CIKM4
2024 Embedding Knowledge Graphs in Function Spaces
abstract
We introduce a novel embedding method diverging from conventional approaches by operating within function spaces of finite dimension rather than finite vector space, thus departing significantly from standard knowledge graph embedding techniques. Initially employing polynomial functions to compute embeddings, we progress to more intricate representations using neural networks with varying layer complexities. We argue that employing functions for embedding computation enhances expressiveness and allows for more degrees of freedom, enabling operations such as composition, derivatives and primitive of entities representation. Additionally, we meticulously outline the step-by-step construction of our approach and provide code for reproducibility, thereby facilitating further exploration and application in the field.
Louis Mozart Kamdem Teyou, Caglar Demir, Axel-Cyrille Ngonga Ngomo
CIKM3
2024 FaVEL: Fact Validation Ensemble Learning
Umair Qudus, Franck Lionel Tatkeu Pekarou, Ana Alexandra Morim da Silva, Michael Röder, Axel-Cyrille Ngonga Ngomo
EKAW5
2024 UniQ-Gen: Unified Query Generation Across Multiple Knowledge Graphs
Daniel Vollmers, Nikit Srivastava, Hamada M. Zahera, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
EKAW5
2024 ESLM: Improving Entity Summarization by Leveraging Language Models
Asep Fajar Firmansyah, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
ESWC (1)3
2024 Efficient Evaluation of Conjunctive Regular Path Queries Using Multi-way Joins
Nikolaos Karalis, Alexander Bigerl, Liss Heidrich, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
ESWC (1)5
2024 Enhancing Relation Extraction Through Augmented Data: Large Language Models Unleashed
Manzoor Ali, Muhammad Sohail Nisar, Muhammad Saleem 0002, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
NLDB (2)5
2024 IndEL: Indonesian Entity Linking Benchmark Dataset for General and Specific Domains
Ria Hari Gusmita, Muhammad Faruq Amiral Abshar, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
NLDB (1)4
2024 Evaluating Negation with Multi-way Joins Accelerates Class Expression Learning
Nikolaos Karalis, Alexander Bigerl, Caglar Demir, Liss Heidrich, Axel-Cyrille Ngonga Ngomo
ECML/PKDD (6)5
2024 Blink: Blank Node Matching Using Embeddings
Alexander Becker 0001, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
ISWC (1)3
2023 RELD: A Knowledge Graph of Relation Extraction Datasets
Manzoor Ali, Muhammad Saleem 0002, Diego Moussallem, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
ESWC5
2023 Neural Class Expression Synthesis
N'Dah Jean Kouagou, Stefan Heindorf, Caglar Demir, Axel-Cyrille Ngonga Ngomo
ESWC4
2023 LauNuts: A Knowledge Graph to Identify and Compare Geographic Regions in the European Union
abstract
The Nomenclature of Territorial Units for Statistics (NUTS) is a classification that represents countries in the European Union (EU). It is published at intervals of several years and organized in a hierarchical system where geographical areas are subdivided according to their population sizes. In addition to NUTS, there is a further subdivided hierarchy level, named Local Administrative Units (LAU), whose data are updated annually by EU member states. While both datasets are published by Eurostat as Excel files, an additional RDF dataset is available for NUTS up to the 2016 scheme. With this work, we provide the Linked Data community with an up-to-date Knowledge Graph in which NUTS and LAU data are linked and which contains population numbers as well as area sizes. We also publish an Open Source generator software for future released versions that will naturally arise due to changes in population numbers. These contributions can be used to enrich other datasets and allow comparisons among regions in the European Union. All resources are available at https://w3id.org/launuts .
Adrian Wilke, Axel-Cyrille Ngonga Ngomo
ESWC2
2023 Learning Permutation-Invariant Embeddings for Description Logic Concepts
Caglar Demir, Axel-Cyrille Ngonga Ngomo
IDA2
2023 Lingua Franca - Entity-Aware Machine Translation Approach for Question Answering over Knowledge Graphs
abstract
This research paper proposes an approach called Lingua Franca that improves machine translation quality by utilizing information from a knowledge graph to translate named entities accurately. The accurate entity translation is crucial when applied to entity-oriented search including Knowledge Graph Question Answering systems. In a nutshell, the approach preserves recognized named entities with an entity-replacement technique during the translation process. It replaces the entities back with their labels found in a knowledge graph for the target language to ensure that questions are translated correctly before answering them using a Knowledge Graph Question Answering system. The paper also introduces an open-source modular framework that enables researchers to design their own named entity-aware machine translation pipelines. The presented experimental results demonstrate the effectiveness of the Lingua Franca approach in comparison to baseline Machine Translation models. The approach shows a statistically significant improvement in the quality provided by several Knowledge Graph Question Answering systems using Lingua Franca on different datasets.
Nikit Srivastava, Aleksandr Perevalov, Denis Kuchelev, Diego Moussallem, Axel-Cyrille Ngonga Ngomo, Andreas Both 0001
K-CAP5
2023 Explainable Integration of Knowledge Graphs Using Large Language Models
Abdullah Fathi Ahmed, Asep Fajar Firmansyah, Mohamed Ahmed Sherif, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
NLDB5
2023 IndQNER: Named Entity Recognition Benchmark Dataset from the Indonesian Translation of the Quran
Ria Hari Gusmita, Asep Fajar Firmansyah, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
NLDB4
2023 Clifford Embeddings - A Generalized Approach for Embedding in Normed Algebras
Caglar Demir, Axel-Cyrille Ngonga Ngomo
ECML/PKDD (3)2
2023 LitCQD: Multi-hop Reasoning in Incomplete Knowledge Graphs with Numeric Literals
Caglar Demir, Michel Wiebesiek, Renzhong Lu, Axel-Cyrille Ngonga Ngomo, Stefan Heindorf
ECML/PKDD (3)4
2023 Neural Class Expression Synthesis in ALCHIQ(D)
N'Dah Jean Kouagou, Stefan Heindorf, Caglar Demir, Axel-Cyrille Ngonga Ngomo
ECML/PKDD (4)4
2023 TemporalFC: A Temporal Fact Checking Approach over Knowledge Graphs
Umair Qudus, Michael Röder, Sabrina Kirrane, Axel-Cyrille Ngonga Ngomo
ISWC4
2022 Learning Concept Lengths Accelerates Concept Learning in ALC
N'Dah Jean Kouagou, Stefan Heindorf, Caglar Demir, Axel-Cyrille Ngonga Ngomo
ESWC4
2022 REBench: Microbenchmarking Framework for Relation Extraction Systems
Manzoor Ali, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo
ISWC3
2022 Hashing the Hypertrie: Space- and Time-Efficient Indexing for SPARQL in Tensors
abstract
Abstract Time-efficient solutions for querying RDF knowledge graphs depend on indexing structures with low response times to answer SPARQL queries rapidly. Hypertries—an indexing structure we recently developed for tensor-based triple stores—have achieved significant runtime improvements over several mainstream storage solutions for RDF knowledge graphs. However, the space footprint of this novel data structure is still often larger than that of many mainstream solutions. In this work, we detail means to reduce the memory footprint of hypertries and thereby further speed up query processing in hypertrie-based RDF storage solutions. Our approach relies on three strategies: (1) the elimination of duplicate nodes via hashing, (2) the compression of non-branching paths, and (3) the storage of single-entry leaf nodes in their parent nodes. We evaluate these strategies by comparing them with baseline hypertries as well as popular triple stores such as Virtuoso, Fuseki, GraphDB, Blazegraph and gStore. We rely on four datasets/benchmark generators in our evaluation: SWDF, DBpedia, WatDiv, and WikiData. Our results suggest that our modifications significantly reduce the memory footprint of hypertries by up to 70% while leading to a relative improvement of up to 39% with respect to average Queries per Second and up to 740% with respect to Query Mixes per Hour.
Alexander Bigerl, Lixi Conrads, Charlotte Behning, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo
ISWC5
2022 HybridFC: A Hybrid Fact-Checking Approach for Knowledge Graphs
Umair Qudus, Michael Röder, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo
ISWC4
2022 MultPAX: Keyphrase Extraction Using Language Models and Knowledge Graphs
Hamada M. Zahera, Daniel Vollmers, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
ISWC4
2022 EvoLearner: Learning Description Logics with Evolutionary Algorithms
abstract
Classifying nodes in knowledge graphs is an important task, e.g., for predicting missing types of entities, predicting which molecules cause cancer, or predicting which drugs are promising treatment candidates. While black-box models often achieve high predictive performance, they are only post-hoc and locally explainable and do not allow the learned model to be easily enriched with domain knowledge. Towards this end, learning description logic concepts from positive and negative examples has been proposed. However, learning such concepts often takes a long time and state-of-the-art approaches provide limited support for literal data values, although they are crucial for many applications. In this paper, we propose EvoLearner—an evolutionary approach to learn concepts in , which is the attributive language with complement () paired with qualified cardinality restrictions () and data properties (). We contribute a novel initialization method for the initial population: starting from positive examples, we perform biased random walks and translate them to description logic concepts. Moreover, we improve support for data properties by maximizing information gain when deciding where to split the data. We show that our approach significantly outperforms the state of the art on the benchmarking framework SML-Bench for structured machine learning. Our ablation study confirms that this is due to our novel initialization method and support for data properties.
Stefan Heindorf, Lukas Blübaum, Nick Düsterhus, Till Werner, Varun Nandkumar Golani, Caglar Demir, Axel-Cyrille Ngonga Ngomo
WWW7
2022 Can Machine Translation be a Reasonable Alternative for Multilingual Question Answering Systems over Knowledge Graphs?
abstract
Providing access to information is the main and most important purpose of the Web. However, despite available easy-to-use tools (e.g., search engines, chatbots, question answering) the accessibility is typically limited by the capability of using the English language. This excludes a huge amount of people. In this work, we discuss Knowledge Graph Question Answering (KGQA) systems that aim at providing natural language access to data stored in Knowledge Graphs (KG). While several KGQA systems have been proposed, only very few have dealt with a language other than English. In this work, we follow our research agenda of enabling speakers of any language to access the knowledge stored in KGs. Because of the lack of native support for many languages, we use machine translation (MT) tools to evaluate KGQA systems regarding questions in languages that are unsupported by a KGQA system. In total, our evaluation is based on 8 different languages (including some that never were evaluated before). For the intensive evaluation, we extend the QALD-9 dataset for KGQA with Wikidata queries and high-quality translations. The extension was done in a crowdsourcing manner by native speakers of the different languages. By using multiple KGQA systems for the evaluation, we were enabled to investigate and answer the main research question: “Can MT be an alternative for multilingual KGQA systems?”. The evaluation results demonstrated that the monolingual KGQA systems can be effectively ported to the new languages with MT tools.
Aleksandr Perevalov, Andreas Both 0001, Dennis Diefenbach, Axel-Cyrille Ngonga Ngomo
WWW4
2022 A survey of RDF stores & SPARQL engines for querying knowledge graphs
Muhammad Saleem 0002, Bin Yao 0002, Aidan Hogan, Axel-Cyrille Ngonga Ngomo
VLDB J.5
2021 Convolutional Complex Knowledge Graph Embeddings
Caglar Demir, Axel-Cyrille Ngonga Ngomo
ESWC2
2021 Applying Grammar-Based Compression to RDF
Michael Röder, Philip Frerk, Lixi Conrads, Axel-Cyrille Ngonga Ngomo
ESWC4
2021 Efficient RDF Knowledge Graph Partitioning Using Querying Workload
abstract
Data partitioning is an effective way to manage large datasets. While a broad range of RDF graph partitioning techniques has been proposed in previous works, little attention has been given to workload-aware RDF graph partitioning. In this paper, we propose two techniques that make use of the querying workload to detect the portions of RDF graphs that are often queried concurrently. Our techniques leverage predicate co-occurrences in SPARQL queries. By detecting highly co-occurring predicates, our techniques can keep data pertaining to these predicates in the same data partition. We evaluate the proposed partitioning techniques using various real-data and query benchmarks generated by the FEASIBLE SPARQL benchmark generation framework. Our evaluation results show the superiority of the proposed techniques in comparison to previous techniques in terms of better query runtime performances.
Adnan Akhter, Muhammad Saleem 0002, Alexander Bigerl, Axel-Cyrille Ngonga Ngomo
K-CAP4
2021 GATES: Using Graph Attention Networks for Entity Summarization
abstract
The sheer size of modern knowledge graphs has led to increased attention being paid to the entity summarization task. Given a knowledge graph T and an entity e found therein, solutions to entity summarization select a subset of the triples from T which summarize e's concise bound description. Presently, the best performing approaches rely on sequence-to-sequence models to generate entity summaries and use little to none of the structure information of T during the summarization process. We hypothesize that this structure information can be exploited to compute better summaries. To verify our hypothesis, we propose GATES, a new entity summarization approach that combines topological information and knowledge graph embeddings to encode triples. The topological information is encoded by means of a Graph Attention Network. Furthermore, ensemble learning is applied to boost the performance of triple scoring. We evaluate GATES on the DBpedia and LMDB datasets from ESBM (version 1.2), as well as on the FACES datasets. Our results show that GATES outperforms the state-of-the-art approaches on 4 of 6 configuration settings and reaches up to 0.574 F-measure. Pertaining to resulted summaries quality, GATES still underperforms the state of the arts as it obtains the highest score only on 1 of 6 configuration settings at 0.697 NDCG score. An open-source implementation of our approach and of the code necessary to rerun our experiments are available at https://github.com/dice-group/GATES.
Asep Fajar Firmansyah, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
K-CAP3
2021 ASSET: A Semi-supervised Approach for Entity Typing in Knowledge Graphs
abstract
Entity typing in knowledge graphs (KGs) aims to infer missing types of entities and might be considered one of the most significant tasks of knowledge graph construction since type information is highly relevant for querying, quality assurance, and KG applications. While supervised learning approaches for entity typing have been proposed, they require large amounts of (manually) labeled data, which can be expensive to obtain. In this paper, we propose a novel approach for KG entity typing that leverages semi-supervised learning from massive unlabeled data. Our approach follows a teacher-student paradigm that allows combining a small amount of labeled data with a large amount of unlabeled data to boost performance. We conduct several experiments on two benchmarking datasets (FB15k-ET and YAGO43k-ET). Our results demonstrate the effectiveness of our approach in improving entity typing in KGs. Given type information for only 1% of entities, our approach ASSET predicts missing types with a F1-score of 0.47 and 0.64 on the datasets FB15k-ET and YAGO43k-ET, respectively, outperforming supervised baselines.
Hamada M. Zahera, Stefan Heindorf, Axel-Cyrille Ngonga Ngomo
K-CAP3
2021 Using Compositional Embeddings for Fact Checking
Ana Alexandra Morim da Silva, Michael Röder, Axel-Cyrille Ngonga Ngomo
ISWC3
2021 Multilingual Verbalization and Summarization for Explainable Link Discovery
Abdullah Fathi Ahmed, Mohamed Ahmed Sherif, Diego Moussallem, Axel-Cyrille Ngonga Ngomo
Data Knowl. Eng.4
2020 CauseNet: Towards a Causality Graph Extracted from the Web
abstract
Causal knowledge is seen as one of the key ingredients to advance artificial intelligence. Yet, few knowledge bases comprise causal knowledge to date, possibly due to significant efforts required for validation. Notwithstanding this challenge, we compile CauseNet, a large-scale knowledge base of claimed causal relations between causal concepts. By extraction from different semi- and unstructured web sources, we collect more than 11 million causal relations with an estimated extraction precision of 83% and construct the first large-scale and open-domain causality graph. We analyze the graph to gain insights about causal beliefs expressed on the web and we demonstrate its benefits in basic causal question answering. Future work may use the graph for causal reasoning, computational argumentation, multi-hop question answering, and more.
Stefan Heindorf, Yan Scholten, Henning Wachsmuth, Axel-Cyrille Ngonga Ngomo, Martin Potthast
CIKM4
2020 Tentris - A Tensor-Based Triple Store
Alexander Bigerl, Lixi Conrads, Charlotte Behning, Mohamed Ahmed Sherif, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo
ISWC (1)6
2020 NABU - Multilingual Graph-Based Neural RDF Verbalizer
Diego Moussallem, Dwaraknath Gnaneshwar, Thiago Castro Ferreira, Axel-Cyrille Ngonga Ngomo
ISWC (1)4
2020 Squirrel - Crawling RDF Knowledge Graphs on the Web
Michael Röder, Geraldo de Souza, Axel-Cyrille Ngonga Ngomo
ISWC (2)3
2020 Revealing Secrets in SPARQL Session Level
Meng Wang 0009, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Guilin Qi, Haofen Wang
ISWC (1)4
2019 Big POI data integration with Linked Data technologies
Spiros Athanasiou, Giorgos Giannopoulos, Damien Graux, Nikos Karagiannakis, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo, Kostas Patroumpas, Mohamed Ahmed Sherif, Dimitrios Skoutas 0001
EDBT6
2019 Dragon: Decision Tree Learning for Link Discovery
Daniel Obraczka, Axel-Cyrille Ngonga Ngomo
ICWE2
2019 Do your Resources Sound Similar?: On the Impact of Using Phonetic Similarity in Link Discovery
abstract
An increasing number of heterogeneous datasets abiding by the Linked Data paradigm is published everyday. Discovering links between these datasets is thus central to achieving the vision behind the Data Web. Declarative Link Discovery (LD) frameworks rely on complex Link Specification (LS) to express the conditions under which two resources should be linked. Complex LS combine similarity measures with thresholds to determine whether a given predicate holds between two resources. State of the art LD frameworks rely mostly on string-based similarity measures such as Levenshtein and Jaccard. However, string-based similarity measures often fail to catch the similarity of resources with phonetically similar property values when these property values are represented using different string representation (e.g., names and street labels). In this paper, we evaluate the impact of using phonetics-based similarities in the process of LD.
Abdullah Fathi Ahmed, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
K-CAP3
2019 Utilizing Knowledge Graphs for Neural Machine Translation Augmentation
abstract
While neural networks have led to substantial progress in machine translation, their success depends heavily on large amounts of training data. However, parallel training corpora are not always readily available. Moreover, out-of-vocabulary words---mostly entities and terminological expressions---pose a difficult challenge to Neural Machine Translation systems. Recent efforts have tried to alleviate the data sparsity problem by augmenting the training data using different strategies, such as external knowledge injection. In this paper, we hypothesize that knowledge graphs enhance the semantic feature extraction of neural models, thus optimizing the translation of entities and terminological expressions in texts and consequently leading to better translation quality. We investigate two different strategies for incorporating knowledge graphs into neural models without modifying the neural network architectures. Additionally, we examine the effectiveness of our augmented models on domain-specific texts and ontologies. Our knowledge-graph-augmented neural translation model, dubbed KG-NMT, achieves significant and consistent improvements of +3 BLEU, METEOR and chrF3 on average on the newstest datasets between 2015 and 2018 for the WMT English-German translation task.
Diego Moussallem, Axel-Cyrille Ngonga Ngomo, Paul Buitelaar, Mihael Arcan
K-CAP2
2019 Congenial Benchmarking of RDF Storage Solutions
abstract
Many SPARQL benchmark generation techniques rely on SPARQL query templates or on selecting representative queries from a set of input queries by inspecting their syntactic features. Hence, prototype queries from such benchmarks mainly capture combinations of SPARQL features, but not the semantics nor the conceptual association between queries. We present congenial benchmarks---a novel type of benchmark that can detect conceptual associations and thus reflect prototypical user intentions when selecting prototype queries. We study SPARROW, an instantiation of congenial benchmarks, where the conceptual associations of SPARQL queries are measured by concept similarity measures. To this end, we transform unary acyclic conjunctive SPARQL queries into ELH-description logic concepts. Our evaluation of three popular triple stores on two datasets shows that the benchmarks generated by SPARROW differ considerably from benchmarks generated using a feature-based approach. Moreover, our evaluation suggests that SPARROW can characterize the performance of common triple stores with respect to user needs by exploiting conceptual associations to detect prototypical user needs.
Axel-Cyrille Ngonga Ngomo, Lixi Conrads, Maximilian Pensel, Anni-Yasmin Turhan
K-CAP1
2019 Jointly Learning from Social Media and Environmental Data for Typhoon Intensity Prediction
abstract
Existing technologies employ different machine learning approachesto predict disasters from historical environmental data. However,for short-term disasters (e.g., earthquakes), historical data alonehas a limited prediction capability. In this work, we consider so-cial media as a supplementary source of knowledge in additionto historical environmental data. Further, we build a joint modelthat learns from disaster-related tweets and environmental data toimprove prediction. We propose the combination of semantically-enriched word embedding to represent entities in tweets with theirsemantics representations computed with the traditionalword2vec.Our experiments show that our proposed approach outperformsthe accuracy of state-of-the-art models in disaster prediction
Hamada M. Zahera, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
K-CAP3
2019 LSVS: Link Specification Verbalization and Summarization
Abdullah Fathi Ahmed, Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo
NLDB3
2019 THOTH: Neural Translation and Enrichment of Knowledge Graphs
Diego Moussallem, Tommaso Soru, Axel-Cyrille Ngonga Ngomo
ISWC (1)3
2019 QaldGen: Towards Microbenchmarking of Question Answering Systems over Knowledge Graphs
Kuldeep Singh 0001, Muhammad Saleem 0002, Abhishek Nadgeri, Lixi Conrads, Jeff Z. Pan, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ISWC (2)6
2019 Unsupervised Discovery of Corroborative Paths for Fact Validation
Zafar Habeeb Syed, Michael Röder, Axel-Cyrille Ngonga Ngomo
ISWC (1)3
2019 How Representative Is a SPARQL Benchmark? An Analysis of RDF Triplestore Benchmarks
abstract
Triplestores are data management systems for storing and querying RDF data. Over recent years, various benchmarks have been proposed to assess the performance of triplestores across different performance measures. However, choosing the most suitable benchmark for evaluating triplestores in practical settings is not a trivial task. This is because triplestores experience varying workloads when deployed in real applications. We address the problem of determining an appropriate benchmark for a given real-life workload by providing a fine-grained comparative analysis of existing triplestore benchmarks. In particular, we analyze the data and queries provided with the existing triplestore benchmarks in addition to several real-world datasets. Furthermore, we measure the correlation between the query execution time and various SPARQL query features and rank those features based on their significance levels. Our experiments reveal several interesting insights about the design of such benchmarks. With this fine-grained evaluation, we aim to support the design and implementation of more diverse benchmarks. Application developers can use our result to analyze their data and queries and choose a data management system.
Muhammad Saleem 0002, Gábor Szárnyas, Lixi Conrads, Syed Ahmad Chan Bukhari, Qaiser Mehmood 0001, Axel-Cyrille Ngonga Ngomo
WWW6
2019 TISCO: Temporal scoping of facts
Anisa Rula, Matteo Palmonari, Simone Rubinacci, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001, Andrea Maurino, Diego Esteves
J. Web Semant.4
2019 Leopard - A baseline approach to attribute prediction and validation for knowledge graph population
René Speck, Axel-Cyrille Ngonga Ngomo
J. Web Semant.2
2018 FactCheck: Validating RDF Triples Using Textual Evidence
abstract
With the increasing uptake of knowledge graphs comes an increasing need for validating the knowledge contained in these graphs. However, the sheer size and number of knowledge bases used in real-world applications makes manual fact checking impractical. In this paper, we employ sentence coherence features gathered from trustworthy source documents to outperform the state of the art in fact checking. Our approach, FactCheck, uses this information to score how likely a fact is to be true and provides the user the evidence used to validate the input facts. We evaluated our approach on two different benchmark datasets and two different corpora. Our results show that FactCheck outperforms the state of the art by up to 13.3% in F-measure and 19.3% AUC. FactCheck is open-source and is available at https://github.com/dice-group/FactCheck.
Zafar Habeeb Syed, Michael Röder, Axel-Cyrille Ngonga Ngomo
CIKM3
2018 An Empirical Evaluation of RDF Graph Partitioning Techniques
Adnan Akhter, Axel-Cyrille Ngonga Ngomo, Muhammad Saleem 0002
EKAW2
2018 On Extracting Relations Using Distributional Semantics and a Tree Generalization
René Speck, Axel-Cyrille Ngonga Ngomo
EKAW2
2018 Dynamic Planning for Link Discovery
Kleanthi Georgala, Daniel Obraczka, Axel-Cyrille Ngonga Ngomo
ESWC3
2018 Where is My URI?
Andre Valdestilhas, Tommaso Soru, Markus Nentwig, Edgard Marx, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo
ESWC6
2018 Efficiently Pinpointing SPARQL Query Containments
Claus Stadler, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ICWE3
2018 SPgen : A Benchmark Generator for Spatial Link Discovery Tools
Tzanina Saveta, Irini Fundulaki, Giorgos Flouris, Axel-Cyrille Ngonga Ngomo
ISWC (1)4
2018 LargeRDFBench: A billion triples benchmark for SPARQL endpoint federation
Muhammad Saleem 0002, Ali Hasnain, Axel-Cyrille Ngonga Ngomo
J. Web Semant.3
2018 Machine Translation using Semantic Web Technologies: A Survey
Diego Moussallem, Matthias Wauer, Axel-Cyrille Ngonga Ngomo
J. Web Semant.3
2017 Implementing scalable structured machine learning for big data in the SAKE project
abstract
Exploration and analysis of large amounts of machine generated data requires innovative approaches. We propose a combination of Semantic Web and Machine Learning to facilitate the analysis. First, data is collected and converted to RDF according to a schema in the Web Ontology Language OWL. Several components can continue working with the data, to interlink, label, augment, or classify. The size of the data poses new challenges to existing solutions, which we solve in this contribution by transitioning from in-memory to database.
Simon Bin, Patrick Westphal, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo
IEEE BigData4
2017 Holistic and scalable ranking of RDF data
abstract
The volume and number of data sources published using Semantic Web standards such as RDF grows continuously. The largest of these data sources now contain billions of facts and are updated periodically. A large number of applications driven by such data sources requires the ranking of entities and facts contained in such knowledge graphs. Hence, there is a need for time-efficient approaches that can compute ranks for entities and facts simultaneously. In this paper, we present the first holistic ranking approach for RDF data. Our approach, dubbed HARE, allows the simultaneous computation of ranks for RDF triples, resources, properties and literals. To this end, HARE relies on the representation of RDF graphs as bi-partite graphs. It then employs a time-efficient extension of the random walk paradigm to bi-partite graphs. We show that by virtue of this extension, the worst-case complexity of HARE is O(n5) while that of PageRank is O(n6). In addition, we evaluate the practical efficiency of our approach by comparing it with PageRank on 6 real and 6 synthetic datasets with sizes up to 108triples. Our results show that HARE is up to 2 orders of magnitude faster than PageRank. We also present a brief evaluation of HARE's ranking accuracy by comparing it with that of PageRank applied directly to RDF graphs. Our evaluation on 19 classes of DBpedia demonstrates that there is no statistical difference between HARE and PageRank. We hence conclude that our approach goes beyond the state of the art by allowing the ranking of all RDF entities and of RDF triples without being worse w.r.t. the ranking quality it achieves on resources. HARE is open-source and is available at http://github.com/dice-group/hare.
Axel-Cyrille Ngonga Ngomo, Michael Hoffmann 0007, Ricardo Usbeck, Kunal Jha
IEEE BigData1
2017 All that Glitters Is Not Gold - Rule-Based Curation of Reference Datasets for Named Entity Recognition and Entity Linking
Kunal Jha, Michael Röder, Axel-Cyrille Ngonga Ngomo
ESWC (1)3
2017 Wombat - A Generalization Approach for Automatic Link Discovery
Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ESWC (1)2
2017 The BigDataEurope Platform - Supporting the Variety Dimension of Big Data
Sören Auer, Simon Scerri, Aad Versteden, Erika Pauwels, Angelos Charalambidis, Stasinos Konstantopoulos, Jens Lehmann 0001, Hajira Jabeen, Ivan Ermilov, Gezim Sejdiu, Andreas Ikonomopoulos, Spyros Andronopoulos, Mandy Vlachogiannis, Charalambos Pappas, Athanasios Davettas, Iraklis A. Klampanos, Efstathios Grigoropoulos, Vangelis Karkaletsis, Victor de Boer, Ronny Siebes, Mohamed Nadjib Mami, Sergio Albani, Michele Lazzarini, Paulo Nunes, Emanuele Angiuli, Nikiforos Pittaras, George Giannakopoulos, Giorgos Argyriou, George Stamoulis 0001, George Papadakis 0001, Manolis Koubarakis, Pythagoras Karampiperis, Axel-Cyrille Ngonga Ngomo, Maria-Esther Vidal
ICWE33
2017 Characterizing mention mismatching problems for improving recognition results
abstract
Mentions to real world things which are recognized by software tools in text often mismatch the ground truth. This paper proposes a formal classification of mention mismatching problems, including partial matching. Then, it depicts evidence that some longer mentions are associated with higher precision and more specific things than shorter mentions that overlap them. Based on this, some algorithms are proposed to automatically improve mentions by increasing their sizes whenever and as much as possible. Experimental results applying a variety of state-of-the-art annotation tools against several datasets made from real world texts show that over-segmentation (returned mention contained in the corresponding one of the ground truth) is the most prevalent partial matching problem among those of the proposed classification. In addition, some of the proposed algorithms for mention enhancing were able to correct most over-segmented mentions returned by tools used in the experiments with prominent benchmarks, leading to gains in precision and recall.
Jean Carlos Oliveira de Abreu, Renato Fileto, Axel-Cyrille Ngonga Ngomo, Michael Röder, Matthias Wittwer, Horacio Saggion
iiWAS3
2017 MAG: A Multilingual, Knowledge-base Agnostic and Deterministic Entity Linking Approach
abstract
Entity linking has recently been the subject of a significant body of research. Currently, the best performing approaches rely on trained mono-lingual models. Porting these approaches to other languages is consequently a difficult endeavor as it requires corresponding training data and retraining of the models. We address this drawback by presenting a novel multilingual, knowledge-base agnostic and deterministic approach to entity linking, dubbed MAG. MAG is based on a combination of context-based retrieval on structured knowledge bases and graph algorithms. We evaluate MAG on 23 data sets and in 7 languages. Our results show that the best approach trained on English datasets (PBOH) achieves a micro F-measure that is up to 4 times worse on datasets in other languages. MAG on the other hand achieves state-of-the-art performance on English datasets and reaches a micro F-measure that is up to 0.6 higher than that of PBOH on non-English languages.
Diego Moussallem, Ricardo Usbeck, Michael Röder, Axel-Cyrille Ngonga Ngomo
K-CAP4
2017 SQCFramework: SPARQL Query Containment Benchmark Generation Framework
abstract
Query containment is a fundamental problem in data management with its main application being in global query optimization. A number of SPARQL query containment solvers for SPARQL have been recently developed. To the best of our knowledge, the Query Containment Benchmark (QC-Bench) is the only benchmark for evaluating these containment solvers. However, this benchmark contains a fixed number of synthetic queries, which were handcrafted by its creators. We propose SQCFramework, a SPARQL query containment benchmark generation framework which is able to generate customized SPARQL containment benchmarks from real SPARQL query logs. The framework is flexible enough to generate benchmarks of varying sizes and according to the user-defined criteria on the most important SPARQL features to be considered for query containment benchmarking. This is achieved using different clustering algorithms. We compare state-of-the-art SPARQL query containment solvers by using different query containment benchmarks generated from DBpedia and Semantic Web Dog Food query logs. In addition, we analyze the quality of the different benchmarks generated by SQCFramework.
Muhammad Saleem 0002, Claus Stadler, Qaiser Mehmood 0001, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo
K-CAP5
2017 Ensemble Learning of Named Entity Recognition Algorithms using Multilayer Perceptron for the Multilingual Web of Data
abstract
Implementing the multilingual Semantic Web vision requires transforming unstructured data in multiple languages from the Document Web into structured data for the multilingual Web of Data. We present the multilingual version of FOX, a knowledge extraction suite which supports this migration by providing named entity recognition based on ensemble learning for five languages. Our evaluation results show that our approach goes beyond the performance of existing named entity recognition systems on all five languages. In our best run, we outperform the state of the art by a gain of 32.38% F1-Score points on a Dutch dataset. More information and a demo can be found at http://fox.aksw.org as well as an extended version of the paper descriping the evaluation in detail.
René Speck, Axel-Cyrille Ngonga Ngomo
K-CAP2
2017 Iguana: A Generic Framework for Benchmarking the Read-Write Performance of Triple Stores
Lixi Conrads, Jens Lehmann 0001, Muhammad Saleem 0002, Mohamed Morsey, Axel-Cyrille Ngonga Ngomo
ISWC (2)5
2017 Distributed Semantic Analytics Using the SANSA Stack
Jens Lehmann 0001, Gezim Sejdiu, Lorenz Bühmann, Patrick Westphal, Claus Stadler, Ivan Ermilov, Simon Bin, Nilesh Chakraborty, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Hajira Jabeen
ISWC (2)10
2017 SIGIR 2017 Workshop on Open Knowledge Base and Question Answering (OKBQA2017)
abstract
Over the past years, several challenges and calls for research projects have pointed out the dire need for pushing natural language interfaces. In this context, the importance of Semantic Web data as a premier knowledge source is rapidly increasing. But we are still far from having accurate natural language interfaces that allow handling complex information needs in a user-centric and highly performant manner. The development of such interfaces requires collaboration of a range of different fields, including natural language processing, information extraction, knowledge base construction and population, reasoning, and question answering. With the goal to join forces in the collaborative development of natural language QA systems, the second OKBQA workshop is organized within the 40th SIGIR conference.
Key-Sun Choi, Teruko Mitamura, Piek Vossen, Jin-Dong Kim, Axel-Cyrille Ngonga Ngomo
SIGIR5
2017 GENESIS: a generic RDF data access interface
abstract
The availability of billions of facts represented in RDF on the Web provides novel opportunities for data discovery and access. In particular, keyword search and question answering approaches enable even lay people to access this data. However, the interpretation of the results of these systems, as well as the navigation through these results, remains challenging. In this paper, we present Genesis, a generic RDF data access interface. Genesis can be deployed on top of any knowledge base and search engine with minimal effort and allows for the representation of RDF data in a layperson-friendly way. This is facilitated by the modular architecture for reusable components underlying our framework. Currently, these include a generic search back-end, together with corresponding interactive user interface components based on a service for similar and related entities as well as verbalization services to bridge between RDF and natural language.
Timofey Ermilov, Diego Moussallem, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo
WI4
2017 LOG4MEX: a library to export machine learning experiments
abstract
A choice of the best computational solution for a particular task is increasingly reliant on experimentation. Even though experiments are often described through text, tables, and figures, their descriptions are often incomplete or confusing. Thus, researchers often have to perform lengthy web searches for reproducing and understanding the results. In order to minimize this gap, vocabularies and ontologies have been proposed for representing data mining and machine learning (ML) experiments. However, we still lack proper tools to export properly these metadata. To this end, we present an open-source library dubbed LOG4MEX which aims at supporting the scientific community to fulfill this gap.
Diego Esteves, Diego Moussallem, Tommaso Soru, Ciro Baron, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo, Julio C. Duarte
WI6
2017 An evaluation of models for runtime approximation in link discovery
abstract
Time-efficient link discovery is of central importance to implement the vision of the Semantic Web. Some of the most rapid Link Discovery approaches rely internally on planning to execute link specifications. In newer works, linear models have been used to estimate the runtime of the fastest planners. However, no other category of models has been studied for this purpose so far. In this paper, we study non-linear runtime estimation functions for runtime estimation. In particular, we study exponential and mixed models for the estimation of the runtimes of planners. To this end, we evaluate three different models for runtime on six datasets using 500 link specifications. We show that exponential and mixed models achieve better fits when trained but are only to be preferred in some cases. Our evaluation also shows that the use of better runtime approximation models has a positive impact on the overall execution of link specifications.
Kleanthi Georgala, Michael Hoffmann 0007, Axel-Cyrille Ngonga Ngomo
WI3
2017 CEDAL: time-efficient detection of erroneous links in large-scale link repositories
abstract
More than 500 million facts on the Linked Data Web are statements across knowledge bases. These links are of crucial importance for the Linked Data Web as they make a large number of tasks possible, including cross-ontology, question answering and federated queries. However, a large number of these links are erroneous and can thus lead to these applications producing absurd results. We present a time-efficient and complete approach for the detection of erroneous links for properties that are transitive. To this end, we make use of the semantics of URIs on the Data Web and combine it with an efficient graph partitioning algorithm. We then apply our algorithm to the LinkLion repository and show that we can analyze 19,200,114 links in 4.6 minutes. Our results show that at least 13% of the owl :sameAs links we considered are erroneous. In addition, our analysis of the provenance of links allows discovering agents and knowledge bases that commonly display poor linking. Our algorithm can be easily executed in parallel and on a GPU. We show that these implementations are up to two orders of magnitude faster than classical reasoners and a non-parallel implementation.
Andre Valdestilhas, Tommaso Soru, Axel-Cyrille Ngonga Ngomo
WI3
2016 TAIPAN: Automatic Property Mapping for Tabular Data
Ivan Ermilov, Axel-Cyrille Ngonga Ngomo
EKAW2
2016 The Lazy Traveling Salesman - Memory Management for Large-Scale Link Discovery
Axel-Cyrille Ngonga Ngomo, Mofeed Mohamed Hassan
ESWC1
2016 Detecting Similar Linked Datasets Using Topic Modelling
Michael Röder, Axel-Cyrille Ngonga Ngomo, Ivan Ermilov, Andreas Both 0001
ESWC2
2015 Automating RDF Dataset Transformation and Enrichment
Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ESWC2
2015 HAWK - Hybrid Question Answering Using Linked Data
Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Lorenz Bühmann, Christina Unger
ESWC2
2015 Using Caching for Local Link Discovery on Large Data Sets
Mofeed Mohamed Hassan, René Speck, Axel-Cyrille Ngonga Ngomo
ICWE3
2015 ASSESS - Automatic Self-Assessment Using Linked Data
Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo
ISWC (2)3
2015 LSQ: The Linked SPARQL Queries Dataset
Muhammad Saleem 0002, Muhammad Intizar Ali, Aidan Hogan, Qaiser Mehmood 0001, Axel-Cyrille Ngonga Ngomo
ISWC (2)5
2015 FEASIBLE: A Feature-Based SPARQL Benchmark Generation Framework
Muhammad Saleem 0002, Qaiser Mehmood 0001, Axel-Cyrille Ngonga Ngomo
ISWC (1)3
2015 LANCE: Piercing to the Heart of Instance Matching Tools
Tzanina Saveta, Evangelia Daskalaki, Giorgos Flouris, Irini Fundulaki, Melanie Herschel, Axel-Cyrille Ngonga Ngomo
ISWC (1)6
2015 ROCKER: A Refinement Operator for Key Discovery
abstract
The Linked Data principles provide a decentral approach for publishing structured data in the RDF format on the Web. In contrast to structured data published in relational databases where a key is often provided explicitly, finding a set of properties that allows identifying a resource uniquely is a non-trivial task. Still, finding keys is of central importance for manifold applications such as resource deduplication, link discovery, logical data compression and data integration. In this paper, we address this research gap by specifying a refinement operator, dubbed ROCKER, which we prove to be finite, proper and non-redundant. We combine the theoretical characteristics of this operator with two monotonicities of keys to obtain a time-efficient approach for detecting keys, i.e., sets of properties that describe resources uniquely. We then utilize a hash index to compute the discriminability score efficiently. Therewith, we ensure that our approach can scale to very large knowledge bases. Results show that ROCKER yields more accurate results, has a comparable runtime, and consumes less memory w.r.t. existing state-of-the-art techniques.
Tommaso Soru, Edgard Marx, Axel-Cyrille Ngonga Ngomo
WWW3
2015 GERBIL: General Entity Annotator Benchmarking Framework
abstract
We present GERBIL, an evaluation framework for semantic entity annotation. The rationale behind our framework is to provide developers, end users and researchers with easy-to-use interfaces that allow for the agile, fine-grained and uniform evaluation of annotation tools on multiple datasets. By these means, we aim to ensure that both tool developers and end users can derive meaningful insights pertaining to the extension, integration and use of annotation applications. In particular, GERBIL provides comparable results to tool developers so as to allow them to easily discover the strengths and weaknesses of their implementations with respect to the state of the art. With the permanent experiment URIs provided by our framework, we ensure the reproducibility and archiving of evaluation results. Moreover, the framework generates data in machine-processable format, allowing for the efficient querying and post-processing of evaluation results. Finally, the tool diagnostics provided by GERBIL allows deriving insights pertaining to the areas in which tools should be further refined, thus allowing developers to create an informed agenda for extensions and end users to detect the right tools for their purposes. GERBIL aims to become a focal point for the state of the art, driving the research agenda of the community by presenting comparable objective evaluation results.
Ricardo Usbeck, Michael Röder, Axel-Cyrille Ngonga Ngomo, Ciro Baron, Andreas Both 0001, Martin Brümmer, Diego Ceccarelli, Marco Cornolti, Didier Cherix, Bernd Eickmann, Paolo Ferragina, Christiane Lemke, Andrea Moro 0001, Roberto Navigli, Francesco Piccinno, Giuseppe Rizzo 0002, Harald Sack, René Speck, Raphaël Troncy, Jörg Waitelonis, Lars Wesemann
WWW3
2015 DeFacto - Temporal and multilingual Deep Fact Validation
Daniel Gerber, Diego Esteves, Jens Lehmann 0001, Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, René Speck
J. Web Semant.6
2015 SINA: Semantic interpretation of user queries for question answering on interlinked data
Saeedeh Shekarpour, Edgard Marx, Axel-Cyrille Ngonga Ngomo, Sören Auer
J. Web Semant.3
2014 conTEXT - Lightweight Text Analytics Using Linked Data
Ali Khalili, Sören Auer, Axel-Cyrille Ngonga Ngomo
ESWC3
2014 Unsupervised Link Discovery through Knowledge Base Repair
Axel-Cyrille Ngonga Ngomo, Mohamed Ahmed Sherif, Klaus Lyko
ESWC1
2014 Hybrid Acquisition of Temporal Scopes for RDF Data
Anisa Rula, Matteo Palmonari, Axel-Cyrille Ngonga Ngomo, Daniel Gerber, Jens Lehmann 0001, Lorenz Bühmann
ESWC3
2014 HiBISCuS: Hypergraph-Based Source Selection for SPARQL Endpoint Federation
Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo
ESWC2
2014 Web-Scale Extension of RDF Knowledge Bases from Templated Websites
Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Muhammad Saleem 0002, Andreas Both 0001, Valter Crescenzi, Paolo Merialdo, Disheng Qiu
ISWC (1)3
2014 HELIOS - Execution Optimization for Link Discovery
Axel-Cyrille Ngonga Ngomo
ISWC (1)1
2014 Ensemble Learning for Named Entity Recognition
René Speck, Axel-Cyrille Ngonga Ngomo
ISWC (1)2
2014 AGDISTIS - Graph-Based Disambiguation of Named Entities Using Linked Data
Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, Michael Röder, Daniel Gerber, Sandro A. Coelho, Sören Auer, Andreas Both 0001
ISWC (1)2
2014 Big linked cancer data: Integrating linked TCGA and PubMed
Muhammad Saleem 0002, Maulik R. Kamdar, Aftab Iqbal, Shanmukha S. Padmanabhuni, Helena F. Deus, Axel-Cyrille Ngonga Ngomo
J. Web Semant.6
2013 When to Reach for the Cloud: Using Parallel Hardware for Link Discovery
Axel-Cyrille Ngonga Ngomo, Lars Kolb, Norman Heino, Michael Hartung, Sören Auer, Erhard Rahm
ESWC1
2013 COALA - Correlation-Aware Active Learning of Link Specifications
Axel-Cyrille Ngonga Ngomo, Klaus Lyko, Victor Christen
ESWC1
2013 Real-Time RDF Extraction from Unstructured Data Streams
Daniel Gerber, Sebastian Hellmann 0001, Lorenz Bühmann, Tommaso Soru, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo
ISWC (1)6
2013 ORCHID - Reduction-Ratio-Optimal Computation of Geo-spatial Distances for Link Discovery
Axel-Cyrille Ngonga Ngomo
ISWC (1)1
2013 DAW: Duplicate-AWare Federated Query Processing over the Web of Data
Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Josiane Xavier Parreira, Helena F. Deus, Manfred Hauswirth
ISWC (1)2
2013 Sorry, i don't speak SPARQL: translating SPARQL queries into natural language
abstract
Over the past years, Semantic Web and Linked Data technologies have reached the backend of a considerable number of applications. Consequently, large amounts of RDF data are constantly being made available across the planet. While experts can easily gather information from this wealth of data by using the W3C standard query language SPARQL, most lay users lack the expertise necessary to proficiently interact with these applications. Consequently, non-expert users usually have to rely on forms, query builders, question answering or keyword search tools to access RDF data. However, these tools have so far been unable to explicate the queries they generate to lay users, making it difficult for these users to i) assess the correctness of the query generated out of their input, and ii) to adapt their queries or iii) to choose in an informed manner between possible interpretations of their input. This paper addresses this drawback by presenting SPARQL2NL, a generic approach that allows verbalizing SPARQL queries, i.e., converting them into natural language. Our framework can be integrated into applications where lay users are required to understand SPARQL or to generate SPARQL queries in a direct (forms, query builders) or an indirect (keyword search, question answering) manner. We evaluate our approach on the DBpedia question set provided by QALD-2 within a survey setting with both SPARQL experts and lay users. The results of the 115 filled surveys show that SPARQL2NL can generate complete and easily understandable natural language descriptions. In addition, our results suggest that even SPARQL experts can process the natural language representation of SPARQL queries computed by our approach more efficiently than the corresponding SPARQL queries. Moreover, non-experts are enabled to reliably understand the content of SPARQL queries.
Axel-Cyrille Ngonga Ngomo, Lorenz Bühmann, Christina Unger, Jens Lehmann 0001, Daniel Gerber
WWW1
2013 Question answering on interlinked data
abstract
The Data Web contains a wealth of knowledge on a large number of domains. Question answering over interlinked data sources is challenging due to two inherent characteristics. First, different datasets employ heterogeneous schemas and each one may only contain a part of the answer for a certain question. Second, constructing a federated formal query across different datasets requires exploiting links between the different datasets on both the schema and instance levels. We present a question answering system, which transforms user supplied queries (i.e. natural language sentences or keywords) into conjunctive SPARQL queries over a set of interlinked data sources. The contribution of this paper is two-fold: Firstly, we introduce a novel approach for determining the most suitable resources for a user-supplied query from different datasets (disambiguation). We employ a hidden Markov model, whose parameters were bootstrapped with different distribution functions. Secondly, we present a novel method for constructing a federated formal queries using the disambiguated resources and leveraging the linking structure of the underlying datasets. This approach essentially relies on a combination of domain and range inference as well as a link traversal method for constructing a connected graph which ultimately renders a corresponding SPARQL query. The results of our evaluation with three life-science datasets and 25 benchmark queries demonstrate the effectiveness of our approach.
Saeedeh Shekarpour, Axel-Cyrille Ngonga Ngomo, Sören Auer
WWW2
2012 Extracting Multilingual Natural-Language Patterns for RDF Predicates
Daniel Gerber, Axel-Cyrille Ngonga Ngomo
EKAW2
2012 EAGLE: Efficient Active Learning of Link Specifications Using Genetic Programming
Axel-Cyrille Ngonga Ngomo, Klaus Lyko
ESWC1
2012 deqa: Deep Web Extraction for Question Answering
Jens Lehmann 0001, Tim Furche, Giovanni Grasso 0001, Axel-Cyrille Ngonga Ngomo, Christian Schallhart, Andrew Jon Sellers, Christina Unger, Lorenz Bühmann, Daniel Gerber, Konrad Höffner, Sören Auer
ISWC (2)4
2012 DeFacto - Deep Fact Validation
Jens Lehmann 0001, Daniel Gerber, Mohamed Morsey, Axel-Cyrille Ngonga Ngomo
ISWC (1)4
2012 Link Discovery with Guaranteed Reduction Ratio in Affine Spaces with Minkowski Measures
Axel-Cyrille Ngonga Ngomo
ISWC (1)1
2012 Template-based question answering over RDF data
abstract
As an increasing amount of RDF data is published as Linked Data, intuitive ways of accessing this data become more and more important. Question answering approaches have been proposed as a good compromise between intuitiveness and expressivity. Most question answering systems translate questions into triples which are matched against the RDF data to retrieve an answer, typically relying on some similarity metric. However, in many cases, triples do not represent a faithful representation of the semantic structure of the natural language question, with the result that more expressive queries can not be answered. To circumvent this problem, we present a novel approach that relies on a parse of the question to produce a SPARQL template that directly mirrors the internal structure of the question. This template is then instantiated using statistical entity identification and predicate detection. We show that this approach is competitive and discuss cases of questions that can be answered with our approach but not with competing approaches.
Christina Unger, Lorenz Bühmann, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo, Daniel Gerber, Philipp Cimiano
WWW4
2011 DBpedia SPARQL Benchmark - Performance Assessment with Real Queries on Real Data
Mohamed Morsey, Jens Lehmann 0001, Sören Auer, Axel-Cyrille Ngonga Ngomo
ISWC (1)4
2011 SCMS - Semantifying Content Management Systems
Axel-Cyrille Ngonga Ngomo, Norman Heino, Klaus Lyko, René Speck, Martin Kaltenböck
ISWC (2)1
2011 Keyword-Driven SPARQL Query Generation Leveraging Background Knowledge
abstract
The search for information on the Web of Data is becoming increasingly difficult due to its dramatic growth. Especially novice users need to acquire both knowledge about the underlying ontology structure and proficiency in formulating formal queries (e. g. SPARQL queries) to retrieve information from Linked Data sources. So as to simplify and automate the querying and retrieval of information from such sources, we present in this paper a novel approach for constructing SPARQL queries based on user-supplied keywords. Our approach utilizes a set of predefined basic graph pattern templates for generating adequate interpretations of user queries. This is achieved by obtaining ranked lists of candidate resource identifiers for the supplied keywords and then injecting these identifiers into suitable positions in the graph pattern templates. The main advantages of our approach are that it is completely agnostic of the underlying knowledge base and ontology schema, that it scales to large knowledge bases and is simple to use. We evaluate17 possible valid graph pattern templates by measuring their precision and recall on 53 queries against DBpedia. Our results show that 8 of these basic graph pattern templates return results with a precision above 70%. Our approach is implemented as a Web search interface and performs sufficiently fast to return instant answers to the user even with large knowledge bases.
Saeedeh Shekarpour, Sören Auer, Axel-Cyrille Ngonga Ngomo, Daniel Gerber, Sebastian Hellmann 0001, Claus Stadler
Web Intelligence3