Jens Lehmann 0001

dblp:71/4882 · DBLP profile ↗
← Back
111ranked-venue papers in the field
7as first author
23since 2021 · last 2026
0000-0001-9108-4278ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 79 (7 first)Information Retrieval & Web Search · 18Data Mining & Knowledge Discovery · 6Database Systems & Data Management · 5Other / Interdisciplinary · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 Graph Query Generation with Constraint-Guided Large Language Agents
abstract
Knowledge Graph Question Answering (KGQA) has advanced through structured query generation, yet most efforts target RDF/SPARQL, leaving Cypher and property graphs underexplored, despite increasing demand for unified KGQA in industry settings. We propose UniQGen, a novel constraint-based framework that employs LLM agents to dynamically extract and refine representative graph query clauses into executable, intent-aligned graph queries across query languages. The foundation of our method is a variant of Chase & Backchase, a family of algorithms for query optimization and reformulation. We extend Chase & Backchase with a dynamic reasoning process over query constraints that also interact with LLMs for query quality estimation. With a Cypher-supported Freebase graph deployed on Amazon Neptune, we extensively evaluate our approach on popular KGQA benchmarks (GraphQ, GrailQA, and WebQSP). We demonstrate that UniQGen outperforms state-of-the-art graph query generation techniques in both accuracy and efficiency, with F1 gains of 31.6% on GraphQ and 4.9% on GrailQA. Unlike prior methods, our framework does not require fine-tuning for schema matching, making it more extensible to schema-less graphs and semantics in query workloads, and is more suitable for enterprise-grade KGQA. We release Cypher outputs and a Neptune-ready Freebase snapshot to support reproducible, cross-language KGQA research.
Mengying Wang 0001, Nicolaas Paul Jedema, Rahul Pandey, RaviKiran Krishnan, Jens Lehmann 0001, Yinghui Wu 0001
ICDE5
2025 ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
Riccardo Pozzi, Matteo Palmonari, Andrea Coletta, Luigi Bellomarini, Jens Lehmann 0001, Sahar Vahdati
ISWC (1)5
2024 Large Language Models for Scientific Question Answering: An Extensive Analysis of the SciQA Benchmark
Jens Lehmann 0001, Antonello Meloni, Enrico Motta, Francesco Osborne, Diego Reforgiato Recupero, Angelo A. Salatino, Sahar Vahdati
ESWC (1)1
2023 Retention is All You Need
abstract
Skilled employees are the most important pillars of an organization. Despite this, most organizations face high attrition and turnover rates. While several machine learning models have been developed to analyze attrition and its causal factors, the interpretations of those models remain opaque. In this paper, we propose the HR-DSS approach, which stands for Human Resource (HR) Decision Support System, and uses explainable AI for employee attrition problems. The system is designed to assist HR departments in interpreting the predictions provided by machine learning models. In our experiments, we employ eight machine learning models to provide predictions. We further process the results achieved by the best-performing model by the SHAP explainability process and use the SHAP values to generate natural language explanations which can be valuable for HR. Furthermore, using "What-if-analysis", we aim to observe plausible causes for attrition of an individual employee. The results show that by adjusting the specific dominant features of each individual, employee attrition can turn into employee retention through informative business decisions.
Karishma Mohiuddin, Mirza Ariful Alam, Mirza Mohtashim Alam, Pascal Welke, Michael Martin 0001, Jens Lehmann 0001, Sahar Vahdati
CIKM6
2023 Distinct Geometrical Representations for Temporal and Relational Structures in Knowledge Graphs
Chengjin Xu, Kossi Amouzouvi, Maocai Wang, Jens Lehmann 0001, Sahar Vahdati
ECML/PKDD (3)5
2023 Integrating Knowledge Graph Embeddings and Pre-trained Language Models in Hypercomplex Spaces
Mojtaba Nayyeri, Mst. Mahfuja Akter, Mirza Mohtashim Alam, Md. Rashad Al Hasan Rony, Jens Lehmann 0001, Steffen Staab
ISWC6
2023 Geometric Algebra Based Embeddings for Static and Temporal Knowledge Graph Completion
abstract
Recent years, Knowledge Graph Embeddings (KGEs) have shown promising performance on link prediction tasks by mapping the entities and relations from a Knowledge Graph (KG) into a geometric space and thus have gained increasing attentions. In addition, many recent Knowledge Graphs involve evolving data, e.g., the fact (Obama, PresidentOf, USA) is valid only from 2009 to 2017. This introduces important challenges for knowledge representation learning since such temporal KGs change over time. In this work, we strive to move beyond the complex or hypercomplex space for KGE and propose a novel geometric algebra based embedding approach, GeomE, which uses multivector representations and the geometric product to model entities and relations. GeomE subsumes several state-of-the-art KGE models and is able to model diverse relations patterns. On top of this, we extend GeomE to TGeomE for temporal KGE, which performs 4th-order tensor factorization of a temporal KG and devises a new linear temporal regularization for time representation learning. Moreover, we study the effect of time granularity on the performance of TGeomE models. Experimental results show that our proposed models achieve the state-of-the-art performances on link prediction over four commonly-used static KG datasets and four well-established temporal KG datasets across various metrics.
Chengjin Xu, Mojtaba Nayyeri, Yung-Yu Chen, Jens Lehmann 0001
IEEE Trans. Knowl. Data Eng.4
2022 Contrastive Representation Learning for Conversational Question Answering over Knowledge Graphs
abstract
This paper addresses the task of conversational question answering (ConvQA) over knowledge graphs (KGs). The majority of existing ConvQA methods rely on full supervision signals with a strict assumption of the availability of gold logical forms of queries to extract answers from the KG. However, creating such a gold logical form is not viable for each potential question in a real-world scenario. Hence, in the case of missing gold logical forms, the existing information retrieval-based approaches use weak supervision via heuristics or reinforcement learning, formulating ConvQA as a KG path ranking problem. Despite missing gold logical forms, an abundance of conversational contexts, such as entire dialog history with fluent responses and domain information, can be incorporated to effectively reach the correct KG path. This work proposes a contrastive representation learning-based approach to rank KG paths effectively. Our approach solves two key challenges. Firstly, it allows weak supervision-based learning that omits the necessity of gold annotations. Second, it incorporates the conversational context (entire dialog history and domain information) to jointly learn its homogeneous representation with KG paths to improve contrastive representations for effective path ranking. We evaluate our approach on standard datasets for ConvQA, on which it significantly outperforms existing baselines on all domains and overall. Specifically, in some cases, the Mean Reciprocal Rank (MRR) and [email protected] ranking metrics improve by absolute 10 and 18 points, respectively, compared to the state-of-the-art performance.
Endri Kacupaj, Kuldeep Singh 0001, Maria Maleshkova, Jens Lehmann 0001
CIKM4
2022 Dihedron Algebraic Embeddings for Spatio-Temporal Knowledge Graph Completion
Mojtaba Nayyeri, Sahar Vahdati, Md Tansen Khan, Mirza Mohtashim Alam, Lisa Wenige, Andreas Behrend, Jens Lehmann 0001
ESWC7
2022 Time-aware Entity Alignment using Temporal Relational Attention
abstract
Knowledge graph (KG) alignment is to match entities in different KGs, which is important to knowledge fusion and integration. Temporal KGs (TKGs) extend traditional Knowledge Graphs (KGs) by associating static triples with specific timestamps (e.g., temporal scopes or time points). Moreover, open-world KGs (OKGs) are dynamic with new emerging entities and timestamps. While entity alignment (EA) between KGs has drawn increasing attention from the research community, EA between TKGs and OKGs still remains unexplored. In this work, we propose a novel Temporal Relational Entity Alignment method (TREA) which is able to learn alignment-oriented TKG embeddings and represent new emerging entities. We first map entities, relations and timestamps into an embedding space, and the initial feature of each entity is represented by fusing the embeddings of its connected relations and timestamps as well as its neighboring entities. A graph neural network (GNN) is employed to capture intra-graph information and a temporal relational attention mechanism is utilized to integrate relation and time features of links between nodes. Finally, a margin-based full multi-class log-loss is used for efficient training and a sequential time regularizer is used to model unobserved timestamps. We use three well-established TKG datasets, as references for evaluating temporal and non-temporal EA methods. Experimental results show that our method outperforms the state-of-the-art EA methods.
Chengjin Xu, Fenglong Su, Bo Xiong 0001, Jens Lehmann 0001
WWW4
2021 DistRDF2ML - Scalable Distributed In-Memory Machine Learning Pipelines for RDF Knowledge Graphs
abstract
This paper presents DistRDF2ML, the generic, scalable, and distributed framework for creating in-memory data preprocessing pipelines for Spark-based machine learning on RDF knowledge graphs. This framework introduces software modules that transform large-scale RDF data into ML-ready fixed-length numeric feature vectors. The developed modules are optimized to the multi-modal nature of knowledge graphs. DistRDF2ML provides aligned software design and usage principles as common data science stacks that offer an easy-to-use package for creating machine learning pipelines. The modules used in the pipeline, the hyper-parameters and the results are exported as a semantic structure that can be used to enrich the original knowledge graph. The semantic representation of metadata and machine learning results offers the advantage of increasing the machine learning pipelines' reusability, explainability, and reproducibility. The entire framework of DistRDF2ML is open source, integrated into the holistic SANSA stack, documented in scala-docs, and covered by unit tests. DistRDF2ML demonstrates its scalable design across different processing power configurations and (hyper-)parameter setups within various experiments. The framework brings the three worlds of knowledge graph engineers, distributed computation developers, and data scientists closer together and offers all of them the creation of explainable ML pipelines using a few lines of code.
Carsten Draschner, Claus Stadler, Farshad Bakhshandegan Moghaddam, Jens Lehmann 0001, Hajira Jabeen
CIKM4
2021 Pattern-Aware and Noise-Resilient Embedding Models
Mojtaba Nayyeri, Sahar Vahdati, Emanuel Sallinger, Mirza Mohtashim Alam, Hamed Shariat Yazdi, Jens Lehmann 0001
ECIR (1)6
2021 Grounding Dialogue Systems via Knowledge Graph Aware Decoding with Pre-trained Transformers
Debanjan Chaudhuri, Md. Rashad Al Hasan Rony, Jens Lehmann 0001
ESWC3
2021 ParaQA: A Question Answering Dataset with Paraphrase Responses for Single-Turn Conversation
Endri Kacupaj, Barshana Banerjee, Kuldeep Singh 0001, Jens Lehmann 0001
ESWC4
2021 Context Transformer with Stacked Pointer Networks for Conversational Question Answering over Knowledge Graphs
Joan Plepi, Endri Kacupaj, Kuldeep Singh 0001, Harsh Thakkar, Jens Lehmann 0001
ESWC5
2021 A Scalable Approach for Distributed Reasoning over Large-scale OWL Datasets
abstract
51
Heba Aamer, Said Fathalla, Jens Lehmann 0001, Hajira Jabeen
KEOD3
2021 HORUS-NER: A Multimodal Named Entity Recognition Framework for Noisy Data
Diego Esteves, José Marcelino, Piyush Chawla, Asja Fischer, Jens Lehmann 0001
IDA5
2021 Loss-Aware Pattern Inference: A Correction on the Wrongly Claimed Limitations of Embedding Models
Mojtaba Nayyeri, Chengjin Xu, Yadollah Yaghoobzadeh, Sahar Vahdati, Mirza Mohtashim Alam, Hamed Shariat Yazdi, Jens Lehmann 0001
PAKDD (3)7
2021 VOGUE: Answer Verbalization Through Multi-Task Learning
Endri Kacupaj, Shyamnath Premnadh, Kuldeep Singh 0001, Jens Lehmann 0001, Maria Maleshkova
ECML/PKDD (3)4
2021 Embedding Knowledge Graphs Attentive to Positional and Centrality Qualities
Afshin Sadeghi, Diego Collarana, Damien Graux, Jens Lehmann 0001
ECML/PKDD (2)4
2021 Improving Inductive Link Prediction Using Hyper-relational Facts
Mehdi Ali, Max Berrendorf, Michael Galkin, Veronika Thost, Tengfei Ma 0001, Volker Tresp, Jens Lehmann 0001
ISWC7
2021 GeoWINE: Geolocation based Wiki, Image, News and Event Retrieval
abstract
In the context of social media, geolocation inference on news or events has become a very important task. In this paper, we present the GeoWINE (Geolocation-based Wiki-Image-News-Event retrieval) demonstrator, an effective modular system for multimodal retrieval which expects only a single image as input. The GeoWINE system consists of five modules in order to retrieve related information from various sources. The first module is a state-of-the-art model for geolocation estimation of images. The second module performs a geospatial-based query for entity retrieval using the Wikidata knowledge graph. The third module exploits four different image embedding representations, which are used to retrieve most similar entities compared to the input image. The last two modules perform news and event retrieval from EventRegistry and the Open Event Knowledge Graph (OEKG). GeoWINE provides an intuitive interface for end-users and is insightful for experts for reconfiguration to individual setups. The GeoWINE achieves promising results in entity label prediction for images on Google Landmarks dataset. The demonstrator is publicly available at http://cleopatra.ijs.si/geowine/.
Golsa Tahmasebzadeh, Endri Kacupaj, Eric Müller-Budack, Sherzod Hakimov, Jens Lehmann 0001, Ralph Ewerth
SIGIR5
2021 Towards holistic Entity Linking: Survey and directions
Italo Lopes Oliveira, Renato Fileto, René Speck, Luís Paulo F. Garcia, Diego Moussallem, Jens Lehmann 0001
Inf. Syst.6
2020 MLM: A Benchmark Dataset for Multitask Learning with Multiple Languages and Modalities
abstract
In this paper, we introduce the MLM (Multiple Languages and Modalities) dataset - a new resource to train and evaluate multitask systems on samples in multiple modalities and three languages. The generation process and inclusion of semantic data provide a resource that further tests the ability for multitask systems to learn relationships between entities. The dataset is designed for researchers and developers who build applications that perform multiple tasks on data encountered on the web and in digital archives. A second version of MLM provides a geo-representative subset of the data with weighted samples for countries of the European Union. We demonstrate the value of the resource in developing novel applications in the digital humanities with a motivating use case and specify a benchmark set of tasks to retrieve modalities and locate entities in the dataset. Evaluation of baseline multitask and single task systems on the full and geo-representative versions of MLM demonstrate the challenges of generalising on diverse data. In addition to the digital humanities, we expect the resource to contribute to research in multimodal representation learning, location estimation, and scene understanding.
Jason Armitage, Endri Kacupaj, Golsa Tahmasebzadeh, Maria Maleshkova, Ralph Ewerth, Jens Lehmann 0001
CIKM7
2020 Evaluating the Impact of Knowledge Graph Context on Entity Disambiguation Models
abstract
Pretrained Transformer models have emerged as state-of-the-art approaches that learn contextual information from the text to improve the performance of several NLP tasks. These models, albeit powerful, still require specialized knowledge in specific scenarios. In this paper, we argue that context derived from a knowledge graph (in our case: Wikidata) provides enough signals to inform pretrained transformer models and improve their performance for named entity disambiguation (NED) on Wikidata KG. We further hypothesize that our proposed KG context can be standardized for Wikipedia, and we evaluate the impact of KG context on the state of the art NED model for the Wikipedia knowledge base. Our empirical results validate that the proposed KG context can be generalized (for Wikipedia), and providing KG context in transformer architectures considerably outperforms the existing baselines, including the vanilla transformer models.
Isaiah Onando Mulang', Kuldeep Singh 0001, Chaitali Prabhu, Abhishek Nadgeri, Johannes Hoffart, Jens Lehmann 0001
CIKM6
2020 Unveiling Relations in the Industry 4.0 Standards Landscape Based on Knowledge Graph Embeddings
Ariam Rivas, Irlán Grangel-González, Diego Collarana, Jens Lehmann 0001, Maria-Esther Vidal
DEXA (2)4
2020 Ontology Design for Pharmaceutical Research Outcomes
Zeynep Say, Said Fathalla, Sahar Vahdati, Jens Lehmann 0001, Sören Auer
TPDL4
2020 VQuAnDa: Verbalization QUestion ANswering DAtaset
abstract
Question Answering (QA) systems over Knowledge Graphs (KGs) aim to provide a concise answer to a given natural language question. Despite the significant evolution of QA methods over the past years, there are still some core lines of work, which are lagging behind. This is especially true for methods and datasets that support the verbalization of answers in natural language. Specifically, to the best of our knowledge, none of the existing Question Answering datasets provide any verbalization data for the question-query pairs. Hence, we aim to fill this gap by providing the first QA dataset VQuAnDa that includes the verbalization of each answer. We base VQuAnDa on a commonly used large-scale QA dataset – LC-QuAD, in order to support compatibility and continuity of previous work. We complement the dataset with baseline scores for measuring future training and evaluation work, by using a set of standard sequence to sequence models and sharing the results of the experiments. This resource empowers researchers to train and evaluate a variety of models to generate answer verbalizations.
Endri Kacupaj, Hamid Zafar, Jens Lehmann 0001, Maria Maleshkova
ESWC3
2020 Embedding-Based Recommendations on Scholarly Knowledge Graphs
Mojtaba Nayyeri, Sahar Vahdati, Hamed Shariat Yazdi, Jens Lehmann 0001
ESWC5
2020 A Distributed Approach for Parsing Large-scale OWL Datasets
abstract
227
Heba Aamer, Said Fathalla, Jens Lehmann 0001, Hajira Jabeen
KEOD3
2020 Semantic Representation of Physics Research Data
abstract
Improvements in web technologies and artificial intelligence enable novel, more data-driven research practices for scientists. However, scientific knowledge generated from data-intensive research practices is disseminated with unstructured formats, thus hindering the scholarly communication in various respects. The traditional document-based representation of scholarly information hampers the reusability of research contributions. To address this concern, we developed the Physics Ontology (PhySci) to represent physics-related scholarly data in a machine-interpretable format. PhySci facilitates knowledge exploration, comparison, and organization of such data by representing it as knowledge graphs. It establishes a unique conceptualization to increase the visibility and accessibility to the digital content of physics publications. We present the iterative design principles by outlining a methodology for its development and applying three different evaluation approaches: data-driven and criteria-based evaluation, as well as ontology testing.
Aysegul Say, Said Fathalla, Sahar Vahdati, Jens Lehmann 0001, Sören Auer
KEOD4
2020 PNEL: Pointer Network Based End-To-End Entity Linking over Knowledge Graphs
Debayan Banerjee, Debanjan Chaudhuri, Mohnish Dubey, Jens Lehmann 0001
ISWC (1)4
2020 Fantastic Knowledge Graph Embeddings and How to Find the Right Space for Them
Mojtaba Nayyeri, Chengjin Xu, Sahar Vahdati, Nadezhda Vassilyeva, Emanuel Sallinger, Hamed Shariat Yazdi, Jens Lehmann 0001
ISWC (1)7
2020 CASQAD - A New Dataset for Context-Aware Spatial Question Answering
Jewgeni Rose, Jens Lehmann 0001
ISWC (2)2
2020 Temporal Knowledge Graph Completion Based on Time Series Gaussian Embedding
Chenjin Xu, Mojtaba Nayyeri, Fouad Alkhoury, Hamed Shariat Yazdi, Jens Lehmann 0001
ISWC (1)5
2020 No one is perfect: Analysing the performance of question answering components over the DBpedia knowledge graph
Kuldeep Singh 0001, Ioanna Lytra, Arun Sethupat Radhakrishna, Saeedeh Shekarpour, Maria-Esther Vidal, Jens Lehmann 0001
J. Web Semant.6
2020 Schema-agnostic SPARQL-driven faceted search benchmark generation
Claus Stadler, Simon Bin, Lisa Wenige, Lorenz Bühmann, Jens Lehmann 0001
J. Web Semant.5
2020 IQA: Interactive query construction in semantic question answering systems
Hamid Zafar, Mohnish Dubey, Jens Lehmann 0001, Elena Demidova
J. Web Semant.3
2019 Big POI data integration with Linked Data technologies
Spiros Athanasiou, Giorgos Giannopoulos, Damien Graux, Nikos Karagiannakis, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo, Kostas Patroumpas, Mohamed Ahmed Sherif, Dimitrios Skoutas 0001
EDBT5
2019 Incorporating Joint Embeddings into Goal-Oriented Dialogues with Multi-task Learning
abstract
Attention-based encoder-decoder neural network models have recently shown promising results in goal-oriented dialogue systems. However, these models struggle to reason over and incorporate state-full knowledge while preserving their end-to-end text generation functionality. Since such models can greatly benefit from user intent and knowledge graph integration, in this paper we propose an RNN-based end-to-end encoder-decoder architecture which is trained with joint embeddings of the knowledge graph and the corpus as input. The model provides an additional integration of user intent along with text generation, trained with multi-task learning paradigm along with an additional regularization technique to penalize generating the wrong entity as output. The model further incorporates a Knowledge Graph entity lookup during inference to guarantee the generated output is state-full based on the local knowledge graph provided. We finally evaluated the model using the BLEU score, empirical evaluation depicts that our proposed architecture can aid in the betterment of task-oriented dialogue system’s performance.
Firas Kassawat, Debanjan Chaudhuri, Jens Lehmann 0001
ESWC3
2019 Uniform Access to Multiform Data Lakes using Semantic Technologies
abstract
Increasing data volumes have extensively increased application possibilities. However, accessing this data in an ad hoc manner remains an unsolved problem due to the diversity of data management approaches, formats and storage frameworks, resulting in the need to effectively access and process distributed heterogeneous data at scale. For years, Semantic Web techniques have addressed data integration challenges with practical knowledge representation models and ontology-based mappings. Leveraging these techniques, we provide a solution enabling uniform access to large, heterogeneous data sources, without enforcing centralization; thus realizing the vision of a Semantic Data Lake. In this paper, we define the core concepts underlying this vision and the architectural requirements that systems implementing it need to fulfill. Squerall, an example of such a system, is an extensible framework built on top of state-of-the-art Big Data technologies. We focus on Squerall's distributed query execution techniques and strategies, empirically evaluating its performance throughout its various sub-phases.
Mohamed Nadjib Mami, Damien Graux, Simon Scerri, Hajira Jabeen, Sören Auer, Jens Lehmann 0001
iiWAS6
2019 The KEEN Universe - An Ecosystem for Knowledge Graph Embeddings with a Focus on Reproducibility and Transferability
Mehdi Ali, Hajira Jabeen, Charles Tapley Hoyt, Jens Lehmann 0001
ISWC (2)4
2019 Using a KG-Copy Network for Non-goal Oriented Dialogues
Debanjan Chaudhuri, Md. Rashad Al Hasan Rony, Simon Jordan, Jens Lehmann 0001
ISWC (1)4
2019 LC-QuAD 2.0: A Large Dataset for Complex Question Answering over Wikidata and DBpedia
Mohnish Dubey, Debayan Banerjee, Abdelrahman Abdelkawi, Jens Lehmann 0001
ISWC (2)4
2019 DBpedia FlexiFusion the Best of Wikipedia > Wikidata > Your Data
Johannes Frey, Marvin Hofer, Daniel Obraczka, Jens Lehmann 0001, Sebastian Hellmann 0001
ISWC (2)4
2019 Incorporating Literals into Knowledge Graph Embeddings
Agustinus Kristiadi, Mohammad Asif Khan 0001, Denis Lukovnikov, Jens Lehmann 0001, Asja Fischer
ISWC (1)4
2019 Pretrained Transformers for Simple Question Answering over Knowledge Graphs
Denis Lukovnikov, Asja Fischer, Jens Lehmann 0001
ISWC (1)3
2019 Learning to Rank Query Graphs for Complex Question Answering over Knowledge Graphs
Gaurav Maheshwari 0001, Priyansh Trivedi, Denis Lukovnikov, Nilesh Chakraborty, Asja Fischer, Jens Lehmann 0001
ISWC (1)6
2019 Squerall: Virtual Ontology-Based Access to Heterogeneous and Large Data Sources
Mohamed Nadjib Mami, Damien Graux, Simon Scerri, Hajira Jabeen, Sören Auer, Jens Lehmann 0001
ISWC (2)6
2019 A Scalable Framework for Quality Assessment of RDF Datasets
Gezim Sejdiu, Anisa Rula, Jens Lehmann 0001, Hajira Jabeen
ISWC (2)3
2019 QaldGen: Towards Microbenchmarking of Question Answering Systems over Knowledge Graphs
Kuldeep Singh 0001, Muhammad Saleem 0002, Abhishek Nadgeri, Lixi Conrads, Jeff Z. Pan, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ISWC (2)7
2019 Sparklify: A Scalable Software Component for Efficient Evaluation of SPARQL Queries over Distributed RDF Datasets
Claus Stadler, Gezim Sejdiu, Damien Graux, Jens Lehmann 0001
ISWC (2)4
2019 TISCO: Temporal scoping of facts
Anisa Rula, Matteo Palmonari, Simone Rubinacci, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001, Andrea Maurino, Diego Esteves
J. Web Semant.5
2018 Divided We Stand Out! Forging Cohorts fOr Numeric Outlier Detection in Large Scale Knowledge Graphs (CONOD)
Hajira Jabeen, Rajjat Dadwal, Gezim Sejdiu, Jens Lehmann 0001
EKAW4
2018 KnIGHT: Mapping Privacy Policies to GDPR
Najmeh Mousavi Nejad, Simon Scerri, Jens Lehmann 0001
EKAW3
2018 Formal Query Generation for Question Answering over Knowledge Bases
Hamid Zafar, Giulio Napolitano, Jens Lehmann 0001
ESWC3
2018 Efficiently Pinpointing SPARQL Query Containments
Claus Stadler, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ICWE4
2018 Deep Query Ranking for Question Answering over Knowledge Bases
Hamid Zafar, Giulio Napolitano, Jens Lehmann 0001
ECML/PKDD (3)3
2018 EARL: Joint Entity and Relation Linking for Question Answering over Knowledge Graphs
Mohnish Dubey, Debayan Banerjee, Debanjan Chaudhuri, Jens Lehmann 0001
ISWC (1)4
2018 DistLODStats: Distributed Computation of RDF Dataset Statistics
Gezim Sejdiu, Ivan Ermilov, Jens Lehmann 0001, Mohamed Nadjib Mami
ISWC (2)3
2018 Why Reinvent the Wheel: Let's Build Question Answering Systems Together
abstract
Modern question answering (QA) systems need to flexibly integrate a number of components specialised to fulfil specific tasks in a QA pipeline. Key QA tasks include Named Entity Recognition and Disambiguation, Relation Extraction, and Query Building. Since a number of different software components exist that implement different strategies for each of these tasks, it is a major challenge to select and combine the most suitable components into a QA system, given the characteristics of a question. We study this optimisation problem and train classifiers, which take features of a question as input and have the goal of optimising the selection of QA components based on those features. We then devise a greedy algorithm to identify the pipelines that include the suitable components and can effectively answer the given question. We implement this model within Frankenstein, a QA framework able to select QA components and compose QA pipelines. We evaluate the effectiveness of the pipelines generated by Frankenstein using the QALD and LC-QuAD benchmarks. These results not only suggest that Frankenstein precisely solves the QA optimisation problem but also enables the automatic composition of optimised QA pipelines, which outperform the static Baseline QA pipeline. Thanks to this flexible and fully automated pipeline generation process, new QA components can be easily included in Frankenstein, thus improving the performance of the generated pipelines.
Kuldeep Singh 0001, Arun Sethupat Radhakrishna, Andreas Both 0001, Saeedeh Shekarpour, Ioanna Lytra, Ricardo Usbeck, Akhilesh Vyas, Akmal Khikmatullaev, Dharmen Punjani, Christoph Lange 0002, Maria-Esther Vidal, Jens Lehmann 0001, Sören Auer
WWW12
2017 Implementing scalable structured machine learning for big data in the SAKE project
abstract
Exploration and analysis of large amounts of machine generated data requires innovative approaches. We propose a combination of Semantic Web and Machine Learning to facilitate the analysis. First, data is collected and converted to RDF according to a schema in the Web Ontology Language OWL. Several components can continue working with the data, to interlink, label, augment, or classify. The size of the data poses new challenges to existing solutions, which we solve in this contribution by transitioning from in-memory to database.
Simon Bin, Patrick Westphal, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo
IEEE BigData3
2017 Wombat - A Generalization Approach for Automatic Link Discovery
Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ESWC (1)3
2017 The BigDataEurope Platform - Supporting the Variety Dimension of Big Data
Sören Auer, Simon Scerri, Aad Versteden, Erika Pauwels, Angelos Charalambidis, Stasinos Konstantopoulos, Jens Lehmann 0001, Hajira Jabeen, Ivan Ermilov, Gezim Sejdiu, Andreas Ikonomopoulos, Spyros Andronopoulos, Mandy Vlachogiannis, Charalambos Pappas, Athanasios Davettas, Iraklis A. Klampanos, Efstathios Grigoropoulos, Vangelis Karkaletsis, Victor de Boer, Ronny Siebes, Mohamed Nadjib Mami, Sergio Albani, Michele Lazzarini, Paulo Nunes, Emanuele Angiuli, Nikiforos Pittaras, George Giannakopoulos, Giorgos Argyriou, George Stamoulis 0001, George Papadakis 0001, Manolis Koubarakis, Pythagoras Karampiperis, Axel-Cyrille Ngonga Ngomo, Maria-Esther Vidal
ICWE7
2017 SimDoc: Topic Sequence Alignment based Document Similarity Framework
abstract
Document similarity is the problem of estimating the degree to which a given pair of documents has similar semantic content. An accurate document similarity measure can improve several enterprise relevant tasks such as document clustering, text mining, and question-answering. In this paper, we show that a document's thematic flow, which is often disregarded by bag-of-word techniques, is pivotal in estimating their similarity. To this end, we propose a novel semantic document similarity framework, called SimDoc. We model documents as topic-sequences, where topics represent latent generative clusters of related words. Then, we use a sequence alignment algorithm to estimate their semantic similarity. We further conceptualize a novel mechanism to compute topic-topic similarity to fine tune our system. In our experiments, we show that SimDoc outperforms many contemporary bag-of-words techniques in accurately computing document similarity, and on practical applications such as document clustering.
Gaurav Maheshwari 0001, Priyansh Trivedi, Harshita Sahijwani, Kunal Jha, Sourish Dasgupta, Jens Lehmann 0001
K-CAP6
2017 SQCFramework: SPARQL Query Containment Benchmark Generation Framework
abstract
Query containment is a fundamental problem in data management with its main application being in global query optimization. A number of SPARQL query containment solvers for SPARQL have been recently developed. To the best of our knowledge, the Query Containment Benchmark (QC-Bench) is the only benchmark for evaluating these containment solvers. However, this benchmark contains a fixed number of synthetic queries, which were handcrafted by its creators. We propose SQCFramework, a SPARQL query containment benchmark generation framework which is able to generate customized SPARQL containment benchmarks from real SPARQL query logs. The framework is flexible enough to generate benchmarks of varying sizes and according to the user-defined criteria on the most important SPARQL features to be considered for query containment benchmarking. This is achieved using different clustering algorithms. We compare state-of-the-art SPARQL query containment solvers by using different query containment benchmarks generated from DBpedia and Semantic Web Dog Food query logs. In addition, we analyze the quality of the different benchmarks generated by SQCFramework.
Muhammad Saleem 0002, Claus Stadler, Qaiser Mehmood 0001, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo
K-CAP4
2017 Iguana: A Generic Framework for Benchmarking the Read-Write Performance of Triple Stores
Lixi Conrads, Jens Lehmann 0001, Muhammad Saleem 0002, Mohamed Morsey, Axel-Cyrille Ngonga Ngomo
ISWC (2)2
2017 Distributed Semantic Analytics Using the SANSA Stack
Jens Lehmann 0001, Gezim Sejdiu, Lorenz Bühmann, Patrick Westphal, Claus Stadler, Ivan Ermilov, Simon Bin, Nilesh Chakraborty, Muhammad Saleem 0002, Axel-Cyrille Ngonga Ngomo, Hajira Jabeen
ISWC (2)1
2017 Sustainable Linked Data Generation: The Case of DBpedia
Wouter Maroy, Anastasia Dimou, Dimitris Kontokostas, Ben De Meester, Ruben Verborgh, Jens Lehmann 0001, Erik Mannens, Sebastian Hellmann 0001
ISWC (2)6
2017 LC-QuAD: A Corpus for Complex Question Answering over Knowledge Graphs
abstract
Being able to access knowledge bases in an intuitive way has been an active area of research over the past years. In particular, several question answering (QA) approaches which allow to query RDF datasets in natural language have been developed as they allow end users to access knowledge without needing to learn the schema of a knowledge base and learn a formal query language. To foster this research area, several training datasets have been created, e.g. in the QALD (Question Answering over Linked Data) initiative. However, existing datasets are insufficient in terms of size, variety or complexity to apply and evaluate a range of machine learning based QA approaches for learning complex SPARQL queries. With the provision of the Large-Scale Complex Question Answering Dataset (LC-QuAD), we close this gap by providing a dataset with 5000 questions and their corresponding SPARQL queries over the DBpedia dataset. In this article, we describe the dataset creation process and how we ensure a high variety of questions, which should enable to assess the robustness and accuracy of the next generation of QA systems for knowledge graphs. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Priyansh Trivedi, Gaurav Maheshwari 0001, Mohnish Dubey, Jens Lehmann 0001
ISWC (2)4
2017 LOG4MEX: a library to export machine learning experiments
abstract
A choice of the best computational solution for a particular task is increasingly reliant on experimentation. Even though experiments are often described through text, tables, and figures, their descriptions are often incomplete or confusing. Thus, researchers often have to perform lengthy web searches for reproducing and understanding the results. In order to minimize this gap, vocabularies and ontologies have been proposed for representing data mining and machine learning (ML) experiments. However, we still lack proper tools to export properly these metadata. To this end, we present an open-source library dubbed LOG4MEX which aims at supporting the scientific community to fulfill this gap.
Diego Esteves, Diego Moussallem, Tommaso Soru, Ciro Baron, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo, Julio C. Duarte
WI5
2017 Neural Network-based Question Answering over Knowledge Graphs on Word and Character Level
abstract
Question Answering (QA) systems over Knowledge Graphs (KG) automatically answer natural language questions using facts contained in a knowledge graph. Simple questions, which can be answered by the extraction of a single fact, constitute a large part of questions asked on the web but still pose challenges to QA systems, especially when asked against a large knowledge resource. Existing QA systems usually rely on various components each specialised in solving different sub-tasks of the problem (such as segmentation, entity recognition, disambiguation, and relation classification etc.). In this work, we follow a quite different approach: We train a neural network for answering simple questions in an end-to-end manner, leaving all decisions to the model. It learns to rank subject-predicate pairs to enable the retrieval of relevant facts given a question. The network contains a nested word/character-level question encoder which allows to handle out-of-vocabulary and rare word problems while still being able to exploit word-level semantics. Our approach achieves results competitive with state-of-the-art end-to-end approaches that rely on an attention mechanism.
Denis Lukovnikov, Asja Fischer, Jens Lehmann 0001, Sören Auer
WWW3
2016 Integrating New Refinement Operators in Terminological Decision Trees Learning
Giuseppe Rizzo 0001, Nicola Fanizzi, Jens Lehmann 0001, Lorenz Bühmann
EKAW3
2016 ACRyLIQ: Leveraging DBpedia for Adaptive Crowdsourcing in Linked Data Quality Assessment
Umair ul Hassan, Amrapali Zaveri, Edgard Marx, Edward Curry, Jens Lehmann 0001
EKAW5
2016 AskNow: A Framework for Natural Language Query Formalization in SPARQL
Mohnish Dubey, Sourish Dasgupta, Konrad Höffner, Jens Lehmann 0001
ESWC5
2016 Semantically Enhanced Quality Assurance in the JURION Business Use Case
Dimitris Kontokostas, Christian Mader, Christian Dirschl, Katja Eck, Michael Leuthold, Jens Lehmann 0001, Sebastian Hellmann 0001
ESWC6
2016 LODStats: The Data Web Census Dataset
Ivan Ermilov, Jens Lehmann 0001, Michael Martin 0001, Sören Auer
ISWC (2)2
2016 CubeQA - Question Answering on RDF Data Cubes
Konrad Höffner, Jens Lehmann 0001, Ricardo Usbeck
ISWC (1)2
2016 DL-Learner - A framework for inductive learning on the Semantic Web
Lorenz Bühmann, Jens Lehmann 0001, Patrick Westphal
J. Web Semant.2
2015 Automating RDF Dataset Transformation and Enrichment
Mohamed Ahmed Sherif, Axel-Cyrille Ngonga Ngomo, Jens Lehmann 0001
ESWC3
2015 Assessing and Refining Mappings to RDF to Improve Dataset Quality
Anastasia Dimou, Dimitris Kontokostas, Markus Freudenberg, Ruben Verborgh, Jens Lehmann 0001, Erik Mannens, Sebastian Hellmann 0001, Rik Van de Walle
ISWC (2)5
2015 DBpedia Commons: Structured Multimedia Metadata from the Wikimedia Commons
abstract
The Wikimedia Commons is an online repository of over twenty-five million freely usable audio, video and still image files, including scanned books, historically significant photographs, animal recordings, illustrative figures and maps. Being volunteer-contributed, these media files have different amounts of descriptive metadata with varying degrees of accuracy. The DBpedia Information Extraction Framework is capable of parsing unstructured text into semi-structured data from Wikipedia and transforming it into RDF for general use, but so far it has only been used to extract encyclopedia-like content. In this paper, we describe the creation of the DBpedia Commons ( DBc ) dataset, which was achieved by an extension of the Extraction Framework to support knowledge extraction from Wikimedia Commons as a media repository. To our knowledge, this is the first complete RDFization of the Wikimedia Commons and the largest media metadata RDF database in the LOD cloud.
Gaurav Vaidya, Dimitris Kontokostas, Magnus Knuth, Jens Lehmann 0001, Sebastian Hellmann 0001
ISWC (2)4
2015 DeFacto - Temporal and multilingual Deep Fact Validation
Daniel Gerber, Diego Esteves, Jens Lehmann 0001, Lorenz Bühmann, Ricardo Usbeck, Axel-Cyrille Ngonga Ngomo, René Speck
J. Web Semant.3
2014 Inductive Lexical Learning of Class Expressions
Lorenz Bühmann, Daniel Fleischhacker, Jens Lehmann 0001, André Melo, Johanna Völker
EKAW3
2014 NLP Data Cleansing Based on Linguistic Ontology Constraints
Dimitris Kontokostas, Martin Brümmer, Sebastian Hellmann 0001, Jens Lehmann 0001, Lazaros Ioannidis
ESWC4
2014 Hybrid Acquisition of Temporal Scopes for RDF Data
Anisa Rula, Matteo Palmonari, Axel-Cyrille Ngonga Ngomo, Daniel Gerber, Jens Lehmann 0001, Lorenz Bühmann
ESWC5
2014 Test-driven evaluation of linked data quality
abstract
Linked Open Data (LOD) comprises an unprecedented volume of structured data on the Web. However, these datasets are of varying quality ranging from extensively curated datasets to crowdsourced or extracted data of often relatively low quality. We present a methodology for test-driven quality assessment of Linked Data, which is inspired by test-driven software development. We argue that vocabularies, ontologies and knowledge bases should be accompanied by a number of test cases, which help to ensure a basic level of quality. We present a methodology for assessing the quality of linked data resources, based on a formalization of bad smells and data quality problems. Our formalization employs SPARQL query templates, which are instantiated into concrete quality test case queries. Based on an extensive survey, we compile a comprehensive library of data quality test case patterns. We perform automatic test case instantiation based on schema constraints or semi-automatically enriched schemata and allow the user to generate specific test case instantiations that are applicable to a schema or dataset. We provide an extensive evaluation of five LOD datasets, manual test case instantiation for five schemas and automatic test case instantiations for all available schemata registered with Linked Open Vocabularies (LOV). One of the main advantages of our approach is that domain specific semantics can be encoded in the data quality test cases, thus being able to discover data quality problems beyond conventional quality heuristics.
Dimitris Kontokostas, Patrick Westphal, Sören Auer, Sebastian Hellmann 0001, Jens Lehmann 0001, Roland Cornelissen, Amrapali Zaveri
WWW5
2013 Crowdsourcing Linked Data Quality Assessment
Maribel Acosta, Amrapali Zaveri, Elena Simperl, Dimitris Kontokostas, Sören Auer, Jens Lehmann 0001
ISWC (2)6
2013 Pattern Based Knowledge Base Enrichment
Lorenz Bühmann, Jens Lehmann 0001
ISWC (1)2
2013 Integrating NLP Using Linked Data
Sebastian Hellmann 0001, Jens Lehmann 0001, Sören Auer, Martin Brümmer
ISWC (2)2
2013 Sorry, i don't speak SPARQL: translating SPARQL queries into natural language
abstract
Over the past years, Semantic Web and Linked Data technologies have reached the backend of a considerable number of applications. Consequently, large amounts of RDF data are constantly being made available across the planet. While experts can easily gather information from this wealth of data by using the W3C standard query language SPARQL, most lay users lack the expertise necessary to proficiently interact with these applications. Consequently, non-expert users usually have to rely on forms, query builders, question answering or keyword search tools to access RDF data. However, these tools have so far been unable to explicate the queries they generate to lay users, making it difficult for these users to i) assess the correctness of the query generated out of their input, and ii) to adapt their queries or iii) to choose in an informed manner between possible interpretations of their input. This paper addresses this drawback by presenting SPARQL2NL, a generic approach that allows verbalizing SPARQL queries, i.e., converting them into natural language. Our framework can be integrated into applications where lay users are required to understand SPARQL or to generate SPARQL queries in a direct (forms, query builders) or an indirect (keyword search, question answering) manner. We evaluate our approach on the DBpedia question set provided by QALD-2 within a survey setting with both SPARQL experts and lay users. The results of the 115 filled surveys show that SPARQL2NL can generate complete and easily understandable natural language descriptions. In addition, our results suggest that even SPARQL experts can process the natural language representation of SPARQL queries computed by our approach more efficiently than the corresponding SPARQL queries. Moreover, non-experts are enabled to reliably understand the content of SPARQL queries.
Axel-Cyrille Ngonga Ngomo, Lorenz Bühmann, Christina Unger, Jens Lehmann 0001, Daniel Gerber
WWW4
2012 LODStats - An Extensible Framework for High-Performance Dataset Analytics
Sören Auer, Jan Demter, Michael Martin 0001, Jens Lehmann 0001
EKAW4
2012 Universal OWL Axiom Enrichment for Large Knowledge Bases
Lorenz Bühmann, Jens Lehmann 0001
EKAW2
2012 Linked-Data Aware URI Schemes for Referencing Text Fragments
Sebastian Hellmann 0001, Jens Lehmann 0001, Sören Auer
EKAW2
2012 NIF Combinator: Combining NLP Tool Output
Sebastian Hellmann 0001, Jens Lehmann 0001, Sören Auer, Marcus Nitzschke
EKAW2
2012 Assessing Linked Data Mappings Using Network Measures
Christophe Guéret, Paul Groth, Claus Stadler, Jens Lehmann 0001
ESWC4
2012 Managing the Life-Cycle of Linked Data with the LOD2 Stack
Sören Auer, Lorenz Bühmann, Christian Dirschl, Orri Erling, Michael Hausenblas, Robert Isele, Jens Lehmann 0001, Michael Martin 0001, Pablo N. Mendes, Bert Van Nuffelen, Claus Stadler, Sebastian Tramp, Hugh Williams
ISWC (2)7
2012 deqa: Deep Web Extraction for Question Answering
Jens Lehmann 0001, Tim Furche, Giovanni Grasso 0001, Axel-Cyrille Ngonga Ngomo, Christian Schallhart, Andrew Jon Sellers, Christina Unger, Lorenz Bühmann, Daniel Gerber, Konrad Höffner, Sören Auer
ISWC (2)1
2012 DeFacto - Deep Fact Validation
Jens Lehmann 0001, Daniel Gerber, Mohamed Morsey, Axel-Cyrille Ngonga Ngomo
ISWC (1)1
2012 Template-based question answering over RDF data
abstract
As an increasing amount of RDF data is published as Linked Data, intuitive ways of accessing this data become more and more important. Question answering approaches have been proposed as a good compromise between intuitiveness and expressivity. Most question answering systems translate questions into triples which are matched against the RDF data to retrieve an answer, typically relying on some similarity metric. However, in many cases, triples do not represent a faithful representation of the semantic structure of the natural language question, with the result that more expressive queries can not be answered. To circumvent this problem, we present a novel approach that relies on a parse of the question to produce a SPARQL template that directly mirrors the internal structure of the question. This template is then instantiated using statistical entity identification and predicate detection. We show that this approach is competitive and discuss cases of questions that can be answered with our approach but not with competing approaches.
Christina Unger, Lorenz Bühmann, Jens Lehmann 0001, Axel-Cyrille Ngonga Ngomo, Daniel Gerber, Philipp Cimiano
WWW3
2011 AutoSPARQL: Let Users Query Your Knowledge Base
Jens Lehmann 0001, Lorenz Bühmann
ESWC (1)1
2011 DBpedia SPARQL Benchmark - Performance Assessment with Real Queries on Real Data
Mohamed Morsey, Jens Lehmann 0001, Sören Auer, Axel-Cyrille Ngonga Ngomo
ISWC (1)2
2011 ReDD-Observatory: Using the Web of Data for Evaluating the Research-Disease Disparity
abstract
It is widely accepted that there is a large disparity between the availability of treatment options and the prevalence of diseases all over the world, thus placing individuals in danger. This disparity is partially caused by the restricted access to information that would allow health care and research policy makers to formulate more appropriate measures to mitigate it. Specifically, this shortage of information is caused by the difficulty in reliably obtaining and integrating data regarding the disease burden and the respective research investments. In response to these challenges, the Linked Data paradigm provides a simple mechanism for publishing and interlinking structured information on the Web. In conjunction with the ever increasing data on diseases and health care research available as Linked Data, an opportunity is created to reduce this information gap that would allow for better policy in response to these disparities. In this paper, we present the ReDD-Observatory, an approach for evaluating the Research-Disease Disparity based on the interlinking and integrating of various biomedical data sources. Specifically, we devise a method for representing statistical information as Linked Data and adopt interlinking algorithms for integrating relevant datasets (mainly GHO, Linked CT and PubMed). The assessment of the disparity is then performed with a number of parametrized SPARQL queries on the integrated data substrate. As a consequence, we are for the first time able to provide reliable indicators for the extent of the research-disease disparity in a semi-automated fashion, thus enabling health care professionals and policy makers to make more informed decisions.
Amrapali Zaveri, Ricardo Pietrobon, Sören Auer, Jens Lehmann 0001, Michael Martin 0001, Timofey Ermilov
Web Intelligence4
2011 Class expression learning for ontology engineering
Jens Lehmann 0001, Sören Auer, Lorenz Bühmann, Sebastian Tramp
J. Web Semant.1
2010 I18n of Semantic Web Applications
Sören Auer, Matthias Weidl, Jens Lehmann 0001, Amrapali Zaveri, Key-Sun Choi
ISWC (2)3
2010 ORE - A Tool for Repairing and Enriching Knowledge Bases
Jens Lehmann 0001, Lorenz Bühmann
ISWC (2)1
2009 LinkedGeoData: Adding a Spatial Dimension to the Web of Data
Sören Auer, Jens Lehmann 0001, Sebastian Hellmann 0001
ISWC2
2009 Triplify: light-weight linked data publication from relational databases
abstract
In this paper we present Triplify - a simplistic but effective approach to publish Linked Data from relational databases. Triplify is based on mapping HTTP-URI requests onto relational database queries. Triplify transforms the resulting relations into RDF statements and publishes the data on the Web in various RDF serializations, in particular as Linked Data. The rationale for developing Triplify is that the largest part of information on the Web is already stored in structured form, often as data contained in relational databases, but usually published by Web applications only as HTML mixing structure, layout and content. In order to reveal the pure structured information behind the current Web, we have implemented Triplify as a light-weight software component, which can be easily integrated into and deployed by the numerous, widely installed Web applications. Our approach includes a method for publishing update logs to enable incremental crawling of linked data sources. Triplify is complemented by a library of configurations for common relational schemata and a REST-enabled data source registry. Triplify configurations containing mappings are provided for many popular Web applications, including osCommerce, WordPress, Drupal, Gallery, and phpBB. We will show that despite its light-weight architecture Triplify is usable to publish very large datasets, such as 160GB of geo data from the OpenStreetMap project.
Sören Auer, Sebastian Tramp, Jens Lehmann 0001, Sebastian Hellmann 0001, David Aumüller
WWW3
2009 Learning of OWL Class Descriptions on Very Large Knowledge Bases
abstract
The vision of the Semantic Web is to make use of semantic representations on the largest possible scale - the Web. Large knowledge bases such as DBpedia, OpenCyc, GovTrack, and others are emerging and are freely available as Linked Data and SPARQL endpoints. Exploring and analysing such knowledge bases is a significant hurdle for Semantic Web research and practice. As one possible direction for tackling this problem, the authors present an approach for obtaining complex class descriptions from objects in knowledge bases by using Machine Learning techniques. They describe in detail how we leverage existing techniques to achieve scalability on large knowledge bases available as SPARQL endpoints or Linked Data. Their algorithms are made available in the open source DL-Learner project and we present several real-life scenarios in which they can be used by Semantic Web applications.
Sebastian Hellmann 0001, Jens Lehmann 0001, Sören Auer
Int. J. Semantic Web Inf. Syst.2
2009 DBpedia - A crystallization point for the Web of Data
Christian Bizer, Jens Lehmann 0001, Georgi Kobilarov, Sören Auer, Christian Becker 0001, Richard Cyganiak, Sebastian Hellmann 0001
J. Web Semant.2
2007 What Have Innsbruck and Leipzig in Common? Extracting Semantics from Wiki Content
Sören Auer, Jens Lehmann 0001
ESWC2