Maria-Esther Vidal

dblp:92/654 · DBLP profile ↗
← Back
86ranked-venue papers in the field
4as first author
22since 2021 · last 2026
0000-0003-1160-8727ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 36 (3 first)Database Systems & Data Management · 25 (1 first)Information Retrieval & Web Search · 18Data Mining & Knowledge Discovery · 7
YearPublicationVenuePosition
2026 Tool4Boxology: A Semantic Toolbox for Constructing and Analysing Neuro-Symbolic Architectures
Johannes E. Bendler, Yashrajsinh Chudasama, Mahsa Forghani, Enrique Iglesias, Disha Purohit, Jacquiline Roney, Annette ten Teije, Frank van Harmelen, Maria-Esther Vidal
ESWC (2)9
2026 Towards a Symbolic Representation of the MEMS Development Domain
Florian Diehl, Ivan Marevic, Irlán Grangel-González, Simon Blattner, Maria-Esther Vidal
ESWC (2)5
2026 Joint Graph Learning for Robust Causal Inference over Knowledge Graphs
abstract
Causal inference is critical for understanding cause-effect relationships in real-world domains. However, applying it over knowledge graphs (KGs) poses unique challenges due to two key issues: missing attributes caused by the Open-World Assumption and interference effects arising from complex relational dependencies among entities. Existing methods often assume fully observed data or fail to model inter-unit dependencies, leading to biased or unreliable effect estimates. We introduce BaLu, a joint graph learning framework that addresses both challenges through an end-to-end solution. BaLu reformulates the causal inference over KGs as two interconnected tasks: (1) attribute imputation as edge prediction between units (entities) and their attributes, and (2) treatment effect estimation as node prediction that accounts for interference through representation learning. BaLu employs Graph Neural Networks (GNNs) to capture attribute similarity and relational structure, enabling both accurate imputation and interference-aware message passing. Experiments on four benchmark datasets show that BaLu consistently outperforms state-of-the-art baselines—even when enhanced with strong imputation techniques—demonstrating robust performance in incomplete and relationally complex KGs. These results demonstrate that BaLu offers a principled and practical solution for robust causal inference in knowledge-driven domains, empowering data-driven decision-making under real-world conditions of incompleteness and relational complexity.
Hao Huang 0014, Maria-Esther Vidal
WSDM2
2025 Capturing Symbolic Knowledge of Constraints and Incompleteness to Guide Inductive Learning in Neuro-Symbolic Knowledge Graph Completion
abstract
Knowledge Graphs (KGs) are widely used to represent structured knowledge. However, their incompleteness under the Open World Assumption (OWA) limits their effectiveness for reasoning and inference. Neural link prediction models can recover missing links. Yet, these models often overlook the distinction between semantically valid and invalid inferences and lack mechanisms to validate predictions against domain-specific constraints. In sensitive domains such as healthcare, predicting a plausible but contraindicated relation can have harmful consequences. This work addresses semantically grounded KG completion by extending the Partial Completeness Assumption (PCA) with two metrics—PCAvalid and PCAinvalid. These metrics distinguish constraint-compliant from constraint-violating predictions using SHACL validation. They guide the selection of symbolic rules and the generation of labeled training data for neural models. As a result, link prediction systems can assess both plausibility and semantic validity. Experiments on 315 testbeds demonstrate that incorporating constraint-aware symbolic knowledge enhances MRR and Hits@K across multiple KG embedding models, including TransE and TransH. Thus, this approach supports interpretable and trustworthy KG completion.
Disha Purohit, Yashrajsinh Chudasama, Maria-Esther Vidal
K-CAP3
2025 From Legal Texts to Structured Knowledge: A Comprehensive Pipeline for Legal Text Summarization
Ahmad Sakor, Kuldeep Singh 0001, Maria-Esther Vidal
ISWC (2)3
2025 HyKG-CF: A Hybrid Approach for Counterfactual Prediction using Domain Knowledge
Hao Huang 0014, Maria-Esther Vidal
WSDM2
2025 Enhancing Medical Knowledge Discovery: A Neuro-symbolic System for Inductive Learning over Medical KGs
abstract
Medical knowledge graphs (KGs) excel at integrating heterogeneous healthcare data with domain knowledge, but face challenges due to incompleteness. While Knowledge Graph Embedding (KGE) models show promise in link prediction, they often fail to incorporate crucial semantic constraints from medical ontologies and clinical guidelines. We propose a neuro-symbolic system that enhances medical knowledge discovery by combining symbolic learning from medical ontologies, inductive learning through KGE, and semantic constraint validation. Applied to lung cancer care, our system demonstrates enhanced performance in predicting novel medical relationships while maintaining semantic consistency with medical knowledge. Experimental results show our approach enhances the KGE model's performance while ensuring clinical validity and the implementation is publicly accessible on GitHub https://github.com/SDM-TIB/KOSMOS.
Disha Purohit, Yashrajsinh Chudasama, Maria-Esther Vidal
WSDM3
2025 BioLinkerAI: Leveraging LLMs to Improve Biomedical Entity Linking and Knowledge Capture
Ahmad Sakor, Kuldeep Singh 0001, Maria-Esther Vidal
WSDM3
2025 Integrating Knowledge Graphs and Neuro-Symbolic AI: LDM Enables FAIR and Federated Research Data Management
abstract
Managing research digital objects (RDOs) in compliance with FAIR principles is crucial for ensuring accessibility, interoperability, and reusability across scientific domains. The Leibniz Data Manager (LDM) is a state-of-the-art framework that integrates Knowledge Graphs (KGs) and Neuro-Symbolic AI, combining the reasoning power of Large Language Models (LLMs) with structured metadata. LDM supports the management and enhancement of RDOs through entity linking, connecting datasets to external KGs like Wikidata and the Open Research Knowledge Graph (ORKG). Additionally, LDM offers federated query processing across KGs, enabling users to explore related papers, datasets, and resources through natural language questions. This demo showcases LDM's capabilities to explore RDOs, compare existing datasets, and extend metadata. By blending Neuro-Symbolic AI with FAIR and federated research data management, LDM offers a powerful tool for accelerating data-driven discovery in science. LDM is publicly accessible at https://service.tib.eu/ldmservice/.
Ahmad Sakor, Mauricio Brunet, Enrique Iglesias, Ariam Rivas, Philipp D. Rohde, Angelina Kraft, Maria-Esther Vidal
WSDM7
2025 From Genesis to Maturity: Managing Knowledge Graph Ecosystems Through Life Cycles
abstract
Knowledge graphs (KGs) play a crucial role in the integration and organization of heterogeneous data and knowledge, enabling advanced data analytics and decision-making across various industries. This vision paper addresses critical challenges in managing KGs, emphasizing their relevance in integrating information from disparate sources. We propose the concept of knowledge graph ecosystems and life cycles to systematically manage tasks, e.g., data integration, standardization, continuous updates, efficient querying, and provenance tracking. By adopting our approach, organizations can enhance the accuracy, consistency, and reliability of KGs, thus improving knowledge management, enabling the extraction of valuable insights, and ensuring transparency and accountability.
Sandra Geisler, Cinzia Cappiello, Irene Celino, David Fraga 0001, Anastasia Dimou, Ana Iglesias-Molina, Maurizio Lenzerini, Anisa Rula, Dylan Van Assche, Sascha Welten, Maria-Esther Vidal
Proc. VLDB Endow.11
2025 Integrating Knowledge Graphs with Symbolic AI: The Path to Interpretable Hybrid AI Systems in Medicine
abstract
Knowledge Graphs (KGs) are graph-based structures that integrate heterogeneous data, capture domain knowledge, and enable explainable AI through symbolic reasoning. This position paper examines the challenges and research opportunities in integrating KGs with neuro-symbolic AI, highlighting their potential to enhance explainability, scalability, and context-aware reasoning in hybrid AI systems. Using a lung cancer use case, we illustrate how hybrid approaches address tasks such as link prediction—uncovering hidden relationships in medical data—and counterfactual reasoning—analyzing alternative scenarios to understand causal factors. The discussion is framed around TrustKG, which demonstrates how constraint validation, causal reasoning, and user-centric communication can support transparent and reliable decision-making. Additionally, we identify current limitations of KGs, including gaps in knowledge coverage, evolving data integration challenges, and the need for improved usability and impact assessment. These insights are not limited to healthcare but extend to other domains like energy, manufacturing, and mobility, showcasing the broad applicability of KGs. Finally, we propose research directions to unlock their full potential in building robust, transparent, and widely adopted real-world applications.
Maria-Esther Vidal, Yashrajsinh Chudasama, Hao Huang 0014, Disha Purohit, Maria Torrente
J. Web Semant.1
2024 Discovering Relationships Among Properties in Wikidata Knowledge Graph
Emetis Niazmand, Maria-Esther Vidal
DaWaK2
2024 Completing Predicates Based on Alignment Rules from Knowledge Graphs
Emetis Niazmand, Maria-Esther Vidal
DEXA (1)2
2024 SemMatch: Semantics-Aware Matching for Causal Inference over Knowledge Graphs
Hao Huang 0014, Maria-Esther Vidal
WISE (2)2
2024 BioLinkerAI: Capturing Knowledge Using LLMs to Enhance Biomedical Entity Linking
Ahmad Sakor, Kuldeep Singh 0001, Maria-Esther Vidal
WISE (4)3
2023 Unraveling the Hepatitis B Cure: A Hybrid AI Approach for Capturing Knowledge about the Immune System's Impact
abstract
Chronic hepatitis B virus (HBV) infection is still a global health problem, with over 296 million chronically HBV-infected individuals worldwide. The merging data about clinical parameters, immune phenotyping data, and genetic information, together with AI models reliant on this integrated information, holds promise in effectively predicting the likelihood of functional cure in HBV-infected patients. Yet, the limited size of multidimensional datasets and characteristic of HBV cases poses a challenge for machine learning (ML) systems that typically require substantial data for pattern recognition. This paper addresses this challenge by introducing HyAI, a hybrid AI framework. HyAI employs knowledge graphs (KGs) and inductive learning to unearth meaningful patterns. HyAI relies on KG embedding models to learn a numerical representation of the HyAI KG in a k-dimensional vector space. Through community detection methods, closely related HBV patients are clustered using similarity metrics formulated from the acquired embeddings. HyAI is studied in a population of HBV patients integrated with multidimensional datasets. Our empirical analysis shows that HyAI uncovers immune markers that, together with clinical and demographic parameters, correspond to good predictors for forecasting the cure of chronic HBV infection.
Shahi Dost, Ariam Rivas, Hanan Begali, Annett Ziegler, Elimira Aliabadi, Markus Cornberg, Anke Rm Kraft, Maria-Esther Vidal
K-CAP8
2023 SPaRKLE : Symbolic caPtuRing of knowledge for Knowledge graph enrichment with LEarning
abstract
Knowledge graphs (KGs) naturally capture the convergence of data and knowledge, making them expressive frameworks for describing and integrating heterogeneous data in a coherent and interconnected manner. However, based on the Open World Assumption (OWA), the absence of information within KGs does not indicate falsity or non-existence; it merely reflects incompleteness. Inductive learning over KGs involves predicting new relationships based on existing statements in the KG, using either numerical or symbolic learning models. The Partial Completeness Assumption (PCA) heuristic efficiently guides inductive learning methods for Link Prediction (LP) by refining predictions about absent KG relationships. Nevertheless, numeric techniques– like KG embedding models– alone may fall short in accurately predicting missing information, particularly when it comes to capturing implicit knowledge and complex relationships. We propose a hybrid method named SPaRKLE that seamlessly integrates symbolic and numerical techniques, leveraging the PCA heuristic to capture implicit knowledge and enrich KGs. We empirically compare SPaRKLE with state-of-the-art KG embedding and symbolic models, using established benchmarks. Our experimental outcomes underscore the efficacy of this hybrid approach, as it harnesses the strengths of both paradigms. SPaRKLE is publicly available on GitHub1.
Disha Purohit, Yashrajsinh Chudasama, Ariam Rivas, Maria-Esther Vidal
K-CAP4
2023 Scaling up knowledge graph creation to large and heterogeneous data sources
Enrique Iglesias, Samaneh Jozashoori, Maria-Esther Vidal
J. Web Semant.3
2023 Knowledge4COVID-19: A semantic-based approach for constructing a COVID-19 related knowledge graph from various sources and analyzing treatments' toxicities
Ahmad Sakor, Samaneh Jozashoori, Emetis Niazmand, Ariam Rivas, Konstantinos Bougiatiotis, Fotis Aisopos, Enrique Iglesias, Philipp D. Rohde, Trupti Padiya, Anastasia Krithara, Georgios Paliouras, Maria-Esther Vidal
J. Web Semant.12
2021 Capturing Knowledge about Drug-Drug Interactions to Enhance Treatment Effectiveness
abstract
Capturing knowledge about Drug-Drug Interactions (DDI) is a crucial factor to support clinicians in better treatments. Nowadays, public drug databases provide a wealth of information on drugs that can be exploited to enhance tasks, e.g., data mining, ranking, and query answering. However, all the interactions in the public database are focused on pairs of drugs. Since current treatments are composed of multi-drugs, it is extremely challenging to know which potential drugs affect the effectiveness of the treatment. In this work, we tackle the problem of discovering DDIs and reduce this problem to link prediction over a property graph represented in RDF-star. A deductive system captures knowledge about the conditions that define when a group of drugs interacts as Datalog rules. Extensional statements represent the property graph. Lastly, the intensional rules guide the deduction process to discover relationships in the graph and their properties. As a proof concept, we have implemented a graph traversal method on top of the property graph and the deduced edges. The technique aims to identify the combination of drugs whose interactions may reduce the effectiveness of a treatment or increase the number of toxicities. This traversal method relies on the computation of wedges in the property graph. Albeit illustrated in the context of DDI, this method could be generalized to other link traversal tasks. We conduct an experimental study on a DDIs property graph for different treatments. The results suggest that by capturing knowledge about DDIs, our approach can discover the drugs that decrease the effectiveness of the treatment. Our results are promising and suggest that clinicians can better understand the DDIs in treatment and prescribe improved treatments through the knowledge captured by our approach.
Ariam Rivas, Maria-Esther Vidal
K-CAP2
2021 Trav-SHACL: Efficiently Validating Networks of SHACL Constraints
abstract
Knowledge graphs have emerged as expressive data structures for Web data. Knowledge graph potential and the demand for ecosystems to facilitate their creation, curation, and understanding, is testified in diverse domains, e.g., biomedicine. The Shapes Constraint Language (SHACL) is the W3C recommendation language for integrity constraints over RDF knowledge graphs. Enabling quality assements of knowledge graphs, SHACL is rapidly gaining attention in real-world scenarios. SHACL models integrity constraints as a network of shapes, where a shape contains the constraints to be fullfiled by the same entities. The validation of a SHACL shape schema can face the issue of tractability during validation. To facilitate full adoption, efficient computational methods are required. We present Trav-SHACL, a SHACL engine capable of planning the traversal and execution of a shape schema in a way that invalid entities are detected early and needless validations are minimized. Trav-SHACL reorders the shapes in a shape schema for efficient validation and rewrites target and constraint queries for fast detection of invalid entities. Trav-SHACL is empirically evaluated on 27 testbeds executed against knowledge graphs of up to 34M triples. Our experimental results suggest that Trav-SHACL exhibits high performance gradually and reduces validation time by a factor of up to 28.93 compared to the state of the art.
Mónica Figuera, Philipp D. Rohde, Maria-Esther Vidal
WWW3
2021 Compact representations for efficient storage of semantic sensor data
Farah Karim, Maria-Esther Vidal, Sören Auer
J. Intell. Inf. Syst.2
2020 SDM-RDFizer: An RML Interpreter for the Efficient Creation of RDF Knowledge Graphs
abstract
In recent years, the amount of data has increased exponentially, and knowledge graphs have gained attention as data structures to integrate data and knowledge harvested from myriad data sources. However, data complexity issues like large volume, high-duplicate rate, and heterogeneity usually characterize these data sources, being required data management tools able to address the negative impact of these issues on the knowledge graph creation process. In this paper, we propose the SDM-RDFizer, an interpreter of the RDF Mapping Language (RML), to transform raw data in various formats into an RDF knowledge graph. SDM-RDFizer implements novel algorithms to execute the logical operators between mappings in RML, allowing thus to scale up to complex scenarios where data is not only broad but has a high-duplication rate. We empirically evaluate the SDM-RDFizer performance against diverse testbeds with diverse configurations of data volume, duplicates, and heterogeneity. The observed results indicate that SDM-RDFizer is two orders of magnitude faster than state of the art, thus, meaning that SDM-RDFizer an interoperable and scalable solution for knowledge graph creation. SDM-RDFizer is publicly available as a resource through a Github repository and a DOI.
Enrique Iglesias, Samaneh Jozashoori, David Fraga 0001, Diego Collarana, Maria-Esther Vidal
CIKM5
2020 Falcon 2.0: An Entity and Relation Linking Tool over Wikidata
abstract
The Natural Language Processing (NLP) community has significantly contributed to the solutions for entity and relation recognition from a natural language text, and possibly linking them to proper matches in Knowledge Graphs (KGs). Considering Wikidata as the background KG, there are still limited tools to link knowledge within the text to Wikidata. In this paper, we present Falcon 2.0, the first joint entity and relation linking tool over Wikidata. It receives a short natural language text in the English language and outputs a ranked list of entities and relations annotated with the proper candidates in Wikidata. The candidates are represented by their Internationalized Resource Identifier (IRI) in Wikidata. Falcon 2.0 resorts to the English language model for the recognition task (e.g., N-Gram tiling and N-Gram splitting), and then an optimization approach for the linking task. We have empirically studied the performance of Falcon 2.0 on Wikidata and concluded that it outperforms all the existing baselines. Falcon 2.0 is open source and can be reused by the community; all the required instructions of Falcon 2.0 are well-documented at our GitHub repository (https://github.com/SDM-TIB/falcon2.0). We also demonstrate an online API, which can be run without any technical expertise. Falcon 2.0 and its background knowledge bases are available as resources at https://labs.tib.eu/falcon/falcon2/.
Ahmad Sakor, Kuldeep Singh 0001, Anery Patel, Maria-Esther Vidal
CIKM4
2020 Unveiling Relations in the Industry 4.0 Standards Landscape Based on Knowledge Graph Embeddings
Ariam Rivas, Irlán Grangel-González, Diego Collarana, Jens Lehmann 0001, Maria-Esther Vidal
DEXA (2)5
2020 A Knowledge Graph for Industry 4.0
Sebastian R. Bader, Irlán Grangel-González, Priyanka Nanjappa, Maria-Esther Vidal, Maria Maleshkova
ESWC4
2020 Creating and Capturing Artificial Emotions in Autonomous Robots and Software Agents
Claus Hoffmann, Maria-Esther Vidal
ICWE2
2020 FunMap: Efficient Execution of Functional Mappings for Knowledge Graph Creation
Samaneh Jozashoori, David Fraga 0001, Enrique Iglesias, Maria-Esther Vidal, Óscar Corcho
ISWC (1)4
2020 Encoding Knowledge Graph Entity Aliases in Attentive Neural Network for Wikidata Entity Linking
Isaiah Onando Mulang', Kuldeep Singh 0001, Akhilesh Vyas, Saeedeh Shekarpour, Maria-Esther Vidal, Sören Auer
WISE (1)5
2020 Compacting frequent star patterns in RDF graphs
Farah Karim, Maria-Esther Vidal, Sören Auer
J. Intell. Inf. Syst.2
2020 No one is perfect: Analysing the performance of question answering components over the DBpedia knowledge graph
Kuldeep Singh 0001, Ioanna Lytra, Arun Sethupat Radhakrishna, Saeedeh Shekarpour, Maria-Esther Vidal, Jens Lehmann 0001
J. Web Semant.5
2019 Ontario: Federated Query Processing Against a Semantic Data Lake
Kemele M. Endris, Philipp D. Rohde, Maria-Esther Vidal, Sören Auer
DEXA (1)3
2019 PURE: A Privacy Aware Rule-Based Framework over Knowledge Graphs
Marlene Goncalves, Maria-Esther Vidal, Kemele M. Endris
DEXA (1)2
2019 COMET: A Contextualized Molecule-Based Matching Technique
Mayesha Tasnim, Diego Collarana, Damien Graux, Michael Galkin, Maria-Esther Vidal
DEXA (1)5
2019 Semantic Representation of Scientific Publications
Sahar Vahdati, Said Fathalla, Sören Auer, Christoph Lange 0002, Maria-Esther Vidal
TPDL5
2019 Ranking Knowledge Graphs By Capturing Knowledge about Languages and Labels
abstract
Capturing knowledge about the mulitilinguality of a knowledge graph is of supreme importance to understand its applicability across multiple languages. Several metrics have been proposed for describing mulitilinguality at the level of a whole knowledge graph. Albeit enabling the understanding of the ecosystem of knowledge graphs in terms of the utilized languages, they are unable to capture a fine-grained description of the languages in which the different entities and properties of the knowledge graph are represented. This lack of representation prevents the comparison of existing knowledge graphs in order to decide which are the most appropriate for a multilingual application.
Lucie-Aimée Kaffee, Kemele M. Endris, Elena Simperl, Maria-Esther Vidal
K-CAP4
2019 Managing the evolution and preservation of the data web
Jeremy Debattista, Javier D. Fernández, Maria-Esther Vidal, Jürgen Umbrich
J. Web Semant.3
2018 BOUNCER: Privacy-Aware Query Processing over Federations of RDF Datasets
Kemele M. Endris, Zuhair Almhithawi, Ioanna Lytra, Maria-Esther Vidal, Sören Auer
DEXA (1)4
2018 Knowledge Graphs for Semantically Integrating Cyber-Physical Systems
Irlán Grangel-González, Lavdim Halilaj, Maria-Esther Vidal, Omar Rana, Steffen Lohmann, Sören Auer, Andreas W. Müller
DEXA (1)3
2018 GARUM: A Semantic Similarity Measure Based on Machine Learning and Entity Characteristics
Ignacio Traverso Ribón, Maria-Esther Vidal
DEXA (1)2
2018 Unveiling Scholarly Communities over Knowledge Graphs
Sahar Vahdati, Guillermo Palma, Rahul Jyoti Nath, Christoph Lange 0002, Sören Auer, Maria-Esther Vidal
TPDL6
2018 Intelligent Clients for Replicated Triple Pattern Fragments
Thomas Minier, Hala Skaf-Molli, Pascal Molli, Maria-Esther Vidal
ESWC4
2018 OpenBudgets.eu: A Platform for Semantically Representing and Analyzing Open Fiscal Data
Fathoni A. Musyaffa, Lavdim Halilaj, Fabrizio Orlandi, Hajira Jabeen, Sören Auer, Maria-Esther Vidal
ICWE7
2018 Synthesizing Knowledge Graphs from Web Sources with the MINTE ^+ + Framework
Diego Collarana, Michael Galkin, Christoph Lange 0002, Simon Scerri, Sören Auer, Maria-Esther Vidal
ISWC (2)6
2018 Dynamic Composition of Question Answering Pipelines with FRANKENSTEIN
abstract
Question answering (QA) systems provide user-friendly interfaces for retrieving answers from structured and unstructured data given natural language questions. Several QA systems, as well as related components, have been contributed by the industry and research community in recent years. However, most of these efforts have been performed independently from each other and with different focuses, and their synergies in the scope of QA have not been addressed adequately. FRANKENSTEIN is a novel framework for developing QA systems over knowledge bases by integrating existing state-of-the-art QA components performing different tasks. It incorporates several reusable QA components, employs machine learning techniques to predict best performing components and QA pipelines for a given question, and generates static and dynamic executable QA pipelines. In this paper, we illustrate different functionalities of FRANKENSTEIN for performing independent QA component execution, QA component prediction, given an input question as well as the static and dynamic composition of different QA pipelines.
Kuldeep Singh 0001, Ioanna Lytra, Arun Sethupat Radhakrishna, Akhilesh Vyas, Maria-Esther Vidal
SIGIR5
2018 Why Reinvent the Wheel: Let's Build Question Answering Systems Together
abstract
Modern question answering (QA) systems need to flexibly integrate a number of components specialised to fulfil specific tasks in a QA pipeline. Key QA tasks include Named Entity Recognition and Disambiguation, Relation Extraction, and Query Building. Since a number of different software components exist that implement different strategies for each of these tasks, it is a major challenge to select and combine the most suitable components into a QA system, given the characteristics of a question. We study this optimisation problem and train classifiers, which take features of a question as input and have the goal of optimising the selection of QA components based on those features. We then devise a greedy algorithm to identify the pipelines that include the suitable components and can effectively answer the given question. We implement this model within Frankenstein, a QA framework able to select QA components and compose QA pipelines. We evaluate the effectiveness of the pipelines generated by Frankenstein using the QALD and LC-QuAD benchmarks. These results not only suggest that Frankenstein precisely solves the QA optimisation problem but also enables the automatic composition of optimised QA pipelines, which outperform the static Baseline QA pipeline. Thanks to this flexible and fully automated pipeline generation process, new QA components can be easily included in Frankenstein, thus improving the performance of the generated pipelines.
Kuldeep Singh 0001, Arun Sethupat Radhakrishna, Andreas Both 0001, Saeedeh Shekarpour, Ioanna Lytra, Ricardo Usbeck, Akhilesh Vyas, Akmal Khikmatullaev, Dharmen Punjani, Christoph Lange 0002, Maria-Esther Vidal, Jens Lehmann 0001, Sören Auer
WWW11
2017 MULDER: Querying the Linked Data Web by Bridging RDF Molecule Templates
Kemele M. Endris, Michael Galkin, Ioanna Lytra, Mohamed Nadjib Mami, Maria-Esther Vidal, Sören Auer
DEXA (1)5
2017 SJoin: A Semantic Join Operator to Integrate Heterogeneous RDF Graphs
Michael Galkin, Diego Collarana, Ignacio Traverso Ribón, Maria-Esther Vidal, Sören Auer
DEXA (1)4
2017 QAestro - Semantic-Based Composition of Question Answering Pipelines
Kuldeep Singh 0001, Ioanna Lytra, Maria-Esther Vidal, Dharmen Punjani, Harsh Thakkar, Christoph Lange 0002, Sören Auer
DEXA (1)3
2017 Towards an Integrated Graph Algebra for Graph Pattern Matching with Gremlin
Harsh Thakkar, Dharmen Punjani, Sören Auer, Maria-Esther Vidal
DEXA (1)4
2017 Integration of Scholarly Communication Metadata Using Knowledge Graphs
Afshin Sadeghi, Christoph Lange 0002, Maria-Esther Vidal, Sören Auer
TPDL3
2017 The BigDataEurope Platform - Supporting the Variety Dimension of Big Data
Sören Auer, Simon Scerri, Aad Versteden, Erika Pauwels, Angelos Charalambidis, Stasinos Konstantopoulos, Jens Lehmann 0001, Hajira Jabeen, Ivan Ermilov, Gezim Sejdiu, Andreas Ikonomopoulos, Spyros Andronopoulos, Mandy Vlachogiannis, Charalambos Pappas, Athanasios Davettas, Iraklis A. Klampanos, Efstathios Grigoropoulos, Vangelis Karkaletsis, Victor de Boer, Ronny Siebes, Mohamed Nadjib Mami, Sergio Albani, Michele Lazzarini, Paulo Nunes, Emanuele Angiuli, Nikiforos Pittaras, George Giannakopoulos, Giorgos Argyriou, George Stamoulis 0001, George Papadakis 0001, Manolis Koubarakis, Pythagoras Karampiperis, Axel-Cyrille Ngonga Ngomo, Maria-Esther Vidal
ICWE34
2017 MateTee: A Semantic Similarity Metric Based on Translation Embeddings for Knowledge Graphs
Camilo Morales, Diego Collarana, Maria-Esther Vidal, Sören Auer
ICWE3
2017 Capturing Knowledge in Semantically-typed Relational Patterns to Enhance Relation Linking
abstract
Transforming natural language questions into formal queries is an integral task in Question Answering (QA) systems. QA systems built on knowledge graphs like DBpedia, require a step after natural language processing for linking words, specifically including named entities and relations, to their corresponding entities in a knowledge graph. To achieve this task, several approaches rely on background knowledge bases containing semantically-typed relations, e.g., PATTY, for an extra disambiguation step. Two major factors may affect the performance of relation linking approaches whenever background knowledge bases are accessed: a) limited availability of such semantic knowledge sources, and b) lack of a systematic approach on how to maximize the benefits of the collected knowledge. We tackle this problem and devise SIBKB, a semantic-based index able to capture knowledge encoded on background knowledge bases like PATTY. SIBKB represents a background knowledge base as a bi-partite and a dynamic index over the relation patterns included in the knowledge base. Moreover, we develop a relation linking component able to exploit SIBKB features. The benefits of SIBKB are empirically studied on existing QA benchmarks and observed results suggest that SIBKB is able to enhance the accuracy of relation linking by up to three times.
Kuldeep Singh 0001, Isaiah Onando Mulang', Ioanna Lytra, Mohamad Yaser Jaradeh, Ahmad Sakor, Maria-Esther Vidal, Christoph Lange 0002, Sören Auer
K-CAP6
2017 Diefficiency Metrics: Measuring the Continuous Efficiency of Query Processing Approaches
abstract
During empirical evaluations of query processing techniques, metrics like execution time, time for the first answer, and throughput are usually reported. Albeit informative, these metrics are unable to quantify and evaluate the efficiency of a query engine over a certain time period – or diefficiency –, thus hampering the distinction of cutting-edge engines able to exhibit high-performance gradually. We tackle this issue and devise two experimental metrics named dief@t and dief@k , which allow for measuring the diefficiency during an elapsed time period t or while k answers are produced, respectively. The dief@t and dief@k measurement methods rely on the computation of the area under the curve of answer traces, and thus capturing the answer concentration over a time interval. We report experimental results of evaluating the behavior of a generic SPARQL query engine using both metrics. Observed results suggest that dief@t and dief@k are able to measure the performance of SPARQL query engines based on both the amount of answers produced by an engine and the time required to generate these answers.
Maribel Acosta, Maria-Esther Vidal, York Sure-Vetter
ISWC (2)2
2017 Enhancing answer completeness of SPARQL queries via crowdsourcing
Maribel Acosta, Elena Simperl, Fabian Flöck, Maria-Esther Vidal
J. Web Semant.4
2017 Decomposing federated queries in presence of replicated fragments
Gabriela Montoya, Hala Skaf-Molli, Pascal Molli, Maria-Esther Vidal
J. Web Semant.4
2016 Towards Semantification of Big Data Technology
Mohamed Nadjib Mami, Simon Scerri, Sören Auer, Maria-Esther Vidal
DaWaK4
2016 Alligator: A Deductive Approach for the Integration of Industry 4.0 Standards
Irlán Grangel-González, Diego Collarana, Lavdim Halilaj, Steffen Lohmann, Christoph Lange 0002, Maria-Esther Vidal, Sören Auer
EKAW6
2016 Considering Semantics on the Discovery of Relations in Knowledge Graphs
Ignacio Traverso Ribón, Guillermo Palma, Alejandro Flores-Velazco, Maria-Esther Vidal
EKAW4
2016 Proactive Prevention of False-Positive Conflicts in Distributed Ontology Development
abstract
S.43-51
Lavdim Halilaj, Irlán Grangel-González, Maria-Esther Vidal, Steffen Lohmann, Sören Auer
KEOD3
2016 SerVCS: Serialization Agnostic Ontology Development in Distributed Settings
Lavdim Halilaj, Irlán Grangel-González, Maria-Esther Vidal, Steffen Lohmann, Sören Auer
IC3K3
2016 Co-evolution of RDF Datasets
Sidra Faisal, Kemele M. Endris, Saeedeh Shekarpour, Sören Auer, Maria-Esther Vidal
ICWE5
2015 D-FOPA: A Dynamic Final Object Pruning Algorithm to Efficiently Produce Skyline Points Over Data Streams
Stephanie Alibrandi, Sofia Bravo, Marlene Goncalves, Maria-Esther Vidal
DEXA (2)4
2015 HARE: A Hybrid SPARQL Engine to Enhance Query Answers via Crowdsourcing
abstract
Due to the semi-structured nature of RDF data, missing values affect answer completeness of queries that are posed against RDF. To overcome this limitation, we present HARE, a novel hybrid query processing engine that brings together machine and human computation to execute SPARQL queries. We propose a model that exploits the characteristics of RDF in order to estimate the completeness of portions of a data set. The completeness model complemented by crowd knowledge is used by the HARE query engine to on-the-fly decide which parts of a query should be executed against the data set or via crowd computing. To evaluate HARE, we created and executed a collection of 50 SPARQL queries against the DBpedia data set. Experimental results clearly show that our solution accurately enhances answer completeness.
Maribel Acosta, Elena Simperl, Fabian Flöck, Maria-Esther Vidal
K-CAP4
2015 Networks of Linked Data Eddies: An Adaptive Web Query Processing Engine for RDF Data
Maribel Acosta, Maria-Esther Vidal
ISWC (1)2
2015 Federated SPARQL Queries Processing with Replicated Fragments
Gabriela Montoya, Hala Skaf-Molli, Pascal Molli, Maria-Esther Vidal
ISWC (1)4
2014 Drug-Target Interaction Prediction Using Semantic Similarity and Edge Partitioning
Guillermo Palma, Maria-Esther Vidal, Louiqa Raschid
ISWC (1)2
2013 FOPA: A Final Object Pruning Algorithm to Efficiently Produce Skyline Points
Ana Alvarado, Oriana Baldizan, Marlene Goncalves, Maria-Esther Vidal
DEXA (2)4
2013 GUN: An Efficient Execution Strategy for Querying the Web of Data
Gabriela Montoya, Luis-Daniel Ibáñez, Hala Skaf-Molli, Pascal Molli, Maria-Esther Vidal
DEXA (1)5
2012 Benchmarking Federated SPARQL Query Engines: Are Existing Testbeds Enough?
Gabriela Montoya, Maria-Esther Vidal, Óscar Corcho, Edna Ruckhaus, Carlos Buil-Aranda
ISWC (2)2
2012 PAnG: finding patterns in annotation graphs
abstract
Annotation graph datasets are a natural representation of scientific knowledge. They are common in the life sciences and health sciences, where concepts such as genes, proteins or clinical trials are annotated with controlled vocabulary terms from ontologies. We present a tool, PAnG (Patterns in Annotation Graphs), that is based on a complementary methodology of graph summarization and dense subgraphs. The elements of a graph summary correspond to a pattern and its visualization can provide an explanation of the underlying knowledge. Scientists can use PAnG to develop hypotheses and for exploration.
Philip Anderson 0003, Andreas Thor, Joseph Benik, Louiqa Raschid, Maria-Esther Vidal
SIGMOD Conference5
2011 ANAPSID: An Adaptive Query Processing Engine for SPARQL Endpoints
Maribel Acosta, Maria-Esther Vidal, Tomas Lampo, Julio Castillo, Edna Ruckhaus
ISWC (1)2
2010 Efficiently Joining Group Patterns in SPARQL Queries
Maria-Esther Vidal, Edna Ruckhaus, Tomas Lampo, Amadís Martínez, Javier Sierra, Axel Polleres
ESWC (1)1
2010 BioNav: An Ontology-Based Framework to Discover Semantic Links in the Cloud of Linked Data
Maria-Esther Vidal, Louiqa Raschid, Natalia Marquez, Jean Carlo Rivera, Edna Ruckhaus
ESWC (2)1
2010 A sampling-based approach to identify QoS for web service orchestrations
abstract
QoS parameters are used to describe services in terms of their behavior and can be used to rank services according to non-functional criteria. To provide an accurate characterization of the quality of a service, we propose a sampling-based technique. The proposed technique uses Adaptive and Sequential Sampling strategies to estimate the QoS parameters that satisfy the required confidence levels while the size of the sample remains small. QoS estimates are used by a hybrid composer, named PT-SAM, to identify the service compositions that satisfy a functional condition and best meet non-functional criteria of a user query. PT-SAM adapts a Petri-Net unfolding algorithm to find a desired marking from an initial state by using a utility function defined on QoS estimates and functional properties of the available services. PT-SAM uses a QoS-based utility function to guide the search into portions of good quality service compositions; thus, PT-SAM is able to scale up to large-scale search spaces of services. We report on the quality of the sampling techniques and the performance of the composer. First, we show correlation between the estimates and the real values of the QoS parameters; then, we report on the benefits of using these estimates to traverse large search spaces of service compositions (e.g., in the range of 1,000 to 100,000 services). Our experiments show that the quality of the compositions identified by our algorithm is close to the optimal solution produced by the exhaustive algorithm.
Eduardo Blanco 0001, Yudith Cardinale, Maria-Esther Vidal
iiWAS3
2010 An Expressive and Efficient Solution to the Service Selection Problem
Maria-Esther Vidal, Blai Bonet
ISWC (1)2
2009 Reaching the Top of the Skyline: An Efficient Indexed Algorithm for Top-k Skyline Queries
Marlene Goncalves, Maria-Esther Vidal
DEXA2
2009 Flexible and efficient querying and ranking on hyperlinked data sources
abstract
There has been an explosion of hyperlinked data in many domains, e.g., the biological Web. Expressive query languages and effective ranking techniques are required to convert this data into browsable knowledge. We propose the Graph Information Discovery (GID) framework to support sophisticated user queries on a rich web of annotated and hyperlinked data entries, where query answers need to be ranked in terms of some customized ranking criteria, e.g., PageRank or ObjectRank. GID has a data model that includes a schema graph and a data graph, and an intuitive query interface. The GID framework allows users to easily formulate queries consisting of sequences of hard filters (selection predicates) and soft filters (ranking criteria); it can also be combined with other specialized graph query languages to enhance their ranking capabilities. GID queries have a well-defined semantics and are implemented by a set of physical operators, each of which produces a ranked result graph. We discuss rewriting opportunities to provide an efficient evaluation of GID queries. Soft filters are a key feature of GID and they are implemented using authority flow ranking techniques; these are query dependent rankings and are expensive to compute at runtime. We present approximate optimization techniques for GID soft filter queries based on the properties of random walks, and using novel path-length-bound and graph-sampling approximation techniques. We experimentally validate our optimization techniques on large biological and bibliographic datasets. Our techniques can produce high quality (Top K) answers with a savings of up to an order of magnitude, in comparison to the evaluation time for the exact solution.
Ramakrishna Varadarajan, Vagelis Hristidis, Louiqa Raschid, Maria-Esther Vidal, Luis-Daniel Ibáñez, Héctor Rodríguez-Drumond
EDBT4
2008 BiOnMap: a deductive approach for resource discovery
abstract
We present a deductive approach that supports resource discovery. The BiOnMap Web service is designed to support the selection of resources suitable to implement specific tasks. The BiOnMap service is comprised of a metadata catalog and a reasoning engine. The metadata catalog uses domain ontologies to annotate resources semantically and express domain rules that capture path equivalences at the level of the ontology graph. The BiOnMap reasoning engine is able to infer new properties of services to support service discovery, composition, and mapping. We illustrate our approach with an application case from the domain of bioinformatics.
Nadia Yaacoubi Ayadi, Zoé Lacroix, Maria-Esther Vidal
iiWAS3
2005 Preferred Skyline: A Hybrid Approach Between SQLf and Skyline
Marlene Goncalves, Maria-Esther Vidal
DEXA2
2005 A Data Model and Query Language to Explore Enhanced Links and Paths in Life Science Sources
George A. Mihaila, Felix Naumann, Louiqa Raschid, Maria-Esther Vidal
WebDB4
2004 Exploiting Multiple Paths to Express Scientific Queries
Zoé Lacroix, Tiffany Morris, Kaushal Parekh, Louiqa Raschid, Maria-Esther Vidal
SSDBM5
2004 Challenges in Selecting Paths for Navigational Queries: Trade-Off of Benefit of Path versus Cost of Plan
abstract
Life sciences sources are characterized by a complex graph of overlapping sources, and multiple alternate links between sources. A (navigational) query may be answered by traversing multiple alternate paths between a start source and a target source. Each of these paths may have dissimilar benefit, e.g., the cardinality of result objects that are reached in the target source. Paths may also have dissimilar costs of evaluation, i.e., the execution cost of a query evaluation plan for a path. In prior research, we developed ESearch, an algorithm based on a Deterministic Finite Automaton (DFA), which exhaustively enumerates all paths to answer a navigational query. The challenge is to develop heuristics that improve on the exhaustive ESearch solution and identify good utility functions that can rank the sources, the links between sources, and the sub-paths that are already visited, in order to quickly produce paths that have the highest benefit and the least cost. In this paper, we present a heuristic that uses local utility functions to rank sources, using either the benefit attributed to the source, the cost of a plan using the source, or both. The heuristic will limit its search to some Top XX% of the ranked sources. To compare ESearch and the heuristic, we construct a Pareto surface of all dominant solutions produced by ESearch, with respect to benefit and cost. We choose the Top 25% of the ESearch solutions that are in the Pareto surface. We compare the paths produced by the heuristic to this Top 25% of ESearch solutions with respect to precision and recall. This motivates the need for further research on developing a more efficient algorithm and better utility functions.
Maria-Esther Vidal, Louiqa Raschid, Julián Mestre
WebDB1
2002 Efficient evaluation of queries in a mediator for WebSources
abstract
We consider an architecture of mediators and wrappers for Internet accessible WebSources of limited query capability. Each call to a source is a WebSource Implementation (WSI) and it is associated with both a capability and (a possibly dynamic) cost. The multiplicity of WSIs with varying costs and capabilities increases the complexity of a traditional optimizer that must assign WSIs for each remote relation in the query while generating an (optimal) plan. We present a two-phase Web Query Optimizer (WQO). In a pre-optimization phase, the WQO selects one or more WSIs for a pre-plan; a pre-plan represents a space of query evaluation plans (plans) based on this choice of WSIs. The WQO uses cost-based heuristics to evaluate the choice of WSI assignment in the pre-plan and to choose a good pre-plan. The WQO uses the pre-plan to drive the extended relational optimizer to obtain the best plan for a pre-plan. A prototype of the WQO has been developed. We compare the effectiveness of the WQO, i.e., its ability to efficiently search a large space of plans and obtain a low cost plan, in comparison to a traditional optimizer. We also validate the cost-based heuristics by experimental evaluation of queries in the noisy Internet environment.
Vladimir Zadorozhny, Louiqa Raschid, Maria-Esther Vidal, Tolga Urhan, Laura Bright
SIGMOD Conference3
2000 Web Query Optimizer
abstract
We demonstrate a Web Query Optimizer (WQO) within an architecture of mediators and wrappers, for WebSources of limited capability in a wide area environment. The WQO has several innovative features, including a CBR (capability based rewriting) tool, an enhanced randomized relational optimizer extended to a Web environment, and a WebWrapper cost model that can provide relevant metrics for accessing WebSources. The prototype has been tested against a number of WebSources.
Vladimir Zadorozhny, Laura Bright, Louiqa Raschid, Tolga Urhan, Maria-Esther Vidal
ICDE5