EDBT 2026 Demo / reviewers in the wild / expert
Emanuel Sallinger
dblp:20/9117
· DBLP profile ↗
34ranked-venue papers in the field
0as first author
19since 2021 · last 2026
0000-0001-7441-129XORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 30Information Retrieval & Web Search · 2Knowledge Engineering, Semantic Web & Information Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-aware query answering with Large Language ModelsabstractIn the modern data-driven world, answering queries over heterogeneous and semantically inconsistent data remains a significant challenge. Modern datasets originate from diverse sources, such as relational databases, semi-structured repositories, and unstructured documents, leading to substantial variability in schemas, terminologies, and data formats. Traditional systems, constrained by rigid syntactic matching and strict data binding, struggle to capture critical semantic connections and schema ambiguities, failing to meet the growing demand among data scientists for advanced forms of flexibility and context-awareness in query answering. In parallel, the advent of Large Language Models (LLMs) has introduced new capabilities in natural language interpretation, making them highly promising for addressing such challenges. However, LLMs alone lack the systematic rigor and explainability required for robust query processing and decision-making in high-stakes domains. In this paper, we propose Soft Query Answering (Soft QA), a novel hybrid approach that integrates LLMs as an intermediate semantic layer within the query processing pipeline. Soft QA enhances query answering adaptability and flexibility by injecting semantic understanding through context-aware, schema-informed prompts, and leverages LLMs to semantically link entities, resolve ambiguities, and deliver accurate query results in complex settings. We demonstrate its practical effectiveness through real-world examples, highlighting its ability to resolve semantic mismatches and improve query outcomes without requiring extensive data cleaning or restructuring. Paolo Atzeni, Teodoro Baldazzi, Luigi Bellomarini, Eleonora Laurenza, Emanuel Sallinger |
Data Knowl. Eng. | 5 |
| 2026 | Using ontologies to facilitate healthcare process mining and analysisabstractHealthcare organisations collect detailed data on the care that they deliver. This data can be used to identify issues, including deviations from care standards and recommendations, and opportunities for improvement; it can be used also to support the development of new technologies and treatments. The volume and complexity of the data means that automated techniques such as process mining are needed to support the extraction and analysis of relevant information. This paper explains how the ontological information held in clinical terminologies can be used to facilitate process extraction and analysis, by connecting and aggregating clinical events through the classification of diagnoses made and treatments performed. The approach is demonstrated through application to data collected on care delivered to patients with cancer in a major hospital. The results are compared with those obtained from benchmark datasets using approaches in which connections and aggregations are proposed and curated by domain experts. This comparison highlights the potential, and the shortcomings, of ontology-based extraction and analysis in healthcare process mining. Supplementary Information: The online version contains supplementary material available at 10.1007/s10844-025-00942-8. Owen P. Dwyer, Lara Chammas, Emanuel Sallinger, Jim Davies |
J. Intell. Inf. Syst. | 3 |
| 2025 | Template-based Explainable Inference over High-Stakes Financial Knowledge Graphs
Andrea Colombo, Teodoro Baldazzi, Luigi Bellomarini, Emanuel Sallinger, Stefano Ceri |
EDBT | 4 |
| 2025 | Vadacode: A Logician-friendly IDE for Datalog+/-abstractLanguages, namely, fragments, of the Datalog+/- family are attracting interest in both academia and industry because of their possibility to balance high expressive power and computational complexity. However, understanding the differences among the fragments, mastering them to achieve scalable industrial applications, and communicating their peculiarities to a non-expert audience is challenging for researchers, developers, logicians, and educators. In this demo, we introduce Vadacode, an IDE for Datalog+/- designed to support a broad category of users. The tool offers advanced features, including fragment detection, syntax highlighting, code completion, error diagnostics, schema inference, debugging support, and AI-assisted coding capabilities. Thanks to our experience in the financial context, our demo will guide the audience in modeling financial Datalog+/- programs, showcasing a seamless and effective coding experience. Luigi Bellomarini, Andrea Gentili 0005, Davide Magnanimi, Emanuel Sallinger |
Proc. VLDB Endow. | 4 |
| 2024 | "Please, Vadalog, tell me why": Interactive Explanation of Datalog-based Reasoning
Teodoro Baldazzi, Luigi Bellomarini, Stefano Ceri, Andrea Colombo, Andrea Gentili 0005, Emanuel Sallinger |
EDBT | 6 |
| 2024 | The Vadalog Parallel System: Distributed Reasoning with Datalog+/-abstractOver the past years, there has been a growing demand for ontological reasoning systems based on languages of the Datalog+/- family, such as Vadalog, for their ability to effectively model a wide range of real-world problems with powerful features such as existential quantification. As the scale and complexity of data analysis tasks continue to grow, the ability to distribute the computational workload across multiple non-communicating processors has become vital for these systems to achieve scalable performance. The joint presence of existential quantification and recursion poses new challenges, currently unsolved by existing distributed systems, which only concentrate on Datalog and are therefore unsuitable for ontological reasoning. When working across multiple processors, generating all the facts to answer a specific reasoning query, avoiding duplication, and guaranteeing termination are non-trivial tasks as infinitely many new symbols and facts can be generated by existential quantification and recursion. In this paper, we address such challenges and introduce the first distributed framework in the Datalog+/- space. We propose the condition of homomorphic decomposability, which identifies sets of Datalog+/- rules with good distribution properties. We put homomorphic decomposability into action with a distributed reasoning algorithm for Warded Datalog+/-, the core of Vadalog. We implement Vadalog Parallel, a distributed reasoner for Vadalog and provide experimental evaluation against state-of-the-art systems. Luigi Bellomarini, Davide Benedetto, Matteo Brandetti, Emanuel Sallinger, Adriano Vlad-Starrabba |
Proc. VLDB Endow. | 4 |
| 2023 | Reasoning over Financial Scenarios with the Vadalog System
Teodoro Baldazzi, Luigi Bellomarini, Emanuel Sallinger |
EDBT | 3 |
| 2023 | SparqLog: A System for Efficient Evaluation of SPARQL 1.1 Queries via DatalogabstractOver the past decade, Knowledge Graphs have received enormous interest both from industry and from academia. Research in this area has been driven, above all, by the Database (DB) community and the Semantic Web (SW) community. However, there still remains a certain divide between approaches coming from these two communities. For instance, while languages such as SQL or Datalog are widely used in the DB area, a different set of languages such as SPARQL and OWL is used in the SW area. Interoperability between such technologies is still a challenge. The goal of this work is to present a uniform and consistent framework meeting important requirements from both, the SW and DB field. Renzo Angles, Georg Gottlob, Aleksandar Pavlovic 0002, Reinhard Pichler, Emanuel Sallinger |
Proc. VLDB Endow. | 5 |
| 2023 | KG-Roar: Interactive Datalog-based Reasoning on Virtual Knowledge GraphsabstractLogic-based Knowledge Graphs (KGs) are gaining momentum in academia and industry thanks to the rise of expressive and efficient languages for Knowledge Representation and Reasoning (KRR). These languages accurately express business rules, through which valuable new knowledge is derived. A versatile and scalable backend reasoner, like Vadalog, a state-of-the-art system for logic-based KGs---based on an extension of Datalog---executes the reasoning. In this demo, we present KG-Roar, a web-based interactive development and navigation environment for logical KGs. The system lets the user augment an input graph database with intensional definitions of new nodes and edges and turn it into a KG, via the metaphor of reasoning widgets---user-defined or off-the-shelf code snippets that capture business definitions in the Vadalog language. Then, the user can seamlessly browse the original and the derived nodes and edges within a "Virtual Knowledge Graph", which is reasoned upon and generated interactively at runtime, thanks to the scalability and responsiveness of Vadalog. KG-Roar is domain-independent but domain aware, as exploration controls are contextually generated based on the intensional definitions. We walk the audience through KG-Roar showcasing the construction of certain business definitions and putting it into action on a real-world financial KG, from our work with the Bank of Italy. Luigi Bellomarini, Marco Benedetti, Andrea Gentili 0005, Davide Magnanimi, Emanuel Sallinger |
Proc. VLDB Endow. | 5 |
| 2022 | Model-Independent Design of Knowledge Graphs - Lessons Learnt From Complex Financial Graphs
Luigi Bellomarini, Andrea Gentili 0005, Eleonora Laurenza, Emanuel Sallinger |
EDBT | 4 |
| 2022 | iTemporal: An Extensible Generator of Temporal BenchmarksabstractDatalogMTL is an extension of the fundamental rule language Datalog with metric temporal operators over rational numbers, whose adoption is soaring within an increasing number of communities (semantic web, databases, stream data processing, temporal logic, knowledge graphs, etc.), which are more and more willing to handle temporal data and deal with temporal database queries as a consequence. Despite the rising research efforts towards new extensions of DatalogMTL, such as the fundamental support for aggregations, and the uprising systems, we still lack a corpus of benchmarks for temporal reasoning. This paper contributes iTemporal, an extensible generator of temporal benchmarks. Our system is able to generate a very broad set of benchmarks, thanks to a white-box configuration mechanism to control and stimulate the theoretical underpinnings of DatalogMTL and its extensions, such as temporal operators in the presence of full recursion and aggregations. We provide a comprehensive presentation of the system as well as an empirical evaluation of the benchmarks within two reference reasoners. Luigi Bellomarini, Markus Nissl, Emanuel Sallinger |
ICDE | 3 |
| 2022 | Rule Learning over Knowledge Graphs with Genetic Logic ProgrammingabstractDeclarative rules such as Prolog and Datalog rules are common formalisms to express expert knowledge and facts. They play an important role in Knowledge Graph (KG) construction and completion. Such rules not only encode the expert background knowledge and the relational patterns among the data, but also infer new knowledge and insights from them. Formalizing rules is often a laborious manual process, while learning them from data automatically can ease this process. Within the rule hypothesis space, current approaches resort to exhaustive search with a number of heuristics and syntactic restrictions on the rule language, which impacts the efficiency and quality of the outcome rules. In this paper, we extend the rule hypothesis space from usual path rules to general Datalog rule space by proposing a novel Genetic Logic Programming algorithm named Evoda. It is an iterative process to learn high-quality rules over large scale KG for a matter of seconds. We have performed experiments over multiple real-world KGs and various evaluation metrics to show its mining capabilities for higher quality rules and more precise predictions. Additionally, we have applied it on the KG completion tasks to illustrate its competitiveness with several state-of-the-art embedding or neural-based models. The experiments demonstrate the feasibility, effectiveness and efficiency of the Evoda algorithm. Lianlong Wu, Emanuel Sallinger, Evgeny Sherkhonov, Sahar Vahdati, Georg Gottlob |
ICDE | 2 |
| 2022 | Reasoning on company takeovers: From tactic to strategy
Luigi Bellomarini, Lorenzo Bencivelli, Claudia Biancotti, Livia Blasi, Francesco Paolo Conteduca, Andrea Gentili 0005, Rosario Laurendi, Davide Magnanimi, Michele Savini Zangrandi, Flavia Tonelli, Stefano Ceri, Davide Benedetto, Markus Nissl, Emanuel Sallinger |
Data Knowl. Eng. | 14 |
| 2022 | Vadalog: A modern architecture for automated reasoning with large knowledge graphs
Luigi Bellomarini, Davide Benedetto, Georg Gottlob, Emanuel Sallinger |
Inf. Syst. | 4 |
| 2022 | Exploiting the Power of Equality-generating Dependencies in Ontological ReasoningabstractEquality-generating dependencies (EGDs) allow to fully exploit the power of existential quantification in ontological reasoning settings modeled via Tuple-Generating Dependencies (TGDs), by enabling value-assignment or forcing the equivalence of fresh symbols. These capabilities are at the core of many common reasoning tasks, including graph traversals, clustering, data matching and data fusion, and many more related real-world scenarios. However, the interplay of TGDs and EGDs is known to lead to undecidability or intractability of query answering in tractable Datalog+/- fragments, like Warded Datalog+/-, for which, in the sole presence of TGDs, query answering is PTIME in data complexity. Restrictions of equality constraints, like separable EGDs, have been studied, but all achieve decidability at the cost of limited expressive power, which makes them unsuitable for the mentioned tasks. This paper introduces the class of "harmless" EGDs, that subsume separable EGDs and allow to model a very broad class of tasks. We contribute a sufficient syntactic condition for testing harmlessness, an undecidable task in general. We argue that in Warded Datalog+/- with harmless EGDs, ontological reasoning is decidable and PTIME. From such theoretical underpinnings, we develop novel chase-based techniques for reasoning with harmless EGDs and present an implementation within the Vadalog system, a state-of-the-art Datalog-based reasoner. We provide full-scale experimental evaluation and comparative analysis. Luigi Bellomarini, Davide Benedetto, Matteo Brandetti, Emanuel Sallinger |
Proc. VLDB Endow. | 4 |
| 2022 | The Space-Efficient Core of VadalogabstractVadalog is a system for performing complex reasoning tasks such as those required in advanced knowledge graphs. The logical core of the underlying Vadalog language is the warded fragment of tuple-generating dependencies (TGDs). This formalism ensures tractable reasoning in data complexity, while a recent analysis focusing on a practical implementation led to the reasoning algorithm around which the Vadalog system is built. A fundamental question that has emerged in the context of Vadalog is whether we can limit the recursion allowed by wardedness in order to obtain a formalism that provides a convenient syntax for expressing useful recursive statements, and at the same time achieves space-efficiency. After analyzing several real-life examples of warded sets of TGDs provided by our industrial partners, as well as recent benchmarks, we observed that recursion is often used in a restricted way: the body of a TGD contains at most one atom whose predicate is mutually recursive with a predicate in the head. We show that this type of recursion, known as piece-wise linear in the Datalog literature, is the answer to our main question. We further show that piece-wise linear recursion alone, without the wardedness condition, is not enough as it leads to undecidability. We also study the relative expressiveness of the query languages based on (piece-wise linear) warded sets of TGDs. Finally, we give preliminary experimental evidence for the practical effect of piece-wise linearity on Vadalog. Gerald Berger, Georg Gottlob, Andreas Pieris, Emanuel Sallinger |
ACM Trans. Database Syst. | 4 |
| 2021 | Pattern-Aware and Noise-Resilient Embedding Models
Mojtaba Nayyeri, Sahar Vahdati, Emanuel Sallinger, Mirza Mohtashim Alam, Hamed Shariat Yazdi, Jens Lehmann 0001 |
ECIR (1) | 3 |
| 2021 | Financial Data Exchange with Statistical Confidentiality: A Reasoning-based Approach
Luigi Bellomarini, Livia Blasi, Rosario Laurendi, Emanuel Sallinger |
EDBT | 4 |
| 2021 | Distributed Company Control in Company Shareholding GraphsabstractThe Company Control Problem is of central importance to banks, financial intermediaries, financial intelligence units, regulatory and supervisory authorities such as the Central Banks. It consists in understanding who takes decisions in a large company network, that is, who controls the majority of votes for each single company. This has an impact on a large number of business areas, with examples including evaluation of creditworthiness, economic analysis of the control dispersion, anti-money laundering, prevention of potentially hostile takeovers, evaluation of risks, and shock propagation.This paper is based on our experience with the Central Bank of Italy and presents an approach to the solution of the company control problem in distributed settings, especially relevant, as large and distributed ownership graphs reflect European-size applications where scalability is paramount.In particular, we formalize the problem as query answering on a large distributed database. We study how independent subqueries can be executed in each partition and the partial results assembled at a master site to produce the answer. We study the formal properties of the problem, that is not easily parallelizable, and then present a method that supports parallelism at best.We present a thorough experimental evaluation of our approach with the Italian company graph of the Bank of Italy and the European Register of Financial Intermediaries and Affiliates as well as many artificial graphs to fully assess scalability. Andrea Gulino, Stefano Ceri, Georg Gottlob, Emanuel Sallinger, Luigi Bellomarini |
ICDE | 4 |
| 2020 | Weaving Enterprise Knowledge Graphs: The Case of Company Ownership GraphsabstractMotivated by our experience in building the Enterprise Knowledge Graph of Italian companies for the Central Bank of Italy, in this paper we present an in-depth case analysis of company ownership graphs, graphs having company ownership as a central concept. In particular, we study and introduce three industrially relevant problems related to such graphs: company control, asset eligibility and detection of personal links. We formally characterize the problems and present Vada-Link, a framework based on state-of-the-art approaches for knowledge representation and reasoning. With our methodology and system, we solve the problems at hand in a scalable, model-independent and generalizable way. We illustrate the favourable architectural properties of Vada-Link and give experimental evaluation of the approach. Paolo Atzeni, Luigi Bellomarini, Michela Iezzi, Emanuel Sallinger, Adriano Vlad-Starrabba |
EDBT | 4 |
| 2020 | Fantastic Knowledge Graph Embeddings and How to Find the Right Space for Them
Mojtaba Nayyeri, Chengjin Xu, Sahar Vahdati, Nadezhda Vassilyeva, Emanuel Sallinger, Hamed Shariat Yazdi, Jens Lehmann 0001 |
ISWC (1) | 5 |
| 2020 | On the Language of Nested Tuple Generating DependenciesabstractDuring the past 15 years, schema mappings have been extensively used in formalizing and studying such critical data interoperability tasks as data exchange and data integration. Much of the work has focused on GLAV mappings, i.e., schema mappings specified by source-to-target tuple-generating dependencies (s-t tgds), and on schema mappings specified by second-order tgds (SO tgds), which constitute the closure of GLAV mappings under composition. In addition, nested GLAV mappings have also been considered, i.e., schema mappings specified by nested tgds, which have expressive power intermediate between s-t tgds and SO tgds. Even though nested GLAV mappings have been used in data exchange systems, such as IBM’s Clio, no systematic investigation of this class of schema mappings has been carried out so far. In this article, we embark on such an investigation by focusing on the basic reasoning tasks, algorithmic problems, and structural properties of nested GLAV mappings. One of our main results is the decidability of the implication problem for nested tgds. We also analyze the structure of the core of universal solutions with respect to nested GLAV mappings and develop useful tools for telling apart SO tgds from nested tgds. By discovering deeper structural properties of nested GLAV mappings, we show that also the following problem is decidable: Given a nested GLAV mapping, is it logically equivalent to a GLAV mapping? Phokion G. Kolaitis, Reinhard Pichler, Emanuel Sallinger, Vadim Savenkov |
ACM Trans. Database Syst. | 3 |
| 2019 | Knowledge Graphs and Enterprise AI: The Promise of an Enabling TechnologyabstractAdopting a mature AI strategy is fundamental for modern knowledge companies to govern the proliferation of smart AI-driven applications and to coordinate them within coherent knowledge workflows. We propose knowledge graphs as the reference technology for the enterprise AI context, i.e., the complex of entities, properties and relationships that shape a business domain and constitute a common backbone for all AI-driven applications. We contribute and discuss principles to design software architectures for AI-driven applications based on knowledge graphs. We focus on the Vadalog system, a successful knowledge graph middleware from the University of Oxford and show knowledge graphs in action in a number of use cases from the financial domain. Luigi Bellomarini, Daniele Fakhoury, Georg Gottlob, Emanuel Sallinger |
ICDE | 4 |
| 2019 | The Space-Efficient Core of VadalogabstractVadalog is a system for performing complex reasoning tasks such as those required in advanced knowledge graphs. The logical core of the underlying Vadalog language is the warded fragment of tuple-generating dependencies (TGDs). This formalism ensures tractable reasoning in data complexity, while a recent analysis focusing on a practical implementation led to the reasoning algorithm around which the Vadalog system is built. A fundamental question that has emerged in the context of Vadalog is the following: can we limit the recursion allowed by wardedness in order to obtain a formalism that provides a convenient syntax for expressing useful recursive statements, and at the same time achieves space-efficiency? After analyzing several real-life examples of warded sets of TGDs provided by our industrial partners, as well as recent benchmarks, we observed that recursion is often used in a restricted way: the body of a TGD contains at most one atom whose predicate is mutually recursive with a predicate in the head. We show that this type of recursion, known as piece-wise linear in the Datalog literature, is the answer to our main question. We further show that piece-wise linear recursion alone, without the wardedness condition, is not enough as it leads to the undecidability of reasoning. We finally study the relative expressiveness of the query languages based on (piece-wise linear) warded sets of TGDs. Gerald Berger, Georg Gottlob, Andreas Pieris, Emanuel Sallinger |
PODS | 4 |
| 2018 | Data Science with Vadalog: Bridging Machine Learning and Reasoning
Luigi Bellomarini, Ruslan R. Fayzrakhmanov, Georg Gottlob, Andrey Kravchenko, Eleonora Laurenza, Yavor Nenov, Stéphane Reissfelder, Emanuel Sallinger, Evgeny Sherkhonov, Lianlong Wu |
MEDI | 8 |
| 2018 | Browserless Web Data Extraction: Challenges and OpportunitiesabstractMost modern web scrapers use an embedded browser to render web pages and to simulate user actions. Such scrapers (or wrappers) are therefore expensive to execute, in terms of time and network traffic. In contrast, it is magnitudes more resource-efficient to use a "browserless" wrapper which directly accesses a web server through HTTP requests, and takes the desired data directly from the raw replies. However, creating and maintaining browserless wrappers of high precision requires specialists, and is prohibitively labor-intensive at scale. In this paper, we demonstrate the principal feasibility of automatically translating browser-based wrappers into "browserless" wrappers. We present the first algorithm and system performing such an automated translation on suitably restricted types of web sites. This system works in the vast majority of test cases and produces very fast and extremely resource-efficient wrappers. We discuss research challenges for extending our approach to a general method applicable to a yet larger number of cases. Ruslan R. Fayzrakhmanov, Emanuel Sallinger, Ben Spencer, Tim Furche, Georg Gottlob |
WWW | 2 |
| 2018 | The Vadalog System: Datalog-based Reasoning for Knowledge GraphsabstractOver the past years, there has been a resurgence of Datalog-based systems in the database community as well as in industry. In this context, it has been recognized that to handle the complex knowledge-based scenarios encountered today, such as reasoning over large knowledge graphs, Datalog has to be extended with features such as existential quantification. Yet, Datalog-based reasoning in the presence of existential quantification is in general undecidable. Many efforts have been made to define decidable fragments. Warded Datalog+/- is a very promising one, as it captures PTIME complexity while allowing ontological reasoning. Yet so far, no implementation of Warded Datalog+/- was available. In this paper we present the Vadalog system, a Datalog-based system for performing complex logic reasoning tasks, such as those required in advanced knowledge graphs. The Vadalog system is Oxford's contribution to the VADA research programme, a joint effort of the universities of Oxford, Manchester and Edinburgh and around 20 industrial partners. As the main contribution of this paper, we illustrate the first implementation of Warded Datalog+/-, a high-performance Datalog+/- system utilizing an aggressive termination control strategy. We also provide a comprehensive experimental evaluation. Luigi Bellomarini, Emanuel Sallinger, Georg Gottlob |
Proc. VLDB Endow. | 2 |
| 2017 | The VADA Architecture for Cost-Effective Data WranglingabstractData wrangling, the multi-faceted process by which the data required by an application is identified, extracted, cleaned and integrated, is often cumbersome and labor intensive. In this paper, we present an architecture that supports a complete data wrangling lifecycle, orchestrates components dynamically, builds on automation wherever possible, is informed by whatever data is available, refines automatically produced results in the light of feedback, takes into account the user's priorities, and supports data scientists with diverse skill sets. The architecture is demonstrated in practice for wrangling property sales and open government data. Nikolaos Konstantinou 0001, Martin Koehler, Edward Abel, Cristina Civili, Bernd Neumayr, Emanuel Sallinger, Alvaro A. A. Fernandes, Georg Gottlob, John A. Keane, Leonid Libkin, Norman W. Paton |
SIGMOD Conference | 6 |
| 2016 | Complexity of Repair Checking and Consistent Query AnsweringabstractInconsistent databases (i.e., databases violating some given set of integrity constraints) may arise in many applications such as, for instance, data integration. Hence, the handling of inconsistent data has evolved as an active field of research. In this paper, we consider two fundamental problems in this context: Repair Checking (RC) and Consistent Query Answering (CQA). So far, these problems have been mainly studied from the point of view of data complexity (where all parts of the input except for the database are considered as fixed). While for some kinds of integrity constraints, also combined complexity (where all parts of the input are allowed to vary) has been considered, for several other kinds of integrity constraints, combined complexity has been left unexplored. Moreover, a more detailed analysis (keeping other parts of the input fixed - e.g., the constraints only) is completely missing. The goal of our work is a thorough analysis of the complexity of the RC and CQA problems. Our contribution is a complete picture of the complexity of these problems for a wide range of integrity constraints. Our analysis thus allows us to get a better understanding of the true sources of complexity. Sebastian Arming, Reinhard Pichler, Emanuel Sallinger |
ICDT | 3 |
| 2016 | Limits of Schema MappingsabstractSchema mappings have been extensively studied in the context of data exchange and data integration, where they have turned out to be the right level of abstraction for formalizing data inter-operability tasks. Up to now and for the most part, schema mappings have been studied as static objects, in the sense that each time the focus has been on a single schema mapping of interest or, in the case of composition, on a pair of schema mappings of interest. In this paper, we adopt a dynamic viewpoint and embark on a study of sequences of schema mappings and of the limiting behavior of such sequences. To this effect, we first introduce a natural notion of distance on sets of finite target instances that expresses how "close" two sets of target instances are as regards the certain answers of conjunctive queries on these sets. Using this notion of distance, we investigate pointwise limits and uniform limits of sequences of schema mappings, as well as the companion notions of pointwise Cauchy and uniformly Cauchy sequences of schema mappings. We obtain a number of results about the limits of sequences of GAV schema mappings and the limits of sequences of LAV schema mappings that reveal striking differences between these two classes of schema mappings. We also consider the completion of the metric space of sets of target instances and obtain concrete representations of limits of sequences of schema mappings in terms of generalized schema mappings, i.e., schema mappings with infinite target instances as solutions to (finite) source instances. Phokion G. Kolaitis, Reinhard Pichler, Emanuel Sallinger, Vadim Savenkov |
ICDT | 3 |
| 2015 | Function Symbols in Tuple-Generating Dependencies: Expressive Power and ComputabilityabstractTuple-generating dependencies -- for short tgds -- have been a staple of database research throughout most of its history. Yet one of the central aspects of tgds, namely the role of existential quantifiers, has not seen much investigation so far. When studying dependencies, existential quantifiers and -- in their Skolemized form -- function symbols are often viewed as two ways to express the same concept. But in fact, tgds are quite restrictive in the way that functional terms can occur. Georg Gottlob, Reinhard Pichler, Emanuel Sallinger |
PODS | 3 |
| 2015 | On the undecidability of the equivalence of second-order tuple generating dependencies
Ingo Feinerer, Reinhard Pichler, Emanuel Sallinger, Vadim Savenkov |
Inf. Syst. | 3 |
| 2014 | Nested dependencies: structure and reasoningabstractDuring the past decade, schema mappings have been extensively used in formalizing and studying such critical data interoperability tasks as data exchange and data integration. Much of the work has focused on GLAV mappings, i.e., schema mappings specified by source-to-target tuple-generating dependencies (s-t tgds), and on schema mappings specified by second-order tgds (SO tgds), which constitute the closure of GLAV mappings under composition. In addition, nested GLAV mappings have also been considered, i.e., schema mappings specified by nested tgds, which have expressive power intermediate between s-t tgds and SO tgds. Even though nested GLAV mappings have been used in data exchange systems, such as IBM's Clio, no systematic investigation of this class of schema mappings has been carried out so far. In this paper, we embark on such an investigation by focusing on the basic reasoning tasks, algorithmic problems, and structural properties of nested GLAV mappings. One of our main results is the decidability of the implication problem for nested tgds. We also analyze the structure of the core of universal solutions with respect to nested GLAV mappings and develop useful tools for telling apart SO tgds from nested tgds. By discovering deeper structural properties of nested GLAV mappings, we show that also the following problem is decidable: given a nested GLAV mapping, is it logically equivalent to a GLAV mapping? Phokion G. Kolaitis, Reinhard Pichler, Emanuel Sallinger, Vadim Savenkov |
PODS | 3 |
| 2011 | Relaxed notions of schema mapping equivalence revisitedabstractRecently, two relaxed notions of equivalence of schema mappings have been introduced, which provide more potential of optimizing schema mappings than logical equivalence: data exchange (DE) equivalence and conjunctive query (CQ) equivalence. In this work, we systematically investigate these notions of equivalence for mappings consisting of s-t tgds and target egds and/or target tgds. We prove that both CQ- and DE-equivalence are undecidable and so are some important optimization tasks (like detecting if some dependency is redundant). However, we also identify an important difference between the two notions of equivalence: CQ-equivalence remains undecidable even if the schema mappings consist of s-t tgds and target dependencies in the form of key dependencies only. In contrast, DE-equivalence is decidable for schema mappings with s-t tgds and target dependencies in the form of functional and inclusion dependencies with terminating chase property. Reinhard Pichler, Emanuel Sallinger, Vadim Savenkov |
ICDT | 2 |