Michael Schmidt 0002

dblp:156/1960-2 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
3since 2021 · last 2024
0009-0002-3292-0349ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2024 Statement Graphs: Unifying the Graph Data Model Landscape
Ewout Gelling, George Fletcher 0001, Michael Schmidt 0002
DASFAA (7)3
2023 PG-Schema: Schemas for Property Graphs
abstract
Property graphs have reached a high level of maturity, witnessed by multiple robust graph database systems as well as the ongoing ISO standardization effort aiming at creating a new standard Graph Query Language (GQL). Yet, despite documented demand, schema support is limited both in existing systems and in the first version of the GQL Standard. It is anticipated that the second version of the GQL Standard will include a rich DDL. Aiming to inspire the development of GQL and enhance the capabilities of graph database systems, we propose PG-Schema, a simple yet powerful formalism for specifying property graph schemas. It features PG-Schema with flexible type definitions supporting multi-inheritance, as well as expressive constraints based on the recently proposed PG-Keys formalism. We provide the formal syntax and semantics of PG-Schema, which meet principled design requirements grounded in contemporary property graph management scenarios, and offer a detailed comparison of its features with those of existing schema languages and graph database systems.
Renzo Angles, Angela Bonifati, Stefania Dumbrava, George Fletcher 0001, Alastair Green, Jan Hidders, Leonid Libkin, Victor Marsault, Wim Martens, Filip Murlak, Stefan Plantikow, Ognjen Savkovic, Michael Schmidt 0002, Juan F. Sequeda, Slawomir Staworko, Dominik Tomaszuk, Hannes Voigt, Domagoj Vrgoc, Mingxi Wu, Dusan Zivkovic
Proc. ACM Manag. Data14
2021 PG-Keys: Keys for Property Graphs
abstract
We report on a community effort between industry and academia to shape the future of property graph constraints. The standardization for a property graph query language is currently underway through the ISO Graph Query Language (GQL) project. Our position is that this project should pay close attention to schemas and constraints, and should focus next on key constraints. The main purposes of keys are enforcing data integrity and allowing the referencing and identifying of objects. Motivated by use cases from our industry partners, we argue that key constraints should be able to have different modes, which are combinations of basic restriction that require the key to be exclusive, mandatory, and singleton. Moreover, keys should be applicable to nodes, edges, and properties since these all can represent valid real-life entities. Our result is PG-Keys, a flexible and powerful framework for defining key constraints, which fulfills the above goals. PG-Keys is a design by the Linked Data Benchmark Council's Property Graph Schema Working Group, consisting of members from industry, academia, and ISO GQL standards group, intending to bring the best of all worlds to property graph practitioners. PG-Keys aims to guide the evolution of the standardization efforts towards making systems more useful, powerful, and expressive.
Renzo Angles, Angela Bonifati, Stefania Dumbrava, George Fletcher 0001, Keith W. Hare, Jan Hidders, Victor E. Lee, Leonid Libkin, Wim Martens, Filip Murlak, Josh Perryman, Ognjen Savkovic, Michael Schmidt 0002, Juan F. Sequeda, Slawomir Staworko, Dominik Tomaszuk
SIGMOD Conference14
2014 The Social Navigator: A Personalized Learning Platform for Social Media Education
Michael Schmidt 0002, Uta Schwertel, Christina Di Valentin, Andreas Emrich, Dirk Werth
EC-TEL1
2014 Assessing and Training Social Media Skills in Vocational Education Supported by TEL Instruments
Uta Schwertel, Yvonne Kammerer, Clara Oloff, Peter Gerjets, Michael Schmidt 0002
EC-TEL5
2013 Semantic query optimization in the presence of types
Michael Meier 0002, Michael Schmidt 0002, Fang Wei-Kleiner, Georg Lausen
J. Comput. Syst. Sci.2
2011 FedX: A Federation Layer for Distributed Query Processing on Linked Open Data
Andreas Schwarte, Peter Haase 0001, Katja Hose, Ralf Schenkel, Michael Schmidt 0002
ESWC (2)5
2011 FedBench: A Benchmark Suite for Federated Semantic Data Query Processing
Michael Schmidt 0002, Olaf Görlitz, Peter Haase 0001, Günter Ladwig, Andreas Schwarte, Thanh Tran 0001
ISWC (1)1
2011 FedX: Optimization Techniques for Federated Query Processing on Linked Data
Andreas Schwarte, Peter Haase 0001, Katja Hose, Ralf Schenkel, Michael Schmidt 0002
ISWC (1)5
2010 Foundations of SPARQL query optimization
abstract
We study fundamental aspects related to the efficient processing of the SPARQL query language for RDF, proposed by the W3C to encode machine-readable information in the Semantic Web. Our key contributions are (i) a complete complexity analysis for all operator fragments of the SPARQL query language, which -- as a central result -- shows that the SPARQL operator Optional alone is responsible for the PSpace-completeness of the evaluation problem, (ii) a study of equivalences over SPARQL algebra, including both rewriting rules like filter and projection pushing that are well-known from relational algebra optimization as well as SPARQL-specific rewriting schemes, and (iii) an approach to the semantic optimization of SPARQL queries, built on top of the classical chase algorithm. While studied in the context of a theoretically motivated set semantics, almost all results carry over to the official, bag-based semantics and therefore are of immediate practical relevance.
Michael Schmidt 0002, Michael Meier 0002, Georg Lausen
ICDT1
2010 Semantic query optimization in the presence of types
abstract
Both semantic and type-based query optimization rely on the idea that queries often exhibit non-trivial rewritings if the state space of the database is restricted. Despite their close connection, these two problems to date have always been studied separately. We present a unifying, logic-based framework for query optimization in the presence of data dependencies and type information. It builds upon the classical chase algorithm and extends existing query minimization techniques to considerably larger classes of queries and dependencies. In particular, our setting requires chasing conjunctive queries (possibly with union and negation) in the presence of dependencies containing negation and disjunction. We study the applicability of the chase in this setting, develop novel conditions that guarantee its termination, identify fragments for which minimal query computation is always possible (w.r.t. a generic cost function), and investigate the complexity of related decision problems.
Michael Meier 0002, Michael Schmidt 0002, Fang Wei-Kleiner, Georg Lausen
PODS2
2010 Semantic Technologies for Enterprise Cloud Management
Peter Haase 0001, Tobias Mathäß, Michael Schmidt 0002, Andreas Eberhart, Ulrich Walther
ISWC (2)3
2009 SP^2Bench: A SPARQL Performance Benchmark
abstract
Recently, the SPARQL query language for RDF has reached the W3C recommendation status. In response to this emerging standard, the database community is currently exploring efficient storage techniques for RDF data and evaluation strategies for SPARQL queries. A meaningful analysis and comparison of these approaches necessitates a comprehensive and universal benchmark platform. To this end, we have developed SP2Bench, a publicly available, language-specific SPARQL performance benchmark. SP2Bench is settled in the DBLP scenario and comprises both a data generator for creating arbitrarily large DBLP-like documents and a set of carefully designed benchmark queries. The generated documents mirror key characteristics and social-world distributions encountered in the original DBLP data set, while the queries implement meaningful requests on top of this data, covering a variety of SPARQL operator constellationsand RDF access patterns. As a proof of concept, we apply SP2Bench to existing engines and discuss their strengths and weaknesses that follow immediately from the benchmark results.
Michael Schmidt 0002, Thomas Schallhorn, Georg Lausen, Christoph Pinkel
ICDE1
2009 On Chase Termination Beyond Stratification
abstract
We study the termination problem of the chase algorithm, a central tool in various database problems such as the constraint implication problem, Conjunctive Query optimization, rewriting queries using views, data exchange, and data integration. The basic idea of the chase is, given a database instance and a set of constraints as input, to fix constraint violations in the database instance. It is well-known that, for an arbitrary set of constraints, the chase does not necessarily terminate (in general, it is even undecidable if it does or not). Addressing this issue, we review the limitations of existing sufficient termination conditions for the chase and develop new techniques that allow us to establish weaker sufficient conditions. In particular, we introduce two novel termination conditions called safety and inductive restriction , and use them to define the so-called T-hierarchy of termination conditions. We then study the interrelations of our termination conditions with previous conditions and the complexity of checking our conditions. This analysis leads to an algorithm that checks membership in a level of the T -hierarchy and accounts for the complexity of termination conditions. As another contribution, we study the problem of data-dependent chase termination and present sufficient termination conditions w.r.t. fixed instances. They might guarantee termination although the chase does not terminate in the general case. As an application of our techniques beyond those already mentioned, we transfer our results into the field of query answering over knowledge bases where the chase on the underlying database may not terminate, making existing algorithms applicable to broader classes of constraints.
Michael Meier 0002, Michael Schmidt 0002, Georg Lausen
Proc. VLDB Endow.2
2008 SPARQLing constraints for RDF
abstract
The goal of the Semantic Web is to support semantic interoperability between applications exchanging data on the web. The idea heavily relies on data being made available in machine readable format, using semantic markup languages. In this regard, the W3C has standardized RDF as the basic markup language for the Semantic Web. In contrast to relational databases, where data relationships are implicitly given by schema information as well as primary and foreign key constraints, relationships in semantic markup languages are made explicit. When mapping relational data into RDF, it is desirable to maintain the information implied by the origin constraints. As an improvement over existing approaches, our scheme allows for translating conventional databases into RDF without losing general constraints and vital key information. As much as in the relational model, those information are indispensable for data consistency and, as shown by example, can serve as a basis for semantic query optimization. We underline the practicability of our approach by showing that SPARQL, the most popular query language for RDF, can be used as a constraint language, akin to SQL in the relational context. As a theoretical contribution, we also discuss satisfiability for interesting classes of constraints and combinations thereof. 1.
Georg Lausen, Michael Meier 0002, Michael Schmidt 0002
EDBT3
2008 XML Prefiltering as a String Matching Problem
abstract
We propose a new technique for the efficient search and navigation in XML documents and streams. This technique takes string matching algorithms designed for efficient keyword search in flat strings into the second dimension, to navigate in tree structured data. We consider the important XML data management task of prefiltering XML documents (also called XML projection) as an application for our approach. Different from existing prefiltering schemes, we usually process only fractions of the input and get by with very economical consumption of both main memory and processing time. Our experiments reveal that, already on low-complexity problems such as XPath filtering, in-memory query engines can experience speed-ups by two orders of magnitude.
Christoph Koch 0001, Stefanie Scherzinger, Michael Schmidt 0002
ICDE3
2008 An Experimental Comparison of RDF Data Management Approaches in a SPARQL Benchmark Scenario
Michael Schmidt 0002, Thomas Schallhorn, Norbert Küchlin, Georg Lausen, Christoph Pinkel
ISWC1
2007 Combined Static and Dynamic Analysis for Effective Buffer Minimization in Streaming XQuery Evaluation
abstract
Effective buffer management is crucial for efficient in-memory and streaming XQuery processing. We propose a buffer management scheme which combines static and dynamic analysis to keep main memory consumption low. Our approach relies on a technique that we call active garbage collection and which actively purges buffers at runtime based on the current status of query evaluation. We have built a prototype system for a practical fragment of XQuery which employs our buffer management scheme. The experimental results demonstrate the significant impact of combined static and dynamic analysis on reducing main memory consumption and running time.
Michael Schmidt 0002, Stefanie Scherzinger, Christoph Koch 0001
ICDE1
2007 The GCX System: Dynamic Buffer Minimization in Streaming XQuery Evaluation
Christoph Koch 0001, Stefanie Scherzinger, Michael Schmidt 0002
VLDB3