Jérôme Siméon

dblp:15/4061 · DBLP profile ↗
← Back
40ranked-venue papers
2as first author
1since 2021 · last 2022
0000-0002-8622-9716ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 31 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Theory of computation · 2Artificial intelligence and machine learning · 1Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
25 papers
Data models and query languages · 42% Query processing and optimization · 41% Transaction processing and concurrency control · 5%
Software engineering, system software, and programming languages
7 papers
Program verification · 35% Compilers and program optimization · 30% Programming languages and type systems · 24%

Topics — the 30 heaviest of 53, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query compilation
0.932022
Translating canonical SQL to imperative code in Coq · Proc. ACM Program. Lang. 2022
Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler · SIGMOD Conference 2017
A Complete and Efficient Algebraic Compiler for XQuery · ICDE 2006
Data models and query languages › SQL
SQL semantics
0.612022
Translating canonical SQL to imperative code in Coq · Proc. ACM Program. Lang. 2022
Program verification
mechanized verification
0.612022
Translating canonical SQL to imperative code in Coq · Proc. ACM Program. Lang. 2022
Compilers and program optimization
verified compilation
0.612022
Translating canonical SQL to imperative code in Coq · Proc. ACM Program. Lang. 2022
Data models and query languages › relational algebra
nested relational algebra
0.522022
Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler · SIGMOD Conference 2017
Translating canonical SQL to imperative code in Coq · Proc. ACM Program. Lang. 2022
Query processing and optimization
query optimization
0.342017
Q*cert: A Platform for Implementing and Verifying Query Compilers · SIGMOD Conference 2017
Commutativity analysis for XML updates · ACM Trans. Database Syst. 2008
XQuery Streaming à la Carte · ICDE 2007
Query processing and optimization › query compilation
query compiler
0.312017
Q*cert: A Platform for Implementing and Verifying Query Compilers · SIGMOD Conference 2017
Data models and query languages
NoSQL database
0.212016
Virtual lightweight snapshots for consistent analytics in NoSQL stores · ICDE 2016
Transaction processing and concurrency control › isolation levels
snapshot isolation
0.212016
Virtual lightweight snapshots for consistent analytics in NoSQL stores · ICDE 2016
Query processing and optimization
XML query processing
0.242007
Put a Tree Pattern in Your Algebra · ICDE 2007
XQuery Streaming à la Carte · ICDE 2007
Projecting XML Documents · VLDB 2003
Programming languages and type systems
type systems
0.222013
Static and dynamic semantics of NoSQL languages · POPL 2013
The essence of XML · POPL 2003
Data models and query languages › XML query languages
XQuery
0.242007
Highly distributed XQuery with DXQ · SIGMOD Conference 2007
XQuery at your web service · WWW 2004
Yoo-Hoo! Building a Presence Service with XQuery and WSDL · SIGMOD Conference 2004
Data models and query languages
semistructured data
0.212013
Static and dynamic semantics of NoSQL languages · POPL 2013
Programming languages and type systems
type inference
0.212013
Static and dynamic semantics of NoSQL languages · POPL 2013
Data models and query languages
XML query languages
0.132008
Commutativity analysis for XML updates · ACM Trans. Database Syst. 2008
XQuery at your web service · WWW 2004
On Wrapping Query Languages and Efficient XML Integration · SIGMOD Conference 2000
Data models and query languages › XML data management
XML data model
0.132003
The Yin/Yang Web: A Unified Model for XML Syntax and RDF Semantics · IEEE Trans. Knowl. Data Eng. 2003
The Yin/Yang web: XML syntax and RDF semantics · WWW 2002
A unified constraint model for XML · WWW 2001
Query processing and optimization
query rewriting
0.112008
XML query optimization in the presence of side effects · SIGMOD Conference 2008
Query processing and optimization › XML query processing
XML query optimization
0.112008
XML query optimization in the presence of side effects · SIGMOD Conference 2008
Query processing and optimization › query optimization
cost-based optimization
0.122003
Bridging the XML Relational Divide with LegoDB · ICDE 2003
From XML Schema to Relations: A Cost-Based Approach to XML Storage · ICDE 2002
Indexing and storage engines
XML storage
0.122003
Bridging the XML Relational Divide with LegoDB · ICDE 2003
From XML Schema to Relations: A Cost-Based Approach to XML Storage · ICDE 2002
Data integration and cleaning › schema mapping
XML-to-relational mapping
0.122003
Bridging the XML Relational Divide with LegoDB · ICDE 2003
From XML Schema to Relations: A Cost-Based Approach to XML Storage · ICDE 2002
Distributed and cloud data management
distributed query processing
0.112007
Highly distributed XQuery with DXQ · SIGMOD Conference 2007
Data stream processing
XML stream processing
0.112007
XQuery Streaming à la Carte · ICDE 2007
Query processing and optimization
join processing
0.112006
A Complete and Efficient Algebraic Compiler for XQuery · ICDE 2006
Query processing and optimization › query optimization › nested query optimization
query unnesting
0.112006
A Complete and Efficient Algebraic Compiler for XQuery · ICDE 2006
Services computing and microservices › web service interfaces
WSDL
0.122004
Yoo-Hoo! Building a Presence Service with XQuery and WSDL · SIGMOD Conference 2004
XQuery at your web service · WWW 2004
Services computing and microservices
service composition
0.012004
XQuery at your web service · WWW 2004
Services computing and microservices
web services
0.012004
XQuery at your web service · WWW 2004
Services computing and microservices › service composition
web service composition
0.012004
Yoo-Hoo! Building a Presence Service with XQuery and WSDL · SIGMOD Conference 2004
Query processing and optimization › XML query processing
XML projection
0.012003
Projecting XML Documents · VLDB 2003

Methods — techniques the papers use, named apart from their topics

nested relational calculus · 1.4coq · 1.1type inference · 0.3calculus · 0.3nested relational algebra · 0.3coq proof assistant · 0.3XQuery · 0.3versioning · 0.2snapshot isolation · 0.2static analysis · 0.2functional programming · 0.1compilation rules · 0.1WSDL binding · 0.0WSDL · 0.0formal semantics · 0.0model-theoretic semantics · 0.0XPath · 0.0
YearPublicationVenuePosition
2022 Translating canonical SQL to imperative code in Coq
abstract
SQL is by far the most widely used and implemented query language. Yet, on some key features, such as correlated queries and NULL value semantics, many implementations diverge or contain bugs. We leverage recent advances in the formalization of SQL and query compilers to develop DBCert, the first mechanically verified compiler from SQL queries written in a canonical form to imperative code. Building DBCert required several new contributions which are described in this paper. First, we specify and mechanize a complete translation from SQL to the Nested Relational Algebra which can be used for query optimization. Second, we define Imp, a small imperative language sufficient to express SQL and which can target several execution languages including JavaScript. Finally, we develop a mechanized translation from the nested relational algebra to Imp, using the nested relational calculus as an intermediate step.
Véronique Benzaken, Evelyne Contejean, Mohammed Houssem Hachmaoui, Chantal Keller, Louis Mandel, Avraham Shinnar, Jérôme Siméon
Proc. ACM Program. Lang.7
2017 Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler
abstract
Algebras based on combinators, i.e., variable-free, have been proposed as a better representation for query compilation and optimization. A key benefit of combinators is that they avoid the need to handle variable shadowing or accidental capture during rewrites. This simplifies both the optimizer specification and its correctness analysis, but the environment from the source language has to be reified as records, which can lead to more complex query plans.
Joshua S. Auerbach, Martin Hirzel, Louis Mandel, Avraham Shinnar, Jérôme Siméon
SIGMOD Conference5
2017 Q*cert: A Platform for Implementing and Verifying Query Compilers
abstract
We present Q*cert, a platform for the specification, verification, and implementation of query compilers written using the Coq proof assistant. The Q*cert platform is open source and includes some support for SQL and OQL, and for code generation to Spark and Cloudant. It internally relies on familiar database intermediate representations, notably the nested relational algebra and calculus and a novel extension of the nested relational algebra that eases the handling of environments. The platform also comes with simple but functional and extensible query optimizers.
Joshua S. Auerbach, Martin Hirzel, Louis Mandel, Avraham Shinnar, Jérôme Siméon
SIGMOD Conference5
2017 Prototyping a query compiler using Coq (experience report)
abstract
Designing and prototyping new features is important in many industrial projects. Functional programming and formal verification tools can prove valuable for that purpose, but lead to challenges when integrating with existing product code or when planning technology transfer. This article reports on our experience using the Coq proof assistant as a prototyping environment for building a query compiler intended for use in IBM's ODM Insights product. We discuss the pros and cons of using Coq for this purpose and describe our methodology for porting the compiler to Java, as required for product integration.
Joshua S. Auerbach, Martin Hirzel, Louis Mandel, Avraham Shinnar, Jérôme Siméon
Proc. ACM Program. Lang.5
2016 Virtual lightweight snapshots for consistent analytics in NoSQL stores
abstract
Increasingly, applications that deal with big data need to run analytics concurrently with updates. But bridging the gap between big and fast data is challenging: most of these applications require analytics' results that are fresh and consistent, but without impacting system latency and throughput. We propose virtual lightweight snapshots (VLS), a mechanism that enables consistent analytics without blocking incoming updates in NoSQL stores. VLS requires neither native support for database versioning nor a transaction manager. Besides, it is storage-efficient, keeping additional versions of records only when needed to guarantee consistency, and sharing versions across multiple concurrent snapshots. We describe an implementation of VLS in MongoDB and present a detailed experimental evaluation which shows that it supports consistency for analytics with small impact on query evaluation time, update throughput, and latency.
Fernando Seabra Chirigati, Jérôme Siméon, Martin Hirzel, Juliana Freire
ICDE2
2015 A Pattern Calculus for Rule Languages: Expressiveness, Compilation, and Mechanization
abstract
This paper introduces a core calculus for pattern-matching in production rule languages: the Calculus for Aggregating Matching Patterns (CAMP). CAMP is expressive enough to capture modern rule languages such as JRules, including extensions for aggregation. We show how CAMP can be compiled into a nested-relational algebra (NRA), with only minimal extension. This paves the way for applying relational techniques to running rules over large stores. Furthermore, we show that NRA can also be compiled back to CAMP, using named nested-relational calculus (NNRC) as an intermediate step. We mechanize proofs of correctness, program size preservation, and type preservation of the translations using modern theorem-proving techniques. A corollary of the type preservation is that polymorphic type inference for both CAMP and NRA is NP-complete. CAMP and its correspondence to NRA provide the foundations for efficient implementations of rules languages using databases technologies.
Avraham Shinnar, Jérôme Siméon, Martin Hirzel
ECOOP2
2014 Event Processing over a Distributed JSON Store: Design and Performance
Miki Enoki, Jérôme Siméon, Hiroshi Horii, Martin Hirzel
WISE (2)2
2013 Static and dynamic semantics of NoSQL languages
abstract
We present a calculus for processing semistructured data that spans differences of application area among several novel query languages, broadly categorized as "NoSQL". This calculus lets users define their own operators, capturing a wider range of data processing capabilities, whilst providing a typing precision so far typical only of primitive hard-coded operators. The type inference algorithm is based on semantic type checking, resulting in type information that is both precise, and flexible enough to handle structured and semistructured data. We illustrate the use of this calculus by encoding a large fragment of Jaql, including operations and iterators over JSON, embedded SQL expressions, and co-grouping, and show how the encoding directly yields a typing discipline for Jaql as it is, namely without the addition of any type definition or type annotation in the code.
Véronique Benzaken, Giuseppe Castagna, Kim Nguyen 0001, Jérôme Siméon
POPL4
2008 XML query optimization in the presence of side effects
abstract
The emergence of database languages with side effects, notably for XML, raises significant challenges for database compilers and optimizers. In this paper, we extend an algebra for the W3C XML query language with operations that allow data to be immediately updated. We study the impact of that extension on logical optimization, join detection, and pipelining. The main result of this work is to show that, with proper care, a number of important optimizations based on nested relational algebras remain applicable in the presence of side effects. Our approach relies on an analysis of the conditions that must be checked in order for algebraic rewritings to hold. An implementation and experimental results demonstrate the effectiveness of the approach. 1.
Giorgio Ghelli, Nicola Onose, Kristoffer Høgsbro Rose, Jérôme Siméon
SIGMOD Conference4
2008 Commutativity analysis for XML updates
abstract
An effective approach to support XML updates is to use XQuery extended with update operations. This approach results in very expressive languages which are convenient for users but are difficult to optimize or reason about. A crucial question underlying many static analysis problems for such languages, from optimization to view maintenance, is whether two expressions commute. Unfortunately, commutativity is undecidable for most existing XML update languages. In this article, we propose a conservative analysis for an expressive XML update language that can be used to determine commutativity. The approach relies on a form of path analysis that computes upper bounds for the nodes that are accessed or modified in a given expression. Our main result is a theorem that can be used to identify commuting expressions. We illustrate how the technique applies to concrete examples of query optimization in the presence of updates.
Giorgio Ghelli, Kristoffer Høgsbro Rose, Jérôme Siméon
ACM Trans. Database Syst.3
2007 XQuery Streaming à la Carte
abstract
Existing work on XML query evaluation has either focused on algebraic optimization techniques suitable for XML databases, or on algorithms to efficiently process XML messages represented as a stream of parsing events. In practice, complex applications often must handle both. In this paper, we develop a physical algebra that combines streaming operators with other standard relational and XML operators. Our physical model includes marked XML streams, which permit efficient XPath evaluation, but can only be consumed once. This constraint restricts the use of streaming operators to fragments of a query plan that only access data using depth-first traversal. We develop static analysis techniques to decide which fragment of a plan can be streamed. Our experiments demonstrate the benefits of blending streaming with other evaluation techniques.
Mary F. Fernández, Philippe Michiels, Jérôme Siméon, Michael Stark 0003
ICDE3
2007 Put a Tree Pattern in Your Algebra
abstract
To address the needs of data intensive XML applications, a number of efficient tree pattern algorithms have been proposed. Still, most XQuery compilers do not support those algorithms. This is due in part to the lack of support for tree patterns in XML algebras, but also because deciding which part of a query plan should be evaluated as a tree pattern is a hard problem. In this paper, we extend a tuple algebra for XQuery with a tree pattern operator, and present rewrit-ings suitable to introduce that operator in query plans. We demonstrate the robustness of the proposed rewritings under syntactic variations commonly found in queries. The proposed tree pattern operator can be implemented using popular algorithms such as Twig joins and Staircase joins. Our experiments yield useful information to decide which algorithm should be used in a given plan.
Philippe Michiels, George A. Mihaila, Jérôme Siméon
ICDE3
2007 Commutativity Analysis in XML Update Languages
Giorgio Ghelli, Kristoffer Høgsbro Rose, Jérôme Siméon
ICDT3
2007 Highly distributed XQuery with DXQ
abstract
Many modern applications, from Grid computing to RSS handling, need to support data processing in a distributed environment. Currently, most such applications are implemented using a general purpose programming language, which can be expensive to maintain, hard to configure and modify, and require hand optimization of the distributed data processing operations. We present Distributed XQuery (DXQ), a simple, yet powerful, extension of XQuery to support distributed applications. This extension includes the ability to deploy networks of XQuery servers, to remotely invoke XQuery programs on those servers, and to ship code between servers. Our demonstration presents two applications implemented in DXQ: the resolution algorithm of DNS, the Domain Name System, and the Narada overlay-network protocol. We show that our system can flexibly accommodate different patterns of distributed computation and present some simple but essential distributed optimizations.
Mary F. Fernández, Trevor Jim, Kristi Morton, Nicola Onose, Jérôme Siméon
SIGMOD Conference5
2006 A Complete and Efficient Algebraic Compiler for XQuery
abstract
As XQuery nears standardization, more sophisticated XQuery applications are emerging, which often exploit the entire language and are applied to non-trivial XML sources. We propose an algebra and optimization techniques that are suitable for building an XQuery compiler that is complete, correct, and efficient. We describe the compilation rules for the complete language into that algebra and present novel optimization techniques that address the needs of complex queries. These techniques include new query unnesting rewritings and specialized join algorithms that account for XQuery’s complex predicate semantics. The algebra and optimizations are implemented in the Galax XQuery engine, and yield execution plans that are up to three orders of magnitude faster than earlier versions of Galax.
Christopher Ré, Jérôme Siméon, Mary F. Fernández
ICDE2
2005 Optimizing Sorting and Duplicate Elimination in XQuery Path Expressions
Mary F. Fernández, Jan Hidders, Philippe Michiels, Jérôme Siméon, Roel Vercammen
DEXA4
2005 Compiling XSLT 2.0 into XQuery 1.0
abstract
As XQuery is gathering momentum as the standard query language for XML, there is a growing interest in using it as an integral part of the XML application development infrastructure. In that context, one question which is often raised is how well XQuery interoperates with other XML languages, and notably with XSLT. XQuery 1.0 [16] and XSLT 2.0 [7] share a lot in common: they share XPath 2.0 as a common sub-language and have the same expressiveness. However, they are based on fairly different programming paradigms. While XSLT has adopted a highly declarative template based approach, XQuery relies on a simpler, and more operational, functional approach.In this paper, we present an approach to compile XSLT 2.0 into XQuery 1.0, and a working implementation of that approach. The compilation rules explain how XSLT's template-based approach can be implemented using the functional approach of XQuery and underpins the tight connection between the two languages. The resulting compiler can be used to migrate a XSLT code base to XQuery, or to enable the use of XQuery runtimes (e.g., as will soon be provided by most relational database management systems) for XSLT users. We also identify a number of areas where compatibility between the two languages could be improved. Finally, we show experiments on actual XSLT stylesheets, demonstrating the applicability of the approach in practice.
Achille Fokoue, Kristoffer Høgsbro Rose, Jérôme Siméon, Lionel Villard
WWW3
2004 Yoo-Hoo! Building a Presence Service with XQuery and WSDL
abstract
No abstract available.
Mary F. Fernández, Nicola Onose, Jérôme Siméon
SIGMOD Conference3
2004 XQuery at your web service
abstract
XML messaging is at the heart of Web services, providing the flexibility required for their deployment, composition, and maintenance. Yet, current approaches to Web services development hide the messaging layer behind Java or C# APIs, preventing the application to get direct access to the underlying XML information. To address this problem, we advocate the use of a native XML language, namely XQuery, as an integral part of the Web services development infrastructure. The main contribution of the paper is a binding between WSDL, the Web Services Description Language, and XQuery. The approach enables the use of XQuery for both Web services deployment and composition. We present a simple command-line tool that can be used to automatically deploy a Web service from a given XQuery module, and extend the XQuery language itself with a statement for accessing one or more Web services. The binding provides tight-coupling between WSDL and XQuery, yielding additional benefits, notably: the ability to use WSDL as an interface language for XQuery, and the ability to perform static typing on XQuery programs that include Web service calls. Last but not least, the proposal requires only minimal changes to the existing infrastructure. We report on our experience implementing this approach in the Galax XQuery processor.
Nicola Onose, Jérôme Siméon
WWW2
2003 Growing XQuery
Mary F. Fernández, Jérôme Siméon
ECOOP2
2003 Bridging the XML Relational Divide with LegoDB
abstract
We present LegoDB, a cost-based XML storage mapping engine that automatically explores a space of possible XML-to-relational mappings and selects an efficient mapping for a given application.
Philip Bohannon, Juliana Freire, Jayant R. Haritsa, Maya Ramanath, Prasan Roy, Jérôme Siméon
ICDE6
2003 The essence of XML
abstract
The World-Wide Web Consortium (W3C) promotes XML and related standards, including XML Schema, XQuery, and XPath. This paper describes a formalization of XML Schema. A formal semantics based on these ideas is part of the official XQuery and XPath specification, one of the first uses of formal methods by a standards body. XML Schema features both named and structural types, with structure based on tree grammars. While structural types and matching have been studied in other work (notably XDuce, Relax NG, and a previous formalization of XML Schema), this is the first work to study the relation between named types and structural types, and the relation between matching and validation.
Jérôme Siméon, Philip Wadler
POPL1
2003 Implementing Xquery 1.0: The Galax Experience
Mary F. Fernández, Jérôme Siméon, Byron Choi, Amélie Marian, Gargi Sur
VLDB2
2003 Projecting XML Documents
Amélie Marian, Jérôme Siméon
VLDB2
2003 Integrity constraints for XML
Wenfei Fan, Jérôme Siméon
J. Comput. Syst. Sci.2
2003 The Yin/Yang Web: A Unified Model for XML Syntax and RDF Semantics
abstract
XML is the W3C standard document format for writing and exchanging information on the Web. RDF is the W3C standard model for describing the semantics and reasoning about information on the Web. Unfortunately, RDF and XML-although very close to each other-are based on two different paradigms. We argue that, in order to lead the Semantic Web to its full potential, the syntax and the semantics of information need to work together. To this end, we develop a model theory for the XML XQuery 1.0 and XPath 2.0 Data Model, which provides a unified model for both XML and RDF. This unified model can serve as the basis for Web applications that deal with both data and semantics. We illustrate the use of this model on a concrete information integration scenario. Our approach enables each side of the fence to benefit from the other, notably, we show how the RDF world can take advantage of XML Schema description and XML query languages, and how the XML world can take advantage of the reasoning capabilities available for RDF. Our approach can also serve as a foundation for the next layer of the Semantic Web, the ontology layer, and we present a layering of an ontology language on top of our approach.
Peter F. Patel-Schneider, Jérôme Siméon
IEEE Trans. Knowl. Data Eng.2
2002 From XML Schema to Relations: A Cost-Based Approach to XML Storage
abstract
As Web applications manipulate an increasing amount of XML, there is a growing interest in storing XML data in relational databases. Due to the mismatch between the complexity of XML's tree structure and the simplicity of flat relational tables, there are many ways to store the same document in an RDBMS, and a number of heuristic techniques have been proposed. These techniques typically define fixed mappings and do not take application characteristics into account. However, a fixed mapping is unlikely to work well for all possible applications. In contrast, LegoDB is a cost-based XML storage mapping engine that explores a space of possible XML-to-relational mappings and selects the best mapping for a given application. LegoDB leverages current XML and relational technologies: (1) it models the target application with an XML Schema, XML data statistics, and an XQuery workload; (2) the space of configurations is generated through XML-Schema rewritings; and (3) the best among the derived configurations is selected using cost estimates obtained through a standard relational optimizer. We describe the LegoDB storage engine and provide experimental results that demonstrate the effectiveness of this approach.
Philip Bohannon, Juliana Freire, Prasan Roy, Jérôme Siméon
ICDE4
2002 Building the Semantic Web on XML
Peter F. Patel-Schneider, Jérôme Siméon
ISWC2
2002 StatiX: making XML count
abstract
The availability of summary data for XML documents has many applications, from providing users with quick feedback about their queries, to cost-based storage design and query optimization. StatiX is a novel XML Schema-aware statistics framework that exploits the structure derived by regular expressions (which define elements in an XML Schema) to pinpoint places in the schema that are likely sources of structural skew. As we discuss below, this information can be used to build concise, yet accurate, statistical summaries for XML data. StatiX leverages standard XML technology for gathering statistics, notably XML Schema validators, and it uses histograms to summarize both the structure and values in an XML document. In this paper we describe the StatiX system. We develop algorithms that decompose schemas to obtain statistics at different granularities and discuss how statistics can be gathered as documents are validated. We also present an experimental evaluation which demonstrates the accuracy and scalability of our approach and show an application of these statistics to cost-based XML storage design.
Juliana Freire, Jayant R. Haritsa, Maya Ramanath, Prasan Roy, Jérôme Siméon
SIGMOD Conference5
2002 LegoDB: Customizing Relational Storage for XML Documents
Philip Bohannon, Juliana Freire, Jayant R. Haritsa, Maya Ramanath, Prasan Roy, Jérôme Siméon
VLDB6
2002 The Yin/Yang web: XML syntax and RDF semantics
abstract
XML is the W3C standard document format for writing and exchanging information on the Web. RDF is the W3C standard model for describing the semantics and reasoning about information on the Web. Unfortunately, RDF and XML---although very close to each other---are based on two different paradigms. We argue that in order to lead the Semantic Web to its full potential, the syntax and the semantics of information needs to work together. To this end, we develop a model-theoretic semantics for the XML XQuery 1.0 and XPath 2.0 Data Model, which provides a unified model for both XML and RDF. This unified model can serve as the basis for Web applications that deal with both data and semantics. We illustrate the use of this model on a concrete information integration scenario. Our approach enables each side of the fence to benefit from the other, notably, we show how the RDF world can take advantage of XML query languages, and how the XML world can take advantage of the reasoning capabilities available for RDF.
Peter F. Patel-Schneider, Jérôme Siméon
WWW2
2002 A unified constraint model for XML
Wenfei Fan, Gabriel M. Kuper, Jérôme Siméon
Comput. Networks3
2001 A Semi-monad for Semi-structured Data
Mary F. Fernández, Jérôme Siméon, Philip Wadler
ICDT2
2001 Subsumption for XML types
Gabriel M. Kuper, Jérôme Siméon
ICDT2
2001 A unified constraint model for XML
abstract
Integrity constraints are an essential part of modern schema denition languages. They are useful for semantic specication, update consistency control, query optimization, information preservation, etc. In this paper, we propose UCM, a model of integrity constraints for XML that is both simple and expressive. Because it relies on a single notion of keys and foreign keys, the UCM model is easy to use and makes formal reasoning possible. Becauseitreliesonapowerful type system, the UCM model is expressive, capturing in a single framework the constraints found in relational databases, objectoriented schemas and XML DTDs. We study the problem of consistency of UCM constraints, the interaction between constraints and subtyping, and algorithms for implementing these constraints. Keywords XML, XML Schema, Integrity Constraints, Keys, Object Identity, Subtyping, Constraint Reasoning 1.
Wenfei Fan, Gabriel M. Kuper, Jérôme Siméon
WWW3
2000 An Algebra for XML Query
Mary F. Fernández, Jérôme Siméon, Philip Wadler
FSTTCS2
2000 Integrity Constraints for XML
abstract
Integrity constraints are useful for semantic specification, query optimization and data integration. The ID/IDREF mechanism provided by XML DTDs relics on a simple form of constraint to describe references. Yet, this mechanism is not sufficient to express semantic constraints, such as keys or inverse relationships, or stronger, object-style references. In this paper, we investigate integrity constraints for XML, both for semantic purposes and to improve its current reference mechanism. We extend DTDs with several families of constraints, including key, foreign key, inverse constraints and constraints specifying the semantics of object identities. These constraints are useful both for native XML documents and to preserve the semantics of data originating in relational or object databases. Complexity and axiomatization results are established for the (finite) implication problems associated with these constraints. These results also extend relational dependency theory on the interaction between (primary) keys and foreign keys. In addition, we investigate implication of more general constraints, such as functional, inclusion and inverse constraints defined in terms of navigation paths.
Wenfei Fan, Jérôme Siméon
PODS2
2000 On Wrapping Query Languages and Efficient XML Integration
abstract
Modern applications (Web portals, digital libraries, etc.) require integrated access to various information sources (from traditional DBMS to semistructured Web repositories), fast deployment and low maintenance cost in a rapidly evolving environment. Because of its flexibility, there is an increasing interest in using XML as a middleware model for such applications. XML enables fast wrapping and declarative integration. However, query processing in XML-based integration systems is still penalized by the lack of an algebra with adequate optimization properties and the difficulty to understand source query capabilities. In this paper, we propose an algebraic approach to support efficient XML query evaluation. We define a general purpose algebra suitable for semistructured on XML query languages. We show how this algebra can be used, with appropriate type information, to also wrap more structured query languages such as OQL or SQL. Finally, we develop new optimization techniques for XML-based integration systems.
Vassilis Christophides, Sophie Cluet, Jérôme Siméon
SIGMOD Conference3
1998 Your Mediators Need Data Conversion!
abstract
Due to the development of the World Wide Web, the integration of heterogeneous data sources has become a major concern of the database community. Appropriate architectures and query languages have been proposed. Yet, the problem of data conversion which is essential for the development of mediators/wrappers architectures has remained largely unexplored.
Sophie Cluet, Claude Delobel, Jérôme Siméon, Katarzyna Smaga
SIGMOD Conference3
1998 Using YAT to Build a Web Server
Jérôme Siméon, Sophie Cluet
WebDB1