EDBT 2026 Demo / reviewers in the wild / expert
Marcin Wylot
dblp:09/10285
· DBLP profile ↗
8ranked-venue papers
6as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Graph data management · 32% Query processing and optimization · 31% Distributed and cloud data management · 22% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 54% Cloud and datacenter computing · 46% |
Topics — the 8 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
provenance query processing |
0.4 | 2 | 2015 | Executing Provenance-Enabled Queries over Web Data · WWW 2015 TripleProv: efficient processing of lineage queries in a native RDF store · WWW 2014 |
Distributed and cloud data management
distributed RDF processing |
0.2 | 1 | 2016 | DiploCloud: Efficient and Scalable Management of RDF Data in the Cloud · IEEE Trans. Knowl. Data Eng. 2016 |
Distributed and cloud data management › data partitioning
RDF data partitioning |
0.2 | 1 | 2016 | DiploCloud: Efficient and Scalable Management of RDF Data in the Cloud · IEEE Trans. Knowl. Data Eng. 2016 |
Data integration and cleaning › data provenance
provenance management |
0.2 | 1 | 2015 | A Demonstration of TripleProv: Tracking and Querying Provenance over Web Data · Proc. VLDB Endow. 2015 |
Graph data management › graph data model
RDF data model |
0.2 | 1 | 2015 | A Demonstration of TripleProv: Tracking and Querying Provenance over Web Data · Proc. VLDB Endow. 2015 |
Graph data management › RDF data management
RDF triple store |
0.2 | 1 | 2015 | A Demonstration of TripleProv: Tracking and Querying Provenance over Web Data · Proc. VLDB Endow. 2015 |
Cloud and datacenter computing
cloud data management |
0.1 | 1 | 2016 | DiploCloud: Efficient and Scalable Management of RDF Data in the Cloud · IEEE Trans. Knowl. Data Eng. 2016 |
Data models and query languages › semistructured data
RDF data |
0.1 | 1 | 2015 | Executing Provenance-Enabled Queries over Web Data · WWW 2015 |
Methods — techniques the papers use, named apart from their topics
query tailoring · 0.6provenance tracking · 0.6physiological analysis · 0.5graph partitioning · 0.5adaptive query materialization · 0.2empirical evaluation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Storing, Tracking, and Querying Provenance in Linked DataabstractThe proliferation of heterogeneous Linked Data on the Web poses new challenges to database systems. In particular, the capacity to store, track, and query provenance data is becoming a pivotal feature of modern triplestores. We present methods extending a native RDF store to efficiently handle the storage, tracking, and querying of provenance in RDF data. We describe a reliable and understandable specification of the way results were derived from the data and how particular pieces of data were combined to answer a query. Subsequently, we present techniques to tailor queries with provenance data. We empirically evaluate the presented methods and show that the overhead of storing and tracking provenance is acceptable. Finally, we show that tailoring a query with provenance information can also significantly improve the performance of query execution. Marcin Wylot, Philippe Cudré-Mauroux, Manfred Hauswirth, Paul Groth |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | DiploCloud: Efficient and Scalable Management of RDF Data in the CloudabstractDespite recent advances in distributed RDF data management, processing large-amounts of RDF data in the cloud is still very challenging. In spite of its seemingly simple data model, RDF actually encodes rich and complex graphs mixing both instance and schema-level data. Sharding such data using classical techniques or partitioning the graph using traditional min-cut algorithms leads to very inefficient distributed operations and to a high number of joins. In this paper, we describe DiploCloud, an efficient and scalable distributed RDF data management system for the cloud. Contrary to previous approaches, DiploCloud runs a physiological analysis of both instance and schema information prior to partitioning the data. In this paper, we describe the architecture of DiploCloud, its main data structures, as well as the new algorithms we use to partition and distribute data. We also present an extensive evaluation of DiploCloud showing that our system is often two orders of magnitude faster than state-of-the-art systems on standard workloads. Marcin Wylot, Philippe Cudré-Mauroux |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | A Comparison of Data Structures to Manage URIs on the Web of Data
Ruslan Mavlyutov, Marcin Wylot, Philippe Cudré-Mauroux |
ESWC | 2 |
| 2015 | Executing Provenance-Enabled Queries over Web DataabstractThe proliferation of heterogeneous Linked Data on the Web poses new challenges to database systems. In particular, because of this heterogeneity, the capacity to store, track, and query provenance data is becoming a pivotal feature of modern triple stores. In this paper, we tackle the problem of efficiently executing provenance-enabled queries over RDF data. We propose, implement and empirically evaluate five different query execution strategies for RDF queries that incorporate knowledge of provenance. The evaluation is conducted on Web Data obtained from two different Web crawls (The Billion Triple Challenge, and the Web Data Commons). Our evaluation shows that using an adaptive query materialization execution strategy performs best in our context. Interestingly, we find that because provenance is prevalent within Web Data and is highly selective, it can be used to improve query processing performance. This is a counterintuitive result as provenance is often associated with additional overhead. Marcin Wylot, Philippe Cudré-Mauroux, Paul Groth |
WWW | 1 |
| 2015 | A Demonstration of TripleProv: Tracking and Querying Provenance over Web DataabstractThe proliferation of heterogeneous Linked Data on the Web poses new challenges to database systems. In particular, the capacity to store, track, and query provenance data is becoming a pivotal feature of modern triple stores. In this demonstration, we present TripleProv: a new system extending a native RDF store to efficiently handle the storage, tracking and querying of provenance in RDF data. In the following, we give an overview of our approach providing a reliable and understandable specification of the way results were derived from the data and how particular pieces of data were combined to answer the query. Subsequently, we present techniques enabling to tailor queries with provenance data. Finally, we describe our demonstration and how the attendees will be able to interact with our system during the conference. Marcin Wylot, Philippe Cudré-Mauroux, Paul Groth |
Proc. VLDB Endow. | 1 |
| 2014 | TripleProv: efficient processing of lineage queries in a native RDF storeabstractGiven the heterogeneity of the data one can find on the Linked Data cloud, being able to trace back the provenance of query results is rapidly becoming a must-have feature of RDF systems. While provenance models have been extensively discussed in recent years, little attention has been given to the efficient implementation of provenance-enabled queries inside data stores. This paper introduces TripleProv: a new system extending a native RDF store to efficiently handle such queries. TripleProv implements two different storage models to physically co-locate lineage and instance data, and for each of them implements algorithms for tracing provenance at two granularity levels. In the following, we present the overall architecture of our system, its different lineage storage models, and the various query execution strategies we have implemented to efficiently answer provenance-enabled queries. In addition, we present the results of a comprehensive empirical evaluation of our system over two different datasets and workloads. Marcin Wylot, Philippe Cudré-Mauroux, Paul Groth |
WWW | 1 |
| 2013 | NoSQL Databases for RDF: An Empirical Evaluation
Philippe Cudré-Mauroux, Iliya Enchev, Sever Fundatureanu, Paul Groth, Albert Haque, Andreas Harth, Felix Leif Keppmann, Daniel P. Miranker, Juan F. Sequeda, Marcin Wylot |
ISWC (2) | 10 |
| 2011 | dipLODocus[RDF] - Short and Long-Tail RDF Analytics for Massive Webs of Data
Marcin Wylot, Jigé Pont, Mariusz Wisniewski, Philippe Cudré-Mauroux |
ISWC (1) | 1 |