Peter M. Fischer 0001

dblp:f/PeterMFischer1 · also Peter Michael Fischer · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
0since 2021 · last 2018
0009-0001-9794-3472ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 3 first-authorArtificial intelligence and machine learning · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
12 papers
Indexing and storage engines · 29% Spatial and temporal data management · 21% Query processing and optimization · 18%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 57% Performance modeling and evaluation · 20% Electronic design automation · 17%

Topics — the 22 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Spatial and temporal data management
temporal query processing
0.532015
Bi-temporal Timeline Index: A data structure for Processing Queries on bi-temporal data · ICDE 2015
Comprehensive and Interactive Temporal Query Processing with SAP HANA · Proc. VLDB Endow. 2013
Timeline index: a unified data structure for processing queries on temporal data in SAP HANA · SIGMOD Conference 2013
Indexing and storage engines
temporal indexing
0.322013
Comprehensive and Interactive Temporal Query Processing with SAP HANA · Proc. VLDB Endow. 2013
Timeline index: a unified data structure for processing queries on temporal data in SAP HANA · SIGMOD Conference 2013
Indexing and storage engines
in-memory index
0.212015
Bi-temporal Timeline Index: A data structure for Processing Queries on bi-temporal data · ICDE 2015
Query processing and optimization
XML query processing
0.222012
MXQuery with Hardware Acceleration · ICDE 2012
YFilter: Efficient and Scalable Filtering of XML Documents · ICDE 2002
Data models and query languages › XML query languages
XQuery
0.222009
XQuery Reloaded · Proc. VLDB Endow. 2009
Extending XQuery with Window Functions · VLDB 2007
Information retrieval › evaluation
benchmark
0.212013
A generic database benchmarking service · ICDE 2013
Indexing and storage engines › column store
main-memory column store
0.232015
Bi-temporal Timeline Index: A data structure for Processing Queries on bi-temporal data · ICDE 2015
Comprehensive and Interactive Temporal Query Processing with SAP HANA · Proc. VLDB Endow. 2013
Timeline index: a unified data structure for processing queries on temporal data in SAP HANA · SIGMOD Conference 2013
Query processing and optimization › XML query processing
XML projection
0.112012
MXQuery with Hardware Acceleration · ICDE 2012
Reconfigurable computing and FPGAs
FPGA accelerator
0.112012
MXQuery with Hardware Acceleration · ICDE 2012
Data stream processing
continuous query processing
0.112011
Changing flights in mid-air: a model for safely modifying continuous queries · SIGMOD Conference 2011
Query processing and optimization › XML query processing › XML query optimization
XQuery optimization
0.112009
XQuery Reloaded · Proc. VLDB Endow. 2009
Data stream processing › XML stream processing
XML filtering
0.122003
Path sharing and predicate evaluation for high-performance XML filtering · ACM Trans. Database Syst. 2003
YFilter: Efficient and Scalable Filtering of XML Documents · ICDE 2002
Data models and query languages › SQL
window functions
0.112007
Extending XQuery with Window Functions · VLDB 2007
Indexing and storage engines
adaptive indexing
0.112005
AGILE: Adaptive Indexing for Context-Aware Information Filters · SIGMOD Conference 2005
Information retrieval
information filtering
0.112005
Batched Processing for Information Filters · ICDE 2005
Performance modeling and evaluation
benchmarking
0.012013
A generic database benchmarking service · ICDE 2013
Electronic design automation
hardware/software co-design
0.012012
MXQuery with Hardware Acceleration · ICDE 2012
Query processing and optimization › query execution › expression evaluation
predicate evaluation
0.012003
Path sharing and predicate evaluation for high-performance XML filtering · ACM Trans. Database Syst. 2003
Data stream processing
stream processing systems
0.012011
Changing flights in mid-air: a model for safely modifying continuous queries · SIGMOD Conference 2011
Data stream processing › continuous query processing
stream filtering
0.012005
AGILE: Adaptive Indexing for Context-Aware Information Filters · SIGMOD Conference 2005
Distributed systems
publish/subscribe systems
0.012005
Batched Processing for Information Filters · ICDE 2005
Automata and formal languages › finite automata
nondeterministic finite automata
0.012002
YFilter: Efficient and Scalable Filtering of XML Documents · ICDE 2002

Methods — techniques the papers use, named apart from their topics

hardware acceleration · 0.3FPGA offloading · 0.3punctuation-based framework · 0.1message reordering · 0.1batching · 0.1NFA-based execution model · 0.1adaptive index structures · 0.1nondeterministic finite automaton · 0.0finite state machine · 0.0event-based parsing · 0.0
YearPublicationVenuePosition
2018 Web-scale provenance reconstruction of implicit information diffusion on social media
Io Taxidou, Sven Lieber, Peter M. Fischer 0001, Tom De Nies, Ruben Verborgh
Distributed Parallel Databases3
2015 Towards Multi-level Provenance Reconstruction of Information Diffusion on Social Media
abstract
In order to assess the trustworthiness of information on social media, a consumer needs to understand where this information comes from, and which processes were involved in its creation. The entities, agents and activities involved in the creation of a piece of information are referred to as its provenance, which was standardized by W3C PROV. However, current social media APIs cannot always capture the full lineage of every message, leaving the consumer with incomplete or missing provenance, which is crucial for judging the trust it carries. Therefore in this paper, we propose an approach to reconstruct the provenance of messages on social media on multiple levels. To obtain a fine-grained level of provenance, we use an approach from prior work to reconstruct information cascades with high certainty, and map them to PROV using the PROV-SAID extension for social media. To obtain a coarse-grained level of provenance, we adapt our similarity-based, fuzzy provenance reconstruction approach -- previously applied on news. We illustrate the power of the combination by providing the reconstructed provenance of a limited social media dataset gathered during the 2012 Olympics, for which we were able to reconstruct a significant amount of previously unidentified connections.
Tom De Nies, Io Taxidou, Anastasia Dimou, Ruben Verborgh, Peter M. Fischer 0001, Erik Mannens, Rik Van de Walle
CIKM5
2015 Bi-temporal Timeline Index: A data structure for Processing Queries on bi-temporal data
abstract
Following the adoption of basic temporal features in the SQL:2011 standard, there has been a tremendous interest within the database industry in supporting bi-temporal features, as a significant number of real-life workloads would greatly benefit from efficient temporal operations. However, current implementations of bi-temporal storage systems and operators are far from optimal. In this paper, we present the Bi-temporal Timeline Index, which supports a broad range of temporal operators and exploits the special properties of an in-memory column store database system. Comprehensive performance experiments with the TPC-BiH benchmark show that algorithms based on the Bi-temporal Timeline Index outperform significantly both existing commercial database systems and state-of-the-art data structures from research.
Martin Kaufmann, Peter M. Fischer 0001, Norman May, Chang Ge 0002, Anil K. Goel, Donald Kossmann
ICDE2
2015 Indexing bi-temporal windows
abstract
Bi-temporal databases support system (transaction) and application time, enabling users to query the history as recorded today and as it was known in the past. In this paper, we study windows over both system and application time, i.e., bi-temporal windows. We propose a two-dimensional index that supports one-time and continuous queries over fixed and sliding bi-temporal windows, covering static and streaming data. We demonstrate the advantages of the proposed index compared to the state-of-the-art in terms of query performance, index update overhead and space footprint.
Chang Ge 0002, Martin Kaufmann, Lukasz Golab, Peter M. Fischer 0001, Anil K. Goel
SSDBM4
2014 RApID: A System for Real-time Analysis of Information Diffusion in Twitter
abstract
The advent of social media has facilitated the study of information diffusion, expressing information spreading and influence among users on social graphs. In this demo paper, we present a system for real-time analysis of information diffusion on Twitter; it constructs the so-called information cascades that capture how information is being propagated from user to user. We face the challenge of managing and presenting large and fast-evolving graph data. For this purpose, we have developed methods for computing and visualizing information flow dynamically, offering rich structural and temporal information. The interface offers the possibility to interact with the dynamic, evolving cascades and gives valuable insights in terms of how information propagates on real-time and how users are influenced from each other.
Io Taxidou, Peter M. Fischer 0001
CIKM2
2014 Benchmarking Bitemporal Database Systems: Ready for the Future or Stuck in the Past?
abstract
After more than a decade of a virtual standstill, the adoption of temporal data management features has recently picked up speed, driven by customer demand and the inclusion of temporal expressions into SQL:2011. Most of the big commercial DBMS now include support for bitemporal data and operators. In this paper, we perform a thorough analysis of these commercial temporal DBMS: We investigate their architecture, determine their performance and study the impact of performance tuning. This analysis utilizes our recent (TPCTC 2013) benchmark proposal, which includes a comprehensive temporal workload definition. The results of our analysis show that the support for temporal data is still in its infancy: All systems store their data in regular, statically partitioned tables and rely on standard indexes as well as query rewrites for their operations. As shown by our measurements, this causes considerable performance variations on slight workload variations and significant overhead even after extensive tuning.
Martin Kaufmann, Peter M. Fischer 0001, Norman May, Donald Kossmann
EDBT2
2014 Efficient Stream Provenance via Operator Instrumentation
abstract
Managing fine-grained provenance is a critical requirement for data stream management systems (DSMS), not only for addressing complex applications that require diagnostic capabilities and assurance, but also for providing advanced functionality, such as revision processing or query debugging. This article introduces a novel approach that uses operator instrumentation, that is, modifying the behavior of operators, to generate and propagate fine-grained provenance through several operators of a query network. In addition to applying this technique to compute provenance eagerly during query execution, we also study how to decouple provenance computation from query processing to reduce runtime overhead and avoid unnecessary provenance retrieval. Our proposals include computing a concise superset of the provenance (to allow lazily replaying a query and reconstruct its provenance) as well as lazy retrieval (to avoid unnecessary reconstruction of provenance). We develop stream-specific compression methods to reduce the computational and storage overhead of provenance generation and retrieval. Ariadne, our provenance-aware extension of the Borealis DSMS implements these techniques. Our experiments confirm that Ariadne manages provenance with minor overhead and clearly outperforms query rewrite, the current state of the art.
Boris Glavic, Kyumars Sheykh Esmaili, Peter M. Fischer 0001, Nesime Tatbul
ACM Trans. Internet Techn.3
2013 A generic database benchmarking service
abstract
Benchmarks are widely applied for the development and optimization of database systems. Standard benchmarks such as TPC-C and TPC-H provide a way of comparing the performance of different systems. In addition, micro benchmarks can be exploited to test a specific behavior of a system. Yet, despite all the benefits that can be derived from benchmark results, the effort of implementing and executing benchmarks remains prohibitive: Database systems need to be set up, a large number of artifacts such as data generators and queries need to be managed and complex, time-consuming operations have to be orchestrated. In this demo, we introduce a generic benchmarking service that combines a rich meta model, low marginal cost and ease of use, which drastically reduces the time and cost to define, adapt and run a benchmark.
Martin Kaufmann, Peter M. Fischer 0001, Donald Kossmann, Norman May
ICDE2
2013 Timeline index: a unified data structure for processing queries on temporal data in SAP HANA
abstract
Managing temporal data is becoming increasingly important for many applications. Several database systems already support the time dimension, but provide only few temporal operators, which also often exhibit poor performance characteristics. On the academic side, a large number of algorithms and data structures have been proposed, but they often address a subset of these temporal operators only. In this paper, we develop the Timeline Index as a novel, unified data structure that efficiently supports temporal operators such as temporal aggregation, time travel, and temporal joins. As the Timeline Index is independent of the physical order of the data, it provides flexibility in physical design; e.g., it supports any kind of compression scheme, which is crucial for main memory column stores. Our experiments show that the Timeline Index has predictable performance and beats state-of-the-art approaches significantly, sometimes by orders of magnitude.
Martin Kaufmann, Amin Amiri Manjili, Panagiotis Vagenas, Peter M. Fischer 0001, Donald Kossmann, Franz Färber, Norman May
SIGMOD Conference4
2013 Comprehensive and Interactive Temporal Query Processing with SAP HANA
abstract
In this demo, we present a prototype of a main memory database system which provides a wide range of temporal operators featuring predictable and interactive response times. Much of real-life data is temporal in nature, and there is an increasing application demand for temporal models and operations in databases. Nevertheless, SQL:2011 has only recently overcome a decade-long standstill on standardizing temporal features. As a result, few database systems provide any temporal support, and even those only have limited expressiveness and poor performance. Our prototype combines an in-memory column store and a novel, generic temporal index structure named Timeline Index. As we will show on a workload based on real customer use cases, it achieves predictable and interactive query performance for a wide range of temporal query types and data sizes.
Martin Kaufmann, Panagiotis Vagenas, Peter M. Fischer 0001, Donald Kossmann, Franz Färber
Proc. VLDB Endow.3
2012 Transactional stream processing
abstract
Many stream processing applications require access to a multitude of streaming as well as stored data sources. Yet there is no clear semantics for correct continuous query execution over these data sources in the face of concurrent access and failures. Instead, today's Stream Processing Systems (SPSs) hard-code transactional concepts in their execution models, making them both hard to understand and inflexible to use. In this paper, we show that we can successfully reuse the traditional transactional theory (with some minimal extensions) in order to cleanly define the correct interaction of a set of continuous and one-time queries concurrently accessing both streaming and stored data sources. The result is a unified transactional model (UTM) for query processing over streams as well as traditional databases. We present a transaction manager that implements this model on top of an existing storage manager for streams (MXQuery/SMS). Experiments on the Linear Road Benchmark show that our transaction manager flexibly ensures correctness in case of concurrency and failures, without sacrificing from performance. Moreover, this model is powerful enough to express the implicit transactional behaviors of a representative set of state-of-the-art SPSs.
Irina Botan, Peter M. Fischer 0001, Donald Kossmann, Nesime Tatbul
EDBT2
2012 MXQuery with Hardware Acceleration
abstract
We demonstrate MXQuery/H, a modified version of MXQuery that uses hardware acceleration to speed up XML processing. The main goal of this demonstration is to give an interactive example of hardware/software co-design and show how system performance and energy efficiency can be improved by off-loading tasks to FPGA hardware. To this end, we equipped MXQuery/H with various hooks to inspect the different parts of the system. Besides that, our system can finally really leverage the idea of XML projection. Though the idea of projection had been around for a while, its effectiveness remained always limited because of the unavoidable and high parsing overhead. By performing the task in hardware, we relieve the software part from this overhead and achieve processing speed-ups of several factors.
Peter M. Fischer 0001, Jens Teubner
ICDE1
2011 Changing flights in mid-air: a model for safely modifying continuous queries
abstract
Continuous queries can run for unpredictably long periods of time. During their lifetime, these queries may need to be adapted either due to changes in application semantics (e.g., the implementation of a new alert detection policy), or due to changes in the system's behavior (e.g., adapting performance to a changing load). While in previous works query modification has been implicitly utilized to serve specific purposes (e.g., load management), to date no research has been done that defines a general-purpose, reliable, and efficiently implementable model for modifying continuous queries at run-time. In this paper, we introduce a punctuation-based framework that can formally express arbitrary lifecycle operations on the basis of input-output mappings and basic control elements such as start or stop of queries. On top of this foundation, we derive all possible query change methods, each providing different levels of correctness guarantees and performance. We further show how these models can be efficiently realized in a state-of-the-art stream processing engine; we also provide experimental results demonstrating the key performance tradeoffs of the change methods.
Kyumars Sheykh Esmaili, Tahmineh Sanamrad, Peter M. Fischer 0001, Nesime Tatbul
SIGMOD Conference3
2010 Stream schema: providing and exploiting static metadata for data stream processing
abstract
Schemas, and more generally metadata specifying structural and semantic constraints, are invaluable in data management.They facilitate conceptual design and enable checking of data consistency.They also play an important role in permitting semantic query optimization, that is, optimization and processing strategies that are often highly effective, but only correct for data conforming to a given schema.While the use of metadata is well-established in relational and XML databases, the same is not true for data streams.The existing work mostly focuses on the specification of dynamic information.In this paper, we consider the specification of static metadata for streams in a model called Stream Schema.We show how Stream Schema can be used to validate the consistency of streams.By explicitly modeling stream constraints, we show that stream queries can be simplified by removing predicates or subqueries that check for consistency.This can greatly enhance programmability of stream processing systems.We also present a set of semantic query optimization strategies that both permit compiletime checking of queries (for example, to detect empty queries) and new runtime processing options, options that would not have been possible without a Stream Schema specification.Case studies on two stream processing platforms (covering different applications and underlying stream models), along with an experimental evaluation, show the benefits of Stream Schema.
Peter M. Fischer 0001, Kyumars Sheykh Esmaili, Renée J. Miller
EDBT1
2009 Flexible and scalable storage management for data-intensive stream processing
abstract
Data Stream Management Systems (DSMS) operate under strict performance requirements. Key to meeting such requirements is to efficiently handle time-critical tasks such as managing internal states of continuous query operators, traffic on the queues between operators, as well as providing storage support for shared computation and archived data. In this paper, we introduce a general purpose storage management framework for DSMSs that performs these tasks based on a clean, loosely-coupled, and flexible system design that also facilitates performance optimization. An important contribution of the framework is that, in analogy to buffer management techniques in relational database systems, it uses information about the access patterns of streaming applications to tune and customize the performance of the storage manager. In the paper, we first analyze typical application requirements at different granularities in order to identify important tunable parameters and their corresponding values. Based on these parameters, we define a general-purpose storage management interface. Using the interface, a developer can use our SMS (Storage Manager for Streams) to generate a customized storage manager for streaming applications. We explore the performance and potential of SMS through a set of experiments using the Linear Road benchmark.
Irina Botan, Gustavo Alonso, Peter M. Fischer 0001, Donald Kossmann, Nesime Tatbul
EDBT3
2009 XQuery Reloaded
abstract
This paper describes a number of XQuery-related projects. Its goal is to show that XQuery is a useful tool for many different application scenarios. In particular, this paper tries to correct a common myth that XQuery is merely a query language and that SQL is the better query language. Instead, XQuery is a full-fledged programming language for Web applications and services. Furthermore, this paper tries to correct a second myth that XQuery is slow. This paper gives an overview of the state-of-the-art in XQuery implementation and optimization techniques and discusses one particular open-source XQuery processor, Zorba, in more detail. Among others, this paper presents an XQuery Benchmark Service which helps practitioners and XQuery processor vendors to find performance problems in an XQuery processor.
Roger Bamford, Vinayak R. Borkar, Matthias Brantner, Peter M. Fischer 0001, Daniela Florescu, David A. Graf, Donald Kossmann, Tim Kraska, Dan Muresan, Sorin Nasoi, Markos Zacharioudaki
Proc. VLDB Endow.4
2007 Extending XQuery with Window Functions
Irina Botan, Peter M. Fischer 0001, Daniela Florescu, Donald Kossmann, Tim Kraska, Rokas Tamosevicius
VLDB2
2005 Batched Processing for Information Filters
abstract
This paper describes batching, a novel technique in order to improve the throughput of an information filter (e.g. message broker or publish & subscribe system). Rather than processing each message individually, incoming messages are reordered, grouped and a whole group of similar messages is processed. This paper presents alternative strategies to do batching. Extensive performance experiments are conducted on those strategies in order to compare their tradeoffs.
Peter M. Fischer 0001, Donald Kossmann
ICDE1
2005 AGILE: Adaptive Indexing for Context-Aware Information Filters
abstract
Information filtering has become a key technology for modern information systems. The goal of an information filter is to route messages to the right recipients (possibly none) according to declarative rules called profiles. In order to deal with high volumes of messages, several index structures have been proposed in the past. The challenge addressed in this paper is to carry out stateful information filtering in which profiles refer to values in a database or to previous messages. The difficulty is that database update streams need to be processed in addition to messages. This paper presents AGILE, a way to extend existing index structures so that the indexes adapt to the message/update workload and show good performance in all situations. Performance experiments show that AGILE is overall the clear winner as compared to the best existing approaches. In extreme situations in which it is not the winner, the overheads are small.
Jens Dittrich, Peter M. Fischer 0001, Donald Kossmann
SIGMOD Conference2
2003 Path sharing and predicate evaluation for high-performance XML filtering
abstract
XML filtering systems aim to provide fast, on-the-fly matching of XML-encoded data to large numbers of query specifications containing constraints on both structure and content. It is now well accepted that approaches using event-based parsing and Finite State Machines (FSMs) can provide the basis for highly scalable structure-oriented XML filtering systems. The XFilter system [Altinel and Franklin 2000] was the first published FSM-based XML filtering approach. XFilter used a separate FSM per path query and a novel indexing mechanism to allow all of the FSMs to be executed simultaneously during the processing of a document. Building on the insights of the XFilter work, we describe a new method, called "YFilter" that combines all of the path queries into a single Nondeterministic Finite Automaton (NFA). YFilter exploits commonality among queries by merging common prefixes of the query paths such that they are processed at most once. The resulting shared processing provides tremendous improvements in structure matching performance but complicates the handling of value-based predicates.In this article, we first describe the XFilter and YFilter approaches and present results of a detailed performance comparison of structure matching for these algorithms as well as a hybrid approach. The results show that the path sharing employed by YFilter can provide order-of-magnitude performance benefits. We then propose two alternative techniques for extending YFilter's shared structure matching with support for value-based predicates, and compare the performance of these two techniques. The results of this latter study demonstrate some key differences between shared XML filtering and traditional database query processing. Finally, we describe how the YFilter approach is extended to handle more complicated queries containing nested path expressions.
Yanlei Diao, Mehmet Altinel, Michael J. Franklin, Hao Zhang 0003, Peter M. Fischer 0001
ACM Trans. Database Syst.5
2002 YFilter: Efficient and Scalable Filtering of XML Documents
abstract
Much of the data exchanged over the Internet will soon be encoded in XML, allowing for sophisticated filtering and content-based routing. We have built a filtering engine called YFilter, which filters streaming XML documents according to XQuery or XPath queries that involve both path expressions and predicates. Unlike previous work, YFilter uses a novel NFA-based execution model. We present the structures and algorithms underlying YFilter, and show its efficiency and scalability under various workloads.
Yanlei Diao, Peter M. Fischer 0001, Michael J. Franklin, Raymond To
ICDE2