Peter M. G. Apers

dblp:a/PeterMGApers · DBLP profile ↗
← Back
57ranked-venue papers
8as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 46 · 7 first-authorArtificial intelligence and machine learning · 17Software engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
14 papers
Information retrieval · 49% Query processing and optimization · 24% Distributed and cloud data management · 10%
Network and information security
1 paper
Privacy and data protection · 77% Systems and software security · 23%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Distributed systems · 69% Parallel and multicore computing · 27% Performance modeling and evaluation · 4%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 28 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search engines
expert finding
0.112007
Generative modeling of persons and documents for expert search · SIGIR 2007
Information retrieval › document retrieval
digital library search
0.122002
Content-Based Video Indexing for the Support of Digital Library Search · ICDE 2002
Flexible and scalable digital library search · VLDB 2001
Query processing and optimization
selectivity estimation
0.012004
A Selectivity Model for Fragmented Relations: Applied in Information Retrieval · IEEE Trans. Knowl. Data Eng. 2004
Multimedia analysis and retrieval › video indexing
content-based video indexing
0.012002
Content-Based Video Indexing for the Support of Digital Library Search · ICDE 2002
Database system architecture and tuning
main-memory database
0.021999
The Mirror MMDBMS Architecture · VLDB 1999
Parallelism in a Main-Memory DBMS: The Performance of PRISMA/DB · VLDB 1992
Information retrieval
search engines
0.012001
Flexible and scalable digital library search · VLDB 2001
Distributed systems
distributed coordination and fault tolerance
0.012001
Global transaction support for workflow management systems: from formal specification to practical implementation · VLDB J. 2001
Distributed systems
workflow management
0.012001
Global transaction support for workflow management systems: from formal specification to practical implementation · VLDB J. 2001
Parallel and multicore computing
parallel query processing
0.021995
Parallel Evaluation of Multi-Join Queries · SIGMOD Conference 1995
Data fragmentation for parallel transitive closure strategies · ICDE 1993
Query processing and optimization
parallel query processing
0.021992
PRISMA/DB: A Parallel Main Memory Relational DBMS · IEEE Trans. Knowl. Data Eng. 1992
Parallelism in a Main-Memory DBMS: The Performance of PRISMA/DB · VLDB 1992
Data integration and cleaning › interoperability
database interoperability
0.011996
The Role of Integrity Constraints in Database Interoperation · VLDB 1996
Database theory
integrity constraints
0.011996
The Role of Integrity Constraints in Database Interoperation · VLDB 1996
Query processing and optimization › join processing
parallel join
0.011995
Parallel Evaluation of Multi-Join Queries · SIGMOD Conference 1995
Query processing and optimization
join processing
0.011994
From Nested-Loop to Join Queries in OODB · VLDB 1994
Query processing and optimization › join processing › join algorithms
nested-loop join
0.011994
From Nested-Loop to Join Queries in OODB · VLDB 1994
Query processing and optimization › complex data query processing
object-oriented query processing
0.011994
From Nested-Loop to Join Queries in OODB · VLDB 1994
Information retrieval
large-scale retrieval
0.012001
Flexible and scalable digital library search · VLDB 2001
Information retrieval
retrieval models
0.012001
Flexible and scalable digital library search · VLDB 2001
Database system architecture and tuning
parallel database system
0.011992
PRISMA/DB: A Parallel Main Memory Relational DBMS · IEEE Trans. Knowl. Data Eng. 1992
Distributed and cloud data management
distributed query processing
0.021988
Data Allocation in Distributed Database Systems · ACM Trans. Database Syst. 1988
Optimization Algorithms for Distributed Queries · IEEE Trans. Software Eng. 1983
Graph data management › graph algorithms
transitive closure computation
0.011990
Distributed Transitive Closure Computations: The Disconnection Set Approach · VLDB 1990
Distributed systems
distributed algorithms
0.011990
Distributed Transitive Closure Computations: The Disconnection Set Approach · VLDB 1990
Distributed and cloud data management › distributed query processing
communication cost minimization
0.011988
Data Allocation in Distributed Database Systems · ACM Trans. Database Syst. 1988
Distributed and cloud data management
data placement
0.011988
Data Allocation in Distributed Database Systems · ACM Trans. Database Syst. 1988
Distributed and cloud data management › data placement
fragment allocation
0.011988
Data Allocation in Distributed Database Systems · ACM Trans. Database Syst. 1988
Parallel and multicore computing › parallel computing
parallel database systems
0.011992
Parallelism in a Main-Memory DBMS: The Performance of PRISMA/DB · VLDB 1992
Query processing and optimization › query optimization
distributed query optimization
0.011983
Optimization Algorithms for Distributed Queries · IEEE Trans. Software Eng. 1983
Distributed and cloud data management
distributed database design
0.011988
Data Allocation in Distributed Database Systems · ACM Trans. Database Syst. 1988

Methods — techniques the papers use, named apart from their topics

degradation model · 0.1probabilistic modeling · 0.1metadata extraction · 0.1language model · 0.1full-text indexing · 0.1conceptual modeling · 0.1zipfian distribution · 0.0mathematical modeling · 0.0simulation · 0.0comparative performance evaluation · 0.0disconnection set approach · 0.0object-oriented implementation · 0.0
YearPublicationVenuePosition
2014 Comparison of local and global undirected graphical models
Zhemin Zhu, Djoerd Hiemstra, Peter M. G. Apers, Andreas Wombacher
ESANN3
2013 Data Overload: What Can We Do?
Peter M. G. Apers
DASFAA (1)1
2013 Operation Properties: A Representation and their Role in the Propagation of Meta-Data
abstract
To facilitate the sharing and re-use of data in scientific studies we propose an automated technique for annotating operation results. The annotated output has to preserve, as much as possible, the properties of the input annotations. The preservation of properties is achieved by taking into account operation properties. Property preservation is evaluated with information theory metrics.
Juan Amiguet-Vercher, Peter M. G. Apers, Andreas Wombacher
e-Science2
2013 ProvenanceCurious: a tool to infer data provenance from scripts
abstract
The increasing data volume and highly complex models used in different domains make it difficult to debug models in cases of anomalies. Data provenance provides scientists sufficient information to investigate their models. In this paper, we propose a tool which can infer fine-grained data provenance based on a given script. The tool is demonstrated using a hydrological model. The tool is also tested successfully handling other scripts in different contexts.
M. Rezwanul Huq, Peter M. G. Apers, Andreas Wombacher
EDBT2
2013 Empirical Co-occurrence Rate Networks for Sequence Labeling
abstract
Structured prediction has wide applications in many areas. Powerful and popular models for structured prediction have been developed. Despite the successes, they suffer from some known problems: (i) Hidden Markov models are generative models which suffer from the mismatch problem. Also it is difficult to incorporate overlapping, non-independent features into a hidden Markov model explicitly. (ii) Conditional Markov models suffer from the label bias problem. (iii) Conditional Random Fields (CRFs) overcome the label bias problem by global normalization. But the global normalization of CRFs can be expensive which prevents CRFs from applying to big data. In this paper, we propose the Empirical Co-occurrence Rate Networks (ECRNs) for sequence labeling. ECRNs are discriminative models, so ECRNs overcome the problems of HMMs. ECRNs are also immune to the label bias problem even though they are locally normalized. To make the estimation of ECRNs as fast as possible, we simply use the empirical distributions as the estimation of parameters. Experiments on two real-world NLP tasks show that ECRNs reduce the training time radically while obtain competitive accuracy to the state-of-the-art models.
Zhemin Zhu, Djoerd Hiemstra, Peter M. G. Apers, Andreas Wombacher
ICMLA (1)3
2013 An Inference-Based Framework to Manage Data Provenance in Geoscience Applications
abstract
Data provenance allows scientists to validate their model as well as to investigate the origin of an unexpected value. Furthermore, it can be used as a replication recipe for output data products. However, capturing provenance requires enormous effort by scientists in terms of time and training. First, they need to design the workflow of the scientific model, i.e., workflow provenance, which requires both time and training. However, in practice, scientists may not document any workflow provenance before the model execution due to the lack of time and training. Second, they need to capture provenance while the model is running, i.e., fine-grained data provenance. Explicit documentation of fine-grained provenance is not feasible because of the massive storage consumption by provenance data in the applications, including those from the geoscience domain where data are continuously arriving and are processed. In this paper, we propose an inference-based framework, which provides both workflow and fine-grained data provenance at a minimal cost in terms of time, training, and disk consumption. Our proposed framework is applicable to any given scientific model, and is capable of handling different model dynamics, such as variation in the processing time as well as input data products arrival pattern. Our evaluation of the framework in a real use case with geospatial data shows that the proposed framework is relevant and suitable for scientists in geoscientific domain.
M. Rezwanul Huq, Peter M. G. Apers, Andreas Wombacher
IEEE Trans. Geosci. Remote. Sens.2
2012 Probabilistic Inference of Fine-Grained Data Provenance
M. Rezwanul Huq, Peter M. G. Apers, Andreas Wombacher
DEXA (1)2
2012 From scripts towards provenance inference
abstract
Scientists require provenance information either to validate their model or to investigate the origin of an unexpected value. However, they do not maintain any provenance information and even designing the processing workflow is rare in practice. Therefore, in this paper, we propose a solution that can build the workflow provenance graph by interpreting the scripts used for actual processing. Further, scientists can request fine-grained provenance information facilitating the inferred workflow provenance. We also provide a guideline to customize the workflow provenance graph based on user preferences. Our evaluation shows that the proposed approach is relevant and suitable for scientists to manage provenance.
M. Rezwanul Huq, Peter M. G. Apers, Andreas Wombacher, Yoshihide Wada, Ludovicus P. H. van Beek
eScience2
2012 The Identification Problem: A Description
abstract
Scientific data are often annotated based on their properties, which are not maintained during further data processing. Not maintaining annotations results in loss of information. Decisions made on such incomplete information may be wrong. In this paper the problem of propagating annotations along a data processing chain is formulated. In particular, an annotation of a data element is an identification that this data element exhibits a specific property. The propagation of this property from the input of an operation to its output is called the identification problem. In this paper the identification problem is described as a clustering problem.
Juan Amiguet-Vercher, Peter M. G. Apers, Andreas Wombacher
SERVICES2
2012 Fine-Grained Provenance Inference for a Large Processing Chain with Non-materialized Intermediate Views
M. Rezwanul Huq, Peter M. G. Apers, Andreas Wombacher
SSDBM2
2012 Simulating the future of concept-based video retrieval under improved detector performance
abstract
In this paper we address the following important questions for concept-based video retrieval: (1) What is the impact of detector performance on the performance of concept-based retrieval engines, and (2) will these engines be applicable to real-life search tasks if detector performance improves in the future? We use Monte Carlo simulations to answer these questions. To generate the simulation input, we propose to use a probabilistic model of two Gaussians for the confidence scores that concept detectors emit. Modifying the model’s parameters affects the detector performance and the search performance. We study the relation between these two performances on two video collections. For detectors with similar discriminative power and a concept vocabulary of around 100 concepts, the simulation reveals that in order to achieve a search performance of 0.20 mean average precision (MAP)—which is considered sufficient performance for real-life applications—one needs detectors with at least 0.60 MAP . We also find that, given our simulation model and low detector performance, MAP is not always a good evaluation measure for concept detectors since it is not strongly correlated with the search performance.
Robin Aly, Djoerd Hiemstra, Franciska de Jong, Peter M. G. Apers
Multim. Tools Appl.4
2011 Inferring Fine-Grained Data Provenance in Stream Data Processing: Reduced Storage Cost, High Accuracy
M. Rezwanul Huq, Andreas Wombacher, Peter M. G. Apers
DEXA (2)3
2011 Adaptive Inference of Fine-grained Data Provenance to Achieve High Accuracy at Lower Storage Costs
abstract
In stream data processing, data arrives continuously and is processed by decision making, process control and e-science applications. To control and monitor these applications, reproducibility of result is a vital requirement. However, it requires massive amount of storage space to store fine-grained provenance data especially for those transformations with overlapping sliding windows. In this paper, we propose techniques which can significantly reduce storage costs and can achieve high accuracy. Our evaluation shows that adaptive inference technique can achieve almost 100% accurate provenance information for a given dataset at lower storage costs than the other techniques. Moreover, we present a guideline about the usage of different provenance collection techniques described in this paper based on the transformation operation and stream characteristics.
M. Rezwanul Huq, Andreas Wombacher, Peter M. G. Apers
eScience3
2008 Data degradation: making private data less sensitive over time
abstract
Trail disclosure is the leakage of privacy sensitive data, resulting from negligence, attack or abusive scrutinization or usage of personal digital trails. To prevent trail disclosure, data degradation is proposed as an alternative to the limited retention principle. Data degradation is based on the assumption that long lasting purposes can often be satisfied with a less accurate, and therefore less sensi-tive, version of the data. Data will be progressively degraded such that it still serves application purposes, while decreasing accuracy and thus privacy sensitivity.
Nicolas Anciaux, Luc Bouganim, Harold van Heerde, Philippe Pucheral, Peter M. G. Apers
CIKM5
2008 InstantDB: Enforcing Timely Degradation of Sensitive Data
abstract
People cannot prevent personal information from being collected by various actors. Several security measures are implemented on servers to minimize the possibility of a privacy violation. Unfortunately, even the most well defended servers are subject to attacks and however much one trusts a hosting organism/company, such trust does not last forever. We propose a simple and practical degradation model where sensitive data undergoes a progressive and irreversible degradation from an accurate state at collection time, to intermediate but still informative fuzzy states, to complete disappearance. We introduce the data degradation model and identify related technical challenges and open issues.
Nicolas Anciaux, Luc Bouganim, Harold van Heerde, Philippe Pucheral, Peter M. G. Apers
ICDE5
2007 Generative modeling of persons and documents for expert search
abstract
In this paper we address the task of automatically finding an expert within the organization, known as the expert search problem. We present the theoretically-based probabilistic algorithm which models retrieved documents as mixtures of expert candidate language models. Experiments show that our approach outperforms existing theoretically sound solutions.
Pavel Serdyukov, Djoerd Hiemstra, Maarten M. Fokkinga, Peter M. G. Apers
SIGIR4
2006 A Context-Aware Preference Model for Database Querying in an Ambient Intelligent Environment
Arthur H. van Bunningen, Peter M. G. Apers
DEXA3
2005 Score region algebra: building a transparent XML-R database
abstract
A unified database framework that will enable better comprehension of ranked XML retrieval is still a challenge in the XML database field. We propose a logical algebra, named score region algebra, that enables transparent specification of information retrieval (IR) models for XML databases. The transparency is achieved by a possibility to instantiate various retrieval models, using abstract score functions within algebra operators, while logical query plan and operator definitions remain unchanged. Our algebra operators model three important aspects of XML retrieval: element relevance score computation, element score propagation, and element score combination. To illustrate the usefulness of our algebra we instantiate four different, well known IR scoring models, and combine them with different score propagation and combination functions. We implemented the algebra operators in a prototype system on top of a low-level database kernel. The evaluation of the system is performed on a collection of IEEE articles in XML format provided by INEX. We argue that state of the art XML IR models can be transparently implemented using our score region algebra framework on top of any low-level physical database engine or existing RDBMS, allowing a more systematic investigation of retrieval model behavior.
Vojkan Mihajlovic, Henk Ernst Blok, Djoerd Hiemstra, Peter M. G. Apers
CIKM4
2004 Towards Context-Aware Data Management for Ambient Intelligence
Peter M. G. Apers, Willem Jonker
DEXA2
2004 A Selectivity Model for Fragmented Relations: Applied in Information Retrieval
abstract
New application domains cause today's database sizes to grow rapidly, posing great demands on technology. Data fragmentation facilitates techniques (like distribution, parallelization. and main-memory computing) meeting these demands. Also, fragmentation might help to improve efficient processing of query types such as top N. Database design and query optimization require a good notion of the costs resulting from a certain fragmentation. Our mathematically derived selectivity model facilitates this. Once its two parameters have been computed based on the fragmentation, after each (though usually infrequent) update, our model can forget the data distribution, resulting in fast and quite good selectivity estimation. We show experimental verification for Zipfian distributed IR databases.
Henk Ernst Blok, Sunil Choenni, Henk M. Blanken, Peter M. G. Apers
IEEE Trans. Knowl. Data Eng.4
2002 Content-Based Video Indexing for the Support of Digital Library Search
abstract
Presents a digital library search engine that combines efforts of the AMIS and DMW research projects, each covering significant parts of the problem of finding the required information in an enormous mass of data. The most important contributions of our work are the following: (1) We demonstrate a flexible solution for the extraction and querying of meta-data from multimedia documents in general. (2) Scalability and efficiency support are illustrated for full-text indexing and retrieval. (3) We show how, for a more limited domain, like an intranet, conceptual modelling can offer additional and more powerful query facilities. (4) In the limited domain case, we demonstrate how domain knowledge can be used to interpret low-level features into semantic content. In this short description, we focus on the first and fourth items.
Milan Petkovic, Roelof van Zwol, Henk Ernst Blok, Willem Jonker, Peter M. G. Apers, Menzo Windhouwer, Martin L. Kersten
ICDE5
2002 Editorial
Peter M. G. Apers, Stefano Ceri, Richard T. Snodgrass
VLDB J.1
2001 Predicting the Cost-Quality Trade-Off for Information Retrieval Queries: Facilitating Database Design and Query Optimization
abstract
Efficient, flexible, and scalable integration of full text information retrieval (IR) in a DBMS is not a trivial case. This holds in particular for query optimization in such a context. To facilitate the bulk-oriented behavior of database query processing, a priori knowledge of how to limit the data efficiently prior to query evaluation is very valuable at optimization time. The usually imprecise nature of IR querying provides an extra opportunity to limit the data by a trade-off with the quality of the answer. In this paper we present a mathematically derived model to predict the quality implications of neglecting information before query execution. In particular we investigate the possibility to predict the retrieval quality for a document collection for which no training information is available, which is usually the case in practice. Instead, we construct a model that can be trained on other document collections for which the necessary quality information is available, or can be obtained quite easily. We validate our model for several document collections and present the experimental results. These results show that our model performs quite well, even for the case were we did not train it on the test collection itself.
Henk Ernst Blok, Djoerd Hiemstra, Sunil Choenni, Franciska de Jong, Henk M. Blanken, Peter M. G. Apers
CIKM6
2001 Flexible and scalable digital library search
Henk Ernst Blok, Menzo Windhouwer, Roelof van Zwol, Milan Petkovic, Peter M. G. Apers, Martin L. Kersten, Willem Jonker
VLDB5
2001 Global transaction support for workflow management systems: from formal specification to practical implementation
Paul Grefen, Jochem Vonk, Peter M. G. Apers
VLDB J.3
2000 The Webspace Method: On the Integration of Database Technology with Multimedia Retrieval
abstract
Large collections of documents containing various types of multimedia, are made available to the WWW. Unfortunately, due to the un-structuredness of Internet environments it is hard to find specific information when one is looking for it.Search engines available can only rely their results on information retrieval techniques and most of the time they lack the desired power in query formulation.Modelling data on the web, as if it was designed for use within databases, should provide us with the necessary basis for enhancing this query formulation.This of course requires special care for dealing with the included multimedia data and the semi-structured aspects of data on the web.Modelling the entire web would be too ambitious, therefore we focus on a more feasible environment, like the intranet, where one can find large collections of related data.With the webspace method we have already shown how to deal with the various aspects of semi-structured data in large collections of related documents.In this paper we focus on the integration of our webspace method for concept-based search with content-based multimedia information retrieval (IR).A webspace consists of two levels.At the document level, a webspace is considered to be a collection of related documents.At the semantical level, concepts are defined to be used in the documents at the document level.By modelling these concepts using a webspace schema a semantical level of abstraction is gained.This supplies the necessary platform for querying data available within a specific webspace.For the integration with content-based information retrieval an existing IR model is adopted.We will discuss how this is used in the context of Mirror, a Multimedia DBMS, and how this framework is used for the integration with the webspace method for concept-based search.
Roelof van Zwol, Peter M. G. Apers
CIKM2
2000 Modelling the Webspace of an Intranet
abstract
Searching the Internet using the currently available search engines is not satisfactory. The techniques used there focus on the extraction of relevant information directly from the documents available on the World Wide Web. We introduce a new approach, which aims at describing the content of a Web space, formed by a collection of related documents, instead of looking at the single documents. By identifying concepts and the relationships among them, the content of a Web space is described semantically in a schema for the Web space. The main objective is that, by following this approach, we can start querying the content of a collection of related documents rather than the content of a single document. In this paper, we introduce a model for Web spaces that allows us to describe the concepts at a semantic level, in terms of classes, associations over classes and attributes of classes. At the syntactic level, we use XML to describe information as instantiations of the concepts defined in the Web space schema. Dealing with data on the Web implies dealing with semi-structured data. We discuss how this relates to our model for a Web space and show how to deal with these aspects efficiently when moving towards an implementation.
Roelof van Zwol, Peter M. G. Apers
WISE2
1999 Semantics and Architecture of Global Transaction Support in Workflow Environments
abstract
We present an approach to global transaction management in workflow environments. The transaction mechanism is based on the well-known notion of sagas, but extended to deal with arbitrary process structures including cycles and savepoints that allow partial compensation. We present a formal specification of the transaction model and transaction management mechanisms in set and graph theory, providing clear, unambiguous transaction semantics. The specification is straightforwardly mapped to a modular architecture, the implementation of which is applied in the prototype of a commercial workflow management system. The loosely-coupled nature of the resulting system allows easy distribution using middleware technology.
Paul Grefen, Jochem Vonk, Erik M. Boertjes, Peter M. G. Apers
CoopIS4
1999 Distributed Global Transaction Support for Workflow Management Applications
Jochem Vonk, Paul Grefen, Erik M. Boertjes, Peter M. G. Apers
DEXA4
1999 The Mirror MMDBMS Architecture
Arjen P. de Vries, Mark G. L. M. van Doorn, Henk M. Blanken, Peter M. G. Apers
VLDB4
1998 An Architecture for Nested Transaction Support on Standard Database Systems
Erik M. Boertjes, Paul Grefen, Jochem Vonk, Peter M. G. Apers
DEXA4
1998 Specifying Global Behaviour in Database Federations
Mark W. W. Vermeer, Peter M. G. Apers
Inf. Syst.2
1997 Behaviour Specification in Database Interoperation
Mark W. W. Vermeer, Peter M. G. Apers
CAiSE2
1997 Query Modification in Object-Oriented Database Federations
abstract
We discuss the modification of queries against an integrated view in a federation of object-oriented databases. We present a generalisation of existing algorithms for simple global query processing that works for arbitrarily defined integration classes. We then extend this algorithm to deal with object-oriented features such as queries involving path expressions and nesting. We show how properties of the OO-style of modelling relationships through object references can be exploited to reduce the number of subqueries necessary to evaluate such queries.
Mark W. W. Vermeer, Peter M. G. Apers
CoopIS2
1997 Two-Layer Transaction Management for Workflow Management Applications
Paul Grefen, Jochem Vonk, Erik M. Boertjes, Peter M. G. Apers
DEXA4
1997 Modifying Queries on Complex Objects in Database Federations
abstract
We discuss the modification of queries against an integrated view in a federation of object-oriented databases. We present a generalization of existing algorithms for simple global query modification that works for arbitrarily defined integration classes. We then extend this algorithm to deal with object-oriented features that have been neglected in literature, viz. queries involving path expressions and nested queries. We show how properties of the OO-style of modeling relationships through object references can be exploited to reduce the number of subqueries necessary to evaluate such queries.
Mark W. W. Vermeer, Peter M. G. Apers
Int. J. Cooperative Inf. Syst.2
1996 On the Applicability of Schema Integration Techniques to Database Interoperation
Mark W. W. Vermeer, Peter M. G. Apers
ER2
1996 The Role of Integrity Constraints in Database Interoperation
Mark W. W. Vermeer, Peter M. G. Apers
VLDB2
1995 Object-Oriented Views of Relational Databases Incorporating Behaviour
Mark W. W. Vermeer, Peter M. G. Apers
DASFAA2
1995 Parallel Evaluation of Multi-Join Queries
abstract
A number of execution strategies for parallel evaluation of multi-join queries have been proposed in the literature; their performance was evaluated by simulation. In this paper we give a comparative performance evaluation of four execution strategies by implementing all of them on the same parallel database system, PRISMA/DB. Experiments have been done up to 80 processors. The basic strategy is to first determine an execution schedule with minimum total cost and then parallelize this schedule with one of the four execution strategies. These strategies, coming from the literature, are named: Sequential Parallel, Synchronous Execution, Segmented Right-Deep, and Full Parallel. Based on the experiments clear guidelines are given when to use which strategy.
Annita N. Wilschut, Jan Flokstra, Peter M. G. Apers
SIGMOD Conference3
1994 Optimization of Nested Queries in a Complex Object Model
Hennie J. Steenhagen, Peter M. G. Apers, Henk M. Blanken
EDBT2
1994 From Nested-Loop to Join Queries in OODB
Hennie J. Steenhagen, Peter M. G. Apers, Henk M. Blanken, Rolf A. de By
VLDB2
1993 Data fragmentation for parallel transitive closure strategies
abstract
Addresses the problem of fragmenting a relation to make the parallel computation of the transitive closure efficient, based on the disconnection set approach. To better understand this design problem, the authors focus on transportation networks. These are characterized by loosely interconnected clusters of nodes with a high internal connectivity rate. Three requirements that have to be fulfilled by a fragmentation are formulated, and three different fragmentation strategies are presented, each emphasizing one of these requirements. Some test results are presented to show the performance of the various fragmentation strategies.>
Maurice A. W. Houtsma, Peter M. G. Apers, Gideon L. V. Schipper
ICDE2
1993 Integrity Control in Relational Database Systems - An Overview
Paul Grefen, Peter M. G. Apers
Data Knowl. Eng.2
1993 Dataflow Query Execution in a Parallel Main-memory Environment
Annita N. Wilschut, Peter M. G. Apers
Distributed Parallel Databases2
1992 Performance Evaluation of Integrity Control in a Parallel Main-Memory Database System
Paul Grefen, Jan Flokstra, Peter M. G. Apers
DEXA3
1992 Parallelism in a Main-Memory DBMS: The Performance of PRISMA/DB
Annita N. Wilschut, Jan Flokstra, Peter M. G. Apers
VLDB3
1992 PRISMA/DB: A Parallel Main Memory Relational DBMS
abstract
PRISMA/DB, a full-fledged parallel, main memory relational database management system (DBMS) is described. PRISMA/DB's high performance is obtained by the use of parallelism for query processing and main memory storage of the entire database. A flexible architecture for experimenting with functionality and performance is obtained using a modular implementation of the system in an object-oriented programming language. The design and implementation of PRISMA/DB are described in detail. A performance evaluation of the system shows that the system is comparable to other state-of-the-art database machines. The prototype implementation of the system runs on a 100-node parallel multiprocessor.>
Peter M. G. Apers, Carel A. van den Berg, Jan Flokstra, Paul Grefen, Martin L. Kersten, Annita N. Wilschut
IEEE Trans. Knowl. Data Eng.1
1991 Integrity Constraint Enforcement through Transaction Modification
Paul Grefen, Peter M. G. Apers
DEXA2
1991 Algebraic optimization of recursive queries
Maurice A. W. Houtsma, Peter M. G. Apers
Data Knowl. Eng.2
1991 Schema architectures and their relationship to transaction processing in distributed database systems
Peter M. G. Apers, Peter Scheuermann
Inf. Sci.1
1990 Complex Transitive Closure Queries on a Fragmented Graph
Maurice A. W. Houtsma, Peter M. G. Apers, Stefano Ceri
ICDT2
1990 Distributed Transitive Closure Computations: The Disconnection Set Approach
Maurice A. W. Houtsma, Peter M. G. Apers, Stefano Ceri
VLDB2
1989 Future research directions: Evidence from this conference
Peter M. G. Apers
VLDB1
1988 PRISMA Database Machine: A Distributed, Main-Memory Approach
Peter M. G. Apers, Martin L. Kersten, Hans Oerlemans
EDBT1
1988 Data Allocation in Distributed Database Systems
abstract
The problem of allocating the data of a database to the sites of a communication network is investigated. This problem deviates from the well-known file allocation problem in several aspects. First, the objects to be allocated are not known a priori; second, these objects are accessed by schedules that contain transmissions between objects to produce the result. A model that makes it possible to compare the cost of allocations is presented; the cost can be computed for different cost functions and for processing schedules produced by arbitrary query processing algorithms. For minimizing the total transmission cost, a method is proposed to determine the fragments to be allocated from the relations in the conceptual schema and the queries and updates executed by the users. For the same cost function, the complexity of the data allocation problem is investigated. Methods for obtaining optimal and heuristic solutions under various ways of computing the cost of an allocation are presented and compared. Two different approaches to the allocation management problem are presented and their merits are discussed.
Peter M. G. Apers
ACM Trans. Database Syst.1
1983 Optimization Algorithms for Distributed Queries
abstract
The efficiency of processing strategies for queries in a distributed database is critical for system performance. Methods are studied to minimize the response time and the total time for distributed queries. A new algorithm (Algorithm GENERAL) is presented to derive processing strategies for arbitrarily complex queries. Three versions of the algorithm are given: one for minimizing response time and two for minimizing total time. The algorithm is shown to provide optimal solutions under certain conditions.
Peter M. G. Apers, Alan R. Hevner
IEEE Trans. Software Eng.1