EDBT 2026 Demo / reviewers in the wild / expert
Patrick Valduriez
dblp:v/PatrickValduriez
· DBLP profile ↗
117ranked-venue papers in the field
10as first author
5since 2021 · last 2025
0000-0001-6506-7538ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 96 (9 first)Knowledge Engineering, Semantic Web & Information Systems · 7 (1 first)Data Mining & Knowledge Discovery · 6Information Retrieval & Web Search · 4Big Data, Cloud & Distributed Data Systems · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Federated Learning with Heterogeneous Data and Adaptive DropoutabstractFederated Learning (FL) is a promising distributed machine learning approach that enables collaborative training of a global model using multiple edge devices. The data distributed among the edge devices are highly heterogeneous. Thus, FL faces the challenge of data distribution and heterogeneity, where non-Independent and Identically Distributed (non-IID) data across edge devices may yield in significant accuracy drop. Furthermore, the limited computation and communication capabilities of edge devices increase the likelihood of stragglers, thus leading to slow model convergence. In this article, we propose the FedDHAD FL framework, which comes with two novel methods: dynamic heterogeneous model aggregation (FedDH) and adaptive dropout (FedAD). FedDH dynamically adjusts the weights of each local model within the model aggregation process based on the non-IID degree of heterogeneous data to deal with the statistical data heterogeneity. FedAD performs neuron-adaptive operations in response to heterogeneous devices to improve accuracy while achieving superb efficiency. The combination of these two methods makes FedDHAD significantly outperform state-of-the-art solutions in terms of accuracy (up to 6.7% higher), efficiency (up to 2.02 times faster), and computation cost (up to 15.0% smaller). Ji Liu 0003, Beichen Ma, Qiaolin Yu, Ruoming Jin, Jingbo Zhou 0003, Yang Zhou 0001, Huaiyu Dai, Haixun Wang, Dejing Dou, Patrick Valduriez |
ACM Trans. Knowl. Discov. Data | 10 |
| 2024 | AEDFL: Efficient Asynchronous Decentralized Federated Learning with Heterogeneous DevicesabstractFederated Learning (FL) has achieved significant achievements recently, enabling collaborative model training on distributed data over edge devices. Iterative gradient or model exchanges between devices and the centralized server in the standard FL paradigm suffer from severe efficiency bottlenecks on the server. While enabling collaborative training without a central server, existing decentralized FL approaches either focus on the synchronous mechanism that deteriorates FL convergence or ignore device staleness with an asynchronous mechanism, resulting in inferior FL accuracy. In this paper, we propose an Asynchronous Efficient Decentralized FL framework, i.e., AEDFL, in heterogeneous environments with three unique contributions. First, we propose an asynchronous FL system model with an efficient model aggregation method for improving the FL convergence. Second, we propose a dynamic staleness-aware model update approach to achieve superior accuracy. Third, we propose an adaptive sparse training method to reduce communication and computation costs without significant accuracy degradation. Extensive experimentation on four public datasets and four models demonstrates the strength of AEDFL in terms of accuracy (up to 16.3% higher), efficiency (up to 92.9% faster), and computation costs (up to 42.3% lower). Ji Liu 0003, Tianshi Che, Yang Zhou 0001, Ruoming Jin, Huaiyu Dai, Dejing Dou, Patrick Valduriez |
SDM | 7 |
| 2022 | Elastic scalable transaction processing in LeanXcaleabstractInternational audience Ricardo Jiménez-Peris, Diego Burgos-Sancho, Francisco J. Ballesteros, Marta Patiño-Martínez, Patrick Valduriez |
Inf. Syst. | 5 |
| 2021 | Parallel query processing in a polystore
Pavlos Kranas, Boyan Kolev, Oleksandra Levchenko, Esther Pacitti, Patrick Valduriez, Ricardo Jiménez-Peris, Marta Patiño-Martínez |
Distributed Parallel Databases | 5 |
| 2021 | BestNeighbor: efficient evaluation of kNN queries on large time series databases
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas, Dennis E. Shasha, Patrick Valduriez |
Knowl. Inf. Syst. | 8 |
| 2020 | Distributed Caching of Scientific Workflows in Multisite Cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
DEXA (2) | 6 |
| 2020 | Parallel computation of PDFs on big spatial data using Spark
Ji Liu 0003, Noel Moreno Lemus, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez |
Distributed Parallel Databases | 5 |
| 2019 | Pipelined Implementation of a Parallel Streaming Method for Time Series Correlation Discovery on Sliding Windows
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez |
DATA | 7 |
| 2019 | Adaptive Caching for Data-Intensive Scientific Workflows in the Cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
DEXA (2) | 6 |
| 2019 | Distributed Algorithms to Find Similar Time SeriesabstractInternational audience Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Dennis E. Shasha, Themis Palpanas, Patrick Valduriez, Reza Akbarinia, Florent Masseglia |
ECML/PKDD (3) | 6 |
| 2019 | DZI: An air index for spatial queries in one-dimensional channels
KwangJin Park, Alexis Joly, Patrick Valduriez |
Data Knowl. Eng. | 3 |
| 2019 | Efficient Scheduling of Scientific Workflows Using Hot Metadata in a Multisite CloudabstractLarge-scale, data-intensive scientific applications are often expressed as scientific workflows (SWfs). In this paper, we consider the problem of efficient scheduling of a large SWf in a multisite cloud, i.e., a cloud with geo-distributed cloud data centers (sites). The reasons for using multiple cloud sites to run a SWf are that data is already distributed, the necessary resources exceed the limits at a single site, or the monetary cost is lower. In a multisite cloud, metadata management has a critical impact on the efficiency of SWf scheduling as it provides a global view of data location and enables task tracking during execution. Thus, it should be readily available to the system at any given time. While it has been shown that efficient metadata handling plays a key role in performance, little research has targeted this issue in multisite cloud. In this paper, we propose to identify and exploit hot metadata (frequently accessed metadata) for efficient SWf scheduling in a multisite cloud, using a distributed approach. We implemented our approach within a scientific workflow management system, which shows that our approach reduces the execution time of highly parallel jobs up to 64 percent and that of the whole SWfs up to 55 percent. Ji Liu 0003, Luis Pineda-Morales, Esther Pacitti, Alexandru Costan, Patrick Valduriez, Gabriel Antoniu, Marta Mattoso |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Parallel Polyglot Query Processing on Heterogeneous Cloud Data Stores with LeanXcaleabstractThe blooming of different cloud data stores has turned polystore systems to a major topic in the nowadays cloud landscape. Especially, as the amount of processed data grows rapidly each year, much attention is being paid on taking advantage of the parallel processing capabilities of the underlying data stores. To provide data federation, a typical polystore solution defines a common data model and query language with translations to API calls or queries to each data store. However, this may lead to losing important querying capabilities. The polyglot approach of the CloudMdsQL query language allows data store native queries to be expressed as inline scripts and combined with regular SQL statements in ad-hoc integration queries. Moreover, efficient optimization techniques, such as bind join, can still take place to improve the performance of selective joins. In this paper, we introduce the distributed architecture of the LeanXcale query engine that processes polyglot queries in the CloudMdsQL query language, yet allowing native scripts to be handled in parallel at data store shards, so that efficient and scalable parallel joins take place at the query engine level. The experimental evaluation of the LeanXcale parallel query engine on various join queries illustrates well the performance benefits of exploiting the parallelism of the underlying data management technologies in combination with the high expressivity provided by their scripting/querying frameworks. Boyan Kolev, Oleksandra Levchenko, Esther Pacitti, Patrick Valduriez, Ricardo Vilaça, Rui C. Gonçalves, Ricardo Jiménez-Peris, Pavlos Kranas |
IEEE BigData | 4 |
| 2018 | Answering Top-k Queries over Outsourced Sensitive Data in the Cloud
Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez |
DEXA (1) | 3 |
| 2018 | Top-k Query Processing over Distributed Sensitive DataabstractDistributed systems provide users with powerful capabilities to store and process their data in third-party machines. However, the privacy of the outsourced data is not guaranteed. One solution for protecting the user data against privacy attacks is to encrypt the sensitive data before sending to the nodes of the distributed system. Then, the main problem is to evaluate user queries over the encrypted data. Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez |
IDEAS | 3 |
| 2018 | Point pattern search in big dataabstractConsider a set of points P in space with at least some of the pairwise distances specified. Given this set P, consider the following three kinds of queries against a database D of points : (i) pure constellation query: find all sets S in D of size |P| that exactly match the pairwise distances within P up to an additive error ϵ; (ii) isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f for which the distances between pairs in S exactly match f times the distances between corresponding pairs of P up to an additive ϵ; (iii) non-isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f and for at least some pairs of points, a maximum stretch factor mi,j > 1 such that (f X mi,jXdist(pi, pj))+ϵ > dist(si,sj) > (f X dist(pi, pj)) - ϵ. Finding matches to such queries has applications to spatial data in astronomical, seismic, and any domain in which (approximate, scale-independent) geometrical matching is required. Answering the isotropic and non-isotropic queries is challenging because scale factors and stretch factors may take any of an infinite number of values. This paper proposes practically efficient sequential and distributed algorithms for pure, isotropic, and non-isotropic constellation queries. As far as we know, this is the first work to address isotropic and non-isotropic queries. Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Alberto Krone-Martins, Patrick Valduriez, Dennis E. Shasha |
SSDBM | 5 |
| 2018 | ParCorr: efficient parallel methods to identify similar time series pairs across sliding windows
Djamel Edine Yagoubi, Reza Akbarinia, Boyan Kolev, Oleksandra Levchenko, Florent Masseglia, Patrick Valduriez, Dennis E. Shasha |
Data Min. Knowl. Discov. | 6 |
| 2018 | DfAnalyzer: Runtime Dataflow Analysis of Scientific Applications using ProvenanceabstractWe present DfAnalyzer, a tool that enables monitoring, debugging, steering, and analysis of dataflows while being generated by scientific applications. It works by capturing strategic domain data, registering provenance and execution data to enable queries at runtime. DfAnalyzer provides lightweight dataflow monitoring components to be invoked by high performance applications. It can be plugged in scientific code scripts, or Spark applications, in the same way users already plug visualization library components. During this demo, we will show how DfAnalyzer captures the dataflow, provenance, as well as how it provides runtime data analyses of applications. We will also encourage attendees to use DfAnalyzer for their own applications. Vítor Silva 0003, Daniel de Oliveira 0001, Marta Mattoso, Patrick Valduriez |
Proc. VLDB Endow. | 4 |
| 2017 | Pre-processing and Indexing Techniques for Constellation Queries in Big Data
Amir Khatibi, Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Patrick Valduriez, Dennis E. Shasha |
DaWaK | 5 |
| 2017 | Advances in Databases and Information Systems
Ladjel Bellatreche, Patrick Valduriez, Tadeusz Morzy |
Inf. Syst. | 2 |
| 2016 | Benchmarking polystores: The CloudMdsQL experienceabstractThe CloudMdsQL polystore provides integrated access to multiple heterogeneous data stores, such as RDBMS, NoSQL or even HDFS through a big data analytics framework such as MapReduce or Spark. The CloudMdsQL language is a functional SQL-like query language with a flexible nested data model. A major capability is to exploit the full power of each of the underlying data stores by allowing native queries to be expressed as functions and involved in SQL statements. The CloudMdsQL polystore has been validated with a good number of different data stores: HDFS, key-value, document, graph, RDBMS and OLAP engine. In this paper, we introduce the benchmarking of the CloudMdsQL polystore and evaluate the performance benefits of important features enabled by the query language and engine. Boyan Kolev, Raquel Pau, Oleksandra Levchenko, Patrick Valduriez, Ricardo Jiménez-Peris, José Pereira 0001 |
IEEE BigData | 4 |
| 2016 | Managing hot metadata for scientific workflows on multisite cloudsabstractLarge-scale scientific applications are often expressed as workflows that help defining data dependencies between their different components. Several such workflows have huge storage and computation requirements, and so they need to be processed in multiple (cloud-federated) datacenters. It has been shown that efficient metadata handling plays a key role in the performance of computing systems. However, most of this evidence concern only single-site, HPC systems to date. In this paper, we present a hybrid decentralized/distributed model for handling hot metadata (frequently accessed metadata) in multisite architectures. We couple our model with a scientific workflow management system (SWfMS) to validate and tune its applicability to different real-life scientific scenarios. We show that efficient management of hot metadata improves the performance of SWfMS, reducing the workflow execution time up to 50% for highly parallel jobs and avoiding unnecessary cold metadata operations. Luis Pineda-Morales, Ji Liu 0003, Alexandru Costan, Esther Pacitti, Gabriel Antoniu, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 6 |
| 2016 | Spatially Localized Visual Dictionary LearningabstractThis paper addresses the challenge of devising new representation learning algorithms that overcome the lack of interpretability of classical visual models. Therefore, it introduces a new recursive visual patch selection technique built on top of a Shared Nearest Neighbors embedding method. The main contribution of the paper is to drastically reduce the high-dimensionality of such over-complete representation thanks to a recursive feature elimination method. We show that the number of spatial atoms of the representation can be reduced by up to two orders of magnitude without much degrading the encoded information. The resulting representations are shown to provide competitive image classification performance with the state-of-the-art while enabling to learn highly interpretable visual models. Valentin Leveau, Alexis Joly, Olivier Buisson, Patrick Valduriez |
ICMR | 4 |
| 2016 | The CloudMdsQL Multistore SystemabstractThe blooming of different cloud data management infrastructures has turned multistore systems to a major topic in the nowadays cloud landscape. In this demonstration, we present a Cloud Multidatastore Query Language (CloudMdsQL), and its query engine. CloudMdsQL is a functional SQL-like language, capable of querying multiple heterogeneous data stores (relational and NoSQL) within a single query that may contain embedded invocations to each data store's native query interface. The major innovation is that a CloudMdsQL query can exploit the full power of local data stores, by simply allowing some local data store native queries (e.g. a breadth-first search query against a graph database) to be called as functions, and at the same time be optimized. Within our demonstration, we focus on two use cases each involving four diverse data stores (graph, document, relational, and key-value) with its corresponding CloudMdsQL queries. The query execution flows are visualized by an embedded real-time monitoring subsystem. The users can also try out different ad-hoc queries, not necessarily in the context of the use cases. Boyan Kolev, Carlyna Bondiombouy, Patrick Valduriez, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001 |
SIGMOD Conference | 3 |
| 2016 | CloudMdsQL: querying heterogeneous cloud data stores with a common language
Boyan Kolev, Patrick Valduriez, Carlyna Bondiombouy, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001 |
Distributed Parallel Databases | 2 |
| 2016 | FP-Hadoop: Efficient processing of skewed MapReduce jobs
Miguel Liroz-Gistau, Reza Akbarinia, Divyakant Agrawal, Patrick Valduriez |
Inf. Syst. | 4 |
| 2016 | Database System Support of Simulation DataabstractSupported by increasingly efficient HPC infra-structure, numerical simulations are rapidly expanding to fields such as oil and gas, medicine and meteorology. As simulations become more precise and cover longer periods of time, they may produce files with terabytes of data that need to be efficiently analyzed. In this paper, we investigate techniques for managing such data using an array DBMS. We take advantage of multidimensional arrays that nicely models the dimensions and variables used in numerical simulations. However, a naive approach to map simulation data files may lead to sparse arrays, impacting query response time, in particular, when the simulation uses irregular meshes to model its physical domain. We propose efficient techniques to map coordinate values in numerical simulations to evenly distributed cells in array chunks with the use of equi-depth histograms and space-filling curves. We implemented our techniques in SciDB and, through experiments over real-world data, compared them with two other approaches: row-store and column-store DBMS. The results indicate that multidimensional arrays and column-stores are much faster than a traditional row-store system for queries over a larger amount of simulation data. They also help identifying the scenarios where array DBMSs are most efficient, and those where they are outperformed by column-stores. Hermano Lustosa, Fábio Porto 0001, Patrick Valduriez |
Proc. VLDB Endow. | 4 |
| 2015 | An Efficient Solution for Processing Skewed MapReduce Jobs
Reza Akbarinia, Miguel Liroz-Gistau, Divyakant Agrawal, Patrick Valduriez |
DEXA (2) | 4 |
| 2015 | Integrating Big Data and Relational Data with a Functional SQL-like Query Language
Carlyna Bondiombouy, Boyan Kolev, Oleksandra Levchenko, Patrick Valduriez |
DEXA (1) | 4 |
| 2015 | Kernelizing Spatially Consistent Visual Matches for Fine-Grained ClassificationabstractThis paper introduces a new image representation relying on the spatial pooling of geometrically consistent visual matches. We therefore introduce a new match kernel based on the inverse rank of the shared nearest neighbors combined with local geometric constraints. To avoid overfitting and reduce processing costs, the dimensionality of the resulting over-complete representation is further reduced by hierarchically pooling the raw consistent matches according to their spatial position in the training images. The final image representation is obtained by concatenating the resulting feature vectors at several resolutions. Learning from these representations using a logistic regression classifier is shown to provide excellent fine-grained classification performances outperforming the results reported in the literature on several classification tasks. Valentin Leveau, Alexis Joly, Olivier Buisson, Patrick Valduriez |
ICMR | 4 |
| 2015 | OpenAlea: scientific workflows combining data analysis and simulationabstractAnalyzing biological data (e.g., annotating genomes, assembling NGS data...) may involve very complex and interlinked steps where several tools are combined together. Scientific workflow systems have reached a level of maturity that makes them able to support the design and execution of such in-silico experiments, and thus making them increasingly popular in the bioinformatics community. Christophe Pradal, Christian Fournier, Patrick Valduriez, Sarah Cohen Boulakia |
SSDBM | 3 |
| 2015 | Corrigendum to "Best position algorithms for efficient top-k query processing" [Inf. Syst 36(6) (2011) 973-989]
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Inf. Syst. | 3 |
| 2015 | FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed DataabstractBig data parallel frameworks, such as MapReduce or Spark have been praised for their high scalability and performance, but show poor performance in the case of data skew. There are important cases where a high percentage of processing in the reduce side ends up being done by only one node. In this demonstration, we illustrate the use of FP-Hadoop , a system that efficiently deals with data skew in MapReduce jobs. In FP-Hadoop, there is a new phase, called intermediate reduce (IR) , in which blocks of intermediate values, constructed dynamically, are processed by intermediate reduce workers in parallel, by using a scheduling strategy. Within the IR phase, even if all intermediate values belong to only one key, the main part of the reducing work can be done in parallel using the computing resources of all available workers. We implemented a prototype of FP-Hadoop, and conducted extensive experiments over synthetic and real datasets. We achieve excellent performance gains compared to native Hadoop, e.g. more than 10 times in reduce time and 5 times in total execution time. During our demonstration, we give the users the possibility to execute and compare job executions in FP-Hadoop and Hadoop. They can retrieve general information about the job and the tasks and a summary of the phases. They can also visually compare different configurations to explore the difference between the approaches. Miguel Liroz-Gistau, Reza Akbarinia, Patrick Valduriez |
Proc. VLDB Endow. | 3 |
| 2014 | Entity resolution for probabilistic data
Naser Ayat, Reza Akbarinia, Hamideh Afsarmanesh, Patrick Valduriez |
Inf. Sci. | 4 |
| 2014 | Special section on data-intensive cloud infrastructure
Ashraf Aboulnaga, Beng Chin Ooi, Patrick Valduriez |
VLDB J. | 3 |
| 2013 | Algebraic dataflows for big data analysisabstractAnalyzing big data requires the support of dataflows with many activities to extract and explore relevant information from the data. Recent approaches such as Pig Latin propose a high-level language to model such dataflows. However, the dataflow execution is typically delegated to a MapRe-duce implementation such as Hadoop, which does not follow an algebraic approach, thus it cannot take advantage of the optimization opportunities of PigLatin algebra. In this paper, we propose an approach for big data analysis based on algebraic workflows, which yields optimization and parallel execution of activities and supports user steering using provenance queries. We illustrate how a big data processing dataflow can be modeled using the algebra. Through an experimental evaluation using real datasets and the execution of the dataflow with Chiron, an engine that supports our algebra, we show that our approach yields performance gains of up to 19.6% using algebraic optimizations in the dataflow and up to 39.1% of time saved on a user steering scenario. Jonas Dias, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 5 |
| 2013 | What You Pay for Is What You Get
Ruiming Tang, Dongxu Shao, Stéphane Bressan, Patrick Valduriez |
DEXA (2) | 4 |
| 2013 | The Price Is Right - Models and Algorithms for Pricing DataabstractData is a modern commodity. Yet the pricing models in useon electronic data markets either focus on the usage of computing resources,or are proprietary, opaque, most likely ad hoc, and not conduciveof a healthy commodity market dynamics. In this paper we propose ageneric data pricing model that is based on minimal provenance, i.e. minimalsets of tuples contributing to the result of a query.We show that theproposed model fulfills desirable properties such as contribution monotonicity,bounded-price and contribution arbitrage-freedom. We presenta baseline algorithm to compute the exact price of a query based onour pricing model. We show that the problem is NP-hard. We thereforedevise, present and compare several heuristics. We conduct a comprehensiveexperimental study to show their effectiveness and efficiency. Ruiming Tang, Huayu Wu 0001, Zhifeng Bao, Stéphane Bressan, Patrick Valduriez |
DEXA (2) | 5 |
| 2013 | Entity resolution for distributed probabilistic data
Naser Ayat, Reza Akbarinia, Hamideh Afsarmanesh, Patrick Valduriez |
Distributed Parallel Databases | 4 |
| 2013 | A Hierarchical Grid Index (HGI), spatial queries in wireless data broadcasting
KwangJin Park, Patrick Valduriez |
Distributed Parallel Databases | 2 |
| 2013 | Efficient Evaluation of SUM Queries over Probabilistic DataabstractSUM queries are crucial for many applications that need to deal with uncertain data. In this paper, we are interested in the queries, called ALL_SUM, that return all possible sum values and their probabilities. In general, there is no efficient solution for the problem of evaluating ALL_SUM queries. But, for many practical applications, where aggregate values are small integers or real numbers with small precision, it is possible to develop efficient solutions. In this paper, based on a recursive approach, we propose a new solution for those applications. We implemented our solution and conducted an extensive experimental evaluation over synthetic and real-world data sets; the results show its effectiveness. Reza Akbarinia, Patrick Valduriez, Guillaume Verger |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Dynamic Workload-Based Partitioning for Large-Scale Databases
Miguel Liroz-Gistau, Reza Akbarinia, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez |
DEXA (2) | 5 |
| 2012 | Satisfaction-based query replication - An automatic and self-adaptable approach for replicating queries in the presence of autonomous participants
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
Distributed Parallel Databases | 3 |
| 2011 | Scaling Up Query Allocation in the Presence of Autonomous Participants
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Sylvie Cazalens, Patrick Valduriez |
DASFAA (2) | 4 |
| 2011 | Efficient Early Top-k Query Processing in Overloaded P2P Systems
William Kokou Dedzoe, Philippe Lamarre, Reza Akbarinia, Patrick Valduriez |
DEXA (1) | 4 |
| 2011 | Principles of Distributed Data Management in 2020?
Patrick Valduriez |
DEXA (1) | 1 |
| 2011 | Distributed data management in 2020?abstractWork on distributed data management commenced shortly after the introduction of the relational model in the mid-1970's. 1970's and 1980's were very active periods for the development of distributed relational database technology, and claims were made that in the following ten years centralized databases will be an “antique curiosity” and most organizations will move toward distributed database managers [1]. That prediction has certainly become true, and all commercial DBMSs today are distributed. M. Tamer Özsu, Patrick Valduriez, Serge Abiteboul, Bettina Kemme, Ricardo Jiménez-Peris, Beng Chin Ooi |
ICDE | 2 |
| 2011 | ParallelGDB: a parallel graph database based on cache specializationabstractThe need for managing massive attributed graphs is becoming common in many areas such as recommendation systems, proteomics analysis, social network analysis or bibliographic analysis. This is making it necessary to move towards parallel systems that allow managing graph databases containing millions of vertices and edges. Previous work on distributed graph databases has focused on finding ways to partition the graph to reduce network traffic and improve execution time. However, partitioning a graph and keeping the information regarding the location of vertices might be unrealistic for massive graphs. In this paper, we propose Parallel-GDB, a new system based on specializing the local caches of any node in this system, providing a better cache hit ratio. ParallelGDB uses a random graph partitioning, avoiding complex partition methods based on the graph topology, that usually require managing extra data structures. This proposed system provides an efficient environment for distributed graph databases. Luis Barguñó, Victor Muntés-Mulero, David Dominguez-Sal, Patrick Valduriez |
IDEAS | 4 |
| 2011 | Best position algorithms for efficient top-k query processing
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Inf. Syst. | 3 |
| 2011 | An Algebraic Approach for Data-Centric Scientific Workflows
Eduardo S. Ogasawara, Daniel de Oliveira 0001, Patrick Valduriez, Jonas Dias, Fábio Porto 0001, Marta Mattoso |
Proc. VLDB Endow. | 3 |
| 2011 | Energy Efficient Data Access in Mobile P2P NetworksabstractA fundamental problem for peer-to-peer (P2P) applications in mobile-pervasive computing environment is to efficiently identify the node that stores particular data items and download them while preserving battery power. In this paper, we propose a P2P Minimum Boundary Rectangle (PMBR, for short) which is a new spatial index specifically designed for mobile P2P environments. A node that contains desirable data item (s) can be easily identified by reading the PMBR index. Then, we propose a selective tuning algorithm, called Distributed exponential Sequence Scheme (DSS, for short), that provides clients with the ability of selective tuning of data items, thus preserving the scarce power resource. The proposed algorithm is simple but efficient in supporting linear transmission of spatial data and processing of location-aware queries. The results from theoretical analysis and experiments show that the proposed algorithm with the PMBR index is scalable and energy efficient in both range queries and nearest neighbor queries. KwangJin Park, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | SbQA: A Self-Adaptable Query Allocation ProcessabstractWe present a flexible query allocation framework, called satisfaction-based query allocation (SbQA for short), for distributed information systems where both consumers and providers (the participants) have special interests towards queries. A particularity of SbQA is that it allocates queries while considering both query load and participants' interests. To be fair, it dynamically trades consumers' interests for providers' interests based on their satisfaction. In this demo we illustrate the flexibility and efficiency of SbQA to allocate queries on the Berkeley Open Infrastructure for Network Computing (BOINC). We also demonstrate that SbQA is self-adaptable to the participants' expectations. Finally, we demonstrate that SbQA can be adapted to different kinds of applications by varying its parameters. Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
ICDE | 3 |
| 2009 | Parallel OLAP query processing in database clusters with data replication
Alexandre A. B. Lima, Camille Furtado, Patrick Valduriez, Marta Mattoso |
Distributed Parallel Databases | 3 |
| 2009 | DHTJoin: processing continuous join queries using DHT networks
Wenceslao Palma, Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Distributed Parallel Databases | 4 |
| 2009 | A self-adaptable query allocation framework for distributed information systems
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
VLDB J. | 3 |
| 2008 | Efficient Processing of Nearest Neighbor Queries in Parallel Multimedia Databases
Jorge R. Manjarrez Sanchez, José Martinez 0001, Patrick Valduriez |
DEXA | 3 |
| 2008 | Summary management in P2P systemsabstractSharing huge, massively distributed databases in P2P systems is inherently difficult. As the amount of stored data increases, data localization techniques become no longer sufficient. A practical approach is to rely on compact database summaries rather than raw database records, whose access is costly in large P2P systems. In this paper, we consider summaries that are synthetic, multidimensional views with two main virtues. First, they can be directly queried and used to approximately answer a query without exploring the original data. Second, as semantic indexes, they support locating relevant nodes based on data content. Our main contribution is to define a summary model for P2P systems, and the appropriate algorithms for summary management. Our performance evaluation shows that the cost of query routing is minimized, while incurring a low cost of summary maintenance. Rabab Hayek, Guillaume Raschia, Patrick Valduriez, Noureddine Mouaddib |
EDBT | 3 |
| 2008 | Improving Interoperability Using Query Interpretation in Semantic Vector Spaces
Anthony Ventresque, Sylvie Cazalens, Philippe Lamarre, Patrick Valduriez |
ESWC | 4 |
| 2008 | Mobile continuous nearest neighbor queries on airabstractIn this paper, we propose new spatial query processing algorithms to support Mobile Continuous Nearest Neighbor Query (MCNNQ) in wireless broadcast environments. Each client (object) computes individually the function that best captures the locations of the moving objects, while a server is only responsible for delivering information gathered from the clients. Extensive experiments demonstrate that our location-based data dissemination algorithm significantly out-performs index-based solutions. KwangJin Park, Patrick Valduriez, Hyunseung Choo |
GIS | 2 |
| 2008 | P2P logging and timestamping for reconciliationabstractIn this paper, we address data reconciliation in peer-to-peer (P2P) collaborative applications. We propose P2P-LTR (Logging and Timestamping for Reconciliation) which provides P2P logging and timestamping services for P2P reconciliation over a distributed hash table (DHT). While updating at collaborating peers, updates are timestamped and stored in a highly available P2P log. During reconciliation, these updates are retrieved in total order to enforce eventual consistency. In this paper, we first give an overview of P2P-LTR with its model and its main procedures. We then present our prototype used to validate P2P-LTR. To demonstrate P2P-LTR, we propose several scenarios that test our solutions and measure performance. In particular, we demonstrate how P2P-LTR handles the dynamic behavior of peers with respect to the DHT. Mounir Tlili, William Kokou Dedzoe, Esther Pacitti, Patrick Valduriez, Reza Akbarinia, Pascal Molli, Gérôme Canals, Stéphane Laurière |
Proc. VLDB Endow. | 4 |
| 2008 | DiSC: Benchmarking Secure Chip DBMSabstractAbstract—Secure chips, e.g., present in smart cards, USB dongles, i-buttons, are now ubiquitous in applications with strong security requirements. Moreover, they require embedded data management techniques. However, secure chips have severe hardware constraints, which make traditional database techniques irrelevant. The main problem faced by secure chip DBMS designers is to be able to assess various design choices and trade-offs for different applications. Our solution is to use a benchmark for secure chip DBMS in order to 1) compare different database techniques, 2) predict the limits of on-chip applications, and 3) provide codesign hints. In this paper, we propose Data management in Secure Chip (DiSC), a benchmark that reaches these three objectives. This work benefits from our long experience in developing and tuning data management techniques for the smart card. To validate DiSC, we compare the behavior of candidate data management techniques using a cycle-accurate smart-card simulator. Furthermore, we show the applicability of DiSC to future designs involving new hardware platforms and new database techniques. Nicolas Anciaux, Luc Bouganim, Philippe Pucheral, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2007 | Satisfaction balanced mediationabstractWe consider a distributed information system that allows autonomous consumers to query autonomous providers. We focus on the problem of query allocation from a new point of view, by considering consumers and providers' satisfaction in addition to query load. We define satisfaction as a long-run notion based on the consumers and providers' preferences. We propose and validate a mediation process, called SBMediation, which is compared to Capacity based query allocation. The experimental results show that SBMediation significantly outperforms Capacity based when confronted to autonomous participants. Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Sylvie Cazalens, Patrick Valduriez |
CIKM | 4 |
| 2007 | KnBest - A Balanced Request Allocation Method for Distributed Information Systems
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
DASFAA | 3 |
| 2007 | Data currency in replicated DHTsabstractDistributed Hash Tables (DHTs) provide a scalable solution for data sharing in P2P systems. To ensure high data availability, DHTs typically rely on data replication, yet without data currency guarantees. Supporting data currency in replicated DHTs is difficult as it requires the ability to return a current replica despite peers leaving the network or concurrent updates. In this paper, we give a complete solution to this problem. We propose an Update Management Service (UMS) to deal with data availability and efficient retrieval of current replicas based on timestamping. For generating timestamps, we propose a Key-based Timestamping Service (KTS) which performs distributed timestamp generation using local counters. Through probabilistic analysis, we compute the expected number of replicas which UMS must retrieve for finding a current replica. Except for the cases where the availability of current replicas is very low, the expected number of retrieved replicas is typically small, e.g. if at least 35% of available replicas are current then the expected number of retrieved replicas is less than 3. We validated our solution through implementation and experimentation over a 64-node cluster and evaluated its scalability through simulation up to 10,000 peers using SimJava. The results show the effectiveness of our solution. They also show that our algorithm used in UMS achieves major performance gains, in terms of response time and communication cost, compared with a baseline algorithm. Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
SIGMOD Conference | 3 |
| 2007 | Best Position Algorithms for Top-k Queries
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
VLDB | 3 |
| 2007 | SQLB: A Query Allocation Framework for Autonomous Consumers and Providers
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
VLDB | 3 |
| 2007 | A Flexible Mediation Process for Large Distributed Information SystemsabstractWe consider distributed information systems that are open, dynamic and provide access to large numbers of distributed, heterogeneous, autonomous information sources. Most of the work in data mediator systems has dealt with the problem of finding relevant information providers for a request. However, finding relevant requests for information providers is another important side of the mediation problem which has not received much attention. In this paper, we address these two sides of the problem with a flexible mediation process. Once the qualified information providers are identified, our process allows them to express their interest in a request via a bidding mechanism. It also requires to set up a requisition policy, because a request must always be answered if there are qualified providers. This work does not concern pure market mechanisms because we counter-balance the providers' bids by considering their quality wrt a request. We validate our process on a set of simulations in the context of load balancing, which is a good indicator of the system's overall performance. The results show that the mediation process provides a very good long-run regulation of the system, in particular when providers can leave the system. However, load balancing is not the natural application of the flexible mediation and additional testing is required to show the generality of the approach to non-depletable resources. Philippe Lamarre, Sandra Lemp, Sylvie Cazalens, Patrick Valduriez |
Int. J. Cooperative Inf. Syst. | 4 |
| 2007 | The leganet system: Freshness-aware transaction routing in a database cluster
Stéphane Gançarski, Hubert Naacke, Esther Pacitti, Patrick Valduriez |
Inf. Syst. | 4 |
| 2006 | Reducing network traffic in unstructured P2P systems using Top-k queries
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Distributed Parallel Databases | 3 |
| 2005 | Preventive Replication in a Database Cluster
Esther Pacitti, Cédric Coulon, Patrick Valduriez, M. Tamer Özsu |
Distributed Parallel Databases | 3 |
| 2001 | Processing Queries with Expensive Functions and Large Objects in Distributed Mediator SystemsabstractLeSelect is a mediator system which allows scientists to publish their resources (data and programs) so they can be transparently accessed. The scientists can typically issue queries which access distributed published data and involve the execution of expensive functions (corresponding to programs). Furthermore, the queries can involve large objects, such as images (e.g. archived meteorological satellite data). In this context, the costs of transmitting large objects and invoking expensive functions are the dominant factors of execution time. In this paper, we first propose three query execution techniques which minimize these costs by taking full advantage of the distributed architecture of mediator systems like LeSelect. Then we devise parallel processing strategies for queries including expensive functions. Based on experimentation, we show that it is hard to predict the optimal execution order when dealing with several functions. We propose a new hybrid parallel technique to solve this problem and give some experimental results. Luc Bouganim, Françoise Fabret, Fábio Porto 0001, Patrick Valduriez |
ICDE | 4 |
| 2001 | PicoDBMS: Validation and Experience
Nicolas Anciaux, Christophe Bobineau, Luc Bouganim, Philippe Pucheral, Patrick Valduriez |
VLDB | 5 |
| 2001 | User-Optimizer Communication using Abstract Plans in Sybase ASE
Mihnea Andrei, Patrick Valduriez |
VLDB | 2 |
| 2001 | PicoDBMS: Scaling down database techniques for the smartcard
Philippe Pucheral, Luc Bouganim, Patrick Valduriez, Christophe Bobineau |
VLDB J. | 3 |
| 2000 | Dynamic Query Scheduling in Data Integration SystemsabstractExecution plans produced by traditional query optimizers for data integration queries may yield poor performance for several reasons. The cost estimates may be inaccurate, the memory available at run-time may be insufficient, or data delivery rate can be unpredictable. We address the problem of unpredictable data arrival rate. We propose to dynamically schedule queries in order to deal with irregular data delivery rate and gracefully adapt to the available memory. Our approach performs careful step-by-step scheduling of several query fragments and processes these fragments based on data arrivals. We describe a performance evaluation that shows important performance gains in several configurations. Luc Bouganim, Françoise Fabret, C. Mohan 0001, Patrick Valduriez |
ICDE | 4 |
| 2000 | PicoDMBS: Scaling Down Database Techniques for the Smartcard
Christophe Bobineau, Luc Bouganim, Philippe Pucheral, Patrick Valduriez |
VLDB | 4 |
| 2000 | Caching Strategies for Data-Intensive Web Sites
Khaled Yagoub, Daniela Florescu, Valérie Issarny, Patrick Valduriez |
VLDB | 4 |
| 2000 | Building and Customizing Data-Intensive Web Sites Using Weave
Khaled Yagoub, Daniela Florescu, Valérie Issarny, Patrick Valduriez |
VLDB | 4 |
| 1999 | Load Balancing for Parallel Query Execution on NUMA Multiprocessors
Luc Bouganim, Daniela Florescu, Patrick Valduriez |
Distributed Parallel Databases | 3 |
| 1998 | Memory-Adaptive Scheduling for Large Query ExecutionabstractArticle Free Access Share on Memory-adaptive scheduling for large query execution Authors: Luc Bouganim PRiSM, Versailles, France PRiSM, Versailles, FranceView Profile , Olga Kapitskaia INRIA, Rocquencourt, France INRIA, Rocquencourt, FranceView Profile , Patrick Valduriez INRIA, Rocquencourt, France INRIA, Rocquencourt, FranceView Profile Authors Info & Claims CIKM '98: Proceedings of the seventh international conference on Information and knowledge managementNovember 1998 Pages 105–115https://doi.org/10.1145/288627.288646Online:01 November 1998Publication History 13citation347DownloadsMetricsTotal Citations13Total Downloads347Last 12 Months9Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Luc Bouganim, Olga Kapitskaia, Patrick Valduriez |
CIKM | 3 |
| 1998 | Scaling Access to Heterogeneous Data Sources with DISCOabstractAccessing many data sources aggravates problems for users of heterogeneous distributed databases. Database administrators must deal with fragile mediators, that is, mediators with schemas and views that must be significantly changed to incorporate a new data source. When implementing translators of queries from mediators to data sources, database implementers must deal with data sources that do not support all the functionality required by mediators. Application programmers must deal with graceless failures for unavailable data sources. Queries simply return failure and no further information when data sources are unavailable for query processing. The Distributed Information Search COmponent (Disco) addresses these problems. Data modeling techniques manage the connections to data sources, and sources can be added transparently to the users and applications. The interface between mediators and data sources flexibly handles different query languages and different data source functionality. Query rewriting and optimization techniques rewrite queries so they are efficiently evaluated by sources. Query processing and evaluation semantics are developed to process queries over unavailable data sources. In this article, we describe: 1) the distributed mediator architecture of Disco; 2) the data model and its modeling of data source connections; 3) the interface to underlying data sources and the query rewriting process; and 4) query processing semantics. We describe several advantages of our system. Anthony Tomasic, Louiqa Raschid, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1997 | Concurrent Garbage Collection in O2
Marcin Skubiszewski, Patrick Valduriez |
VLDB | 2 |
| 1996 | Adaptive Parallel Query Execution in DBS3
Luc Bouganim, Benoît Dageville, Patrick Valduriez |
EDBT | 3 |
| 1996 | Dynamic Load Balancing in Hierarchical Parallel Database Systems
Luc Bouganim, Daniela Florescu, Patrick Valduriez |
VLDB | 3 |
| 1996 | A Methodology for Query Reformulation in CIS Using Semantic KnowledgeabstractWe consider Cooperative Information Systems (CIS) that are multidatabase systems (MDBMS), with a common object-oriented model, based on the ODMG standard, together with local databases that may be relational, object-oriented, or dedicated data servers. The MDBMS interface (or mediator interface) that describes this CIS could be different from the union of the local interfaces that describe each local database. In particular, the mediator interface may be defined by semantic knowledge that includes views over particular local databases, integrity constraints, and knowledge about data replication in local databases. We present a methodology for query reformulation which is based on the uniform representation of all semantic knowledge in the form of integrity assertions and mapping rules. A reformulation algorithm exploits this semantic knowledge, and performs semantic rewriting based on pattern-matching, to obtain a query on the union of the local interfaces. A decomposition algorithm then produces a composite query, and local sub-queries, one for each local interface. The reformulation is general enough to re-use the results of previously computed queries in the CIS. We have implemented this reformulation technique in our Flora compiler prototype which we used for validation and experimentation with O2 databases. Daniela Florescu, Louiqa Raschid, Patrick Valduriez |
Int. J. Cooperative Inf. Syst. | 3 |
| 1995 | Locking in OODBMS Client Supported Nestd TransactionsabstractNested transactions facilitate the control of complex persistent applications by enabling both fine-tuning of the scope of rollback and safe intra-transaction parallelism. We are concerned with supporting concurrent nested transactions on client workstations of an OODBMS. Use of the traditional design and implementation of a lock manager results in a high CPU overhead: in-cache traversals of the 007 benchmark perform, at best, 4.5 times slower than the same traversal achieved in virtual memory by a nonpersistent programming language. We propose a new design and implementation of a lock manager which cuts that factor down to 1.8. This lock manager supports nested transactions with both sibling and parent/child parallelisms, and provides object locking at a cost comparable to page locking. Object locking is therefore a better alternative due to its higher functionality.> Laurent Daynès, Olivier Gruber, Patrick Valduriez |
ICDE | 3 |
| 1995 | Design and Implementation of Flora, A Language for Object Algebra
Daniela Florescu, Jean-Robert Gruser, Michael Novak, Patrick Valduriez, Mikal Ziane |
Inf. Sci. | 4 |
| 1995 | Transaction Chopping: Algorithms and Performance StudiesabstractChopping transactions into pieces is good for performance but may lead to nonserializable executions. Many researchers have reacted to this fact by either inventing new concurrency-control mechanisms, weakening serializability, or both. We adopt a different approach. We assume a user who —has access only to user-level tools such as (1) choosing isolation degrees 1ndash;4, (2) the ability to execute a portion of a transaction using multiversion read consistency, and (3) the ability to reorder the instructions in transaction programs; and —knows the set of transactions that may run during a certain interval (users are likely to have such knowledge for on-line or real-time transactional applications). Given this information, our algorithm finds the finest chopping of a set of transactions TranSet with the following property: If the pieces of the chopping execute serializably, then TranSet executes serializably . This permits users to obtain more concurrency while preserving correctness. Besides obtaining more intertransaction concurrency, chopping transactions in this way can enhance intratransaction parallelism. The algorithm is inexpensive, running in O(n×(e+m)) time, once conflicts are identified, using a naive implementation, where n is the number of concurrent transactions in the interval, e is the number of edges in the conflict graph among the transactions, and m is the maximum number of accesses of any transaction. This makes it feasible to add as a tuning knob to real systems. Dennis E. Shasha, François Llirbat, Eric Simon, Patrick Valduriez |
ACM Trans. Database Syst. | 4 |
| 1994 | Flora: A Functional-Style Language for Object and relational Algebra
Michael Novak, Georges Gardarin, Patrick Valduriez |
DEXA | 3 |
| 1994 | Invited Project Review: Industrial-strength parallel query optimization: issues and lessons
Rosana S. G. Lanzelotte, Patrick Valduriez, Mohamed Zaït, Mikal Ziane |
Inf. Syst. | 2 |
| 1993 | Parallel Database Systems: the case for shared-somethingabstractParallel database systems are becoming the primary application of multiprocessor computers. The reason for this is that they can provide high-performance and high-availability database support at a much lower price than do equivalent mainframe computers. The traditional shared-memory, shared-disk, and shared-nothing architectures of parallel database systems are compared, based on the following dimensions: simplicity, cost, performance, availability and extensibility. Based on these comparisons, the case is made for the shared-something architecture, which can provide a better trade-off between the various objectives.> Patrick Valduriez |
ICDE | 1 |
| 1993 | On the Effectiveness of Optimization Search Strategies for Parallel Execution Spaces
Rosana S. G. Lanzelotte, Patrick Valduriez, Mohamed Zaït |
VLDB | 2 |
| 1993 | Parallel Database Systems: Open Problems and New Issues
Patrick Valduriez |
Distributed Parallel Databases | 1 |
| 1992 | Schema Extensions in OODB
Marie-Jo Bellosta, Patrick Valduriez, Fabienne Viallet |
DEXA | 2 |
| 1992 | ESQL2: An Object-Oriented SQL with F-Logic SemanticsabstractESQL2 is an SQL2 upward-compatible database language that integrates the essential concepts of relational, object-oriented, and deductive databases. ESQL2's salient features are a rich and extendible type system based on abstract data types (ADTs) implemented in various programming languages, complex objects with object sharing by combining generic ADTs and object identity, the capability of querying and updating relations containing simple or complex objects using SQL-compatible syntax and semantics, and a DATALOG-like deductive capability provided as an extension of the SQL view mechanism. A declarative semantics is proposed for ESQL2 retrieval statements using F-Logic, which provides a solid basis for understanding the integration of objects and relations.> Georges Gardarin, Patrick Valduriez |
ICDE | 2 |
| 1992 | Optimization of Object-Oriented Recursive Queries using Cost-Controlled StrategiesabstractObject-oriented data models are being extended with recursion to gain expressive power. This complicates the optimization problem which has to deal with recursive queries on complex objects. Because unary operations invoking methods or path expressions on objects may be costly to execute, traditional heuristics for optimizing recursive queries are no longer valid. In this paper we propose a cost-based optimization method which handles object-oriented recursive queries. In particular, it is able to delay the decision of pushing selective operations through recursion until the effect of such a transformation can be measured by a cost model. The approach integrates rewriting and increases the optimization opportunities for recursive queries on objects while allowing for efficient optimization. Rosana S. G. Lanzelotte, Patrick Valduriez, Mohamed Zaït |
SIGMOD Conference | 2 |
| 1992 | Simple Rational Guidance for Chopping Up TransactionsabstractChopping transactions into pieces is good for performance but may lead to non-serializable executions. Many researchers have reacted to this fact by either inventing new concurrency control mechanisms, weakening serializability, or both. We adopt a different approach. Dennis E. Shasha, Eric Simon, Patrick Valduriez |
SIGMOD Conference | 3 |
| 1992 | SVP: A Model Capturing Sets, Lists, Streams, and Parallelism
Douglas Stott Parker Jr., Eric Simon, Patrick Valduriez |
VLDB | 3 |
| 1992 | Object-Oriented Database Systems
Patrick Valduriez |
VLDB | 1 |
| 1992 | The data model of FAD, a database programming language
Scott Danforth, Patrick Valduriez |
Inf. Sci. | 2 |
| 1992 | Functional SOL (FSOL), an SQL upward-compatible database programming language
Patrick Valduriez, Scott Danforth |
Inf. Sci. | 1 |
| 1992 | A FAD for Data Intensive ApplicationsabstractFAD is a strongly typed database programming language designed for uniformly manipulating transient and persistent data on Bubba, a parallel database system developed at MCC. The paper provides an overall description of FAD, and discusses the design rationale behind a number of its distinguishing features. Comparisons with other database programming languages are provided.> Stan Danforth, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1991 | Extending the Search Strategy in a Query Optimizer
Rosana S. G. Lanzelotte, Patrick Valduriez |
VLDB | 2 |
| 1990 | ESQL: An Extended SQL with Object and Deductive Capabilities
Georges Gardarin, Patrick Valduriez |
DEXA | 2 |
| 1990 | Efficient Main Memory Data Management Using the DBGraph Storage Model
Philippe Pucheral, Jean-Marc Thévenin, Patrick Valduriez |
VLDB | 3 |
| 1990 | Prototyping Bubba, A Highly Parallel Database SystemabstractBubba is a highly parallel computer system for data-intensive applications. The basis of the Bubba design is a scalable shared-nothing architecture which can scale up to thousands of nodes. Data are declustered across the nodes (i.e. horizontally partitioned via hashing or range partitioning) and operations are executed at those nodes containing relevant data. In this way, parallelism can be exploited within individual transactions as well as among multiple concurrent transactions to improve throughput and response times for data-intensive applications. The current Bubba prototype runs on a commercial 40-node multicomputer and includes a parallelizing compiler, distributed transaction management, object management, and a customized version of Unix. The current prototype is described and the major design decisions that went into its construction are discussed. The lessons learned from this prototype and its predecessors are presented.> Haran Boral, William Alexander, Larry Clay, George P. Copeland, Scott Danforth, Michael J. Franklin, Brian E. Hart, Marc G. Smith, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 9 |
| 1988 | Parallel Query Processing for Complex ObjectsabstractThe authors investigate a direct storage scheme for complex objects called FIHSM (Fully Inverted Hierarchical Storage Model) and propose a novel parallel-query-processing strategy (QPS) for it. The QPS has four phases (select, pivot, value materialize and compose). With a declustered placement strategy, each of these phases provides for both inter and intra-operation parallelism. Furthermore, partial results of one phase could be pipelined to the subsequent phase. The proposed four-phase structured algorithm is based on heuristics and thus avoids the prohibitive exhaustive searches which are needed for optimizing query executions in parallel environments.> Setrag Khoshafian, Patrick Valduriez, George P. Copeland |
ICDE | 2 |
| 1987 | A Query Processing Strategy for the Decomposed Storage ModelabstractHandling parallelism in database systems involves the specification of a storage model, a placement strategy, and a query processing strategy. An important goal is to determine the appropriate combination of these three strategies in order to obtain the best performance advantages. In this paper we present a novel and promising query processing strategy for a decomposed storage model. We discuss some of the qualitative advantages of the scheme. We also compare the performance of the proposed “pivot” strategy with conventional query processing for the n-ary storage model. The comparison is performed using the Wisconsin Benchmarks. Setrag Khoshafian, George P. Copeland, Thomas Jagodis, Haran Boral, Patrick Valduriez |
ICDE | 5 |
| 1987 | FAD, a Powerful and Simple Database Language
François Bancilhon, Ted Briggs, Setrag Khoshafian, Patrick Valduriez |
VLDB | 4 |
| 1987 | Join IndicesabstractIn new application areas of relational database systems, such as artificial intelligence, the join operator is used more extensively than in conventional applications. In this paper, we propose a simple data structure, called a join index, for improving the performance of joins in the context of complex queries. For most of the joins, updates to join indices incur very little overhead. Some properties of a join index are (i) its efficient use of memory and adaptiveness to parallel execution, (ii) its compatibility with other operations (including select and union), (iii) its support for abstract data type join predicates, (iv) its support for multirelation clustering, and (v) its use in representing directed graphs and in evaluating recursive queries. Finally, the analysis of the join algorithm using join indices shows its excellent performance. Patrick Valduriez |
ACM Trans. Database Syst. | 1 |
| 1986 | Buffering Schemes for Permanent DataabstractThe availability of larger RAM spaces for DBMSs provides interesting opportunities for performance enhancements, especially in buffer management. In this paper we propose and compare two alternative strategies for the buffer management of permanent data (i.e., the data committed by transactions) called block buffering and attribute buffering. These strategies use statistics to capture the changing locality of a reference string. We model and demonstrate the impact of locality on the performance of buffering. We also analyze and compare the effect of both the attribute and predicate dimensions of locality on buffering, varying a number of parameters including the degree of locality, RAM size, and RAM utilization. George P. Copeland, Setrag Khoshafian, Marc G. Smith, Patrick Valduriez |
ICDE | 4 |
| 1986 | Implementation Techniques of Complex Objects
Patrick Valduriez, Setrag Khoshafian, George P. Copeland |
VLDB | 1 |
| 1984 | Predicate Trees: An Approach to Optimize Relational Query OperationsabstractWith the advent of relational database systems, multi-key searching problems have became the focus of a great deal of research. In this paper, we present a new data structure called predicate trees for clustering tuples of a relation in a way that allows the system to accelerate a large number of multi-dimensional queries. The directories used for the implementation of predicate trees in SABRE are organized as relations which are searched efficiently by filters. One of the most significant advantages of predicate trees is the possibility of defining logical addresses based on content, called signatures, and to use filters to manage directories. Georges Gardarin, Patrick Valduriez, Yann Viémont |
ICDE | 2 |
| 1984 | Design and Implementation of an Extendible Integrity SubsystemabstractTnls paper presents a powerful integrrty subsystem, which is implemented in the SABRE database system. The specification language is simple. Tne enforcement algorithm is general, in particular, it handles referential dependency and temporal assertions. Specialized strategies efficiently treat each class of assertions. The system automatically manages integrity checkpoints. Also, an efficient method is described for processing assertions involving aggregates. An analysis exhibits the value of the algorithms. It 1s shown that, in general, this method is better than the query modification method for domain assertions. Measures have also been done for giving the cost added for controlling integrity in comparison with the cost of the request itself. Eric Simon, Patrick Valduriez |
SIGMOD Conference | 2 |
| 1984 | A Multikey Hashing Scheme Using Predicate TreesabstractA new method for multikey access suitable for dynamic files is proposed that transforms multiple key values into a logical address This method is based on a new structure, called predicate tree, that represents the function applied to several keys A predicate tree permits to specify in a unified way various hashing schemes by allowing for different definitions of predicates A logical address qualifies a space partition of a file according to its predicate tree This address is seen as a single key by a digital hashing method which transforms it into a physical address This method is used to address records in a file and to transform a retrieval qualification on a file into a set of partitions to access Finally, a qualitative analysis of the behavior of the method is given which exhibits its value Patrick Valduriez, Yann Viémont |
SIGMOD Conference | 1 |
| 1984 | Join and Semijoin Algorithms for a Multiprocessor Database MachineabstractThis paper presents and analyzes algorithms for computing joins and semijoins of relations in a multiprocessor database machine. First, a model of the multiprocessor architecture is described, incorporating parameters defining I/O, CPU, and message transmission times that permit calculation of the execution times of these algorithms. Then, three join algorithms are presented and compared. It is shown that, for a given configuration, each algorithm has an application domain defined by the characteristics of the operand and result relations. Since a semijoin operator is useful for decreasing I/O and transmission times in a multiprocessor system, we present and compare two equi-semijoin algorithms and one non-equi-semijoin algorithm. The execution times of these algorithms are generally linearly proportional to the size of the operand and result relations, and inversely proportional to the number of processors. We then compare a method which consists of joining two relations to a method whereby one joins their semijoins. Finally, it is shown that the latter method, using semijoins, is generally better. The various algorithms presented are implemented in the SABRE database system; an evaluation model selects the best algorithm for performing a join according to the results presented here. A first version of the SABRE system is currently operational at INRIA. Patrick Valduriez, Georges Gardarin |
ACM Trans. Database Syst. | 1 |
| 1982 | Semi-Join Algorithms for Multiprocessor SystemsabstractSemi-join is a relational operator that decreases the cost of processing queries involving binary operations. This is accomplished by initially selecting the data relevant to answer the queries and thereby reducing the size of the operand relations. This paper presents and analyzes algorithms for computing semi-joins in a multiprocessor database machine. First, an architecture model of a multiprocessor system is described. The model incorporates 1-0, CPU and messages transmission cost parameters to enable the evaluation of these algorithms in terms of their execution costs. Then two equi-semi-join algorithms are presented and compared and one inequi-semi-join algorithm is proposed. The execution cost of these algorithms are generally lineary proportional to the size of the operand and result relations and inversely proportional to the number of processors. Then, the method by joining two relations and the method by joining their semi-joins are compared. Finally it is shown that the method using semi-joins is generally better. Patrick Valduriez |
SIGMOD Conference | 1 |