VLDB 2026 Research / reviewers in the wild / expert
Patrick Valduriez
dblp:v/PatrickValduriez
· DBLP profile ↗
165ranked-venue papers
11as first author
13since 2021 · last 2025
0000-0001-6506-7538ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 117 · 10 first-author · 5 since 2021Systems, architecture and hardware · 28 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Software engineering, systems software and programming languages · 7Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 3Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Federated Learning with Heterogeneous Data and Adaptive DropoutabstractFederated Learning (FL) is a promising distributed machine learning approach that enables collaborative training of a global model using multiple edge devices. The data distributed among the edge devices are highly heterogeneous. Thus, FL faces the challenge of data distribution and heterogeneity, where non-Independent and Identically Distributed (non-IID) data across edge devices may yield in significant accuracy drop. Furthermore, the limited computation and communication capabilities of edge devices increase the likelihood of stragglers, thus leading to slow model convergence. In this article, we propose the FedDHAD FL framework, which comes with two novel methods: dynamic heterogeneous model aggregation (FedDH) and adaptive dropout (FedAD). FedDH dynamically adjusts the weights of each local model within the model aggregation process based on the non-IID degree of heterogeneous data to deal with the statistical data heterogeneity. FedAD performs neuron-adaptive operations in response to heterogeneous devices to improve accuracy while achieving superb efficiency. The combination of these two methods makes FedDHAD significantly outperform state-of-the-art solutions in terms of accuracy (up to 6.7% higher), efficiency (up to 2.02 times faster), and computation cost (up to 15.0% smaller). Ji Liu 0003, Beichen Ma, Qiaolin Yu, Ruoming Jin, Jingbo Zhou 0003, Yang Zhou 0001, Huaiyu Dai, Haixun Wang, Dejing Dou, Patrick Valduriez |
ACM Trans. Knowl. Discov. Data | 10 |
| 2024 | Fisher Information-based Efficient Curriculum Federated Learning with Large Language ModelsabstractAs a promising paradigm to collaboratively train models with decentralized data, Federated Learning (FL) can be exploited to fine-tune Large Language Models (LLMs).While LLMs correspond to huge size, the scale of the training data significantly increases, which leads to tremendous amounts of computation and communication costs.The training data is generally non-Independent and Identically Distributed (non-IID), which requires adaptive data processing within each device.Although Low-Rank Adaptation (LoRA) can significantly reduce the scale of parameters to update in the fine-tuning process, it still takes unaffordable time to transfer the low-rank parameters of all the layers in LLMs.In this paper, we propose a Fisher Information-based Efficient Curriculum Federated Learning framework (FibecFed) with two novel methods, i.e., adaptive federated curriculum learning and efficient sparse parameter update.First, we propose a fisher informationbased method to adaptively sample data within each device to improve the effectiveness of the FL fine-tuning process.Second, we dynamically select the proper layers for global aggregation and sparse parameters for local update with LoRA so as to improve the efficiency of the FL fine-tuning process.Extensive experimental results based on 10 datasets demonstrate that FibecFed yields excellent performance (up to 45.35% in terms of accuracy) and superb fine-tuning speed (up to 98.61% faster) compared with 17 baseline approaches).Our code will be publicly available. Ji Liu 0003, Jiaxiang Ren 0001, Ruoming Jin, Zijie Zhang 0001, Yang Zhou 0001, Patrick Valduriez, Dejing Dou |
EMNLP | 6 |
| 2024 | AEDFL: Efficient Asynchronous Decentralized Federated Learning with Heterogeneous DevicesabstractFederated Learning (FL) has achieved significant achievements recently, enabling collaborative model training on distributed data over edge devices. Iterative gradient or model exchanges between devices and the centralized server in the standard FL paradigm suffer from severe efficiency bottlenecks on the server. While enabling collaborative training without a central server, existing decentralized FL approaches either focus on the synchronous mechanism that deteriorates FL convergence or ignore device staleness with an asynchronous mechanism, resulting in inferior FL accuracy. In this paper, we propose an Asynchronous Efficient Decentralized FL framework, i.e., AEDFL, in heterogeneous environments with three unique contributions. First, we propose an asynchronous FL system model with an efficient model aggregation method for improving the FL convergence. Second, we propose a dynamic staleness-aware model update approach to achieve superior accuracy. Third, we propose an adaptive sparse training method to reduce communication and computation costs without significant accuracy degradation. Extensive experimentation on four public datasets and four models demonstrates the strength of AEDFL in terms of accuracy (up to 16.3% higher), efficiency (up to 92.9% faster), and computation costs (up to 42.3% lower). Ji Liu 0003, Tianshi Che, Yang Zhou 0001, Ruoming Jin, Huaiyu Dai, Dejing Dou, Patrick Valduriez |
SDM | 7 |
| 2023 | ProvLight: Efficient Workflow Provenance Capture on the Edge-to-Cloud ContinuumabstractModern scientific workflows require hybrid infrastructures combining numerous decentralized resources on the IoT/Edge interconnected to Cloud/HPC systems (aka the Computing Continuum) to enable their optimized execution. Understanding and optimizing the performance of such complex Edge-to-Cloud workflows is challenging. Capturing the provenance of key performance indicators, with their related data and processes, may assist in understanding and optimizing workflow executions. However, the capture overhead can be prohibitive, particularly in resource-constrained devices, such as the ones on the IoT/Edge.To address this challenge, based on a performance analysis of existing systems, we propose ProvLight, a tool to enable efficient provenance capture on the IoT/Edge. We leverage simplified data models, data compression and grouping, and lightweight transmission protocols to reduce overheads. We further integrate ProvLight into the E2Clab framework to enable workflow provenance capture across the Edge-to-Cloud Continuum. This integration makes E2Clab a promising platform for the performance optimization of applications through reproducible experiments.We validate ProvLight at a large scale with synthetic workloads on 64 real-life IoT/Edge devices in the FIT IoT LAB testbed. Evaluations show that ProvLight outperforms state-of-the-art systems like ProvLake and DfAnalyzer in resource-constrained devices. ProvLight is 26—37x faster to capture and transmit provenance data; uses 5—7x less CPU; 2x less memory; transmits 2x less data; and consumes 2—2.5x less energy. ProvLight [1] and E2Clab [2] are available as open-source tools. Daniel Rosendo, Marta Mattoso, Alexandru Costan, Renan Souza 0001, Débora B. Pina, Patrick Valduriez, Gabriel Antoniu |
CLUSTER | 6 |
| 2023 | Large-scale knowledge distillation with elastic heterogeneous computing resourcesabstractAbstract Although more layers and more parameters generally improve the accuracy of the models, such big models generally have high computational complexity and require big memory, which exceed the capacity of small devices for inference and incurs long training time. In addition, it is difficult to afford long training time and inference time of big models even in high performance servers, as well. As an efficient approach to compress a large deep model (a teacher model) to a compact model (a student model), knowledge distillation emerges as a promising approach to deal with the big models. Existing knowledge distillation methods cannot exploit the elastic available computing resources and correspond to low efficiency. In this paper, we propose an Elastic Deep Learning framework for knowledge Distillation, that is, EDL‐Dist. The advantages of EDL‐Dist are threefold. First, the inference and the training process is separated. Second, elastic available computing resources can be utilized to improve the efficiency. Third, fault‐tolerance of the training and inference processes is supported. We take extensive experimentation to show that the throughput of EDL‐Dist is up to 3.125 times faster than the baseline method (online knowledge distillation) while the accuracy is similar or higher. Ji Liu 0003, Daxiang Dong, An Qin 0001, Xingjian Li 0002, Patrick Valduriez, Dejing Dou, Dianhai Yu |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Workflow provenance in the lifecycle of scientific machine learningabstractAbstract Machine learning (ML) has already fundamentally changed several businesses. More recently, it has also been profoundly impacting the computational science and engineering domains, like geoscience, climate science, and health science. In these domains, users need to perform comprehensive data analyses combining scientific data and ML models to provide for critical requirements, such as reproducibility, model explainability, and experiment data understanding. However, scientific ML is multidisciplinary, heterogeneous, and affected by the physical constraints of the domain, making such analyses even more challenging. In this work, we leverage workflow provenance techniques to build a holistic view to support the lifecycle of scientific ML. We contribute with (i) characterization of the lifecycle and taxonomy for data analyses; (ii) design decisions to build this view, with a W3C PROV compliant data representation and a reference system architecture; and (iii) lessons learned after an evaluation in an Oil & Gas case using an HPC cluster with 393 nodes and 946 GPUs. The experiments show that the decisions enable queries that integrate domain semantics with ML models while keeping low overhead (<1%), high scalability, and an order of magnitude of query acceleration under certain workloads against without our representation. Renan Souza 0001, Leonardo Guerreiro Azevedo, Vítor N. Lourenço, Elton F. S. Soares, Raphael Thiago, Rafael Brandão 0001, Daniel Civitarese, Emilio Vital Brazil, Márcio Ferreira Moreno, Patrick Valduriez, Marta Mattoso, Renato Cerqueira, Marco Aurélio Stelmar Netto |
Concurr. Comput. Pract. Exp. | 10 |
| 2022 | Elastic scalable transaction processing in LeanXcaleabstractInternational audience Ricardo Jiménez-Peris, Diego Burgos-Sancho, Francisco J. Ballesteros, Marta Patiño-Martínez, Patrick Valduriez |
Inf. Syst. | 5 |
| 2022 | Distributed intelligence on the Edge-to-Cloud Continuum: A systematic literature review
Daniel Rosendo, Alexandru Costan, Patrick Valduriez, Gabriel Antoniu |
J. Parallel Distributed Comput. | 3 |
| 2022 | Two-Phase Scheduling for Efficient Vehicle SharingabstractCooperative Intelligent Transport Systems (C-ITS) is a promising technology to make transportation safer and more efficient. Ridesharing for long-distance is becoming a key means of transportation in C-ITS. In this paper, we focus on private long-distance ridesharing, which reduces the total cost of vehicle utilization for long-distance journeys. In this context, we investigate journey scheduling problem with shared vehicles to reduce the total cost of vehicle utilization. Most of the existing works directly schedule journeys to vehicles with long scheduling time and only consider the cost of driving travellers instead of the total cost. In contrast, to reduce the total cost and scheduling time, we propose a comprehensive cost model and a two-phase journey scheduling approach, which includes path generation and path scheduling. On this basis, we propose two path generation methods: a simple near optimal method and a reset near optimal method as well as a greedy based path scheduling method. Finally, we present an experimental evaluation with different path generation and path scheduling methods with synthetic data generated based on real-world data. The results reveal that the proposed scheduling approach significantly outperforms baseline methods in terms of total cost (up to 69.8%) and scheduling time (up to 84.0%) and the scheduling time is reasonable (up to 0.16s). The results also show that our approach has higher efficiency (up to 141.7%) than baseline methods. Ji Liu 0003, Carlyna Bondiombouy, Lei Mo, Patrick Valduriez |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Reproducible Performance Optimization of Complex Applications on the Edge-to-Cloud ContinuumabstractIn more and more application areas, we are witnessing the emergence of complex workflows that combine computing, analytics and learning. They often require a hybrid execution infrastructure with IoT devices interconnected to cloud/HPC systems (aka Computing Continuum). Such workflows are subject to complex constraints and requirements in terms of performance, resource usage, energy consumption and financial costs. This makes it challenging to optimize their configuration and deployment. We propose a methodology to support the optimization of real-life applications on the Edge-to-Cloud Continuum. We implement it as an extension of E2Clab, a previously proposed framework supporting the complete experimental cycle across the Edge-to-Cloud Continuum. Our approach relies on a rigorous analysis of possible configurations in a controlled testbed environment to understand their behaviour and related performance tradeoffs. We illustrate our methodology by optimizing Pl@ntNet, a world-wide plant identification application. Our methodology can be generalized to other applications in the Edge-to-Cloud Continuum. Daniel Rosendo, Alexandru Costan, Gabriel Antoniu, Matthieu Simonin, Jean-Christophe Lombardo, Alexis Joly, Patrick Valduriez |
CLUSTER | 7 |
| 2021 | Parallel query processing in a polystore
Pavlos Kranas, Boyan Kolev, Oleksandra Levchenko, Esther Pacitti, Patrick Valduriez, Ricardo Jiménez-Peris, Marta Patiño-Martínez |
Distributed Parallel Databases | 5 |
| 2021 | Cache-aware scheduling of scientific workflows in a multisite cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
Future Gener. Comput. Syst. | 6 |
| 2021 | BestNeighbor: efficient evaluation of kNN queries on large time series databases
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas, Dennis E. Shasha, Patrick Valduriez |
Knowl. Inf. Syst. | 8 |
| 2020 | Distributed Caching of Scientific Workflows in Multisite Cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
DEXA (2) | 6 |
| 2020 | Parallel computation of PDFs on big spatial data using Spark
Ji Liu 0003, Noel Moreno Lemus, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez |
Distributed Parallel Databases | 5 |
| 2020 | Data reduction in scientific workflows using provenance monitoring and user steering
Renan Souza 0001, Vítor Silva 0003, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 4 |
| 2019 | Parallel Streaming Implementation of Online Time Series Correlation Discovery on Sliding Windows with Regression CapabilitiesabstractInternational audience Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez |
CLOSER | 7 |
| 2019 | Pipelined Implementation of a Parallel Streaming Method for Time Series Correlation Discovery on Sliding Windows
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez |
DATA | 7 |
| 2019 | Adaptive Caching for Data-Intensive Scientific Workflows in the Cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
DEXA (2) | 6 |
| 2019 | Efficient Runtime Capture of Multiworkflow Data Using ProvenanceabstractComputational Science and Engineering (CSE) projects are typically developed by multidisciplinary teams. Despite being part of the same project, each team manages its own workflows, using specific execution environments and data processing tools. Analyzing the data processed by all workflows globally is a core task in a CSE project. However, this analysis is hard because the data generated by these workflows are not integrated. In addition, since these workflows may take a long time to execute, data analysis needs to be done at runtime to reduce cost and time of the CSE project. A typical solution in scientific data analysis is to capture and relate the data in a provenance database while the workflows run, thus allowing for data analysis at runtime. However, the main problem is that such data capture competes with the running workflows, adding significant overhead to their execution. To mitigate this problem, we introduce in this paper a system called ProvLake, which adopts design principles for providing efficient distributed data capture from the workflows. While capturing the data, ProvLake logically integrates and ingests them into a provenance database ready for analyses at runtime. We validated ProvLake in a real use case in the O&G industry encompassing four workflows that process 5 TB datasets for a deep learning classifier. Compared with Komadu, the closest solution that meets our goals, our approach enables runtime multiworkflow data analysis with much smaller overhead, such as 0.1%. Renan Souza 0001, Marta Mattoso, Leonardo Guerreiro Azevedo, Raphael Thiago, Elton F. S. Soares, Marcelo Nery dos Santos, Marco Aurélio Stelmar Netto, Emilio Vital Brazil, Renato Cerqueira, Patrick Valduriez |
eScience | 10 |
| 2019 | Distributed Algorithms to Find Similar Time SeriesabstractInternational audience Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Dennis E. Shasha, Themis Palpanas, Patrick Valduriez, Reza Akbarinia, Florent Masseglia |
ECML/PKDD (3) | 6 |
| 2019 | DZI: An air index for spatial queries in one-dimensional channels
KwangJin Park, Alexis Joly, Patrick Valduriez |
Data Knowl. Eng. | 3 |
| 2019 | Keeping track of user steering actions in dynamic workflows
Renan Souza 0001, Vítor Silva 0003, José J. Camata, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 5 |
| 2019 | Efficient Scheduling of Scientific Workflows Using Hot Metadata in a Multisite CloudabstractLarge-scale, data-intensive scientific applications are often expressed as scientific workflows (SWfs). In this paper, we consider the problem of efficient scheduling of a large SWf in a multisite cloud, i.e., a cloud with geo-distributed cloud data centers (sites). The reasons for using multiple cloud sites to run a SWf are that data is already distributed, the necessary resources exceed the limits at a single site, or the monetary cost is lower. In a multisite cloud, metadata management has a critical impact on the efficiency of SWf scheduling as it provides a global view of data location and enables task tracking during execution. Thus, it should be readily available to the system at any given time. While it has been shown that efficient metadata handling plays a key role in performance, little research has targeted this issue in multisite cloud. In this paper, we propose to identify and exploit hot metadata (frequently accessed metadata) for efficient SWf scheduling in a multisite cloud, using a distributed approach. We implemented our approach within a scientific workflow management system, which shows that our approach reduces the execution time of highly parallel jobs up to 64 percent and that of the whole SWfs up to 55 percent. Ji Liu 0003, Luis Pineda-Morales, Esther Pacitti, Alexandru Costan, Patrick Valduriez, Gabriel Antoniu, Marta Mattoso |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Parallel Polyglot Query Processing on Heterogeneous Cloud Data Stores with LeanXcaleabstractThe blooming of different cloud data stores has turned polystore systems to a major topic in the nowadays cloud landscape. Especially, as the amount of processed data grows rapidly each year, much attention is being paid on taking advantage of the parallel processing capabilities of the underlying data stores. To provide data federation, a typical polystore solution defines a common data model and query language with translations to API calls or queries to each data store. However, this may lead to losing important querying capabilities. The polyglot approach of the CloudMdsQL query language allows data store native queries to be expressed as inline scripts and combined with regular SQL statements in ad-hoc integration queries. Moreover, efficient optimization techniques, such as bind join, can still take place to improve the performance of selective joins. In this paper, we introduce the distributed architecture of the LeanXcale query engine that processes polyglot queries in the CloudMdsQL query language, yet allowing native scripts to be handled in parallel at data store shards, so that efficient and scalable parallel joins take place at the query engine level. The experimental evaluation of the LeanXcale parallel query engine on various join queries illustrates well the performance benefits of exploiting the parallelism of the underlying data management technologies in combination with the high expressivity provided by their scripting/querying frameworks. Boyan Kolev, Oleksandra Levchenko, Esther Pacitti, Patrick Valduriez, Ricardo Vilaça, Rui C. Gonçalves, Ricardo Jiménez-Peris, Pavlos Kranas |
IEEE BigData | 4 |
| 2018 | Answering Top-k Queries over Outsourced Sensitive Data in the Cloud
Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez |
DEXA (1) | 3 |
| 2018 | Privacy-Preserving Top-k Query Processing in Distributed Systems
Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez |
Euro-Par | 3 |
| 2018 | Top-k Query Processing over Distributed Sensitive DataabstractDistributed systems provide users with powerful capabilities to store and process their data in third-party machines. However, the privacy of the outsourced data is not guaranteed. One solution for protecting the user data against privacy attacks is to encrypt the sensitive data before sending to the nodes of the distributed system. Then, the main problem is to evaluate user queries over the encrypted data. Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez |
IDEAS | 3 |
| 2018 | Point pattern search in big dataabstractConsider a set of points P in space with at least some of the pairwise distances specified. Given this set P, consider the following three kinds of queries against a database D of points : (i) pure constellation query: find all sets S in D of size |P| that exactly match the pairwise distances within P up to an additive error ϵ; (ii) isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f for which the distances between pairs in S exactly match f times the distances between corresponding pairs of P up to an additive ϵ; (iii) non-isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f and for at least some pairs of points, a maximum stretch factor mi,j > 1 such that (f X mi,jXdist(pi, pj))+ϵ > dist(si,sj) > (f X dist(pi, pj)) - ϵ. Finding matches to such queries has applications to spatial data in astronomical, seismic, and any domain in which (approximate, scale-independent) geometrical matching is required. Answering the isotropic and non-isotropic queries is challenging because scale factors and stretch factors may take any of an infinite number of values. This paper proposes practically efficient sequential and distributed algorithms for pure, isotropic, and non-isotropic constellation queries. As far as we know, this is the first work to address isotropic and non-isotropic queries. Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Alberto Krone-Martins, Patrick Valduriez, Dennis E. Shasha |
SSDBM | 5 |
| 2018 | ParCorr: efficient parallel methods to identify similar time series pairs across sliding windows
Djamel Edine Yagoubi, Reza Akbarinia, Boyan Kolev, Oleksandra Levchenko, Florent Masseglia, Patrick Valduriez, Dennis E. Shasha |
Data Min. Knowl. Discov. | 6 |
| 2018 | DfAnalyzer: Runtime Dataflow Analysis of Scientific Applications using ProvenanceabstractWe present DfAnalyzer, a tool that enables monitoring, debugging, steering, and analysis of dataflows while being generated by scientific applications. It works by capturing strategic domain data, registering provenance and execution data to enable queries at runtime. DfAnalyzer provides lightweight dataflow monitoring components to be invoked by high performance applications. It can be plugged in scientific code scripts, or Spark applications, in the same way users already plug visualization library components. During this demo, we will show how DfAnalyzer captures the dataflow, provenance, as well as how it provides runtime data analyses of applications. We will also encourage attendees to use DfAnalyzer for their own applications. Vítor Silva 0003, Daniel de Oliveira 0001, Marta Mattoso, Patrick Valduriez |
Proc. VLDB Endow. | 4 |
| 2017 | Pre-processing and Indexing Techniques for Constellation Queries in Big Data
Amir Khatibi, Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Patrick Valduriez, Dennis E. Shasha |
DaWaK | 5 |
| 2017 | InfraPhenoGrid: A scientific workflow infrastructure for plant phenomics on the Grid
Christophe Pradal, Simon Artzet, Jérôme Chopard, Dimitri Dupuis, Christian Fournier, Michael Mielewczik, Vincent Nègre, Pascal Neveu, Didier Parigot, Patrick Valduriez, Sarah Cohen Boulakia |
Future Gener. Comput. Syst. | 10 |
| 2017 | Raw data queries during data-intensive parallel workflow execution
Vítor Silva 0003, José Leite, José J. Camata, Daniel de Oliveira 0001, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 6 |
| 2017 | Advances in Databases and Information Systems
Ladjel Bellatreche, Patrick Valduriez, Tadeusz Morzy |
Inf. Syst. | 2 |
| 2016 | Benchmarking polystores: The CloudMdsQL experienceabstractThe CloudMdsQL polystore provides integrated access to multiple heterogeneous data stores, such as RDBMS, NoSQL or even HDFS through a big data analytics framework such as MapReduce or Spark. The CloudMdsQL language is a functional SQL-like query language with a flexible nested data model. A major capability is to exploit the full power of each of the underlying data stores by allowing native queries to be expressed as functions and involved in SQL statements. The CloudMdsQL polystore has been validated with a good number of different data stores: HDFS, key-value, document, graph, RDBMS and OLAP engine. In this paper, we introduce the benchmarking of the CloudMdsQL polystore and evaluate the performance benefits of important features enabled by the query language and engine. Boyan Kolev, Raquel Pau, Oleksandra Levchenko, Patrick Valduriez, Ricardo Jiménez-Peris, José Pereira 0001 |
IEEE BigData | 4 |
| 2016 | Managing hot metadata for scientific workflows on multisite cloudsabstractLarge-scale scientific applications are often expressed as workflows that help defining data dependencies between their different components. Several such workflows have huge storage and computation requirements, and so they need to be processed in multiple (cloud-federated) datacenters. It has been shown that efficient metadata handling plays a key role in the performance of computing systems. However, most of this evidence concern only single-site, HPC systems to date. In this paper, we present a hybrid decentralized/distributed model for handling hot metadata (frequently accessed metadata) in multisite architectures. We couple our model with a scientific workflow management system (SWfMS) to validate and tune its applicability to different real-life scientific scenarios. We show that efficient management of hot metadata improves the performance of SWfMS, reducing the workflow execution time up to 50% for highly parallel jobs and avoiding unnecessary cold metadata operations. Luis Pineda-Morales, Ji Liu 0003, Alexandru Costan, Esther Pacitti, Gabriel Antoniu, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 6 |
| 2016 | Design and Implementation of the CloudMdsQL Multistore SystemabstractInternational audience Boyan Kolev, Carlyna Bondiombouy, Oleksandra Levchenko, Patrick Valduriez, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001 |
CLOSER (1) | 4 |
| 2016 | Spatially Localized Visual Dictionary LearningabstractThis paper addresses the challenge of devising new representation learning algorithms that overcome the lack of interpretability of classical visual models. Therefore, it introduces a new recursive visual patch selection technique built on top of a Shared Nearest Neighbors embedding method. The main contribution of the paper is to drastically reduce the high-dimensionality of such over-complete representation thanks to a recursive feature elimination method. We show that the number of spatial atoms of the representation can be reduced by up to two orders of magnitude without much degrading the encoded information. The resulting representations are shown to provide competitive image classification performance with the state-of-the-art while enabling to learn highly interpretable visual models. Valentin Leveau, Alexis Joly, Olivier Buisson, Patrick Valduriez |
ICMR | 4 |
| 2016 | The CloudMdsQL Multistore SystemabstractThe blooming of different cloud data management infrastructures has turned multistore systems to a major topic in the nowadays cloud landscape. In this demonstration, we present a Cloud Multidatastore Query Language (CloudMdsQL), and its query engine. CloudMdsQL is a functional SQL-like language, capable of querying multiple heterogeneous data stores (relational and NoSQL) within a single query that may contain embedded invocations to each data store's native query interface. The major innovation is that a CloudMdsQL query can exploit the full power of local data stores, by simply allowing some local data store native queries (e.g. a breadth-first search query against a graph database) to be called as functions, and at the same time be optimized. Within our demonstration, we focus on two use cases each involving four diverse data stores (graph, document, relational, and key-value) with its corresponding CloudMdsQL queries. The query execution flows are visualized by an embedded real-time monitoring subsystem. The users can also try out different ad-hoc queries, not necessarily in the context of the use cases. Boyan Kolev, Carlyna Bondiombouy, Patrick Valduriez, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001 |
SIGMOD Conference | 3 |
| 2016 | Analyzing related raw data files through dataflowsabstractSummary Computer simulations may ingest and generate high numbers of raw data files. Most of these files follow a de facto standard format established by the application domain, for example, Flexible Image Transport System for astronomy. Although these formats are supported by a variety of programming languages, libraries, and programs, analyzing thousands or millions of files requires developing specific programs. Database management systems (DBMS) are not suited for this, because they require loading the raw data and structuring it, which becomes heavy at large scale. Systems like NoDB, RAW, and FastBit have been proposed to index and query raw data files without the overhead of using a database management system. However, these solutions are focused on analyzing one single large file instead of several related files. In this case, when related files are produced and required for analysis, the relationship among elements within file contents must be managed manually, with specific programs to access raw data. Thus, this data management may be time‐consuming and error‐prone. When computer simulations are managed by a scientific workflow management system (SWfMS), they can take advantage of provenance data to relate and analyze raw data files produced during workflow execution. However, SWfMS registers provenance at a coarse grain, with limited analysis on elements from raw data files. When the SWfMS is dataflow‐aware, it can register provenance data and the relationships among elements of raw data files altogether in a database, which is useful to access the contents of a large number of files. In this paper, we propose a dataflow approach for analyzing element data from several related raw data files. Our approach is complementary to the existing single raw data file analysis approaches. We use the Montage workflow from astronomy and a workflow from Oil and Gas domain as data‐intensive case studies. Our experimental results for the Montage workflow explore different types of raw data flows like showing all linear transformations involved in projection simulation programs, considering specific mosaic elements from input repositories. The cost for raw data extraction is approximately 3.7% of the total application execution time. Copyright © 2015 John Wiley & Sons, Ltd. Vítor Silva 0003, Daniel de Oliveira 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | CloudMdsQL: querying heterogeneous cloud data stores with a common language
Boyan Kolev, Patrick Valduriez, Carlyna Bondiombouy, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001 |
Distributed Parallel Databases | 2 |
| 2016 | Multi-objective scheduling of Scientific Workflows in multisite clouds
Ji Liu 0003, Esther Pacitti, Patrick Valduriez, Daniel de Oliveira 0001, Marta Mattoso |
Future Gener. Comput. Syst. | 3 |
| 2016 | FP-Hadoop: Efficient processing of skewed MapReduce jobs
Miguel Liroz-Gistau, Reza Akbarinia, Divyakant Agrawal, Patrick Valduriez |
Inf. Syst. | 4 |
| 2016 | Database System Support of Simulation DataabstractSupported by increasingly efficient HPC infra-structure, numerical simulations are rapidly expanding to fields such as oil and gas, medicine and meteorology. As simulations become more precise and cover longer periods of time, they may produce files with terabytes of data that need to be efficiently analyzed. In this paper, we investigate techniques for managing such data using an array DBMS. We take advantage of multidimensional arrays that nicely models the dimensions and variables used in numerical simulations. However, a naive approach to map simulation data files may lead to sparse arrays, impacting query response time, in particular, when the simulation uses irregular meshes to model its physical domain. We propose efficient techniques to map coordinate values in numerical simulations to evenly distributed cells in array chunks with the use of equi-depth histograms and space-filling curves. We implemented our techniques in SciDB and, through experiments over real-world data, compared them with two other approaches: row-store and column-store DBMS. The results indicate that multidimensional arrays and column-stores are much faster than a traditional row-store system for queries over a larger amount of simulation data. They also help identifying the scenarios where array DBMSs are most efficient, and those where they are outperformed by column-stores. Hermano Lustosa, Fábio Porto 0001, Patrick Valduriez |
Proc. VLDB Endow. | 4 |
| 2015 | An Efficient Solution for Processing Skewed MapReduce Jobs
Reza Akbarinia, Miguel Liroz-Gistau, Divyakant Agrawal, Patrick Valduriez |
DEXA (2) | 4 |
| 2015 | Integrating Big Data and Relational Data with a Functional SQL-like Query Language
Carlyna Bondiombouy, Boyan Kolev, Oleksandra Levchenko, Patrick Valduriez |
DEXA (1) | 4 |
| 2015 | Kernelizing Spatially Consistent Visual Matches for Fine-Grained ClassificationabstractThis paper introduces a new image representation relying on the spatial pooling of geometrically consistent visual matches. We therefore introduce a new match kernel based on the inverse rank of the shared nearest neighbors combined with local geometric constraints. To avoid overfitting and reduce processing costs, the dimensionality of the resulting over-complete representation is further reduced by hierarchically pooling the raw consistent matches according to their spatial position in the training images. The final image representation is obtained by concatenating the resulting feature vectors at several resolutions. Learning from these representations using a logistic regression classifier is shown to provide excellent fine-grained classification performances outperforming the results reported in the literature on several classification tasks. Valentin Leveau, Alexis Joly, Olivier Buisson, Patrick Valduriez |
ICMR | 4 |
| 2015 | OpenAlea: scientific workflows combining data analysis and simulationabstractAnalyzing biological data (e.g., annotating genomes, assembling NGS data...) may involve very complex and interlinked steps where several tools are combined together. Scientific workflow systems have reached a level of maturity that makes them able to support the design and execution of such in-silico experiments, and thus making them increasingly popular in the bioinformatics community. Christophe Pradal, Christian Fournier, Patrick Valduriez, Sarah Cohen Boulakia |
SSDBM | 3 |
| 2015 | Data-centric iteration in dynamic workflows
Jonas Dias, Gabriel Guerra, Fernando Rochinha, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 5 |
| 2015 | A Survey of Data-Intensive Scientific Workflow Management
Ji Liu 0003, Esther Pacitti, Patrick Valduriez, Marta Mattoso |
J. Grid Comput. | 3 |
| 2015 | Corrigendum to "Best position algorithms for efficient top-k query processing" [Inf. Syst 36(6) (2011) 973-989]
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Inf. Syst. | 3 |
| 2015 | FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed DataabstractBig data parallel frameworks, such as MapReduce or Spark have been praised for their high scalability and performance, but show poor performance in the case of data skew. There are important cases where a high percentage of processing in the reduce side ends up being done by only one node. In this demonstration, we illustrate the use of FP-Hadoop , a system that efficiently deals with data skew in MapReduce jobs. In FP-Hadoop, there is a new phase, called intermediate reduce (IR) , in which blocks of intermediate values, constructed dynamically, are processed by intermediate reduce workers in parallel, by using a scheduling strategy. Within the IR phase, even if all intermediate values belong to only one key, the main part of the reducing work can be done in parallel using the computing resources of all available workers. We implemented a prototype of FP-Hadoop, and conducted extensive experiments over synthetic and real datasets. We achieve excellent performance gains compared to native Hadoop, e.g. more than 10 times in reduce time and 5 times in total execution time. During our demonstration, we give the users the possibility to execute and compare job executions in FP-Hadoop and Hadoop. They can retrieve general information about the job and the tasks and a summary of the phases. They can also visually compare different configurations to explore the difference between the approaches. Miguel Liroz-Gistau, Reza Akbarinia, Patrick Valduriez |
Proc. VLDB Endow. | 3 |
| 2014 | Recognizing Thousands of Legal Entities through Instance-based Visual ClassificationabstractThis paper considers the problem of recognizing legal entities in visual contents in a similar way to named-entity recognizers for text documents. Whereas previous works were restricted to the recognition of a few tens of logotypes, we generalize the problem to the recognition of thousands of legal persons, each being modeled by a rich corporate identity automatically built from web images. We introduce a new geometrically-consistent instance-based classification method that is shown to outperform state-of-the-art techniques on several challenging datasets while being much more scalable. Further experiments performed on an automatic web crawl of 5,824 legal entities demonstrates the scalability of the approach. Valentin Leveau, Alexis Joly, Olivier Buisson, Pierre Letessier, Patrick Valduriez |
ACM Multimedia | 5 |
| 2014 | Entity resolution for probabilistic data
Naser Ayat, Reza Akbarinia, Hamideh Afsarmanesh, Patrick Valduriez |
Inf. Sci. | 4 |
| 2014 | Special section on data-intensive cloud infrastructure
Ashraf Aboulnaga, Beng Chin Ooi, Patrick Valduriez |
VLDB J. | 3 |
| 2013 | Algebraic dataflows for big data analysisabstractAnalyzing big data requires the support of dataflows with many activities to extract and explore relevant information from the data. Recent approaches such as Pig Latin propose a high-level language to model such dataflows. However, the dataflow execution is typically delegated to a MapRe-duce implementation such as Hadoop, which does not follow an algebraic approach, thus it cannot take advantage of the optimization opportunities of PigLatin algebra. In this paper, we propose an approach for big data analysis based on algebraic workflows, which yields optimization and parallel execution of activities and supports user steering using provenance queries. We illustrate how a big data processing dataflow can be modeled using the algebra. Through an experimental evaluation using real datasets and the execution of the dataflow with Chiron, an engine that supports our algebra, we show that our approach yields performance gains of up to 19.6% using algebraic optimizations in the dataflow and up to 39.1% of time saved on a user steering scenario. Jonas Dias, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 5 |
| 2013 | What You Pay for Is What You Get
Ruiming Tang, Dongxu Shao, Stéphane Bressan, Patrick Valduriez |
DEXA (2) | 4 |
| 2013 | The Price Is Right - Models and Algorithms for Pricing DataabstractData is a modern commodity. Yet the pricing models in useon electronic data markets either focus on the usage of computing resources,or are proprietary, opaque, most likely ad hoc, and not conduciveof a healthy commodity market dynamics. In this paper we propose ageneric data pricing model that is based on minimal provenance, i.e. minimalsets of tuples contributing to the result of a query.We show that theproposed model fulfills desirable properties such as contribution monotonicity,bounded-price and contribution arbitrage-freedom. We presenta baseline algorithm to compute the exact price of a query based onour pricing model. We show that the problem is NP-hard. We thereforedevise, present and compare several heuristics. We conduct a comprehensiveexperimental study to show their effectiveness and efficiency. Ruiming Tang, Huayu Wu 0001, Zhifeng Bao, Stéphane Bressan, Patrick Valduriez |
DEXA (2) | 5 |
| 2013 | Chiron: a parallel engine for algebraic scientific workflowsabstractSUMMARY Large‐scale scientific experiments based on computer simulations are typically modeled as scientific workflows, which eases the chaining of different programs. These scientific workflows are defined, executed, and monitored by scientific workflow management systems (SWfMS). As these experiments manage large amounts of data, it becomes critical to execute them in high‐performance computing environments, such as clusters, grids, and clouds. However, few SWfMS provide parallel support. The ones that do so are usually labor‐intensive for workflow developers and have limited primitives to optimize workflow execution. To address these issues, we developed workflow algebra to specify and enable the optimization of parallel execution of scientific workflows. In this paper, we show how the workflow algebra is efficiently implemented in Chiron, an algebraic based parallel scientific workflow engine. Chiron has a unique native distributed provenance mechanism that enables runtime queries in a relational database. We developed two studies to evaluate the performance of our algebraic approach implemented in Chiron; the first study compares Chiron with different approaches, whereas the second one evaluates the scalability of Chiron. By analyzing the results, we conclude that Chiron is efficient in executing scientific workflows, with the benefits of declarative specification and runtime provenance support. Copyright © 2013 John Wiley & Sons, Ltd. Eduardo S. Ogasawara, Jonas Dias, Vítor Silva 0003, Fernando Seabra Chirigati, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 7 |
| 2013 | Entity resolution for distributed probabilistic data
Naser Ayat, Reza Akbarinia, Hamideh Afsarmanesh, Patrick Valduriez |
Distributed Parallel Databases | 4 |
| 2013 | A Hierarchical Grid Index (HGI), spatial queries in wireless data broadcasting
KwangJin Park, Patrick Valduriez |
Distributed Parallel Databases | 2 |
| 2013 | Efficient Evaluation of SUM Queries over Probabilistic DataabstractSUM queries are crucial for many applications that need to deal with uncertain data. In this paper, we are interested in the queries, called ALL_SUM, that return all possible sum values and their probabilities. In general, there is no efficient solution for the problem of evaluating ALL_SUM queries. But, for many practical applications, where aggregate values are small integers or real numbers with small precision, it is possible to develop efficient solutions. In this paper, based on a recursive approach, we propose a new solution for those applications. We implemented our solution and conducted an extensive experimental evaluation over synthetic and real-world data sets; the results show its effectiveness. Reza Akbarinia, Patrick Valduriez, Guillaume Verger |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Dynamic Workload-Based Partitioning for Large-Scale Databases
Miguel Liroz-Gistau, Reza Akbarinia, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez |
DEXA (2) | 5 |
| 2012 | Satisfaction-based query replication - An automatic and self-adaptable approach for replicating queries in the presence of autonomous participants
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
Distributed Parallel Databases | 3 |
| 2012 | StreamCloud: An Elastic and Scalable Data Streaming SystemabstractMany applications in several domains such as telecommunications, network security, large-scale sensor networks, require online processing of continuous data flows. They produce very high loads that requires aggregating the processing capacity of many nodes. Current Stream Processing Engines do not scale with the input load due to single-node bottlenecks. Additionally, they are based on static configurations that lead to either under or overprovisioning. In this paper, we present StreamCloud, a scalable and elastic stream processing engine for processing large data stream volumes. StreamCloud uses a novel parallelization technique that splits queries into subqueries that are allocated to independent sets of nodes in a way that minimizes the distribution overhead. Its elastic protocols exhibit low intrusiveness, enabling effective adjustment of resources to the incoming load. Elasticity is combined with dynamic load balancing to minimize the computational resources used. The paper presents the system design, implementation, and a thorough evaluation of the scalability and elasticity of the fully implemented system. Vincenzo Gulisano, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Claudio Soriente, Patrick Valduriez |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2011 | Scaling Up Query Allocation in the Presence of Autonomous Participants
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Sylvie Cazalens, Patrick Valduriez |
DASFAA (2) | 4 |
| 2011 | Efficient Early Top-k Query Processing in Overloaded P2P Systems
William Kokou Dedzoe, Philippe Lamarre, Reza Akbarinia, Patrick Valduriez |
DEXA (1) | 4 |
| 2011 | Principles of Distributed Data Management in 2020?
Patrick Valduriez |
DEXA (1) | 1 |
| 2011 | Distributed data management in 2020?abstractWork on distributed data management commenced shortly after the introduction of the relational model in the mid-1970's. 1970's and 1980's were very active periods for the development of distributed relational database technology, and claims were made that in the following ten years centralized databases will be an “antique curiosity” and most organizations will move toward distributed database managers [1]. That prediction has certainly become true, and all commercial DBMSs today are distributed. M. Tamer Özsu, Patrick Valduriez, Serge Abiteboul, Bettina Kemme, Ricardo Jiménez-Peris, Beng Chin Ooi |
ICDE | 2 |
| 2011 | ParallelGDB: a parallel graph database based on cache specializationabstractThe need for managing massive attributed graphs is becoming common in many areas such as recommendation systems, proteomics analysis, social network analysis or bibliographic analysis. This is making it necessary to move towards parallel systems that allow managing graph databases containing millions of vertices and edges. Previous work on distributed graph databases has focused on finding ways to partition the graph to reduce network traffic and improve execution time. However, partitioning a graph and keeping the information regarding the location of vertices might be unrealistic for massive graphs. In this paper, we propose Parallel-GDB, a new system based on specializing the local caches of any node in this system, providing a better cache hit ratio. ParallelGDB uses a random graph partitioning, avoiding complex partition methods based on the graph topology, that usually require managing extra data structures. This proposed system provides an efficient environment for distributed graph databases. Luis Barguñó, Victor Muntés-Mulero, David Dominguez-Sal, Patrick Valduriez |
IDEAS | 4 |
| 2011 | Best position algorithms for efficient top-k query processing
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Inf. Syst. | 3 |
| 2011 | An Algebraic Approach for Data-Centric Scientific Workflows
Eduardo S. Ogasawara, Daniel de Oliveira 0001, Patrick Valduriez, Jonas Dias, Fábio Porto 0001, Marta Mattoso |
Proc. VLDB Endow. | 3 |
| 2011 | Energy Efficient Data Access in Mobile P2P NetworksabstractA fundamental problem for peer-to-peer (P2P) applications in mobile-pervasive computing environment is to efficiently identify the node that stores particular data items and download them while preserving battery power. In this paper, we propose a P2P Minimum Boundary Rectangle (PMBR, for short) which is a new spatial index specifically designed for mobile P2P environments. A node that contains desirable data item (s) can be easily identified by reading the PMBR index. Then, we propose a selective tuning algorithm, called Distributed exponential Sequence Scheme (DSS, for short), that provides clients with the ability of selective tuning of data items, thus preserving the scarce power resource. The proposed algorithm is simple but efficient in supporting linear transmission of spatial data and processing of location-aware queries. The results from theoretical analysis and experiments show that the proposed algorithm with the PMBR index is scalable and energy efficient in both range queries and nearest neighbor queries. KwangJin Park, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | StreamCloud: A Large Scale Data Streaming SystemabstractData streaming has become an important paradigm for the real-time processing of continuous data flows in domains such as finance, telecommunications, networking, Some applications in these domains require to process massive data flows that current technology is unable to manage, that is, streams that, even for a single query operator, require the capacity of potentially many machines. Research efforts on data streaming have mainly focused on scaling in the number of queries or query operators, but overlooked the scalability issue with respect to the stream volume. In this paper, we present StreamCloud a large scale data streaming system for processing large data stream volumes. We focus on how to parallelize continuous queries to obtain a highly scalable data streaming infrastructure. StreamCloud goes beyond the state of the art by using a novel parallelization technique that splits queries into subqueries that are allocated to independent sets of nodes in a way that minimizes the distribution overhead. StreamCloud is implemented as a middleware and is highly independent of the underlying data streaming engine. We explore and evaluate different strategies to parallelize data streaming and tackle with the main bottlenecks and overheads to achieve scalability. The paper presents the system design, implementation and a thorough evaluation of the scalability of the fully implemented system. Vincenzo Gulisano, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Patrick Valduriez |
ICDCS | 4 |
| 2010 | PeerUnit: a framework for testing peer-to-peer systemsabstractTesting distributed systems is challenging. Peer-to-peer (P2P) systems are composed of a high number of concurrent nodes distributed across the network. The nodes are also highly volatile (i.e., free to join and leave the system at any time). In this kind of system, a great deal of control should be carried out by the test harness, including: volatility of nodes, test case deployment and coordination. In this demonstration we present the PeerUnit framework for testing P2P systems. The original characteristic of this framework is the individual control of nodes, allowing test cases to precisely control their volatility during execution. We validated this framework through implementation and experimentation on two popular open-source P2P systems. Eduardo C. de Almeida, João Eugenio Marynowski, Gerson Sunyé, Patrick Valduriez |
ASE | 4 |
| 2010 | ASAP Top-k Query Processing in Unstructured P2P SystemsabstractTop-k query processing techniques are useful in unstructured peer-to-peer (P2P) systems, to avoid overwhelming users with too many results. However, existing approaches suffer from long waiting times. This is because top-k results are returned only when all queried peers have finished processing the query. As a result, query response time is dominated by the slowest queried peer. In this paper, we address this users' waiting time problem. For this, we revisit top-k query processing in P2P systems by introducing two novel notions in addition to response time: the stabilization time and the cumulative quality gap. Using these notions, we formally define the as-soon-as-possible (ASAP) top-k processing problem. Then, we propose a family of algorithms called ASAP to deal with this problem. We validate our solution through implementation and extensive experimentation. The results show that ASAP significantly outperforms baseline algorithms by re- turning final top-k result to users in much better times William Kokou Dedzoe, Philippe Lamarre, Reza Akbarinia, Patrick Valduriez |
Peer-to-Peer Computing | 4 |
| 2010 | Efficient Distributed Test Architectures for Large-Scale Systems
Eduardo C. de Almeida, João Eugenio Marynowski, Gerson Sunyé, Yves Le Traon, Patrick Valduriez |
ICTSS | 5 |
| 2010 | Testing peer-to-peer systems
Eduardo C. de Almeida, Gerson Sunyé, Yves Le Traon, Patrick Valduriez |
Empir. Softw. Eng. | 4 |
| 2010 | A semantic information system for services and traded resources in Grid e-markets
George A. Vouros, Andreas Papasalouros, Konstantinos Tzonas, Alexandros G. Valarakos, Konstantinos Kotis, Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
Future Gener. Comput. Syst. | 8 |
| 2010 | A scalable energy-efficient continuous nearest neighbor search in wireless broadcast systems
KwangJin Park, Hyunseung Choo, Patrick Valduriez |
Wirel. Networks | 3 |
| 2009 | SbQA: A Self-Adaptable Query Allocation ProcessabstractWe present a flexible query allocation framework, called satisfaction-based query allocation (SbQA for short), for distributed information systems where both consumers and providers (the participants) have special interests towards queries. A particularity of SbQA is that it allocates queries while considering both query load and participants' interests. To be fair, it dynamically trades consumers' interests for providers' interests based on their satisfaction. In this demo we illustrate the flexibility and efficiency of SbQA to allocate queries on the Berkeley Open Infrastructure for Network Computing (BOINC). We also demonstrate that SbQA is self-adaptable to the participants' expectations. Finally, we demonstrate that SbQA can be adapted to different kinds of applications by varying its parameters. Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
ICDE | 3 |
| 2009 | Real-Time Support for Software Transactional MemoryabstractTransactional memory is a hot research topic, having attracted the focus of both academic researchers and development groups at companies. Indeed, the concept of transactional memory has attracted much interest for multicore systems as it eases programming and avoids the problems of lock-based methods. However, up to now, the scheduling of real-time transactions within software transactional memories has not been studied. To address this issue, we present in this paper a real-time software transactional memory, namely RT-STM. We focus on the scheduling of concurrent soft real-time transactions. In particular, we explore a new heuristic for conflict resolution that reduces the number of deadline violations when scheduling soft real-time transactions. After having discussed the scalability of various classical STMs under a real-time operating system, we present experimental results that show that RT-STM can improve the performance of transactional memory-based applications on multicore platforms. Toufik Sarni, Audrey Queudet, Patrick Valduriez |
RTCSA | 3 |
| 2009 | Parallel OLAP query processing in database clusters with data replication
Alexandre A. B. Lima, Camille Furtado, Patrick Valduriez, Marta Mattoso |
Distributed Parallel Databases | 3 |
| 2009 | DHTJoin: processing continuous join queries using DHT networks
Wenceslao Palma, Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Distributed Parallel Databases | 4 |
| 2009 | Towards the efficient development of model transformations using model weaving and matching transformations
Marcos Didonet Del Fabro, Patrick Valduriez |
Softw. Syst. Model. | 2 |
| 2009 | A self-adaptable query allocation framework for distributed information systems
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
VLDB J. | 3 |
| 2008 | Efficient Processing of Nearest Neighbor Queries in Parallel Multimedia Databases
Jorge R. Manjarrez Sanchez, José Martinez 0001, Patrick Valduriez |
DEXA | 3 |
| 2008 | Summary management in P2P systemsabstractSharing huge, massively distributed databases in P2P systems is inherently difficult. As the amount of stored data increases, data localization techniques become no longer sufficient. A practical approach is to rely on compact database summaries rather than raw database records, whose access is costly in large P2P systems. In this paper, we consider summaries that are synthetic, multidimensional views with two main virtues. First, they can be directly queried and used to approximately answer a query without exploring the original data. Second, as semantic indexes, they support locating relevant nodes based on data content. Our main contribution is to define a summary model for P2P systems, and the appropriate algorithms for summary management. Our performance evaluation shows that the cost of query routing is minimized, while incurring a low cost of summary maintenance. Rabab Hayek, Guillaume Raschia, Patrick Valduriez, Noureddine Mouaddib |
EDBT | 3 |
| 2008 | Improving Interoperability Using Query Interpretation in Semantic Vector Spaces
Anthony Ventresque, Sylvie Cazalens, Philippe Lamarre, Patrick Valduriez |
ESWC | 4 |
| 2008 | Efficient Processing of Continuous Join Queries Using Distributed Hash Tables
Wenceslao Palma, Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Euro-Par | 4 |
| 2008 | Mobile continuous nearest neighbor queries on airabstractIn this paper, we propose new spatial query processing algorithms to support Mobile Continuous Nearest Neighbor Query (MCNNQ) in wireless broadcast environments. Each client (object) computes individually the function that best captures the locations of the moving objects, while a server is only responsible for delivering information gathered from the clients. Extensive experiments demonstrate that our location-based data dissemination algorithm significantly out-performs index-based solutions. KwangJin Park, Patrick Valduriez, Hyunseung Choo |
GIS | 2 |
| 2008 | A Framework for Testing Peer-to-Peer SystemsabstractDeveloping peer-to-peer (P2P) systems is hard because they must be deployed on a high number of nodes, which can be autonomous, refusing to answer to some requests or even unexpectedly leaving the system. Such volatility of nodes is a common behavior in P2P system and can be interpreted as fault during tests.In this paper, we propose a framework for testing P2P systems. This framework is based on the individual control of nodes, allowing test cases to precisely control the volatility of nodes during execution.We validated this framework through implementation and experimentation on an open-source P2P system. Eduardo C. de Almeida, Gerson Sunyé, Yves Le Traon, Patrick Valduriez |
ISSRE | 4 |
| 2008 | Testing Peers' VolatilityabstractPeer-to-peer (P2P) is becoming a key technology for software development, but still lacks integrated solutions to build trust in the final software, in terms of correctness and security. Testing such systems is difficult because of the high numbers of nodes which can be volatile. In this paper, we present a framework for testing volatility of P2P systems. The framework is based on the individual control of peers, allowing test cases to precisely control the volatility of peers during execution. We validated our framework through implementation and experimentation on two open-source P2P systems. Through experimentation, we analyze the behavior of both systems on different conditions of volatility and show how the framework is able to detect implementation problems. Eduardo C. de Almeida, Gerson Sunyé, Yves Le Traon, Patrick Valduriez |
ASE | 4 |
| 2008 | Parallel query processing for OLAP in gridsabstractAbstract OLAP query processing is critical for enterprise grids. Capitalizing on our experience with the ParGRES database cluster, we propose a middleware solution, GParGRES, which exploits database replication and inter‐ and intra‐query parallelism to efficiently support OLAP queries in a grid. GParGRES is designed as a wrapper that enables the use of ParGRES in PC clusters of a grid (in our case, Grid5000). Our approach has two levels of query splitting: grid‐level splitting, implemented by GParGRES, and node‐level splitting, implemented by ParGRES. GParGRES has been partially implemented as database grid services compatible with existing grid solutions such as the open grid service architecture and the Web services resource framework. We give preliminary experimental results obtained with two clusters of Grid5000 using queries of the TPC‐H Benchmark. The results show linear or almost linear speedup in query execution, as more nodes are added in all tested configurations. Copyright © 2008 John Wiley & Sons, Ltd. Nelson Kotowski, Alexandre A. B. Lima, Esther Pacitti, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 4 |
| 2008 | P2P logging and timestamping for reconciliationabstractIn this paper, we address data reconciliation in peer-to-peer (P2P) collaborative applications. We propose P2P-LTR (Logging and Timestamping for Reconciliation) which provides P2P logging and timestamping services for P2P reconciliation over a distributed hash table (DHT). While updating at collaborating peers, updates are timestamped and stored in a highly available P2P log. During reconciliation, these updates are retrieved in total order to enforce eventual consistency. In this paper, we first give an overview of P2P-LTR with its model and its main procedures. We then present our prototype used to validate P2P-LTR. To demonstrate P2P-LTR, we propose several scenarios that test our solutions and measure performance. In particular, we demonstrate how P2P-LTR handles the dynamic behavior of peers with respect to the DHT. Mounir Tlili, William Kokou Dedzoe, Esther Pacitti, Patrick Valduriez, Reza Akbarinia, Pascal Molli, Gérôme Canals, Stéphane Laurière |
Proc. VLDB Endow. | 4 |
| 2008 | DiSC: Benchmarking Secure Chip DBMSabstractAbstract—Secure chips, e.g., present in smart cards, USB dongles, i-buttons, are now ubiquitous in applications with strong security requirements. Moreover, they require embedded data management techniques. However, secure chips have severe hardware constraints, which make traditional database techniques irrelevant. The main problem faced by secure chip DBMS designers is to be able to assess various design choices and trade-offs for different applications. Our solution is to use a benchmark for secure chip DBMS in order to 1) compare different database techniques, 2) predict the limits of on-chip applications, and 3) provide codesign hints. In this paper, we propose Data management in Secure Chip (DiSC), a benchmark that reaches these three objectives. This work benefits from our long experience in developing and tuning data management techniques for the smart card. To validate DiSC, we compare the behavior of candidate data management techniques using a cycle-accurate smart-card simulator. Furthermore, we show the applicability of DiSC to future designs involving new hardware platforms and new database techniques. Nicolas Anciaux, Luc Bouganim, Philippe Pucheral, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2007 | Satisfaction balanced mediationabstractWe consider a distributed information system that allows autonomous consumers to query autonomous providers. We focus on the problem of query allocation from a new point of view, by considering consumers and providers' satisfaction in addition to query load. We define satisfaction as a long-run notion based on the consumers and providers' preferences. We propose and validate a mediation process, called SBMediation, which is compared to Capacity based query allocation. The experimental results show that SBMediation significantly outperforms Capacity based when confronted to autonomous participants. Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Sylvie Cazalens, Patrick Valduriez |
CIKM | 4 |
| 2007 | KnBest - A Balanced Request Allocation Method for Distributed Information Systems
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
DASFAA | 3 |
| 2007 | Processing Top-k Queries in Distributed Hash Tables
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Euro-Par | 3 |
| 2007 | Design of PeerSum: A Summary Service for P2P Applications
Rabab Hayek, Guillaume Raschia, Patrick Valduriez, Noureddine Mouaddib |
GPC | 3 |
| 2007 | Data currency in replicated DHTsabstractDistributed Hash Tables (DHTs) provide a scalable solution for data sharing in P2P systems. To ensure high data availability, DHTs typically rely on data replication, yet without data currency guarantees. Supporting data currency in replicated DHTs is difficult as it requires the ability to return a current replica despite peers leaving the network or concurrent updates. In this paper, we give a complete solution to this problem. We propose an Update Management Service (UMS) to deal with data availability and efficient retrieval of current replicas based on timestamping. For generating timestamps, we propose a Key-based Timestamping Service (KTS) which performs distributed timestamp generation using local counters. Through probabilistic analysis, we compute the expected number of replicas which UMS must retrieve for finding a current replica. Except for the cases where the availability of current replicas is very low, the expected number of retrieved replicas is typically small, e.g. if at least 35% of available replicas are current then the expected number of retrieved replicas is less than 3. We validated our solution through implementation and experimentation over a 64-node cluster and evaluated its scalability through simulation up to 10,000 peers using SimJava. The results show the effectiveness of our solution. They also show that our algorithm used in UMS achieves major performance gains, in terms of response time and communication cost, compared with a baseline algorithm. Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
SIGMOD Conference | 3 |
| 2007 | Best Position Algorithms for Top-k Queries
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
VLDB | 3 |
| 2007 | SQLB: A Query Allocation Framework for Autonomous Consumers and Providers
Jorge-Arnulfo Quiané-Ruiz, Philippe Lamarre, Patrick Valduriez |
VLDB | 3 |
| 2007 | Preface to the Special Issue on Grid Data Management
Esther Pacitti, Marta Mattoso, Patrick Valduriez |
J. Grid Comput. | 3 |
| 2007 | Grid Data Management: Open Problems and New Issues
Esther Pacitti, Patrick Valduriez, Marta Mattoso |
J. Grid Comput. | 2 |
| 2007 | A Flexible Mediation Process for Large Distributed Information SystemsabstractWe consider distributed information systems that are open, dynamic and provide access to large numbers of distributed, heterogeneous, autonomous information sources. Most of the work in data mediator systems has dealt with the problem of finding relevant information providers for a request. However, finding relevant requests for information providers is another important side of the mediation problem which has not received much attention. In this paper, we address these two sides of the problem with a flexible mediation process. Once the qualified information providers are identified, our process allows them to express their interest in a request via a bidding mechanism. It also requires to set up a requisition policy, because a request must always be answered if there are qualified providers. This work does not concern pure market mechanisms because we counter-balance the providers' bids by considering their quality wrt a request. We validate our process on a set of simulations in the context of load balancing, which is a good indicator of the system's overall performance. The results show that the mediation process provides a very good long-run regulation of the system, in particular when providers can leave the system. However, load balancing is not the natural application of the flexible mediation and additional testing is required to show the generality of the approach to non-depletable resources. Philippe Lamarre, Sandra Lemp, Sylvie Cazalens, Patrick Valduriez |
Int. J. Cooperative Inf. Syst. | 4 |
| 2007 | The leganet system: Freshness-aware transaction routing in a database cluster
Stéphane Gançarski, Hubert Naacke, Esther Pacitti, Patrick Valduriez |
Inf. Syst. | 4 |
| 2006 | Topic 5: Parallel and Distributed Databases, Data Mining and Knowledge Discovery
Patrick Valduriez, Wolfgang Lehner, Domenico Talia, Paul Watson 0001 |
Euro-Par | 1 |
| 2006 | Reducing network traffic in unstructured P2P systems using Top-k queries
Reza Akbarinia, Esther Pacitti, Patrick Valduriez |
Distributed Parallel Databases | 3 |
| 2005 | Topic 5 - Parallel and Distributed Databases, Data Mining and Knowledge Discovery
Domenico Talia, Hillol Kargupta, Patrick Valduriez, Rui Camacho |
Euro-Par | 3 |
| 2005 | Physical and Virtual Partitioning in OLAP Database ClustersabstractOn-line analytical processing (OLAP) applications require high performance database support to achieve good response time (crucial for decision making). Database clusters provide a cost-effective alternative to parallel database systems. For OLAP applications, that typically use heavy weight queries, intra-query parallelism yields better performance as it reduces the execution time of individual queries. Intra-query parallelism is based on processing the same query on different subsets of the query table. Combining physical and virtual partitioning to define table subsets provides flexibility in intra-query parallelism while optimizing disk space usage and data availability. Experiments with our partitioning technique using TPC-H benchmark queries on a 32-dual node cluster gave linear and super-linear speedup, thereby reducing significantly the time of typical OLAP heavy weight queries. Camille Furtado, Alexandre A. B. Lima, Esther Pacitti, Patrick Valduriez, Marta Mattoso |
SBAC-PAD | 4 |
| 2005 | Preventive Replication in a Database Cluster
Esther Pacitti, Cédric Coulon, Patrick Valduriez, M. Tamer Özsu |
Distributed Parallel Databases | 3 |
| 2004 | OLAP Query Processing in a Database Cluster
Alexandre A. B. Lima, Marta Mattoso, Patrick Valduriez |
Euro-Par | 3 |
| 2001 | Processing Queries with Expensive Functions and Large Objects in Distributed Mediator SystemsabstractLeSelect is a mediator system which allows scientists to publish their resources (data and programs) so they can be transparently accessed. The scientists can typically issue queries which access distributed published data and involve the execution of expensive functions (corresponding to programs). Furthermore, the queries can involve large objects, such as images (e.g. archived meteorological satellite data). In this context, the costs of transmitting large objects and invoking expensive functions are the dominant factors of execution time. In this paper, we first propose three query execution techniques which minimize these costs by taking full advantage of the distributed architecture of mediator systems like LeSelect. Then we devise parallel processing strategies for queries including expensive functions. Based on experimentation, we show that it is hard to predict the optimal execution order when dealing with several functions. We propose a new hybrid parallel technique to solve this problem and give some experimental results. Luc Bouganim, Françoise Fabret, Fábio Porto 0001, Patrick Valduriez |
ICDE | 4 |
| 2001 | PicoDBMS: Validation and Experience
Nicolas Anciaux, Christophe Bobineau, Luc Bouganim, Philippe Pucheral, Patrick Valduriez |
VLDB | 5 |
| 2001 | User-Optimizer Communication using Abstract Plans in Sybase ASE
Mihnea Andrei, Patrick Valduriez |
VLDB | 2 |
| 2001 | PicoDBMS: Scaling down database techniques for the smartcard
Philippe Pucheral, Luc Bouganim, Patrick Valduriez, Christophe Bobineau |
VLDB J. | 3 |
| 2000 | Dynamic Query Scheduling in Data Integration SystemsabstractExecution plans produced by traditional query optimizers for data integration queries may yield poor performance for several reasons. The cost estimates may be inaccurate, the memory available at run-time may be insufficient, or data delivery rate can be unpredictable. We address the problem of unpredictable data arrival rate. We propose to dynamically schedule queries in order to deal with irregular data delivery rate and gracefully adapt to the available memory. Our approach performs careful step-by-step scheduling of several query fragments and processes these fragments based on data arrivals. We describe a performance evaluation that shows important performance gains in several configurations. Luc Bouganim, Françoise Fabret, C. Mohan 0001, Patrick Valduriez |
ICDE | 4 |
| 2000 | PicoDMBS: Scaling Down Database Techniques for the Smartcard
Christophe Bobineau, Luc Bouganim, Philippe Pucheral, Patrick Valduriez |
VLDB | 4 |
| 2000 | Caching Strategies for Data-Intensive Web Sites
Khaled Yagoub, Daniela Florescu, Valérie Issarny, Patrick Valduriez |
VLDB | 4 |
| 2000 | Building and Customizing Data-Intensive Web Sites Using Weave
Khaled Yagoub, Daniela Florescu, Valérie Issarny, Patrick Valduriez |
VLDB | 4 |
| 1999 | Load Balancing for Parallel Query Execution on NUMA Multiprocessors
Luc Bouganim, Daniela Florescu, Patrick Valduriez |
Distributed Parallel Databases | 3 |
| 1998 | Memory-Adaptive Scheduling for Large Query ExecutionabstractArticle Free Access Share on Memory-adaptive scheduling for large query execution Authors: Luc Bouganim PRiSM, Versailles, France PRiSM, Versailles, FranceView Profile , Olga Kapitskaia INRIA, Rocquencourt, France INRIA, Rocquencourt, FranceView Profile , Patrick Valduriez INRIA, Rocquencourt, France INRIA, Rocquencourt, FranceView Profile Authors Info & Claims CIKM '98: Proceedings of the seventh international conference on Information and knowledge managementNovember 1998 Pages 105–115https://doi.org/10.1145/288627.288646Online:01 November 1998Publication History 13citation347DownloadsMetricsTotal Citations13Total Downloads347Last 12 Months9Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Luc Bouganim, Olga Kapitskaia, Patrick Valduriez |
CIKM | 3 |
| 1998 | Scaling Access to Heterogeneous Data Sources with DISCOabstractAccessing many data sources aggravates problems for users of heterogeneous distributed databases. Database administrators must deal with fragile mediators, that is, mediators with schemas and views that must be significantly changed to incorporate a new data source. When implementing translators of queries from mediators to data sources, database implementers must deal with data sources that do not support all the functionality required by mediators. Application programmers must deal with graceless failures for unavailable data sources. Queries simply return failure and no further information when data sources are unavailable for query processing. The Distributed Information Search COmponent (Disco) addresses these problems. Data modeling techniques manage the connections to data sources, and sources can be added transparently to the users and applications. The interface between mediators and data sources flexibly handles different query languages and different data source functionality. Query rewriting and optimization techniques rewrite queries so they are efficiently evaluated by sources. Query processing and evaluation semantics are developed to process queries over unavailable data sources. In this article, we describe: 1) the distributed mediator architecture of Disco; 2) the data model and its modeling of data source connections; 3) the interface to underlying data sources and the query rewriting process; and 4) query processing semantics. We describe several advantages of our system. Anthony Tomasic, Louiqa Raschid, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1997 | Concurrent Garbage Collection in O2
Marcin Skubiszewski, Patrick Valduriez |
VLDB | 2 |
| 1996 | Adaptive Parallel Query Execution in DBS3
Luc Bouganim, Benoît Dageville, Patrick Valduriez |
EDBT | 3 |
| 1996 | Scaling Heterogeneous Databases and the Design of DiscoabstractAccess to large numbers of data sources introduces new problems for users of heterogeneous distributed databases. End users and application programmers must deal with unavailable data sources. Database administrators must deal with incorporating new sources into the model. Database implementers must deal with the translation of queries between query languages and schemas. The Distributed Information Search COmponent (Disco) addresses these problems. Query processing semantics are developed to process queries over data sources which do not return answers. Data modeling techniques manage connections to data sources. The component interface to data sources flexibly handles different query languages and translates queries. This paper describes (a) the distributed mediator architecture of Disco, (b) its query processing semantics, (C) the data model and its modeling of data source connections, and (d) the interface to underlying data sources. Anthony Tomasic, Louiqa Raschid, Patrick Valduriez |
ICDCS | 3 |
| 1996 | Dynamic Load Balancing in Hierarchical Parallel Database Systems
Luc Bouganim, Daniela Florescu, Patrick Valduriez |
VLDB | 3 |
| 1996 | A Methodology for Query Reformulation in CIS Using Semantic KnowledgeabstractWe consider Cooperative Information Systems (CIS) that are multidatabase systems (MDBMS), with a common object-oriented model, based on the ODMG standard, together with local databases that may be relational, object-oriented, or dedicated data servers. The MDBMS interface (or mediator interface) that describes this CIS could be different from the union of the local interfaces that describe each local database. In particular, the mediator interface may be defined by semantic knowledge that includes views over particular local databases, integrity constraints, and knowledge about data replication in local databases. We present a methodology for query reformulation which is based on the uniform representation of all semantic knowledge in the form of integrity assertions and mapping rules. A reformulation algorithm exploits this semantic knowledge, and performs semantic rewriting based on pattern-matching, to obtain a query on the union of the local interfaces. A decomposition algorithm then produces a composite query, and local sub-queries, one for each local interface. The reformulation is general enough to re-use the results of previously computed queries in the CIS. We have implemented this reformulation technique in our Flora compiler prototype which we used for validation and experimentation with O2 databases. Daniela Florescu, Louiqa Raschid, Patrick Valduriez |
Int. J. Cooperative Inf. Syst. | 3 |
| 1995 | Using Heterogeneous Equivalences for Query Rewriting in Multidatabase Systems
Daniela Florescu, Louiqa Raschid, Patrick Valduriez |
CoopIS | 3 |
| 1995 | Locking in OODBMS Client Supported Nestd TransactionsabstractNested transactions facilitate the control of complex persistent applications by enabling both fine-tuning of the scope of rollback and safe intra-transaction parallelism. We are concerned with supporting concurrent nested transactions on client workstations of an OODBMS. Use of the traditional design and implementation of a lock manager results in a high CPU overhead: in-cache traversals of the 007 benchmark perform, at best, 4.5 times slower than the same traversal achieved in virtual memory by a nonpersistent programming language. We propose a new design and implementation of a lock manager which cuts that factor down to 1.8. This lock manager supports nested transactions with both sibling and parent/child parallelisms, and provides object locking at a cost comparable to page locking. Object locking is therefore a better alternative due to its higher functionality.> Laurent Daynès, Olivier Gruber, Patrick Valduriez |
ICDE | 3 |
| 1995 | Design and Implementation of Flora, A Language for Object Algebra
Daniela Florescu, Jean-Robert Gruser, Michael Novak, Patrick Valduriez, Mikal Ziane |
Inf. Sci. | 4 |
| 1995 | Transaction Chopping: Algorithms and Performance StudiesabstractChopping transactions into pieces is good for performance but may lead to nonserializable executions. Many researchers have reacted to this fact by either inventing new concurrency-control mechanisms, weakening serializability, or both. We adopt a different approach. We assume a user who —has access only to user-level tools such as (1) choosing isolation degrees 1ndash;4, (2) the ability to execute a portion of a transaction using multiversion read consistency, and (3) the ability to reorder the instructions in transaction programs; and —knows the set of transactions that may run during a certain interval (users are likely to have such knowledge for on-line or real-time transactional applications). Given this information, our algorithm finds the finest chopping of a set of transactions TranSet with the following property: If the pieces of the chopping execute serializably, then TranSet executes serializably . This permits users to obtain more concurrency while preserving correctness. Besides obtaining more intertransaction concurrency, chopping transactions in this way can enhance intratransaction parallelism. The algorithm is inexpensive, running in O(n×(e+m)) time, once conflicts are identified, using a naive implementation, where n is the number of concurrent transactions in the interval, e is the number of edges in the conflict graph among the transactions, and m is the maximum number of accesses of any transaction. This makes it feasible to add as a tuning knob to real systems. Dennis E. Shasha, François Llirbat, Eric Simon, Patrick Valduriez |
ACM Trans. Database Syst. | 4 |
| 1994 | Towards Persistent Object Systems for Desktop Computing
Olivier Gruber, Patrick Valduriez |
CoopIS | 2 |
| 1994 | Flora: A Functional-Style Language for Object and relational Algebra
Michael Novak, Georges Gardarin, Patrick Valduriez |
DEXA | 3 |
| 1994 | Invited Project Review: Industrial-strength parallel query optimization: issues and lessons
Rosana S. G. Lanzelotte, Patrick Valduriez, Mohamed Zaït, Mikal Ziane |
Inf. Syst. | 2 |
| 1993 | Parallel Database Systems: the case for shared-somethingabstractParallel database systems are becoming the primary application of multiprocessor computers. The reason for this is that they can provide high-performance and high-availability database support at a much lower price than do equivalent mainframe computers. The traditional shared-memory, shared-disk, and shared-nothing architectures of parallel database systems are compared, based on the following dimensions: simplicity, cost, performance, availability and extensibility. Based on these comparisons, the case is made for the shared-something architecture, which can provide a better trade-off between the various objectives.> Patrick Valduriez |
ICDE | 1 |
| 1993 | On the Effectiveness of Optimization Search Strategies for Parallel Execution Spaces
Rosana S. G. Lanzelotte, Patrick Valduriez, Mohamed Zaït |
VLDB | 2 |
| 1993 | Overview of Parallel Architectures for DatabasesabstractWe present and compare hardware and system architectures for databases that take advantage of parallelism. First, we state the problem and identify the main comparison criterions (price, performance, extensibility, data availability). Then, we review the major architecture classes (shared nothing, shared everything, shared disks, hybrid) and discuss their advantages and drawbacks for different kinds of workloads. Björn Bergsten, Michel Couprie, Patrick Valduriez |
Comput. J. | 3 |
| 1993 | Parallel Database Systems: Open Problems and New Issues
Patrick Valduriez |
Distributed Parallel Databases | 1 |
| 1992 | Schema Extensions in OODB
Marie-Jo Bellosta, Patrick Valduriez, Fabienne Viallet |
DEXA | 2 |
| 1992 | ESQL2: An Object-Oriented SQL with F-Logic SemanticsabstractESQL2 is an SQL2 upward-compatible database language that integrates the essential concepts of relational, object-oriented, and deductive databases. ESQL2's salient features are a rich and extendible type system based on abstract data types (ADTs) implemented in various programming languages, complex objects with object sharing by combining generic ADTs and object identity, the capability of querying and updating relations containing simple or complex objects using SQL-compatible syntax and semantics, and a DATALOG-like deductive capability provided as an extension of the SQL view mechanism. A declarative semantics is proposed for ESQL2 retrieval statements using F-Logic, which provides a solid basis for understanding the integration of objects and relations.> Georges Gardarin, Patrick Valduriez |
ICDE | 2 |
| 1992 | Optimization of Object-Oriented Recursive Queries using Cost-Controlled StrategiesabstractObject-oriented data models are being extended with recursion to gain expressive power. This complicates the optimization problem which has to deal with recursive queries on complex objects. Because unary operations invoking methods or path expressions on objects may be costly to execute, traditional heuristics for optimizing recursive queries are no longer valid. In this paper we propose a cost-based optimization method which handles object-oriented recursive queries. In particular, it is able to delay the decision of pushing selective operations through recursion until the effect of such a transformation can be measured by a cost model. The approach integrates rewriting and increases the optimization opportunities for recursive queries on objects while allowing for efficient optimization. Rosana S. G. Lanzelotte, Patrick Valduriez, Mohamed Zaït |
SIGMOD Conference | 2 |
| 1992 | Simple Rational Guidance for Chopping Up TransactionsabstractChopping transactions into pieces is good for performance but may lead to non-serializable executions. Many researchers have reacted to this fact by either inventing new concurrency control mechanisms, weakening serializability, or both. We adopt a different approach. Dennis E. Shasha, Eric Simon, Patrick Valduriez |
SIGMOD Conference | 3 |
| 1992 | SVP: A Model Capturing Sets, Lists, Streams, and Parallelism
Douglas Stott Parker Jr., Eric Simon, Patrick Valduriez |
VLDB | 3 |
| 1992 | Object-Oriented Database Systems
Patrick Valduriez |
VLDB | 1 |
| 1992 | The data model of FAD, a database programming language
Scott Danforth, Patrick Valduriez |
Inf. Sci. | 2 |
| 1992 | Functional SOL (FSOL), an SQL upward-compatible database programming language
Patrick Valduriez, Scott Danforth |
Inf. Sci. | 1 |
| 1992 | A FAD for Data Intensive ApplicationsabstractFAD is a strongly typed database programming language designed for uniformly manipulating transient and persistent data on Bubba, a parallel database system developed at MCC. The paper provides an overall description of FAD, and discusses the design rationale behind a number of its distinguishing features. Comparisons with other database programming languages are provided.> Stan Danforth, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1991 | Extending the Search Strategy in a Query Optimizer
Rosana S. G. Lanzelotte, Patrick Valduriez |
VLDB | 2 |
| 1990 | ESQL: An Extended SQL with Object and Deductive Capabilities
Georges Gardarin, Patrick Valduriez |
DEXA | 2 |
| 1990 | Efficient Main Memory Data Management Using the DBGraph Storage Model
Philippe Pucheral, Jean-Marc Thévenin, Patrick Valduriez |
VLDB | 3 |
| 1990 | Prototyping Bubba, A Highly Parallel Database SystemabstractBubba is a highly parallel computer system for data-intensive applications. The basis of the Bubba design is a scalable shared-nothing architecture which can scale up to thousands of nodes. Data are declustered across the nodes (i.e. horizontally partitioned via hashing or range partitioning) and operations are executed at those nodes containing relevant data. In this way, parallelism can be exploited within individual transactions as well as among multiple concurrent transactions to improve throughput and response times for data-intensive applications. The current Bubba prototype runs on a commercial 40-node multicomputer and includes a parallelizing compiler, distributed transaction management, object management, and a customized version of Unix. The current prototype is described and the major design decisions that went into its construction are discussed. The lessons learned from this prototype and its predecessors are presented.> Haran Boral, William Alexander, Larry Clay, George P. Copeland, Scott Danforth, Michael J. Franklin, Brian E. Hart, Marc G. Smith, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 9 |
| 1988 | Parallel Query Processing for Complex ObjectsabstractThe authors investigate a direct storage scheme for complex objects called FIHSM (Fully Inverted Hierarchical Storage Model) and propose a novel parallel-query-processing strategy (QPS) for it. The QPS has four phases (select, pivot, value materialize and compose). With a declustered placement strategy, each of these phases provides for both inter and intra-operation parallelism. Furthermore, partial results of one phase could be pipelined to the subsequent phase. The proposed four-phase structured algorithm is based on heuristics and thus avoids the prohibitive exhaustive searches which are needed for optimizing query executions in parallel environments.> Setrag Khoshafian, Patrick Valduriez, George P. Copeland |
ICDE | 2 |
| 1987 | A Query Processing Strategy for the Decomposed Storage ModelabstractHandling parallelism in database systems involves the specification of a storage model, a placement strategy, and a query processing strategy. An important goal is to determine the appropriate combination of these three strategies in order to obtain the best performance advantages. In this paper we present a novel and promising query processing strategy for a decomposed storage model. We discuss some of the qualitative advantages of the scheme. We also compare the performance of the proposed “pivot” strategy with conventional query processing for the n-ary storage model. The comparison is performed using the Wisconsin Benchmarks. Setrag Khoshafian, George P. Copeland, Thomas Jagodis, Haran Boral, Patrick Valduriez |
ICDE | 5 |
| 1987 | FAD, a Powerful and Simple Database Language
François Bancilhon, Ted Briggs, Setrag Khoshafian, Patrick Valduriez |
VLDB | 4 |
| 1987 | Join IndicesabstractIn new application areas of relational database systems, such as artificial intelligence, the join operator is used more extensively than in conventional applications. In this paper, we propose a simple data structure, called a join index, for improving the performance of joins in the context of complex queries. For most of the joins, updates to join indices incur very little overhead. Some properties of a join index are (i) its efficient use of memory and adaptiveness to parallel execution, (ii) its compatibility with other operations (including select and union), (iii) its support for abstract data type join predicates, (iv) its support for multirelation clustering, and (v) its use in representing directed graphs and in evaluating recursive queries. Finally, the analysis of the join algorithm using join indices shows its excellent performance. Patrick Valduriez |
ACM Trans. Database Syst. | 1 |
| 1986 | Buffering Schemes for Permanent DataabstractThe availability of larger RAM spaces for DBMSs provides interesting opportunities for performance enhancements, especially in buffer management. In this paper we propose and compare two alternative strategies for the buffer management of permanent data (i.e., the data committed by transactions) called block buffering and attribute buffering. These strategies use statistics to capture the changing locality of a reference string. We model and demonstrate the impact of locality on the performance of buffering. We also analyze and compare the effect of both the attribute and predicate dimensions of locality on buffering, varying a number of parameters including the degree of locality, RAM size, and RAM utilization. George P. Copeland, Setrag Khoshafian, Marc G. Smith, Patrick Valduriez |
ICDE | 4 |
| 1986 | Implementation Techniques of Complex Objects
Patrick Valduriez, Setrag Khoshafian, George P. Copeland |
VLDB | 1 |
| 1984 | Predicate Trees: An Approach to Optimize Relational Query OperationsabstractWith the advent of relational database systems, multi-key searching problems have became the focus of a great deal of research. In this paper, we present a new data structure called predicate trees for clustering tuples of a relation in a way that allows the system to accelerate a large number of multi-dimensional queries. The directories used for the implementation of predicate trees in SABRE are organized as relations which are searched efficiently by filters. One of the most significant advantages of predicate trees is the possibility of defining logical addresses based on content, called signatures, and to use filters to manage directories. Georges Gardarin, Patrick Valduriez, Yann Viémont |
ICDE | 2 |
| 1984 | Design and Implementation of an Extendible Integrity SubsystemabstractTnls paper presents a powerful integrrty subsystem, which is implemented in the SABRE database system. The specification language is simple. Tne enforcement algorithm is general, in particular, it handles referential dependency and temporal assertions. Specialized strategies efficiently treat each class of assertions. The system automatically manages integrity checkpoints. Also, an efficient method is described for processing assertions involving aggregates. An analysis exhibits the value of the algorithms. It 1s shown that, in general, this method is better than the query modification method for domain assertions. Measures have also been done for giving the cost added for controlling integrity in comparison with the cost of the request itself. Eric Simon, Patrick Valduriez |
SIGMOD Conference | 2 |
| 1984 | A Multikey Hashing Scheme Using Predicate TreesabstractA new method for multikey access suitable for dynamic files is proposed that transforms multiple key values into a logical address This method is based on a new structure, called predicate tree, that represents the function applied to several keys A predicate tree permits to specify in a unified way various hashing schemes by allowing for different definitions of predicates A logical address qualifies a space partition of a file according to its predicate tree This address is seen as a single key by a digital hashing method which transforms it into a physical address This method is used to address records in a file and to transform a retrieval qualification on a file into a set of partitions to access Finally, a qualitative analysis of the behavior of the method is given which exhibits its value Patrick Valduriez, Yann Viémont |
SIGMOD Conference | 1 |
| 1984 | Join and Semijoin Algorithms for a Multiprocessor Database MachineabstractThis paper presents and analyzes algorithms for computing joins and semijoins of relations in a multiprocessor database machine. First, a model of the multiprocessor architecture is described, incorporating parameters defining I/O, CPU, and message transmission times that permit calculation of the execution times of these algorithms. Then, three join algorithms are presented and compared. It is shown that, for a given configuration, each algorithm has an application domain defined by the characteristics of the operand and result relations. Since a semijoin operator is useful for decreasing I/O and transmission times in a multiprocessor system, we present and compare two equi-semijoin algorithms and one non-equi-semijoin algorithm. The execution times of these algorithms are generally linearly proportional to the size of the operand and result relations, and inversely proportional to the number of processors. We then compare a method which consists of joining two relations to a method whereby one joins their semijoins. Finally, it is shown that the latter method, using semijoins, is generally better. The various algorithms presented are implemented in the SABRE database system; an evaluation model selects the best algorithm for performing a join according to the results presented here. A first version of the SABRE system is currently operational at INRIA. Patrick Valduriez, Georges Gardarin |
ACM Trans. Database Syst. | 1 |
| 1982 | Semi-Join Algorithms for Multiprocessor SystemsabstractSemi-join is a relational operator that decreases the cost of processing queries involving binary operations. This is accomplished by initially selecting the data relevant to answer the queries and thereby reducing the size of the operand relations. This paper presents and analyzes algorithms for computing semi-joins in a multiprocessor database machine. First, an architecture model of a multiprocessor system is described. The model incorporates 1-0, CPU and messages transmission cost parameters to enable the evaluation of these algorithms in terms of their execution costs. Then two equi-semi-join algorithms are presented and compared and one inequi-semi-join algorithm is proposed. The execution cost of these algorithms are generally lineary proportional to the size of the operand and result relations and inversely proportional to the number of processors. Then, the method by joining two relations and the method by joining their semi-joins are compared. Finally it is shown that the method using semi-joins is generally better. Patrick Valduriez |
SIGMOD Conference | 1 |