EDBT 2026 Demo / reviewers in the wild / expert
Jean-Marc Pierson
dblp:08/2911
· DBLP profile ↗
52ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-8948-0474ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 2Computer networks · 2Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sufficiency in Data Centers: Energy Aware Resource Recommendation System
Eyvaz Ahmadzada, Patricia Stolf, Jean-Marc Pierson, Laurent Lefèvre |
Euro-Par (2) | 3 |
| 2026 | MPI malleability validation under replayed real-world HPC conditions
Sergio Iserte, Maël Madon, Georges Da Costa, Jean-Marc Pierson, Antonio J. Peña |
Future Gener. Comput. Syst. | 4 |
| 2025 | Sufficiency power consideration to run a workload on renewable energy operated datacenter
Damien Landré, Laurent Philippe 0001, Jean-Marc Pierson |
Future Gener. Comput. Syst. | 3 |
| 2024 | Seasonal study of user demand and IT system usage in datacentersabstractTo limit the impact of datacenters on climate change various sustainable and effective solutions are being developed. The Datazero2 project aims to design datacenters running solely on local renewable energy combined with storage sources. A multi-hour prevision of user demand (workload) or of IT system usage, using time series, are possible ways to plan energy usage, with the aim of reducing energy consumption while maintaining a required quality of service. The seasonal component of time series is a key element, since it indicates one or more periodic phenomena that can be estimated and implemented in forecasting methods. In this paper, we propose a seasonal study of multiple time series from different workloads. Seasonality is studied using periodograms, a highly effective Fast Fourier Transform technique for detecting seasonal patterns. In addition, since the detected seasonal patterns can be irregular (unexpected, delayed and/or extended user demands or IT system usage), we propose a distribution-based clustering using the Kruskal-Wallis test and a seasonal peak appearances analysis. The results show that most of these time series are multi-seasonal (daily and weekly seasonal patterns) but highly irregular, which can reduce the performance of seasonal forecasting methods. Damien Landré, Laurent Philippe 0001, Jean-Marc Pierson |
ICPADS | 3 |
| 2024 | A Review of Data Placement and Replication Strategies Based on Machine LearningabstractThe global increase in data volumes has brought forth the need for scalable distributed systems that can provide satisfactory quality of service. Data placement and replication are well known techniques that provide increased performance, improved fault tolerance and higher availability. These techniques often require threshold-based activation mechanisms that can vary due to the nature of the workload and the underlying system architecture. Hence, setting and adjusting those thresholds usually require human intervention. In this context, machine learning presents a promising facet to automatically define such thresholds to adapt to different workloads and architectures. In this paper, we study the data placement and replication strategies proposed in the literature that employ machine learning. We classify such strategies based on the machine learning method, the platform on which they are deployed, the dynamicity and the achieved objectives. We describe the approach applied by each strategy as well as possible limitations. In addition, we provide insights into metrics used to evaluate the strategies. We highlight the need to design data placement and replication strategies that respond better to modern needs for distributed systems. We also motivate the use of machine learning to achieve autonomy in distributed systems. Amir Najjar, Riad Mokadem, Jean-Marc Pierson |
ICPADS | 3 |
| 2024 | Replay with Feedback: How does the performance of HPC system impact user submission behavior?abstractHigh Performance Computing (HPC) is a key infrastructure to solve large scale scientific problems, from weather to quantum simulations. Scheduling jobs in HPC infrastructures is complex due to their scale, the different behaviors of their users, and the multiple objectives, from performance to ecological impact. Schedulers are evaluated on data center simulations, due to the complexity and cost of evaluating them in-situ. One key element for this evaluation is the behavioral model of users. Most studies are limited to replaying past workload of existing data centers. This reduces the realism of performance evaluation in cases where the scheduler and the hardware infrastructure are not exactly the same. Any such change would potentially impact the behavior of the users. In this article we introduce a novel model “Replay with Feedback” accounting for the impact of HPC system performances on user submission behavior in simulations. Instead of keeping the original timestamps of job submissions, we exhibit and use the relationships between each user jobs. We propose an open-source implementation of this model along with an extensive and reproducible set of experiments to assess the impact of the scheduler and infrastructure changes. We also provide new metrics adapted to the flexibility of user submission behaviors. Results show that using this model, we advance towards more realistic simulations of schedulers in HPC systems. Maël Madon, Georges Da Costa, Jean-Marc Pierson |
Future Gener. Comput. Syst. | 3 |
| 2023 | Assessing Power Needs to Run a Workload with Quality of Service on Green Datacenters
Louis-Claude Canon, Damien Landré, Laurent Philippe 0001, Jean-Marc Pierson, Paul Renaud-Goud |
Euro-Par | 4 |
| 2022 | Dynamic Energy and Expenditure Aware Data Replication StrategyabstractNowadays, data are generated and accessed from all over the world. Applications and companies store these data on geo-distributed cloud providers that have to be profitable while reducing their environmental impact. As Cloud providers aim to satisfy their service level objectives in terms of availability and response time, they often rely on data replication. In this paper, we propose a dynamic data replication strategy (DE2ARS) that adapts the number of replicas according to the workload and addresses energy consumption and expenditure issues. It follows an initial placement and is triggered by a Control Chart. We first compare different parameter choices in order to provide better analysis for the proposed strategy. We compare DE2ARS with strategies from the literature. Results highlight the fact that DE2ARS reaches its goal to reduce both energy consumption and expenditure while having good performance in a strong workload context. Morgan Séguéla, Riad Mokadem, Jean-Marc Pierson |
CLOUD | 3 |
| 2022 | Characterization of Different User Behaviors for Demand Response in Data Centers
Maël Madon, Georges Da Costa, Jean-Marc Pierson |
Euro-Par | 3 |
| 2021 | Energy and Expenditure Aware Data Replication StrategyabstractEnergy saving is a major challenge for Information Technology (IT) companies that aim to reduce their carbon footprint while providing large scale cloud services. These companies often rely on data replication technique in order to satisfy tenant's objectives, e.g., performance, especially with the increasing volume of data distributed throughout the world. In this paper, we propose a static and multi objective data replication strategy (E2ARS) that aims to reduce both energy consumption and expenditure of the provider. E2ARS leverages on cloud heterogeneity and energy efficient technologies. We first compare different policies of our strategy, from only taking energy consumption into account to only taking expenditure into account. Unsurprisingly, the more you want to reduce the energy consumption, the less you replicate. Then, we compare E2ARS with strategies from the literature. E2ARS reduces both energy consumption and expenditure where those strategies satisfy only one of the two objectives. Morgan Séguéla, Riad Mokadem, Jean-Marc Pierson |
CLOUD | 3 |
| 2021 | Horizontal Scaling in Cloud Using Contextual Bandits
David Delande, Patricia Stolf, Raphaël Féraud, Jean-Marc Pierson, André Bottaro |
Euro-Par | 4 |
| 2020 | Fast maximum coverage of system behavior from a performance and power point of view
Georges Da Costa, Jean-Marc Pierson, Leandro F. Cupertino |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Negotiation game for joint IT and energy management in green datacenters
Minh-Thuyen Thi, Jean-Marc Pierson, Georges Da Costa, Patricia Stolf, Jean-Marc Nicod, Gustavo Rostirolla, Marwa Haddad |
Future Gener. Comput. Syst. | 2 |
| 2019 | A Discrete Particle Swarm Optimization Approach for Energy-Efficient IoT Services Placement Over Fog InfrastructuresabstractThe Internet of Things (IoT) encompasses both large-scale deployed physical infrastructures and software layers that enable intuitive and transparent creation of applications. This highly distributed, energy-greedy environment must ensure the quality of deployed services while taking into account the heterogeneity of capabilities and protocols as well as users and objects mobility. Deployment infrastructure has been redesigned to provide the necessary features, including paradigms such as software-defined networks and Fog computing. The purpose of this article is to study IoT services placement in a Fog architecture. We propose a model of the infrastructure and IoT applications as well as a placement strategy taking into account system's energy consumption and applications delay violations minimization with a Discrete Particles Swarm Optimization algorithm (DPSO). Simulations have been done with iFogSim simulator. Results have been compared with heuristics coming from the literature: Binary Partical Swarm optimization (BPSO), Dicothomous Module Mapping (DCT), CloudOnly, IoTFogOnly, IoTCloud (IC) and FogCloud (FC) placement approaches. Tanissia Djemai, Patricia Stolf, Thierry Monteil 0001, Jean-Marc Pierson |
ISPDC | 4 |
| 2019 | Hot-N-Cold model for energy aware cloud databases
Chaopeng Guo, Jean-Marc Pierson, Jie Song 0001, Christina Herzog |
J. Parallel Distributed Comput. | 2 |
| 2018 | A Generic Learning Multi-agent-System Approach for Spatio-Temporal-, Thermal- and Energy-Aware SchedulingabstractThis paper proposes an agent based approach to the scheduling of jobs in data centers under thermal constraints. The model encompasses both temporal and spatial aspects of the temperature evolution using a unified model, taking into account the dynamics of heat production and dissipation. Agents coordinate to eventually move jobs to the best suitable place and to adapt dynamically the frequency settings of the nodes to the best combination. Several objectives of the agents are compared under different circumstances by an extensive set of experiments. Christina Herzog, Jean-Marc Pierson |
PDP | 2 |
| 2018 | Frequency Selection Approach for Energy Aware Cloud DatabaseabstractA lot of cloud systems are adopted in industry and academia to face the explosion of the data volume and the arrival of the big data era. Meanwhile, energy efficiency and energy saving become major concerns for data centers where massive cloud systems are deployed. However, energy waste is quite common due to resource over-provisioning. In this paper, using Dynamic Voltage and Frequency Scaling (DVFS), a frequency selection approach is introduced to improve the energy efficiency of cloud systems in terms of resource over-provisioning. In the approach, two algorithms, Genetic Algorithm (GA) and Monte Carlo Tree Search Algorithm (MCTS), are proposed. Cloud database system is taken as an example to evaluate the approach. The results of the experiments show that the algorithms have great scalability which can be applied to a 120-nodes case with high accuracy compared to optimal solutions (up to 99.9% and 99.6% for GA and MCTS respectively). According to an optimality bound analysis, 21 % of energy can be saved at most using our frequency selection approach. Chaopeng Guo, Jean-Marc Pierson |
SBAC-PAD | 2 |
| 2018 | Modulo Based Data Placement Algorithm for Energy Consumption Optimization of MapReduce System
Jie Song 0001, HongYan He, Ge Yu 0001, Jean-Marc Pierson |
J. Grid Comput. | 5 |
| 2017 | Spatio-temporal thermal-aware scheduling for homogeneous high-performance computing datacenters
Hongyang Sun 0001, Patricia Stolf, Jean-Marc Pierson |
Future Gener. Comput. Syst. | 3 |
| 2016 | Dynamically Building Energy Proportional Data Centers with Heterogeneous Computing ResourcesabstractAs the number of data centers increases, it is urgent to reduce their energy consumption. Although servers are becoming more energy-efficient, their idle consumption remains high, which is an issue as data centers are often over-provisioned. This work proposes a novel approach for building data centers with heterogeneous machines carefully chosen for their performance and energy efficiency ratios. We focus on web applications whose load varies over time, and design a scheduler that dynamically reconfigures the infrastructure, by migrating applications and switching machines on or off, so that the energy consumed by the data center is proportional to the load. Experiments evaluate the approach and show the energy savings achieved by our heterogeneous data center design and management while satisfying Quality of Service (QoS) constraints. Violaine Villebonnet, Georges Da Costa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf |
CLUSTER | 4 |
| 2016 | Energy Proportionality in Heterogeneous Data Center Supporting Applications with Variable LoadabstractThe increasing number of data centers raises serious concerns regarding their energy consumption. Although servers have become more energy-efficient over time, their idle consumption remains high, which is an issue as resources in data centers are often over-provisioned. This work proposes a novel approach for building data centers so that their energy consumption is proportional to load. A data center hence comprises heterogeneous machines carefully chosen for their performance and energy efficiency ratios. We focus on web applications whose load varies over time and design a scheduler that dynamically reconfigures the infrastructure to minimize its energy consumption according to current load and application requirements. Based on load forecasts, it takes reconfiguration decisions and performs actions such as migrating applications and switching machines on or off. The approach is evaluated considering a data center with heterogeneous resources, and the experiments show how to adjust the parameters of scheduling policies to save the most energy while satisfying Quality of Service (QoS) constraints. Violaine Villebonnet, Georges Da Costa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf |
ICPADS | 4 |
| 2016 | A Multi Agent System for Understanding the Impact of Technology Transfer Offices in Green-IT
Christina Herzog, Jean-Marc Pierson, Laurent Lefèvre |
PRIMA | 2 |
| 2016 | Energy Aware Dynamic Provisioning for Heterogeneous Data CentersabstractThe huge amount of energy consumed by data centers represents a limiting factor in their operation. Many of these infrastructures are over-provisioned, thus a significant portion of this energy is consumed by inactive servers staying powered on even if the load is low. Although servers have become more energy-efficient over time, their idle power consumption remains still high. To tackle this issue, we consider a data center with an heterogeneous infrastructure composed of different machine types - from low power processors to classical powerful servers - in order to enhance its energy proportionality. We develop a dynamic provisioning algorithm which takes into account the various characteristics of the architectures composing the infrastructure: their performance, energy consumption and on/off reactivity. Based on future load information, it makes intelligent decisions of resource reconfiguration that impact the infrastructure at multiple terms. Our algorithm is reactive to load evolutions and is able to respect a perfect Quality of Service (QoS) while being energy-efficient. We evaluate our original approach with profiling data from real hardware and the experiments show that our dynamic provisioning brings significant energy savings compared to classical data centers operation. Violaine Villebonnet, Georges Da Costa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf |
SBAC-PAD | 4 |
| 2015 | Application-Agnostic Framework for Improving the Energy Efficiency of Multiple HPC SubsystemsabstractThe subsystems that compose a HPC platform (e.g. CPU, memory, storage and network) are often designed and configured to deliver exceptional performance to a wide range of workloads. As a result, a large part of the power that these subsystems consume is dissipated as heat even when executing workloads that do not require maximum performance. Attempts to tackle this problem include technologies whereby operating systems and applications can reconfigure subsystems dynamically, such as by using DVFS for CPUs, LPI for network components, and variable disk spinning for HDDs. Most previous work has explored these technologies individually to optimise workload execution and reduce energy consumption. We propose a framework that performs on-line analysis of an HPC system in order to identify application execution patterns without a priori information of their workload. The framework takes advantage of reoccurring patterns to reconfigure multiple subsystems dynamically and reduce overall energy consumption. Performance evaluation was carried out on Grid'5000 considering both traditional HPC benchmarks and real-life applications. Ghislain Landry Tsafack Chetsa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf, Georges Da Costa |
PDP | 3 |
| 2015 | DVFS Governor for HPC: Higher, Faster, GreenerabstractIn High Performance Computing, being respectful of the environment is usually secondary compared to performance: The faster, the better. As Exascale computing is in the spotlight, electric power concerns arise as current exascale projects might need too much power to even boot. A recent incentive (Exascale at maximum 20MW) shows that reality is catching up with HPC center designers. Beyond classical works on hardware infrastructure or at the middleware level, we do believe that system-level solutions have great potential for energy reduction. Moreover energy-reduction has often been neglected by the HPC community that focus mainly on raw computing performance. In the literature, energy savings is achieved mainly by two means: Either processor load is the only metric taken into account to reduce processors frequency and to ensure no impact on raw performances, Or processor frequency is managed only at task level outside the critical path. In this article we show that designing and implementing a DVFS (Dynamic Voltage and Frequency Scaling) mechanism based on instantaneous system values (here network activity) can save up to 25% of energy consumption while reducing marginally performance. In several cases, reducing energy consumption also leads to an increase in performances because of the thermal budget of recent processors. This work is validated with real experiments on a Linux cluster using the NAS Parallel Benchmark (NPB). Georges Da Costa, Jean-Marc Pierson |
PDP | 2 |
| 2015 | Energy-efficient, thermal-aware modeling and simulation of data centers: The CoolEmAll approach and evaluation results
Leandro F. Cupertino, Georges Da Costa, Ariel Oleksiak, Wojciech Piatek, Jean-Marc Pierson, Jaume Salom, Laura Siso, Patricia Stolf, Hongyang Sun 0001, Thomas Zilio |
Ad Hoc Networks | 5 |
| 2015 | HaoLap: A Hadoop based OLAP system for big data
Jie Song 0001, Chaopeng Guo, Yichan Zhang, Ge Yu 0001, Jean-Marc Pierson |
J. Syst. Softw. | 6 |
| 2014 | Multi-objective Scheduling for Heterogeneous Server Systems with Machine PlacementabstractHeterogeneous servers are becoming prevalent in many high-performance computing environments, including clusters and data enters. In this paper, we consider multi-objective scheduling for heterogeneous server systems to optimize simultaneously the application performance, energy consumption and thermal imbalance. First, a greedy online framework is presented to allow the scheduling decisions to be made based on any well-defined cost function. To tackle the possibly conflicting objectives, we propose a fuzzy-based priority approach for exploring the tradeoffs of two or more objectives at the same time. Moreover, we present a heuristic algorithm for the static placement of physical machines in order to reduce the maximum temperature at the server outlets. Extensive simulations based on an emerging class of high-density server system have demonstrated the effectiveness of our proposed approach and heuristics in optimizing multiple objectives while achieving better thermal balance. Hongyang Sun 0001, Patricia Stolf, Jean-Marc Pierson, Georges Da Costa |
CCGRID | 3 |
| 2014 | Introduction to special issue on selected papers from Energy Efficiency in Large-Scale Distributed Systems 2013 conferenceabstractWe are happy to present you this special issue dedicated to Energy Efficiency in Large-Scale Distributed Systems conference (EE-LSDS2013) of Concurrency and Communications: Practice and Experience. Jean-Marc Pierson, Lars Dittmann, Georges Da Costa |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Exploiting performance counters to predict and improve energy performance of HPC systems
Ghislain Landry Tsafack Chetsa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf, Georges Da Costa |
Future Gener. Comput. Syst. | 3 |
| 2013 | Cooperative Scheduling Anti-load Balancing Algorithm for Cloud: CSAACabstractIn the past decade, more and more attention focuses on job scheduling strategies in a variety of scenarios. Due to the characteristics of clouds, meta-scheduling turns out to be an important scheduling pattern because it is responsible for orchestrating resources managed by independent local schedulers and bridges the gap between participating nodes. Likewise, to overcome issues such as bottleneck, overloading, under loading and impractical unique administrative management, which are normally led by conventional centralized or hierarchical schemes, the distributed scheduling scheme is emerging as a promising approach because of its capability with regards to scalability and flexibility. In this paper, we introduce a decentralized dynamic scheduling approach entitled Cooperative scheduling Anti-load balancing Algorithm for cloud (CSAAC). To validate CSAAC we used a simulator which extends the MaGateSim simulator and provides better support to energy aware scheduling algorithms. CSAAC goal is to achieve optimized scheduling performance and energy gain over the scope of overall cloud, instead of individual participating nodes. The extensive experimental evaluation with a real workload dataset shows that, when compared to the centralized scheduling scheme with Best Fit as the meta-scheduling policy, the use of CSAAC can lead to a 30%61% energy gain, and a 20%30% shorter average job execution time in a decentralized scheduling manner without requiring detailed real-time processing information from participating nodes. Cheikhou Thiam, Georges Da Costa, Jean-Marc Pierson |
CloudCom (1) | 3 |
| 2012 | A Runtime Framework for Energy Efficient HPC Systems without a Priori Knowledge of ApplicationsabstractThe rising computing demands of scientific endeavors often require the creation and management of High Performance Computing (HPC) systems for running experiments and processing vast amounts of data. These HPC systems generally operate at peak performance, consuming a large quantity of electricity, even though their workload varies over time. Understanding the behavioral patterns (i.e., phases) of HPC systems during their use is key to adjust performance to resource demand and hence improve the energy efficiency. In this paper, we describe (i) a method to detect phases of an HPC system based on its workload, and (ii) a partial phase recognition technique that works cooperatively with on-the-fly dynamic management. We implement a prototype that guides the use of energy saving capabilities to demonstrate the benefits of our approach. Experimental results reveal the effectiveness of the phase detection method under real-life workload and benchmarks. A comparison with baseline unmanaged execution shows that the partial phase recognition technique saves up to 15% of energy with less than 1% performance degradation. Ghislain Landry Tsafack Chetsa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf, Georges Da Costa |
ICPADS | 3 |
| 2012 | Beyond CPU Frequency Scaling for a Fine-grained Energy Control of HPC SystemsabstractModern high performance computing subsystems (HPC) – including processor, network, memory, and IO – are provided with power management mechanisms. These include dynamic speed scaling and dynamic resource sleeping. Understanding the behavioral patterns of high performance computing systems at runtime can lead to a multitude of optimization opportunities including controlling and limiting their energy usage. In this paper, we present a general purpose methodology for optimizing energy performance of HPC systems considering processor, disk and network. We rely on the concept of execution vector along with a partial phase recognition technique for on-the-fly dynamic management without any a priori knowledge of the workload. We demonstrate the effectiveness of our management policy under two real-life workloads. Experimental results show that our management policy in comparison with baseline unmanaged execution saves up to 24% of energy with less than 4% performance overhead for our real-life workloads. Ghislain Landry Tsafack Chetsa, Laurent Lefèvre, Jean-Marc Pierson, Patricia Stolf, Georges Da Costa |
SBAC-PAD | 3 |
| 2012 | Energy-aware service allocation
Damien Borgetto, Henri Casanova, Georges Da Costa, Jean-Marc Pierson |
Future Gener. Comput. Syst. | 4 |
| 2011 | On the Utility of DVFS for Power-Aware Job Placement in Clusters
Jean-Marc Pierson, Henri Casanova |
Euro-Par (1) | 1 |
| 2011 | 1st International Workshop on Sustainable Internet and Internet for Sustainability (SustaInet 2011abstractWe are pleased to present the proceedings of the First International Workshop on Sustainable Internet and Internet for Sustainability (SustaInet 2011), held in conjunction with WoWMoM 2011. Marco Conti, Giuseppe Anastasi, Jean-Marc Pierson, Marco Ortolani |
WOWMOM | 3 |
| 2009 | The GREEN-NET framework: Energy efficiency in large scale distributed systemsabstractThe question of energy savings has been a matter of concern since a long time in the mobile distributed systems and battery-constrained systems. However, for large-scale non-mobile distributed systems, which nowadays reach impressive sizes, the energy dimension (electrical consumption) just starts to be taken into account. In this paper, we present the GREEN-NET1framework which is based on 3 main components: an ON/OFF model based on an Energy Aware Resource Infrastructure (EARI), an adapted Resource Management System (OAR) for energy efficiency and a trust delegation component to assume network presence of sleeping nodes. Georges Da Costa, Jean-Patrick Gelas, Yiannis Georgiou 0002, Laurent Lefèvre, Anne-Cécile Orgerie, Jean-Marc Pierson, Olivier Richard, K. Sharma |
IPDPS | 6 |
| 2008 | Special Issue: Selection of Best Papers of the VLDB Data Management in Grids Workshop (VLDB DMG 2007)abstractGrid computing has greatly matured since the last decade. It exists now at a widely distributed scale. After 10 years of internationally combined efforts to develop from a vision to existing middlewares, we now face a large number of applications being deployed and taking benefit from the Grids. The need to handle data properly came with the applications, managing data at different semantic levels, from raw data coming from sensors in particle physics to rich data in Healthgrids. Most of the time, data management was ad hoc: The need to handle these data has been mainly seen as a constraint, and little effort has been put in their smart management in the early stages of the grid evolution. Data were present in raw files, or in databases, maybe distributed databases, but with little concerns from the application developer who focused on the core development of the process of the data. In the close past, the Grid community has been developing specialized services to handle data in a simpler and more smart way: Pieces of middlewares have been developed, for instance, to access several data sources with a common programming interface, giving the developer also the possibility of adding treatment on the data retrieved, anonymizing it, caching it, or replicating it. Works on data caching, replication, data integration, and security, to name but a few, have been seen. Efforts have been put to access and process the data, not in planning optimization or distributed balanced queries, which are core distributed database services. The database community has been investigating for a long time the issues related to the distribution of the data sources, data queries, query plan optimization in parallel systems, etc. Among the challenges rising in grid environments, we can cite, among others: dynamicity, reliability, security, data availability and transport, data indexing, search and access, autonomics, etc. These are existing distributed database challenges revisited with respect to the grid paradigm. This special issue of ‘Concurrency and Computation: Practice and Experience’ includes a selection of revised papers presented at the VLDB Data Management in Grids workshop, Vienna, Austria, held on 23rd September 2007. The 2007 edition of the workshop was the third in a raw, after the success of the 1st edition in Trondheim in 2005 and the second in Seoul in 2006. The idea of the workshop collocated with one of the most important conferences in databases (VLDB: Very Large Data Bases) is to bring together the experts from the database and grid communities, discuss, and argue the above-stated challenges. In that sense, the objective of the workshop is fulfilled, with several interesting emerging discussions during the event that may be reflected in this Special Issue. Efficient grid data management and integration is achieved only through a perfect teamwork of all levels in grid databases, i.e. the conceptual level, the middleware level for service integration, and finally the low level for data processing and redistribution. The following papers give a good overview of methods and implementations for efficient data management and integration at these levels. The paper authored by Göres and Dessloch 1 presents the PALADIN Integration Framework that allows the integrated representation of arbitrary data, metadata, and operations. On top of it different integration services can be more efficiently built, as their describing metadata have already been integrated. Rabl et al. 2 deal with the low-level aspects of data integration and management. In dynamic grid environments, automatic data management is one key issue for good performance. The paper proposes a new algorithm for load balancing which adapts the data management to the automatic scaling in the number of nodes used. The second and third papers go ‘middleware’: Kotowski et al. 3 present the GParGRES Grid database, which originally implements a middleware solution among cluster databases to enable data grid services. The architecture scales perfectly. Bassi et al. 4 present a multi-layer solution to tackle the outstanding data to be produced and managed in the forthcoming years. Their solution named Intelligent Network Caching Architecture (INCA) combines new networking entities, dedicated transport protocols, peer-to-peer middleware placement, and access functions in a comprehensive framework. Two papers dealt with replication or file management and access. In 5, Akal et al. presented a new replication scheme for grid databases following a publish/subscribe model. Principally, a middleware over the different grid nodes thus allows one to dynamically control the creation and maintenance of replicas. In 6 Hupfeld et al. analyze how an object-based file system can be adapted to Grid environment to provide an abstraction layer to the distributed file management procedure. Then they describe the XtreemFS architecture and explain how it can serve as an improvement of current data management in Grids. Santos and Koblitz present in 7 some security issues related to the management of replicated metadata catalog in the EGEE European project. They explain how security is handled in the AMGA metadata catalog, and its link and use of common security mechanisms in EGEE and Globus-based Grids (VOMS, GSI). In particular, the authors discuss the authorization policies with replicated data, illustrated with a usecase in the healthgrid field. The last paper in this special issue is authored by Norman W. Paton, invited speaker at the VLDB DMG'07 workshop. In 8, he provides an interesting position paper on the state of the art on functions already automated in the current approaches to adaptive data management (for instance, in database administration, query processing, data integration). He argues about some limitations in the current approaches of autonomic computing for data management (predictability, methodology, composability, and semantics) and highlights new challenging researches to handle automation in all aspects related to data management. These eight articles reflect the common trends of research in Grid-related projects linked with data management. We have seen a particular interest this year on adaptability and data management to make the data management in grids more sensitive to the changes in the infrastructure. We hope you will enjoy this selection of papers as much as we enjoyed them during the workshop and during the last reviewing phase. We would like to take the opportunities to thank all our Program Committee members who helped a lot during the whole process as well as G. Fox for his help to make this special issue possible. Finally, we invite you to participate in the 4th edition of the workshop that will be held at VLDB in Auckland in 2008. Jean-Marc Pierson, Harald Kosch |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Adaptable Distance-Based Decision-Making Support in Dynamic Cross-Grid Environment
Julien Gossa, Jean-Marc Pierson, Lionel Brunie |
Euro-Par | 2 |
| 2007 | Authentication and Access Control Using Trust Collaboration in Pervasive Grid Environments
Rachid Saadi, Jean-Marc Pierson, Lionel Brunie |
GPC | 2 |
| 2007 | X316 Security Toolbox for New Generation of Certificate
Rachid Saadi, Jean-Marc Pierson, Lionel Brunie |
TrustBus | 2 |
| 2007 | Management of a cooperative cache in grids with grid cache servicesabstractAbstract Distributed systems like grids support diverse models of distributed computation and need to operate large data entities in a distributed way. A significant quantity of this data are used only for a limited period of time. Caching is recognized as one of the most effective techniques to manage temporary data and collaborative cache is traditionally proposed to scale cache capabilities in distributed environments. Grid needs to manage dynamically different models of computation with different data access patterns. In this paper, we propose a basic infrastructure for the management of collaborative caches that permits to operate and control dynamically different cache mechanisms and cache schemes in grid. Beside traditional collaborative caching where the cooperation is often limited to data resolution, in our infrastructure the collaborative cache capacities are extended to operate and manage these distributed temporal data. Our proposition is composed of a reference cache model that defines four layers for the management of collaborative cache; an information model that represents the main cache elements and their activity; and a set of operations to request specific tasks to monitor, operate, and coordinate a generic collaborative cache system. Implementation issues of a prototype in Globus Toolkit 4 are discussed. Copyright © 2007 John Wiley & Sons, Ltd. Yonny Cardenas, Jean-Marc Pierson, Lionel Brunie |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Special Issue: Selection of Best Papers of the VLDB Data Management in Grids Workshop (VLDB DMG 2006)abstractGrid computing exists now at a widely distributed scale.After 10 years of internationally combined efforts to develop from a vision to existing middlewares, we face now a large number of applications being deployed and taking benefit from the Grids.Almost all these applications handle data, at different semantic levels, from raw data coming from sensors in particle physics to rich data in the healthgrids.Most of the time, the procedure is ad hoc: The need to handle these data has been mainly seen as a constraint, and little effort has been put in their smart management in the early stages of the grid evolution.Data were present in raw files, or in databases, may be distributed databases, but with little concerns from the application developer who focused (and that is normal) on the core development of the process of the data.In the close past, the Grid community has been developing specialized services to handle data in a simpler and more smart way: OGSA-DAI allows for instance to access several data sources with a common programming interface, giving the developer also the possibility to add treatment on the data retrieved, to anonymize it, to cache it, or to replicate it.Efforts have been put to access and process the data, not in planning optimization or distributed balanced queries, which are core distributed databases services.Leading databases companies have publicized grid-aware databases that mainly are 'old parallel databases' in clusters, not taking much into account the challenges of the grid environment.The database community has been investigating for a long time issues related to the distribution of the data sources, data queries, query plan optimization in parallel systems but has not been really involved up to know in the grid community and development.Among the challenges rising in these environments, we can cite, among others: dynamicity, reliability, security, data availability and transport, data indexing, search and access, etc.These are existing distributed databases challenges revisited with respect to the grid paradigm.This special issue of 'Concurrency and Computation: Practice and Experience' includes a selection of revised papers presented at the VLDB Data Management in Grids workshop, Seoul, Korea, held on 11th September 2006.They illustrate some of the challenges together with current trends of solutions.The 2006 edition of the workshop is the second in a raw, after the success of the 1st edition in Trondheim in 2005.The idea of the workshop collocated with one of the most important conference in databases (VLDB: Very Large Data Bases) is to bring together the experts from the communities to meet, discuss and argue the above stated challenges.In that sense, the objective of the workshop Jean-Marc Pierson, Lionel Brunie |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Special issue: International Conference on Pervasive Services (ICPS 2006)
Laurent Lefèvre, Jean-Marc Pierson |
J. Syst. Softw. | 2 |
| 2006 | Temporal Storage Space for Grids
Yonny Cardenas, Jean-Marc Pierson, Lionel Brunie |
HPCC | 2 |
| 2005 | Grid for Geno-Medicine: a glimpse on the GGM projectabstractThis paper presents briefly the aims and challenges addressed in the GGM (Grid for Geno-Medicine) project. The idea behind the project is to offer a software infrastructure able to analyze and discover links between distributed medical and genetic data. Jean-Marc Pierson, Lionel Brunie, Clarisse Dhaenens, Abdelkader Hameurlain, Nouredine Melab, Maryvonne Miquel, Franck Morvan, El-Ghazali Talbi, Anne Tchounikine |
CCGRID | 1 |
| 2004 | Medical Images Simulation, Storage, and Processing on the European DataGrid Testbed
Johan Montagnat, Fabrice Bellet, Hugues Benoit-Cattin, Vincent Breton, Lionel Brunie, Hector Duque, Yannick Legré, Isabelle E. Magnin, Lydia Maigne, Serge Miguet, Jean-Marc Pierson, Ludwig Seitz, Tiffany Tweed |
J. Grid Comput. | 11 |
| 2003 | DM2: A Distributed Medical Data Manager for GridsabstractMedical data represent tremendous amount of data for which automatic analysis is increasingly needed. Grids are very promising to face today's challenging health issues such as epidemiological studies through large image data sets. However, the sensitive nature of medical data makes it difficult to widely distribute medical applications over computational grids. In this paper, we review fundamental medical data manipulation requirements and we propose a distributed data management architecture that addresses the medical data security and high performance constraints. A prototype is currently being developed inside our laboratories to demonstrate the architecture capability to face realistic distributed medical data manipulation situations. Hector Duque, Johan Montagnat, Jean-Marc Pierson, Lionel Brunie, Isabelle E. Magnin |
CCGRID | 3 |
| 2003 | Semantic Access Control for Medical Applications in Grid Environments
Ludwig Seitz, Jean-Marc Pierson, Lionel Brunie |
Euro-Par | 2 |
| 2002 | Semantic collaborative web cachingabstractToo much information is no information. When the quantity of data becomes too large, users need help and tools to find their path in the information space. Users can browse through Terabytes of data and more, and the difficulty is finding the best way through all possible documents. Our feeling is that a user is not generic and has particular interests. Thus, the percentage of relevant documents is small, so the time to access these documents must be reduced. We propose a collaborative proxy architecture, close to the user, based on "hot" relevance topics, and the dynamic construction of virtual communities. In this framework, the temperature of a document or a subject reflects its current interest among a community. We propose an architecture that allows proxies not only to manage efficiently cached documents but also to exchange documents with other proxies. Furthermore, these proxies manage meta-data about their partial views of the information space. We show how such an approach fits to enhance global efficiency of the Web both in terms of access time to relevant documents and in terms of Web content indexing, adding semantic value where raw byte arrays are often considered. Lionel Brunie, Jean-Marc Pierson, David Coquil |
WISE | 2 |
| 2000 | Quality and Complexity Bounds of Load Balancing Algorithms for Parallel Image ProcessingabstractThe parallel implementation of image processing algorithms implies an important choice of data distribution strategy. In order to handle the specific constraints associated with images, data distribution must take into account not only the locality of the data and its geometrical regularity but also the possible irregular computation costs associated with different image elements. A widely studied field to tackle this problem is the family of methods related to rectilinear partitioning. We introduce two fully parallel heuristics that compute suboptimal partitions, with a better complexity than the best known algorithms that compute optimal partitions. In this paper, we compare our heuristics to an optimal partitioning, both in terms of execution time and accuracy of the partition. We give some theoretical bounds on the quality of these heuristics that are corroborated by results of random numerical experiments and real applications. Serge Miguet, Jean-Marc Pierson |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1999 | Using Network of Workstations to Support a Web-Based Visualization Service
Wilfrid Lefer, Jean-Marc Pierson |
Euro-Par | 2 |