Adrien Lèbre

dblp:25/6482 · also Adrien Lebre · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-0305-4130ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 8 first-author · 4 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Boosting Task-Driven Applications from Cloud to Edge: Leveraging Utility for Effective Data Replication
abstract
The ongoing shift from Cloud Computing to Edge Computing paradigm brings geo-distributed data management challenges back to the forefront, particularly in the context of large-scale data processing.This article explores data replication, a key strategy for improving system performance in geo-distributed computing environments. While replication can improve efficiency, naive strategies often result in excessive and unnecessary data transfers, leading to inefficient resource utilization, particularly in infrastructures interconnected via heterogeneous network links. To address these limitations, we propose a replication strategy that estimates the utility of each potential replica before deployment. The approach is first evaluated in homogeneous environments and then extended to heterogeneous settings with varying network characteristics. Simulation results show that our method significantly reduces data transfers while maintaining high execution efficiency, achieving a balanced trade-off between performance and resource consumption.
Cherif Si Mohammed, Adrien Lèbre, Alexandre van Kempen
IC2E2
2024 Thinking out of replication for geo-distributing applications: the sharding case
abstract
To promote the adoption of the edge paradigm, our community needs innovative approaches for geo-distributing cloud applications across multiple locations without modifying existing business logic. While recent efforts propose using external services to orchestrate REST operations and achieve geo-distribution, relying solely on resource sharing and replication has limitations in finely distributing manipulated resources. This paper introduces a novel collaboration method that extends resources across multiple instances, going beyond simple replication. Our approach employs a shard-like strategy, enabling the creation of a distributed resource with a unified state view while mitigating a synchronization overhead. The effectiveness of our mechanism is demonstrated through a proof-of-concept implemented on top of the Kubernetes ecosystem.
Geo Johns Antony, Marie Delavergne, Adrien Lèbre, Matthieu Rakotojaona Rainimangavelo
ICFEC3
2022 AS-cast: Lock Down the Traffic of Decentralized Content Indexing at the Edge
Adrien Lèbre, Brice Nédelec, Alexandre van Kempen
ICA3PP1
2022 Cheops, a Service to Blow Away Cloud Applications to the Edge
Marie Delavergne, Geo Johns Antony, Adrien Lèbre
ICSOC3
2022 EnosLib: A Library for Experiment-Driven Research in Distributed Computing
abstract
Despite the importance of experiment-driven research in the distributed computing community, there has been little progress in helping researchers conduct their experiments. In most cases, they have to achieve tedious and time-consuming development and instrumentation activities to deal with the specifics of testbeds and the system under study. In order to relieve researchers of the burden of those efforts, we have developedEnosLib: a Python library that takes into account best experimentation practices and leverages modern toolkits on automatic deployment and configuration systems.EnosLibhelps researchers not only in the process of developing their experimental artifacts, but also in running them over different infrastructures. To demonstrate the relevance of our library, we discuss three experimental engines built on top ofEnosLib, and used to conduct empirical studies on complex software stacks between 2016 and 2019 (database systems, communication buses and OpenStack). By introducingEnosLib, our goal is to gather academic and industrial actors of our community around a library that aggregates everyday experiment-driven research operations. A library that has been already adopted by open-source projects and members of the scientific community thanks to its ease of use and extension.
Ronan-Alexandre Cherrueau, Marie Delavergne, Alexandre van Kempen, Adrien Lèbre, Dimitri Pertin, Javier Rojas Balderrama, Anthony Simonet, Matthieu Simonin
IEEE Trans. Parallel Distributed Syst.4
2022 Estimating Energy Consumption of Cloud, Fog, and Edge Computing Infrastructures
abstract
In order to improve locality aspects, new Cloud-related architectures such as Edge Computing have been proposed. Despite the growing popularity of these new architectures, their energy consumption has not been well investigated yet. To move forward on such a critical question, we first introduce a taxonomy of different Cloud-related architectures. From this taxonomy, we then present an energy model to evaluate their consumption. Unlike previous proposals, our model comprises the full energy consumption of the computing facilities, including cooling systems, and the energy consumption of network devices linking end users to Cloud resources. Finally, we instantiate our model on different Cloud-related architectures, ranging from fully centralized to completely distributed ones, and compare their energy consumption. The results show that a completely distributed architecture, because of not using intra-data center network and large-size cooling systems, consumes between 14 and 25 percent less energy than fully centralized and partly distributed architectures, respectively. To the best of our knowledge, our work is the first one to propose a model that enables researchers to analyze and compare energy consumption of different Cloud-related architectures.
Ehsan Ahvar, Anne-Cécile Orgerie, Adrien Lèbre
IEEE Trans. Sustain. Comput.3
2021 Geo-distribute Cloud Applications at the Edge
Ronan-Alexandre Cherrueau, Marie Delavergne, Adrien Lèbre
Euro-Par3
2021 Kubernetes WANWide: a Deployment Scenario to Expose and Use Edge Computing Resources?
abstract
Cloud Computing have highlighted the importance of container orchestration to manage distributed applicationy's life-cycle. With the advent of Edge Computing, DevOps expect to find the features of containers in the cloud, also at the edge. However, orchestration systems have not been designed to deal with geo-distribution aspects such as latency, intermittent networks, etc. In other words, it is unclear whether they could be directly used on top of massively distributed edge infrastructures without revision. In this paper, we provide an evaluation of Kubernetes in a WANWide context. Precisely, we present and discuss results we obtained during an experimental campaign to analyze the impact of WAN links on its behaviour. While there exist initiatives investigating Kubernetes revisions to deal with distribution aspects, there is, to the best of our knowledge, no rigorous performance evaluations to disqualify the vanilla code.
Karim Manaouil, Adrien Lèbre
PDP2
2020 Multi-site Connectivity for Edge Infrastructures : DIMINET: DIstributed Module for Inter-site NETworking
abstract
The deployment of a geo-distributed cloud infrastructure, leveraging for instance Point-of-Presences at the edge of the network, could better fit the requirements of Network Function Virtualization services and Internet of Things applications. The envisioned architecture to operate such a widely distributed infrastructure relies on executing one instance of a Virtual Infrastructure Manager (VIM) per location and implement appropriate code to enable collaborations between them when needed. However, delivering the mechanisms that allow the collaborations is complex and error prone task. This is particularly true for the one in charge of establishing connectivity among VIM instances on-demand. Besides the reconfiguration of the network equipment, the main challenge is to design a mechanism that can offer usual network virtualization operations to the users while dealing with scalability and intermittent network properties of geo-distributed infrastructures.In this paper, we present how such a challenge can be tackled in the context of OpenStack. More precisely, we introduce DIMINET, a DIstributed Module for Inter-site NETworking services capable to interconnect independent networking resources in an automatized and transparent manner. DIMINET relies on a decentralized architecture where each agent communicates with others only if needed. Moreover, there is no global view of all networking resources but each agent is in charge of interconnecting resources that have been created locally. This approach enables us to mitigate management traffic and keep each site operational in case of network partitions. A promising approach to make other cloud-services collaborative on-demand.
David Espinel Sarmiento, Adrien Lèbre, Lucas Nussbaum, Abdelhadi Chari
CCGRID2
2020 Evaluating Computation and Data Placements in Edge Infrastructures through a Common Simulator
abstract
Scheduling computational jobs with data-sets dependencies is an important challenge of edge computing infrastructures. Although several strategies have been proposed, they have been evaluated through ad-hoc simulator extensions that are, when available, usually not maintained. This is a critical problem because it prevents researchers to -easily- perform fair comparisons between different proposals. In this paper, we propose to address this limitation by presenting a simulation engine dedicated to the evaluation and comparison of scheduling and data movement policies for edge computing use-cases. Built upon the Batsim/SimGrid toolkit, our tool includes an injector that allows the simulator to replay a series of events captured in real infrastructures. It also includes a controller that supervises storage entities and data transfers during the simulation, and a plug-in system that allows researchers to add new models to cope with the diversity of edge computing devices. We demonstrate the relevance of such a simulation toolkit by studying two scheduling strategies with four data movement policies on top of a simulated version of the Qarnot Computing platform, a production edge infrastructure based on smart heaters. We chose this use-case as it illustrates the heterogeneity as well as the uncertainties of edge infrastructures. Our ultimate goal is to gather industry and academics around a common simulator so that efforts made by one group can be factorised by others.
Anderson Andrei Da Silva, Clément Mommessin, Pierre Neyron, Denis Trystram, Adwait Bauskar, Adrien Lèbre, Alexandre van Kempen, Yanik Ngoko, Yoann Ricordel
SBAC-PAD6
2019 YOLO: Speeding Up VM and Docker Boot Time by Reducing I/O Operations
Thuy-Linh Nguyen 0001, Ramon Nou, Adrien Lèbre
Euro-Par3
2019 Efficient Resource Allocation for Multi-Tenant Monitoring of Edge Infrastructures
abstract
By relying on small sized and massively distributed infrastructures, the Edge computing paradigm aims at supporting the low latency and high bandwidth requirements of the next generation services that will leverage IoT devices (e.g., video cameras, sensors). To favor the advent of this paradigm, management services, similar to the ones that made the success of Cloud computing platforms, should be proposed. However, they should be designed in order to cope with the limited capabilities of the resources that are located at the edge. In that sense, they should mitigate as much as possible their footprint. Among the different management services that need to be revisited, we investigate in this paper the monitoring one. Monitoring functions tend to become compute-, storage- and network-intensive, in particular because they will be used by a large part of applications that rely on real-time data. To reduce as much as possible the footprint of the whole monitoring service, we propose to mutualize identical processing functions among different tenants while ensuring their quality-of-service (QoS) expectations. We formalize our approach as a constraint satisfaction problem and show through micro-benchmarks its relevance to mitigate compute and network footprints.
Mohamed Abderrahim 0002, Meryem Ouzzif, Karine Guillouard, Jérôme François, Adrien Lèbre, Charles Prud'homme, Xavier Lorca
PDP5
2019 Data Location Management Protocol for Object Stores in a Fog Computing Infrastructure
abstract
Fog computing infrastructures have been proposed as an alternative to cloud computing to provide low latency computing for the Internet of Things (IoT). But no storage solutions have been proposed to work specifically in this environment. Existing solutions, relying on a distributed Hash Table to locate the data, are not efficient because location record may be placed far away from the object replicas. In this paper, we propose to use a tree-based approach to locate the data, inspired by the domain name system (DNS) protocol. In our protocol, servers look for the location of an object by requesting successively their ancestors in a tree built with a modified version of the Dijkstra's algorithm applied to the physical topology. Location records are replicated close to the object replicas to limit the network traffic when requesting an object. We evaluate our approach on the Grid'5000 testbed using micro experiments with simple network topologies and a macro experiment using the topology of the French National Research and Education Network (RENATER). In this macro benchmark, we show that the time to locate an object in our approach is less than 15 ms on average which is around 20% shorter than using a traditional distributed Hash Table (DHT).
Bastien Confais, Benoît Parrein, Adrien Lèbre
IEEE Trans. Netw. Serv. Manag.3
2019 Putting the Next 500 VM Placement Algorithms to the Acid Test: The Infrastructure Provider Viewpoint
abstract
Most current infrastructures for cloud computing leverage static and greedy policies for the placement of virtual machines. Such policies impede the optimal allocation of resources from the infrastructure provider viewpoint. Over the last decade, more dynamic and often more efficient policies based, e.g., on consolidation and load balancing techniques, have been developed. Due to the underlying complexity of cloud infrastructures, these policies are evaluated either using limited scale testbeds/in-vivo experiments or ad-hoc simulators. These validation methodologies are unsatisfactory for two important reasons: they (i) do not model precisely enough real production platforms (size, workload variations, failure, etc.) and (ii) do not enable the fair comparison of different approaches. More generally, new placement algorithms are thus continuously being proposed without actually identifying their benefits with respect to the state of the art. In this article, we show how VMPlaceS, a dedicated simulation framework enables researchers (i) to study and compare VM placement algorithms from the infrastructure perspective, (ii) to detect possible limitations at large scale and (iii) to easily investigate different design choices. Built on top of the SimGrid simulation platform, VMPlaceS provides programming support to ease the implementation of placement algorithms and runtime support dedicated to load injection and execution trace analysis. To illustrate the relevance of VMPlaceS, we first discuss a few experiments that enabled us to study in details three well known VM placement strategies. Diving into details, we also identify several modifications that can significantly increase their performance in terms of reactivity. Second, we complete this overall presentation of VMPlaceS by focusing on the energy efficiency of the well-know FFD strategy. We believe that VMPlaceS will allow researchers to validate the benefits of new placement algorithms, thus accelerating placement research and favouring the transfer of results to IaaS production platforms.
Adrien Lèbre, Jonathan Pastor, Anthony Simonet, Mario Südholt
IEEE Trans. Parallel Distributed Syst.1
2018 A Tree-Based Approach to Locate Object Replicas in a Fog Storage Infrastructure
abstract
Fog Computing infrastructures have been proposed as an alternative to Cloud Computing to provide computing with low latency for the Internet of Things (IoT). A few storage systems have been proposed to store data in those infrastructures. Most of them are relying on a Distributed Hash Table (DHT) to store the location of objects which is not efficient because the node storing the location of the data may be placed far away from the object replicas. In this paper, we propose to replace the DHT by a tree-based approach mapping the physical topology. Servers look for the location of an object by requesting successively their ancestors in the tree. New location records are also added close to the object replicas not only to limit the network traffic when requesting an object, but also to improve the access times. We also propose to modify the Dijkstra's algorithm to compute the tree used. Finally, we evaluate our approach using the object store InterPlanetary FileSystem (IPFS) on Grid'5000 using both a micro experiment with a simple network topology and a macro experiment using the topology of the French National Research and Education Network (RENATER). We show the time to locate an object in our approach is less than 15 ms on average which is around 20% better than using a DHT.
Bastien Confais, Benoît Parrein, Adrien Lèbre
GLOBECOM3
2018 SimGrid VM: Virtual Machine Support for a Simulation Framework of Distributed Systems
abstract
As real systems become larger and more complex, the use of simulator frameworks grows in our research community. By leveraging them, users can focus on the major aspects of their algorithm, run in-siclo experiments (i.e., simulations), and thoroughly analyze results, even for a large-scale environment without facing the complexity of conducting in-vivo studies (i.e., on real testbeds). Since nowadays the virtual machine (VM) technology has become a fundamental building block of distributed computing environments, in particular in cloud infrastructures, our community needs a full-fledged simulation framework that enables us to investigate large-scale virtualized environments through accurate simulations. To be adopted, such a framework should provide easy-to-use APIs as well as accurate simulation results. In this paper, we present a highly-scalable and versatile simulation framework supporting VM environments. By leveraging SimGrid, a widely-used open-source simulation toolkit, our simulation framework allows users to launch hundreds of thousands of VMs on their simulation programs and control VMs in the same manner as in the real world (e.g., suspend/resume and migrate). Users can execute computation and communication tasks on physical machines (PMs) and VMs through the same SimGrid API, which will provide a seamless migration path to IaaS simulations for hundreds of SimGrid users. Moreover, SimGrid VM includes a live migration model implementing the precopy migration algorithm. This model correctly calculates the migration time as well as the migration traffic, taking account of resource contention caused by other computations and data exchanges within the whole system. This allows user to obtain accurate results of dynamic virtualized systems. We confirmed accuracy of both the VM and the live migration models by conducting several micro-benchmarks under various conditions. Finally, we conclude the article by presenting a first use-case of one consolidation algorithm dealing with a significant number of VMs/PMs. In addition to confirming the accuracy and scalability of our framework, this first scenario illustrates the main interest of SimGrid VM: investigating through in-siclo experiments pros/cons of new algorithms in order to limit expensive in-vivo experiments only to the most promising ones.
Takahiro Hirofuchi, Adrien Lèbre, Laurent Pouilloux
IEEE Trans. Cloud Comput.2
2017 Toward a Holistic Framework for Conducting Scientific Evaluations of OpenStack
abstract
By massively adopting OpenStack for operating small to large private and public clouds, the industry has made it one of the largest running software project, overgrowing the Linux kernel. However, with success comes increased complexity, facing technical and scientific challenges, developers are in great difficulty when testing the impact of individual changes on the performance of such a large codebase, which will likely slow down the evolution of OpenStack. Thus, we claim it is now time for the scientific community to join the effort and get involved in the development of OpenStack, like it has been once done for Linux. In this spirit, we developed Enos, an integrated framework that relies on container technologies for deploying and evaluating OpenStack on any testbed. Enos allows researchers to easily express different configurations, enabling fine-grained investigations of OpenStack services. Enos collects performance metrics at runtime and stores them for post-mortem analysis and sharing. The relevance of the Enos approach to reproducible research is illustrated by evaluating different OpenStack scenarios on the Grid'5000 testbed.
Ronan-Alexandre Cherrueau, Dimitri Pertin, Anthony Simonet, Adrien Lèbre, Matthieu Simonin
CCGrid4
2017 Revising OpenStack to Operate Fog/Edge Computing Infrastructures
abstract
Academic and industry experts are now advocating for going from large-centralized Cloud Computing infrastructures to smaller ones massively distributed at the edge of the network. Among the obstacles to the adoption of this model is the development of a convenient and powerful IaaS system capable of managing a significant number of remote data-centers in a unified way. In this paper, we introduce the premises of such a system by revising the OpenStack software, a leading IaaS manager in the industry. The novelty of our solution is to operate such an Internet-scale IaaS platform in a fully decentralized manner, using P2P mechanisms to achieve high flexibility and avoid single points of failure. More precisely, we describe how we revised the OpenStack Nova service by leveraging a distributed key/value store instead of the centralized SQL backend. We present experiments that validate the correct behavior and gives performance trends of our prototype through an emulation of several data-centers using Grid'5000 testbed. In addition to paving the way to the first large-scale and Internet-wide IaaS manager, we expect this work will attract a community of specialists from both distributed system and network areas to address the Fog/Edge Computing challenges within the OpenStack ecosystem.
Adrien Lèbre, Jonathan Pastor, Anthony Simonet, Frédéric Desprez
IC2E1
2017 An Object Store Service for a Fog/Edge Computing Infrastructure Based on IPFS and a Scale-Out NAS
abstract
Fog and Edge Computing infrastructures have been proposed to address the latency issue of the current Cloud Computing platforms. While a couple of works illustrated the advantages of these infrastructures in particular for the Internet of Things (IoT) applications, elementary Cloud services that can take advantage of the geo-distribution of resources have not been proposed yet. In this paper, we propose a first-class object store service for Fog/Edge facilities. Our proposal is built with Scale-out Network Attached Storage systems (NAS) and IPFS, a BitTorrent-based object store spread throughout the Fog/Edge infrastructure. Without impacting the IPFS advantages particularly in terms of data mobility, the use of a Scale-out NAS on each site reduces the inter-site exchanges that are costly but mandatory for the metadata management in the original IPFS implementation. Several experiments conducted on Grid'5000 testbed are analyzed and confirmed, first, the benefit of using an object store service spread at the Edge and second, the importance of mitigating inter-site accesses. The paper concludes by giving a few directions to improve the performance and fault tolerance criteria of our Fog/Edge Object Store Service.
Bastien Confais, Adrien Lèbre, Benoît Parrein
ICFEC2
2017 Virtual Machine Boot Time Model
abstract
Cloud computing by far brings a lot of undeniable advantages. Accordingly, many research works aim to evaluate the characteristics of cloud systems on many aspects such as performance, workload, cost, provisioning policies, and resources management. In order to setup a cloud system for running rigorous experiments, ones have to overcome a huge amount of challenges and obstacles to build, deploy, and manage systems and applications. Cloud simulation tools help researchers to focus only on the parts they are interesting about without facing the aforementioned challenges. However cloud simulators still do not provide accurate models for Virtual Machine (VM) operations. This leads to incorrect results in evaluating real cloud systems. Following previous works on live-migration, we present in this paper an experiment study we conducted in order to propose a first-class VM boot time model. Most cloud simulators often ignore the VM boot time or give a naive model to represent it. After studying the relationship between the VM boot time and different system parameters such as CPU utilization, memory usage, I/O and network bandwidth, we introduce a first boot time model that could be integrated in current cloud simulators. Through experiments, we also confirmed that our model correctly reproduced the boot time of a VM under different resources contention.
Thuy-Linh Nguyen 0001, Adrien Lèbre
PDP2
2016 Performance Analysis of Object Store Systems in a Fog/Edge Computing Infrastructures
abstract
Fog/Edge computing infrastructures have been proposed as an alternative of current Cloud Computing facilities to address the latency issue that prevents the development of several applications. The main idea is to deploy smaller data-centers at the edge of the backbone in order to bring Cloud computing resources closer to the end-usages. While coupleof works illustrated the advantages of such infrastructures in particular for the Internet of Things (IoT) applications, the way of designing elementary services that can take advantage of such massively distributed infrastructures has not been yet discussed. In this paper, we propose to deal with such a question from the storage point of view. First, we propose a list of properties a storage system should meet in this context. Second, we evaluate through performance analysis three "off-the-shelf" object store solutions, namely Rados, Cassandra and InterPlanetary File System (IPFS). In particular, we focused (i) on access times to push and get objects under different scenarios and (ii) on the amount of network traffic that is exchanged between the different sites during such operations. Experiments have been conducted using the Yahoo Cloud System Benchmark (YCSB) on top of the Grid'5000 testbed. We show that among the three tested solutions IPFS fills most of the criteria expected for a Fog/Edge computing infrastructure.
Bastien Confais, Adrien Lèbre, Benoît Parrein
CloudCom2
2015 Adding Storage Simulation Capacities to the SimGrid Toolkit: Concepts, Models, and API
abstract
For each kind of distributed computing infrastructures, i.e., clusters, grids, clouds, data centers, or supercomputers, storage is a essential component to cope with the tremendous increase in scientific data production and the ever-growing need for data analysis and preservation. Understanding the performance of a storage subsystem or dimensioning it properly is an important concern for which simulation can help by allowing for fast, fully repeatable, and configurable experiments for arbitrary hypothetical scenarios. However, most simulation frameworks tailored for the study of distributed systems offer no or little abstractions or models of storage resources. In this paper, we detail the extension of SimGrid, a versatile toolkit for the simulation of large-scale distributed computing systems, with storage simulation capacities. We first define the required abstractions and propose anew API to handle storage components and their contents in SimGrid-based simulators. Then we characterize the performance of the fundamental storage component that are disks and derive models of these resources. Finally we list several concrete use cases of storage simulations in clusters, grids, clouds, and data centers for which the proposed extension would be beneficial.
Adrien Lèbre, Arnaud Legrand, Frédéric Suter, Pierre Veyre
CCGRID1
2015 VMPlaceS: A Generic Tool to Investigate and Compare VM Placement Algorithms
Adrien Lèbre, Jonathan Pastor, Mario Südholt
Euro-Par1
2014 Locality-Aware Cooperation for VM Scheduling in Distributed Clouds
Jonathan Pastor, Marin Bertier, Frédéric Desprez, Adrien Lèbre, Flavien Quesnel, Cédric Tedeschi
Euro-Par4
2013 Adding a Live Migration Model into SimGrid: One More Step Toward the Simulation of Infrastructure-as-a-Service Concerns
abstract
Although virtual machine (VM) placement problem has been an active research area over the past decade, the research community is still looking for an open simulation framework that can simulate in an accurate as well as scalable manner VM operations including live migrations. Existing frameworks, however, leverage a naive migration model that considers neither memory update operations nor resource sharing contention, resulting in an underestimate of both the duration of a live migration and the size of migration traffic. In this paper, we propose a simulation framework of virtualized distributed systems with the first class support of live migration operations. We developed a resource share calculation mechanism for VMs and a live migration model implementing the precopy migration algorithm of Qemu/KVM. We extended a widely used simulation toolkit, SimGrid, which allows users to simulate large-scale distributed systems by using user-friendly programming API. Through experiments, we confirmed that our simulation framework correctly reproduced live migration behaviors of the real world under various conditions. Through a first use case, we also confirmed that it is possible to conduct large-scale simulations of complex virtualized workloads upon hundred thousands of VMs upon thousands of physical machines (PMs).
Takahiro Hirofuchi, Adrien Lèbre, Laurent Pouilloux
CloudCom (1)2
2013 Using the EXECO Toolkit to Perform Automatic and Reproducible Cloud Experiments
abstract
This paper describes EXECO, a library that provides easy and efficient control of local or remote, standalone or parallel, processes execution, as well as tools designed for scripting distributed computing experiments on any computing platform. After discussing the EXECO internals, we illustrate its interest by presenting two experiments dealing with virtualization technologies on the Grid'5000 testbed.
Matthieu Imbert, Laurent Pouilloux, Jonathan Rouzaud-Cornabas, Adrien Lèbre, Takahiro Hirofuchi
CloudCom (2)4
2013 Topic 5: Parallel and Distributed Data Management - (Introduction)
María S. Pérez 0001, André Brinkmann, Stergios V. Anastasiadis, Sandro Fiore, Adrien Lèbre, Kostas Magoutis
Euro-Par5
2013 Cooperative and reactive scheduling in large-scale virtualized platforms with DVMS
abstract
SUMMARY One of the principal goals of cloud computing is the outsourcing of the hosting of data and applications, thus enabling a per‐usage model of computation. Data and applications may be packaged in virtual machines (VM), which are themselves hosted by nodes, that is, physical machines. Several frameworks have been designed to manage VMs on pools of physical machines; most of them, however, do not efficiently address a major objective of cloud providers: maximizing system utilization while ensuring the QoS. Several approaches promote virtualization capabilities to improve this trade‐off. However, the dynamic scheduling of a large number of VMs as part of a large distributed infrastructure is subject to important and hard scalability problems that become even worse when VM image transfers have to be managed. Consequently, most current frameworks schedule VMs statically using a centralized control strategy. In this article, we present distributed VM scheduler, a framework that enables VMs to be scheduled cooperatively and dynamically in large‐scale distributed systems. We describe, in particular, how several VM reconfigurations can be dynamically calculated in parallel and applied simultaneously. Reconfigurations are enabled by partitioning the system (i.e., nodes and VMs) on the fly. Partitions are created with a minimum of resources necessary to find a solution to the reconfiguration problem. Moreover, we propose an algorithm to handle deadlocks that may appear because of the partitioning policy. We have evaluated our prototype through simulations and compared our approach with a centralized one. The results show that our scheduler permits VMs to be reconfigured more efficiently: the time needed to manage thousands of VMs on hundreds of machines is typically reduced to a tenth or less. Copyright © 2012 John Wiley & Sons, Ltd.
Flavien Quesnel, Adrien Lèbre, Mario Südholt
Concurr. Comput. Pract. Exp.2
2012 Architecture for the next generation system management tools
Jérôme Gallard, Adrien Lèbre, Christine Morin, Thomas J. Naughton, Stephen L. Scott, Geoffroy Vallée
Future Gener. Comput. Syst.2
2011 Operating Systems and Virtualization Frameworks: From Local to Distributed Similarities
abstract
Virtualization technologies radically changed the way in which distributed architectures are exploited. With the contribution of VM capabilities and with the emergence of IaaS platforms, more and more frameworks tend to manage VMs across distributed architectures like operating systems handle processes on a single node. Taking into account that most of these frameworks follow a centralized model -- where roughly one node is in charge of the management of VMs -- and considering the growing size of infrastructures in terms of nodes and VMs, new proposals relying on more autonomic and decentralized approaches should be submitted. Designing and implementing such models is a tedious and complex task. However, as well as research studies on OSes and hyper visors are complementary at the node level, we advocate that virtualization frameworks can benefit from lessons learnt from distributed operating system proposals. In this article, we motivate such a position by analyzing similarities between OSes and virtualization frameworks. More precisely, we focus on the management of processes and VMs, first at the node level and then on a cluster scale. From our point of view, such investigations can guide the community to design and implement new proposals in a more autonomic and distributed way.
Flavien Quesnel, Adrien Lèbre
PDP2
2010 Cluster-wide context switch of virtualized jobs
abstract
Clusters are mostly used through Resources Management Systems (RMS) with a static allocation of resources for a bounded amount of time. Those approaches are known to be insufficient for an efficient use of clusters. To provide a finer RMS, job preemption, migration and dynamic allocation of resources are required. However due to the complexity of developing and using such mechanisms, advanced scheduling strategies have rarely been deployed. This trend is currently evolving thanks to the use of migration and preemption capabilities of Virtual Machines (VMs). However, although the manipulation of jobs composed of VM enables to change the state of the jobs according to the scheduling objective, changing the state and the location of numerous VMs at each decision is tedious and degrades the overall performance. In addition to the scheduling policy implementation, developers have to focus on the feasibility of the actions while executing them in the most efficient way.
Fabien Hermenier, Adrien Lèbre, Jean-Marc Menaud
HPDC2
2010 Saline: Improving Best-Effort Job Management in Grids
abstract
Virtualization technologies have recently gained a lot of interest in Grid computing as they allow flexible resource management. However, the most common way to exploit grids relies on dedicated services like resource management systems (RMSs) to get resources at a particular time. To improve resource usage, most of these systems provide a best-effort mode where lowest priority jobs can be executed when resources are idle. This particular mode does not provide any guarantee of service and jobs may be killed at any time by the RMS when the nodes they use are subject to higher priority reservations. This behaviour potentially leads to a huge waste of computation time or at least requires users to deal with checkpoints of their jobs. In this paper we present Saline, a generic and non-intrusive framework to manage best-effort jobs at grid level through virtual machines (VMs) usage. We discuss the main challenges concerning the design of such a grid system, focusing on VM snapshot management and network configuration. Results of experiments show our proposal ensures an efficient execution of best-effort jobs through the whole grid.
Jérôme Gallard, Adrien Lèbre, Christine Morin
PDP2
2009 Handling Persistent States in Process Checkpoint/Restart Mechanisms for HPC Systems
abstract
Computer clusters are today the reference architecture for high-performance computing. The large number of nodes in these systems induces a high failure rate. This makes fault tolerance mechanisms, e.g. process checkpoint/restart, a required technology to effectively exploit clusters. Most of the process checkpoint/restart implementations only handle volatile states and do not take into account persistent states of applications, which can lead to incoherent application restarts. In this paper, we introduce an efficient persistent state checkpoint/restoration approach that can be interconnected with a large number of file systems. To avoid the performance issues of a stable support relying on synchronous replication mechanisms, we present a failure resilience scheme optimized for such persistent state checkpointing techniques in a distributed environment. First evaluations of our implementation in the kDFS distributed file system show the negligible performance impact of our proposal.
Pierre Riteau, Adrien Lèbre, Christine Morin
CCGRID2
2009 Refinement Proposal of the Goldberg's Theory
Jérôme Gallard, Adrien Lèbre, Geoffroy Vallée, Christine Morin, Pascal Gallard, Stephen L. Scott
ICA3PP2
2008 Reducing Kernel Development Complexity in Distributed Environments
Adrien Lèbre, Renaud Lottiaux, Erich Focht, Christine Morin
Euro-Par1
2006 I/O Scheduling Service for Multi-Application Clusters
abstract
Distributed applications, especially the ones being I/O intensive, often access the storage subsystem in a nonsequential way (stride requests). Since such behaviours lower the overall system performance, many applications use parallel I/O libraries such as ROMIO to gather and reorder requests. In the meantime, as cluster usage grows, several applications are often executed concurrently, competing for access to storage subsystems and, thus, potentially canceling optimisations brought by parallel I/O libraries. The aIOLi project aims at optimising the I/O accesses within the cluster and providing a simple POSIX API. This article presents an extension of aIOLi to address the issue of disjoint accesses generated by different concurrent applications in a cluster. In such a context, good trade-off has to be assessed between performance, fairness and response time. To achieve this, an I/O scheduling algorithm together with a "requests aggregator" that considers both application access patterns and global system load, have been designed and merged into aIOLi This improvement led to the implementation of a new generic framework pluggable into any I/O file system layer. A test composed of two concurrent IOR benchmarks has shown improvements on read accesses by a factor ranging from 3.5 to 35 with POSIX calls and from 3.3 to 5 with ROMIO, both reference benchmarks have been executed on a traditional NFS server without any additional optimisations
Adrien Lèbre, Guillaume Huard, Yves Denneulin, Przemyslaw Sowa
CLUSTER1
2006 Adaptive I/O Scheduling for Distributed Multi-applications Environments
abstract
The aIOLi project aims at optimizing the I/O accesses within the cluster by providing a simple POSIX API, thus avoiding the constraints to use a dedicated parallel I/O library. This paper introduces an extension of aIOLi to address the issue of disjoint accesses generated by different concurrent applications in a cluster. In such a context, performance, fairness and response time are the criteria for which good tradeoffs have to be assessed. A test composed of two concurrent IOR benchmarks showed improvements on read accesses by a factor ranging from 3.5 to 35 with POSIX calls and from 3.3 to 5 with ROMIO
Adrien Lèbre, Yves Denneulin, Guillaume Huard, Przemyslaw Sowa
HPDC1
2005 aIOLi: an input/output library for cluster of SMP
abstract
As clusters use grows, lots of scientific applications (biology, climatology, nuclear physics...) have been rewritten to fully exploit this extra CPU power and storage capacity. This kind of software uses and creates huge amounts of data with typical parallel I/O access patterns. Several issues, like "out-of-core limitation" or "efficient parallel input/output access" already known in a local context (on SMP nodes for example), have to be handled in a distributed environment such as a cluster. Several solutions have been proposed by the scientific community to handle these issues, like parallel file systems or parallel I/O libraries, but their specific API limits portability and requires good knowledge of their internal mechanisms. This paper presents an efficient I/O library, aIOLi, for parallel access to remote storages in SMP clusters. Thanks to the SMP kernel features, our framework provides parallel I/O without inter-processes synchronization mechanisms as well as a simple interface based on the classic UNIX system calls (creat/open/read/write/close). Our experiments show that the aIOLi solution allows achieving performance close to the limits of the remote storage system.
Adrien Lèbre, Yves Denneulin
CCGRID1
2004 Performance Evaluation of a Prototype Distributed NFS Server
abstract
A high-performance file system is normally a key point for large cluster installations, where hundreds or even thousands of nodes frequently need to manage large volumes of data. While most solutions usually make use of dedicated hardware and/or specific distribution and replication protocols, the NFSP (NFS Parallel) project aims at improving performance within a standard NFS client/server system. In this paper we investigate the possibilities of a replication model for the NFS server, which is based on Lasy Release Consistency (LRC). A prototype has been built upon the user-level NFSv2 server and a performance evaluation is carried out.
Rafael Bohrer Ávila, Philippe Olivier Alexandre Navaux, Pierre Lombard, Adrien Lèbre, Yves Denneulin
SBAC-PAD4