Guillaume Pierre

dblp:63/6649 · DBLP profile ↗
← Back
47ranked-venue papers
2as first author
10since 2021 · last 2027
0000-0003-1418-8174ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 1 first-author · 2 since 2021Computer networks · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4
YearPublicationVenuePosition
2027 PACE: Dynamic node consolidation for near-energy-proportionality in edge clusters
abstract
Edge computing deployments in remote locations are often powered exclusively by renewable energy sources such as solar panels coupled with batteries. Reducing energy consumption is therefore crucial as excessive power draw may result in depleting the batteries and forcing the system to stop. These settings generate a difficult dilemma in the presence of bursty workloads: on the one hand, edge platforms operate a majority of the time under low-intensity workload, a situation where hardware is the least energy efficient. On the other hand, the system must remain able to deliver high processing capacity during workload peaks. Addressing these two challenges simultaneously mandates the usage of platforms that are as close as possible from energy proportionality. We show in this paper that static provisioning strategies, whether based on powerful monolithic nodes or fixed-sized clusters of Single-Board Computers (SBCs), fail to address these challenges due to high baseline power demand under idle load. Specifically, monolithic nodes have high idle power draw, while static SBC clusters suffer from cumulative idle power across all active nodes, a problem that conventional methods such as Dynamic Voltage and Frequency Scaling (DVFS) cannot solve. We therefore propose Proportional Adaptive Cluster for Edge (PACE), a novel architectural approach using an adaptive SBC cluster governed by a computationally lightweight node consolidation algorithm. PACE’s node consolidation algorithm dynamically powers cluster nodes on or off to precisely match the real workload demand, effectively eliminating the idle power draw of active but unused resources. Through comprehensive performance-to-power characterization and evaluation using real-world traces, we demonstrate that PACE achieves up to a 33% reduction of total energy usage and a 2 × better energy proportionality score compared to static configurations.
Ammar Kazem, Guillaume Pierre, Laurent Longuevergne
Future Gener. Comput. Syst.2
2026 Energy-Efficient Right-Sizing of Kafka-like Message Brokers for IoT Workloads
abstract
IoT data pipelines rely on message brokers, such as Apache Kafka and Redpanda, for continuous telemetry ingestion. When it comes to capacity planning of these systems, the absence of clear sizing guidance often leads to conservative over-provisioning and unnecessary energy use. We present a calibration-based methodology for energy-efficient right-sizing of Kafka-compatible clusters for IoT ingest. Using a small set of initial experiments on 3--4 nodes, we fit a performance model that predicts maximum sustainable throughput and per-node power, enabling operators to choose the smallest cluster that satisfies a target ingest rate with headroom while minimizing energy consumption. We substantiate the approach with an experimental study of Kafka and Redpanda across three hardware generations (HDD, SATA SSD, NVMe), varying partition counts, node counts, and resource limits. We find that storage technology is the primary determinant of throughput, horizontal scaling is near-linear, and vertical CPU scaling yields diminishing returns; the two brokers exhibit distinct energy proportionality properties. On previously unseen hardware, the model predicts throughput and power with median errors below 10% and 7%, respectively. Our results provide a practical, reproducible capacity-planning workflow that maps IoT workload requirements (message size and rate) to concrete, energy-aware deployment decisions.
Govind KP, Guillaume Pierre, Romain Rouvoy
ICPE2
2024 Toward Stream Processing Elasticity in Realistic Geo-Distributed Environments
abstract
Stream data processing is a widely used technology for analysing IoT-generated data shortly after being produced, and delivering timely insights about them. Executing such analysis in geo-distributed platforms enables shorter delays between data production and processing and fewer disturbances due to potential instability of long-distance networks, while retaining the ability to scale the processing capacity up and down according to the demand. However, current stream processing systems were designed for environments made of homogeneous servers connected together using high-speed network links. We experimentally study the performance of Apache Flink coupled with the Gesscale auto-scaler in conditions which resemble those of geo-distributed platforms. We demonstrate that Flink’s backpressure mechanism should not be used as the only trigger for rescaling operations in heterogeneous network conditions. Raw performance, as well as performance predictability, also degrade quickly in the presence of stateful data processing operators and/or high network latency between the processing nodes.
Khaled Arsalane, Guillaume Pierre, Shadi Ibrahim
IC2E2
2024 Aggregate Monitoring for Geo-Distributed Kubernetes Cluster Federations
abstract
Distributed monitoring is an essential functionality to allow large cluster federations to efficiently schedule applications on a set of available geo-distributed resources. However, periodically reporting the precise status of each available server is both unnecessary to allow accurate scheduling and unscalable when the number of servers grows. This paper proposes Acala, an aggregate monitoring framework for geo-distributed Kubernetes cluster federations which aims to provide the management cluster with aggregated information about the entire cluster instead of individual servers. Based on actual deployment under a controlled environment in the geo-distributed Grid’5000 testbed, our evaluations show that Acala reduces the cross-cluster network traffic by up to 97% and the scrape duration by up to 55% in the single member cluster experiment. Our solution also decreases cross-cluster network traffic by 95% and memory resource consumption by 83% in multiple member cluster scenarios. A comparison of scheduling efficiency with and without data aggregation shows that aggregation has minimal effects on the system’s scheduling function. These results indicate that our approach is superior to the existing solution and is suitable to handle large-scale geo-distributed Kubernetes cluster federation environments.
Chih-Kai Huang 0001, Guillaume Pierre
IEEE Trans. Cloud Comput.2
2023 Studying the Energy Consumption of Stream Processing Engines in the Cloud
abstract
Reducing the energy consumption of the global IT industry requires one to understand and optimize the large software infrastructures the modern data economy relies on. Among them are the data stream processing systems that are deployed in cloud data centers by companies, such as Twitter, to process billion of events per day in real time. However, studying the energy consumption of such infrastructures is difficult because they rely on a complex virtualized software ecosystem where attributing energy consumption to individual software components is a challenge, and because the space of possible configurations is large. We present GreenFlow, a principled methodology and tool designed to automate the deployment of energy measurement experiments for data stream processing systems in cloud environments. GreenFlow is designed to deliver reproducible results while remaining flexible enough to support a wide range of experiments. We illustrate its usage and show in particular that consolidating a DSP system in the smallest number of servers that are capable of processing it is an effective way to reduce energy consumption.
Govind KP, Guillaume Pierre, Romain Rouvoy
IC2E2
2023 AdapPF: Self-Adaptive Scrape Interval for Monitoring in Geo-Distributed Cluster Federations
abstract
Monitoring plays a vital role in geo-distributed cluster federation environments to accurately schedule applications across geographically dispersed computing resources. However, using a fixed frequency for collecting monitoring data from clusters may waste network bandwidth and is not necessary for ensuring accurate scheduling. In this paper, we propose Adaptive Prometheus Federation (AdapPF), an extension of the widely-used open-source monitoring tool, Prometheus, and its feature, Prometheus Federation. AdapPF aims to dynamically adjust the collection frequency of monitoring data for each cluster in geo-distributed cluster federations. Based on actual deployment in the geo-distributed Grid'5000 testbed, our evaluations demonstrate that AdapPF can achieve comparable results to Prometheus Federation with 5-seconds scrape interval while reducing cross-cluster network traffic by 36%.
Chih-Kai Huang 0001, Guillaume Pierre
ISCC2
2023 VioLinn: Proximity-aware Edge Placementwith Dynamic and Elastic Resource Provisioning
abstract
Deciding where to handle services and tasks, as well as provisioning an adequate amount of computing resources for this handling, is a main challenge of edge computing systems. Moreover, latency-sensitive services constrain the type and location of edge devices that can provide the needed resources. When available resources are scarce there is a possibility that some resource allocation requests are denied. In this work, we propose the VioLinn system to tackle the joint problems of task placement, service placement, and edge device provisioning. Dealing with latency-sensitive services is achieved through proximity-aware algorithms that ensure the tasks are handled close to the end-user. Moreover, the concept of spare edge device is introduced to handle sudden load variations in time and space without having to continuously overprovision. Several spare device selection algorithms are proposed with different cost/performance tradeoffs. Evaluations are performed both in a Kubernetes-based testbed and using simulations and show the benefit of using spare devices for handling localized load spikes with higher quality of service (QoS) and lower computing resource usage. The study of the different algorithms shows that it is possible to achieve this increase in QoS with different tradeoffs against cost and performance.
Klervie Toczé, Ali J. Fahs, Guillaume Pierre, Simin Nadjm-Tehrani
ACM Trans. Internet Things3
2022 Good Shepherds Care For Their Cattle: Seamless Pod Migration in Geo-Distributed Kubernetes
abstract
Container technology has become a very popular choice for easing and managing the deployment of cloud applications and services. Container orchestration systems such as Kubernetes can automate to a large extent the deployment, scaling, and operations for containers across clusters of nodes, reducing human errors and saving cost and time. Designed with "traditional" cloud environments in mind (i.e., large datacenters with close-by machines connected by high-speed networks), systems like Kubernetes present some limitations in geo-distributed environments where computational workloads are moved to the edges of the network, close to where data is being generated/consumed. In geo-distributed environments, moving around containers, either to follow moving data sources/sinks or due to unpredictable changes in the network substrate, is a rather common operation. We present MyceDrive, a stateful resource migration solution natively integrated with the Kubernetes orchestrator. We show that geo-distributed Kubernetes pod migration is feasible while remaining fully transparent to the migrated application as well as its clients, while reducing downtimes up to 7x compared to state-of-the-art solutions.
Paulo Souza Junior, Daniele Miorandi, Guillaume Pierre
ICFEC3
2021 Model-based Stream Processing Auto-scaling in Geo-Distributed Environments
abstract
Data stream processing is an attractive paradigm for analyzing IoT data at the edge of the Internet before transmitting processed results to a cloud. However, the relative scarcity of fog computing resources combined with the workloads’ non-stationary properties make it impossible to allocate a static set of resources for each application. We propose Gesscale, a resource auto-scaler which guarantees that a stream processing application maintains a sufficient Maximum Sustainable Throughput to process its incoming data with no undue delay, while not using more resources than strictly necessary. Gesscale derives its decisions about when to rescale and which geo-distributed resource(s) to add or remove on a performance model that gives precise predictions about the future maximum sustainable throughput after reconfiguration. We show that this auto-scaler uses 17% less resources, generates 52% fewer reconfigurations, and processes more input data than baseline auto-scalers based on threshold triggers or a simpler performance model.
HamidReza Arkian, Guillaume Pierre, Johan Tordsson, Erik Elmroth
ICCCN2
2021 mck8s: An orchestration platform for geo-distributed multi-cluster environments
abstract
Following the adoption of cloud computing, the proliferation of cloud data centers in multiple regions, and the emergence of computing paradigms such as fog computing, there is a need for integrated and efficient management of geo-distributed clusters. Geo-distributed deployments suffer from resource fragmentation, as the resources in certain locations are over-allocated while others are under-utilized. Orchestration platforms such as Kubernetes and Kubernetes Federation offer the conceptual models and building blocks that can be used to build integrated solutions that address the resource fragmentation challenge. In this work, we propose mck8s – an orchestration platform for multi-cluster applications on multiple geo-distributed Kubernetes clusters. It offers controllers that automatically place, scale, and burst multi-cluster applications across multiple geo-distributed Kubernetes clusters. mck8s allocates the requested resources to all incoming applications while making efficient use of resources. We designed mck8s to be easy to use by development and operation teams by adopting Kubernetes’ design principles and manifest files. We evaluated mck8s in a geo-distributed experimental testbed in Grid’5000. Our results show that mck8s balances the resource allocation across multiple clusters and reduces the fraction of pending pods to 6% as opposed to 65% in the case of Kubernetes Federation for the same workload.
Mulugeta Ayalew Tamiru, Guillaume Pierre, Johan Tordsson, Erik Elmroth
ICCCN2
2020 Stateful Container Migration in Geo-Distributed Environments
abstract
Container migration is an essential functionality in large-scale geo-distributed platforms such as fog computing infrastructures. Contrary to migration within a single data center, long-distance migration requires that the container's disk state should be migrated together with the container itself. However, this state may be arbitrarily large, so its transfer may create long periods of unavailability for the container. We propose to exploit the layered structure provided by the OverlayFS file system to transparently snapshot the volumes' contents and transfer them prior to the actual container migration. We implemented this mechanism within Kubernetes. Our evaluations based on a real fog computing test-bed show that our techniques reduce the container's downtime during migration by a factor 4 compared to a baseline with no volume checkpoint.
Paulo Souza Junior, Daniele Miorandi, Guillaume Pierre
CloudCom3
2020 An Experimental Evaluation of the Kubernetes Cluster Autoscaler in the Cloud
abstract
International audience
Mulugeta Ayalew Tamiru, Johan Tordsson, Erik Elmroth, Guillaume Pierre
CloudCom4
2020 Tail-Latency-Aware Fog Application Replica Placement
Ali J. Fahs, Guillaume Pierre
ICSOC2
2020 Voilà: Tail-Latency-Aware Fog Application Replicas Autoscaler
abstract
Latency-sensitive fog computing applications may use replication both to scale their capacity and to place application instances as close as possible to their end users. In such geo-distributed environments, a good replica placement should maintain the tail network latency between end-user devices and their closest replica within acceptable bounds while avoiding overloaded replicas. When facing non-stationary workloads it is essential to dynamically adjust the number and locations of a fog application's replicas. We propose Voilà, a tail-Iatency-aware auto-scaler integrated in the Kubernetes orchestration system. Voila maintains a fine-grained view of the volumes of traffic generated from different user locations, and uses simple yet highly-effective procedures to maintain suitable application resources in terms of size and location.
Ali J. Fahs, Guillaume Pierre, Erik Elmroth
MASCOTS2
2020 Instability in Geo-Distributed Kubernetes Federation: Causes and Mitigation
abstract
As resources in geo-distributed environments are typically located in remote sites characterized by high latency and intermittent network connectivity, delays and transient network failures are common between the management layer and the remote resources. In this paper, we show that delays and transient network failures coupled with static configuration, including the default configuration parameter values, can lead to instability of application deployments in Kubernetes Federation, making applications unavailable for long periods of time. Leveraging on the benefits of configuration tuning, we propose a feedback controller to dynamically adjust the concerned configuration parameter to improve the stability of application deployments without slowing down the detection of hard failures. We show the effectiveness of our approach in a geo-distributed setup across five sites of Grid'5000, bringing system stability from 83-92% with no controller to 99.5-100% using the controller.
Mulugeta Ayalew Tamiru, Guillaume Pierre, Johan Tordsson, Erik Elmroth
MASCOTS2
2020 A systematic approach toward security in Fog computing: Assets, vulnerabilities, possible countermeasures
abstract
Summary Fog computing is an emerging paradigm in the Internet of Things (IoT) space, consisting of a middle computation layer, sitting between IoT devices and Cloud servers. Fog computing provides additional computing, storage, and networking resources in close proximity to where data is being generated and/or consumed. As the Fog layer has direct access to data streams generated by IoT devices and responses/commands sent from the Cloud, it is in a critical position in terms of security of the entire IoT system. Currently, there is no specific tool or methodology for analysing the security of Fog computing systems in a comprehensive way. Generic security evaluation procedures applicable to most information technology products are time consuming, costly, and badly suited to the Fog context. In this article, we introduce a methodology for evaluating the security of Fog computing systems in a systematic way. We also apply our methodology to a generic Fog computing system, showcasing how it can be purposefully used by security analysts and system designers.
Mozhdeh Farhadi, Jean-Louis Lanet, Guillaume Pierre, Daniele Miorandi
Softw. Pract. Exp.3
2019 Proximity-Aware Traffic Routing in Distributed Fog Computing Platforms
abstract
Container orchestration engines such as Kubernetes do not take into account the geographical location of application replicas when deciding which replica should handle which request. This makes them ill-suited to act as a general-purpose fog computing platforms where the proximity between end users and the replica serving them is essential. We present proxy-mity, a proximity-aware traffic routing system for distributed fog computing platforms. It seamlessly integrates in Kubernetes, and provides very simple control mechanisms to allow system administrators to address the necessary trade-off between reducing the user-to-replica latencies and balancing the load equally across replicas. proxy-mity is very lightweight and it can reduce average user-to-replica latencies by as much as 90% while allowing the system administrators to control the level of load imbalance in their system.
Ali J. Fahs, Guillaume Pierre
CCGRID2
2019 Docker Image Sharing in Distributed Fog Infrastructures
abstract
Fog computing platforms offer virtualized resources located in the vicinity of their end users. Their broad geographical distribution force them to split physical resources in large numbers of relatively weak machines. The limited available disk space per fog node however creates problems for Docker-based systems which locally cache a copy of every container image they execute: first, caches may fill very quickly, whereas standard Docker never automatically evicts any image from its cache. This motivates us to implement automatic cache replacement in Docker. Second, splitting cache space in multiple disjoint partitions negatively impacts the hit rates. This motivates us to propose allowing multiple co-located Docker servers to share their caches. Our trace-based evaluations show that the proposed design achieves significant cache hit improvements, leading to reductions of average container deployment times between 37% and 78% depending on the scenarios.
Arif Ahmed 0001, Guillaume Pierre
CloudCom2
2015 Heterogeneous Resource Selection for Arbitrary HPC Applications in the Cloud
Anca Iordache, Eliya Buyukkaya, Guillaume Pierre
DAIS3
2015 Kangaroo: A Tenant-Centric Software-Defined Cloud Infrastructure
abstract
Applications on cloud infrastructures acquire virtual machines (VMs) from providers when necessary. The current interface for acquiring VMs from most providers, however, is too limiting for the tenants, in terms of granularity in which VMs can be acquired (e.g., small, medium, large, etc.), while giving very limited control over their placement. The former leads to VM underutilization, and the latter has performance implications, both translating into higher costs for the tenants. In this work, we leverage nested virtualization and a networking overlay to tackle these problems. We present Kangaroo, an Open Stack-based virtual infrastructure provider, and IPOPsm, a virtual networking switch for communication between nested VMs over different infrastructure VMs. In addition, we design and implement Skippy, the realization of our proposed virtual infrastructure API for programming Kangaroo. Our benchmarks show that through careful mapping of nested VMs to infrastructure VMs, Kangaroo achieves up to an order of magnitude better performance, with only half the cost on Amazon EC2. Further, Kangaroo's unified Open Stack API allows us to migrate an entire application between Amazon EC2 and our local Open Nebula deployment within a few minutes, without any downtime or modification to the application code.
Kaveh Razavi, Ana Ion, Genc Tato, Kyuho Jeong, Renato J. O. Figueiredo, Guillaume Pierre, Thilo Kielmann
IC2E6
2014 Robust Performance Control for Web Applications in the Cloud
abstract
Abstract: With the success of Cloud computing, more and more websites have been moved to cloud platforms. The elas-ticity and high availability of cloud solutions are attractive features for hosting web applications. In particular, the elasticity is supported through trigger-based provisioning systems that dynamically add/release resources when certain conditions are met. However, when dealing with websites, this operation becomes more prob-lematic, as the workload demand fluctuates following an irregular pattern. An excessive reactiveness turns these systems into imprecise and wasteful in terms of SLA fulfillment and resource consumption. In this pa-per, we propose three different provisioning techniques that expose the limitations of traditional systems, and overcomes their drawbacks without overly increasing complexity. Our experiments conducted on both public and private infrastructures show significant reductions in SLA violations while offering performance stability. 1
Hector Fernandez, Corina Stratan, Guillaume Pierre
CLOSER3
2014 Autoscaling Web Applications in Heterogeneous Cloud Infrastructures
abstract
Improving resource provisioning of heterogeneous cloud infrastructures is an important research challenge. The wide diversity of cloud-based applications and customers with different QoS requirements have recently exhibited the weaknesses of current provisioning systems. Today's cloud infrastructures provide provisioning systems that dynamically adapt the computational power of applications by adding or releasing resources. Unfortunately, these scaling systems are fairly limited:(i) They restrict themselves to a single type of resource, (ii)they are unable to fulfill QoS requirements in face of spiky workload, and (iii) they offer the same QoS level to all their customers, independent of customer preferences such as different levels of service availability and performance. In this paper, we present an autoscaling system that overcomes these limitations by exploiting heterogeneous types of resources, and by defining multiple levels of QoS requirements. The proposed system selects a resource scaling plan according to both workload and customer requirements. Our experiments conducted on both public and private infrastructures show significant reductions in QoS-level violations when faced with highly variable workloads.
Hector Fernandez, Guillaume Pierre, Thilo Kielmann
IC2E2
2013 Failure Analysis and Modeling in Large Multi-site Infrastructures
Tran Ngoc Minh, Guillaume Pierre
DAIS2
2012 Scalable Join Queries in Cloud Data Stores
abstract
Cloud data stores provide scalability and high availability properties for Web applications, but do not support complex queries such as joins. Web application developers must therefore design their programs according to the peculiarities of No SQL data stores rather than established software engineering practice. This results in complex and error-prone code, especially with respect to subtle issues such as data consistency under concurrent read/write queries. We present join query support in Cloud TPS, a middleware layer which stands between a Web application and its data store. The system enforces strong data consistency and scales linearly under a demanding workload composed of join queries and read-write transactions. In large-scale deployments, Cloud TPS outperforms replicated Postgre SQL up to three times.
Guillaume Pierre, Chihung Chi
CCGRID2
2012 The XtreemOS Resource Selection Service
abstract
Many large-scale utility computing infrastructures comprise heterogeneous hardware and software resources. This raises the need for scalable resource selection services that identify resources that match application requirements. Such a service must provide an efficient lookup in spite of changing resource attributes such as disk size, changing application requirements such as installed software libraries, and changing system composition as resources join or leave. We present a fully decentralized, self-managing Resource Selection Service (RSS) algorithm by which resources autonomously select themselves when their attributes match a query. An application specifies what it expects from a resource by means of a conjunction of (attribute,value-range) pairs, which are matched against the attribute values of resources. The set of search attributes can also be updated online to reflect new requirements. We show that our solution scales in the number of resources and in the number of attributes, while being relatively insensitive to churn and other membership changes like node failures. Our RSS continuously self-adapts its routing structure in response to variations in the distribution of node attributes and queries. We show that this autonomous optimization maintains performance and availability in a long-lived service even when the set of application requirements used to select resources changes.
Corina Stratan, Jan Sacha, Jeff Napper, Paolo Costa, Guillaume Pierre
ACM Trans. Auton. Adapt. Syst.5
2012 CloudTPS: Scalable Transactions for Web Applications in the Cloud
abstract
NoSQL cloud data stores provide scalability and high availability properties for web applications, but at the same time they sacrifice data consistency. However, many applications cannot afford any data inconsistency. CloudTPS is a scalable transaction manager which guarantees full ACID properties for multi-item transactions issued by web applications, even in the presence of server failures and network partitions. We implement this approach on top of the two main families of scalable data layers: Bigtable and SimpleDB. Performance evaluation on top of HBase (an open-source version of Bigtable) in our local cluster and Amazon SimpleDB in the Amazon cloud shows that our system scales linearly at least up to 40 nodes in our local cluster and 80 nodes in the Amazon cloud.
Guillaume Pierre, Chihung Chi
IEEE Trans. Serv. Comput.2
2011 Introduction
Dariusz R. Kowalski, Pierre Sens 0001, Antonio Fernández 0001, Guillaume Pierre
Euro-Par (1)4
2010 Secure Data Aggregation through Proactive Defense
abstract
Gossip based aggregation protocols are a promising approach to monitoring large-scale decentralized IT infrastructures. Compared to traditional approaches they exhibit good properties of scalability, tolerance of churn, and communication overhead. Gossip-based protocols can compute statistical aggregates such as the average, sum or statistical distribution of an attribute across a large system. However, such protocols are extremely vulnerable to malicious attacks, and even a small number of attackers in the system can largely undermine aggregation results. This paper presents a secure protocol for computing attribute averages. In this system, each node autonomously judges whether its neighbors are malicious, and may subsequently stop any interaction with them. A node appearing malicious to its neighbors quickly gets excluded from the system. Instead of defining malicious behavior (and excluding nodes that follow the definition of maliciousness), our system defines correct behavior (and excludes any node that behaves differently). This allows in principle our system to address arbitrary types of attacks. Simulations based on real-world attribute data demonstrate that our system offers good resistance against four different types of attacks.
Guillaume Pierre, Chihung Chi
ICCCN2
2010 Decentralized As-Soon-As-Possible Grid Scheduling: A Feasibility Study
abstract
NA
Xenofon Vasilakos, Jan Sacha, Guillaume Pierre
ICCCN3
2010 Adam2: Reliable Distribution Estimation in Decentralised Environments
abstract
To enable decentralised actions in very large distributed systems, it is often important to provide the nodes with global knowledge about the values of attributes across all nodes. This paper shows how, given an attribute whose values are distributed across a large decentralised system, each node can efficiently estimate the statistical distribution of these values. Simulations using heavily skewed real-world node attribute distributions show that our estimation methods outperform the state-of-the-art heuristics by an order of magnitude with an average error of 0.05% and a maximum error of 2%. To obtain this accuracy, each node sends on average just 120 kB of data independent of the system size. Our algorithms also achieve this accuracy in the presence of heavy churn of system membership. Furthermore, our algorithm enables self-tuning by continuously estimating the accuracy of its own distribution approximation.
Jan Sacha, Jeff Napper, Corina Stratan, Guillaume Pierre
ICDCS4
2010 Autonomous resource provisioning for multi-service web applications
abstract
Dynamic resource provisioning aims at maintaining the end-to-end response time of a web application within a pre-defined SLA. Although the topic has been well studied for monolithic applications, provisioning resources for applications composed of multiple services remains a challenge. When the SLA is violated, one must decide which service(s) should be reprovisioned for optimal effect. We propose to assign an SLA only to the front-end service. Other services are not given any particular response time objectives. Services are autonomously responsible for their own provisioning operations and collaboratively negotiate performance objectives with each other to decide the provisioning service(s). We demonstrate through extensive experiments that our system can add/remove/shift both servers and caches within an entire multi-service application under varying workloads to meet the SLA target and improve resource utilization.
Dejun Jiang 0003, Guillaume Pierre, Chihung Chi
WWW2
2010 Corrigendum to "Wikipedia workload analysis for decentralized hosting" [Computer Networks 53 (11) (2009) 1830-1845]
Guido Urdaneta, Guillaume Pierre, Maarten van Steen
Comput. Networks2
2009 Introduction
Dejan Kostic, Guillaume Pierre, Flavio Paiva Junqueira, Peter R. Pietzuch
Euro-Par2
2009 Zero-Day Reconciliation of BitTorrent Users with Their ISPs
Marco Slot, Paolo Costa, Guillaume Pierre, Vivek Rai
Euro-Par3
2009 Scalable Transactions for Web Applications in the Cloud
Guillaume Pierre, Chihung Chi
Euro-Par2
2009 Autonomous Resource Selection for Decentralized Utility Computing
abstract
Many large-scale utility computing infrastructures comprise heterogeneous hardware and software resources. This raises the need for scalable resource selection services, which identify resources that match application requirements, and can potentially be assigned to these applications. We present a fully decentralized resource selection algorithm by which resources autonomously select themselves when their attributes match a query. An application specifies what it expects from a resource by means of a conjunction of (attribute, value-range) pairs, which are matched against the attribute values of resources. We show that our solution scales in the number of resources as well as in the number of attributes, while being relatively insensitive to churn and other membership changes such as node failures.
Paolo Costa, Jeff Napper, Guillaume Pierre, Maarten van Steen
ICDCS3
2009 Wikipedia workload analysis for decentralized hosting
Guido Urdaneta, Guillaume Pierre, Maarten van Steen
Comput. Networks2
2008 Service-oriented data denormalization for scalable web applications
abstract
Many techniques have been proposed to scale web applications. However, the data interdependencies between the database queries and transactions issued by the applications limit their efficiency. We claim that major scalability improvements can be gained by restructuring the web application data into multiple independent data services with exclusive access to their private data store. While this restructuring does not provide performance gains by itself, the implied simplification of each database workload allows a much more efficient use of classical techniques. We illustrate the data denormalization process on three benchmark applications: TPC-W, RUBiS and RUBBoS. We deploy the resulting service-oriented implementation of TPC-W across an 85-node cluster and show that restructuring its data can provide at least an order of magnitude improvement in the maximum sustainable throughput compared to master-slave database replication, while preserving strong consistency and transactional properties.
Dejun Jiang 0003, Guillaume Pierre, Chihung Chi, Maarten van Steen
WWW3
2008 Practical large-scale latency estimation
Michal Szymaniak, David L. Presotto, Guillaume Pierre, Maarten van Steen
Comput. Networks3
2007 A Decentralized Wiki Engine for Collaborative Wikipedia Hosting
Guido Urdaneta, Guillaume Pierre, Maarten van Steen
WEBIST (1)2
2007 Globetp: template-based database replication for scalable web applications
abstract
Generic database replication algorithms do not scale linearly in throughput as all update, deletion and insertion (UDI) queries must be applied to every database replica. The throughput is therefore limited to the point where the number of UDI queries alone is sufficient to overload one server. In such scenarios, partial replication of a database can help, as UDI queries are executed only by a subset of all servers. In this paper we propose GlobeTP, a system that employs partial replication to improve database throughput. GlobeTP exploits the fact that a Web application's query workload is composed of a small set of read and write templates. Using knowledge of these templates and their respective execution costs, GlobeTP provides database table placements that produce significant improvements in database throughput. We demonstrate the efficiency of this technique using two different industry standard benchmarks. In our experiments, GlobeTP increases the throughput by 57% to 150% compared to full replication, while using identical hardware configuration. Furthermore, adding a single query cache improves the throughput by another 30% to 60%.
Tobias Groothuyse, Swaminathan Sivasubramanian, Guillaume Pierre
WWW3
2007 Enabling service adaptability with versatile anycast
abstract
Abstract We present versatile anycast, which allows a service running on a varying collection of nodes scattered over a wide‐area network to present itself to the clients as one running on a single node. Providing a single logical address enables the client‐side software to preserve the traditional service access model based on single access points. At the same time, the dynamic composition of anycast groups implemented by versatile anycast enables the server‐side service infrastructure to evolve and adapt to changing network conditions. We implement versatile anycast using Mobile IPv6, which decouples the logical addresses of mobile nodes from their physical location. We exploit that decoupling to implement logical service addresses that are not bound to any physical nodes, and employ standard MIPv6 mechanisms to dynamically map each such address onto individual service nodes. Our solution enables a service to transparently hand off clients among the service nodes at the network level while preserving optimal routing between the clients and the service nodes. We demonstrate that the overhead of versatile anycasting is very low. In particular, the client‐perceived handoff time is shown to be a linear function of the latencies among the client and the service nodes participating in the handoff. Copyright © 2007 John Wiley & Sons, Ltd.
Michal Szymaniak, Guillaume Pierre, Mariana Simons-Nikolova, Maarten van Steen
Concurr. Comput. Pract. Exp.2
2005 GlobeDB: autonomic data replication for web applications
abstract
We present GlobeDB, a system for hosting Web applications that performs autonomic replication of application data. GlobeDB offers data-intensive Web applications the benefits of low access latencies and reduced update traffic. The major distinction in our system compared to existing edge computing infrastructures is that the process of distribution and replication of application data is handled by the system automatically with very little manual administration. We show that significant performance gains can be obtained this way. Performance evaluations with the TPC-W benchmark over an emulated wide-area network show that GlobeDB reduces latencies by a factor of 4 compared to non-replicated systems and reduces update traffic by a factor of 6 compared to fully replicated systems.
Swaminathan Sivasubramanian, Gustavo Alonso, Guillaume Pierre, Maarten van Steen
WWW3
2004 Scalable Cooperative Latency Estimation
Michal Szymaniak, Guillaume Pierre, Maarten van Steen
ICPADS2
2002 Dynamically Selecting Optimal Distribution Strategies for Web Documents
abstract
To improve the scalability of the Web, it is common practice to apply caching and replication techniques. Numerous strategies for placing and maintaining multiple copies of Web documents at several sites have been proposed. These approaches essentially apply a global strategy by which a single family of protocols is used to choose replication sites and keep copies mutually consistent. We propose a more flexible approach by allowing each distributed document to have its own associated strategy. We propose a method for assigning an optimal strategy to each document separately and prove that it generates a family of optimal results. Using trace-based simulations, we show that optimal assignments clearly outperform any global strategy. We have designed an architecture for supporting documents that can dynamically select their optimal strategy and evaluate its feasibility.
Guillaume Pierre, Maarten van Steen, Andrew S. Tanenbaum
IEEE Trans. Computers1
2001 Differentiated strategies for replicating Web documents
Guillaume Pierre, Ihor Kuz, Maarten van Steen, Andrew S. Tanenbaum
Comput. Commun.1
1999 Replicated Directory Service for Weakly Consistent Distributed Caches
abstract
Relais is a replicated directory service that improves Web caching within an organization. Relais connects a distributed set of caches and mirrors, providing the abstraction of a single consistent, shared cache. Relais is based on a replication protocol that exploits the semantics of user requests to guarantee cache coherence. The protocol also exploits the semantics of origin servers to optimize bandwidth requirements. At the entry point for a client, Relais maintains a directory of the contents of each Relais data provider; this minimizes look-up latency, by making it a local operation. The cost of maintaining the directories consistent is minimized thanks to a weak consistency protocol. Relais has been prototyped on top of Squid; it has been in daily use for over a year.
Mesaac Makpangou, Guillaume Pierre, Christian Khoury, Neilze Dorta
ICDCS2