VLDB 2026 Research / reviewers in the wild / expert
Francisco Vilar Brasileiro
dblp:b/FranciscoVilarBrasileiro · also Francisco V. Brasileiro
· DBLP profile ↗
69ranked-venue papers
7as first author
0since 2021 · last 2020
0000-0001-9631-0190ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 5 first-authorComputer networks · 7Security and privacy · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 3Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Distributed systems · 82% Cloud and datacenter computing · 7% Hardware reliability and fault tolerance · 6% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 77% Usability and user experience research · 23% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 100% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Collaborative and social computing
online communities |
0.2 | 1 | 2013 | Contributor profiles, their dynamics, and their importance in five q&a sites · CSCW 2013 |
Distributed systems
fault tolerance |
0.1 | 3 | 2004 | A Timeout-Based Message Ordering Protocol for a Lightweight Software Implementation of TMR Systems · IEEE Trans. Parallel Distributed Syst. 2004 Solving the Group Priority Inversion Problem in a Timed Asynchronous System · IEEE Trans. Computers 2002 Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 |
Distributed systems › group communication
message ordering protocols |
0.1 | 2 | 2004 | A Timeout-Based Message Ordering Protocol for a Lightweight Software Implementation of TMR Systems · IEEE Trans. Parallel Distributed Syst. 2004 Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 |
Software testing › test execution
distributed test execution |
0.1 | 1 | 2006 | GridUnit: software testing on the grid · ICSE 2006 |
Software testing
test execution |
0.1 | 1 | 2006 | GridUnit: software testing on the grid · ICSE 2006 |
Distributed systems
replication |
0.1 | 2 | 2002 | Solving the Group Priority Inversion Problem in a Timed Asynchronous System · IEEE Trans. Computers 2002 Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 |
Usability and user experience research › evaluation methodology
longitudinal study |
0.0 | 1 | 2013 | Contributor profiles, their dynamics, and their importance in five q&a sites · CSCW 2013 |
Distributed systems
distributed coordination |
0.0 | 1 | 2004 | A Timeout-Based Message Ordering Protocol for a Lightweight Software Implementation of TMR Systems · IEEE Trans. Parallel Distributed Syst. 2004 |
Distributed systems
peer-to-peer systems |
0.0 | 1 | 2004 | Discouraging Free Riding in a Peer-to-Peer CPU-Sharing Grid · HPDC 2004 |
Algorithmic game theory and mechanism design
incentive mechanism |
0.0 | 1 | 2004 | Discouraging Free Riding in a Peer-to-Peer CPU-Sharing Grid · HPDC 2004 |
Distributed systems › replication › state machine replication
active replication |
0.0 | 1 | 2002 | Solving the Group Priority Inversion Problem in a Timed Asynchronous System · IEEE Trans. Computers 2002 |
Distributed systems › group communication
atomic broadcast |
0.0 | 1 | 2002 | Solving the Group Priority Inversion Problem in a Timed Asynchronous System · IEEE Trans. Computers 2002 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2006 | GridUnit: software testing on the grid · ICSE 2006 |
Distributed systems
grid computing |
0.0 | 1 | 2006 | GridUnit: software testing on the grid · ICSE 2006 |
Hardware reliability and fault tolerance
self-checking systems |
0.0 | 1 | 1996 | Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 |
Hardware reliability and fault tolerance › redundancy › modular redundancy
triple modular redundancy |
0.0 | 1 | 2004 | A Timeout-Based Message Ordering Protocol for a Lightweight Software Implementation of TMR Systems · IEEE Trans. Parallel Distributed Syst. 2004 |
Embedded and real-time systems › real-time scheduling
priority inversion |
0.0 | 1 | 2002 | Solving the Group Priority Inversion Problem in a Timed Asynchronous System · IEEE Trans. Computers 2002 |
Embedded and real-time systems
real-time scheduling |
0.0 | 1 | 2002 | Solving the Group Priority Inversion Problem in a Timed Asynchronous System · IEEE Trans. Computers 2002 |
Methods — techniques the papers use, named apart from their topics
longitudinal analysis · 0.2behavioral profiling · 0.2JUnit extension · 0.1worst-case delay analysis · 0.0timeout-based ordering · 0.0timed asynchronous system model · 0.0failure detector · 0.0message ordering · 0.0comparison protocol · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Community-governed Services on the Edge
João Mafra, Francisco Vilar Brasileiro, Raquel Lopes 0001 |
CLOSER | 2 |
| 2020 | Federated and secure cloud services for building medical image classifiers on an intercontinental infrastructure
Ignacio Blanquer, Francisco Vilar Brasileiro, Andrey Brito, Amanda Calatrava, Christof Fetzer, Flavio Figueiredo, Ronny Petterson Guimarães, Leandro Bezerra Marinho, Wagner Meira Jr., Altigran S. da Silva, Angel Alberich-Bayarri, Eduardo Camacho-Ramos, Ana Jimenez-Pastor, Antonio Luiz L. Ribeiro, Bruno Ramos Nascimento |
Future Gener. Comput. Syst. | 2 |
| 2019 | On the Impacts of Transitive Indirect Reciprocity on P2P Cloud FederationsabstractSeveral P2P systems of resource sharing use cooperation incentive mechanisms to identify and punish free riders, i.e., non-reciprocal individuals. A widespread approach is to use the levels of reciprocity, either directly or indirectly, to decide the extent to which an individual should trust other partners. One restriction of direct reciprocity mechanisms is the inability to foster cooperation between individuals with asymmetrical resources or availability incompatibility. In this work, we evaluate the performance of cloud federations ruled by the combination of the well-known direct reciprocity with transitive reciprocity, a strategy that allows direct reciprocity mechanisms to deal with asymmetry between individuals, while still keeping the benefits of direct reciprocity. For this, we implemented a simulator of resource bartering in cloud federations and experimented it with workloads synthesized from traces of real systems. Our best results showed an average increase of 12.83% and 26.38% on the sharing level of the federation, in an optimistic but unrealistic mechanism setup. When configured in a feasible and realistic manner, the transitive reciprocity was able to increase the sharing level up to an average of 6.02% and 7.53%. Eduardo De Lucena Falcão, Antônio A. Neto, Francisco Vilar Brasileiro, Andrey Brito |
CLOSER | 3 |
| 2019 | DistributedFaaS: Execution of Containerized Serverless Applications in Multi-Cloud Infrastructures
Adbys Vasconcelos, Ítalo Batista, Rodolfo Silva, Francisco Vilar Brasileiro |
CLOSER | 5 |
| 2019 | BioClimate: A Science Gateway for Climate Change and Biodiversity research in the EUBrazilCloudConnect project
Sandro Fiore, Donatello Elia, Ignacio Blanquer, Francisco Vilar Brasileiro, Alessandra Nuzzo, Paola Nassisi, Iana A. A. Rufino, Arie C. Seijmonsbergen, Niels S. Anders, Carlos de Oliveira Galvao, John E. de B. L. Cunha, Miguel Caballer, Mariane S. Sousa-Baena, Vanderlei Perez Canhos, Giovanni Aloisio |
Future Gener. Comput. Syst. | 4 |
| 2018 | Supporting Mixed Workloads in OpenStack-Based CloudsabstractCurrently available open-source cloud management middlewares provide a single service class to allocate computing resources on-demand. The allocation of resources is constrained only by the actual capacity of the infrastructure and usage quotas applied on a per-user basis. Typically, quotas are defined for different users allowing oversubscribed resources, potentially impacting the Quality of Service (QoS) delivered. This is particularly undesirable when mixed workloads with different QoS requirements are submitted to the cloud. We propose a new scheduler for the popular OpenStack middleware which supports multiple service classes. We compared the performance of the proposed scheduler with that of OpenStack's standard scheduler that supports a single service class. Different from the latter, our scheduler handles mixed workloads in a way that the resource utilization is increased, without significantly affecting the QoS guarantees offered to the different service classes. Fábio Morais 0001, Giovanni Farias da Silva, Marcus Carvalho, Francisco Vilar Brasileiro, João Mafra, Alessandro Fook, Raquel Lopes 0001, Daniel Turull |
IEEE CLOUD | 4 |
| 2018 | Agreement-based credibility assessment and task replication in human computation systems
Lesandro Ponciano, Francisco Vilar Brasileiro |
Future Gener. Comput. Syst. | 2 |
| 2017 | Multi-dimensional admission control and capacity planning for IaaS clouds with multiple service classesabstractInfrastructure as a Service (IaaS) providers typically offer multiple service classes to deal with the wide variety of users adopting this cloud computing model. In this scenario, IaaS providers need to perform efficient admission control and capacity planning in order to minimize infrastructure costs, while fulfilling the different Service Level Objectives (SLOs) defined for all service classes offered. However, most of the previous work on this field consider a single resource dimension - typically CPU - when making such management decisions. We show that this approach will either increase infrastructure costs due to over-provisioning, or violate SLOs due to lack of capacity for the resource dimensions being ignored. To fill this gap, we propose admission control and capacity planning methods that consider multiple service classes and multiple resource dimensions. Our results show that our admission control method can guarantee a high availability SLO fulfillment in scenarios where both CPU and memory can become the bottleneck resource. Moreover, we show that our capacity planning method can find the minimum capacity required for both CPU and memory to meet SLOs with good accuracy. We also analyze how the load variation on one resource dimension can affect another, highlighting the need to manage resources for multiple dimensions simultaneously. Marcus Carvalho, Francisco Vilar Brasileiro, Raquel Lopes 0001, Giovanni Farias da Silva, Alessandro Fook, João Mafra, Daniel Turull |
CCGrid | 2 |
| 2017 | On the Efficiency Gains of Using Disaggregated Hardware to Build Warehouse-Scale ClustersabstractEfficiently scheduling the workloads that are submitted to warehouse-scale clusters is not a trivial task. In these systems, the scheduler needs to deal with the heterogeneity in both the jobs that compose the workload, as well as in the servers that comprise the clusters. Moreover, placement constraints that either prevent or force jobs to be allocated in particular servers, makes the scheduler's task even harder. A number of strategies have been proposed to increase the efficiency of these schedulers, however, all of them assume the nowadays prevalent server-based architecture to build clusters. In this paper we assess the possible efficiency gains that can be attained considering that the underlying infrastructure is based on a disaggregated hardware (DH) architecture. This novel paradigm allows the dynamic assembling of logical servers from pools of system-widere sources, providing a way to shape the computing infrastructure while the workload is being allocated to it. Our simulation results, fed with publicly available data from relevant production systems, show that, on average, 5% more CPU demand, and 6% more RAM demand can be allocated, when we compare the fraction of a workload that a state-of-the-art scheduler is able to schedule on a server-based infrastructure with the one that it can allocate on an equivalent DH-based infrastructure. Moreover, in addition to allocating a larger workload, it can do so using less resources. For the workload we studied, on average, 8% of RAM capacity may be kept unused in the resource pools, where it can be switched off to save energy. Although at first sight these numbers might seem small, given the size of these systems, even small percentage improvements can lead to a very large economical impact. Giovanni Farias da Silva, Francisco Vilar Brasileiro, Raquel Lopes 0001, Marcus Carvalho, Fábio Morais 0001, Daniel Turull |
CloudCom | 2 |
| 2017 | An Analysis of the Use of Qualifications on the Amazon Mechanical Turk Online Labor Market
Ianna Sodré, Francisco Vilar Brasileiro |
Comput. Support. Cooperative Work. | 2 |
| 2017 | Capacity planning for IaaS cloud providers offering multiple service classes
Marcus Carvalho, Daniel A. Menascé, Francisco Vilar Brasileiro |
Future Gener. Comput. Syst. | 3 |
| 2016 | Fogbow: A Middleware for the Federation of IaaS CloudsabstractThis paper presents a new middleware, called Fogbow, designed to support large federations of Infrastructure-as-a-service (IaaS) cloud providers. Fogbow follows a novel approach that implements federation functionalities outside the cloud orchestrator. This approach provides great flexibility, since it can use plug-ins that allow for the definition of precise interaction points between the federation middleware and the underlying cloud orchestrator. The resulting architecture, which relies on standards for conciliating different orchestrators' peculiarities, is thereby able to provide a common API to decouple federation functionalities from the orchestrator functionalities. In the demonstration we will showcase how Fogbow has been used to implement several cloud federations, with different requirements. Francisco Vilar Brasileiro, Giovanni Farias da Silva, Francisco Araujo, Marcos Nobrega, Igor Silva, Gustavo Rocha |
CCGrid | 1 |
| 2016 | Instance Type Selection in Proactive Horizontal Auto-ScalingabstractHorizontally scalable applications can potentially run very efficiently over IaaS environments. For that, application providers need to appropriately plan the resource capacity that is to be acquired from the cloud providers, such that, at any point in time, they allocate the smallest infrastructure that is needed to provide the required quality of service for their applications. Since the workload of these applications typically vary widely over time, proactive auto-scaling of the infrastructure is a must. In this paper, we study the impact that an efficient instance type selection based on demands of multidimensional resources has on the performance of proactive auto-scaling. This issue has been mostly overlooked in the related literature. Our results show that suitable selection of the most cost-effective instance type yields a potential cost saving of as much as 50% when compared to the case where the auto-scaling mechanism is oblivious to the instance type selection. We also show evidences that a large portion of applications can benefit of this selection. Finally, we propose a simple selection mechanism that can lead to reasonable cost savings at the expenses of a small number of SLO violations. Fábio Morais 0001, Raquel Lopes 0001, Francisco Vilar Brasileiro |
CloudCom | 3 |
| 2016 | Federation of Private IaaS Cloud Providers through the Barter of ResourcesabstractThis paper presents the Fogbow middleware. This middlewareaddresses cloud federation challenges by providing an additional layer for federation atop each local private IaaS that wants to join the federation. It uses designed-to-federate, internet-friendly technologies like XMPP and it is flexible enough to deal with a wide range of cloud technologies, since it provides a plugin framework for allowing simplerinteroperation with a wide range of orchestrators. The pluginframework is also useful to define the behaviour of each member of the federation, providing greater autonomy tothese members. A lightweight business model based onbarter has been implemented and provides a cheap way tofederate small and medium size private IaaS cloud providers. Francisco Vilar Brasileiro, Eduardo De Lucena Falcão |
ICDCS | 1 |
| 2016 | File system trace replay methods through the lens of metrologyabstractThere are various methods to evaluate the performance of file systems through the replay of file system traces. Despite this diversity, little attention was given on comparing the alternatives, thus bringing some skepticism about the results attained using these methods. In this paper, to fill this understanding gap, we analyze two popular trace replay methods through the lens of metrology. This case study indicates that the evaluated methods provide similar, good precision but are biased in some scenarios. Our results identified limitations in the implementation of the replay tool as well as flaws in the established practices to experiment with trace replayers as the root causes of the measurement bias. After improving the implementation of the trace replayer and discarding inappropriate experimental practices, we were able to reduce the bias, leading to lower measurement uncertainty. Finally, our case study also shows that, in some cases, collecting only the file system activity is not enough to accurately replay the traces; in these cases, collecting resource consumption information, such as the amount of allocated memory, can improve the quality of trace replay methods. Thiago Emmanuel Pereira, Francisco Vilar Brasileiro, Lívia M. R. Sampaio |
MSST | 2 |
| 2016 | A study on the errors and uncertainties of file system trace capture methodsabstractDespite the popularity of trace-based file system performance evaluation, there is currently no accepted methodology to capture and accurately replay file system traces. The less we know about the limitations of trace-based methodologies, the less we can rely on the obtained results. In this paper, we present a case study analyzing the two most popular trace capture methods. The results of the case study indicate that the two evaluated methods provide good precision, but significant bias in some cases. In addition to providing guidelines on how to improve the quality of trace capture, our results allow us to draw important observations about current practice. We show that bias can be corrected by the execution of a calibration procedure, a practice that is mostly absent in the methodology used in the area. Our results also revealed that, to achieve correct calibration, it is crucial to collect information about the background activity and about the operating system layer running in the experimental environment. Finally, our results allow to quantify the overall trace capture uncertainty, providing means for researches to decide on the suitability of trace capture tools for their purposes. Thiago Emmanuel Pereira, Francisco Vilar Brasileiro, Lívia M. R. Sampaio |
SYSTOR | 2 |
| 2015 | Prediction-Based Admission Control for IaaS Clouds with Multiple Service ClassesabstractThere is a growing adoption of cloud computing services, attracting users with different requirements and budgets to run their applications in cloud infrastructures. In order to match users' needs, cloud providers can offer multiple service classes with different pricing and Service Level Objective (SLO) guarantees. Admission control mechanisms can help providers to meet target SLOs by limiting the demand at peak periods. This paper proposes a prediction-based admission control model for IaaS clouds with multiple service classes, aiming to maximize request admission rates while fulfilling availability SLOs defined for each class. We evaluate our approach with trace-driven simulations fed with data from production systems. Our results show that admission control can reduce SLO violations significantly, specially in underprovisioned scenarios. Moreover, our predictive heuristics are less sensitive to different capacity planning and SLO decisions, as they fulfill availability SLOs for more than 91% of requests even in the worst case scenario, for which only 56% of SLOs are fulfilled by a simpler greedy heuristic and as little as 0.2% when admission control is not used. Marcus Carvalho, Daniel A. Menascé, Francisco Vilar Brasileiro |
CloudCom | 3 |
| 2015 | Incentivising Resource Sharing in Federated Clouds
Eduardo De Lucena Falcão, Francisco Vilar Brasileiro, Andrey Brito, José Luis Vivas |
DAIS | 2 |
| 2014 | Long-term SLOs for reclaimed cloud computing resourcesabstractThe elasticity promised by cloud computing does not come for free. Providers need to reserve resources to allow users to scale on demand, and cope with workload variations, which results in low utilization. The current response to this low utilization is to re-sell unused resources with no Service Level Objectives (SLOs) for availability. In this paper, we show how to make some of these reclaimable resources more valuable by providing strong, long-term availability SLOs for them. These SLOs are based on forecasts of how many resources will remain unused during multi-month periods, so users can do capacity planning for their long-running services. By using confidence levels for the predictions, we give service providers control over the risk of violating the availability SLOs, and allow them trade increased risk for more resources to make available. We evaluated our approach using 45 months of workload data from 6 production clusters at Google, and show that 6--17% of the resources can be re-offered with a long-term availability of 98.9% or better. A conservative analysis shows that doing so may increase the profitability of selling reclaimed resources by 22--60%. Marcus Carvalho, Walfredo Cirne, Francisco Vilar Brasileiro, John Wilkes |
SoCC | 3 |
| 2013 | Autoflex: Service Agnostic Auto-scaling Framework for IaaS Deployment ModelsabstractElasticity is a key property to reduce the costs associated with running services in cloud systems that employ an infrastructure-as-a-service (IaaS) deployment model. However, to be able to exploit this property, users of IaaS systems need to be able to anticipate the short-term future demand of their own services, so that only the required infrastructure is requested at any instant in time. This guarantees that service level objectives (SLOs) are always honored by the users of IaaS systems, while over provisioning is avoided. The process of automatically change the amount of resources used to run a service in an IaaS system is named auto-scaling, and the state-of-the-practice uses simple reactive approaches. Although these approaches can successfully reduce the costs due to over provisioning, they are frequently insufficient to minimize the costs due to SLO violations. For that, proactive approaches are required. In this paper we propose a framework for the implementation of auto-scaling services that follows both reactive and proactive approaches. The latter is based on the use of a set of predictors of the future demand of services deployed over IaaS resources, and a selection mechanism that chooses, over time, what is the best predictor to be used. We have also proposed a correction method that uses historical data on the predictors' errors to reduce the probability of under provisioning the services, thus further diminishing the number of SLO violations. We have evaluated the performance of the proposed approach using production traces of HP customers. Our results show that costs savings of as much as 37% can be achieved, while the probability of an SLO violation can be kept, on average, as small as 0.008%, and never larger than 0.036%. Fábio Morais 0001, Francisco Vilar Brasileiro, Raquel Lopes 0001, Ricardo Araújo Santos, Wade Satterfield, Leandro Rosa |
CCGRID | 2 |
| 2013 | Contributor profiles, their dynamics, and their importance in five q&a sitesabstractQ&A sites currently enable large numbers of contributors to collectively build valuable knowledge bases. Naturally, these sites are the product of contributors acting in different ways - creating questions, answers or comments and voting in these - contributing in diverse amounts, and creating content of varying quality. This paper advances present knowledge about Q&A sites using a multifaceted view of contributors that accounts for diversity of behavior, motivation and expertise to characterize their profiles in five sites. This characterization resulted in the definition of ten behavioral profiles that group users according to the quality and quantity of their contributions. Using these profiles, we find that the five sites have remarkably similar distributions of contributor profiles. We also conduct a longitudinal study of contributor profiles in one of the sites, identifying common profile transitions, and finding that although users change profiles with some frequency, the site composition is mostly stable over time. Adabriand Furtado, Nazareno Andrade, Nigini Oliveira, Francisco Vilar Brasileiro |
CSCW | 4 |
| 2013 | On the Accuracy of Trace Replay Methods for File System EvaluationabstractCurrent trace replay methods for file system evaluation fail to represent traced workloads accurately. When using a misrepresented workload one may take wrong conclusions about the system evaluation. For example, a system designer can miss performance problems if the replay of a trace produces an under loaded representation of the real workload. Even worse, one can take wrong design decisions, leading to optimization of untypical workloads. In this study, we captured and replayed traces from standard file systems using methods proposed in the literature, to exemplify the inaccuracy of state-of-art trace replay methods. We also exposed a shortcoming of current methodologies, in a replay of a general purpose workload trace, we observed a difference of up to 100% on request response time, caused by the choice of trace replay method. Thiago Emmanuel Pereira, Lívia M. R. Sampaio, Francisco Vilar Brasileiro |
MASCOTS | 3 |
| 2013 | Analyzing the impact of elasticity on the profit of cloud computing providers
Rostand Costa, Francisco Vilar Brasileiro, Guido Lemos de Souza Filho, Dênio Mariz Sousa |
Future Gener. Comput. Syst. | 2 |
| 2013 | Assessing Green Strategies in Peer-to-Peer Opportunistic Grids
Lesandro Ponciano, Francisco Vilar Brasileiro |
J. Grid Comput. | 2 |
| 2012 | Perspectives of UnaCloud: An Opportunistic Cloud Computing Solution for Facilitating ResearchabstractThis paper presents UnaCloud, an opportunistic Infraestructure as a Service implementation oriented to academic and research institutions, where the IaaS model is supported through the opportunistic use of idle computing resources available in the institution campus, providing researchers with significant and low cost computing capabilities. Also presented are the existing perspectives for current and future work on this project, related work about different concepts of opportunistic or volunteer cloud computing, and the significance and potential impact of this project in the regional research community. Juan D. Osorio, Harold E. Castro, Francisco Vilar Brasileiro |
CCGRID | 3 |
| 2012 | Using Broadcast Networks to Create On-demand Extremely Large Scale High-throughput Computing Infrastructures
Rostand Costa, Francisco Vilar Brasileiro, Guido Lemos de Souza Filho, Dênio Mariz Sousa |
J. Grid Comput. | 2 |
| 2012 | Business-driven short-term management of a hybrid IT infrastructure
Paulo Ditarso Maciel Jr., Francisco Vilar Brasileiro, Ricardo Araújo Santos, David Candeia, Raquel Lopes 0001, Marcus Carvalho, Renato Miceli, Nazareno Andrade, Miranda Mowbray |
J. Parallel Distributed Comput. | 2 |
| 2011 | Evaluating the impact of planning long-term contracts on the management of a hybrid IT infrastructureabstractThe cloud computing market has emerged as an alternative for the provisioning of resources on a pay-as-you-go basis. This flexibility potentially allows clients of cloud computing solutions to reduce the total cost of ownership of their Information Technology infrastructures. On the other hand, this market-based model is not the only way to reduce costs. Among other solutions proposed, peer-to-peer (P2P) grid computing has been suggested as a way to enable a simpler economy for the trading of idle resources. In this paper, we consider an IT infrastructure which benefits from both of these strategies. In such a hybrid infrastructure, computing power can be obtained from in-house dedicated resources, from resources acquired from cloud computing providers, and from resources received as donations from a P2P grid. We take a business-driven approach to the problem and try to maximise the profit that can be achieved by running applications in this hybrid infrastructure. The execution of applications yields utility, while costs may be incurred when resources are used to run the applications, or even when they sit idle. We assume that resources made available from cloud computing providers can be either reserved in advance, or bought on-demand. We study the impact that long-term contracts established with the cloud computing providers have on the profit achieved. Anticipating the optimal contracts is not possible due to the many uncertainties in the system, which stem from the prediction error on the workload demand, the lack of guarantees on the quality of service of the P2P grid, and fluctuations in the future prices of on-demand resources. However, we show that the judicious planning of long term contracts can lead to profits close to those given by an optimal contract set. In particular, we model the planning problem as an optimisation problem and show that the planning performed by solving this optimization problem is robust to the inherent uncertainties of the system, producing profits that for some scenarios can be more than double those achieved by following some common rule-of-thumb approaches to choosing reservation contracts. Paulo Ditarso Maciel Jr., Francisco Vilar Brasileiro, Raquel Lopes 0001, Marcus Carvalho, Miranda Mowbray |
Integrated Network Management | 2 |
| 2011 | Using a Simple Prioritisation Mechanism to Effectively Interoperate Service and Opportunistic Grids in the EELA-2 e-Infrastructure
Francisco Vilar Brasileiro, Matheus Gaudencio, Alexandre Duarte, Diego Carvalho 0001, Diego Scardaci, Leandro Neumann Ciuffo, Rafael Mayo 0001, Herbert Hoeger, Michael Stanton, Raul Ramos, Roberto Barbera, Bernard Marechal, Philippe Gavillet |
J. Grid Comput. | 1 |
| 2010 | Predicting the Quality of Service of a Peer-to-Peer Desktop GridabstractPeer-to-peer (P2P) desktop grids have been proposed as an economical way to increase the processing capabilities of information technology (IT) infrastructures. In a P2P grid, a peer donates its idle resources to the other peers in the system, and, in exchange, can use the idle resources of other peers when its processing demand surpasses its local computing capacity. Despite their cost-effectiveness, scheduling of processing demands on IT infrastructures that encompass P2P desktop grids is more difficult. At the root of this difficulty is the fact that the quality of the service provided by P2P desktop grids varies significantly over time. The research we report in this paper tackles the problem of estimating the quality of service of P2P desktop grids. We base our study on the OurGrid system, which implements an autonomous incentive mechanism based on reciprocity, called the Network of Favours (NoF). In this paper we propose a model for predicting the quality of service of a P2P desktop grid that uses the NoF incentive mechanism. The model proposed is able to estimate the amount of resources that is available for a peer in the system at future instants of time. We also evaluate the accuracy of the model by running simulation experiments fed with field data. Our results show that in the worst scenario the proposed model is able to predict how much of a given demand for resources a peer is going to obtain from the grid with a mean prediction error of only 7.2%. Marcus Carvalho, Renato Miceli, Paulo Ditarso Maciel Jr., Francisco Vilar Brasileiro, Raquel Lopes 0001 |
CCGRID | 4 |
| 2010 | Investigating Business-Driven Cloudburst Schedulers for E-Science Bag-of-Tasks ApplicationsabstractThe new ways of doing science, rooted on the unprecedented processing, communication and storage infrastructures that became available to scientists, are collectively called e-Science. Many research labs now need non-trivial computational power to run e-Science applications. Grid and voluntary computing are well-established solutions that cater to this need, but are not accessible for all labs and institutions. Besides, there is an uncertainty about the future amount of resources that will be available in such infrastructures, which prevents the researchers from planning their activities to guarantee that deadlines will be met. With the emergence of the cloud computing paradigm come new opportunities. One possibility is to run e-Science activities at resources acquired on-demand from cloud providers. However, although very low, there is a cost associated with the usage of cloud resources. Besides that, the amount of resources that can be simultaneously acquired is, in practice, limited. Another possibility is the not new idea of composing hybrid infrastructures in which the huge amount of computational resources shared by the grid infrastructures are used whenever possible and extra capacity is acquired from cloud computing providers. We here investigate how to schedule e-Science activities in such hybrid infrastructures so that deadlines are met and costs are reduced. David Candeia, Ricardo Araújo Santos, Raquel Lopes 0001, Francisco Vilar Brasileiro |
CloudCom | 4 |
| 2009 | Using heuristics to improve service portfolio selection in P2P gridsabstractIn this paper we consider a peer-to-peer grid system which provides multiple services to its users. An incentive mechanism promotes collaboration among peers. It has been shown that the use of a reciprocation-based incentive mechanism in such a system prevents free-riding and, at the same time, promotes the clustering of peers that have mutually profitable interactions. On the other hand, an issue that has not been sufficiently studied in this context is that of service portfolio selection. Normally, peers are subject to resource limitations, which force them to provide only a subset of all services that can be possibly provided. Clearly, the subset of selected services impacts the profit that the grid yields to the peers, since each service will have a different cost and will return a different utility. Moreover, the utility generated by a service is strongly influenced by the behavior of the other peers, which in turn may change over time. In this paper we explore the use of heuristics to select the portfolio of services to be offered by peers in such a grid. The main contributions of this work are the use of heuristics to improve the average profit of peers and a study on the impact of some system characteristics on the heuristics behavior. Alvaro Coelho, Francisco Vilar Brasileiro, Paulo Ditarso Maciel Jr. |
Integrated Network Management | 2 |
| 2009 | Analytical Study of Adversarial Strategies in Cluster-based OverlaysabstractAwerbuch and Scheideler have shown that peer-to-peer overlays networks can survive Byzantine attacks only if malicious nodes are not able to predict what will be the topology of the network for a given sequence of join and leave operations. In this paper we investigate adversarial strategies by following specific protocols. Our analysis demonstrates first that an adversary can very quickly subvert DHT-based overlays by simply never triggering leave operations. We then show that when all nodes (honest and malicious ones) are imposed on a limited lifetime, the system eventually reaches a stationary regime where the ratio of polluted clusters is bounded, independently from the initial amount of corruption in the system. Emmanuelle Anceaume, Francisco Vilar Brasileiro, Romaric Ludinard, Bruno Sericola, Frédéric Tronel |
PDCAT | 2 |
| 2009 | Brief Announcement: Induced Churn to Face Adversarial Behavior in Peer-to-Peer Systems
Emmanuelle Anceaume, Francisco Vilar Brasileiro, Romaric Ludinard, Bruno Sericola, Frédéric Tronel |
SSS | 2 |
| 2009 | Resource demand and supply in BitTorrent content-sharing communities
Nazareno Andrade, Elizeu Santos-Neto, Francisco Vilar Brasileiro, Matei Ripeanu |
Comput. Networks | 3 |
| 2009 | NodeWiz: Fault-tolerant grid information service
Sujoy Basu, Lauro Beltrão Costa, Francisco Vilar Brasileiro, Sujata Banerjee, Puneet Sharma 0001, Sung-Ju Lee 0001 |
Peer-to-Peer Netw. Appl. | 3 |
| 2008 | Scheduling CPU-Intensive Grid Applications Using Partial InformationabstractScheduling parallel applications on computational grids is a difficult task. In order to map the parallel application's tasks onto resources in a efficient way, grid schedulers apply scheduling heuristics. The existing scheduling heuristics can be broadly classified in two approaches: i) bin-packingschedulers,andii)replicationschedulers. The first approach requires complete and accurate information about the applications and the grid environment. The second approach does not use any information but, instead, applies the principle of task replication to achieve good performance. Each of these approaches have drawbacks; attaining accurate and complete information about resources and applications is not always possible in a grid environment, while the redundancy of replication schedulers yield an extra consumption of resources. In this work, we investigate the trade-off between these two approaches. We propose scheduling heuristics that use any available information to perform efficient scheduling ofbag-of-tasksapplications, a subclass of parallel applications. Our results show that judicious use of whatever information is available leads to a reduction on resource consumption, without compromising the application's performance. Nelson Nobrega, Leonardo De Assis, Francisco Vilar Brasileiro |
ICPP | 3 |
| 2008 | Improving Automated Testing of Multi-threaded SoftwareabstractThis paper discusses an approach to avoid incorrect results in the execution of automatic tests of multi-threaded systems. We argue that such incorrect results have two main sources. First, it is typically difficult to determine when all threads have finished processing and thus when it is safe to perform the test assertions. Second, background threads can change the system state while assertions are being performed, thus producing non-deterministic results. The main contributions of this work are: (i) a generic approach that ensures that test assertions are performed in a safe moment; (ii) implementation details of such an approach using aspect-oriented programming (AOP); and (Hi) an evaluation of the proposed approach. Ayla Débora Dantas de Souza Rebouças, Francisco Vilar Brasileiro, Walfredo Cirne |
ICST | 2 |
| 2008 | TorrentLab: investigating BitTorrent through simulation and live experimentsabstractBitTorrent is probably the most popular file sharing protocol nowadays. Since it is a complex protocol which was created mostly as an engineering effort, there have been attempts to evaluate and understand the behavior of BitTorrent, exposing drawbacks and identifying opportunities for improvement. Simulation(al) and experimental evaluation are two important methodologies for the investigation of BitTorrent implementations and protocol variations, but using them with BitTorrent and P2P in general represents a challenge. In this paper, we introduce TorrentLab, a testbed in which BitTorrent simulations as well as live experiments can be performed under both controlled and uncontrolled settings. For a given swarm, we compare results obtained by means of simulation and live experiments, and show that both are in line with the expected behavior. Marinho P. Barcellos, Rodrigo B. Mansilha, Francisco Vilar Brasileiro |
ISCC | 3 |
| 2008 | On the planning of a hybrid IT infrastructureabstractWith the emergence of utility computing and the continuous search for reducing the cost of running information technology (IT) infrastructures, we will soon experience an important change on the way these infrastructures are assembled, configured and managed. In this paper we consider the problem of managing a hybrid high-performance computing infrastructure whose processing elements comprise in-house dedicated machines, a utility computing service provider, and idle machines from a best-effort peer-to-peer grid. This infrastructure supports the execution of both best-effort and real-time applications. Realtime applications use primarily computing power from the in- house machines and any processing power that can be attained from the best-effort grid. Extra capacity required to meet deadlines is purchased from the utility computing service provider. This extra capacity is reserved for future use through short term contracts which are negotiated with no human intervention. We take a business-driven approach for the management of this hybrid infrastructure and propose heuristics that can be used by a contract planner agent to reduce the cost of running the applications at the same time that guarantees that deadlines are met. In particular, we show that constructing an estimation for the behavior of the grid is essential for making contracts that lead to high efficiency in the use of the hybrid infrastructure. Paulo Ditarso Maciel Jr., Flavio Figueiredo, D. Maia, Francisco Vilar Brasileiro, Alvaro Coelho |
NOMS | 4 |
| 2008 | Scalable Resource Annotation in Peer-to-Peer GridsabstractPeer-to-peer grids are large-scale, dynamic environments where autonomous sites share computing resources. Producing and maintaining relevant and up-to-date resource information in such environments is a challenging problem, due to the grid scale, the resource heterogeneity, and the variety of user demand. This work proposes a peer-to-peer annotation approach where users can freely annotate available resources as a solution to this problem. We advocate that the proposed approach (i) is scalable, as the job of updating the resource information is divided among users; (ii) will improve resources' utilization, by reducing the amount of resources which are allocated to users without matching their applications constraints; and (iii) will allow resource allocators to increase users' utility, leveraging access to more detailed preference descriptions. The paper also discusses the challenges in implementing and deploying such approach and present solutions to tackle these challenges. Nazareno Andrade, Elizeu Santos-Neto, Francisco Vilar Brasileiro |
Peer-to-Peer Computing | 3 |
| 2007 | On the Efficiency and Cost of Introducing QoS in BitTorrentabstractBitTorrent is currently a de facto standard for scalable content-distribution. However, its peer-to-peer model for resource allocation does not provide high availability and its performance depends on best-effort contributions given by peers. This has motivated several content-providers to use a hybrid model in which they operate a superpeer in order to attain a higher quality of service. In this paper, we use BitTorrent traces and analytical modelling to investigate the cost incurred by such an entity in relation to the benefits it can provide to the system. Nazareno Andrade, Jaindson Santana, Francisco Vilar Brasileiro, Walfredo Cirne |
CCGRID | 3 |
| 2007 | TVGrid: A Grid Architecture to use the idle resources on a Digital TV networkabstractSystems such as SETI@home [8] have proven that it is possible to make good use of massive amounts of communication and computing power that would otherwise be wasted in the leaves of the Internet. In this paper we explore this idea in a different setting. We present the TVGrid architecture, which brings parallel application execution to the digital TV domain, by using the idle communication and processing capacity of a digital TV network. In the proposed architecture, a TV station runs a scheduler that uses unused bandwidth in the digital TV broadcast channel to send tasks to digital TV receivers, which may run them using their unused capacity. Task's outputs are later sent back to the TV station through a return channel (a broadband Internet connection, for example). TV Grid architecture was developed during the studies of the Brazilian Digital TV project. Carlos Eduardo Coelho Freire Batista, Tiago Maritan Ugulino de Araújo, Derzu Omaia, Thiago Curvelo dos Anjos, Giuliano Maia Lins de Castro, Francisco Vilar Brasileiro, Guido Lemos de Souza Filho |
CCGRID | 6 |
| 2007 | Bridging the High Performance Computing Gap: the OurGrid ExperienceabstractHigh performance computing is currently not affordable for those users that cannot rely on having a highly qualified computing support team. To cater for these users' needs we have proposed, implemented and deployed OurGrid. OurGrid is a peer-to-peer grid middleware that supports the automatic creation of large computational grids for the execution of embarrassingly parallel applications. It has been used to support the OurGrid Community - a public free-to-join grid that is in production since December 2004. In this paper we show how the OurGrid Community has been used to support the execution of a number of applications. Further we discuss the main benefits brought up by the system and the difficulties that have been faced by the system developers and the users and managers of the OurGrid Community. Francisco Vilar Brasileiro, Eliane Araújo, William Voorsluys, Milena Oliveira, Flavio Figueiredo |
CCGRID | 1 |
| 2007 | Evaluating the Impact of Simultaneous Round Participation and Decentralized Decision on the Performance of ConsensusabstractConsensus services have been recognized as fundamental building blocks for fault-tolerant distributed systems. Many different protocols to implement such a service have been proposed, however, not a lot of effort has been placed in evaluating their performance. In particular, in the context of round-based consensus protocols for asynchronous systems augmented with failure detectors, there has been some work on evaluating how the QoS of the failure detector impacts the performance of the protocols, as well as on the trade-off between having faster decentralized decision at the expenses of generating more network load. These studies, however, focus on protocols that have no mechanism to deal with an eventual bad QoS provided by the failure detector, and have a decision pattern that is either completely centralized - only one process being able to autonomously decide - or completely decentralized - all processes being able to autonomously decide. This paper reports a thorough evaluation of the performance of a consensus protocol that has two unique features. Firstly, it mitigates the problems due to bad QoS delivered by the failure detector by allowing processes to simultaneously participate in multiple rounds. Secondly, it allows its decision pattern to be configured to have different numbers of processors allowed to autonomously decide. We have measured the decision latency of the protocol to conduct the performance analysis. The results, obtained by means of simulation, highlight the advantages and limitations of the two mechanisms and allow one to understand in a comprehensive framework how the protocol's parameters should be set, such that the best performance is achieved depending on the application's requirements. Lívia M. R. Sampaio, Michel Hurfin, Francisco Vilar Brasileiro, Fabíola Greve |
DSN | 3 |
| 2007 | Relative autonomous accounting for peer-to-peer GridsabstractAbstract Here we present and evaluate relative accounting, an autonomous accounting scheme that provides accurate results even when the parties (consumer and provider) do not trust each other. Relative accounting relies on the observed relative performance amongst the parties. As such, the basic requirement to use it is that resource consumers must also be resource providers. Relative accounting is totally autonomous in the sense that it uses only local information, i.e. there is no exchange of information between the parties. This allows for the deployment of the autonomous accounting without requiring any sort of identification infrastructure, such as certificate authorities. Not requiring trust or sophisticated infrastructure makes relative accounting a perfect fit for peer‐to‐peer Grids, which aim to scale much further than traditional Grids by allowing free unidentified entry into the Grid. Our results show that relative accounting performs very close to a perfect accounting, whose implementation is infeasible in most systems, including those we target. Relative accounting was developed to work with OurGrid, a peer‐to‐peer Grid in production since December 2004, but it can also be used in other peer‐to‐peer Grids. Copyright © 2006 John Wiley & Sons, Ltd. Robson Santos, Alisson Andrade, Walfredo Cirne, Francisco Vilar Brasileiro, Nazareno Andrade |
Concurr. Comput. Pract. Exp. | 4 |
| 2007 | Automatic grid assembly by promoting collaboration in peer-to-peer grids
Nazareno Andrade, Francisco Vilar Brasileiro, Walfredo Cirne, Miranda Mowbray |
J. Parallel Distributed Comput. | 2 |
| 2007 | On the efficacy, efficiency and emergent behavior of task replication in large distributed systems
Walfredo Cirne, Francisco Vilar Brasileiro, Daniel Paranhos da Silva, Fabrício Góes, William Voorsluys |
Parallel Comput. | 2 |
| 2006 | Collaborative Fault Diagnosis in Grids through Automated TestsabstractGrids have the potential to revolutionize computing by providing ubiquitous, on demand access to computational services and resources. However, grid systems are extremely large, complex and prone to failures. A survey we have conducted reveals that fault diagnosis is still a major problem for grid users. When a failure appears at the user screen, it becomes very difficult for the user to identify whether the problem is in his application, somewhere in the grid middleware, or even lower in the fabric that comprises the grid. To overcome this problem, we argue that current grid platforms must be augmented with a collaborative diagnosis mechanism. We propose for such mechanism to use automated tests to identify the root cause of a failure and propose the appropriate fix. We also present a Java-based implementation of the proposed mechanism, which provides a simple and flexible framework that eases the development and maintenance of the automated tests. Alexandre Duarte, Francisco Vilar Brasileiro, Walfredo Cirne, Jose Alencar Filho |
AINA (1) | 2 |
| 2006 | GridUnit: software testing on the gridabstractSoftware testing is a fundamental part of system development. As software grows, its test suite becomes larger and its execution time may become a problem to software developers. This is especially the case for agile methodologies, which preach a short develop/test cycle. Moreover, due to the increasing complexity of systems, there is the need to test software in a variety of environments. In this paper, we introduce GridUnit, an extension of the widely adopted JUnit testing framework, able to automatically distribute the execution of software tests on a computational grid with minimum user intervention. Experiments conducted with this solution have showed a speed-up of almost 70x, reducing the duration of the test phase of a synthetic application from 24 hours to less than 30 minutes. The solution does not require any source-code modification, hides the grid complexity from the user and provides a cost-effectiveness improvement to the software testing experience. Alexandre Duarte, Walfredo Cirne, Francisco Vilar Brasileiro, Patrícia Duarte de Lima Machado |
ICSE | 3 |
| 2006 | UsingWeb Services for Configuration and Deployment according to the CDDLM StandardabstractAs Web services and service oriented architectures are adopted, it is increasingly important to have standard and interoperable means to deploy and configure Web services. Within the Global Grid Forum, HP, NEC, and Softricity have been developing a standard for configuration description, deployment, and lifecycle management (CDDLM). In order to prove its feasibility, reference implementations are being developed. This paper describes an independent reference implementation of CDDLM and the experience in using Web Services for deployment in a standardized manner. Our main contributions are: the lessons learned in implementing this WS-based standard and an architecture for implementing CDDLM Ayla Débora Dantas de Souza Rebouças, Guilherme Germoglio, Flavio Santos, Marcelo Iury S. Oliveira, Walfredo Cirne, Francisco Vilar Brasileiro, Dejan S. Milojicic, Sandro Rafaeli, Katia Barbosa Saikoski |
ICWS | 6 |
| 2006 | A Reciprocation-Based Economy for Multiple Services in Peer-to-Peer GridsabstractIn this paper we study reciprocation-based mechanisms to encourage donation in peer-to-peer grids in which multiple services, such as processing power and data transfers, are shared explicitly. We have modeled such a system and established how peers should assess whether it is profitable to exchange services with another peer, an issue that is not present in the single service case. Unfortunately, this assessment relies on information provided by untrustworthy peers. As an alternative, we have extended, to the case of multiple services, a reciprocation-based mechanism which uses only reliable information gathered locally. We have assessed this mechanism by simulating scenarios in which services are exchanged that are combinations of two different basic services. In the explored scenarios the mechanism performs very well, and can marginalize free riders even when the cost to peers of donating a service is nearly as large as the utility gained by receiving it Miranda Mowbray, Francisco Vilar Brasileiro, Nazareno Andrade, Jaindson Santana, Walfredo Cirne |
Peer-to-Peer Computing | 2 |
| 2006 | Labs of the World, Unite!!!
Walfredo Cirne, Francisco Vilar Brasileiro, Nazareno Andrade, Lauro Beltrão Costa, Alisson Andrade, Reynaldo Novaes, Miranda Mowbray |
J. Grid Comput. | 2 |
| 2005 | Adaptive Indulgent ConsensusabstractDue to their fundamental role in the design of fault-tolerant distributed systems, consensus protocols have been widely studied. In particular, design and performance issues of indulgent consensus are a research topic that has gained considerable attention. Most of these protocols are asymmetric in the sense that different participants can assume different roles during the execution of the protocol. Usually, there is a process that assumes a "special" role and the others cooperate with it to finish the computation. However, the asymmetric structure of indulgent consensus protocols has a performance pitfall, specially when processes and communication channels are subject to considerable variability in load. The problem is that such protocols use an a priori agreed process ordering to select the process to perform the "special" role. We advocate that adaptive indulgent consensus protocols can be constructed by the introduction of an adaptive process ordering module. In this sense, it is proposed a generic implementation for this module. Based on this generic module we provide implementations of both /spl diams/S- and /spl Omega/-based adaptive indulgent consensus protocols. Further, we investigate their performance by means of simulation and real experiments over a widely distributed system. The experimental results obtained show that the adaptive consensus protocols can outperform their non-adaptive counterparts in as much as 50%. Lívia M. R. Sampaio, Francisco Vilar Brasileiro |
DSN | 2 |
| 2004 | When can an autonomous reputation scheme discourage free-riding in a peer-to-peer system?abstractWe investigate the circumstances under which it is possible to discourage free-riding in a peer-to-peer system for resource-sharing by prioritizing resource allocation to peers with higher reputation. We use a model to predict conditions necessary for any reputation scheme to succeed in discouraging free-riding by this method. We show with simulations that for representative cases, a very simple autonomous reputation scheme works nearly as well at discouraging free-riding as an ideal reputation scheme. Finally, we investigate the expected dynamic behavior of the system. Nazareno Andrade, Miranda Mowbray, Walfredo Cirne, Francisco Vilar Brasileiro |
CCGRID | 4 |
| 2004 | Discouraging Free Riding in a Peer-to-Peer CPU-Sharing Grid
Nazareno Andrade, Francisco Vilar Brasileiro, Walfredo Cirne, Miranda Mowbray |
HPDC | 2 |
| 2004 | Exploiting Replication and Data Reuse to Efficiently Schedule Data-Intensive Applications on Grids
Elizeu Santos-Neto, Walfredo Cirne, Francisco Vilar Brasileiro, Aliandro Lima |
JSSPP | 3 |
| 2004 | Scheduling in Bag-of-Task Grids: The PAUÁ CaseabstractIn this paper we discuss the difficulties involved in the scheduling of applications on computational grids. We highlight two main sources of difficulties: 1) the size of the grid rules out the possibility of using a centralized scheduler; 2) since resources are managed by different parties, the scheduler must consider several different policies. Thus, we argue that scheduling applications on a grid require the orchestration of several schedulers, with possibly conflicting goals. We discuss how we have addressed this issue in the context of PAUA, a grid for Bag-of-Tasks applications (i.e. parallel applications whose tasks are independent) that we are currently deploying throughout Brazil. Walfredo Cirne, Francisco Vilar Brasileiro, Lauro Beltrão Costa, Daniel Paranhos da Silva, Elizeu Santos-Neto, Nazareno Andrade, César A. F. De Rose, Tiago Ferreto, Miranda Mowbray, Roque Scheer, João Jornada |
SBAC-PAD | 2 |
| 2004 | A Timeout-Based Message Ordering Protocol for a Lightweight Software Implementation of TMR SystemsabstractReplicated processing with majority voting is a well-known method for achieving reliability and availability. Triple modular redundant (TMR) processing is the most commonly used version of that method. Replicated processing requires that the replicas reach agreement on the order in which input requests are to be processed. Almost all synchronous and deterministic ordering protocols published in the literature are time-based in the sense that they require replicas' clocks to be kept synchronized within some known bound. We present a protocol for TMR systems that is based on timeouts and does not require clocks to be kept in bounded synchronism. Our design efforts focus on keeping the ordering delays small, without an unnecessary increase in message overhead. Consequently, we are able to show that no symmetric protocol that works only with unsynchronized clocks can provide a smaller worst-case delay. We also demonstrate through analysis and experiments that our protocol is faster than a time-based one of identical message complexity in certain situations which can prevail in many application settings. Paul D. Ezhilchelvan, Francisco Vilar Brasileiro, Neil A. Speirs |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2003 | How Bad Are Wrong Suspicions? Towards Adaptive Distributed ProtocolsabstractIn this paper, we analyze the performance of consensus protocols based on the rotating coordinator paradigm. We consider a simulated production environment, on which processing and communication resources available for the different processes running the protocols are not necessarily the same. Firstly, we show that, in some scenarios, the performance of the consensus protocol is enhanced when there is an increase in the number and duration of the wrong suspicions periods of the failure detection service used. Since it is well known that wrong suspicions may also decrease the performance of the consensus protocol, a new dilemma is posed to the designers of such protocols. We then propose a new approach to address performance issues in the design of crash-detection based distributed protocols for asynchronous systems. We argue that they must be designed to adapt themselves to the variations on the availability of resources. The concept of slowness oracles is proposed to achieve this goal. Finally, we present a slowness oracle that can be used to transform a non-adaptive consensus protocol into an adaptive one. Simulations show that the adaptive protocol outperforms its conventional non-adaptive counterpart in a number of scenarios, having an equivalent performance in the other scenarios. Lívia M. R. Sampaio, Francisco Vilar Brasileiro, Walfredo Cirne, Jorge C. A. de Figueiredo |
DSN | 2 |
| 2003 | Trading Cycles for Information: Using Replication to Schedule Bag-of-Tasks Applications on Computational Grids
Daniel Paranhos da Silva, Walfredo Cirne, Francisco Vilar Brasileiro |
Euro-Par | 3 |
| 2003 | Running Bag-of-Tasks Applications on Computational Grids: The MyGrid ApproachabstractWe here discuss how to run Bag-of-Tasks applications on computational grids. Bag-of-Tasks applications (those parallel applications whose tasks are independent) are both relevant and amendable for execution on grids. However, few users currently execute their Bag-of-Tasks applications on grids. We investigate the reason for this state of affairs and introduce MyGrid, a system designed to overcome the identified difficulties. MyGrid provides a simple, complete and secure way for a user to run Bag-of-Tasks applications on all resources she has access to. Besides putting together a complete solution useful for real users, MyGrid embeds two important research contributions to grid computing. First, we introduce some simple working environment abstractions that hide machine configuration heterogeneity from the user. Second, we introduce work queue with replication (WQR), a scheduling heuristics that attains good performance without relying on information about the grid or the application, although consuming a few more cycles. Note that not depending on information makes WQR much easier to deploy in practice. Walfredo Cirne, Daniel Paranhos da Silva, Lauro Beltrão Costa, Elizeu Santos-Neto, Francisco Vilar Brasileiro, Jacques Philippe Sauvé, Fabrício Alves Barbosa da Silva, Carla O. Barros, Cirano Silveira |
ICPP | 5 |
| 2003 | OurGrid: An Approach to Easily Assemble Grids with Equitable Resource Sharing
Nazareno Andrade, Walfredo Cirne, Francisco Vilar Brasileiro, Paulo Roisenberg |
JSSPP | 3 |
| 2002 | Solving the Group Priority Inversion Problem in a Timed Asynchronous SystemabstractConsiders the priority inversion problem in an actively replicated system. Priority inversion was originally defined in the context of nonreplicated systems. Therefore, we first introduce the concept of group priority inversion, which extends the concept of (local) priority inversion to the context of a group of processors that perform an actively replicated processing. We then present the properties of a request scheduling protocol to enforce a total ordering for the processing of requests while avoiding group priority inversions. These properties have been implemented in a protocol that relies on a timed asynchronous system model equipped with a failure detector of the class /spl diams/S. The proposed solution allows us to replicate a critical server while ensuring that the processing of all the incoming requests is consistent (mechanisms for solving the atomic broadcast problem) and predictable (mechanisms for solving the group priority inversion problem). Thus, the described request scheduling protocol is a key component which can be used to develop fault-tolerant real-time applications in a timed asynchronous system. Emmanuelle Anceaume, Francisco Vilar Brasileiro, Fabíola Greve, Michel Hurfin |
IEEE Trans. Computers | 3 |
| 2001 | Avoiding Priority Inversion on the Processing of Requests by Active Replicated ServersabstractWe consider the priority inversion problem in an actively replicated system. Priority inversion was originally defined in the context of non-replicated systems. Therefore we first introduce the concept of group priority inversion, which extends the concept of (local) priority inversion to the context of a group of processors that perform an actively replicated processing. We then present the properties of a request scheduling protocol to enforce a total ordering for the processing of requests while avoiding group priority inversions. These properties have been implemented in a protocol that relies on a timed asynchronous system model equipped with a failure detector of the class /spl square/S. The proposed solution allows one to replicate a critical server while ensuring that the processing of all the incoming requests is consistent (mechanisms for solving the atomic broadcast problem) and predictable (mechanisms for solving the group priority inversion problem). Thus, the described request scheduling protocol is a key component which can be used to develop fault tolerant real time applications in a timed asynchronous system. Francisco Vilar Brasileiro, Emmanuelle Anceaume, Fabíola Greve, Michel Hurfin |
DSN | 2 |
| 2001 | Eva: An Event-Based Framework for Developing Specialized Communication ProtocolsabstractPresents a framework for the development of higher level communication protocols that provides extra functionalities not supplied by standard off-the-shelf lower level communication protocols. The framework is based on the event channel abstraction which allows circumventing the main drawbacks of the layered-based approach traditionally used to develop such protocols, whilst at the same time providing a flexible, simple and well structured way to implement them. The event channel service provided by EVA establishes how entities that share the same address space interact. Then, the application designer has the opportunity to define the most appropriate lower level communication protocols that control the way entities that execute within different processes will interact. The framework specifies a way to accommodate these protocols and provides several standard protocol implementations. Further a development methodology is described for constructing applications on top of the framework. In designing the framework, we have followed the approach of using, whenever possible, well established concepts, thus the paper also discusses the utilisation of such concepts in improving both the efficiency and the structuring of the framework and of the applications to be built on top of it. Francisco Vilar Brasileiro, Fabíola Greve, Frédéric Tronel, Michel Hurfin, Jean-Pierre Le Narzul |
NCA | 1 |
| 1998 | Applying coloured Petri nets to analyze fail silent nodes in distributed systemsabstractA fail-silent node is a self-checking node composed of a number of conventional fail-uncontrolled processors that work together to provide the following fail-controlled behavior: the node either functions correctly or stops functioning after an internal failure is detected. In a software implemented fail-silent node, the non-faulty processors of the node need to execute message order and comparison protocols to keep in step and check each other respectively. In this paper we present a Petri net model for a software implemented fail-silent node specification. Formal analysis by means of occurrence graph is also shown. Lívia M. R. Sampaio, Jorge C. A. de Figueiredo, Francisco Vilar Brasileiro |
SMC | 3 |
| 1996 | Implementing Fail-Silent Nodes for Distributed SystemsabstractA fail-silent node is a self-checking node that either functions correctly or stops functioning after an internal failure is detected. Such a node can be constructed from a number of conventional processors. In a software-implemented fail-silent node, the nonfaulty processors of the node need to execute message order and comparison protocols to "keep in step" and check each other, respectively. In this paper, the design and implementation of efficient protocols for a two processor fail-silent node are described in detail. The performance figures obtained indicate that in a wide class of applications requiring a high degree of fault tolerance, software-implemented fail-silent nodes constructed simply by utilizing standard "off-the-shelf" components are an attractive alternative to their hardware-implemented counterparts that do require special-purpose hardware components, such as fault-tolerant clocks, comparator, and bus interface circuits. Francisco Vilar Brasileiro, Paul D. Ezhilchelvan, Santosh K. Shrivastava, Neil A. Speirs |
IEEE Trans. Computers | 1 |
| 1995 | TMR Processing without Explicit Clock SynchronizationabstractReplicated processing with majority voting is a well known method for achieving fault tolerance. Triple Modular Redundant (TMR) processing is the most commonly used version of that method. Replicated processing requires that the replicas reach agreement on the order in which messages are to be processed. Synchronous and deterministic ordering protocols published in the literature require that the replicas maintain an abstraction of clocks that are kept in known and bounded synchronism. We present a protocol for TMR systems that does not require this abstraction of synchronised clocks. We analyse the protocol performance and show that this protocol in practice can be at least as fast as any synchronised clock based ordering protocol. We also derive a faster protocol that has an improved performance in the absence of processor failures. We then build a TMR node and measure its performance to illustrate that the protocols developed here provide faster ordering and are easier to implement. Francisco Vilar Brasileiro, Paul D. Ezhilchelvan, Neil A. Speirs |
SRDS | 1 |