Filipe Araújo

dblp:18/5512 · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
11since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Security and privacy · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorComputer networks · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2024 Efficient Write Operations in Event Sourcing with Replication
abstract
Event Sourcing (ES) is an architectural pattern where applications maintain a record of all events that alter their state. While replication of the event $\log$ can enhance dependability, existing ES frameworks do not exploit replication to accelerate write operations. The challenge lies in enabling write operations on replicas without disrupting the sequential ordering of the log or incurring significant communication overhead between geographically distributed replicas. Our approach addresses this by partitioning resources across replicas, ensuring that most operations require only local ordering, reserving interreplica communication for scenarios where resources are limited. We implemented this approach using the Axon framework and evaluated its performance against MongoDB Replica Sets, where writes must go through the primary replica. The results demonstrate that our method significantly improves the speed of local write operations, thereby fully leveraging the benefits of replication.
Tiago Rolo, Nuno M. Preguiça, Filipe Araújo
NCA3
2023 Cost-Availability Aware Scaling: Towards Optimal Scaling of Cloud Services
abstract
Abstract Cloud services have become increasingly popular for developing large-scale applications due to the abundance of resources they offer. The scalability and accessibility of these resources have made it easier for organizations of all sizes to develop and implement sophisticated and demanding applications to meet demand instantly. As monetary fees are involved in the use of the cloud, one of the challenges for application developers and operators is to balance their budget constraints with crucial quality attributes, such as availability. Industry standards usually default to simplified solutions that cannot simultaneously consider competing objectives. Our research addresses this challenge by proposing a Cost-Availability Aware Scaling (CAAS) approach that uses multi-objective optimization of availability and cost. We evaluate CAAS using two open-source microservices applications, yielding improved results compared to the industry standard CPU-based Autoscaler (AS). CAAS can find optimal system configurations with higher availability, between 1 and 2 nines on average, and reduced costs, 6% on average, with the first application, and 1 nine of availability on average, and reduced costs up to 18% on average, with the second application. The gap in the results between our model and the default AS suggests that operators can significantly improve the operation of their applications.
André Bento, Filipe Araújo, Raul Barbosa
J. Grid Comput.2
2023 Efficient Causal Access in Geo-Replicated Storage Systems
abstract
Abstract We consider a setting where applications, such as websites or games, need causal access to objects available in geo-replicated cloud data stores. Common ways of implementing causal consistency involve hiding objects while waiting for their dependencies or waiting for server replicas to synchronize. To minimize delays and retrieve objects faster, applications may try to reach different server replicas at once. This entails a cost because providers charge for each reading request, including reading misses where the causal copy of the object is unavailable. Therefore, latency and cost are conflicting goals, which we control by selecting where to read and when. We formulate this challenge as a multi-criteria optimization problem and propose five non-dominated reading strategies, four of which are Pareto optimal, in a setting constrained to two server replicas. We validate these solutions on the following real cloud storage services: AWS S3, DynamoDB and MongoDB. Savings of as much as 50% on reading costs, with no significant or even a positive impact on latency, demonstrate that both clients and cloud providers could benefit from richer services compatible with these retrieval strategies.
Stanley Lima, Filipe Araújo, Miguel de Oliveira Guerreiro, Jaime Correia, André Bento, Raul Barbosa
J. Grid Comput.2
2022 Bi-objective optimization of availability and cost for cloud services
abstract
Cloud-based services are a current approach for developing large-scale applications with advantages such as flexibility, access to on-demand resources, and business agility. The overall application functionality results from complex interactions of many decoupled services, each having its operational specificity. Due to this complexity, the manual configuration of these systems is very arduous, error-prone and likely to impair the quality of service, leading to malfunctioning services, lowering availability and accruing costs. Identifying the optimal solution to simultaneously optimize availability and costs, whilst meeting service level objectives remains a challenge for professionals developing solutions using cloud services. This paper proposes a mathematical formulation of a bi-objective problem to identify the optimal set of solutions for the system configuration. Empirical evaluation of the proposed approach in a case study of a real industrial scenario results in an R-Squared of 0.85, an MSE of 0.021 and an optimization accuracy of 0.928. These methods can help practitioners to keep services at an optimum configuration enabling autonomic service operation, whilst improving availability and cost.
André Bento, João Durães, José Ferreira, Rita Carreira, Filipe Araújo, Raul Barbosa
NCA7
2022 The Effects of Soft Errors and Mitigation Strategies for Virtualization Servers
abstract
Virtualized servers compose the majority of cloud computing environments, where these nodes are used to host multiple clients over the same hardware. Many organizations run online applications by hiring elastic computing resources in order to match demand while reducing fixed costs. However, such organizations are unlikely to take advantage of these benefits for critical applications, as it would expose them to several risks. Among other threats, soft errors are a concern in large-scale reliable servers and are expected to become more frequent as a consequence of smaller transistors and lower operating voltages of integrated circuits. This article characterizes virtualized servers of cloud environments in presence of soft errors. Using fault injection, we collect experimental data to determine the failure modes of applications, operating systems, VMs, and hypervisor. The analysis exposes distinct failure modes, ranging from crash failures of a single virtual machine to silent data corruption in permanent storage. The most frequent failure mode, observed in 10–30 percent of injected errors, consists of a hang affecting multiple virtual machines. Given that such failures are a primary cause of downtime, we develop and evaluate a recovery mechanism which uses online testing and recovers a server from all hangs by rebooting its hypervisor.
Frederico Cerveira, Raul Barbosa, Henrique Madeira, Filipe Araújo
IEEE Trans. Cloud Comput.4
2021 μ Viz: Visualization of Microservices
abstract
Microservice architectures have become very popular and widely adopted by the industry, because of the benefits they bring to the software development process and resulting systems, such as parallel development, modularity and scalability. However, as interfaces become more fine-grained and systems grown in size, complexity is moved from the component services to their interactions, eventually leading to intricate workflows that are hard to observe, visualize, and understand. This problem is compounded by the typically high workloads that produce intractable amounts of observation data. To deal with these challenges, operators need support from tools able to take in observation data, in particular tracing, and provide a fast and intuitive understanding of which components or workflows require attention and how are they affecting a module, service, instance, or the whole application. In this paper, we present the design of a microservice visualization application that can fill a gap that exists in leveraging tracing data, aggregating and navigating it in ways that are actionable for operators. Our application provides multiple views of the system and uses spatial and hierarchical navigation using flip zoom to simplify their exploration, while preserving context. Our application can provide a better understanding of the system than existing applications that lack navigability and do not preserve context when switching between different services, layers or views.
Sara Silva, Jaime Correia, André Bento, Filipe Araújo, Raul Barbosa
IV4
2021 A layered framework for root cause diagnosis of microservices
abstract
Microservice-based architectures feature function-ally independent, well-defined and fine-grained components suit-able for loosely coupled deployments and for building reli-able cloud-native applications. Despite the advantages of this approach, component interactions introduce complexity, thus turning boundary -spanning service operation into a daunting challenge. As systems grow in size, complexity can easily outgrow the cognitive capacity of human operators, who are unable to effectively diagnose faulty microservices. We address this problem by proposing a novel framework to diagnose faulty microservices. Through failure injection and an experimental assessment, our layered diagnosis framework using service response analysis, timing constraints, causality and a ranking algorithm from traces, is able to effectively diagnose faulty microservices. Empirical evaluation of the proposed approach, by examining 130 experi-ments in a representative microservice application in the presence of faults, shows that it can achieve approximately 89% specificity and 77% recall.
André Bento, Jaime Correia, João Durães, Luís Ribeiro, Rita Carreira, Filipe Araújo, Raul Barbosa
NCA8
2021 Automated Analysis of Distributed Tracing: Challenges and Research Directions
André Bento, Jaime Correia, Ricardo Filipe, Filipe Araújo, Jorge Cardoso 0001
J. Grid Comput.4
2021 End-to-end secure group communication for the Internet of Things
André Lizardo, Raul Barbosa, Samuel Neves, Jaime Correia, Filipe Araújo
J. Inf. Secur. Appl.5
2021 Improving observability in Event Sourcing systems
Stanley Lima, Jaime Correia, Filipe Araújo, Jorge Cardoso 0001
J. Syst. Softw.3
2021 Reductions and abstractions for formal verification of distributed round-based algorithms
Raul Barbosa, Alcides Fonseca, Filipe Araújo
Softw. Qual. J.3
2020 Towards optimal convergecast in wireless ad hoc networks
Filipe Araújo, André Gomes, Rui P. Rocha
Ad Hoc Networks1
2020 Benchmarking Serverless Computing Platforms
Horácio Martins, Filipe Araújo, Paulo Rupino da Cunha
J. Grid Comput.2
2019 Client-Side Monitoring of HTTP Clusters Using Machine Learning Techniques
abstract
Large online web sites are supported in the back-end by a cluster of servers behind a load balancer. Ensuring proper operation of the cluster with minimal monitoring efforts from the load balancer is necessary to ensure performance. Previous monitoring efforts require extensive data from the system and fail to include the client perspective. We monitor the cluster using machine learning techniques that process data collected and uploaded by web clients, an approach that might complement system-side information. To experiment our solution, we trained the machine learning algorithms in a cluster of 10 machines with a load balancer and evaluated the results of these algorithms when one of the machines is overloaded. While a fine-grained view of the state of the machines, may require much effort to accomplish, given the compensation effect of the remaining healthy machines, the results show that we can achieve a coarse grained view of the entire system, to produce relevant insight about the cluster.
Ricardo Filipe, Filipe Araújo
ICMLA2
2019 Towards Occupation Inference in Non-instrumented Services
abstract
Measuring the capacity and modeling the response to load of a real distributed system and its components requires painstaking instrumentation. Even though it greatly improves observability, instrumentation may not be desirable, due to cost, or possible due to legacy constraints. To model how a component responds to load and estimate its maximum capacity, and in turn act in time to preserve quality of service, we need a way to measure component occupation. Hence, recovering the occupation of internal non-instrumented components is extremely useful for system operators, as they need to ensure responsiveness of each one of these components and ways to plan resource provisioning. Unfortunately, complex systems will often exhibit non-linear responses that resist any simple closed-form decomposition. To achieve this decomposition in small subsets of non-instrumented components, we propose training a neural network that computes their respective occupations. We consider a subsystem comprised of two simple sequential components and resort to simulation, to evaluate the neural network against an optimal baseline solution. Results show that our approach can indeed infer the occupation of the layers with high accuracy, thus showing that the sampled distribution preserves enough information about the components. Hence, neural networks can improve the observability of online distributed systems in parts that lack instrumentation.
Ricardo Filipe, Jaime Correia, Filipe Araújo, Jorge Cardoso 0001
NCA3
2019 An observation on NORX, BLAKE2, and ChaCha
Samuel Neves, Filipe Araújo
Inf. Process. Lett.2
2018 Response Time Characterization of Microservice-Based Systems
abstract
In pursuit of faster development cycles, companies have favored small decoupled services over monoliths. Following this trend, distributed systems made of microservices have grown in scale and complexity, giving rise to a new set of operational problems. Even though this paradigm simplifies development, deployment, management of individual services, it hinders system observability. In particular, performance monitoring and analysis becomes more challenging, especially for critical production systems that have grown organically, operate continuously, and cannot afford the availability cost of online benchmarking. Additionally, these systems are often very large and expensive, thus being bad candidates for full-scale development replicas. Creating models of services and systems for characterization and formal analysis can alleviate the aforementioned issues. Since performance, namely response time, is the main interest of this work, we focused on bottleneck detection and optimal resource scheduling. We propose a method for modeling production services as queuing systems from request traces. Additionally, we provide analytical tools for response time characterization and optimal resource allocation. Our results show that a simple queuing system with a single queue and multiple homogeneous servers has a small parameter space that can be estimated in production. The resulting model can be used to accurately predict response time distribution and the necessary number of instances to maintain a desired service level, under a given load.
Jaime Correia, Fabio Ribeiro, Ricardo Filipe, Filipe Araújo, Jorge Cardoso 0001
NCA4
2018 On Black-Box Monitoring Techniques for Multi-Component Services
abstract
Despite the advantages of microservice and function-oriented architectures, there is an increase in complexity to monitor such highly dynamic systems. In this paper, we analyze two distinct methods to tackle the monitoring problem in a system with reduced instrumentation. Our goal is to understand the feasibility of such approach with one specific driver: simplicity. We aim to determine the extent to which it is possible to characterize the state of two generic tandem processes, using as little information as possible. To answer this question, we resorted to a simulation approach. Using a queue system, we simulated two services, that we could manipulate with distinct operation sets for each module. We used the total response time seen upstream of the system. Having this setup and metric, we applied two distinct methods to analyze the results. First, we used supervised machine learning algorithms to identify where the bottleneck is happening. Secondly, we used an exponential decomposition to identify the occupation in the two components in a more black-box fashion. Results show that both methodologies have their advantages and limitations. The separation of the signal more accurately identifies occupation in low occupied resources, but when a service is totally dominating the overall time, it lacks precision. The machine learning has a more stable error, but needs the training set. This study suggest that a black-box occupation approach with both techniques is possible and very useful.
Ricardo Filipe, Jaime Correia, Filipe Araújo, Jorge Cardoso 0001
NCA3
2018 Nonintrusive Monitoring of Microservice-Based Systems
abstract
Breaking large software systems into smaller functionally interconnected components is a trend on the rise. This architectural style, known as “microservices”, simplifies development, deployment and management at the expense of complexity and observability. In fact, in large scale systems, it is particularly difficult to determine the set of microservices responsible for delaying a client's request, when one module impacts several other microservices in a cascading effect. Components cannot be analyzed in isolation, and without instrumenting their source code extensively, it is difficult to find the bottlenecks and trace their root causes. To mitigate this problem, we propose a much simpler approach: log gateway activity, to register all calls to and between microservices, as well as their responses, thus enabling the extraction of topology and performance metrics, without changing source code. For validation, we implemented the proposed platform, with a microservices-based application that we observe under load. Our results show that we can extract relevant performance information with a negligible effort, even in legacy systems, where instrumenting modules may be a very expensive task.
Fabio Pina, Jaime Correia, Ricardo Filipe, Filipe Araújo, Jorge Cardroom
NCA4
2018 A survey on reliable distributed communication
Naghmeh Ramezani Ivaki, Nuno Laranjeiro, Filipe Araújo
J. Syst. Softw.3
2017 Client-side black-box monitoring for web sites
abstract
In spite of their growing maturity, current web monitoring tools are unable to observe all operating conditions. For example, clients in different geographical locations might get very diverse latencies to the server; the network between client and server might be slow; or third-party servers with external page resources might underperform. Ultimately, only the clients can determine whether a site is up and running in good conditions. In this paper, we use the response times experienced by clients, to infer about server and network performance. The goal is to detect internal and external bottlenecks doing black-box monitoring, in particular CPU (internal) and network (external). We aim to determine to what extent are the clients able to tell one type of bottleneck from the other, i.e., what kind of information do the server and network leak, regarding their operating conditions. To answer this question, we resort to an empirical approach. We submit an HTTP server and network to a large number of operating conditions and train two machine learning algorithms, a linear and a non-linear one, to identify the cause of the congestion affecting the system. Results show that the server and network leak information to a level of detail that allows sorting out CPU from network bottlenecks, or even a combination of the two, in a large spectrum of cases. This suggests that a black-box monitoring approach is not only possible, but promising, as it may complement traditional white-box approaches.
Ricardo Filipe, Rui Pedro Paiva, Filipe Araújo
NCA3
2016 Client-side monitoring techniques for web sites
abstract
Ensuring the correct presentation and execution of web sites is a major concern for system developers and administrators. Unfortunately, only end users can determine which resources are available and working properly. For example, some internal or external addresses might be unavailable or unreachable for specific clients, while seemingly available resources, like JavaScript, might run with errors in some browsers. While standard monitoring and analytic tools certainly provide valuable information on web pages, problems might still escape such measures, to reach end web users. To demonstrate the limitations of current tools, we ran an experiment to count web page errors in a sample of 3,000 web sites, including network and JavaScript errors. Our results are significant: as many as 16% of the top 1,000 sites have errors in their own resources; less popular sites have even more. Based on these results, we make a review of three client-side monitoring approaches to mitigate such errors: stand-alone applications, browser extensions and JavaScript snippets with analytic tools. Interestingly, even the latter approach, which requires no software installation, and involves no security changes, can cover a large fraction of existing web errors.
Ricardo Filipe, Filipe Araújo
NCA2
2016 Towards designing reliable messaging patterns
abstract
Reliable communication is nowadays pervasively supported by TCP, which is poorly adapted for message-based communications, because it offers a streaming channel with no mechanisms to encapsulate messages. Moreover, TCP does not tolerate connection crashes. Thus, whenever reliable message-based communication is needed, developers either use heavy-weight middleware, like Java Message Service (JMS), or develop their own custom error-prone solutions for recovering from crashes. In this paper, we introduce two TCP-based design patterns that address these limitations, and facilitate the development of light-weight and reliable message-based applications. Our design solutions are modular, in the sense that they build on top of each other.
Naghmeh Ramezani Ivaki, Nuno Laranjeiro, Filipe Araújo
NCA3
2015 Straight-line programs for fast sparse matrix-vector multiplication
abstract
Summary Sparse matrix‐vector multiplication dominates the performance of many scientific and industrial problems. For example, iterative methods for solving linear systems often rely on the performance of this critical operation. The particular case of binary matrices shows up in several important areas of computing, such as graph theory and cryptography. Unfortunately, irregular memory access patterns cause poor memory throughput, slowing down this operation. To maximize memory throughput, we translate the matrix into a straight‐line program that takes advantage of the CPU's instruction cache and hardware prefetchers. The regular loopless pattern of the program reduces cache misses, thus decreasing the latency for most instructions. We focus on the widely used x86_64 architecture and on binary matrices, to explore several possible tradeoffs regarding memory access policies and code size. We also consider matrices with elements over various mathematical structures, such as floating‐point reals and integers modulo m . When compared to a Compressed Row Storage implementation, we obtain significant speedups. Copyright © 2014 John Wiley & Sons, Ltd.
Samuel Neves, Filipe Araújo
Concurr. Comput. Pract. Exp.2
2014 Taking an electronic ticketing system to the cloud: Design and discussion
abstract
In this paper we address the challenge of creating an electronic ticketing system for transportation systems that can partially or completely run on the cloud. This challenge is defined within the scope of an industrial project. The resulting system should be able to reach a large spectrum of customers and should provide two key advantages: lower operational costs, especially for small clients without IT departments, and faster execution of queries for monthly or other sorts of analysis, using the elasticity of cloud-based resources. To fulfill the goals of the project, we propose very standard technologies and procedures: a three-tiered architecture; a separation of the online and analysis databases; and an Enterprise Service Bus to get the input from very diverse hardware and software stacks. In this paper we discuss several options regarding the location of these facilities on the cloud and we also evaluate the costs involved. While this work already defines many features of the system, it must be considered as preliminary, as some open details remain for future work.
Filipe Araújo, Marília Curado, Pedro Furtado 0001, Raul Barbosa
IEEE BigData1
2014 Fault-Tolerant bi-directional communications in web-based applications
abstract
The Hypertext Transfer Protocol (HTTP) and the Transmission Control Protocol (TCP) are the most popular protocols used in the development of web-based applications. Despite their popularity, the use of these protocols brings two limitations to applications and systems that require reliable interactive real-time communications: 1) HTTP forces applications to work in a request-response paradigm, even if a reply is not necessary, not allowing the server to send anything to a client without the client explicitly requesting it; 2) TCP provides no recovery options for network outages, thus forcing developers to write their own error-prone, complex, and ad hoc solutions. In this paper we introduce a solution that offers both bi-directional and reliable communication to web-based applications, even in presence of connection failures. To make this possible, we combine the idea behind WebSockets and a Session-Based Fault-Tolerant design pattern.
Naghmeh Ramezani Ivaki, Filipe Araújo
ICPADS2
2014 Session-based fault-tolerant design patterns
abstract
Despite offering reliability against dropped and reordered packets, the widely adopted Transmission Control Protocol (TCP) provides nearly no recovery options for longterm network outages. When the network fails, developers must rollback the application to some coherent state on their own, using error-prone solutions. Overcoming this limitation is, therefore, a deeply investigated and challenging problem. Existing solutions range from transport-layer to application-layer protocols, including additions to TCP, usually transparent to the application. None of these solutions is perfect, because they all impact TCP's simplicity, performance or ubiquity, if not all. To avoid these shortcomings, we contain TCP connection crashes inside a single session layer exposed as a sockets interface. Based on this interface, we create a blocking and a non-blocking fault-tolerant design pattern. We explore the blocking design in an open source File Transfer Protocol (FTP) server and perform a thorough evaluation of performance, complexity and overhead of both designs. Our results show that using one of the patterns to tolerate TCP connection crashes, in new or existing applications, involves a very limited effort and negligible penalties.
Naghmeh Ramezani Ivaki, Filipe Araújo, Fernando J. Barros
ICPADS2
2014 Design of Multi-threaded Fault-Tolerant Connection-Oriented Communication
abstract
Fault-tolerance is vital for dependable distributed applications that can deliver service, even in the presence of faults. Over the last few decades, above all protocols proposed to offer reliability and fault-tolerance, TCP grew to become one of the cornerstones of the Internet. However, despite emulating reliable communication in distributed environments, TCP does not handle connection failures when the connectivity is lost for some time, even if both endpoints are still running. When this occurs, developers must rollback the peers to some coherent state, many times with error-prone, ad hoc, or custom application-level solutions. In this paper, we refine the Acceptor-Connector design pattern to tackle the TCP unreliability problem. The pattern decouples the failure-related processing from the connection and service processing, efficiently handling different connections and their possible crashes concurrently, thereby yielding more reusable, extensible, and efficient distributed communication. The solution we propose incorporates proven multi-threaded solutions and a buffering scheme that discards the need for an application-layer acknowledgment scheme. This simplifies the development of reliable connection-oriented applications using the ubiquitous TCP protocol.
Naghmeh Ramezani Ivaki, Filipe Araújo, Fernando J. Barros
PRDC2
2014 CloudBFT: Elastic Byzantine Fault Tolerance
abstract
Cloud computing is increasingly important, with the industry moving towards outsourcing computational resources as a means to reduce investment and management costs, while improving security, dependability and performance. Cloud operators use multi-tenancy, by grouping virtual machines (VMs) into a few physical machines (PMs), to pool computing resources, thus offering elasticity to clients. Although cloud-based fault tolerance schemes impose communication and synchronization overheads, the cloud offers excellent facilities for critical applications, as it can host varying numbers of replicas in independent resources. Given these contradictory forces, determining whether the cloud can host elastic critical services is a major research question. We address this challenge from the perspective of a standard three-tiered system with relational data. We propose to tolerate Byzantine faults using groups of replicas placed on distinct physical machines, as a means to avoid exposing applications to correlated failures. To improve the scalability of our system, we divide data to enable parallel accesses. Using a realistic setup, this setting can reach speedups largely exceeding the number of partitions. Even for a wide variation of the load, the system preserves latency and throughput within reasonable bounds. We believe that the elasticity we observe demonstrates the feasibility of tolerating Byzantine faults in a cloud-based server using a relational database.
Filipe Araújo, Raul Barbosa
PRDC2
2012 Dynamic Interaction Models for Web Enabled Wireless Sensor Networks
abstract
Wireless Sensor Networks (WSNs) are a fundamental technology for science in many domains, and their inclusion in Web environments, e.g. through Web services, allows for their open/standard access and integration. Although such Web enabled WSNs simplify data access, network parametrisation and aggregation, the available interaction models and run-time adaptation mechanisms are still scarce. Nevertheless, applications increasingly demand richer and more flexible interface accesses, for instance, the interaction model dynamic adaptation according to contextual information. To this extent, this paper discusses the relevance of the session and pattern abstractions on the design of a middleware prototype providing richer interaction models, as well as a few context-based dynamic adaptation mechanisms for Web enabled WSNs.
Maria Cecilia Gomes, Hervé Paulino, Adérito Baptista, Filipe Araújo
ISPA4
2012 A Middleware for Exactly-Once Semantics in Request-Response Interactions
abstract
Although the need for the exactly-once request-response interaction pattern is ubiquitous in distributed systems, making it work in practice is anything but simple. Ensuring the at-most-once part of the invocation is relatively easy. Unfortunately, the same is not true for the at-least-once guarantee, which depends on the recovery from crashes of the client, the server and the network. This is what makes the exactly-once interaction so difficult in practice: client and server must log their actions into stable storage, and they must be able to restart the network connections. In this paper, we present a middleware that implements the exactly-once request-response pattern, in presence of network and endpoints crashes. The main contribution of our work is to release the programmer from the complex tasks of recovering from message losses and network crashes.
Naghmeh Ramezani Ivaki, Filipe Araújo, Raul Barbosa
PRDC2
2011 On the performance of GPU public-key cryptography
abstract
Graphics processing units (GPUs) have become increasingly popular over the last years as a cost-effective means of accelerating various computationally intensive tasks. We study the particular case of modular exponentiation, the crucial operation behind most modern public-key cryptography algorithms. We focus our attention on the NVIDIA GT200 architecture, currently one of the most popular for general purpose GPU computation. We report our efforts to run modular exponentiation faster than any other method we were aware of for GPUs. Part of our performance advantage results from a different interleaving of the Montgomery multiplication, which was neglected in previous literature. The other part comes from carefully exploring general techniques, like loop unrolling and inline PTX assembly. Our throughput results, at over 20000 RSA-1024 decryptions per second or 41426 512-bit modular exponentiations per second, present a significant speedup over previous GPU implementations, without any significant latency penalty. Lastly, we evaluate our results in light of several popular metrics, namely performance/price and performance/watt ratios. We find that, while current GPUs generally perform better than CPUs, they show worse performance/watt ratios.
Samuel Neves, Filipe Araújo
ASAP2
2011 A maximum independent set approach for collusion detection in voting pools
Filipe Araújo, Jorge Farinha, Patrício Domingues, Gheorghe Cosmin Silaghi, Derrick Kondo
J. Parallel Distributed Comput.1
2010 Extending the EGEE Grid with XtremWeb-HEP Desktop Grids
abstract
Desktop Grids and Service Grids are widely used by scientific communities to execute high throughput application. The European EDGeS project aims at developing the technologies to bridge these two kinds of Grid technologies together. In this paper, we present the development and the application of these new technology to extend the EGEE Grid with XtremWeb-HEP based Desktop Grids. We present the setup of the distributed infrastructure which enables EGEE users' jobs to run on one of the three various XtremWeb-HEP DGs. We describe the new volunteer computing project EGEE@Home based on XtremWeb-HEP middleware. To evaluate the capacity of our technology, we present a performance measurement of the main components and conclude that the overhead is kept reasonable despite the fact that the infrastructure is highly distributed.
Haiwu He, Gilles Fedak, Péter Kacsuk, Zoltán Farkas, Zoltán Balaton, Oleg Lodygensky, Etienne Urbah, Gabriel Caillat, Filipe Araújo, Ad Emmen
CCGRID9
2009 Monitoring the EDGeS project infrastructure
abstract
EDGeS is an European funded Framework Program 7 project that aims to connect desktop and service grids together. While in a desktop grid, personal computers pull jobs when they are idle, in service grids there is a scheduler that pushes jobs to available resources. The work in EDGeS goes well beyond conceptual solutions to bridge these grids together: it reaches as far as actual implementation, standardization, deployment, application porting and training. One of the work packages of this project concerns monitoring the overall EDGeS infrastructure. Currently, this infrastructure includes two types of desktop grids, BOINC and XtremWeb, the EGEE service grid, and a couple of bridges to connect them. In this paper, we describe the monitoring effort in EDGeS: our technical approaches, the goals we achieved, and the plans for future work.
Filipe Araújo, Diogo Ferreira, Jorge Farinha, Patrício Domingues, Luís Moura Silva, Etienne Urbah, Oleg Lodygensky, Haiwu He, Attila Csaba Marosi, Gabor Gombás, Zoltán Balaton, Zoltán Farkas, Péter Kacsuk
IPDPS1
2009 Evaluating the performance and intrusiveness of virtual machines for desktop grid computing
abstract
We experimentally evaluate the performance overhead of the virtual environments VMware Player, QEMU, VirtualPC and VirtualBox on a dual-core machine. Firstly, we assess the performance of a Linux guest OS running on a virtual machine by separately benchmarking the CPU, file I/O and the network bandwidth. These values are compared to the performance achieved when applications are run on a Linux OS directly over the physical machine. Secondly, we measure the impact that a virtual machine running a volunteer @home project worker causes on a host OS. Results show that performance attainable on virtual machines depends simultaneously on the virtual machine software and on the application type, with CPU-bound applications much less impacted than IO-bound ones. Additionally, the performance impact on the host OS caused by a virtual machine using all the virtual CPU, ranges from 10% to 35%, depending on the virtual environment.
Patrício Domingues, Filipe Araújo, Luís Moura Silva
IPDPS2
2009 Defeating Colluding Nodes in Desktop Grid Computing Platforms
Gheorghe Cosmin Silaghi, Filipe Araújo, Luís Moura Silva, Patrício Domingues, Álvaro Enrique Arenas
J. Grid Comput.2
2009 Single-step creation of localized Delaunay triangulations
Filipe Araújo, Luís E. T. Rodrigues
Wirel. Networks1
2008 Defeating colluding nodes in Desktop Grid computing platforms
abstract
Desktop grid systems reached a preeminent place among the most powerful computing platforms in the planet. Unfortunately, they are extremely vulnerable to mischief because volunteers can output bad results, for reasons ranging from faulty hardware (like over-clocked CPUs) to intentional sabotage. To mitigate this problem, desktop grid projects replicate work units and apply majority voting, typically on 2 or 3 results. In this paper, we observe that this form of replication is powerless against malicious volunteers that have the intention and the (simple) means to ruin the project using some form of collusion. We argue that each work unit needs at least 3 voters and that voting pools with conflicts enable the master to spot colluding malicious nodes. Hence, we post- process the voting pools in two steps: i) we use a statistical approach to identify nodes that were not colluding, but submitted bad results; ii) we use a rather simple principle to go after malicious nodes which acted together: they might have won conflicting voting pools against nodes that were not identified in step i. We use simulation to show that our heuristic can be quite effective against colluding nodes, in scenarios where honest nodes form a majority.
Gheorghe Cosmin Silaghi, Patrício Domingues, Filipe Araújo, Luís Moura Silva, Álvaro Enrique Arenas
IPDPS3
2007 Characterizing Result Errors in Internet Desktop Grids
Derrick Kondo, Filipe Araújo, Paul Malecot, Patrício Domingues, Luís Moura Silva, Gilles Fedak, Franck Cappello
Euro-Par2
2006 A DHT-Based Infrastructure for Sharing Checkpoints in Desktop Grid Computing
abstract
In this paper we present Chkpt2Chkpt, a desktop grid system that aims to reduce turnaround times of applications by replicating checkpoints. We target desktop computing projects with applications that are comprised of longrunning independent tasks, executed in hundreds or thousands of computers spread over the Internet. While these applications typically do local checkpointing to deal with failures, we propose to replicate those checkpoints in remote places to make them available to other worker nodes. The main idea is to organize the worker nodes of a desktop grid into a peer-to-peer Distributed Hash Table. Worker nodes can take advantage of this P2P network to keep track, share, manage and reclaim the space of the checkpoint files. We used simulation to validate our system and we show that remotely storing replicas of checkpoints can considerably reduce the turnaround times of the tasks, when compared to the traditional approaches where nodes manage their own checkpoints locally. These results make us conclude that the application of P2P techniques seems to be quite helpful in wide-scale desktop grid environments.
Patrício Domingues, Filipe Araújo, Luís Moura Silva
e-Science2
2005 Long Range Contacts in Overlay Networks
Filipe Araújo, Luís E. T. Rodrigues
Euro-Par1
2005 Scalable QoS-Based Event Routing in Publish-Subscribe Systems
abstract
This paper proposes a distributed and scalable publish-subscribe broker with support for QoS. The broker, called IndiQoS, leverages on existing mechanisms to reserve resources in the underlying network and on an overlay network of peer-to-peer rendezvous nodes, to automatically select QoS-capable paths. By avoiding flooding of either QoS reservations or link-state information, IndiQoS is able to scale with respect to network size and number of reservations. Experimental results show the validity of our approach
Nuno Carvalho, Filipe Araújo, Luís E. T. Rodrigues
NCA2
2004 GeoPeer: A Location-Aware Peer-to-Peer System
abstract
This work presents a novel peer-to-peer system that is particularly well suited to support context-aware computing. The system, called GeoPeer, aims to combine the advantages of peer-to-peer systems that implement distributed hash tables with the suitability of geographical routing for supporting location-constrained queries and information dissemination. GeoPeer is comprised of two fundamental components: a Delaunay triangulation used to build a connected lattice of nodes and a mechanism to manage long range contacts that allows good routing performance, despite unbalanced distribution of nodes.
Filipe Araújo, Luís E. T. Rodrigues
NCA1
2004 Fast Localized Delaunay Triangulation
Filipe Araújo, Luís E. T. Rodrigues
OPODIS1
2001 A neural network for shortest path computation
abstract
This paper presents a new neural network to solve the shortest path problem for inter-network routing. The proposed solution extends the traditional single-layer recurrent Hopfield architecture introducing a two-layer architecture that automatically guarantees an entire set of constraints held by any valid solution to the shortest path problem. This new method addresses some of the limitations of previous solutions, in particular the lack of reliability in what concerns successful and valid convergence. Experimental results show that an improvement in successful convergence can be achieved in certain classes of graphs. Additionally, computation performance is also improved at the expense of slightly worse results.
Filipe Araújo, Bernardete Ribeiro, Luís E. T. Rodrigues
IEEE Trans. Neural Networks1