Raul Barbosa

dblp:19/5720 · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-2916-7571ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 3 first-author · 8 since 2021Security and privacy · 14 · 4 first-author · 2 since 2021Systems, architecture and hardware · 8 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A GenAI-Driven Multi-Agent Framework for Explainable Intent-Based Slice Recommendation
Rui Ferreira 0001, Raul Barbosa, Marco Araújo, Petia Georgieva, Susana Sargento, Anabela Tereso, Paulo Novais, Pedro Rito, Bruno Mendes
NetSoft2
2026 Automated formalisation of informal specifications by combining LLMs and grammar-based language processing
abstract
Abstract Informal specifications of software artefacts need to be formalised to facilitate rigorous analysis, including deductive verification. Automated formalisation of informal specifications can significantly reduce the effort of manual translation. However, a major challenge lies in bridging the gap between the rich syntax of natural language and the need for precise semantics in formal specification languages. Building on the empirical observation that modern pre-trained large language models (LLMs) can effectively handle the breadth of natural language, while symbolic natural language processing (NLP) is more efficient in formal languages, this paper proposes an approach that combines both methodologies. The proposed solution, Hybrid Automated Formalisation of Informal Specifications (HAFIS), uses an LLM to restrict the syntax of informal specifications into the language of a formal grammar, and proposes the concept of Typed Semantic Interpretation to enforce the semantics of the resulting formal specifications. Using a public set of real-world informal specifications, we evaluated HAFIS and compared it with a purely symbolic approach and the direct use of LLMs as baselines. Results show that HAFIS increases language acceptance from 23% to 100% and accurately translates 88% of the cases into Java Modeling Language (JML). Furthermore, mutation analysis shows that the HAFIS-generated JML effectively finds defects in programs. These results substantiate that the proposed approach is effective and contributes to software analysis using natural language.
Iat Tou Leong, Raul Barbosa
Empir. Softw. Eng.2
2026 A software architecture for verifiable and explainable classification
abstract
Abstract In the context of machine learning, classification is the procedure of predicting the class to which each element of a population belongs to. Most classification functions, for real world problems, are imperfect and thus require rigorous analysis for use in safety-critical applications such as health care. This paper proposes a software architecture for improving the trustworthiness and explainability of AI-based classifiers. The architecture combines a search-based approach with machine-learned explanations and satisfiability solving, to provide an indication of classification confidence and counterfactual explanation rules that are deductively verified to be consistent with the classifier. An implementation of the proposed architecture is evaluated on a medical case study of prognosis of Acute Coronary Syndrome (ACS). The evaluation shows that the proposed architecture is consistently able to complement each individual classification with an indication of confidence and an explanation, which is formally verified for consistency with the classifier. This contributes to foster trustworthy and explainable classification.
Raul Barbosa, Salvatore Rinzivillo, Jacques Robin, Andrea Beretta, Henrique Madeira
Mach. Learn.1
2024 Translating meaning representations to behavioural interface specifications
abstract
Higher-order logic can be used for meaning representation in natural language processing to encode the semantic relationships in text. Alternatively, using a formal specification language for meaning representation is more precise for specifying programs and widely supported by automatic theorem provers, while deductive verification based on higher order logic is less common for mainstream programming languages. This paper addresses the research question of translating higher-order logic meaning representations generated from method-level code comments into a formal specification language that extends first-order logic. Doing so requires resolving possible ambiguities in determining the appropriate semantics for predicates. This is an open challenge in the path toward using natural language processing with formal methods. To address this, the paper proposes an approach and constructs a compiler for translating meaning representations, generated from Java programs with method-level comments, into Java Modeling Language. We evaluate the compiler on a set of representative benchmarks, including programs and specifications from the Java API, by generating Java Modeling Language specifications and statically checking them with a theorem prover. Results show that in 94% of the cases Java Modeling Language is accurately generated and in 97% of those cases it can be automatically checked with a state-of-the-art theorem prover.
Iat Tou Leong, Raul Barbosa
J. Syst. Softw.2
2023 Demo: Object detection under 5G-edge mobility
abstract
In the mid-term future, vehicles will generate large amounts of data for both standalone usage (e.g., to recognize road features and external elements such as lanes, signs, and pedestrians) and cooperative usage (e.g., lane merging). However, processing the captured video and image data results comes with significant computational requirements (e.g., GPUs). Computer vision tasks, such as feature extraction, are unfeasible from a business perspective if performed directly in the User Equipment (UE), as automotive manufacturers are unwilling to increase the end-product’s costs. Thus, the logical solution is to collect and upload this data to be processed elsewhere. Nonetheless, processing the data as close to the vehicle is important due to latency constraints, thus calling for the use of Mobile Edge Computing (MEC). An additional benefit of this scenario, in which 5G connectivity enables data to be offloaded to the edge, is that the data from our car is not processed alone. Data from several sources, e.g., multiple vehicles and fixed cameras, can be offloaded to the edge node and processed together, enhancing its quality as more sources of data enhance the prediction output of machine-learning models. This demo showcases a video recording from a vehicle uploaded to an edge node via 5G software-defined-radio FPGA devices. There, a YOLO application to detect objects processes the video and communicates this information to the vehicle, ensuring QoS metrics even when the UE performs handover to a different cell or geographical area.
Marco Araújo, Pedro M. Santos 0002, Deepak Gunjal, João Pedro Fonseca 0001, Paulo Duarte, Bruno Mendes, Raul Barbosa, Peter Steenkiste, Saeid Sabamoniri, Luis Lam, Harrison Kurunathan
WoWMoM9
2023 Demo: Enhancing Network Performance based on 5G Network Function and Slice Load Analysis
abstract
The Fifth Generation Mobile Networks has transformed the paradigm of mobile network communications. In Beyond Fifth Generation Networks networks, Machine Learning (ML) and Artificial Intelligence (AI) are crucial components, optimizing network resource management to improve the network performance as well as end-users Quality of Service while lowering the network operating costs. This work makes use of an End-to-End 5G architecture to validate three demonstrations: 1) Radio Access Network monitoring using a Flexible RIC’s xApp; 2) 5G Core Network’s metrics collection via Capgemini Engineering’s Network Data Analytics Function; 3) Analysis of the Core Network’s collected data to predict Network Function load and Network Slice Instance load through the Capgemini Engineering’s NetAnticipate AI/ML engine.
Rui Ferreira 0001, João Pedro Fonseca 0001, Mayuri Tendulkar, Paulo Duarte, Marco Araújo, Raul Barbosa, Bruno Mendes, Adriano Almeida Góes
WoWMoM7
2023 Cost-Availability Aware Scaling: Towards Optimal Scaling of Cloud Services
abstract
Abstract Cloud services have become increasingly popular for developing large-scale applications due to the abundance of resources they offer. The scalability and accessibility of these resources have made it easier for organizations of all sizes to develop and implement sophisticated and demanding applications to meet demand instantly. As monetary fees are involved in the use of the cloud, one of the challenges for application developers and operators is to balance their budget constraints with crucial quality attributes, such as availability. Industry standards usually default to simplified solutions that cannot simultaneously consider competing objectives. Our research addresses this challenge by proposing a Cost-Availability Aware Scaling (CAAS) approach that uses multi-objective optimization of availability and cost. We evaluate CAAS using two open-source microservices applications, yielding improved results compared to the industry standard CPU-based Autoscaler (AS). CAAS can find optimal system configurations with higher availability, between 1 and 2 nines on average, and reduced costs, 6% on average, with the first application, and 1 nine of availability on average, and reduced costs up to 18% on average, with the second application. The gap in the results between our model and the default AS suggests that operators can significantly improve the operation of their applications.
André Bento, Filipe Araújo, Raul Barbosa
J. Grid Comput.3
2023 Efficient Causal Access in Geo-Replicated Storage Systems
abstract
Abstract We consider a setting where applications, such as websites or games, need causal access to objects available in geo-replicated cloud data stores. Common ways of implementing causal consistency involve hiding objects while waiting for their dependencies or waiting for server replicas to synchronize. To minimize delays and retrieve objects faster, applications may try to reach different server replicas at once. This entails a cost because providers charge for each reading request, including reading misses where the causal copy of the object is unavailable. Therefore, latency and cost are conflicting goals, which we control by selecting where to read and when. We formulate this challenge as a multi-criteria optimization problem and propose five non-dominated reading strategies, four of which are Pareto optimal, in a setting constrained to two server replicas. We validate these solutions on the following real cloud storage services: AWS S3, DynamoDB and MongoDB. Savings of as much as 50% on reading costs, with no significant or even a positive impact on latency, demonstrate that both clients and cloud providers could benefit from richer services compatible with these retrieval strategies.
Stanley Lima, Filipe Araújo, Miguel de Oliveira Guerreiro, Jaime Correia, André Bento, Raul Barbosa
J. Grid Comput.6
2023 Quality Evaluation of Modern Code Reviews Through Intelligent Biometric Program Comprehension
abstract
Code review is an essential practice in software engineering to spot code defects in the early stages of software development. Modern code reviews (e.g., acceptance or rejection of pull requests with Git) have become less formal than classic Fagan's inspections, lightweight, and more reliant on individuals (i.e., reviewers). However, reviewers may encounter mentally demanding challenges during the code review, such as code comprehension difficulties or distractions that might affect the code review quality. This work proposes a novel approach that evaluates the quality of code reviews in terms of bug-finding effectiveness and provides the reviewers with a clear message of whether the review should be repeated, indicating the code regions that may not have been well-reviewed. The proposed approach utilizes biometric information collected from the reviewer during the review process using non-intrusive biofeedback devices (e.g., smartwatches). Biometric measures such as Heart Rate Variability (HRV) and task-evoked pupillary response are captured as a surrogate of the cognitive state of the reviewer (e.g., mental workload) and inexpensive desktop eye-trackers compatible with the software development settings. This work uses Artificial Intelligence techniques to predict the cognitive load from the extracted biomarkers and classify each code region according to a set of features. The final evaluation considers various factors such as code complexity, time of the code review, the experience level of the reviewer, and other factors. Our experimental results show the approach could predict the review quality with 87.77%±4.65 accuracy and a Spearman correlation coefficient of 0.85 (p-value < 0.001) between the predicted and the actual review performance. This evaluation validates the cognitive load measurement using electroencephalography (EEG) signals as ground truth for the HRV and pupil signals.
Haytham Hijazi, João Durães, Ricardo Couceiro, João Castelhano, Raul Barbosa, Júlio Medeiros, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
IEEE Trans. Software Eng.5
2022 AI-driven Human-centric Control Interfaces for Industry 4.0 with Role-based Access
abstract
Nowadays, we are accustomed to using voice virtual assistants in Smart Home contexts or to ask for instructions. The potential that this technology presents has been growing, however, and even with the transition to Industry 4.0, these capabilities are significantly unexplored in industrial environments. In these contexts of high automation, human operators will have more sophisticated interventions, being supported and collaborating with intelligent systems to execute operations and receive information dynamically and proactively. The goal of this work is to introduce areas such as Machine Learning, in order to recognize the user by his or her voice, and the Intent Based Networking area, which allows the orchestration of the network through intents, a type of policy that expresses objectives without mentioning how they are implemented. The system will use Machine Learning models to recognise and provide permissions to users that, depending on the permissions, will implement policies over the network. Since specific knowledge about networks is not generally known, this technology has made it easier for the user to orchestrate the network using voice commands.
Raul Barbosa, Marco Araújo
INISTA1
2022 Bi-objective optimization of availability and cost for cloud services
abstract
Cloud-based services are a current approach for developing large-scale applications with advantages such as flexibility, access to on-demand resources, and business agility. The overall application functionality results from complex interactions of many decoupled services, each having its operational specificity. Due to this complexity, the manual configuration of these systems is very arduous, error-prone and likely to impair the quality of service, leading to malfunctioning services, lowering availability and accruing costs. Identifying the optimal solution to simultaneously optimize availability and costs, whilst meeting service level objectives remains a challenge for professionals developing solutions using cloud services. This paper proposes a mathematical formulation of a bi-objective problem to identify the optimal set of solutions for the system configuration. Empirical evaluation of the proposed approach in a case study of a real industrial scenario results in an R-Squared of 0.85, an MSE of 0.021 and an optimization accuracy of 0.928. These methods can help practitioners to keep services at an optimum configuration enabling autonomic service operation, whilst improving availability and cost.
André Bento, João Durães, José Ferreira, Rita Carreira, Filipe Araújo, Raul Barbosa
NCA8
2022 ucXception: A Framework for Evaluating Dependability of Software Systems
abstract
Fault injection is a well-established technique in the research community that consists of emulating faults in order to obtain dependability-related data. Despite its potential, fault injection has been less widely adopted outside of academia, due to the expertise required to effectively conduct fault injection campaigns and to the lack of tools that can be easily adapted to different systems. This paper presents ucXception, an easy-to-install, extendable, open-source framework for orchestrating the entire lifecycle of fault injection campaigns without requiring expert knowledge and using a graphical interface. ucXception supports injection of software and hardware faults using realistic fault models and can be applied to a variety of target systems, including virtualized systems and complex cloud computing deployments. This brings fault injection to modern environments of cloud computing. As a use case, a preliminary analysis on the usage of failure models as a valid alternative to fault models is performed.
Pedro David Almeida, Frederico Cerveira, Raul Barbosa, Henrique Madeira
QRS3
2022 Strategies for Improving the Error Robustness of Convolutional Neural Networks
abstract
The error robustness of Convolutional Neural Networks (CNNs) is an important attribute requiring attention due to their growing application in safety-critical domains such as autonomous driving and medical devices. Hardware errors affecting the execution of such models may lead to system failures and, therefore, fault tolerance techniques are necessary to improve dependability. This paper proposes an approach to improve the robustness of CNNs and experimentally compares it with three other existing techniques. Fault injection is used to emulate hardware faults affecting CNNs targeting four distinct datasets. Results indicate that the ranger technique globally provides the best robustness closely followed by the stimulated training technique, although the former provides much lower temporal overhead than the latter. Architectural redundancy and dropout provide varying results. In all cases, caution through final evaluation of any CNN is required, because there are corner cases in which the robustness decreases, contrary to the intended outcome.
António Morais, Raul Barbosa, Nuno Lourenço 0002, Frederico Cerveira, Michele Lombardi 0001, Henrique Madeira
QRS2
2022 Public Policies Vectors for Urban Greening Technological Strategies
Maria José Sousa, Waleska Campos, Luciana B. da Rosa, Raul Barbosa, M. Carolina Rodrigues, Miguel Sousa, Álvaro Rocha 0001
WorldCIST (3)4
2022 The Effects of Soft Errors and Mitigation Strategies for Virtualization Servers
abstract
Virtualized servers compose the majority of cloud computing environments, where these nodes are used to host multiple clients over the same hardware. Many organizations run online applications by hiring elastic computing resources in order to match demand while reducing fixed costs. However, such organizations are unlikely to take advantage of these benefits for critical applications, as it would expose them to several risks. Among other threats, soft errors are a concern in large-scale reliable servers and are expected to become more frequent as a consequence of smaller transistors and lower operating voltages of integrated circuits. This article characterizes virtualized servers of cloud environments in presence of soft errors. Using fault injection, we collect experimental data to determine the failure modes of applications, operating systems, VMs, and hypervisor. The analysis exposes distinct failure modes, ranging from crash failures of a single virtual machine to silent data corruption in permanent storage. The most frequent failure mode, observed in 10–30 percent of injected errors, consists of a hang affecting multiple virtual machines. Given that such failures are a primary cause of downtime, we develop and evaluate a recovery mechanism which uses online testing and recovers a server from all hangs by rebooting its hypervisor.
Frederico Cerveira, Raul Barbosa, Henrique Madeira, Filipe Araújo
IEEE Trans. Cloud Comput.2
2021 A bag of nodes primer on weightless graph classification
abstract
This paper proposes a weightless architecture for graph classification scenarios.This architecture is a three-headed arrangement composed of graph hand-picked features, a quantization method and a final classifier.Although multiple new strategies for graph classification have been proposed in recent years, it is still necessary to settle comparable studies with respect to weightless neural networks.The proposed architecture is evaluated along with other baseline classifiers and independent strategies, showing that weightless architectures are able to compete with other well-established methods such as graph kernels.* This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior -Brasil (CAPES) -Finance Code 001, CNPq, FAPERJ and DIPPG -CE-FET/RJ.
Raul Barbosa, Diego Carvalho 0001, Priscila M. V. Lima, Felipe M. G. França
ESANN1
2021 μ Viz: Visualization of Microservices
abstract
Microservice architectures have become very popular and widely adopted by the industry, because of the benefits they bring to the software development process and resulting systems, such as parallel development, modularity and scalability. However, as interfaces become more fine-grained and systems grown in size, complexity is moved from the component services to their interactions, eventually leading to intricate workflows that are hard to observe, visualize, and understand. This problem is compounded by the typically high workloads that produce intractable amounts of observation data. To deal with these challenges, operators need support from tools able to take in observation data, in particular tracing, and provide a fast and intuitive understanding of which components or workflows require attention and how are they affecting a module, service, instance, or the whole application. In this paper, we present the design of a microservice visualization application that can fill a gap that exists in leveraging tracing data, aggregating and navigating it in ways that are actionable for operators. Our application provides multiple views of the system and uses spatial and hierarchical navigation using flip zoom to simplify their exploration, while preserving context. Our application can provide a better understanding of the system than existing applications that lack navigability and do not preserve context when switching between different services, layers or views.
Sara Silva, Jaime Correia, André Bento, Filipe Araújo, Raul Barbosa
IV5
2021 A layered framework for root cause diagnosis of microservices
abstract
Microservice-based architectures feature function-ally independent, well-defined and fine-grained components suit-able for loosely coupled deployments and for building reli-able cloud-native applications. Despite the advantages of this approach, component interactions introduce complexity, thus turning boundary -spanning service operation into a daunting challenge. As systems grow in size, complexity can easily outgrow the cognitive capacity of human operators, who are unable to effectively diagnose faulty microservices. We address this problem by proposing a novel framework to diagnose faulty microservices. Through failure injection and an experimental assessment, our layered diagnosis framework using service response analysis, timing constraints, causality and a ranking algorithm from traces, is able to effectively diagnose faulty microservices. Empirical evaluation of the proposed approach, by examining 130 experi-ments in a representative microservice application in the presence of faults, shows that it can achieve approximately 89% specificity and 77% recall.
André Bento, Jaime Correia, João Durães, Luís Ribeiro, Rita Carreira, Filipe Araújo, Raul Barbosa
NCA9
2021 Measuring lead times for failure prediction
abstract
Failure prediction anticipates system failures before they occur so that preemptive action can be taken, thus improving the dependability of the system. For effective failure prediction, the lead time, i.e., the time between the occurrence of a fault and the appearance of a system failure, must accommodate both the prediction step and the preemptive action that is triggered after it. Lead time is intrinsically related to complex error propagation phenomena, which depends on the software architecture of the target system (i.e., the system where failures are predicted) and on the dynamics of such software. For this reason, lead time is highly dependent on the specific nature and intrinsic details of the target system, which means that determining the distribution of lead time for a particular target system should be the very first step in developing failure prediction models. Furthermore, this step is of utmost importance, as it may decide whether failure prediction is viable for a given target system or not. For example, if lead time in a given target system is very short, it means that failure prediction is not viable in such system and classic (and expensive) fault tolerance should be applied. This paper proposes a method for obtaining the lead time distribution of a system using fault injection and presents a practical experiment illustrating such method for a virtualized system. The results suggest that the lead times of failures caused by software faults are usually much larger than those of failures caused by hardware faults.
Frederico Cerveira, Jomar Domingos, Raul Barbosa, Henrique Madeira
PRDC3
2021 End-to-end secure group communication for the Internet of Things
André Lizardo, Raul Barbosa, Samuel Neves, Jaime Correia, Filipe Araújo
J. Inf. Secur. Appl.2
2021 Reductions and abstractions for formal verification of distributed round-based algorithms
Raul Barbosa, Alcides Fonseca, Filipe Araújo
Softw. Qual. J.1
2020 The VALU3S ECSEL Project: Verification and Validation of Automated Systems Safety and Security
abstract
Manufacturers of automated systems and their components have been allocating an enormous amount of time and effort in R&D activities. This effort translates into an overhead on the V&V (verification and validation) process making it time-consuming and costly. In this paper, we present an ECSEL JU project (VALU3S) that aims to evaluate the state-of-the-art V&V methods and tools, and design a multi-domain framework to create a clear structure around the components and elements needed to conduct the V&V process. The main expected benefit of the framework is to reduce time and cost needed to verify and validate automated systems with respect to safety, cyber-security, and privacy requirements. This is done through identification and classification of evaluation methods, tools, environments and concepts for V&V of automated systems with respect to the mentioned requirements. To this end, VALU3S brings together a consortium with partners from 10 different countries, amounting to a mix of 25 industrial partners, 6 leading research institutes, and 10 universities to reach the project goal.
Raul Barbosa, Stylianos Basagiannis, Georgios Giantamidis, H. Becker, Enrico Ferrari, J. Jahic, Alper Kanak, Mikel Labayen, Vanessa Orani, David Pereira, Luigi Pomante, Rupert Schlick, Ales Smrcka, Ahmet Yazici, Peter Folkesson, Behrooz Sangchoolie
DSD1
2020 Evaluation of RESTful frameworks under soft errors
abstract
RESTful frameworks provide a platform for easy deployment, and management of enterprise-level microservices in a scalable and maintainable manner. Like any computer system and its components, RESTful frameworks are susceptible to soft errors, a subset of transient hardware faults that are caused by cosmic rays and package impurities, which can lead to unexpected behaviour from the services that use the frameworks. Failures in these platforms can cause unavailability and unreliability which can lead to major damages including financial or reputation losses to the service providers, and frustration to users who rely on the service provided. Despite soft errors and their impact being a well-studied problem in some fields, such as aeronautics and safety-critical systems, their effect on service frameworks is still uncharacterized. This paper employs fault injection and fuzzing to evaluate how 5 different frameworks behave when affected by soft errors. The obtained results show that using a framework increases the probability of experiencing a failure by an amount that varies from framework to framework and suggest that most failures pose an issue for service availability, which can be relatively easily handled by standard fault tolerance techniques.
Frederico Cerveira, Rui André Oliveira, Raul Barbosa, Henrique Madeira
ISSRE3
2020 Intrusion Detection Systems for Mitigating SQL Injection Attacks: Review and State-of-Practice
abstract
Databases are widely used by organizations to store business-critical information, which makes them one of the most attractive targets for security attacks. SQL Injection is the most common attack to webpages with dynamic content. To mitigate it, organizations use Intrusion Detection Systems (IDS) as part of the security infrastructure, to detect this type of attack. However, the authors observe a gap between the comprehensive state-of-the-art in detecting SQL Injection attacks and the state-of-practice regarding existing tools capable of detecting such attacks. The majority of IDS implementations provide little or no protection against SQL Injection attacks, with exceptions like the tools Bro and ModSecurity. In this article, the authors compare these tools using the CSIC dataset in order to examine the state-of-practice in database protection from SQL Injection attacks, identifying the main characteristics and implementation details needed for IDSs to successfully detect such attacks. The experiments indicate that signature-based IDS provide the greatest coverage against SQL Injection.
Raul Barbosa, Jorge Bernardino
Int. J. Inf. Secur. Priv.2
2019 Spotting Problematic Code Lines using Nonintrusive Programmers' Biofeedback
abstract
Recent studies have shown that programmers' cognitive load during typical code development activities can be assessed using wearable and low intrusive devices that capture peripheral physiological responses driven by the autonomic nervous system. In particular, measures such as heart rate variability (HRV) and pupillography can be acquired by nonintrusive devices and provide accurate indication of programmers' cognitive load and attention level in code related tasks, which are known elements of human error that potentially lead to software faults. This paper presents an experimental study designed to evaluate the possibility of using HRV and pupillography together with eye tracking to identify and annotate specific code lines (or even finer grain lexical tokens) of the program under development (or under inspection) with information on the cognitive load of the programmer while dealing with such lines of code. The experimental data is discussed in the paper to assess different alternatives for using code annotations representing programmers' cognitive load while producing or reading code. In particular, we propose the use of biofeedback code highlighting techniques to provide online programmer's warnings for potentially problematic code lines that may need a second look at (to remove possible bugs), and biofeedback-driven software testing to optimize testing effort, focusing the tests on code areas with higher bug probability.
Ricardo Couceiro, Paulo Carvalho 0001, Miguel Castelo-Branco, Henrique Madeira, Raul Barbosa, João Durães, Gonçalo Duarte, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Nuno Laranjeiro, Júlio Medeiros
ISSRE5
2018 Virtualization: Past and Present Challenges
Frederico Cerveira, Raul Barbosa, Jorge Bernardino
ICSOFT3
2018 Exploratory Data Analysis of Fault Injection Campaigns
abstract
Fault injection (FI) is an experimental methodology used in a wide range of scenarios for validating the fault resilience of applications, especially safety-critical ones. A sufficiently thoroughgoing evaluation produces a significant amount of data regarding the behavior of software components or entire systems in the presence of faults. The core questions that practitioners using fault injection face are 1) how to extract and represent information, 2) how to effectively analyze that data and how to utilize the gained knowledge to improve the FI process. Previous works addressing these questions relied mainly on ad hoc approaches. The current paper presents a modern view of these problems, preparing and executing the knowledge extraction by exploratory (big) data analysis, methods, and tools. A real use-case based on FI campaigns composed of thousands of fault injections into a virtualized system indicates the huge potential of the approach. The outcome is the discovery of an opportunity for a drastic speed-up of the FI process unrevealed by the traditional methodology.
Frederico Cerveira, Imre Kocsis, Raul Barbosa, Henrique Madeira, András Pataricza
QRS3
2018 Weightless neuro-symbolic GPS trajectory classification
Raul Barbosa, Douglas de O. Cardoso, Diego Carvalho 0001, Felipe M. G. França
Neurocomputing1
2018 Language-Based Expression of Reliability and Parallelism for Low-Power Computing
abstract
Improving the energy-efficiency of computing systems while ensuring reliability is a challenge in all domains, ranging from low-power embedded devices to large-scale servers. In this context, a key issue is that many techniques aiming to reduce power consumption negatively affect reliability, while fault tolerance techniques require computation or state redundancy that increases power consumption, thereby leading to systematic tradeoffs. Managing these tradeoffs requires a combination of techniques involving both the hardware and the software, as it is impractical to focus on a single component or level of the system to reach adequate power consumption and reliability. In this paper, we adopt a language-based approach to express reliability and parallelism, in which programs remain adaptable after compilation and may be executed with different strategies concerning reliability and energy consumption. We implement the proposed programming model, which is named MISO, and perform an experimental analysis aiming to improve the reliability of programs, through fault injection experiments conducted at compile-time, as well as an experimental measurement of power consumption. The results obtained indicate that it is feasible to write programs that remain adaptable after compilation in order to improve the ability to balance reliability, power, and performance.
Alcides Fonseca, Frederico Cerveira, Bruno Cabral 0001, Raul Barbosa
IEEE Trans. Sustain. Comput.4
2017 A neuro-symbolic approach to GPS trajectory classification
Diego Carvalho 0001, Felipe M. G. França, Raul Barbosa, Douglas de O. Cardoso
ESANN3
2017 The Ability of Cloud Computing Performance Benchmarks to Measure Dependability
Eduardo Carvalho, Raul Barbosa, Jorge Bernardino
ICSOFT2
2017 Experience Report: On the Impact of Software Faults in the Privileged Virtual Machine
abstract
Cloud computing is revolutionizing how organizations treat computing resources. The privileged virtual machine is a key component in systems that use virtualization, but poses a dependability risk for several reasons. The activation of residual software faults that exist in every software project is a real threat and can impact the correct operation of the entire virtualized system. To study this question, we begin by performing a detailed analysis of the privileged virtual machine and its components, followed by software fault injection campaigns that target two of those important components - toolstack and a device driver. The obstacles faced during this experimental phase and how they were overcome is herein described with practitioners in mind. The results show that software faults in those components can have either no impact or lead to drastic failures, showing that the privileged virtual machine is a single point of failure that must be protected (for 4-9% of the faults). Most of the failures are detectable by monitoring basic functionalities, but some faults caused inconsistent states that manifest later on. No silent data failures (SDF) have been observed, but the number of faults injected so far only allows to conclude that SDF are not very frequent.
Frederico Cerveira, Raul Barbosa, Henrique Madeira
ISSRE2
2017 Soft Errors Susceptibility of Virtualization Servers
abstract
Virtualization is essential in supporting today's information infrastructure, and in particular the Cloud Computing area. However, the move to a virtualized architecture implies the addition of a new single point of failure: the hypervisor. Attempts to characterize and compare the susceptibility of systems (including virtualized systems) are often limited to the study of failure modes and their probabilities. Although undoubtedly useful, in isolation it is not enough to accurately depict the susceptibility of a system, and much less to enable comparison. In this paper, the failure mode analysis of a new and promising virtualization mode of the leading hypervisor in cloud computing deployments (Xen) is performed, followed by the presentation of a general approach to evaluate and compare the susceptibility of systems to soft errors. Exemplifying the approach, a comparison between the susceptibility of three virtualization modes (PVH, HVM and PV), for soft errors in processor registers that occur in a privileged virtual machine (Domain-0), ensues.
Frederico Cerveira, Raul Barbosa, Henrique Madeira
PRDC2
2017 A Probabilistic Analysis of a Leader Election Protocol for Virtual Traffic Lights
abstract
Vehicle-to-vehicle communication systems support diverse cooperative applications such as virtual traffic lights. However, in order to harness the potential benefits of those applications, one must address the challenges faced by distributed algorithms in environments based on unreliable wireless communications. In this paper we address the problem of leader election among the nodes of a cooperative system, under the assumption that the network is unreliable and any number of messages may be lost, and that the number of participating nodes is initially unknown. It is known from the literature that a deterministic solution to this problem does not exist. Nevertheless, one may devise probabilistic solutions that provide arbitrarily high probability of success. In the proposed solution, nodes are enhanced with a local oracle that provides minimal information on the state of the system. Such local oracles, even if unreliable, are shown to increase the probability of success in reaching consensus on the result of the leader election.
Negin Fathollahnejad, Raul Barbosa
PRDC2
2016 Improving self-adaptation planning through software architecture-based stochastic modeling
João Miguel Franco, Francisco Correia, Raul Barbosa, Mário Zenha Rela, Bradley R. Schmerl, David Garlan
J. Syst. Softw.3
2014 Taking an electronic ticketing system to the cloud: Design and discussion
abstract
In this paper we address the challenge of creating an electronic ticketing system for transportation systems that can partially or completely run on the cloud. This challenge is defined within the scope of an industrial project. The resulting system should be able to reach a large spectrum of customers and should provide two key advantages: lower operational costs, especially for small clients without IT departments, and faster execution of queries for monthly or other sorts of analysis, using the elasticity of cloud-based resources. To fulfill the goals of the project, we propose very standard technologies and procedures: a three-tiered architecture; a separation of the online and analysis databases; and an Enterprise Service Bus to get the input from very diverse hardware and software stacks. In this paper we discuss several options regarding the location of these facilities on the cloud and we also evaluate the costs involved. While this work already defines many features of the system, it must be considered as preliminary, as some open details remain for future work.
Filipe Araújo, Marília Curado, Pedro Furtado 0001, Raul Barbosa
IEEE BigData4
2014 CloudBFT: Elastic Byzantine Fault Tolerance
abstract
Cloud computing is increasingly important, with the industry moving towards outsourcing computational resources as a means to reduce investment and management costs, while improving security, dependability and performance. Cloud operators use multi-tenancy, by grouping virtual machines (VMs) into a few physical machines (PMs), to pool computing resources, thus offering elasticity to clients. Although cloud-based fault tolerance schemes impose communication and synchronization overheads, the cloud offers excellent facilities for critical applications, as it can host varying numbers of replicas in independent resources. Given these contradictory forces, determining whether the cloud can host elastic critical services is a major research question. We address this challenge from the perspective of a standard three-tiered system with relational data. We propose to tolerate Byzantine faults using groups of replicas placed on distinct physical machines, as a means to avoid exposing applications to correlated failures. To improve the scalability of our system, we divide data to enable parallel accesses. Using a realistic setup, this setting can reach speedups largely exceeding the number of partitions. Even for a wide variation of the load, the system preserves latency and throughput within reasonable bounds. We believe that the elasticity we observe demonstrates the feasibility of tolerating Byzantine faults in a cloud-based server using a relational database.
Filipe Araújo, Raul Barbosa
PRDC3
2013 Evaluating Xilinx SEU Controller Macro for fault injection
abstract
This paper presents a preliminary evaluation of the SEU Controller Macro, a VHDL component developed by Xilinx for the detection and recovery of single event upsets, as a building block of an FPGA fault-injector. We found that this SEU Controller Macro is extremely effective for injecting faults into the FPGA configuration memory, as single and double bit-flips, with precise location, virtually no intrusiveness, and coarse timing accuracy. We present some clues on how to extend its functionalities to build a fully-fledge FPGA fault injector.
Jose Luis Nunes, João Carlos Cunha, Raul Barbosa, Mário Zenha Rela
DSN3
2013 Probabilistic Analysis of a 1-of-n Selection Algorithm Using a Moderately Pessimistic Decision Criterion
abstract
In this paper we are concerned with the fundamental problem of reaching agreement among a set of distributed processes in presence of an unbounded number of communication failures. We present a probabilistic analysis of a family of synchronous consensus algorithms that aim to solve the 1-of-n selection problem. In this problem, a set of n nodes are to select one common value among a set of n proposed values. There are two possible outcomes of each node's selection process: it can decide either to select a value, or to abort. Agreement implies that all nodes select the same value, or all nodes decide to abort. We know from previous research that it is impossible to guarantee agreement if there is no upper bound on the number of communication failures that can occur. Our aim is to study how the probability of disagreement varies for different decision criteria. The decision criterion consists of the logical expressions that determine whether a process will select a value or decide to abort based on its view of the system state. In this paper we propose and analyse a moderately pessimistic decision criterion. We compared this decision criterion with an optimistic and a pessimistic decision criterion, which we have investigated in our previous work. Our results show that the moderately pessimistic decision criterion for most configurations has a lower maximum probability of disagreement compared with the two other decision criteria. Furthermore, it provides a compromise between the optimistic and the pessimistic approaches since it reduces the probability of disagreement without increasing excessively the probability of agreeing to abort.
Negin Fathollahnejad, Emília Villani, Risat Mahmud Pathan, Raul Barbosa
PRDC4
2012 A Middleware for Exactly-Once Semantics in Request-Response Interactions
abstract
Although the need for the exactly-once request-response interaction pattern is ubiquitous in distributed systems, making it work in practice is anything but simple. Ensuring the at-most-once part of the invocation is relatively easy. Unfortunately, the same is not true for the at-least-once guarantee, which depends on the recovery from crashes of the client, the server and the network. This is what makes the exactly-once interaction so difficult in practice: client and server must log their actions into stable storage, and they must be able to restart the network connections. In this paper, we present a middleware that implements the exactly-once request-response pattern, in presence of network and endpoints crashes. The main contribution of our work is to release the programmer from the complex tasks of recovering from message losses and network crashes.
Naghmeh Ramezani Ivaki, Filipe Araújo, Raul Barbosa
PRDC3
2011 Toward dependability benchmarking of partitioning operating systems
abstract
This paper describes a dependability benchmark intended to evaluate partitioning operating systems. The benchmark includes both hardware and software faultloads and measures the spatial as well as the temporal isolation among tasks, provided by a given real-time operating system. To validate the benchmark, a prototype implementation is carried out and three targets are benchmarked according to the specified process. The results substantiate that the proposed benchmark is able to compare and rank the targets in an objective way, and that it provides the ability to identify aspects of the target systems that need improvement.
Raul Barbosa, Qiu Yu, Xiaozhen Mao
DSN1
2010 GOOFI-2: A tool for experimental dependability assessment
abstract
This paper presents GOOFI-2, a comprehensive fault injection tool for experimental dependability assessment of embedded systems. The tool includes a large number of extensions and improvements over its predecessor, GOOFI. These include support for three widely used fault injection techniques, two target processors, and a variety of new features for storing, disseminating and analyzing experimental data. We report on our experiences and lessons learned from the use and development of GOOFI-2. In particular, we compare and discuss properties of three fault injection techniques: Nexus-based, exception-based and instrumentation-based injection. The comparison relies on several sets of experiments with two target processors, Freescale's MPC565 and MPC5554.
Daniel Skarin, Raul Barbosa
DSN2
2010 Monitoring Local Progress with Watchdog Timers Deduced from Global Properties
abstract
Distributed systems are used in numerous applications where failures can be costly. Due to concerns that some of the nodes may become faulty, critical services are usually replicated across several nodes, which execute distributed algorithms to ensure correct service in spite of failures. To prevent replica-exhaustion, it is fundamental to detect errors and trigger appropriate recovery actions. In particular, it is important to detect situations in which nodes cease to execute the intended algorithm, e.g., when a replica is compromised by an attacker or when a hardware fault causes the node to behave erratically. This paper proposes a method for monitoring the local execution of nodes using watchdog timers. The approach consists in deducing, from the global system properties, local states that must be visited periodically by nodes that execute the intended algorithm correctly. When a node fails to trigger a watchdog before the time limit, an appropriate response can be initiated. The approach is applied to a well-known Byzantine consensus algorithm. The algorithm is modeled in the Promela language and the Spin model checker is used to identify local states that must be visited periodically by correct nodes. Such states are suitable for online monitoring using watchdog timers.
Raul Barbosa
SRDS1
2007 Implementation of a Flexible Membership Protocol on a Real-Time Ethernet Prototype
abstract
This paper describes the implementation of a processor- group membership protocol in an experimental real-time network. The protocol is appropriate for fault-tolerant distributed systems using TDMA for scheduling messages. It allows nodes to maintain a consensus on the operational state of all nodes, in the presence of node failures and restarts, as well as network failures. The protocol is based on the principle that, in a system of n nodes, each node must acknowledge the messages from k other nodes in the membership group, where k can assume values between 2 and n - 1. Membership agreement is guaranteed if f les k - 1 failures occur during n consecutive transmission slots. We have implemented the membership protocol in a time-triggered network based on COTS Ethernet hardware, programmed to schedule messages according to the TDMA method.
Raul Barbosa
PRDC1
2006 Flexible, Cost-EffectiveMembership Agreement in Synchronous Systems
abstract
This paper presents a processor group membership protocol for fault-tolerant distributed real-time systems that utilize periodic, time-triggered scheduling for sending messages over the system's communication network. The protocol allows fault-free nodes to reach agreement on the operational state of all nodes in the presence of fail-silent or fail-reporting node failures as well as network failures (lost or corrupted messages). The protocol is based on the principle that each message sent by a node in the membership is acknowledged by k other nodes in a system of n nodes, where k can be set to any number between 2 and n - 1. Agreement on node failure (membership departure) and agreement on node recovery (membership reintegration) are handled by two different mechanisms. Agreement on departure is guaranteed if no more than f = k - 1 failures occur in the same communication round, while at most one node can be reintegrated into the membership per communication round
Raul Barbosa
PRDC1