VLDB 2026 Research / reviewers in the wild / expert
Frederico Cerveira
dblp:174/2172
· DBLP profile ↗
12ranked-venue papers
6as first author
6since 2021 · last 2023
0000-0002-0180-4815ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 5 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Network Failures in Cloud Management Platforms: A Study on OpenStack
Hassan Mahmood Khan, Frederico Cerveira, Tiago Cruz 0001, Henrique Madeira |
CLOSER | 2 |
| 2022 | TESRAC: A Framework for Test Suite Reduction Assessment at ScaleabstractRegression testing is an important task in any large software project, however as codebase increases, test suites grow and become composed of highly redundant test cases, thus greatly increasing the time required for testing. To solve this problem various test suite reduction tools have been proposed, however their absolute and relative performance are unclear to their prospective users, since there is a lack of a standardized evaluation or approach for choosing the best reduction tool. This work proposes TESRAC, a framework for assessing and comparing test suite reduction tools, which allows users to evaluate and rank a customizable set of tools in terms of reduction performance according to criteria (coverage, dimension, and execution time), and which can be configured to prioritize specific criteria. We used TESRAC to assess and compare three test suite reduction tools and one test suite prioritization tool that has been adapted to perform test suite reduction, across eleven projects of various dimensions and characteristics. Results show that a test suite prioritization tool can be adapted to perform a adequate test suite reduction, and a subset of tools outperforms the remaining tools for the majority of the projects. However, the project and test suite being reduced can have a strong impact on a tool's performance. João Becho, Frederico Cerveira, João Leitão 0001, Rui André Oliveira |
ICST | 2 |
| 2022 | ucXception: A Framework for Evaluating Dependability of Software SystemsabstractFault injection is a well-established technique in the research community that consists of emulating faults in order to obtain dependability-related data. Despite its potential, fault injection has been less widely adopted outside of academia, due to the expertise required to effectively conduct fault injection campaigns and to the lack of tools that can be easily adapted to different systems. This paper presents ucXception, an easy-to-install, extendable, open-source framework for orchestrating the entire lifecycle of fault injection campaigns without requiring expert knowledge and using a graphical interface. ucXception supports injection of software and hardware faults using realistic fault models and can be applied to a variety of target systems, including virtualized systems and complex cloud computing deployments. This brings fault injection to modern environments of cloud computing. As a use case, a preliminary analysis on the usage of failure models as a valid alternative to fault models is performed. Pedro David Almeida, Frederico Cerveira, Raul Barbosa, Henrique Madeira |
QRS | 2 |
| 2022 | Strategies for Improving the Error Robustness of Convolutional Neural NetworksabstractThe error robustness of Convolutional Neural Networks (CNNs) is an important attribute requiring attention due to their growing application in safety-critical domains such as autonomous driving and medical devices. Hardware errors affecting the execution of such models may lead to system failures and, therefore, fault tolerance techniques are necessary to improve dependability. This paper proposes an approach to improve the robustness of CNNs and experimentally compares it with three other existing techniques. Fault injection is used to emulate hardware faults affecting CNNs targeting four distinct datasets. Results indicate that the ranger technique globally provides the best robustness closely followed by the stimulated training technique, although the former provides much lower temporal overhead than the latter. Architectural redundancy and dropout provide varying results. In all cases, caution through final evaluation of any CNN is required, because there are corner cases in which the robustness decreases, contrary to the intended outcome. António Morais, Raul Barbosa, Nuno Lourenço 0002, Frederico Cerveira, Michele Lombardi 0001, Henrique Madeira |
QRS | 4 |
| 2022 | The Effects of Soft Errors and Mitigation Strategies for Virtualization ServersabstractVirtualized servers compose the majority of cloud computing environments, where these nodes are used to host multiple clients over the same hardware. Many organizations run online applications by hiring elastic computing resources in order to match demand while reducing fixed costs. However, such organizations are unlikely to take advantage of these benefits for critical applications, as it would expose them to several risks. Among other threats, soft errors are a concern in large-scale reliable servers and are expected to become more frequent as a consequence of smaller transistors and lower operating voltages of integrated circuits. This article characterizes virtualized servers of cloud environments in presence of soft errors. Using fault injection, we collect experimental data to determine the failure modes of applications, operating systems, VMs, and hypervisor. The analysis exposes distinct failure modes, ranging from crash failures of a single virtual machine to silent data corruption in permanent storage. The most frequent failure mode, observed in 10–30 percent of injected errors, consists of a hang affecting multiple virtual machines. Given that such failures are a primary cause of downtime, we develop and evaluate a recovery mechanism which uses online testing and recovers a server from all hangs by rebooting its hypervisor. Frederico Cerveira, Raul Barbosa, Henrique Madeira, Filipe Araújo |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | Measuring lead times for failure predictionabstractFailure prediction anticipates system failures before they occur so that preemptive action can be taken, thus improving the dependability of the system. For effective failure prediction, the lead time, i.e., the time between the occurrence of a fault and the appearance of a system failure, must accommodate both the prediction step and the preemptive action that is triggered after it. Lead time is intrinsically related to complex error propagation phenomena, which depends on the software architecture of the target system (i.e., the system where failures are predicted) and on the dynamics of such software. For this reason, lead time is highly dependent on the specific nature and intrinsic details of the target system, which means that determining the distribution of lead time for a particular target system should be the very first step in developing failure prediction models. Furthermore, this step is of utmost importance, as it may decide whether failure prediction is viable for a given target system or not. For example, if lead time in a given target system is very short, it means that failure prediction is not viable in such system and classic (and expensive) fault tolerance should be applied. This paper proposes a method for obtaining the lead time distribution of a system using fault injection and presents a practical experiment illustrating such method for a virtualized system. The results suggest that the lead times of failures caused by software faults are usually much larger than those of failures caused by hardware faults. Frederico Cerveira, Jomar Domingos, Raul Barbosa, Henrique Madeira |
PRDC | 1 |
| 2020 | Evaluation of RESTful frameworks under soft errorsabstractRESTful frameworks provide a platform for easy deployment, and management of enterprise-level microservices in a scalable and maintainable manner. Like any computer system and its components, RESTful frameworks are susceptible to soft errors, a subset of transient hardware faults that are caused by cosmic rays and package impurities, which can lead to unexpected behaviour from the services that use the frameworks. Failures in these platforms can cause unavailability and unreliability which can lead to major damages including financial or reputation losses to the service providers, and frustration to users who rely on the service provided. Despite soft errors and their impact being a well-studied problem in some fields, such as aeronautics and safety-critical systems, their effect on service frameworks is still uncharacterized. This paper employs fault injection and fuzzing to evaluate how 5 different frameworks behave when affected by soft errors. The obtained results show that using a framework increases the probability of experiencing a failure by an amount that varies from framework to framework and suggest that most failures pose an issue for service availability, which can be relatively easily handled by standard fault tolerance techniques. Frederico Cerveira, Rui André Oliveira, Raul Barbosa, Henrique Madeira |
ISSRE | 1 |
| 2018 | Virtualization: Past and Present Challenges
Frederico Cerveira, Raul Barbosa, Jorge Bernardino |
ICSOFT | 2 |
| 2018 | Exploratory Data Analysis of Fault Injection CampaignsabstractFault injection (FI) is an experimental methodology used in a wide range of scenarios for validating the fault resilience of applications, especially safety-critical ones. A sufficiently thoroughgoing evaluation produces a significant amount of data regarding the behavior of software components or entire systems in the presence of faults. The core questions that practitioners using fault injection face are 1) how to extract and represent information, 2) how to effectively analyze that data and how to utilize the gained knowledge to improve the FI process. Previous works addressing these questions relied mainly on ad hoc approaches. The current paper presents a modern view of these problems, preparing and executing the knowledge extraction by exploratory (big) data analysis, methods, and tools. A real use-case based on FI campaigns composed of thousands of fault injections into a virtualized system indicates the huge potential of the approach. The outcome is the discovery of an opportunity for a drastic speed-up of the FI process unrevealed by the traditional methodology. Frederico Cerveira, Imre Kocsis, Raul Barbosa, Henrique Madeira, András Pataricza |
QRS | 1 |
| 2018 | Language-Based Expression of Reliability and Parallelism for Low-Power ComputingabstractImproving the energy-efficiency of computing systems while ensuring reliability is a challenge in all domains, ranging from low-power embedded devices to large-scale servers. In this context, a key issue is that many techniques aiming to reduce power consumption negatively affect reliability, while fault tolerance techniques require computation or state redundancy that increases power consumption, thereby leading to systematic tradeoffs. Managing these tradeoffs requires a combination of techniques involving both the hardware and the software, as it is impractical to focus on a single component or level of the system to reach adequate power consumption and reliability. In this paper, we adopt a language-based approach to express reliability and parallelism, in which programs remain adaptable after compilation and may be executed with different strategies concerning reliability and energy consumption. We implement the proposed programming model, which is named MISO, and perform an experimental analysis aiming to improve the reliability of programs, through fault injection experiments conducted at compile-time, as well as an experimental measurement of power consumption. The results obtained indicate that it is feasible to write programs that remain adaptable after compilation in order to improve the ability to balance reliability, power, and performance. Alcides Fonseca, Frederico Cerveira, Bruno Cabral 0001, Raul Barbosa |
IEEE Trans. Sustain. Comput. | 2 |
| 2017 | Experience Report: On the Impact of Software Faults in the Privileged Virtual MachineabstractCloud computing is revolutionizing how organizations treat computing resources. The privileged virtual machine is a key component in systems that use virtualization, but poses a dependability risk for several reasons. The activation of residual software faults that exist in every software project is a real threat and can impact the correct operation of the entire virtualized system. To study this question, we begin by performing a detailed analysis of the privileged virtual machine and its components, followed by software fault injection campaigns that target two of those important components - toolstack and a device driver. The obstacles faced during this experimental phase and how they were overcome is herein described with practitioners in mind. The results show that software faults in those components can have either no impact or lead to drastic failures, showing that the privileged virtual machine is a single point of failure that must be protected (for 4-9% of the faults). Most of the failures are detectable by monitoring basic functionalities, but some faults caused inconsistent states that manifest later on. No silent data failures (SDF) have been observed, but the number of faults injected so far only allows to conclude that SDF are not very frequent. Frederico Cerveira, Raul Barbosa, Henrique Madeira |
ISSRE | 1 |
| 2017 | Soft Errors Susceptibility of Virtualization ServersabstractVirtualization is essential in supporting today's information infrastructure, and in particular the Cloud Computing area. However, the move to a virtualized architecture implies the addition of a new single point of failure: the hypervisor. Attempts to characterize and compare the susceptibility of systems (including virtualized systems) are often limited to the study of failure modes and their probabilities. Although undoubtedly useful, in isolation it is not enough to accurately depict the susceptibility of a system, and much less to enable comparison. In this paper, the failure mode analysis of a new and promising virtualization mode of the leading hypervisor in cloud computing deployments (Xen) is performed, followed by the presentation of a general approach to evaluate and compare the susceptibility of systems to soft errors. Exemplifying the approach, a comparison between the susceptibility of three virtualization modes (PVH, HVM and PV), for soft errors in processor registers that occur in a privileged virtual machine (Domain-0), ensues. Frederico Cerveira, Raul Barbosa, Henrique Madeira |
PRDC | 1 |