Henrique Madeira

dblp:33/2879 · DBLP profile ↗
← Back
116ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-8146-4664ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 53 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 53 · 17 since 2021Systems, architecture and hardware · 27 · 6 first-author · 3 since 2021Databases, data management, data science and information retrieval · 22Artificial intelligence and machine learning · 10 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 From Centralized Learning to Federated Setting: Keeping Reliability on Track
Junjian Yan, Paulo Carvalho 0001, Jorge Henriques, João Loureiro, Chan-Tong Lam, Henrique Madeira
DSN6
2026 Complementarity in software code complexity metrics
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira
J. Syst. Softw.7
2026 A software architecture for verifiable and explainable classification
abstract
Abstract In the context of machine learning, classification is the procedure of predicting the class to which each element of a population belongs to. Most classification functions, for real world problems, are imperfect and thus require rigorous analysis for use in safety-critical applications such as health care. This paper proposes a software architecture for improving the trustworthiness and explainability of AI-based classifiers. The architecture combines a search-based approach with machine-learned explanations and satisfiability solving, to provide an indication of classification confidence and counterfactual explanation rules that are deductively verified to be consistent with the classifier. An implementation of the proposed architecture is evaluated on a medical case study of prognosis of Acute Coronary Syndrome (ACS). The evaluation shows that the proposed architecture is consistently able to complement each individual classification with an indication of confidence and an explanation, which is formally verified for consistency with the classifier. This contributes to foster trustworthy and explainable classification.
Raul Barbosa, Salvatore Rinzivillo, Jacques Robin, Andrea Beretta, Henrique Madeira
Mach. Learn.5
2026 Safety Assessment of UAV Operations in U-Space: A Comprehensive Study on Key Safety Metrics
abstract
Unmanned Aircraft Systems Traffic Management (UTM) and its European version, U-Space, are regulatory frameworks designed to ensure safe, efficient, and secure integration of Unmanned Aerial Vehicles (UAVs) into urban airspace by providing services such as monitoring, conflict resolution, and traffic management. To ensure the safety of UAVs' operations, a comprehensive safety assessment framework is crucial. To build such a framework, it is necessary to identify appropriate safety metrics and develop an approach to measure them, enabling the measurement and management of associated safety risks. In this work, we identify and analyze two categories of safety metrics: collision metrics and surveillance performance metrics. We present an approach grounded in U-space regulatory framework concepts to design and conduct a comprehensive experimental study investigating the impact of several factors that can affect UAV safety, including GPS and IMU failures of varying duration at different UAV speeds, update intervals, traffic densities, and weather conditions, through quantitative assessment of the identified safety metrics. The results reveal key insights into: 1) Identifying the metrics most affected by variations in factors in the presence of GPS or IMU failures, 2) Determination of metrics most correlated to safety risk level under varying conditions, 3) Establishment of risk thresholds for selected metrics under erroneous or varying conditions, contributing to the identification of reliable risk indicators, and 4) Evaluation of the performance and limitations of preventive mechanisms, such as the fail-safe system, under erroneous behavior of GPS and IMU and across different operational and environmental conditions.
Omid Asghari, Naghmeh Ramezani Ivaki, Henrique Madeira
IEEE Trans. Dependable Secur. Comput.3
2026 Enhancing Task In-Progress Time Predictions through Affective and Personality Factors
abstract
Software developers’ personality traits, emotional states, and stress levels are crucial in their task performance. This study aims to enhance the prediction of task in-progress time by integrating traditional features, such as developers’ experience and task estimates, with affective states and personality traits. This article reports a long-term empirical study across seven agile projects in four software development companies, applying various machine learning algorithms to assess the predictive power of these combined features, evaluating them primarily through validation accuracy score. Also, we investigated the impact of weighting developers’ affective states based on their personality traits on model performance. Incorporating developers’ affective states and personality traits improved the task in-progress time prediction in 42.59% of classifier, dataset, and scaling combinations, with oversampled combinations achieving up to 8.4% higher validation accuracy than traditional feature models. The innovative weighting strategy improved 31.48% of the combinations. Our best model achieved a validation accuracy of 0.85. These findings suggest that integrating affective and personality data can significantly improve task in-progress time predictions, with significant implications for project planning and task allocation in software development.
Leo Silva, Cephas A. S. Barreto, Margarida Lima, Henrique Madeira
ACM Trans. Softw. Eng. Methodol.4
2025 NRevisit: A Cognitive Behavioral Metric for Code Understandability Assessment
abstract
Measuring code understandability is both highly relevant and exceptionally challenging. This paper proposes a dynamic code understandability assessment method, which estimates a personalized code understandability score from the perspective of the specific programmer handling the code. The method consists of dynamically dividing the code unit under development or review in code regions (invisible to the programmer) and using the number of revisits (NRevisit) to each region as the primary feature for estimating the code understandability score. This approach removes the uncertainty related to the concept of a "typical programmer" assumed by static software code complexity metrics and can be easily implemented using a simple, low-cost, and non-intrusive desktop eye tracker or even a standard computer camera. This metric was evaluated using cognitive load measured through electroencephalography (EEG) in a controlled experiment with 35 programmers. Results show a very high correlation ranging from rs = 0.9067 to rs = 0.9860 (with p nearly 0) between the scores obtained with different alternatives of NRevisit and the ground truth represented by the EEG measurements of programmers’ cognitive load, demonstrating the effectiveness of our approach in reflecting the cognitive effort required for code comprehension. The paper also discusses possible practical applications of NRevisit, including its use in the context of AI-generated code, which is already widely used today.
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira
EASE7
2025 No Vibe Without Comprehension: Measuring Code Understanding in Modern Coding Workflows Using Neurophysiological Signals
abstract
Code comprehension assessment is crucial in modern software engineering contexts, such as the emerging LLM-supported programming paradigm, where evaluating and adjusting LLM-generated code to ensure suitability, correctness, and readability is mandatory. Recent literature offers various code comprehension solutions, ranging from subjective surveys to neurophysiological-based approaches that are more personalized and operational. However, existing proposals often estimate the cognitive load experienced by programmers during code handling, using this measure as a surrogate for code comprehension. This approach has limitations: it is indirect, as other factors influence cognitive load, and a high cognitive load does not necessarily indicate a lack of code understanding. In this paper, we propose a neurophysiological and AI-based solution using a multimodal set of biosensors, including EEG and eyetracking, along with other contextual features to measure the level of code comprehension. Instead of using cognitive load as a surrogate for code comprehension, this work tackles the challenge of assessing code comprehension by employing performance-annotated ground truth. The solution is customizable, allowing adaptation to different industrial requirements, such as stringent safety and reliability needs in mission-critical software or less critical contexts. We analyze various application scenarios to minimize the intrusiveness of the solution while maintaining acceptable performance. Evaluated in a controlled experiment with 50 programmers and 7 code comprehensions tasks, the porposed solution achieved an accuracy of 69% in the prediction of correct code comprehension. This binary modelling achieve an AUC of 75%, demonstrating its viability for measuring code comprehension in modern software development. We believe that such comprehension assessment methods are essential in current “vibe coding” workflows, where AI tools assist programmers interactively, and code understanding levels must be monitored in real-time to ensure effective human-AI collaboration.
Ricardo Saraiva, João Durães, Paulo Carvalho 0001, Henrique Madeira, Haytham Hijazi
ISSRE4
2024 A Comprehensive Study on Drones Resilience in the Presence of Inertial Measurement Unit Faults
abstract
Unmanned aerial vehicles (UAVs) have gained immense popularity for their versatility and diverse applications. However, this increased usage has raised concerns about the safety and security of UAVs, emphasizing the critical role of their Inertial Measurement Units (IMUs) in ensuring accurate orientation and position data. IMU faults, including both Accelerometer faults and Gyrometer faults, can lead to severe consequences, such as mission failures, collisions, or loss of control. This study addresses the need to enhance UAV resilience in urban airspace by exploring the impact of various IMU faults. A comprehensive fault model is introduced in this paper, covering a range of faults from hardware malfunctions to external attacks. Through extensive fault injection experiments in a simulated environment, the study assesses the effects of different fault types and durations on mission outcomes, providing valuable insights for developing resilient and fault-tolerant UAV systems. Evaluation metrics, including inner and outer bubble violations, missions completed, flight duration, and distance traveled, offer a comprehensive understanding of IMU fault impacts in dynamic operational scenarios. Results reveal that longer injection durations, particularly at 30 seconds, increase bubble violations and significantly reduce mission completion rates. Accelerometer faults, such as “Accelerometer Freeze” and “Accelerometer Random” exhibit reduced mission completion rates of 42.5% and 5%, respectively. Gyrometer faults, especially “Gyrometer Minimum” and “Gyrometer Random” lead to the lowest mission completion rates (2.5%). Additionally, IMU faults (where the fault affects both the Accelerometer and Gyrometer), notably “IMU Minimum”, “IMU Freeze”, and “IMU Random” result in complete mission failures, highlighting the importance of understanding specific fault characteristics. These insights can contribute to developing fault tolerance mechanisms and resilient UAV systems in complex and dynamic environments.
Anamta Khan, Naghmeh Ramezani Ivaki, Henrique Madeira
DSN3
2024 Advancing modern code review effectiveness through human error mechanisms
abstract
Modern code reviews tend to take a lightweight process, in which the accuracy and efficiency of identifying defects rely heavily on code reviewers’ experience. The human errors of developers, as a significant cause of software defects, is a key to identifying defects. However, there is a lack of understanding of the human error mechanisms underlying defects in code. This paper proposes an innovative code review method for identifying defects by pinpointing the scenarios that developers tend to commit errors. The method was validated by two experimental studies that involved 40 participants of about 5 years’ programming experience and modest code review experience. The experiment shows that the proposed method has significantly improved True Positives and Sensitivity by about 400%, improved Precision by approximately 200%, and reduced around one-third of False Positives. The effects were consistent across different tasks and different code reviewers.
Fuqun Huang, Henrique Madeira
J. Syst. Softw.2
2023 Network Failures in Cloud Management Platforms: A Study on OpenStack
Hassan Mahmood Khan, Frederico Cerveira, Tiago Cruz 0001, Henrique Madeira
CLOSER4
2023 Lead Time Analysis for UAVs' Failure Prediction in U-space
abstract
In recent years, UAVs have been increasingly used in urban environments due to agility in movement, simplicity in mechanics, low price, and ability to access locations that are difficult or impossible to reach by humans. A significant number of drones are expected to fly in the urban sky shortly. The profitable nature of commercial UAVs/drone applications in urban space will imply a high density of drones; therefore, avoiding mid-air collisions will be critical for the safe operation of the UAVs. In Europe, U-space services are being created to guarantee the safe operations of UAVs in urban Very Low Level (VLL) airspace. To avoid collisions, U-space considers a separation minima (i.e., the minimum safe distance between UAVs) surrounding each UAV. Thus, violating the separation minima, which might be caused by abnormal conditions (e.g., bad weather conditions), failure conditions (e.g., GPS failure in UAVs), or unreliable behavior of the system (e.g., inaccurate GPS positioning data or erratic position estimation by flight controller), could potentially result in conflicts that require immediate mitigation measures to avoid mid-air collisions. Failure prediction is a promising method for preventing separation minima violations in U-space services. However, in order to have effective failure prediction, the lead time, which is the time between the activation of a fault and its manifestation in a system as a failure, must account for both the prediction step and the subsequent mitigation actions. This paper aims to evaluate the lead time in UAV systems in the presence of positioning-related issues (as being critical for the safe operation of UAVs) from a U-space perspective. We used fault injection to inject 18 different types of faults (or emulating failures) in 28 different UAV missions. The results show that the lead time for 17 types of faults injected is at least 14 seconds (in some cases, no failure occurred). Thus, U-space has at least 14 seconds to predict and mitigate such faults. In the case of GPS failure (i.e., GPS signal is entirely missing), lead time is about 5 seconds, requiring faster strategies for failure prediction and mitigation plans.
Omid Asghari, Naghmeh Ramezani Ivaki, Henrique Madeira
PRDC3
2023 A Machine Learning driven Fault Tolerance Mechanism for UAVs' Flight Controller
abstract
Unmanned Aerial Vehicles (UAVs) are susceptible to various hazards (e.g., software or hardware failures, communication failures, or security attacks) that may hinder mission completion or compromise safety by violating the separation minima (i.e., the minimum distance that must be maintained between UAVs in order to ensure safe and efficient operations). To address this issue, this paper proposes a new machine learning-based fault-tolerant mechanism for UAV flight controllers that tolerates GPS-related faults. These faults are of paramount importance (i.e., accidental faults and/or security attacks that eventually cause failures in the GPS function/data), as accurate positioning and tracking are essential to assure safe operation in UAVs. The proposed machine learning models were built using 884,410 data records from 1,985 flight logs collected from the PX4 public repository. The trained models are used to predict the expected position of the UAV during a mission, and separation minima are used as a threshold to detect the GPS hazards by comparing it with the distance between two consecutive position values. When a hazard is detected (i.e., the distance is higher than separation minima), the predicted values by machine learning models are fed into the flight controller’s position estimator (i.e., an Extended Kalman Filter (EKF)). To evaluate the effectiveness of this approach, validation experiments were conducted on several realistically defined missions while being exposed to different types of failure conditions (e.g., GPS signal loss or GPS Spoofing), both with and without using the proposed fault-tolerant mechanism. The results show a remarkable reduction in safety violations (the number of separation minima violations was reduced from 94 to 1). Additionally, the proposed mechanism demonstrated a notable improvement in the distance traveled by UAV and the duration of the flight mission in failure conditions, showing its ability to mitigate faults effectively. These findings support the effectiveness of the proposed fault tolerance mechanism in enhancing UAV safety in the presence of issues caused by GPS.
Anamta Khan, João R. Campos, Naghmeh Ramezani Ivaki, Henrique Madeira
PRDC4
2023 Quality Evaluation of Modern Code Reviews Through Intelligent Biometric Program Comprehension
abstract
Code review is an essential practice in software engineering to spot code defects in the early stages of software development. Modern code reviews (e.g., acceptance or rejection of pull requests with Git) have become less formal than classic Fagan's inspections, lightweight, and more reliant on individuals (i.e., reviewers). However, reviewers may encounter mentally demanding challenges during the code review, such as code comprehension difficulties or distractions that might affect the code review quality. This work proposes a novel approach that evaluates the quality of code reviews in terms of bug-finding effectiveness and provides the reviewers with a clear message of whether the review should be repeated, indicating the code regions that may not have been well-reviewed. The proposed approach utilizes biometric information collected from the reviewer during the review process using non-intrusive biofeedback devices (e.g., smartwatches). Biometric measures such as Heart Rate Variability (HRV) and task-evoked pupillary response are captured as a surrogate of the cognitive state of the reviewer (e.g., mental workload) and inexpensive desktop eye-trackers compatible with the software development settings. This work uses Artificial Intelligence techniques to predict the cognitive load from the extracted biomarkers and classify each code region according to a set of features. The final evaluation considers various factors such as code complexity, time of the code review, the experience level of the reviewer, and other factors. Our experimental results show the approach could predict the review quality with 87.77%±4.65 accuracy and a Spearman correlation coefficient of 0.85 (p-value < 0.001) between the predicted and the actual review performance. This evaluation validates the cognitive load measurement using electroencephalography (EEG) signals as ground truth for the HRV and pupil signals.
Haytham Hijazi, João Durães, Ricardo Couceiro, João Castelhano, Raul Barbosa, Júlio Medeiros, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
IEEE Trans. Software Eng.9
2022 Are UAVs' Flight Controller Software Reliable?
abstract
Unmanned Ariel Vehicles (UAVs) are recently being studied and worked upon to make them safe and secure for the upcoming expected growth of UAVs in civilian airspace. These efforts resulted in services such as Unmanned Aircraft System Traffic Management (UTM) or U-space in Europe, providing services to regularize and organize (pre-flight), monitor/track (during the flight) drones in civilian airspace while avoiding collisions. The primary source of information for tracking drones during flight is GPS positioning data, which is used and filtered (after being fused with the other sensors' data) by the flight controller software to estimate the vehicle position, velocity, and orientation. Extended Kalman Filter (EKF), which is used in most open-source flight controllers such as PX4, is responsible for doing this estimation. This makes EKF a critical component of the whole system. This paper aims to study the reliability of flight controllers and their core component, namely EKF, in the presence of GPS-related failures. To do so, we injected faults (i.e., we emulated failures indeed) on GPS raw data ranging from small noises to complete failure (missing GPS signals) and GPS spoofing to study their impact on EKF estimation and on the system as a whole. We observed that for small faults (e.g., Fixed Small Noise or Freeze Values), EKF is efficient and can tolerate/compensate the faults, whereas there is a gap in the filter for handling bigger anomalies (e.g., Invalid Values or Random Values) in the GPS data. Our research also clearly demonstrates that GPS faults lasting 30 seconds or more have a noticeable effect, which represents a clear vulnerability since GPS can be subject of cyber attacks such as spoofing. The quantification of the impact of GPS-related failures in the PX4 is an essential step to measure and improve the reliability of UAVs' flight controller software.
Anamta Khan, Naghmeh Ramezani Ivaki, Henrique Madeira
PRDC3
2022 Enhanced software development process for CubeSats to cope with space radiation faults
abstract
CubeSats are an established trend in the space industry. The CubeSat standard opens opportunities for rapid and low-cost access to space. The use of COTS components instead of space-hardened hardware greatly reduces the cost of CubeSat-based missions and provides the additional benefit of increasing software functionalities at a low power consumption. However, COTS components are not designed for the space environment, making CubeSats sensitive to space radiation. This means that CubeSats need additional software mechanisms to guarantee resilient behavior in the presence of space radiation. Our proposal is that such software implemented fault tolerance mechanisms must be tailored to the specific code running in each CubeSat and the logical way to achieve that is to extend the software development process for CubeSats to include the systematic resilience evaluation of software as part of the CubeSats software lifecycle process. This paper proposes a set of structured steps to enhance the classic software development process used in CubeSats, focusing particularly on the Verification and Validation (V&V) phase. The approach uses fault injection as an integral part of the development environment for CubeSats software and includes three major steps: a) sensitivity evaluation (verification) of software in the presence of faults caused by space radiation, b) strengthen of the software with targeted software implemented fault tolerance (SWIFT) mechanisms and c) validation of the effectiveness of the SWIFT mechanisms to confirm that the software is immune to space radiation faults. These added steps to the V&V process must be carried out during software development, as well as every time the CubeSat software has an update, or even a minor change, to ensure that the impact of faults caused by space radiation is tolerated by the CubeSat software. The paper demonstrates the proposed approach using three different embedded software running in the EDC (Environment Data Collection) CubeSat board, which is part (payload) of a constellation of satellites being developed by the Brazilian National Institute for Space Research (INPE). EDC use case provides a realistic insight on the effectiveness of the proposed steps. Our results show that the proposed approach can reduce the percentage of silent data corruption (the most problematic failure mode) from the range of 15% to less than 1% and even to 0% in some embedded software, meaning that the CubeSat software becomes immune to space radiation.
David Paiva, Raffael S. C. G. de Lima, Manoel J. M. Carvalho, Fátima Mattiello-Francisco, Henrique Madeira
PRDC5
2022 ucXception: A Framework for Evaluating Dependability of Software Systems
abstract
Fault injection is a well-established technique in the research community that consists of emulating faults in order to obtain dependability-related data. Despite its potential, fault injection has been less widely adopted outside of academia, due to the expertise required to effectively conduct fault injection campaigns and to the lack of tools that can be easily adapted to different systems. This paper presents ucXception, an easy-to-install, extendable, open-source framework for orchestrating the entire lifecycle of fault injection campaigns without requiring expert knowledge and using a graphical interface. ucXception supports injection of software and hardware faults using realistic fault models and can be applied to a variety of target systems, including virtualized systems and complex cloud computing deployments. This brings fault injection to modern environments of cloud computing. As a use case, a preliminary analysis on the usage of failure models as a valid alternative to fault models is performed.
Pedro David Almeida, Frederico Cerveira, Raul Barbosa, Henrique Madeira
QRS4
2022 A Functional FMECA Approach for the Assessment of Critical Infrastructure Resilience
abstract
The damage or destruction of Critical Infrastructures (CIs) affect societies’ sustainable functioning. Therefore, it is crucial to have effective methods to assess the risk and resilience of CIs. Failure Mode and Effects Analysis (FMEA) and Failure Mode Effects and Criticality Analysis (FMECA) are two approaches to risk assessment and criticality analysis. However, these approaches are complex to apply to intricate CIs and associated Cyber-Physical Systems (CPS). We provide a top-down strategy, starting from a high abstraction level of the system and progressing to cover the functional elements of the infrastructures. This approach develops from FMECA but estimates risks and focuses on assessing resilience. We applied the proposed technique to a real-world CI, predicting how possible improvement scenarios may influence the overall system resilience. The results show the effectiveness of our approach in benchmarking the CI resilience, providing a cost-effective way to evaluate plausible alternatives concerning the improvement of preventive measures.
Gonçalo Carvalho, Nádia Medeiros, Henrique Madeira, Bruno Cabral 0001
QRS3
2022 A New Code Review Method based on Human Errors
abstract
Modern code reviews tend to take a lightweight process, in which the accuracy and efficiency of identifying defects rely heavily on code reviewers’ experience. The human errors of developers, as a significant cause of software defects, is a key to identifying defects. However, there is a lack of understanding of the human error mechanisms underlying defects in code. This paper proposes an innovative code review method for identifying defects by pinpointing the scenarios that developers tend to commit errors. The method was validated by a comprehensive experimental study that involved 49 code reviewers organized in two independent groups, i.e. experimental group vs. controlled group for each other. Forty reviewers have completed the whole experiment and provided the data for statistical analysis on the effects of the approach. The experiment shows that the proposed method has significantly improved True Positives and Sensitivity by about 400%, improved Precision by approximately 200%, and reduced around one-third of False Positives. The effects were consistent across different tasks and different code reviewers.
Fuqun Huang, Henrique Madeira
QRS3
2022 Strategies for Improving the Error Robustness of Convolutional Neural Networks
abstract
The error robustness of Convolutional Neural Networks (CNNs) is an important attribute requiring attention due to their growing application in safety-critical domains such as autonomous driving and medical devices. Hardware errors affecting the execution of such models may lead to system failures and, therefore, fault tolerance techniques are necessary to improve dependability. This paper proposes an approach to improve the robustness of CNNs and experimentally compares it with three other existing techniques. Fault injection is used to emulate hardware faults affecting CNNs targeting four distinct datasets. Results indicate that the ranger technique globally provides the best robustness closely followed by the stimulated training technique, although the former provides much lower temporal overhead than the latter. Architectural redundancy and dropout provide varying results. In all cases, caution through final evaluation of any CNN is required, because there are corner cases in which the robustness decreases, contrary to the intended outcome.
António Morais, Raul Barbosa, Nuno Lourenço 0002, Frederico Cerveira, Michele Lombardi 0001, Henrique Madeira
QRS6
2022 Emotional Dashboard: a Non-Intrusive Approach to Monitor Software Developers' Emotions and Personality Traits
abstract
Developers' emotions are crucial elements that influence the overall job satisfaction of software engineers, including motivation, productivity, and quality of the work, affecting the software development lifecycle. Existing approaches to assess and monitor developers' emotions, such as facial expressions, self-assessed surveys, and biometric sensors, imply considerable intrusiveness on developers' routines and tend to be used only during limited periods. This paper proposes a new non-intrusive and automatable tool (Emotional Dashboard) to assess, monitor, and visualize software developers' emotions during long periods, providing team leaders and project managers with an overview of teams' and software developers' emotional statuses. The idea is to use posts shared by developers on social media to assess their emotions' polarity and visualize the emotional situation on a dashboard, allowing the identification of potentially abnormal emotional periods that may affect the software development. A first evaluation of the tool’s accuracy, done by comparing the emotion polarity (negative, positive, or neutral) of posts done by our tool with the manual classification of a set of posts done by three psychologists, has shown an accuracy of 77%. The tool is available for analysis at this link: https://emotional-dashboard.herokuapp.com.
Leo Silva, Marília Gurgel Castro, Miriam Bernardino Silva, Milena Santos, Uirá Kulesza, Margarida Lima, Henrique Madeira
QRS7
2022 Assessment of the Impact of U-space Faulty Conditions on Drones Conflict Rate
Anamta Khan, Carlos A. Chuquitarco Jiménez, Morcillo-Pallarés Pablo, Naghmeh Ramezani Ivaki, Juan Vicente Balbastre-Tejedor, Henrique Madeira
SAFECOMP6
2022 The Effects of Soft Errors and Mitigation Strategies for Virtualization Servers
abstract
Virtualized servers compose the majority of cloud computing environments, where these nodes are used to host multiple clients over the same hardware. Many organizations run online applications by hiring elastic computing resources in order to match demand while reducing fixed costs. However, such organizations are unlikely to take advantage of these benefits for critical applications, as it would expose them to several risks. Among other threats, soft errors are a concern in large-scale reliable servers and are expected to become more frequent as a consequence of smaller transistors and lower operating voltages of integrated circuits. This article characterizes virtualized servers of cloud environments in presence of soft errors. Using fault injection, we collect experimental data to determine the failure modes of applications, operating systems, VMs, and hypervisor. The analysis exposes distinct failure modes, ranging from crash failures of a single virtual machine to silent data corruption in permanent storage. The most frequent failure mode, observed in 10–30 percent of injected errors, consists of a hang affecting multiple virtual machines. Given that such failures are a primary cause of downtime, we develop and evaluate a recovery mechanism which uses online testing and recovers a server from all hangs by rebooting its hypervisor.
Frederico Cerveira, Raul Barbosa, Henrique Madeira, Filipe Araújo
IEEE Trans. Cloud Comput.3
2021 iReview: an Intelligent Code Review Evaluation Tool using Biofeedback
abstract
Code reviews and software inspections are essential for building reliable software. However, current code reviews practice in the software industry (e.g., acceptance or rejection of pull requests with Git) deviates considerably from classic (and expensive) Fagan's inspections. Modern code reviews are lightweight and asynchronous and do not rely on a group of inspectors and inspection meetings any longer. The modern style of code reviews is much more flexible and cost-effective. Still, these advantages come with the price of reducing the quality of code reviews, as a single reviewer generally makes them with all the inherent and the very human limitations of one single look. The reviewer could be distracted, overloaded, under stress, or not even fully understand the code under review. This paper proposes a new tool (iReview) that evaluates the code review quality using biometric measures gathered from code reviewers (often called Biofeedback). Biometric measures such as Heart Rate Variability (HRV) and eye movement dynamics are used to assess the reviewer's comprehension of the code under review. iReview evaluates the quality of each review globally and indicates the code regions that have not been well-reviewed, explaining why those code regions should be reviewed again. The tool uses Artificial Intelligence techniques to classify the code regions into good and bad reviews based on various biometric and non-biometric features. The first results show that iReview can predict the review quality of medium or complex programs with an accuracy ranging from 75% to 87% in detecting bad reviews (i.e., code regions classified as bad reviewed still have undetected bugs). This tool is expected to improve software reliability by ensuring that good reviews have been carried out despite the current lightweight reviewing processes.
Haytham Hijazi, José Cruz, João Castelhano, Ricardo Couceiro, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
ISSRE7
2021 Measuring lead times for failure prediction
abstract
Failure prediction anticipates system failures before they occur so that preemptive action can be taken, thus improving the dependability of the system. For effective failure prediction, the lead time, i.e., the time between the occurrence of a fault and the appearance of a system failure, must accommodate both the prediction step and the preemptive action that is triggered after it. Lead time is intrinsically related to complex error propagation phenomena, which depends on the software architecture of the target system (i.e., the system where failures are predicted) and on the dynamics of such software. For this reason, lead time is highly dependent on the specific nature and intrinsic details of the target system, which means that determining the distribution of lead time for a particular target system should be the very first step in developing failure prediction models. Furthermore, this step is of utmost importance, as it may decide whether failure prediction is viable for a given target system or not. For example, if lead time in a given target system is very short, it means that failure prediction is not viable in such system and classic (and expensive) fault tolerance should be applied. This paper proposes a method for obtaining the lead time distribution of a system using fault injection and presents a practical experiment illustrating such method for a virtualized system. The results suggest that the lead times of failures caused by software faults are usually much larger than those of failures caused by hardware faults.
Frederico Cerveira, Jomar Domingos, Raul Barbosa, Henrique Madeira
PRDC4
2020 Evaluation of RESTful frameworks under soft errors
abstract
RESTful frameworks provide a platform for easy deployment, and management of enterprise-level microservices in a scalable and maintainable manner. Like any computer system and its components, RESTful frameworks are susceptible to soft errors, a subset of transient hardware faults that are caused by cosmic rays and package impurities, which can lead to unexpected behaviour from the services that use the frameworks. Failures in these platforms can cause unavailability and unreliability which can lead to major damages including financial or reputation losses to the service providers, and frustration to users who rely on the service provided. Despite soft errors and their impact being a well-studied problem in some fields, such as aeronautics and safety-critical systems, their effect on service frameworks is still uncharacterized. This paper employs fault injection and fuzzing to evaluate how 5 different frameworks behave when affected by soft errors. The obtained results show that using a framework increases the probability of experiencing a failure by an amount that varies from framework to framework and suggest that most failures pose an issue for service availability, which can be relatively easily handled by standard fault tolerance techniques.
Frederico Cerveira, Rui André Oliveira, Raul Barbosa, Henrique Madeira
ISSRE4
2019 Pupillography as Indicator of Programmers' Mental Effort and Cognitive Overload
abstract
Our research explores a recent paradigm called Biofeedback Augmented Software Engineering (BASE) that introduces a strong new element in the software development process: the programmers' biofeedback. In this Practical Experience Report we present the results of an experiment to evaluate the possibility of using pupillography to gather biofeedback from the programmers. The idea is to use pupillography to get meta information about the programmers' cognitive and emotional states (stress, attention, mental effort level, cognitive overload,...) during code development to identify conditions that may precipitate programmers making bugs or bugs escaping human attention, and tag the corresponding code locations in the software under development to provide online warnings to the programmer or identify code snippets that will need more intensive testing. The experiments evaluate the use of pupillography as cognitive load predictor, compare the results with the mental effort perceived by programmers using NASATLX, and discuss different possibilities for the use of pupillography as biofeedback sensor in real software development scenarios.
Ricardo Couceiro, Gonçalo Duarte, João Durães, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
DSN9
2019 Spotting Problematic Code Lines using Nonintrusive Programmers' Biofeedback
abstract
Recent studies have shown that programmers' cognitive load during typical code development activities can be assessed using wearable and low intrusive devices that capture peripheral physiological responses driven by the autonomic nervous system. In particular, measures such as heart rate variability (HRV) and pupillography can be acquired by nonintrusive devices and provide accurate indication of programmers' cognitive load and attention level in code related tasks, which are known elements of human error that potentially lead to software faults. This paper presents an experimental study designed to evaluate the possibility of using HRV and pupillography together with eye tracking to identify and annotate specific code lines (or even finer grain lexical tokens) of the program under development (or under inspection) with information on the cognitive load of the programmer while dealing with such lines of code. The experimental data is discussed in the paper to assess different alternatives for using code annotations representing programmers' cognitive load while producing or reading code. In particular, we propose the use of biofeedback code highlighting techniques to provide online programmer's warnings for potentially problematic code lines that may need a second look at (to remove possible bugs), and biofeedback-driven software testing to optimize testing effort, focusing the tests on code areas with higher bug probability.
Ricardo Couceiro, Paulo Carvalho 0001, Miguel Castelo-Branco, Henrique Madeira, Raul Barbosa, João Durães, Gonçalo Duarte, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Nuno Laranjeiro, Júlio Medeiros
ISSRE4
2018 Effects of GPS Spoofing on Unmanned Aerial Vehicles
abstract
Unmanned Aerial Vehicles (UAVs) are no longer exclusively military and scientific solutions. These vehicles have been growing in popularity among hobbyist and also as industrial solutions for specific activities. The flying characteristics and the absence of a crew on board of these devices allow them to perform a wide variety of activities, which can be inaccessible to humans or may threat their life. Despite the advantages, they also bring up major concerns regarding security breaches in the flight controller software, which may lead to security (e.g., vehicle hijacking by attackers), safety (e.g., crashing the vehicle into a planned area or building), or privacy (e.g., eavesdropping or stealing video footage) problems. GPS spoofing is one the main threat of UAVs. The predictability and knowledge of GPS signal properties, create conditions to attackers to assume control of the UAV and use it for their own objectives. In this paper the GPS spoofing effect on UAV is analyzed through a series of tests, under a simulation environment. The results are shown as deviation from the original trajectory and attack success, and analyzed over time and by attack type.
Daniel Mendes, Naghmeh Ramezani Ivaki, Henrique Madeira
PRDC3
2018 Exploratory Data Analysis of Fault Injection Campaigns
abstract
Fault injection (FI) is an experimental methodology used in a wide range of scenarios for validating the fault resilience of applications, especially safety-critical ones. A sufficiently thoroughgoing evaluation produces a significant amount of data regarding the behavior of software components or entire systems in the presence of faults. The core questions that practitioners using fault injection face are 1) how to extract and represent information, 2) how to effectively analyze that data and how to utilize the gained knowledge to improve the FI process. Previous works addressing these questions relied mainly on ad hoc approaches. The current paper presents a modern view of these problems, preparing and executing the knowledge extraction by exploratory (big) data analysis, methods, and tools. A real use-case based on FI campaigns composed of thousands of fault injections into a virtualized system indicates the huge potential of the approach. The outcome is the discovery of an opportunity for a drastic speed-up of the FI process unrevealed by the traditional methodology.
Frederico Cerveira, Imre Kocsis, Raul Barbosa, Henrique Madeira, András Pataricza
QRS4
2017 Experience Report: On the Impact of Software Faults in the Privileged Virtual Machine
abstract
Cloud computing is revolutionizing how organizations treat computing resources. The privileged virtual machine is a key component in systems that use virtualization, but poses a dependability risk for several reasons. The activation of residual software faults that exist in every software project is a real threat and can impact the correct operation of the entire virtualized system. To study this question, we begin by performing a detailed analysis of the privileged virtual machine and its components, followed by software fault injection campaigns that target two of those important components - toolstack and a device driver. The obstacles faced during this experimental phase and how they were overcome is herein described with practitioners in mind. The results show that software faults in those components can have either no impact or lead to drastic failures, showing that the privileged virtual machine is a single point of failure that must be protected (for 4-9% of the faults). Most of the failures are detectable by monitoring basic functionalities, but some faults caused inconsistent states that manifest later on. No silent data failures (SDF) have been observed, but the number of faults injected so far only allows to conclude that SDF are not very frequent.
Frederico Cerveira, Raul Barbosa, Henrique Madeira
ISSRE3
2017 Resilience Benchmarking of Transactional Systems: Experimental Study of Alternative Metrics
abstract
Assessing and comparing computer systems under changing contexts is becoming crucial due to the dynamic characteristics of modern computing environments. This is especially relevant for database management systems, as the behavior of the DBMS when immersed in today's volatile environments is determinant for the success of a multitude of commercial, industrial and scientific endeavors. This paper presents an example of a complete resilience benchmarking scenario for database centric systems and concrete benchmark results, showing that the concept of resilience benchmarking is sound and highly applicable to real transactional systems. The key elements of the approach are discussed, and we define a procedure and a changeload, which includes a set of typical changes that affect the available resources (e.g. memory, CPU) and variations in the type of transactions executed. We then use three distinct quantification approaches to evaluate the resilience of the considered transactional systems concerning their performance, and draw conclusions on the suitableness of the considered metrics for resilience benchmarking.
Raquel Almeida 0002, Afonso Araújo Neto, Henrique Madeira
PRDC3
2017 Soft Errors Susceptibility of Virtualization Servers
abstract
Virtualization is essential in supporting today's information infrastructure, and in particular the Cloud Computing area. However, the move to a virtualized architecture implies the addition of a new single point of failure: the hypervisor. Attempts to characterize and compare the susceptibility of systems (including virtualized systems) are often limited to the study of failure modes and their probabilities. Although undoubtedly useful, in isolation it is not enough to accurately depict the susceptibility of a system, and much less to enable comparison. In this paper, the failure mode analysis of a new and promising virtualization mode of the leading hypervisor in cloud computing deployments (Xen) is performed, followed by the presentation of a general approach to evaluate and compare the susceptibility of systems to soft errors. Exemplifying the approach, a comparison between the susceptibility of three virtualization modes (PVH, HVM and PV), for soft errors in processor registers that occur in a privileged virtual machine (Domain-0), ensues.
Frederico Cerveira, Raul Barbosa, Henrique Madeira
PRDC3
2016 WAP: Understanding the Brain at Software Debugging
abstract
We propose that understanding functional patterns of activity in mapped brain regions associated with code comprehension tasks and, more specifically, to the activity of finding bugs in traditional code inspections could reveal useful insights to improve software reliability and to improve the software development process in general. This includes helping to select the best professionals for the debugging effort, improving the conditions for code inspections, and identify new directions to follow for training code reviewers. This paper presents an interdisciplinary study to analyze the brain activity during code inspection tasks using functional magnetic resonance imaging (fMRI), which is a well-established tool in cognitive neuroscience research. We used several programs where realistic bugs representing the most frequent types of software faults found in the field were injected. The code inspectors involved in the research include programmers with different levels of expertise and experience in real code reviews. The goal is to understand brain activity patterns associated with code comprehension tasks and, more specifically, the brain activity when the code reviewer identifies a bug in the code ('eureka' moment), which can be a true positive or a false positive. Our results confirmed that brain areas associated with language processing and mathematics are highly active during code reviewing and shows that there are specific brain activity patterns that can be related to the decision-making moment of suspicion/bug detection. Importantly, the activity at the anterior insula region that we find to play a relevant role in the process of identifying software bugs is positively correlated to the precision of bug detection by the inspectors. This finding provides a new perspective on the role of this region on error awareness and monitoring and of its potential predictive value in predicting the quality of bug removing.
João Durães, Henrique Madeira, João Castelhano, Isabel Catarina Duarte, Miguel Castelo-Branco
ISSRE2
2015 Temporal Analysis of CHAVE Collection
Olga Craveiro, Joaquim Macedo 0001, Henrique Madeira
SPIRE3
2015 Practical and representative faultloads for large-scale software systems
Pedro Costa 0002, João Gabriel Silva, Henrique Madeira
J. Syst. Softw.3
2015 A benchmarking process to assess software requirements documentation for space applications
Paulo C. Véras, Emília Villani, Ana Maria Ambrosio, Marco Vieira, Henrique Madeira
J. Syst. Softw.5
2014 Query Expansion with Temporal Segmented Texts
Olga Craveiro, Joaquim Macedo 0001, Henrique Madeira
ECIR3
2014 Time-Aware Focused Web Crawling
Joaquim Macedo 0001, Olga Craveiro, Henrique Madeira
ECIR4
2014 Security Benchmarks for Web Serving Systems
abstract
The security of software-based systems is one of the most difficult issues when accessing the suitability of systems to most application scenarios. However, security is very hard to evaluate and quantify, and there are no standard methods to benchmark the security of software systems. This work proposes a novel methodology for benchmarking the security of software-based systems. This methodology uses the notion of risk in a quantifiable way and allows the comparison of functionally-equivalent systems (or different configurations of the same system) to enable users and system integrators to identify and select the most secure one. The benchmark methodology is based on both analytical and experimental steps and can be applicable to any software system. The benchmark procedures and rules guide users on how to instantiate the methodology to specific scenarios and how to execute the benchmark. In this paper we also present an instantiation of the methodology to a case study of web-serving systems and show how to use the results to identify the most secure system under benchmark.
Naaliel Mendes, Henrique Madeira, João Durães
ISSRE2
2014 Analysis of Field Data on Web Security Vulnerabilities
abstract
Most web applications have critical bugs (faults) affecting their security, which makes them vulnerable to attacks by hackers and organized crime. To prevent these security problems from occurring it is of utmost importance to understand the typical software faults. This paper contributes to this body of knowledge by presenting a field study on two of the most widely spread and critical web application vulnerabilities: SQL Injection and XSS. It analyzes the source code of security patches of widely used Web applications written in weak and strong typed languages. Results show that only a small subset of software fault types, affecting a restricted collection of statements, is related to security. To understand how these vulnerabilities are really exploited by hackers, this paper also presents an analysis of the source code of the scripts used to attack them. The outcomes of this study can be used to train software developers and code inspectors in the detection of such faults and are also the foundation for the research of realistic vulnerability and attack injectors that can be used to assess security mechanisms, such as intrusion detection systems, vulnerability scanners, and static code analyzers.
José Fonseca 0002, Nuno Seixas, Marco Vieira, Henrique Madeira
IEEE Trans. Dependable Secur. Comput.4
2014 Evaluation of Web Security Mechanisms Using Vulnerability & Attack Injection
abstract
In this paper we propose a methodology and a prototype tool to evaluate web application security mechanisms. The methodology is based on the idea that injecting realistic vulnerabilities in a web application and attacking them automatically can be used to support the assessment of existing security mechanisms and tools in custom setup scenarios. To provide true to life results, the proposed vulnerability and attack injection methodology relies on the study of a large number of vulnerabilities in real web applications. In addition to the generic methodology, the paper describes the implementation of the Vulnerability & Attack Injector Tool (VAIT) that allows the automation of the entire process. We used this tool to run a set of experiments that demonstrate the feasibility and the effectiveness of the proposed methodology. The experiments include the evaluation of coverage and false positives of an intrusion detection system for SQL Injection attacks and the assessment of the effectiveness of two top commercial web application vulnerability scanners. Results show that the injection of vulnerabilities and attacks is indeed an effective way to evaluate security mechanisms and to point out not only their weaknesses but also ways for their improvement.
José Fonseca 0002, Marco Vieira, Henrique Madeira
IEEE Trans. Dependable Secur. Comput.3
2014 A Technique for Deploying Robust Web Services
abstract
Developing robust web services is a difficult task. Field studies show that a large number of web services are deployed with robustness problems (i.e., presenting unexpected behaviors in the presence of invalid inputs). Although several techniques for the identification of robustness problems have been proposed in the past, there is no practical approach to automatically fix those problems. This paper proposes a mechanism that automatically fixes robustness problems in web services. The approach consists of using robustness testing to detect robustness issues and then mitigate those issues by applying inputs verification based on well-defined parameter domains, including domain dependencies between different parameters. This integrated and fully automated methodology has been used to improve three different implementations of the TPC-App web services and several services publicly available on the Internet. Results show that the proposed approach can be easily used to improve the robustness of web services code.
Nuno Laranjeiro, Marco Vieira, Henrique Madeira
IEEE Trans. Serv. Comput.3
2013 On Fault Representativeness of Software Fault Injection
abstract
The injection of software faults in software components to assess the impact of these faults on other components or on the system as a whole, allowing the evaluation of fault tolerance, is relatively new compared to decades of research on hardware fault injection. This paper presents an extensive experimental study (more than 3.8 million individual experiments in three real systems) to evaluate the representativeness of faults injected by a state-of-the-art approach (G-SWFIT). Results show that a significant share (up to 72 percent) of injected faults cannot be considered representative of residual software faults as they are consistently detected by regression tests, and that the representativeness of injected faults is affected by the fault location within the system, resulting in different distributions of representative/nonrepresentative faults across files and functions. Therefore, we propose a new approach to refine the faultload by removing faults that are not representative of residual software faults. This filtering is essential to assure meaningful results and to reduce the cost (in terms of number of faults) of software fault injection campaigns in complex software. The proposed approach is based on classification algorithms, is fully automatic, and can be used for improving fault representativeness of existing software fault injection approaches.
Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira
IEEE Trans. Software Eng.4
2011 Integrating GQM and Data Warehousing for the Definition of Software Reuse Metrics
abstract
Software reuse is the practice of using existing artifacts (code, architecture, requirements, etc.) in new projects. The advantages of using previously developed software in new projects are easily understood. However, reusing artifacts is usually done in an ad-hoc and incipient way, requiring an important effort of adaptation, so developers frequently prefer to develop components from scratch. In this paper we present a strategy that is being adopted by Critical Software, a medium-sized company, to promote software reuse. This strategy starts by assuming that the success of software reuse is dependent on the ability of measuring its advantages. We have thus proposed the use of the Goal-Question-Metric (GQM) technique, extended with Data Warehousing data model design concepts to extract a set of reuse-specific metrics for measuring the gains of reuse. We show that it is very easy to measure the productivity improvement due to code reuse, by simply measuring or estimating the efforts of developing a component for reuse, integrating it a new artifact, and developing this artifact, built with reusing the component.
Marco Vieira, Henrique Madeira, Sérgio Cruz, Marco Costa 0001, João Carlos Cunha
SEW2
2010 Representativeness analysis of injected software faults in complex software
abstract
Despite of the existence of several techniques for emulating software faults, there are still open issues regarding representativeness of the faults being injected. An important aspect, not considered by existing techniques, is the non-trivial activation condition (trigger) of real faults, which causes them to elude testing and remain hidden until operation. In this paper, we investigate how the representativeness of injected software faults can be improved regarding the representativeness of triggers, by proposing a set of generic criteria to select representative faults from afaultload. We used the G-SWFIT technique to inject software faults in a DBMS, resulting in over 40 thousands faults and 2 million runs of a real test suite. We analyzed faults with respect to their triggers, and concluded that a non-negligible share (15%) would not realistically elude testing. Our proposed criteria decreased the percentage of non-elusive faults in the faultload, improving its representativeness.
Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira
DSN4
2010 Leveraging temporal expressions for segmented-based information retrieval
abstract
The extraction of temporal information from text documents is becoming increasingly important in many applications such as natural language processing, information retrieval, question answering, etc. Indeed, the temporal dimension plays a key role on most of these systems, promoting better performance. Our goal is the definition of a temporal document representation, incorporating the time dimension into information retrieval model to improve the quality of the results. Our approach is based on temporal segmentation of documents. Temporal-aware retrieval models may explore a richer temporal document representation, enabled by segmentation. To achieve this, first we must identify temporal expressions and capture, when possible, their normalized time values. Starting from our prior work on temporal expressions recognition, we present in this paper, a resolution tool that achieves promising results in a Portuguese collection. Furthermore, a temporal characterization of the used collection shows enough and suitable information for a meaningful temporal document segmentation.
Olga Craveiro, Joaquim Macedo 0001, Henrique Madeira
ISDA3
2010 The Web Attacker Perspective - A Field Study
abstract
Web applications are a fundamental pillar of today's globalized world. Society depends and relies on them for business and daily life. However, web applications are under constant attack by hackers that exploit their vulnerabilities to access valuable assets and disrupt business. Many studies and reports on web application security problems analyze the victim's perspective by detailing the vulnerabilities publicly disclosed. In this paper we present a field study on the attacker's perspective by looking at over 300 real exploits used by hackers to attack web applications. Results show that SQL injection and Remote File Inclusion are the two most frequently used exploits and that hackers prefer easier rather than complicated attack techniques. Exploit and vulnerability data are also correlated to show that, although there are many types of vulnerabilities out there, only few are interesting enough for attackers to obtain what they want the most: root shell access and admin passwords.
José Fonseca 0002, Marco Vieira, Henrique Madeira
ISSRE3
2010 Errors on Space Software Requirements: A Field Study and Application Scenarios
abstract
This paper presents a field study on real errors found in space software requirements documents. The goal is to understand and characterize the most frequent types of requirement problems in this critical application domain. To classify the software requirement errors analyzed we initially used a well-known existing taxonomy that was later extended in order to allow a more thorough analysis. The results of the study show a high rate of requirement errors (9.5 errors per each 100 requirements), which is surprising if we consider that the focus of the work is critical embedded software. Besides the characterization of the most frequent types of errors, the paper also proposes a set of operators that define how to inject realistic errors in requirement documents. This may be used in several scenarios, including: evaluating and training reviewers, estimating the number of requirement errors in real specifications, defining checklists for quick requirement verification, and defining benchmarks for requirements specifications.
Paulo C. Véras, Emília Villani, Ana Maria Ambrosio, Marco Vieira, Henrique Madeira
ISSRE6
2010 Towards Identifying the Best Variables for Failure Prediction Using Injection of Realistic Software Faults
abstract
Predicting failures at runtime is one of the most promising techniques to increase the availability of computer systems. However, failure prediction algorithms are still far from providing satisfactory results. In particular, the identification of the variables that show symptoms of incoming failures is a difficult problem. In this paper we propose an approach for identifying the most adequate variables for failure prediction. Realistic software faults are injected to accelerate the occurrence of system failures and thus generate a large amount of failure related data that is used to select, among hundreds of system variables, a small set that exhibits a clear correlation with failures. The proposed approach was experimentally evaluated using two configurations based on Windows XP. Results show that the proposed approach is quite effective and easy to use and that the injection of software faults is a powerful tool for improving the state of the art on failure prediction.
Ivano Irrera, João Durães, Marco Vieira, Henrique Madeira
PRDC4
2010 A Learning-Based Approach to Secure Web Services from SQL/XPath Injection Attacks
abstract
Business critical applications are increasingly being deployed as web services that access database systems, and must provide secure operations to its clients. Although the open web environment emphasizes the need for security, several studies show that web services are still being deployed with command injection vulnerabilities. This paper proposes a learning-based approach to secure web services against SQL and XPath Injection attacks. Our approach is able to transparently learn valid request patterns (learning phase) and then detect and abort potentially harmful requests (protection phase). When it is not possible to have a complete learning phase, a set of heuristics can be used to accept/discard doubtful cases. Our mechanism was applied to secure TPC-App services and open source services. It showed to be extremely effective in stopping all tested attacks, while introducing a negligible performance impact.
Nuno Laranjeiro, Marco Vieira, Henrique Madeira
PRDC3
2010 Benchmarking Software Requirements Documentation for Space Application
Paulo C. Véras, Emília Villani, Ana Maria Ambrosio, Rodrigo Pastl Pontes, Marco Vieira, Henrique Madeira
SAFECOMP6
2010 Benchmarking the Resilience of Self-Adaptive Systems: A New Research Challenge
abstract
Self-adaptive systems are widely recognized as the future of computer systems. Due to their dynamic and evolving nature, the characterization of self-adaptation and resilience attributes is of upmost importance. The problem is that nowadays there is no practical way to characterize self-adaptation capabilities or to compare alternative solutions concerning resilience. In this paper we discuss the problem of resilience benchmarking of self-adaptive systems. We start by identifying a set of key challenges and then propose a research roadmap to tackle those challenges.
Raquel Almeida 0002, Henrique Madeira, Marco Vieira
SRDS2
2009 Protecting Database Centric Web Services against SQL/XPath Injection Attacks
Nuno Laranjeiro, Marco Vieira, Henrique Madeira
DEXA3
2009 Vulnerability & attack injection for web applications
abstract
In this paper we propose a methodology to inject realistic attacks in Web applications. The methodology is based on the idea that by injecting realistic vulnerabilities in a Web application and attacking them automatically we can assess existing security mechanisms. To provide true to life results, this methodology relies on field studies of a large number of vulnerabilities in Web applications. The paper also describes a set of tools implementing the proposed methodology. They allow the automation of the entire process, including gathering results and analysis. We used these tools to conduct a set of experiments to demonstrate the feasibility and effectiveness of the proposed methodology. The experiments include the evaluation of coverage and false positives of an intrusion detection system for SQL injection and the assessment of the effectiveness of two Web application vulnerability scanners. Results show that the injection of vulnerabilities and attacks is an effective way to evaluate security mechanisms and tools.
José Fonseca 0002, Marco Vieira, Henrique Madeira
DSN3
2009 From assessment to standardised benchmarking: Will it happen? What could we do about it?
abstract
Cost pressure, short time to market, and increased complexity are responsible for an evident increase of the failure rate of computing systems, while the cost of failures is growing rapidly, as a result of an unprecedented degree of dependence of our society on computing systems. The combination of these factors has created a dependability and security gap that is often perceived by users as a lack of trustworthiness in computer applications, and that is in fact undermining the network and service infrastructures that constitute the very core of the knowledge-based society.
Henrique Madeira, István Majzik
DSN1
2009 Using web security scanners to detect vulnerabilities in web services
abstract
Although Web services are becoming business-critical components, they are often deployed with critical software bugs that can be maliciously explored. Web vulnerability scanners allow detecting security vulnerabilities in Web services by stressing the service from the point of view of an attacker. However, research and practice show that different scanners have different performance on vulnerabilities detection. In this paper we present an experimental evaluation of security vulnerabilities in 300 publicly available Web services. Four well known vulnerability scanners have been used to identify security flaws in Web services implementations. A large number of vulnerabilities has been observed, which confirms that many services are deployed without proper security testing. Additionally, the differences in the vulnerabilities detected and the high number of false-positives (35% and 40% in two cases) and low coverage (less than 20% for two of the scanners) observed highlight the limitations of Web vulnerability scanners on detecting security vulnerabilities in Web services.
Marco Vieira, Nuno Antunes, Henrique Madeira
DSN3
2009 Improving Web Services Robustness
abstract
Developing robust web services is a difficult task. Field studies show that a large number of web services are deployed with robustness problems (i.e., presenting unexpected behaviors in the presence of invalid inputs). Several techniques for the identification of robustness problems have been proposed in the past. This paper proposes a mechanism that automatically fixes the problems detected. The approach consists of using robustness testing to detect robustness issues and then mitigate those issues by applying inputs verification based on well-defined parameter domains, including domain dependencies between different parameters. This integrated and fully automatable methodology has been used to improve three different implementations of the TPC-App web services. Results show that this tool can be easily used by developers to improve the robustness of web services implementations.
Nuno Laranjeiro, Marco Vieira, Henrique Madeira
ICWS3
2009 Looking at Web Security Vulnerabilities from the Programming Language Perspective: A Field Study
abstract
This paper presents a field study on Web security vulnerabilities from the programming language type system perspective. Security patches reported for a set of 11 widely used Web applications written in strongly typed languages (Java, C#, VB.NET) were analyzed in order to understand the fault types that are responsible for the vulnerabilities observed (SQL injection and XSS). The results are analyzed and compared with a similar work on Web applications written using a weakly typed language (PHP). This comparison points out that some of the types of defects that lead to vulnerabilities are programming language independent, while others are strongly related to the language used. Strongly typed languages do reduce the frequency of vulnerabilities, as expected, but there still is a considerable number of vulnerabilities observed in the field. The characterization of those vulnerabilities shows that they are caused by a small number of fault types. This result is relevant to train programmers and code inspectors in the manual detection of such faults, and to improve static code analyzers to automatically detect the most frequent vulnerable program structures found in the field.
Nuno Seixas, José Fonseca 0002, Marco Vieira, Henrique Madeira
ISSRE4
2009 Dependability Benchmarking Using Software Faults: How to Create Practical and Representative Faultloads
abstract
The faultload is one of the most critical components of a dependability benchmark. It should embody a repeatable, portable, representative and generally accepted fault set. Concerning software faults, the definition of that kind of faultloads is particularly difficult, as it requires a much more complex emulation method than the traditional stuck-at or bit-flip used for hardware faults. Although faultloads based on software faults have already been proposed, the choice of adequate fault injection targets (i.e., actual software components where the faults are injected) is still an open and crucial issue. Furthermore, knowing that the number of possible software faults that can be injected in a given system is potentially very large, the problem of defining a faultload made of a small number of representative faults is of utmost importance. This paper proposes a strategy to guide the fault injection target selection and reduce the number of faults required for the faultload and exemplifies the proposed approach with a real Web-server dependability benchmark and a large-scale integer vector sort application.
Pedro Costa 0002, João Gabriel Silva, Henrique Madeira
PRDC3
2009 Use of Co-occurrences for Temporal Expressions Annotation
Olga Craveiro, Joaquim Macedo 0001, Henrique Madeira
SPIRE3
2008 Timing Failures Detection in Web Services
abstract
Current business critical environments increasingly rely on SOA standards to execute business operations. These operations are frequently based on Web service compositions that use several Web services over the internet and have to fulfill specific timing constraints. In these environments, an operation that does not conclude in due time may have a high cost as it can easily turn into service abandonment with financial and prestige losses to the service provider. In fact, at certain points, carrying on with the execution of an operation may be useless as a timely response will be impossible to obtain. This paper proposes a time-aware programming model for Web services that provides transparent timing failure detection. The paper illustrates the proposed model using a set of services specified by the TPC-App performance benchmark.
Nuno Laranjeiro, Marco Vieira, Henrique Madeira
APSCC3
2008 Redundant Array of Inexpensive Nodes for DWS
Jorge Vieira, Marco Vieira, Marco Costa 0001, Henrique Madeira
DASFAA4
2008 RAIN: Always on Data Warehousing
Jorge Vieira, Marco Vieira, Marco Costa 0001, Henrique Madeira
DASFAA4
2008 Efficient Data Distribution for DWS
Raquel Almeida 0002, Jorge Vieira, Marco Vieira, Henrique Madeira, Jorge Bernardino
DaWaK4
2008 Message from the conference general chair and coordinator
abstract
Presents the introductory welcome message from the conference proceedings.
Philip Koopman, Henrique Madeira
DSN2
2008 Training Security Assurance Teams Using Vulnerability Injection
abstract
Writing secure web applications is a complex task. In fact, a vast majority of web applications are likely to have security vulnerabilities that can be exploited using simple tools like a common web browser. This represents a great danger as the attacks may have disastrous consequences to organizations, harming their assets and reputation. To mitigate these vulnerabilities, security code inspections and penetration tests must be conducted by well-trained teams during the development of the application. However, effective code inspections and testing takes time and cost a lot of money, even before any business revenue. Furthermore, software quality assurance teams typically lack the knowledge required to effectively detect security problems. In this paper we propose an approach to quickly and effectively train security assurance teams in the context of web application development. The approach combines a novel Vulnerability Injection Technique with relevant guidance information about the most common security vulnerabilities to provide a realistic training scenario. Our experimental results show that a short training period is sufficient to clearly improve the ability of security assurance teams to detect vulnerabilities during both code inspections and penetration tests.
José Fonseca 0002, Marco Vieira, Henrique Madeira
PRDC3
2008 Assessing and Comparing Security of Web Servers
abstract
This paper presents an approach to assess security of web servers. This method can be used to compare the security features of different web servers installations and to determine how secure a given web server configuration is. The assessment is done by applying a set of tests designed to check if the system under evaluation fulfils a set of security practices defined by an extensive field study. This work targets the most typical issues related to web servers ranging from classic web servers misconfiguration to the absence of a secure network infrastructure and of well-defined security policies to respond to security incidents. The effectiveness and usefulness of the proposed approach is illustrated through the security assessment and comparison of five different real web servers.
Naaliel Mendes, Afonso Araújo Neto, João Durães, Marco Vieira, Henrique Madeira
PRDC5
2007 Towards Timely ACID Transactions in DBMS
Marco Vieira, António Casimiro, Henrique Madeira
DASFAA3
2007 Experimental Risk Assessment and Comparison Using Software Fault Injection
abstract
One important question in component-based software development is how to estimate the risk of using COTS components, as the components may have hidden faults and no source code available. This question is particularly relevant in scenarios where it is necessary to choose the most reliable COTS when several alternative components of equivalent functionality are available. This paper proposes a practical approach to assess the risk of using a given software component (COTS or non-COTS). Although we focus on comparing components, the methodology can be useful to assess the risk in individual modules. The proposed approach uses the injection of realistic software faults to assess the impact of possible component failures and uses software complexity metrics to estimate the probability of residual defects in software components. The proposed approach is demonstrated and evaluated in a comparison scenario using two real off-the-shelf components (the RTEMS and the RTLinux real-time operating system) in a realistic application of a satellite data handling application used by the European Space Agency.
Regina Lúcia de Oliveira Moraes, João Durães, Ricardo Barbosa 0003, Eliane Martins, Henrique Madeira
DSN5
2007 Assessing Robustness of Web-Services Infrastructures
abstract
Web-services are supported by a complex software infrastructure that must provide a robust service to the client applications. This practical experience report presents a practical approach for the evaluation of the robustness of Web-services infrastructures. A set of robustness tests (i.e., invalid web-services call parameters) is applied during Web-services execution in order to reveal possible robustness problems in the Web-services code and in the application server infrastructure. The approach is illustrated using two different implementations of the Web-services specified by the TPC-App performance benchmark running on top of the JBoss application server. The proposed approach is generic and can be used to evaluate the robustness of Web-services implementations (relevant for programmers) and application server infrastructures (relevant for administrators and system integrators).
Marco Vieira, Nuno Laranjeiro, Henrique Madeira
DSN3
2007 Testing and Comparing Web Vulnerability Scanning Tools for SQL Injection and XSS Attacks
abstract
Web applications are typically developed with hard time constraints and are often deployed with security vulnerabilities. Automatic web vulnerability scanners can help to locate these vulnerabilities and are popular tools among developers of web applications. Their purpose is to stress the application from the attacker's point of view by issuing a huge amount of interaction within it. Two of the most widely spread and dangerous vulnerabilities in web applications are SQL injection and cross site scripting (XSS), because of the damage they may cause to the victim business. Trusting the results of web vulnerability scanning tools is of utmost importance. Without a clear idea on the coverage and false positive rate of these tools, it is difficult to judge the relevance of the results they provide. Furthermore, it is difficult, if not impossible, to compare key figures of merit of web vulnerability scanners. In this paper we propose a method to evaluate and benchmark automatic web vulnerability scanners using software fault injection techniques. The most common types of software faults are injected in the web application code which is then checked by the scanners. The results are compared by analyzing coverage of vulnerability detection and false positives. Three leading commercial scanning tools are evaluated and the results show that in general the coverage is low and the percentage of false positives is very high.
José Fonseca 0002, Marco Vieira, Henrique Madeira
PRDC3
2007 Benchmarking the Robustness of Web Services
abstract
This paper proposes an approach for the evaluation of the robustness of web services, which are complex software components that must provide a robust interface to the client applications. However, although web services are becoming business-critical components, there is no practical way to assess the robustness of the code or to compare alternative implementations concerning robustness. The approach proposed is based on a set of robustness tests (i.e., invalid web services call parameters) that is applied in order to discover both programming and design errors. The web services are classified based on the failures observed during the execution of the tests. The approach is illustrated by evaluating several web services publicly available in the Internet and two different implementations of the web services specified by the standard TPC-App performance benchmark. The proposed approach is useful for both web services providers (to assess the robustness of their web services code) and consumers (to select the web services that best fit their requirements).
Marco Vieira, Nuno Laranjeiro, Henrique Madeira
PRDC3
2007 Detecting Malicious SQL
José Fonseca 0002, Marco Vieira, Henrique Madeira
TrustBus3
2006 Software Aging and Rejuvenation in a SOAP-based Server
abstract
Web-services and service-oriented architectures are gaining momentum in the area of distributed systems and Internet applications. However, as we increase the abstraction level of the applications we are also increasing the complexity of the underlying middleware. In this paper, we present a dependability benchmarking study to evaluate and compare the robustness of some of the most popular SOAP-RPC implementations that are intensively used in the industry. The study was focused on Apache Axis where we have observed a high susceptibility of software aging. Building on these results we propose a new SLA-oriented software rejuvenation technique that proved to be a simple way to increase the dependability of the SOAP-server, the degree of self-healing and to maintain a sustained level of performance in the applications
Luís Moura Silva, Henrique Madeira, João Gabriel Silva
NCA2
2006 Monitoring Database Application Behavior for Intrusion Detection
abstract
Database management systems (DBMS) represent the ultimate layer in preventing malicious data access or corruption and implement several security mechanisms to protect data. However these mechanisms cannot always stop malicious users from accessing data by exploiting system vulnerabilities. The aim of this paper is to propose an intrusion detection mechanism for DBMS to fill this gap. Our approach consists of a comprehensive representation of user database utilization profiles to perform concurrent intrusion detection. Prior to the detection it is necessary to define and learn these utilization profiles. Profiles are defined using a three level abstraction and learned directly from monitoring the database utilization in real conditions. The proposed mechanism is generic and can be easily implemented in commercial and open-source DBMS
José Fonseca 0002, Marco Vieira, Henrique Madeira
PRDC3
2006 Towards Timely ACID Transactions in DBMS
abstract
On time data management is becoming a key difficulty faced by organizations. In spite of the importance of timeliness requirements in database applications, commercial DBMS do not assure the detection of the cases when a transaction takes longer than the expected/desired time. This paper discusses the problem of timing failure detection in database applications and proposes a transaction programming approach to help developers in programming database applications with time constraints
Marco Vieira, António Casimiro, Henrique Madeira
PRDC3
2006 Emulation of Software Faults: A Field Data Study and a Practical Approach
abstract
The injection of faults has been widely used to evaluate fault tolerance mechanisms and to assess the impact of faults in computer systems. However, the injection of software faults is not as well understood as other classes of faults (e.g., hardware faults). In this paper, we analyze how software faults can be injected (emulated) in a source-code independent manner. We specifically address important emulation requirements such as fault representativeness and emulation accuracy. We start with the analysis of an extensive collection of real software faults. We observed that a large percentage of faults falls into well-defined classes and can be characterized in a very precise way, allowing accurate emulation of software faults through a small set of emulation operators. A new software fault injection technique (G-SWFIT) based on emulation operators derived from the field study is proposed. This technique consists of finding key programming structures at the machine code-level where high-level software faults can be emulated. The fault-emulation accuracy of this technique is shown. This work also includes a study on the key aspects that may impact the technique accuracy. The portability of the technique is also discussed and it is shown that a high degree of portability can be achieved
João Durães, Henrique Madeira
IEEE Trans. Software Eng.2
2005 Efficient Compression of Text Attributes of Data Warehouse Dimensions
Jorge Vieira, Jorge Bernardino, Henrique Madeira
DaWaK3
2005 Dependability Benchmarking of Computing Systems - Panel Statement
abstract
The importance of benchmarking is increasing as every aspect of human life is relying on correct operation of computing systems. Although considerable efforts have been made, presently there are no widely accepted dependability benchmarks. The panelists, in the order shown above, address benchmarking of computer hardware, operating systems, applications, and systems.
Cristian Constantinescu, Karama Kanoun, Henrique Madeira, Brendan Murphy, Ira Pramanick, Aaron B. Brown
DSN3
2005 Towards a Security Benchmark for Database Management Systems
abstract
One of the main problems faced by organizations is the protection of their data against unauthorized access or corruption due to malicious actions. Database management systems (DBMS) constitute the kernel of the information systems used today to support the daily operations of most organizations and represent the ultimate layer in preventing unauthorized access to data stored in information systems. Nevertheless, in spite of the key role played by the DBMS in the overall data security, no practical way has been proposed so far to characterize the security in such systems or to compare alternative solutions concerning security features. This paper proposes an approach to characterize the security mechanisms in database systems and database applications, according to a set of security classes. The proposed approach is generic and can be applied to both DBMS (relevant for system integrators) and real database installations (relevant for database administrators and end-users).
Marco Vieira, Henrique Madeira
DSN2
2005 Detection of Malicious Transactions in DBMS
abstract
A major difficulty faced by organizations is the protection of data against malicious access or corruption. Database management systems (DBMS) are a key component in the information infrastructure of most organizations and represent the ultimate layer in preventing unauthorized data accesses. Several mechanisms needed to protect data, such as authentication, user privileges, encryption, and auditing, have been implemented in commercial DBMS. However, typical database security mechanisms are not able to detect and handle many data security attacks. In fact, malicious transactions executed by unauthorized users that may gain access to the database by exploring system vulnerabilities and unauthorized database transactions executed by authorized users cannot be detected and stopped by typical security mechanisms. In this paper we propose a new mechanism for the detection of malicious transactions in DBMS. The paper presents a practical example of the implementation of the proposed mechanism in the Oracle 10g DBMS and evaluates the mechanism using the TPC-C benchmark.
Marco Vieira, Henrique Madeira
PRDC2
2004 Portable Faultloads Based on Operator Faults for DBMS Dependability Benchmarking
abstract
Databases play a central role in the information infrastructure of most organizations. The characterization of DBMS (database management systems) dependability is then of utmost importance. Existing performance benchmarks for transactional and database areas include two major components: a workload and a set of performance measures. The definition of a benchmark to characterize dependability needs a new component - the faultload. Operator faults represent a major cause of failures in large DBMS. This paper proposes three approaches for the definition of portable faultloads based on operator faults to benchmark the dependability of DBMS and shows a benchmarking example of a commercial (Oracle) and an open source (PostgreSQL) database.
Marco Vieira, Henrique Madeira
COMPSAC2
2004 Handling big dimensions in distributed data warehouses using the DWS technique
abstract
The DWS (Data Warehouse Striping) technique allows the distribution of large data warehouses through a cluster of computers. The data partitioning approach partition the facts tables through all nodes and replicates the dimension tables. The replication of the dimension tables creates a limitation to the applicability of the DWS technique to data warehouses with big dimensions. This paper proposes a strategy to handle large dimensions in a distributed DWS system and evaluates the proposed strategy experimentally. With the proposed strategy the performance speed up and scale up obtained in the DWS technique are not affected by the presence of big dimensions. Furthermore, it extends the scope of the technique to queries that browse big dimensions that can also benefit of the performance increase of the DWS technique.
Marco Costa 0001, Henrique Madeira
DOLAP2
2004 Generic Faultloads Based on Software Faults for Dependability Benchmarking
abstract
The most critical component of a dependability benchmark is the faultload, as it should represent a repeatable, portable, representative, and generally accepted set of faults. These properties are essential to achieve the desired standardization level required by a dependability benchmark but, unfortunately, are very hard to achieve. This is particularly true for software faults, which surely accounts for the fact that this important class of faults has never been used in known dependability benchmark proposals. This paper proposes a new methodology for the definition of faultloads based on software faults for dependability benchmarking. Faultload properties such as repeatability, portability and scalability are also analyzed and validated through experimentation using a case study of dependability benchmarking of Web-servers. We concluded that software fault-based faultloads generated using our methodology are appropriate and useful for dependability benchmarking. As our methodology is not tied to any specific software vendor or platform, it can be used to generate faultloads for the evaluation of any software product such as OLTP systems.
João Durães, Henrique Madeira
DSN2
2004 Dependability Benchmarking of Web-Servers
João Durães, Marco Vieira, Henrique Madeira
SAFECOMP3
2004 Joint evaluation of recovery and performance of a COTS DBMS in the presence of operator faults
Marco Vieira, Henrique Madeira
Perform. Evaluation2
2003 Definition of Software Fault Emulation Operators: A Field Data Study
abstract
This paper proposes a set of operators for software fault emulation through low-level code mutations. The definition of these operators was based on the analysis of an extensive collection of real software faults. Using the Orthogonal Defect Classification as a starting point, faults were classified in a detailed manner according to the high-level constructs where the faults reside and their effects in the program. We observed that a large percentage of faults fall in well-defined classes and can be characterized in a very precise way, allowing accurate emulation through a small set of mutation operators. The resulting operators closely emulate a broad range of common programmer mistakes. Furthermore, as the mutation is performed directly at the executable code, software faults can be injected in targets for which source code is not available.
João Durães, Henrique Madeira
DSN2
2003 The OLAP and Data Warehousing Approaches for Analysis and Sharing of Results from Dependability Evaluation Experiments
abstract
Two important questions on experimental dependability evaluation remain largely unanswered: 1) how to analyze the usually large amount of raw data produced in dependability evaluation experiments and 2) how to compare results from different experiments or results from similar experiments across different systems. These problems are also common to other dependability evaluation techniques such as the ones based on simulation, or even to the analysis of field data on computer faults. We propose the use of data warehousing technologies to store raw results from different experiments/setups in a common multidimensional structure where raw data can be analyzed and shared world wide by means of web-enabled OLAP (On-Line Analytical Processing) tools. This paper describes how to use the proposed approach in a concrete example of dependability evaluation experiment.
Henrique Madeira, João Pedro Costa, Marco Vieira
DSN1
2003 Benchmarking the Dependability of Different OLTP Systems
abstract
On-Line Transaction Processing (OLTP) systems constitute the kernel of the information systems used today to support the daily operations of most organizations. Although these systems comprise the best examples of complex business-critical systems, no practical way has been proposed so far to characterize the impact of faults in such systems or to compare alternative solutions concerning dependability features. This paper presents a practical example of benchmarking key dependability features of four different transactional systems using a first proposal of dependability benchmark for OLTP application environments. This dependability benchmark is an extension to the TPC-C standard performance benchmark, and specifies the measures and all the steps required to evaluate both the performance and dependability features of OLTP systems. Two different versions of the Oracle transactional engine running over two different operating systems were evaluated and compared. The results show that dependability benchmarking can be successfully applied to OLTP application environments.
Marco Vieira, Henrique Madeira
DSN2
2003 Open Source Software - A Recipe for Vulnerable Software, or The Only Way to Keep the Bugs and the Bad Guys Out?
Saurabh Bagchi, Henrique Madeira
ISSRE2
2003 A Dependability Benchmark for OLTP Application Environments
Marco Vieira, Henrique Madeira
VLDB2
2002 Adding a Performance-Oriented Perspective to Data Warehouse Design
Pedro Bizarro, Henrique Madeira
DaWaK2
2002 Joint Panel - IPDS and Workshop on Dependability Benchmarking
Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Philip Koopman, Henrique Madeira, Gunter Heiner, Karama Kanoun, Haim Levendel, Brendan Murphy, Lawrence G. Votta, Don Wilson
DSN4
2002 Workshop on Dependability Benchmarking
Philip Koopman, Henrique Madeira
DSN2
2002 Experimental Evaluation of a COTS System for Space Application
abstract
This paper evaluates the impact of transient errors in the operating system of a COTS-based system (CETIA board with two PowerPC 750 processors running LynxOS) and quantifies their effects at both the OS and at the application level. The study has been conducted using a Software-Implemented Fault Injection tool (Xception) and both realistic programs and synthetic workloads (to focus on specific OS features) have been used. The results provide a comprehensive picture of the impact of faults on LynxOS key features (process scheduling and the most frequent system calls), data integrity, error propagation, application termination, and correctness of application results.
Henrique Madeira, Raphael R. Some, Francisco Moreira 0001, Diamantino Costa, David A. Rennels
DSN1
2002 Xception? - Enhanced Automated Fault-Injection Environment
abstract
Discusses Xception, an automated fault injection environment that enables accurate and flexible V&V (verification & validation) and evaluation of mission and business critical computer systems using fault injection. Xception is designed to accommodate a variety of fault injection techniques (according to a wide range of configurations of the tool) and emulate in this way different classes of faults, with particular emphasis to hardware and software faults.
Luis Henriques, Diamantino Costa, Henrique Madeira
DSN4
2002 Recovery and Performance Balance of a COTS DBMS in the Presence of Operator Faults
abstract
A major cause of failures in large database management systems (DBMS) is operator faults. Although most of the complex DBMS have comprehensive recovery mechanisms, the effectiveness of these mechanisms is difficult to characterize. On the other hand, the tuning of a large database is very complex and database administrators tend to concentrate on performance tuning and disregard the recovery mechanisms. Above all, database administrators seldom have feedback on how good a given configuration is concerning recovery. This paper proposes an experimental approach to characterize both the performance and the recoverability in DBMS. Our approach is presented through a concrete example of benchmarking the performance and recovery of an Oracle DBMS running the standard TPC-C benchmark, extended to include two new elements: a fault load based on operator faults and measures related to recoverability. A classification of operator faults in DBMS is proposed. The paper ends with the discussion of the results and the proposal of guidelines to help database administrators in finding the balance between performance and recovery tuning.
Marco Vieira, Henrique Madeira
DSN2
2002 DWS-AQA: A Cost Effective Approach for Very Large Data Warehouses
abstract
Data warehousing applications typically involve massive amounts of data that push database management technology to the limit. A scalable architecture is crucial, not only to handle very large amount of data but also to assure interactive response time to the users. Large data warehouses require a very expensive setup, typically based on high-end servers or high-performance clusters. In this paper we propose and evaluate a simple but very effective method to implement a data warehouse using the computers and workstations typically available in large organizations. The proposed approach is called data warehouse striping with approximate query answering (DWS-AQA). The goal is to use the processing and disk capacity normally available in large workstation networks to implement a data warehouse with a very reduced infrastructure cost. As the data warehouse shares computers that are also being used for other purposes, most of the times only a fraction of the computers will be able to execute the partial queries in time. However, as we show in the paper, the approximated answers estimated from partial results have a very small error for most of the plausible scenarios. Moreover, as the data warehouse facts are partitioned in a strict uniform way, it is possible to calculate tight confidence intervals for the approximated answers, providing the user with a measure of the accuracy of the query results. A set of experiments on the TPC-H benchmark database is presented to show the accuracy of DWS-AQA for a large number of scenarios.
Jorge Bernardino, Pedro Furtado 0001, Henrique Madeira
IDEAS3
2002 Emulation of Software Faults by Educated Mutations at Machine-Code Level
abstract
This paper proposes a new technique to emulate software faults by educated mutations introduced at the machine-code level and presents an experimental study on the accuracy of the injected faults. The proposed method consists of finding key programming structures at the machine code-level where high-level software faults can be emulated. The main advantage of emulating software faults at the machine-code level is that software faults can be injected even when the source code of the target application is not available, which is very important for the evaluation of COTS components or for the validation of software fault tolerance techniques in COTS based systems. The technique was evaluated using several real programs and different types of faults and, additionally, it includes our study on the key aspects that may impact on the technique accuracy. The portability of the technique is also addressed. The results show that classes of faults such as assignment, checking, interface, and simple algorithm faults can be directly emulated using this technique.
João Durães, Henrique Madeira
ISSRE2
2002 Characterization of Operating Systems Behavior in the Presence of Faulty Drivers through Software Fault Emulation
abstract
This paper proposes a practical way to evaluate the behavior of commercial-off-the-shelf (COTS) operating systems in the presence of faulty device drivers. The proposed method is based on the emulation of software faults in target device drivers and the observation of the behavior of the system and of a workload regarding a comprehensive set of failure modes analyzed according to different dimensions. The emulation of software faults itself is done through the injection at machine-code level of selected mutations that represent the code produced when typical programming errors are made in the high-level language code. An important aspect of the proposed methodology is the use of simple and established practices to evaluate operating systems failure modes, thus allowing its use as a dependability benchmarking technique. The generalization of the methodology to any software system built of discrete and identifiable components is also discussed.
João Durães, Henrique Madeira
PRDC2
2002 Definition of Faultloads Based on Operator Faults for DMBS Recovery Benchmarking
abstract
The characterization of database management system (DBMS) recovery mechanisms and the comparison of recovery features of different DBMS require a practical approach to benchmark the effectiveness of recovery in the presence of faults. Existing performance benchmarks for transactional and database areas include two major components: a workload and a set of performance measures. The definition of a benchmark to characterize DBMS recovery needs a new component the faultload. A major cause of failures in large DBMS is operator faults, which make them an excellent starting point for the definition of a generic faultload. This paper proposes the steps for the definition of generic faultloads based on operator faults for DBMS recovery benchmarking. A classification for operator faults in DBMS is proposed and a comparative analysis among three commercially DBMS is presented. The paper ends with a practical example of the use of operator faults to benchmark different configurations of the recovery mechanisms of the Oracle 8i DBMS.
Marco Vieira, Henrique Madeira
PRDC2
2002 Approximate Query Answering Using Data Warehouse Striping
Jorge Bernardino, Pedro Furtado 0001, Henrique Madeira
J. Intell. Inf. Syst.3
2001 Approximate Query Answering Using Data Warehouse Striping
Jorge Bernardino, Pedro Furtado 0001, Henrique Madeira
DaWaK3
2001 Experimental Evaluation of a New Distributed Partitioning Technique for Data Warehouses
abstract
Since data warehousing has become a major field of research there has been a lot of interest in reducing the response time of complex queries posed over the very large databases. The problem is that data warehouses store large amounts of data for decision support, requiring a high level of query performance and scalability to the database engines. A novel round-robin data partitioning approach especially designed for relational data warehouse environments is proposed and experimentally evaluated. This approach is specific to data warehouses implemented over relational repositories using the star schema, as it takes advantage of the specific characteristics of star schemas and typical data warehouse query profiles. The proposed approach guarantees optimal load balancing of query execution and assures high scalability. The experimental evaluation presented in the paper, using a comprehensive set of typical queries from the APB-I benchmark running over Oracle 8, shows that an optimal speedup can be obtained with this technique. The proposed technique constitutes an effective and practical way of coping with very large data warehouses and can be applied to existing database technology.
Jorge Bernardino, Henrique Madeira
IDEAS2
2000 Data Cube Compression with QuantiCubes
Pedro Furtado 0001, Henrique Madeira
DaWaK2
2000 Vmhist: Efficient Multidimensional Histograms with Improved Accuracy
Pedro Furtado 0001, Henrique Madeira
DaWaK2
2000 Joint Evaluation of Performance and Robustness of a COTS DBMS through Fault-Injection
abstract
Presents and discusses observed failure modes of a commercial off-the-shelf (COTS) database management system (DBMS) under the presence of transient operational faults induced by SWIFI (software-implemented fault injection). The Transaction Processing Performance Council (TPC) standard TPC-C benchmark and its associated environment is used, together with fault-injection technology, building a framework that discloses both dependability and performance figures. Over 1600 faults were injected in the database server of a client/server computing environment built on the Oracle 8.1.5 database engine and Windows NT running on COTS machines with Intel Pentium processors. A macroscopic view on the impact of faults revealed that: (1) a large majority of the faults caused no observable abnormal impact in the database server (in 96% of hardware faults and 80% of software faults, the database server behaved normally); (2) software faults are more prone to letting the database server hang or to causing abnormal terminations; (3) up to 51% of software faults lead to observable failures in the client processes.
Diamantino Costa, Tiago Rilho, Henrique Madeira
DSN3
2000 On the Emulation of Software Faults by Software Fault Injection
abstract
This paper presents an experimental study on the emulation of software faults by fault injection. In a first experiment, a set of real software faults has been compared with faults injected by a SWIFI tool (Xception) to evaluate the accuracy of the injected faults. Results revealed the limitations of Xception (and other SWIFI tools) in the emulation of different classes of software faults (about 44% of the software faults cannot be emulated). The use of field data about real faults was discussed and software metrics were suggested as an alternative to guide the injection process when field data is nor available. In a second experiment, a set of rules for the injection of errors meant to emulate classes of software faults was evaluated. The fault triggers used seem to be the cause for the observed strong impact of the faults in the target system and in the program results. The results also show the influence in the fault emulation of aspects such as code size, complexity of data structures, and recursive versus sequential execution.
Henrique Madeira, Diamantino Costa, Marco Vieira
DSN1
2000 FCompress: A New Technique for Queriable Compression of Facts and Datacubes
abstract
Decision support applications must analyze information from data warehouses efficiently. For this reason, huge data warehouses must have mechanisms to cope with massive amounts of data. Reducing and compressing fact tables, summary tables and data cubes is important for faster operation and smaller storage overhead. Traditional compression techniques are not useful in this context except for archiving, because they render the data unqueriable. Although data reduction techniques are useful for fast approximate answers to complex queries, their accuracy is not enough to replace the base data. We present FCompress, a new fact compression technique that effectively replaces the base data, compressing it while maintaining queriability. The approach is based on the premise that a very small and adjustable error is acceptable in many fact attributes. The technique is applicable to fact and summary tables and data cubes alike. It has been evaluated, showing that very small errors can be achieved for point reconstruction (typically below 2%) while the original fact table is reduced to about 35% to 60% of its size and the data cube is reduced to about 15% to 30% of the size. The error is even smaller for typical OLAP queries, usually less than 1%, depending on the degree of aggregation.
Pedro Furtado 0001, Henrique Madeira
IDEAS2
1999 Summary Grids: Building Accurate Multidimensional Histograms
abstract
Data summarization is very important for many data analysis tasks. In this paper we propose a simple but efficient data summarization algorithm, which outputs a histogram for multidimensional data, and make a comparative study of its usage with different distributions and with existing algorithms. The idea is to iteratively grow and modify regions of homogeneous data. This is a different strategy from the commonly used strategy of iteratively fracturing subspaces using straight lines. This work compares both strategies and concludes that the new technique is better and helds good results. We also concluded that discriminate handling of outliers is important to provide good approximates.
Pedro Furtado 0001, Henrique Madeira
DASFAA2
1999 Analysis of Accuracy of Data Reduction Techniques
Pedro Furtado 0001, Henrique Madeira
DaWaK2
1999 Experimental Assessment of COTS DBMS Robustness under Transient Faults
abstract
This paper evaluates the behavior of a common off-the-shelf (COTS) database management system (DBMS) in presence of transient faults. Database applications have traditionally been a field with fault-tolerance needs, concerning both data integrity and availability. While most of the commercially available DBMS provide support for data recovery and fault-tolerance, very limited knowledge was available regarding the impact of transient faults in a COTS database system. In this experimental study, a strict off-the-shelf target system is used (Oracle 7.3 server running on top of Wintel platform), combined with a TPC-A based workload and a software implemented fault injection tool, XceptionNT. It was found out that a non-negligible amount of induced faults, 13%, lead to the database server hanging or premature termination. However, the results also show that COTS DBMS products has a reasonable behavior concerning data integrity, none of the injected faults affected end user data.
Diamantino Costa, Henrique Madeira
PRDC2
1998 Xception: A Technique for the Experimental Evaluation of Dependability in Modern Computers
abstract
An important step in the development of dependable systems is the validation of their fault tolerance properties. Fault injection has been widely used for this purpose, however with the rapid increase in processor complexity, traditional techniques are also increasingly more difficult to apply. This paper presents a new software-implemented fault injection and monitoring environment, called Xception, which is targeted at modern and complex processors. Xception uses the advanced debugging and performance monitoring features existing in most modern processors to inject quite realistic faults by software, and to monitor the activation of the faults and their impact on the target system behavior in detail. Faults are injected with minimum interference with the target application. The target application is not modified, no software traps are inserted, and it is not necessary to execute the target application in special trace mode (the application is executed at full speed). Xception provides a comprehensive set of fault triggers, including spatial and temporal fault triggers, and triggers related to the manipulation of data in memory. Faults injected by Xception can affect any process running on the target system (including the kernel), and it is possible to inject faults in applications for which the source code is not available. Experimental, results are presented to demonstrate the accuracy and potential of Xception in the evaluation of the dependability properties of the complex computer systems available nowadays.
Henrique Madeira, João Gabriel Silva
IEEE Trans. Software Eng.2
1995 Experimental Evaluation of the Impact of Processor Faults on Parallel Applications
abstract
This paper addresses the problem of processor faults in distributed memory parallel systems. It shows that transient faults injected at the processor pins of one node of a commercial parallel computer, without any particular fault-tolerant techniques, can cause erroneous application results for up to 43% of the injected faults (depending on the application). In addition to these very subtle faults, up to 19% of the injected faults (almost independent on the application) caused the system to hang up. These results show that fault-tolerant techniques are absolutely required in parallel systems, not only to ensure the completion of long-run applications but, and more important, to achieve confidence in the application results. The benefits of including some fairly simple behaviour based error detection mechanisms in the system were evaluated together with Algorithm Based Fault Tolerance (ABFT) techniques. The inclusion of such Mechanisms in parallel systems seems to be very important for detecting most of those subtle errors without greatly affecting the performance and the cost of these systems.
Diamantino Costa, Francisco Moreira 0001, Henrique Madeira, Mário Zenha Rela, João Gabriel Silva
SRDS3
1990 Experimental evaluation of a set of simple error detection mechanisms
Henrique Madeira, Gonçalo Quadros, João Gabriel Silva
Microprocessing and Microprogramming1
1989 The fault-tolerant architecture of the safe system
Henrique Madeira, Boavida Fernandes, Mário Zenha Rela, João Gabriel Silva
Microprocessing and Microprogramming1