EDBT 2026 Demo / reviewers in the wild / expert
Roberto Pietrantuono
dblp:07/504
· DBLP profile ↗
67ranked-venue papers
15as first author
28since 2021 · last 2026
0000-0003-2449-1724ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 50 · 9 first-author · 22 since 2021Security and privacy · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multivariate anomaly detection and root cause analysis of energy issues in microservice-based systemsabstractContext: Microservice-based systems have become the architecture style of choice for modern applications, offering scalability, flexibility, and resilience. However, their distributed nature leads to increased resource consumption and energy inefficiencies, posing challenges for maintaining sustainable operations. Accurate anomaly detection (AD) and root cause analysis (RCA) tools are critical for diagnosing energy consumption issues in these systems, yet existing solutions often lack focus on energy metrics. Goal: This study aims to evaluate the effectiveness of AD and RCA algorithms in identifying and diagnosing performance-related energy consumption anomalies in microservice-based systems. Method: Two representative systems, Sock Shop and Train Ticket, are deployed under controlled environments. Then, anomalies are deliberately introduced by stressing at the same time CPU, memory, and disk resources. The data collection is conducted using Prometheus for performance metrics and Scaphandre for energy metrics. Once normal and anomalous datasets are constructed for each system, the study evaluates five AD algorithms (Birch, iForest, KNN, LOF, and SVM) and four RCA algorithms (MicroRCA, CausalRCA, CIRCA, and RCD) based on their precision, recall, and scalability across varied scenarios and workloads. Results: The experiment reveals that overall, iForest is the most effective AD algorithms in detecting energy anomalies (0.59 F-Score in Sock Shop and 0.634 F-Score in Train Ticket). In particular, iForest performs better in precision when the user load is high (1000 concurrent users). For RCA, CIRCA performs well in identifying root causes in smaller systems, while RCD is more scalable for larger and more complex systems. Conclusions: The findings of this study provide insights for both researchers and practitioners. In the context of our experiment, AD algorithms tend to perform relatively well, whereas RCA algorithms tend to be imprecise in localizing energy issues. Berta Rodriguez Sanchez, Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Ivano Malavolta |
J. Syst. Softw. | 4 |
| 2026 | Aging-Related Bug Prediction Based on Multi-View Graph Feature Learning and Graph-TransformerabstractSoftware aging, characterized by an increasing failure rate or performance degradation in long-running software systems, poses significant risks, including substantial financial losses and potential threats to human lives. This phenomenon is primarily driven by the accumulation of runtime errors, commonly referred to as aging-related bugs (ARBs). Aging-related bug prediction (ARBP) has been proposed to facilitate the detection and remediation of ARBs prior to software release. However, ARBP’s effectiveness heavily depends on the quality of dataset features used. Previous research has largely relied on a standard set of manually designed metrics, often overlooking that these metrics may fail to distinguish between code segments with different semantics, even when they exhibit identical metric values. While some studies have attempted to develop models that learn semantic features from source code, they typically focus on token-level or graph-level features, neglecting a comprehensive exploration of ARB characteristics within the source code. Specifically, there is insufficient discussion on whether deep semantic features can adequately capture the essential traits that trigger aging phenomena. In this paper, we propose a novel multi-view graph feature learning framework based on Graph-Transformer, which integrates newly proposed ARB features extracted from Abstract Syntax Trees with Code Property Graphs for feature learning. Our approach effectively captures hierarchical structures and variable dependencies, facilitating the identification of complex interactions that contribute to ARBs. Additionally, we implement sub-graph sampling and class imbalance strategies to enhance model performance. Experimental results across three datasets demonstrate that our method surpasses state-of-the-art approaches, a code property graph-based feature extraction method (specifically SGT), achieving precision improvements of 8.2% on Linux, 15.4% on MySQL, and 2.5% on NetBSD, thereby establishing a new benchmark for ARB prediction. Jianwen Xiang, Roberto Natella, Roberto Pietrantuono, Domenico Cotroneo |
IEEE Trans. Software Eng. | 7 |
| 2025 | Adaptive Probabilistic Operational Testing for Large Language Models EvaluationabstractLarge Language Models (LLM) empower many modern software systems, and are required to be highly accurate and reliable. Evaluating LLM poses challenges due to the high costs of manual labeling and of validation of labeled data.This study investigates the suitability of probabilistic operational testing for effective and efficient evaluation of LLM, focusing on a case study with DistilBERT. To this aim, we adopt an existing framework (DeepSample) for Deep Neural Network (DNN) testing and adapt it to the LLM domain by introducing auxiliary variables tailored to LLM and classification tasks.Through a comprehensive evaluation, we demonstrate how sampling-based operational testing can yield reliable LLM accuracy estimates and effectively expose failures, or, under testing budget constraints, it can find a trade off between accuracy estimation and failure exposure. The experimental results, using DistilBERT on three sentiment analysis datasets, show that sampling-based methods can provide cost effective and reliable operational accuracy assessment for LLM. These findings offer practical insights for testers and help address critical gaps in current LLM evaluation practices. Ali Asgari, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
AST | 3 |
| 2025 | Exploring Causal Modeling to Enhance Diabetes Prediction and ManagementabstractGestional Diabetes Mellitus is a common pregnancy complication, affecting up to 10% of pregnancies worldwide. This study applies causal modeling to PIMA dataset to support diabetes prediction and management. Results show that healthy lifestyle interventions reduce glucose level, and counterfactual analysis suggest that modifying key-factors may prevent disease onset. Findings highlight the potential of causal inference for risk identification and prevention. Patrizia Quaranta, Roberto Pietrantuono |
CBMS | 2 |
| 2025 | Log-Driven Testing of Microservice Systems with TransformersabstractRegression testing enhances software reliability by detecting regressions in new versions. Regression test suites often lack awareness of real-world product/service usage, potentially leading to undetected faults and ineffective testing scenarios. We propose LogTest, a transformer-based approach that learns from event logs and system traces to automatically generate service invocation sequences that mimic observed system behaviors. These sequences enhance regression test suites by exposing past real execution patterns. A preliminary experimentation on a realistic benchmark demonstrates its potential. Raffaele Della Corte, Roberto Pietrantuono, Stefano Russo 0001 |
ICWS | 2 |
| 2025 | Learning-based Automated Generation of Critical Workload Configurations for Microservices Performance TestingabstractPerformance testing is an essential activity in the engineering of microservice applications to identify deviations from the specified ranges of relevant metrics and to analyse resources usage. It demands for high automation to fit within the short microservices development-operation cycles. Engineers are often interested in identifying critical workloads - ideally, in the “minimal“ load configurations causing tests to expose performance issues. Triggering performance issues is challenging, requiring proper workload characterization and test design. We present the microWave framework for learning-based automated generation of critical performance testing workloads for microservices. The framework can harness various learning strategies: we analyze a Deep Neural Network, a Large Language Model and a Causal Reasoning strategy. We evaluate them experimentally on four subjects, using a random approach and a manually-crafted ground truth as baselines. The results show that the strategies exhibit different behavior depending on the data they learn from. When inferring from past executions data including performance issues, the causal model performs better. The random predictor is preferable when no data is available; however, it is more costly as it requires more tests. The results allow to draw practical recommendations for testers on how to select the most suitable strategy depending on the needs. Cristian Mascia, Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ICWS | 4 |
| 2025 | Reinforcement learning for online testing of autonomous driving systems: a replication and extension studyabstractIn a recent study, Reinforcement Learning (RL) used in combination with many-objective search, has been shown to outperform alternative techniques (random search and many-objective search) for online testing of Deep Neural Network-enabled systems. The empirical evaluation of these techniques was conducted on a state-of-the-art Autonomous Driving System (ADS). This work is a replication and extension of that empirical study. Our replication shows that RL does not outperform pure random test generation in a comparison conducted under the same settings of the original study, but with no confounding factor coming from the way collisions are measured. Our extension aims at eliminating some of the possible reasons for the poor performance of RL observed in our replication: (1) the presence of reward components providing contrasting feedback to the RL agent; (2) the usage of an RL algorithm (Q-learning) which requires discretization of an intrinsically continuous state space. Results show that our new RL agent is able to converge to an effective policy that outperforms random search. Results also highlight other possible improvements, which open to further investigations on how to best leverage RL for online ADS testing. Luca Giamattei, Matteo Biagiola, Roberto Pietrantuono, Stefano Russo 0001, Paolo Tonella |
Empir. Softw. Eng. | 3 |
| 2025 | Causal reasoning in Software Quality Assurance: A systematic reviewabstractContext: Software Quality Assurance (SQA) is a fundamental part of software engineering to ensure stakeholders that software products work as expected after release in operation. Machine Learning (ML) has proven to be able to boost SQA activities and contribute to the development of quality software systems. In this context, Causal Reasoning is gaining increasing interest as a methodology to go beyond a purely data-driven approach by exploiting the use of causality for more effective SQA strategies. Objective: Provide a broad and detailed overview of the use of causal reasoning for SQA activities, in order to support researchers to access this research field, identifying room for application, main challenges and research opportunities. Methods: A systematic review of the scientific literature on causal reasoning for SQA. The study has found, classified, and analyzed 86 articles, according to established guidelines for software engineering secondary studies. Results: Results highlight the primary areas within SQA where causal reasoning has been applied, the predominant methodologies used, and the level of maturity of the proposed solutions. Fault localization is the activity where causal reasoning is more exploited, especially in the web services/microservices domain, but other tasks like testing are rapidly gaining popularity. Both causal inference and causal discovery are exploited, with the Pearl’s graphical formulation of causality being preferred, likely due to its intuitiveness. Tools to favor their application are appearing at a fast pace — most of them after 2021. Conclusions: The findings show that causal reasoning is a valuable means for SQA tasks with respect to multiple quality attributes , especially during V&V, evolution and maintenance to ensure reliability, while it is not yet fully exploited for phases like requirements engineering and design. We give a picture of the current landscape, pointing out exciting possibilities for future research. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
Inf. Softw. Technol. | 3 |
| 2025 | Automatic Generation of Plausible Co-Occurring Causes for Effects Explanation or PredictionabstractIn numerous contexts, ranging from systems safety assessment to finance and medical diagnosis, a relevant causal inference task is to predict unseen rare events—the so-called black swans . These are plausible, high-impact, but unexpected events for whose prediction a probabilistic-based causal inference falls short. For instance, a safety analyst needs to hypothesize potential rare co-causes that could lead to an accident, so as to manage the most unexpected failures besides the more obvious ones. Given an effect, we use abduction to support the generation of a plausible set of explanatory hypotheses for its causes. We present a generative evolutionary strategy—called Evolutionary Abduction (EVA)—for automating abductive inference by repeatedly constructing hypothetical cause-effect instances, and then automatically assessing their plausibility as well as their novelty with respect to already known instances—a mechanism mimicking the human reasoning employed whenever we need to select the best candidates from a set of hypotheses. Experiments with four datasets confirm that EVA can construct new and realistic multiple-cause hypotheses for a given effect. EVA outperforms alternative strategies based on probabilistic-based causal inference as well as state-of-the-art evolutionary algorithms, generating closer-to-real instances in most settings and datasets. Roberto Pietrantuono, Stefano Russo 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2024 | Identifying Performance Issues in Microservice Architectures through Causal ReasoningabstractEvaluating the performance of Microservices Architectures (MSA) is essential to ensure their proper functioning and meet end-user satisfaction. For MSA performance analysts, one of the most challenging tasks is to determine the cause of any deviation of relevant metrics from the specified range. Luca Giamattei, Antonio Guerriero, Ivano Malavolta, Cristian Mascia, Roberto Pietrantuono, Stefano Russo 0001 |
AST | 5 |
| 2024 | DeepSample: DNN sampling-based testing for operational accuracy assessmentabstractDeep Neural Networks (DNN) are core components for classification and regression tasks of many software systems. Companies incur in high costs for testing DNN with datasets representative of the inputs expected in operation, as these need to be manually labelled. The challenge is to select a representative set of test inputs as small as possible to reduce the labelling cost, while sufficing to yield unbiased high-confidence estimates of the expected DNN accuracy. At the same time, testers are interested in exposing as many DNN mispredictions as possible to improve the DNN, ending up in the need for techniques pursuing a threefold aim: small dataset size, trustworthy estimates, mispredictions exposure. Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 2 |
| 2024 | Anomaly Detection and Root Cause Analysis of Microservices Energy ConsumptionabstractWith the expansion of cloud computing and data centers, the need has arisen to tackle their environmental impact. The increasing adoption of microservice architectures, while offering scalability and flexibility, poses new challenges in the effective management of systems’ energy consumption.This study analyzes experimentally the effectiveness, with respect to energy consumption, of algorithms for Anomaly Detection (AD) and Root Cause Analysis (RCA) for (containerized) microservices systems. The study analyzes five AD and three RCA algorithms. Metrics to assess the effectiveness of AD algorithms are Precision, Recall, and F-Score. For RCA algorithms, the chose metric is Precision at level k. Two subjects of different complexity are used: Sock Shop and UNI-Cloud. Experiments use a cross-over paired comparison design, involving multiple randomized runs for robust measures.The experiments show that AD algorithms exhibit a relatively moderate performance. The mean adjusted Precision for Sock Shop is 61.5%, while it is 75% for the best-performing algorithms (BIRCH, KNN, and SVM) on UNI-Cloud. The Recall and F-Score for UNI-Cloud, for the same algorithms, are 75%, while for Sock Shop KNN yields the best outcome at roughly 45%. MicroRCA and RCD emerge as the top-performing algorithms for RCA.We found that the effectiveness of AD algorithms is strongly influenced by anomaly thresholds, emphasizing the importance of careful tuning such algorithms. RCA algorithms reveal promising results, particularly RCD and MicroRCA, which showed robust performance. However, challenges remain, as seen with the ϵ-diagnosis algorithm, suggesting the need for further refinement.For DevOps engineers, the findings highlight the need to carefully select and tune AD and RCA algorithms for energy, and to take into account system topology and monitoring configurations. Maximilian Stefan Floroiu, Stefano Russo 0001, Luca Giamattei, Antonio Guerriero, Ivano Malavolta, Roberto Pietrantuono |
ICWS | 6 |
| 2024 | Automated functional and robustness testing of microservice architecturesabstractMicroservice Architectures (MSA) are nowadays largely adopted by companies in several domains to provide on-demand services. The reliability of microservices is fundamental to avoid failures compromising the business functionalities. MSA automated testing is possible thanks to well-defined service interfaces specified in open formats like OpenAPI/Swagger. To support automated MSA functional and non-functional testing, we define a framework that: (i) generates test cases with valid and invalid inputs, and executes and monitors tests; (ii) provides coverage and failure information not only on edge, but also on internal microservices; (iii) has the novel feature of identifying causal relations in observed chains of microservices failures. We abstract the testing process of MSA, present the MacroHive framework and its causal inference engine, compare it experimentally to state-of-the-art tools, and discuss its benefits in the MSA testing process. MacroHive exhibits performance comparable to advanced existing tools in terms of edge-level coverage. However, MacroHive has a better failure rate and provides the unique advantages of giving insights about internal coverage and failures, and of inferring causality in failure chains, evidencing microservices to be improved to increase the whole MSA reliability. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 3 |
| 2024 | Monitoring tools for DevOps and microservices: A systematic grey literature reviewabstractMicroservice-based systems are usually developed according to agile practices like DevOps, which enables rapid and frequent releases to promptly react and adapt to changes. Monitoring is a key enabler for these systems, as they allow to continuously get feedback from the field and support timely and tailored decisions for a quality-driven evolution. In the realm of monitoring tools available for microservices in the DevOps-driven development practice, each with different features, assumptions, and performance, selecting a suitable tool is an as much difficult as impactful task. This article presents the results of a systematic study of the grey literature we performed to identify, classify and analyze the available monitoring tools for DevOps and microservices. We selected and examined a list of 71 monitoring tools, drawing a map of their characteristics, limitations, assumptions, and open challenges, meant to be useful to both researchers and practitioners working in this area. Results are publicly available and replicable. Editor's note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001, Ivano Malavolta, Tanjina Islam, Madalina Dinga, Anne Koziolek, Snigdha Singh, Martin Armbruster, Jose-Maria Gutierrez-Martinez, Sergio Caro-Álvaro, Daniel Rodríguez-García, Sebastian Weber 0001, Jörg Henß, Estrella Fernández Vogelin, Fernando Simön Panojo |
J. Syst. Softw. | 3 |
| 2024 | SGT: Aging-related bug prediction via semantic feature learning based on graph-transformer
Jianwen Xiang, Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono |
J. Syst. Softw. | 7 |
| 2024 | Testing the Resilience of MEC-Based IoT Applications Against Resource Exhaustion AttacksabstractMulti-access Edge Computing (MEC) is an emerging computing model that provides the necessary on-demand resources and services to the edge of the network, ensuring powerful computing, storage capacity, mobility, location, and context awareness support to emerging Internet of Things (IoT) applications. Nonetheless, its complex hierarchical model introduces new architectural interdependencies, which can influence the resilience of IoT applications against cyber attacks. Although application resilience has been investigated in the context of cloud computing, existing studies are not directly applicable to such an extended edge-cloud paradigm. The use of different enabling technologies at the edge of the network, such as various wireless access technologies and virtualization, implies several threats and challenges that make the analysis and deployment of resilience mechanisms a technically challenging problem. In this article, we first present an overview of the threat model, describing the threats for the different layers of this paradigm. We then study the impact of resource-exhausting attacks – a particularly relevant class for this paradigm - on three different IoT applications exploiting the services offered by the MEC-based architecture. We adopt a testing-based methodology conceived to characterize the resilience of such applications under attack. A set of most important resilience-related indicators are also identified. The characterization's results are useful to support the analyst in planning proper protection means at individual architectural layers. Roberto Pietrantuono, Massimo Ficco, Francesco Palmieri 0002 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Causality-driven Testing of Autonomous Driving SystemsabstractTesting Autonomous Driving Systems (ADS) is essential for safe development of self-driving cars. For thorough and realistic testing, ADS are usually embedded in a simulator and tested in interaction with the simulated environment. However, their high complexity and the multiple safety requirements lead to costly and ineffective testing. Recent techniques exploit many-objective strategies and ML to efficiently search the huge input space. Despite the indubitable advances, the need for smartening the search keep being pressing. This article presents CART ( CAusal-Reasoning-driven Testing ), a new technique that formulates testing as a causal reasoning task. Learning causation, unlike correlation, allows assessing the effect of actively changing an input on the output, net of possible confounding variables. CART first infers the causal relations between test inputs and outputs, then looks for promising tests by querying the learnt model. Only tests suggested by the model are run on the simulator. An extensive empirical evaluation, using Pylot as ADS and CARLA as simulator, compares CART with state-of-the-art algorithms used recently on ADS. CART shows a significant gain in exposing more safety violations and does so more efficiently. More broadly, the work opens to a wider exploitation of causal learning beside (or on top of) ML for testing-related tasks. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2023 | An Empirical Evaluation of the Energy and Performance Overhead of Monitoring Tools on Docker-Based Systems
Madalina Dinga, Ivano Malavolta, Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono |
ICSOC (1) | 5 |
| 2023 | IFCM: An improved Fuzzy C-means clustering method to handle Class Overlap on Aging-related Software Bug PredictionabstractSoftware aging refers to a problem of performance decay in long-running software systems. This phenomenon is primarily attributed to the accumulation of run-time errors, commonly known as aging-related bugs (ARBs). Detecting ARBs through Aging-related Bug Prediction (ARBP) is crucial in ensuring system reliability. The effectiveness of ARBP heavily relies on the quality of datasets. However, ARB datasets often suffer from class overlap, where instances from different classes exhibit similar feature values. Class overlap poses a significant challenge as it compromises the quality of training data and subsequently impacts ARBP accuracy. To address this issue, we propose an improved Fuzzy C-means clustering method named IFCM, designed to mitigate class overlap in ARBP tasks. IFCM can identify whether an instance occurs overlap, and identify the overlap degree of this instance through the predefined parameters. We evaluate our proposed method on two public datasets Linux and MySQL and one self-collected dataset NetBSD using five different classifiers with five performance metrics (AUC, F1, Balance, PD, PF). Comparison with four existing methods (No clean, NCL, IKMCCA, ROCT) demonstrates that IFCM is effective in alleviating class overlap in ARBP. For Instance, IFCM achieves promising results in terms of AUC blue (which are 0.762, 0.757, and 0.642) and Balance (which are 0.709, 0.736, and 0.595) at the dataset level. Shuo Feng 0003, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang, Roberto Pietrantuono, Roberto Natella, Domenico Cotroneo |
ISSRE | 6 |
| 2023 | DevOpRET: Continuous reliability testing in DevOpsabstractAbstract To enter the production stage, in DevOps practices candidate software releases have to pass quality gates, where they are assessed to meet established target values for key indicators of interest. We believe software reliability should be an important such indicator, as it greatly contributes to the end‐user satisfaction. We propose DevOpRET , an approach for reliability testing as part of the acceptance testing stage in DevOps. DevOpRET relies on operational‐profile–based testing, a common reliability assessment technique. DevOpRET leverages usage and failure data monitored in operations to continuously refine its estimate. We evaluate accuracy and efficiency of DevOpRET through controlled experiments with a real‐world open source platform and with a microservice architectures benchmark. The results show that DevOpRET provides accurate and efficient estimates of the true reliability over subsequent DevOps cycles. Antonia Bertolino, Guglielmo De Angelis, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
J. Softw. Evol. Process. | 5 |
| 2023 | A Comparative Analysis of Software Aging in Image Classifiers on Cloud and EdgeabstractImage classifiers for recognizing real-world objects are widely used in the Internet of Things (IoT) and Cyber-Physical Systems(CPSs). A classifier is trained offline by machine learning algorithms with training data sets, and then it is deployed on a cloud or an edge computing system for online label predictions. As the classifier's performance depends on the underlying software infrastructure, it may degrade over time due to software faults causing software aging. In this paper, we address this issue and experimentally investigate software aging observed in an image classification system that continuously runs on cloud and edge computing environments. We apply several statistical techniques to analyze degradation trends in the systems under stress tests. Our statistical trend analysis confirms the degradation trends in the throughput as well as the available memory resources both in the cloud and the edge environments. Contrary to our expectation, the edge computing environment under test had much less impact on the performance degradation than our cloud environment when the workload is high, although the latter one has four times larger allocated memory resources. We also show that the observed performance degradation trends are associated with the memory usage of specific processes by performing correlation analysis. Ermeson Carneiro de Andrade, Roberto Pietrantuono, Fumio Machida, Domenico Cotroneo |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | Survivability Analysis of IoT Systems Under Resource Exhausting AttacksabstractEssential services in an Internet of Things (IoT)-based critical system should be continuously provided even when undesirable events like failures, attacks, and emergencies happen. In this work, we analyze the system’s ability to survive failures that are caused by resource exhaustion attacks. Such ability to survive means that the system’s services should be provided in compliance with the associated requirements also in presence of failures and other undesired events. Accordingly, we present a hybrid method (i.e., measurements- and model-based) to assess the expected survivability of an IoT system under resource-exhaustion attacks and, based on it, to optimize the preventive maintenance trigger period that maximizes survivability and minimizes the expected downtime cost. A realistic case study is implemented to emulate an IoT scenario and used to estimate the extent of resource consumption at each layer of the IoT stack when the system is subject to a resource-exhaustion attack. A semi-Markov process is then adopted to model the transient behavior of the system during an intrusion. The model is enriched with an additional state that represents a proactive recovery, in which the system is not available for a maintenance action aimed at preventing failure. The model solution gives the optimal maintenance triggering time. Roberto Pietrantuono, Massimo Ficco, Francesco Palmieri 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Automated Grey-Box Testing of Microservice ArchitecturesabstractMicroservices Architectures (MSA) have found large adoption in companies delivering online services, often in conjunction with agile development practices. Microservices are distributed, independent and polyglot entities – all features favouring black-box testing. However, for real-scale MSA, a pure black-box strategy may not be able to exercise the system to properly cover the interactions involving internal microservices.We propose a grey-box strategy (MACROHIVE) for automated testing and monitoring of (internal) microservices interactions. It uses combinatorial testing to generate valid and invalid tests from microservices specification. Tests execution and monitoring are automated by a service mesh infrastructure. MACROHIVE runs the tests and traces the interactions among microservices, to report about internal coverage and failing behaviour.MACROHIVE is experimented on TrainTicket, an open-source MSA benchmark. It performs comparably to state-of-the-art techniques in terms of edge-level coverage, but exposes internal failures undetected by black-box testing, gives detailed internal coverage information, and requires fewer tests. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
QRS | 3 |
| 2022 | An Empirical Study on Software Aging of Long-Running Object Detection AlgorithmsabstractEfficient and effective object detection is a key problem in Computer Vision. Numerous object detection algorithms have been developed, whose aim is to achieve two conflicting goals, namely accuracy and efficiency, while being executed in real-time with high robustness. Many of these algorithms must run for an extended period of time, i.e., in video surveillance or in self-driving cars – a working condition that make them subject to the risk of software aging.In this work, we focus on evaluating several object detection algorithms to understand if and to what extent they are affected by software aging. A measurement-based aging approach was adopted, with a series of long-running tests and subsequent data analysis. The results report significant trends of performance degradation, sometimes leading to aging-related failures, as well as memory consumption trends, which turned out to be the main issue across all the experiments. Roberto Pietrantuono, Domenico Cotroneo, Ermeson Carneiro de Andrade, Fumio Machida |
QRS | 1 |
| 2022 | Software micro-rejuvenation for Android mobile systems
Domenico Cotroneo, Luigi De Simone, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 4 |
| 2021 | Automated Hypotheses Generation via Combinatorial Causal OptimizationabstractA powerful form of causal inference employed in many tasks, such as medical diagnosis, criminology, root cause analysis, biology, is abduction. Given an effect, it aims at generating a plausible and useful set of explanatory hypotheses for its causes. This article formulates the abductive hypotheses generation activity as an optimization problem, introducing a new class called Combinatorial Causal Optimization Problems (CCOP). In a CCOP, solutions are in the form of cause-effect combinations: algorithms are required to construct hypothetical solutions automatically assessed for plausibility - a mechanism mimicking the human reasoning when he skims the best solutions from a set of hypotheses - and for novelty with respect to already known solutions. The paper presents the CCOP formulation and four real-world benchmark problems from various domains, released along with artefacts to implement, run and properly evaluate algorithms for CCOP solutions. Then, for illustrative purpose, four conventional evolutionary algorithms are customized to solve CCOPs. Their application demonstrates the possibility of generating useful solutions (i.e., novel and realistic hypotheses for a given effect), but also evidences a great margin for improvement in terms of ratio of good vs bad solutions. Roberto Pietrantuono |
CEC | 1 |
| 2021 | Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsabstractDeep Neural Networks (DNN) are typically tested for accuracy relying on a set of unlabelled real world data (operational dataset), from which a subset is selected, manually labelled and used as test suite. This subset is required to be small (due to manual labelling cost) yet to faithfully represent the operational context, with the resulting test suite containing roughly the same proportion of examples causing misprediction (i.e., failing test cases) as the operational dataset. However, while testing to estimate accuracy, it is desirable to also learn as much as possible from the failing tests in the operational dataset, since they inform about possible bugs of the DNN. A smart sampling strategy may allow to intentionally include in the test suite many examples causing misprediction, thus providing this way more valuable inputs for DNN improvement while preserving the ability to get trustworthy unbiased estimates. This paper presents a test selection technique (DeepEST) that actively looks for failing test cases in the operational dataset of a DNN, with the goal of assessing the DNN expected accuracy by a small and "informative" test suite (namely with a high number of mispredictions) for subsequent DNN improvement. Experiments with five subjects, combining four DNN models and three datasets, are described. The results show that DeepEST provides DNN accuracy estimates with precision close to (and often better than) those of existing sampling-based DNN testing techniques, while detecting from 5 to 30 times more mispredictions, with the same test suite size. Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 2 |
| 2021 | Adaptive Test Case Allocation, Selection and Generation Using Coverage Spectrum and Operational ProfileabstractWe present an adaptive software testing strategy for test case allocation, selection and generation, based on the combined use of operational profile and coverage spectrum, aimed at achieving high delivered reliability of the program under test. Operational profile-based testing is a black-box technique considered well suited when reliability is a major concern, as it selects the test cases having the largest impact on failure probability in operation. Coverage spectrum is a characterization of a program's behavior in terms of the code entities (e.g., branches, statements, functions) that are covered as the program executes. The proposed strategy - named covrel+ - complements operational profile information with white-box coverage measures, so as to adaptively select/generate the most effective test cases for improving reliability as testing proceeds. We assess covrel+ through experiments with subjects commonly used in software testing research, comparing results with traditional operational testing. The results show that exploiting operational and coverage data in an integrated adaptive way allows generally to outperform operational testing at achieving a given reliability target, or at detecting faults under the same testing budget, and that covrel+ has greater ability than operational testing in detecting hard-to-detect faults. Antonia Bertolino, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
IEEE Trans. Software Eng. | 3 |
| 2020 | Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationabstractIn Continuous Integration (CI), regression testing is constrained by the time between commits. This demands for careful selection and/or prioritization of test cases within test suites too large to be run entirely. To this aim, some Machine Learning (ML) techniques have been proposed, as an alternative to deterministic approaches. Two broad strategies for ML-based prioritization are learning-to-rank and what we call ranking-to-learn (i.e., reinforcement learning). Various ML algorithms can be applied in each strategy. In this paper we introduce ten of such algorithms for adoption in CI practices, and perform a comprehensive study comparing them against each other using subjects from the Apache Commons project. We analyze the influence of several features of the code under test and of the test process. The results allow to draw criteria to support testers in selecting and tuning the technique that best fits their context. Antonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 4 |
| 2020 | A comprehensive study on software aging across android versions and vendors
Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella, Roberto Pietrantuono |
Empir. Softw. Eng. | 4 |
| 2020 | On the testing resource allocation problem: Research trends and perspectives
Roberto Pietrantuono |
J. Syst. Softw. | 1 |
| 2020 | A survey on software aging and rejuvenation in the cloud
Roberto Pietrantuono, Stefano Russo 0001 |
Softw. Qual. J. | 1 |
| 2020 | Testing microservice architectures for operational reliabilityabstractSummary Microservice architectures (MSA) is an emerging software architectural paradigm for service‐oriented applications, well‐suited for dynamic contexts requiring loosely coupled independent services, frequent software releases and decentralized governance. A key problem in the engineering of MSA applications is the estimate of their reliability, which is difficult to perform prior to release due frequent releases/service upgrades, dynamic service interactions, and changes in the way customers use the applications. This paper presents an in vivo testing method, named EMART, to faithfully assess the reliability of an MSA application in operation. EMART is based on an adaptive sampling strategy, leveraging monitoring data about microservices usage and failure/success of user demands. We present results of evaluation of estimation accuracy, confidence and efficiency, through a set of controlled experiments with publicly available subjects. © 2019 John Wiley & Sons, Ltd. Roberto Pietrantuono, Stefano Russo 0001, Antonio Guerriero |
Softw. Test. Verification Reliab. | 1 |
| 2018 | Run-Time Reliability Estimation of Microservice ArchitecturesabstractMicroservices are gaining popularity as an architectural paradigm for service-oriented applications, especially suited for highly dynamic contexts requiring loosely-coupled independent services, frequent software releases, decentralized governance and data management. Because of the high flexibility and evolvability characterizing microservice architectures (MSAs), it is difficult to estimate their reliability at design time, as it changes continuously due to the services' upgrades and/or to the way applications are used by customers. This paper presents a testing method for on-demand reliability estimation of microservice applications in their operational phase. The method allows to faithfully assess, upon request, the reliability of a MSA-based application under a scarce testing budget, at any time when it is in operation, and exploit field data about microservice usage and failing/successful demands. A new in-vivo testing algorithm is developed based on an adaptive web sampling strategy, named Microservice Adaptive Reliability Testing (MART). The method is evaluated by simulation, as well as by experimentation on an example application based on the Netflix Open Source Software MSA stack, with encouraging results in terms of estimation accuracy and, especially, efficiency. Roberto Pietrantuono, Stefano Russo 0001, Antonio Guerriero |
ISSRE | 1 |
| 2018 | Probabilistic Sampling-Based Testing for Accelerated Reliability AssessmentabstractA relevant objective of software reliability assessment is to get unbiased estimates with an acceptable trade-off between the number of tests required and the variance of the estimate. A low variance is desirable to increase the confidence in the estimate, but too many tests may be required by conventional reliability assessment testing techniques based solely on the operational profile. This article presents probabilistic sampling-based testing, a new technique using unequal probability sampling to exploit auxiliary information about the software under test so as to assess reliability unbiasedly and efficiently. The technique expedites the assessment process assuming the availability of some prior belief about input regions failure proneness. The evaluation by simulation and experimentally shows promising results in terms of estimate accuracy and efficiency. Roberto Pietrantuono, Stefano Russo 0001 |
QRS | 1 |
| 2018 | Aging-related performance anomalies in the apache storm stream processing system
Massimo Ficco, Roberto Pietrantuono, Stefano Russo 0001 |
Future Gener. Comput. Syst. | 2 |
| 2018 | A software quality framework for large-scale mission-critical systems engineering
Gabriella Carrozza, Roberto Pietrantuono, Stefano Russo 0001 |
Inf. Softw. Technol. | 2 |
| 2018 | Multiobjective Testing Resource Allocation Under UncertaintyabstractTesting resource allocation is the problem of planning the assignment of resources to testing activities of software components so as to achieve a target goal under given constraints. Existing methods build on software reliability growth models (SRGMs), aiming at maximizing reliability given time/cost constraints, or at minimizing cost given quality/time constraints. We formulate it as a multiobjective debug-aware and robust optimization problem under uncertainty of data, advancing the state-of-the-art in the following ways. Multiobjective optimization produces a set of solutions, allowing to evaluate alternative tradeoffs among reliability, cost, and release time. Debug awareness relaxes the traditional assumptions of SRGMs-in particular the very unrealistic immediate repair of detected faults-and incorporates the bug assignment activity. Robustness provides solutions valid in spite of a degree of uncertainty on input parameters. We show results with a real-world case study. Roberto Pietrantuono, Pasqualina Potena, Antonio Pecchia, Daniel Rodríguez-García, Stefano Russo 0001, Luis Fernández-Sanz |
IEEE Trans. Evol. Comput. | 1 |
| 2017 | Adaptive coverage and operational profile-based testing for reliability improvementabstractWe introduce covrel, an adaptive software testing approach based on the combined use of operational profile and coverage spectrum, with the ultimate goal of improving the delivered reliability of the program under test. Operational profile-based testing is a black-box technique that selects test cases having the largest impact on failure probability in operation, as such, it is considered well suited when reliability is a major concern. Program spectrum is a characterization of a program's behavior in terms of the code entities (e.g., branches, statements, functions) that are covered as the program executes. The driving idea of covrel is to complement operational profile information with white-box coverage measures based on count spectra, so as to dynamically select the most effective test cases for reliability improvement. In particular, we bias operational profile-based test selection towards those entities covered less frequently. We assess the approach by experiments with 18 versions from 4 subjects commonly used in software testing research, comparing results with traditional operational and coverage testing. Results show that exploiting operational and coverage data in a combined adaptive way actually pays in terms of reliability improvement, with covrel overcoming conventional operational testing in more than 80% of the cases. Antonia Bertolino, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 3 |
| 2017 | Optimized task allocation on private cloud for hybrid simulation of large-scale critical systems
Massimo Ficco, Beniamino Di Martino, Roberto Pietrantuono, Stefano Russo 0001 |
Future Gener. Comput. Syst. | 3 |
| 2017 | Debugging-workflow-aware software reliability growth analysisabstractSummary Software reliability growth models support the prediction/assessment of product quality, release time, and testing/debugging cost. Several software reliability growth model extensions take into account the bug correction process. However, their estimates may be significantly inaccurate when debugging fails to fully fit modelling assumptions. This paper proposes debugging‐workflow‐aware software reliability growth method (DWA‐SRGM), a method for reliability growth analysis leveraging the debugging data usually managed by companies in bug tracking systems. On the basis of a characterization of the debugging workflow within the software project under consideration (in terms of bug features and treatment phases), DWA‐SRGM pinpoints the factors impacting the estimates and to spot bottlenecks, thus supporting process improvement decisions. Two industrial case studies are presented, a customer relationship management system and an enterprise resource planning system, whose defects span a period of about 17 and 13 months, respectively. DWA‐SRGM revealed effective to obtain more realistic estimates and to capitalize on the awareness of critical factors for improving debugging. Marcello Cinque, Domenico Cotroneo, Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001 |
Softw. Test. Verification Reliab. | 4 |
| 2017 | Special Section on the 25th IEEE International Symposium on Software Reliability Engineering (ISSRE 2014)abstractThe papers in this special section were presented at the 2014 International Symposium on Software Reliability Engineering (ISSRE) that was held from 3 to 6 November 2014 in Naples, Italy. Roberto Pietrantuono, Katerina Goseva-Popstojanova, Carol S. Smidts |
IEEE Trans. Reliab. | 1 |
| 2016 | Software Aging Analysis of the Android Mobile OSabstractMobile devices are significantly complex, feature-rich, and heavily customized, thus they are prone to software reliability and performance issues. This paper considers the problem of software aging in Android mobile OS, which causes the device to gradually degrade in responsiveness, and to eventually fail. We present a methodology to identify factors (such as workloads and device configurations) and resource utilization metrics that are correlated with software aging. Moreover, we performed an empirical analysis of recent Android devices, finding that software aging actually affects them. The analysis pointed out processes and components of the Android OS affected by software aging, and metrics useful as indicators of software aging to schedule software rejuvenation actions. Domenico Cotroneo, Francesco Fucci, Antonio Ken Iannillo, Roberto Natella, Roberto Pietrantuono |
ISSRE | 5 |
| 2016 | On Adaptive Sampling-Based Testing for Software Reliability AssessmentabstractAssessing reliability of software programs during validation is a challenging task for engineers. The assessment is not only required to be unbiased, but it needs to provide tight variance (hence, tight confidence interval) with as few test cases as possible. Statistical sampling is a theoretically sound approach for reliability testing, but it is often impractical in its current form, because of too many test cases required to achieve desired confidence levels, especially when the software has few residual faults inside. We claim that the potential of statistical sampling methods is largely underestimated. This paper presents an adaptive sampling-based testing (AST) strategy for reliability assessment. A two-stage conceptual framework is defined, where adaptiveness is included to uncover residual faults earlier, while various sampling-based techniques are proposed to improve the efficiency (in terms of variance-test cases tradeoff) by better exploiting the information available to tester. An empirical study is conducted to assess the AST performance and compare the proposed sampling techniques to each other on real programs. Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 1 |
| 2016 | How do bugs surface? A comprehensive study on the characteristics of software bugs manifestation
Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001, Kishor S. Trivedi |
J. Syst. Softw. | 2 |
| 2016 | Using multi-objective metaheuristics for the optimal selection of positioning systems
Massimo Ficco, Roberto Pietrantuono, Stefano Russo 0001 |
Soft Comput. | 2 |
| 2016 | RELAI Testing: A Technique to Assess and Improve Software ReliabilityabstractTesting software to assess or improve reliability presents several practical challenges. Conventional operational testing is a fundamental strategy that simulates the real usage of the system in order to expose failures with the highest occurrence probability. However, practitioners find it unsuitable for assessing/achieving very high reliability levels; also, they do not see the adoption of a “real” usage profile estimate as a sensible idea, being it a source of non-quantifiable uncertainty. Oppositely, debug testing aims to expose as many failures as possible, but regardless of their impact on runtime reliability. These strategies are used either to assess or to improve reliability, but cannot improve and assess reliability in the same testing session. This article proposes Reliability Assessment and Improvement (RELAI) testing, a new technique thought to improve the delivered reliability by an adaptive testing scheme, while providing, at the same time, a continuous assessment of reliability attained through testing and fault removal. The technique also quantifies the impact of a partial knowledge of the operational profile. RELAI is positively evaluated on four software applications compared, in separate experiments, with techniques conceived either for reliability improvement or for reliability assessment, demonstrating substantial improvements in both cases. Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
IEEE Trans. Software Eng. | 2 |
| 2015 | Model-Driven Engineering of a Railway Interlocking SystemabstractModel-Driven Engineering (MDE) promises to enhance system development by reducing development time, and increasing productivity and quality. MDE is gaining popularity in several industry sectors, and is attractive also for critical systems where they can reduce efforts and costs for verification and validation (V&V), and can ease certification. Incorporating model-driven techniques into a legacy well-proven development cycle is not simply a matter of placing models and transformations in the design and implementation phases. We present the experience in the model-driven design and V&V of a safety-critical system in the railway domain, namely the Prolan Block, a railway interlocking system manufactured by the Hungarian company Prolan Co., required to be CENELEC SIL-4 compliant. The experience has been carried out in an industrial-academic partnership within the EU project CECRIS. We discuss the challenges and the lessons learnt in this pilot project of introducing MD design and testing techniques into the company's traditional V-model process. Fabio Scippacercola, Roberto Pietrantuono, Stefano Russo 0001, András Zentai |
MODELSWARD | 2 |
| 2015 | Defect analysis in mission-critical software systems: a detailed investigationabstractThe practice of defect analysis is recognized as an essential task for software process measurement, yet its effective application in the industrial development of large-scale software systems raises several challenges. We report the results of a study conducted at SELEX ES – a large system integrator leader in the market of software-intensive mission-critical systems. The article describes the defect analysis approach that we tailored to evaluate the software development process with respect to the quality of produced software and its relation with the required effort. Three key phases of the process were addressed, regarding the software implementation, the testing phase and the prerelease defect fixing activity, over a set of six computer software configuration items developed from 2009 to 2012 for the naval and maritime domain product line. The analysis highlighted efficiency bottlenecks in each of the monitored phases, providing company engineers with insights about room for process improvement. The implemented approach, the observed phenomena and the inferred conclusions are of support to practitioners coping with systems, development models and industrial environments similar to the considered one. Copyright © 2014 John Wiley & Sons, Ltd. Gabriella Carrozza, Roberto Pietrantuono, Stefano Russo 0001 |
J. Softw. Evol. Process. | 2 |
| 2014 | On the Impact of Debugging on Software Reliability Growth Analysis: A Case Study
Marcello Cinque, Claudio Gaiani, Daniele De Stradis, Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001 |
ICCSA (5) | 5 |
| 2014 | Reproducibility of Environment-Dependent Software Failures: An Experience ReportabstractWe investigate the dependence of software failure reproducibility on the environment in which the software is executed. The existence of such dependence is ascertained in literature, but so far it is not fully characterized. In this paper we pinpoint some of the environmental components that can affect the reproducibility of a failure and show this influence through an experimental campaign conducted on the My SQL Server software system. The set of failures of interest is drawn from My SQL's failure reports database and an experiment is designed for each of these failures. The experiments expose the influence of disk usage and level of concurrency on My SQL failure reproducibility. Furthermore, the results show that high levels of usage of these factors increase the probabilities of failure reproducibility. Davide G. Cavezza, Roberto Pietrantuono, Javier Alonso 0001, Stefano Russo 0001, Kishor S. Trivedi |
ISSRE | 2 |
| 2014 | A survey of software aging and rejuvenation studiesabstractSoftware aging is a phenomenon plaguing many long-running complex software systems, which exhibit performance degradation or an increasing failure rate. Several strategies based on the proactive rejuvenation of the software state have been proposed to counteract software aging and prevent failures. This survey article provides an overview of studies on Software Aging and Rejuvenation (SAR) that have appeared in major journals and conference proceedings, with respect to the statistical approaches that have been used to forecast software aging phenomena and to plan rejuvenation, the kind of systems and aging effects that have been studied, and the techniques that have been proposed to rejuvenate complex software systems. The analysis is useful to identify key results from SAR research, and it is leveraged in this article to highlight trends and open issues. Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2014 | Dynamic test planning: a study in an industrial context
Gabriella Carrozza, Roberto Pietrantuono, Stefano Russo 0001 |
Int. J. Softw. Tools Technol. Transf. | 2 |
| 2013 | A learning-based method for combining testing techniquesabstractThis work presents a method to combine testing techniques adaptively during the testing process. It intends to mitigate the sources of uncertainty of software testing processes, by learning from past experience and, at the same time, adapting the technique selection to the current testing session. The method is based on machine learning strategies. It uses offline strategies to take historical information into account about the techniques performance collected in past testing sessions; then, online strategies are used to adapt the selection of test cases to the data observed as the testing proceeds. Experimental results show that techniques performance can be accurately characterized from features of the past testing sessions, by means of machine learning algorithms, and that integrating this result into the online algorithm allows improving the fault detection effectiveness with respect to single testing techniques, as well as to their random combination. Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 2 |
| 2013 | Analysis and Prediction of Mandelbugs in an Industrial Software SystemabstractMandelbugs are faults that are triggered by complex conditions, such as interaction with hardware and other software, and timing or ordering of events. These faults are considerably difficult to detect with traditional testing techniques, since it can be challenging to control their complex triggering conditions in a testing environment. Therefore, it is necessary to adopt specific verification and/or fault-tolerance strategies for dealing with them in a cost-effective way. In this paper, we investigate how to predict the location of Mandelbugs in complex software systems, in order to focus V&V activities and fault tolerance mechanisms in those modules where Mandelbugs are most likely present. In the context of an industrial complex software system, we empirically analyze Mandelbugs, and investigate an approach for Mandelbug prediction based on a set of novel software complexity metrics. Results show that Mandelbugs account for a noticeable share of faults, and that the proposed approach can predict Mandelbug-prone modules with greater accuracy than the sole adoption of traditional software metrics. Gabriella Carrozza, Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ICST | 4 |
| 2013 | Fault triggers in open-source software: An experience reportabstractWith software systems becoming increasingly large and complex, many difficulties in coping with software bugs arise for developers. Despite good development practices, thorough testing, and proper maintenance policies, a non-negligible number of bugs remain in the released software. Understanding the type of residual bugs is fundamental for adopting proper countermeasures in current and future software releases. Depending on the fault triggering conditions that lead to a failure, developers can introduce fault-tolerance mechanisms and plan verification and validation strategies. In this paper, we analyze bugs in four large open-source software systems during their lifecycle, based on the concept of fault triggers. We first investigate how the type of system affects the bug type proportions, and their evolution over years. Then, an analysis of bug subtypes is performed, so as to better understand their nature, followed by a comparison with respect to attributes such as their average time to fix and severity. Domenico Cotroneo, Michael Grottke, Roberto Natella, Roberto Pietrantuono, Kishor S. Trivedi |
ISSRE | 4 |
| 2013 | Testing techniques selection based on ODC fault types and software metrics
Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 2 |
| 2013 | Predicting aging-related bugs using software complexity metrics
Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono |
Perform. Evaluation | 3 |
| 2013 | A measurement-based ageing analysis of the JVMabstractSUMMARY In this work, a software ageing analysis of Java‐based software systems is conducted. The JVM is the core layer in Java‐based systems, and its dependability greatly affects the overall system quality. Starting from an experimental campaign on a real‐world test bed, this work isolates the contribution of the JVM to the overall ageing trend, and identifies, through statistical methods, which workload parameters are the most relevant to ageing dynamics. Results revealed the presence of several ageing dynamics in the JVM, including (i) a throughput loss trend mainly dependent on the execution unit, (ii) a slow memory depletion drift due to the just‐in‐time‐compiler activity and (iii) a fast memory depletion drift caused by dynamics inside the garbage collector. The outlined procedure and obtained results are useful in order to (i) identify the presence of ageing phenomena, (ii) perform online ageing detection and time‐to‐exhaustion prediction and (iii) define optimal rejuvenation techniques. Copyright © 2011 John Wiley & Sons, Ltd. Domenico Cotroneo, Salvatore Orlando 0002, Roberto Pietrantuono, Stefano Russo 0001 |
Softw. Test. Verification Reliab. | 3 |
| 2013 | Combining Operational and Debug Testing for Improving ReliabilityabstractThis paper addresses the challenge of reliability-driven testing, i.e., of testing software systems with the specific objective of increasing its operational reliability. We first examined the most relevant approach oriented toward this goal, namely operational testing. The main issues that in the past hindered its wide-scale adoption and practical application are first discussed, followed by the analysis of its performance under different conditions and configurations. Then, a new approach conceived to overcome the limits of operational testing in delivering high reliability is proposed. The two testing strategies are evaluated probabilistically, and by simulation. Results report on the performance of operational testing when several involved parameters are taken into account, and on the effectiveness of the new proposed approach in achieving better reliability. At a higher level, the findings of the paper also suggest that a different view of the testing for reliability improvement concept may help to devise new testing approaches for high-reliability, demanding systems. Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
IEEE Trans. Reliab. | 2 |
| 2012 | On the Aging Effects Due to Concurrency Bugs: A Case Study on MySQLabstractThis study investigates software aging effects caused by the activation of concurrency bugs in a wellknown database management system (DBMS), namely MySQL. Experiments with different workloads are performed in order to reproduce the most likely conditions for concurrency bugs activation. Besides the typical aging effects observed in many operational systems (i.e., a gradual degradation over time), results highlight that both available resources and DBMS performance (e.g. service rate, service time, and connection latency) can decrease with time in a hard-to-predict way. We observed that, due to the activation of concurrency bug, the DBMS enters a degraded state in which: i) the estimation of Time-To-Failure (TTF) by means of memory depletion trend analysis is highly inaccurate, and ii) the failure rate does not depend on the instantaneous and/or mean accumulated work. Results suggest that, in such cases, finer-grained indicators and/or different techniques need to be taken into account for properly preventing failures. Antonio Bovenzi, Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 3 |
| 2011 | Workload Characterization for Software Aging AnalysisabstractThe phenomenon of software aging is increasingly recognized as a relevant problem of long-running systems. Numerous experiments have been carried out in the last decade to empirically analyze software aging. Such experiments, besides highlighting the relevance of the phenomenon, have shown that aging is tightly related to the applied workload. However, due to the differences among the experimented applications and among the experimental conditions, results of past studies are not comparable to each other. This prevent from drawing general conclusions (e.g., about the aging-workload relationship), and from comparing systems from the aging perspective. In this paper, we propose a procedure to carry out aging experiments in different applications for: i) assessing aging trend of the individual systems, as well as assessing differences among them (i.e., obtaining comparable results), ii) inferring workload-aging relationships from experiments performed on different applications, by highlighting the most relevant workload parameters. The procedure is applied, through a set of long-running experiments, to three real-scale software applications, namely Apache Web Server, James Mail Server, and CARDAMOM, a middleware for the development of air traffic control (ATC) systems. Antonio Bovenzi, Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 3 |
| 2011 | A Case Study on State-Based Robustness Testing of an Operating System for the Avionic Domain
Domenico Cotroneo, Domenico Di Leo, Roberto Natella, Roberto Pietrantuono |
SAFECOMP | 4 |
| 2011 | Criticality-Driven Component Integration in Complex Software Systems
Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001 |
SAFECOMP | 2 |
| 2010 | Software Aging Analysis of the Linux Operating SystemabstractSoftware systems running continuously for a long time tend to show degrading performance and an increasing failure occurrence rate, due to error conditions that accrue over time and eventually lead the system to failure. This phenomenon is usually referred to as Software Aging. Several long-running mission and safety critical applications have been reported to experience catastrophic aging-related failures. Software aging sources (i.e., aging-related bugs) may be hidden in several layers of a complex software system, ranging from the Operating System (OS) to the user application level. This paper presents a software aging analysis at the Operating System level, investigating software aging sources inside the Linux kernel. Linux is increasingly being employed in critical scenarios; this analysis intends to shed light on its behaviour from the aging perspective. The study is based on an experimental campaign designed to investigate the kernel internal behaviour over long running executions. By means of a kernel tracing tool specifically developed for this study, we collected relevant parameters of several kernel subsystems. Statistical analysis of collected data allowed us to confirm the presence of aging sources in Linux and to relate the observed aging dynamics to the monitored subsystems behaviour. The analysis output allowed us to infer potential sources of aging in the kernel subsystems. Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 3 |
| 2010 | Software Reliability and Testing Time Allocation: An Architecture-Based ApproachabstractWith software systems increasingly being employed in critical contexts, assuring high reliability levels for large, complex systems can incur huge verification costs. Existing standards usually assign predefined risk levels to components in the design phase, to provide some guidelines for the verification. It is a rough-grained assignment that does not consider the costs and does not provide sufficient modeling basis to let engineers quantitatively optimize resources usage. Software reliability allocation models partially address such issues, but they usually make so many assumptions on the input parameters that their application is difficult in practice. In this paper, we try to reduce this gap, proposing a reliability and testing resources allocation model that is able to provide solutions at various levels of detail, depending upon the information the engineer has about the system. The model aims to quantitatively identify the most critical components of software architecture in order to best assign the testing resources to them. A tool for the solution of the model is also developed. The model is applied to an empirical case study, a program developed for the European Space Agency, to verify model's prediction abilities and evaluate the impact of the parameter estimation errors on the prediction accuracy. Roberto Pietrantuono, Stefano Russo 0001, Kishor S. Trivedi |
IEEE Trans. Software Eng. | 1 |
| 2007 | Component airbag: a novel approach to develop dependable component-based applicationsabstractThe increasing use of "commercial off-the-shelf"(COTS) components in safety critical scenarios, arises new issues related to the "dependable" use of third-party software in such contexts. The characteristics of these components, designed for a generic use, are such to make unpredictable the effects of their use whenever they are integrated in the entire system. The author's Ph.D project aim at proposing an approach to improve dependability of COTS based application, which consists of the following phases: i) each component is stimulated by proper workloads in order to learn the failure behavior; ii) from failure behaviors, the component failure model is defined; and iii) once the failure model is known for each component, the "component airbag" is thus created, i.e. a container able of exploiting the failure model in order to monitor and prevent the component from failing. An existent literature analysis, regarding the more used dependability assessment and improvement strategies, is also presented. Roberto Pietrantuono |
ESEC/SIGSOFT FSE | 1 |