EDBT 2026 Demo / reviewers in the wild / expert
Roberto Natella
dblp:63/8166
· DBLP profile ↗
71ranked-venue papers
7as first author
36since 2021 · last 2026
0000-0003-1084-4824ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 41 · 5 first-author · 20 since 2021Security and privacy · 12 · 1 first-author · 6 since 2021Systems, architecture and hardware · 9 · 1 first-author · 3 since 2021Computer networks · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | cAPTure dataset: How fast can you detect APT threats?abstractHigh-quality datasets are essential for machine learning-based intrusion detection systems, which are considered a promising defense for cyber-physical systems against Advanced Persistent Threats (APTs). However, existing datasets often are not built to capture the long, multi-stage nature of real APT campaigns, and they are labeled as general cyber-attacks rather than explicitly as APTs. To address this gap, we propose a methodology for creating a semi-synthetic, labeled dataset that reflects the complex attack paths typical of APTs targeting cyber-physical environments. Our approach integrates realistic network traffic gathered from a real testbed with multi-step APT attack scenarios modeled on the well-established MITRE ATT&CK framework and CVE exploits repository. The cAPTure dataset provides a rich basis for evaluating intrusion detection systems, enabling an evaluation methodology that relates false positive rate and time-to-detection, two metrics that are crucial for practical, real-world NIDS deployment. Tommaso Puccetti, Simona De Vivo, Davide Zhang, Pietro Liguori, Roberto Natella, Andrea Ceccarelli |
Comput. Networks | 5 |
| 2026 | Fuzzbox: blending fuzzing into emulation for binary-only embedded targetsabstractAbstract Coverage-guided fuzzing has been widely applied to address zero-day vulnerabilities in general-purpose software and operating systems. This approach relies on instrumenting the target code at compile time. However, applying it to industrial systems remains challenging, due to proprietary and closed-source compiler toolchains and lack of access to source code. FuzzBox addresses these limitations by integrating emulation with fuzzing: it dynamically instruments code during execution in a virtualized environment, for the injection of fuzz inputs, failure detection, and coverage analysis, without requiring source code recompilation and hardware-specific dependencies. We show the effectiveness of FuzzBox through experiments in the context of a proprietary MILS (Multiple Independent Levels of Security) hypervisor for industrial applications. Additionally, we analyze the applicability of FuzzBox across commercial IoT firmware, showcasing its broad portability. Carmine Cesarano 0002, Roberto Natella |
Cybersecur. | 2 |
| 2026 | GenioSim: A novel simulation platform for edge computing over optical networksabstractThe convergence of Passive Optical Networks (PONs) and edge computing creates new opportunities: Optical Line Terminals (OLTs) and Optical Network Terminals (ONTs) can be repurposed as low-latency edge compute nodes for offloading workloads. However, exploring such design options early in the development cycle is costly and time-consuming, as prototyping requires specialized hardware and realistic traffic conditions. Simulation becomes essential, yet current tools are unable to accurately model this emerging class of systems. To address these gaps, we introduce GenioSim , a simulation platform for hierarchical PON-enabled edge infrastructures. It models OLTs and ONTs with realistic PON behavior, supports hybrid container- and VM-based virtualization, and provides multiple service and execution models. These capabilities enable the evaluation of resource management policies under complex, heterogeneous conditions. We present experiments in the context of use cases of industrial relevance, to show GenioSim can provide insights for capacity planning and for the choice of policies for container placement and task offloading in PON-enabled edge infrastructures. Carmine Cesarano 0002, Alessio Foggia, Roberto Natella |
Future Gener. Comput. Syst. | 3 |
| 2026 | Elevating Cyber Threat Intelligence against disinformation campaigns with LLM-based concept extraction and the FakeCTI datasetabstractThe swift spread of fake news and disinformation campaigns poses a significant threat to public trust, political stability, and cybersecurity. Traditional Cyber Threat Intelligence (CTI) approaches, which rely on low-level indicators such as domain names and social media handles, are easily evaded by adversaries who frequently modify their online infrastructure. To address these limitations, we introduce a novel CTI framework that focuses on high-level, semantic indicators derived from recurrent narratives and relationships of disinformation campaigns. Our approach extracts structured CTI indicators from unstructured disinformation content, capturing key entities and their contextual dependencies within fake news using Large Language Models (LLMs). We further introduce FakeCTI, the first dataset that systematically links fake news to disinformation campaigns and threat actors. To evaluate the effectiveness of our CTI framework, we analyze multiple fake news attribution techniques, spanning from traditional Natural Language Processing (NLP) to fine-tuned LLMs. This work shifts the focus from low-level artifacts to persistent conceptual structures, establishing a scalable and adaptive approach to tracking and countering disinformation campaigns. Domenico Cotroneo, Roberto Natella, Vittorio Orbinato |
J. Syst. Softw. | 2 |
| 2026 | A Survey on Failure Analysis and Fault Injection in AI SystemsabstractThe rapid advancement of AI has led to its integration into various areas, especially with Large Language Models (LLMs) significantly enhancing capabilities in Artificial Intelligence Generated Content (AIGC). However, the complexity of AI systems has also exposed their vulnerabilities, necessitating robust methods for Failure Analysis (FA) and Fault Injection (FI) to ensure resilience and reliability. Despite the importance of these techniques, there lacks a comprehensive review of FA and FI methodologies in AI systems. This study fills this gap by presenting a detailed survey of existing FA and FI approaches across six layers of AI systems. We systematically analyze 142 studies to answer three research questions including (1) what are the prevalent failures in AI systems, (2) what types of faults can current FI tools simulate, (3) what gaps exist between the simulated faults and real-world failures. Our findings reveal a taxonomy of AI system failures, assess the capabilities of existing FI tools, and highlight discrepancies between real-world and simulated failures. Moreover, this survey contributes to the field by providing a framework for fault diagnosis, evaluating the state-of-the-art in FI, and identifying areas for improvement in FI techniques to enhance the resilience of AI systems. Guangba Yu, Gou Tan, Haojia Huang, Pengfei Chen 0002, Roberto Natella, Zibin Zheng, Michael R. Lyu |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2026 | Aging-Related Bug Prediction Based on Multi-View Graph Feature Learning and Graph-TransformerabstractSoftware aging, characterized by an increasing failure rate or performance degradation in long-running software systems, poses significant risks, including substantial financial losses and potential threats to human lives. This phenomenon is primarily driven by the accumulation of runtime errors, commonly referred to as aging-related bugs (ARBs). Aging-related bug prediction (ARBP) has been proposed to facilitate the detection and remediation of ARBs prior to software release. However, ARBP’s effectiveness heavily depends on the quality of dataset features used. Previous research has largely relied on a standard set of manually designed metrics, often overlooking that these metrics may fail to distinguish between code segments with different semantics, even when they exhibit identical metric values. While some studies have attempted to develop models that learn semantic features from source code, they typically focus on token-level or graph-level features, neglecting a comprehensive exploration of ARB characteristics within the source code. Specifically, there is insufficient discussion on whether deep semantic features can adequately capture the essential traits that trigger aging phenomena. In this paper, we propose a novel multi-view graph feature learning framework based on Graph-Transformer, which integrates newly proposed ARB features extracted from Abstract Syntax Trees with Code Property Graphs for feature learning. Our approach effectively captures hierarchical structures and variable dependencies, facilitating the identification of complex interactions that contribute to ARBs. Additionally, we implement sub-graph sampling and class imbalance strategies to enhance model performance. Experimental results across three datasets demonstrate that our method surpasses state-of-the-art approaches, a code property graph-based feature extraction method (specifically SGT), achieving precision improvements of 8.2% on Linux, 15.4% on MySQL, and 2.5% on NetBSD, thereby establishing a new benchmark for ARB prediction. Jianwen Xiang, Roberto Natella, Roberto Pietrantuono, Domenico Cotroneo |
IEEE Trans. Software Eng. | 6 |
| 2025 | KubeFence: Security Hardening of the Kubernetes Attack SurfaceabstractKubernetes (K8s) is widely used to orchestrate containerized applications, including critical services in domains such as finance, healthcare, and government. However, its extensive and feature-rich API interface exposes a broad attack surface, making K8s vulnerable to exploits of software vulnerabilities and misconfigurations. Even if K8s adopts role-based access control (RBAC) to manage access to K8s APIs, this approach lacks the granularity needed to protect specification attributes within API requests. This paper proposes a novel solution, KubeFence, which implements finer-grain API filtering tailored to specific client workloads. KubeFence analyzes Kubernetes Operators from trusted repositories and leverages their configuration files to restrict unnecessary features of the K8s API, to mitigate misconfigurations and vulnerabilities exploitable through the K8s API. The experimental results show that KubeFence can significantly reduce the attack surface and prevent attacks compared to RBAC. Carmine Cesarano 0002, Roberto Natella |
DSN | 2 |
| 2025 | Performability Management of 5G Service Chains with Rejuvenation: The Open5GS Use CaseabstractThis paper presents a stochastic framework for managing the performability (performance and availability) of 5G-based service function chains (SFCs). By integrating an$M / G / m$queueing model for latency estimation and Stochastic Reward Networks (SRNs) for availability assessment, we evaluate the impact of software rejuvenation on 5 G network performability. The final goal is to derive the optimal 5G setting that meets both performance (e.g., delay threshold) and availability (e.g., the “five nines”). Our testbed, based on Open5GS, validates the model and provides insights into optimal 5G settings that balance performance, availability, and resource utilization. Luigi De Simone, Mario Di Mauro, Maurizio Longo, Roberto Natella, Fabio Postiglione |
NetSoft | 4 |
| 2025 | Creation and Use of a Representative Dataset for Advanced Persistent Threats Detection
Tommaso Puccetti, Simona De Vivo, Davide Zhang, Pietro Liguori, Roberto Natella, Andrea Ceccarelli |
SAFECOMP | 5 |
| 2025 | Evaluation of Systems Programming Exercises through Tailored Static AnalysisabstractIn large programming classes, it takes a significant effort from teachers to evaluate exercises and provide detailed feedback. In systems programming, test cases are not sufficient to assess exercises, since concurrency and resource management bugs are difficult to reproduce. This paper presents an experience report on static analysis for the automatic evaluation of systems programming exercises. We design systems programming assignments with static analysis rules that are tailored for each assignment, to provide detailed and accurate feedback. Our evaluation shows that static analysis can identify a significant number of erroneous submissions missed by test cases. Roberto Natella |
SIGCSE (1) | 1 |
| 2025 | Enhancing robustness of AI offensive code generators via data augmentation
Cristina Improta, Pietro Liguori, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
Empir. Softw. Eng. | 3 |
| 2024 | Enhancing AI-based Generation of Software Exploits with Contextual InformationabstractThis practical experience report explores Neural Machine Translation (NMT) models’ capability to generate offensive security code from natural language (NL) descriptions, highlighting the significance of contextual understanding and its impact on model performance. Our study employs a dataset comprising real shellcodes to evaluate the models across various scenarios, including missing information, necessary context, and unnecessary context. The experiments are designed to assess the models’ resilience against incomplete descriptions, their proficiency in leveraging context for enhanced accuracy, and their ability to discern irrelevant information. The findings reveal that the introduction of contextual data significantly improves performance. However, the benefits of additional context diminish beyond a certain point, indicating an optimal level of contextual information for model training. Moreover, the models demonstrate an ability to filter out unnecessary context, maintaining high levels of accuracy in the generation of offensive security code. This study paves the way for future research on optimizing context use in AI-driven code generation, particularly for applications requiring a high degree of technical precision such as the generation of offensive code. Pietro Liguori, Cristina Improta, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
ISSRE | 3 |
| 2024 | Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning AttacksabstractAI-based code generators have become pivotal in assisting developers in writing software starting from natural language (NL). However, they are trained on large amounts of data, often collected from unsanitized online sources (e.g., GitHub, HuggingFace). As a consequence, AI models become an easy target for data poisoning, i.e., an attack that injects malicious samples into the training data to generate vulnerable code. Domenico Cotroneo, Cristina Improta, Pietro Liguori, Roberto Natella |
ICPC | 4 |
| 2024 | DRACO: Distributed Resource-aware Admission Control for large-scale, multi-tier systemsabstractModern distributed systems are designed to manage overload conditions, by throttling the traffic in excess that cannot be served through overload control techniques. However, the adoption of large-scale NoSQL datastores make systems vulnerable to unbalanced overloads, where specific datastore nodes are overloaded because of hot-spot resources and hogs. In this paper, we propose DRACO, a novel overload control solution that is aware of data dependencies between the application and the datastore tiers. DRACO performs selective admission control of application requests, by only dropping the ones that map to resources on overloaded datastore nodes, while achieving high resource utilization on non-overloaded datastore nodes. We evaluate DRACO on two case studies with high availability and performance requirements, a virtualized IP Multimedia Subsystem and a distributed fileserver. Results show that the solution can achieve high performance and resource utilization even under extreme overload conditions, up to 100x the engineered capacity. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
J. Parallel Distributed Comput. | 2 |
| 2024 | Automating the correctness assessment of AI-generated code for security contextsabstractEvaluating the correctness of code generated by AI is a challenging open problem. In this paper, we propose a fully automated method, named ACCA, to evaluate the correctness of AI-generated code for security purposes. The method uses symbolic execution to assess whether the AI-generated code behaves as a reference implementation. We use ACCA to assess four state-of-the-art models trained to generate security-oriented assembly code and compare the results of the evaluation with different baseline solutions, including output similarity metrics, widely used in the field, and the well-known ChatGPT, the AI-powered language model developed by OpenAI. Our experiments show that our method outperforms the baseline solutions and assesses the correctness of the AI-generated code similar to the human-based evaluation, which is considered the ground truth for the assessment in the field. Moreover, ACCA has a very strong correlation with the human evaluation (Pearson’s correlation coefficient r=0.84 on average). Finally, since it is a full y automated solution that does not require any human intervention, the proposed method performs the assessment of every code snippet in ∼0.17 s on average, which is definitely lower than the average time required by human analysts to manually inspect the code, based on our experience. Domenico Cotroneo, Alessio Foggia, Cristina Improta, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 5 |
| 2024 | SGT: Aging-related bug prediction via semantic feature learning based on graph-transformer
Jianwen Xiang, Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono |
J. Syst. Softw. | 6 |
| 2024 | Laccolith: Hypervisor-Based Adversary Emulation With Anti-DetectionabstractAdvanced Persistent Threats (APTs) represent the most threatening form of attack nowadays since they can stay undetected for a long time. Adversary emulation is a proactive approach for preparing against these attacks. However, adversary emulation tools lack the anti-detection abilities of APTs. We introduce Laccolith, a hypervisor-based solution for adversary emulation with anti-detection to fill this gap. We also present an experimental study to compare Laccolith with MITRE CALDERA, a state-of-the-art solution for adversary emulation, against five popular anti-virus products. We found that CALDERA cannot evade detection, limiting the realism of emulated attacks, even when combined with a state-of-the-art anti-detection framework. Our experiments show that Laccolith can hide its activities from all the tested anti-virus products, thus making it suitable for realistic emulations. Vittorio Orbinato, Marco Carlo Feliciano, Domenico Cotroneo, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Performance and Availability Challenges in Designing Resilient 5G ArchitecturesabstractThis work proposes a stochastic characterization of resilient 5G architectures, where attributes such as performance and availability play a crucial role. As regards performance, we focus on the delay associated with the Packet Data Unit session establishment, a 5G procedure recognized as critical for its impact on the Quality of Service and Experience of end-users. To formally characterize this aspect, we employ the non-product-form queueing networks framework where: i) main nodes of a 5G architecture have been realistically modeled as G/G/m queues which do not admit analytical solutions; ii) the decomposition method useful to catch subtle quantities involved in the chain of 5G interconnected nodes has been conveniently customized. The results of performance characterization constitute the input of the availability modeling, where we design a hierarchical scheme to characterize the probabilistic failure/repair behavior of 5G nodes combining two formalisms: i) the Reliability Block Diagrams, useful to capture the high-level interconnections between nodes; ii) the Stochastic Reward Networks to model the internal structure of each node. The final result is an optimal resilient 5G setting that fulfills both a performance constraint (e.g., a temporal threshold) and an availability constraint (e.g., the so-called five nines) at the minimum cost, namely, with the smallest number of redundant elements. The theoretical part is complemented by an empirical assessment carried out through Open5GS, a 5G testbed that we have deployed to realistically estimate main performance and availability metrics. Luigi De Simone, Mario Di Mauro, Roberto Natella, Fabio Postiglione |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | IFCM: An improved Fuzzy C-means clustering method to handle Class Overlap on Aging-related Software Bug PredictionabstractSoftware aging refers to a problem of performance decay in long-running software systems. This phenomenon is primarily attributed to the accumulation of run-time errors, commonly known as aging-related bugs (ARBs). Detecting ARBs through Aging-related Bug Prediction (ARBP) is crucial in ensuring system reliability. The effectiveness of ARBP heavily relies on the quality of datasets. However, ARB datasets often suffer from class overlap, where instances from different classes exhibit similar feature values. Class overlap poses a significant challenge as it compromises the quality of training data and subsequently impacts ARBP accuracy. To address this issue, we propose an improved Fuzzy C-means clustering method named IFCM, designed to mitigate class overlap in ARBP tasks. IFCM can identify whether an instance occurs overlap, and identify the overlap degree of this instance through the predefined parameters. We evaluate our proposed method on two public datasets Linux and MySQL and one self-collected dataset NetBSD using five different classifiers with five performance metrics (AUC, F1, Balance, PD, PF). Comparison with four existing methods (No clean, NCL, IKMCCA, ROCT) demonstrates that IFCM is effective in alleviating class overlap in ARBP. For Instance, IFCM achieves promising results in terms of AUC blue (which are 0.762, 0.757, and 0.642) and Balance (which are 0.709, 0.736, and 0.595) at the dataset level. Shuo Feng 0003, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang, Roberto Pietrantuono, Roberto Natella, Domenico Cotroneo |
ISSRE | 7 |
| 2023 | Who evaluates the evaluators? On automatic metrics for assessing AI-based offensive code generatorsabstractAI-based code generators are an emerging solution for automatically writing programs starting from descriptions in natural language, by using deep neural networks (Neural Machine Translation, NMT). In particular, code generators have been used for ethical hacking and offensive security testing by generating proof-of-concept attacks. Unfortunately, the evaluation of code generators still faces several issues. The current practice uses output similarity metrics, i.e., automatic metrics that compute the textual similarity of generated code with ground-truth references. However, it is not clear what metric to use, and which metric is most suitable for specific contexts. This work analyzes a large set of output similarity metrics on offensive code generators. We apply the metrics on two state-of-the-art NMT models using two datasets containing offensive assembly and Python code with their descriptions in the English language. We compare the estimates from the automatic metrics with human evaluation and provide practical insights into their strengths and limitations. Pietro Liguori, Cristina Improta, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
Expert Syst. Appl. | 3 |
| 2023 | Run-time failure detection via non-intrusive event analysis in a large-scale cloud computing platformabstractCloud computing systems fail in complex and unforeseen ways due to unexpected combinations of events and interactions among hardware and software components. These failures are especially problematic when they are silent, i.e., not accompanied by any explicit failure notification, hindering the timely detection and recovery. In this work, we propose an approach to run-time failure detection tailored for monitoring multi-tenant and concurrent cloud computing systems. The approach uses a non-intrusive form of event tracing, without manual changes to the system’s internals to propagate session identifiers (IDs), and builds a set of lightweight monitoring rules from fault-free executions. We evaluated the effectiveness of the approach in detecting failures in the context of the OpenStack cloud computing platform, a complex and “off-the-shelf” distributed system, by executing a campaign of fault injection experiments in a multi-tenant scenario. Our experiments show that the approach detects the failure with an F1 score (0.85) and accuracy (0.77) higher than the ones provided by the OpenStack failure logging mechanisms (0.53 and 0.50) and two non-session-aware run-time verification approaches (both lower than 0.15). Moreover, the approach significantly decreases the average time to detect failures at run-time (∼114 seconds) compared to the OpenStack logging mechanisms. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 4 |
| 2023 | Multi-Provider IMS Infrastructure With Controlled Redundancy: A Performability EvaluationabstractIn modern telecommunication networks, services are provided through Service Function Chains (SFC), where network resources are implemented by leveraging virtualization and containerization technologies. In particular, the possibility of easily adding or removing network resources has prompted service providers to redefine some concepts including performance and availability. In line with this new trend, we propose a performability study of a multi-provider containerized IP Multimedia Subsystem (cIMS), an SFC-like infrastructure used in the core part of 4G/5G networks to handle multimedia sessions. On the one hand, performance issues are tackled by modeling each cIMS node in terms of a G/G/m queueing system to derive the Call Setup Delay (CSD), a performance metric related to the user-end experience in multimedia communications. On the other hand, availability issues are addressed through the Multi-State System (MSS) formalism, to take into account different performance rates of the system. Then, we devise an algorithm called PE-MUGF (Performability Evaluation through Multidimensional Universal Generating Function) to identify the minimum-redundancy cIMS configuration which meets given performance and availability targets at the same time. Finally, an extensive experimental analysis based on Clearwater, a containerized IMS testbed, allows us to estimate most of system parameters whose robustness is evaluated through a sensitivity analysis. Luigi De Simone, Mario Di Mauro, Maurizio Longo, Roberto Natella, Fabio Postiglione |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | A Latency-Driven Availability Assessment for Multi-Tenant Service ChainsabstractNowadays, most telecommunication services adhere to the Service Function Chain (SFC) paradigm, where network functions are implemented via software. In particular, container virtualization is becoming a popular approach to deploy network functions and to enable resource slicing among several tenants. The resulting infrastructure is a complex system composed by a huge amount of containers implementing different SFC functionalities, along with different tenants sharing the same chain. The complexity of such a scenario lead us to evaluate two critical metrics: the steady-state availability (the probability that a system is functioning in long runs) and the latency (the time between a service request and the pertinent response). Consequently, we propose a latency-driven availability assessment for multi-tenant service chains implemented via Containerized Network Functions (CNFs). We adopt a multi-state system to model single CNFs and the queueing formalism to characterize the service latency. To efficiently compute the availability, we develop a modified version of the Multidimensional Universal Generating Function (MUGF) technique. Finally, we solve an optimization problem to minimize the SFC cost under an availability constraint. As a relevant example of SFC, we consider a containerized version of IP Multimedia Subsystem, whose parameters have been estimated through fault injection techniques and load tests. Luigi De Simone, Mario Di Mauro, Roberto Natella, Fabio Postiglione |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Performability Assessment of Containerized Multi-Tenant IMS through Multidimensional UGFabstractWe advance a performability assessment of a multi- tenant containerized IP Multimedia Subsystem (cIMS), i.e.: one and the same infrastructure is shared among different providers (or tenants). Specifically, we: i) model each cIMS node (a.k.a. Containerized Network Function - CNF) through the Multi-State System (MSS) formalism to capture the dimensionality of the multi-tenant arrangement, and characterize each tenant through queueing theory attributes to catch latency-dependent performance aspects; ii) afford an availability analysis of cIMS by means of an extended version of the Universal Generating Function (UGF) technique, dubbed Multidimensional UGF (MUGF); iii) solve an optimization problem to retrieve the cIMS deployment minimizing costs while guaranteeing high availability requirements. The whole assessment is supported by an experiment based on the containerized IMS platform Clearwater which we deploy to derive some realistic system parameters by means of fault injection techniques. Luigi De Simone, Mario Di Mauro, Maurizio Longo, Roberto Natella, Fabio Postiglione |
CNSM | 4 |
| 2022 | SlowCoach: Mutating Code to Simulate Performance BugsabstractPerformance bugs are unnecessarily inefficient code chunks in software codebases that cause prolonged execution times and degraded computational resource utilization. For performance bug diagnostics, tools that aid in the identification of said bugs, such as benchmarks and profilers, are commonly employed. However, due to factors such as insufficient workloads or ineffective benchmarks, software defects related to code inefficiencies are inherently difficult to diagnose. Hence, the capabilities of performance bug diagnostic tools are limited and performance bug instances may be missed. Traditional mutation testing (MT) is a technique for quantifying a test suite's ability to find functional bugs by mutating the code of the test subject. Similarly, we adopt performance mutation testing (PMT) to evaluate performance bug diagnostic tools and identify where improvements need to be made to a performance testing methodology. We carefully investigate the different performance bug fault models and how synthesized performance bugs based on these models can evaluate benchmarks and workload selection to help improve performance diagnostics. In this paper, we present the design of our PMT framework, SLOWCOACH, and evaluate it with over 1600 mutants from 4 real-world software projects. Oliver Schwahn, Roberto Natella, Matthew Bradbury, Neeraj Suri |
ISSRE | 3 |
| 2022 | Automatic Mapping of Unstructured Cyber Threat Intelligence: An Experimental Study: (Practical Experience Report)abstractProactive approaches to security, such as adversary emulation, leverage information about threat actors and their techniques (Cyber Threat Intelligence, CTI). However, most CTI still comes in unstructured forms (i.e., natural language), such as incident reports and leaked documents. To support proactive security efforts, we present an experimental study on the automatic classification of unstructured CTI into attack techniques using machine learning (ML). We contribute with two new datasets for CTI analysis, and we evaluate several ML models, including both traditional and deep learning-based ones. We present several lessons learned about how ML can perform at this task, which classifiers perform best and under which conditions, which are the main causes of classification errors, and the challenges ahead for CTI analysis. Vittorio Orbinato, Mariarosaria Barbaraci, Roberto Natella, Domenico Cotroneo |
ISSRE | 3 |
| 2022 | Can we generate shellcodes via natural language? An empirical studyabstractAbstract Writing software exploits is an important practice for offensive security analysts to investigate and prevent attacks. In particular, shellcodes are especially time-consuming and a technical challenge, as they are written in assembly language. In this work, we address the task of automatically generating shellcodes, starting purely from descriptions in natural language, by proposing an approach based on Neural Machine Translation (NMT). We then present an empirical study using a novel dataset ( Shellcode_IA32 ), which consists of 3200 assembly code snippets of real Linux/x86 shellcodes from public databases, annotated using natural language. Moreover, we propose novel metrics to evaluate the accuracy of NMT at generating shellcodes. The empirical analysis shows that NMT can generate assembly code snippets from the natural language with high accuracy and that in many cases can generate entire shellcodes with no errors. Pietro Liguori, Erfan Al-Hossami, Domenico Cotroneo, Roberto Natella, Bojan Cukic, Samira Shaikh |
Autom. Softw. Eng. | 4 |
| 2022 | StateAFL: Greybox fuzzing for stateful network serversabstractAbstract Fuzzing network servers is a technical challenge, since the behavior of the target server depends on its state over a sequence of multiple messages. Existing solutions are costly and difficult to use, as they rely on manually-customized artifacts such as protocol models, protocol parsers, and learning frameworks. The aim of this work is to develop a greybox fuzzer ( StateAFL ) for network servers that only relies on lightweight analysis of the target program, with no manual customization, in a similar way to what the AFL fuzzer achieved for stateless programs. The proposed fuzzer instruments the target server at compile-time, to insert probes on memory allocations and network I/O operations. At run-time, it infers the current protocol state of the target server by taking snapshots of long-lived memory areas, and by applying a fuzzy hashing algorithm (Locality-Sensitive Hashing) to map memory contents to a unique state identifier. The fuzzer incrementally builds a protocol state machine for guiding fuzzing. We implemented and released StateAFL as open-source software. As a basis for reproducible experimentation, we integrated StateAFL with a large set of network servers for popular protocols, with no manual customization to accomodate for the protocol. The experimental results show that the fuzzer can be applied with no manual customization on a large set of network servers for popular protocols, and that it can achieve comparable, or even better code coverage and bug detection than customized fuzzing. Moreover, our qualitative analysis shows that states inferred from memory better reflect the server behavior than only using response codes from messages. Roberto Natella |
Empir. Softw. Eng. | 1 |
| 2022 | ThorFI: a Novel Approach for Network Fault Injection as a Service
Domenico Cotroneo, Luigi De Simone, Roberto Natella |
J. Netw. Comput. Appl. | 3 |
| 2022 | Software micro-rejuvenation for Android mobile systems
Domenico Cotroneo, Luigi De Simone, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 3 |
| 2022 | Fault Injection Analytics: A Novel Approach to Discover Failure Modes in Cloud-Computing SystemsabstractCloud computing systems fail in complex and unexpected ways due to unexpected combinations of events and interactions between hardware and software components. Fault injection is an effective means to bring out these failures in a controlled environment. However, fault injection experiments produce massive amounts of data, and manually analyzing these data is inefficient and error-prone, as the analyst can miss severe failure modes that are yet unknown. This article introduces a new paradigm (fault injection analytics) that applies unsupervised machine learning on execution traces of the injected system, to ease the discovery and interpretation of failure modes. We evaluated the proposed approach in the context of fault injection experiments on the OpenStack cloud computing platform, where we show that the approach can accurately identify failure modes with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | EVIL: Exploiting Software via Natural LanguageabstractWriting exploits for security assessment is a challenging task. The writer needs to master programming and obfuscation techniques to develop a successful exploit. To make the task easier, we propose an approach (EVIL) to automatically generate exploits in assembly/Python language from descriptions in natural language. The approach leverages Neural Machine Translation (NMT) techniques and a dataset that we developed for this work. We present an extensive experimental study to evaluate the feasibility of EVIL, using both automatic and manual analysis, and both at generating individual statements and entire exploits. The generated code achieved high accuracy in terms of syntactic and semantic correctness. Pietro Liguori, Erfan Al-Hossami, Vittorio Orbinato, Roberto Natella, Samira Shaikh, Domenico Cotroneo, Bojan Cukic |
ISSRE | 4 |
| 2021 | ProFuzzBench: a benchmark for stateful protocol fuzzingabstractWe present a new benchmark (ProFuzzBench) for stateful fuzzing of network protocols. The benchmark includes a suite of representative open-source network servers for popular protocols, and tools to automate experimentation. We discuss challenges and potential directions for future research based on this benchmark. Roberto Natella, Van-Thuan Pham |
ISSTA | 1 |
| 2021 | Timing covert channel analysis of the VxWorks MILS embedded hypervisor under the common criteria security certification
Domenico Cotroneo, Luigi De Simone, Roberto Natella |
Comput. Secur. | 3 |
| 2021 | Enhancing the analysis of software failures in cloud computing systems with deep learning
Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 4 |
| 2021 | Dependability Assessment of the Android OS Through Fault InjectionabstractThe reliability of mobile devices is a challenge for vendors, since the mobile software stack has significantly grown in complexity. In this article, we study how to assess the impact of faults on the quality of user experience in the Android mobile OS through fault injection. We first address the problem of identifying a realistic fault model for the Android OS, by providing to developers a set of lightweight and systematic guidelines for fault modeling. Then, we present an extensible fault injection tool (AndroFIT) to apply such fault model on actual, commercial Android devices. Finally, we present a large fault injection experimentation on three Android products from major vendors, and point out several reliability issues and opportunities for improving the Android OS. Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella, Stefano Rosiello |
IEEE Trans. Reliab. | 3 |
| 2020 | ProFIPy: Programmable Software Fault Injection as-a-ServiceabstractIn this paper, we present a new fault injection tool (ProFIPy) for Python software. The tool is designed to be programmable, in order to enable users to specify their software fault model, using a domain-specific language (DSL) for fault injection. Moreover, to achieve better usability, ProFIPy is provided as software-as-a-service and supports the user through the configuration of the faultload and workload, failure data analysis, and full automation of the experiments using container- based virtualization and parallelization. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
DSN | 4 |
| 2020 | Dependability Evaluation of Middleware Technology for Large-scale Distributed CachingabstractDistributed caching systems (e.g., Memcached) are widely used by service providers to satisfy accesses by millions of concurrent clients. Given their large-scale, modern distributed systems rely on a middleware layer to manage caching nodes, to make applications easier to develop, and to apply load balancing and replication strategies. In this work, we performed a dependability evaluation of three popular middleware platforms, namely Twemproxy by Twitter, Mcrouter by Facebook, and Dynomite by Netflix, to assess availability and performance under faults, including failures of Memcached nodes and congestion due to unbalanced workloads and network link bandwidth bottlenecks. We point out the different availability and performance trade-offs achieved by the three platforms, and scenarios in which few faulty components cause cascading failures of the whole distributed system. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
ISSRE | 2 |
| 2020 | A comprehensive study on software aging across android versions and vendors
Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella, Roberto Pietrantuono |
Empir. Softw. Eng. | 3 |
| 2020 | Special issue: ISSRE 2018, the 29th IEEE International Symposium on Software Reliability EngineeringabstractThis special issue contains extended versions of five papers from the 29th IEEE International Symposium on Software Reliability Engineering (ISSRE 2018). ISSRE is focused on innovative techniques and tools for assessing, predicting, and improving the reliability, safety, and security of software products. The symposium emphasizes scientific methods, industrial relevance, rigorous empirical validation, and shared value of practical tools and experiences. ISSRE boasts a large industry participation, with authors and participants from international corporations. Based on the reviews from the programme committee members and discussions with the editors-in-chief regarding the relevance of the papers to the journal's topics of interest, we invited the authors of seven papers to extend their work and submit to this special issue. The extended papers went through several rounds of revision during the rigorous peer-review process. The papers were reviewed by a panel of experts that included, but was not limited to, members of the ISSRE 2018 Program Committee. Five papers successfully completed the review process and are included in this special issue. The first paper, Using Mutants to Help Developers Distinguish and Debug (Compiler) Faults by Josie Holmes and Alex Groce, introduces a distance metric for failing test cases based on the intuition that failing tests that kill the same mutants are likely related to the same fault. This issue is especially relevant for very large test suites, as in the ‘compiler fuzzer taming’ problem. The paper evaluates the metric on two widely used real-world compilers by combining the metric with state-of-the-art methods for fault identification and localization. The second paper, Testing Microservice Architectures for Operational Reliability by Roberto Pietrantuono, Stefano Russo, and Antonio Guerriero, proposes a method for quantitatively assessing the probability of failures (‘operational reliability’) in the context of microservice applications, where the usage profile changes often for reasons such as frequent releases. The method achieves significant improvements in terms of accuracy and efficiency of reliability assessment on three open-source applications. The third paper, Model-based Hypothesis Testing of Uncertain Software Systems by Matteo Camilli, Angelo Gargantini, and Patrizia Scandurra, presents a methodology for combining model-based testing with Bayesian reasoning for testing systems with stochastic QoS properties using a model with uncertain parameters. The paper provides a detailed and reproducible case study for demonstrating the methodology. The fourth paper, Fully Automated HTML and Javascript Rewriting for Constructing a Self-healing Web Proxy by Thomas Durieux, Youssef Hamadi, and Martin Monperrus, applies the failure-oblivious computing principle to web applications. Errors are masked through HTML and Javascript code rewriting (e.g., to skip the faulty line) with an HTTP proxy and a browser extension, respectively. The approach is empirically evaluated on a large, publicly available data set of reproducible Javascript errors. A significant share of errors can be automatically self-healed with this simple strategy. The fifth paper, Facilitating Program Performance Profiling via Evolutionary Symbolic Execution by Andrea Aquino, Pietro Braione, Giovanni Denaro, and Pasquale Salza, pursues performance programme profiling by using symbolic execution and evolutionary algorithms to find worst-case execution paths. The paper shows that the combination of these two techniques significantly improves the effectiveness of the search. We express our thanks to the people who contributed to the success of ISSRE 2018 and to this special issue. We would like to express our gratitude to the authors for extending their papers and submitting their valuable work in the area of software reliability engineering and to the panel of experts for providing detailed and constructive reviews to the authors and assuring the quality of the papers. We are also grateful to the ISSRE steering committee, organizing committee, programme committee, and programme board. In particular, we thank the general chair, Bojan Cukic, whose support continued even after the conference and made this special issue possible. Finally, we would like to thank Robert Hierons and Jeff Offutt for their support, guidance, and enthusiasm for this special issue. Roberto Natella, Sudipto Ghosh 0001 |
Softw. Test. Verification Reliab. | 1 |
| 2020 | Analyzing the Effects of Bugs on Software InterfacesabstractCritical systems that integrate software components (e.g., from third-parties) need to address the risk of residual software defects in these components. Software fault injection is an experimental solution to gauge such risk. Many error models have been proposed for emulating faulty components, such as by injecting error codes and exceptions, or by corrupting data with bit-flips, boundary values, and random values. Even if these error models have been able to find breaches in fragile systems, it is unclear whether these errors are in fact representative of software faults. To pursue this open question, we propose a methodology to analyze how software faults in C/C++ software components turn into errors at components' interfaces (interface error propagation), and present an experimental analysis on what, where, and when to inject interface errors. The results point out that the traditional error models, as used so far, do not accurately emulate software faults, but that richer interface errors need to be injected, by: injecting both fail-stop behaviors and data corruptions; targeting larger amounts of corrupted data structures; emulating silent data corruptions not signaled by the component; combining bit-flips, boundary values, and data perturbations. Roberto Natella, Stefan Winter 0001, Domenico Cotroneo, Neeraj Suri |
IEEE Trans. Software Eng. | 1 |
| 2019 | Analyzing the Context of Bug-Fixing Changes in the OpenStack Cloud Computing PlatformabstractMany research areas in software engineering, such as mutation testing, automatic repair, fault localization, and fault injection, rely on empirical knowledge about recurring bug-fixing code changes. Previous studies in this field focus on what has been changed due to bug-fixes, such as in terms of code edit actions. However, such studies did not consider where the bug-fix change was made (i.e., the context of the change), but knowing about the context can potentially narrow the search space for many software engineering techniques (e.g., by focusing mutation only on specific parts of the software). Furthermore, most previous work on bug-fixing changes focused on C and Java projects, but there is little empirical evidence about Python software. Therefore, in this paper we perform a thorough empirical analysis of bug-fixing changes in three OpenStack projects, focusing on both the what and the where of the changes. We observed that all the recurring change patterns are not oblivious with respect to the surrounding code, but tend to occur in specific code contexts. Domenico Cotroneo, Luigi De Simone, Antonio Ken Iannillo, Roberto Natella, Stefano Rosiello, Nematollah Bidokhti |
ISSRE | 4 |
| 2019 | Enhancing Failure Propagation Analysis in Cloud Computing SystemsabstractIn order to plan for failure recovery, the designers of cloud systems need to understand how their system can potentially fail. Unfortunately, analyzing the failure behavior of such systems can be very difficult and time-consuming, due to the large volume of events, non-determinism, and reuse of third-party components. To address these issues, we propose a novel approach that joins fault injection with anomaly detection to identify the symptoms of failures. We evaluated the proposed approach in the context of the OpenStack cloud computing platform. We show that our model can significantly improve the accuracy of failure analysis in terms of false positives and negatives, with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ISSRE | 4 |
| 2019 | How bad can a bug get? an empirical analysis of software failures in the OpenStack cloud computing platformabstractCloud management systems provide abstractions and APIs for programmatically configuring cloud infrastructures. Unfortunately, residual software bugs in these systems can potentially lead to high-severity failures, such as prolonged outages and data losses. In this paper, we investigate the impact of failures in the context widespread OpenStack cloud management system, by performing fault injection and by analyzing the impact of the resulting failures in terms of fail-stop behavior, failure detection through logging, and failure propagation across components. The analysis points out that most of the failures are not timely detected and notified; moreover, many of these failures can silently propagate over time and through components of the cloud management system, which call for more thorough run-time checks and fault containment. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ESEC/SIGSOFT FSE | 4 |
| 2019 | Evolutionary Fuzzing of Android OS Vendor System Services
Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella |
Empir. Softw. Eng. | 3 |
| 2019 | Overload control for virtual network functions under CPU contention
Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
Future Gener. Comput. Syst. | 2 |
| 2018 | Faultprog: Testing the Accuracy of Binary-Level Software Fault InjectionabstractOff-The-Shelf (OTS) software components are the cornerstone of modern systems, including safety-critical ones. However, the dependability of OTS components is uncertain due to the lack of source code, design artifacts and test cases, since only their binary code is supplied. Fault injection in components' binary code is a solution to understand the risks posed by buggy OTS components. In this paper, we consider the problem of the accurate mutation of binary code for fault injection purposes. Fault injection emulates bugs in high-level programming constructs (assignments, expressions, function calls, ...) by mutating their translation in binary code. However, the semantic gap between the source code and its binary translation often leads to inaccurate mutations. We propose Faultprog, a systematic approach for testing the accuracy of binary mutation tools. Faultprog automatically generates synthetic programs using a stochastic grammar, and mutates both their binary code with the tool under test, and their source code as reference for comparisons. Moreover, we present a case study on a commercial binary mutation tool, where Faultprog was adopted to identify code patterns and compiler optimizations that affect its mutation accuracy. Domenico Cotroneo, Anna Lanzaro, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2018 | Run-Time Detection of Protocol Bugs in Storage I/O Device DriversabstractProtocol violation bugs in storage device drivers are a critical threat for data integrity, since these bugs can silently corrupt the commands and data flowing between the OS and storage devices. Due to their nature, these bugs are notoriously difficult to find by traditional testing. In this paper, we propose a run-time monitoring approach for storage device drivers, in order to detect I/O protocol violations that would otherwise silently escalate in corruptions of users' data. The monitoring approach detects violations of I/O protocols by automatically learning a reference model from failure-free execution traces. The approach focuses on selected portions of the storage controller interface, in order to achieve a good tradeoff in terms of low performance overhead and high coverage and accuracy of failure detection. We assess these properties on three real-world storage device drivers from the Linux kernel, through fault injection and stress tests. Moreover, we show that the monitoring approach only requires few minutes of training workload, and that it is robust to differences between the operational and the training workloads. Domenico Cotroneo, Luigi De Simone, Roberto Natella |
IEEE Trans. Reliab. | 3 |
| 2017 | A Fault Correlation Approach to Detect Performance Anomalies in Virtual Network Function ChainsabstractNetwork Function Virtualization is an emerging paradigm to allow the creation, at software level, of complex network services by composing simpler ones. However, this paradigm shift exposes network services to faults and bottlenecks in the complex software virtualization infrastructure they rely on. Thus, NFV services require effective anomaly detection systems to detect the occurrence of network problems. The paper proposes a novel approach to ease the adoption of anomaly detection in production NFV services, by avoiding the need to train a model or to calibrate a threshold. The approach infers the service health status by collecting metrics from multiple elements in the NFV service chain, and by analyzing their (lack of) correlation over the time. We validate this approach on an NFV-oriented Interactive Multimedia System, to detect problems affecting the quality of service, such as the overload, component crashes, avalanche restarts and physical resource contention. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
ISSRE | 2 |
| 2017 | Chizpurfle: A Gray-Box Android Fuzzer for Vendor Service CustomizationsabstractAndroid has become the most popular mobile OS, as it enables device manufacturers to introduce customizations to compete with value-added services. However, customizations make the OS less dependable and secure, since they can introduce software flaws. Such flaws can be found by using fuzzing, a popular testing technique among security researchers.This paper presents Chizpurfle, a novel "gray-box" fuzzing tool for vendor-specific Android services. Testing these services is challenging for existing tools, since vendors do not provide source code and the services cannot be run on a device emulator. Chizpurfle has been designed to run on an unmodified Android OS on an actual device. The tool automatically discovers, fuzzes, and profiles proprietary services. This work evaluates the applicability and performance of Chizpurfle on the Samsung Galaxy S6 Edge, and discusses software bugs found in privileged vendor services. Antonio Ken Iannillo, Roberto Natella, Domenico Cotroneo, Cristina Nita-Rotaru |
ISSRE | 2 |
| 2017 | NFV-Throttle: An Overload Control Framework for Network Function VirtualizationabstractNetwork function virtualization (NFV) aims to provide high-performance network services through cloud computing and virtualization technologies. However, network overloads represent a major challenge. While elastic cloud computing can partially address overloads by scaling on-demand, this mechanism is not quick enough to meet the strict high-availability requirements of “carrier-grade” telecom services. Thus, in this paper we propose a novel overload control framework (NFV-Throttle) to protect NFV services from failures due to an excess of traffic in the short term, by filtering the incoming traffic toward virtual network functions (VNFs) to make the best use of the available capacity, and to preserve the QoS of traffic flows admitted in the network. Moreover, the framework has been designed to fit the service models of NFV, including VNFaaS and NFVIaaS. We present an extensive experimental evaluation on the NFV-oriented Clearwater IMS, showing that the solution is robust and able to sustain severe overload conditions with a very small performance overhead. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2017 | NFV-Bench: A Dependability Benchmark for Network Function Virtualization SystemsabstractNetwork function virtualization (NFV) envisions the use of cloud computing and virtualization technology to reduce costs and innovate network services. However, this paradigm shift poses the question whether NFV will be able to fulfill the strict performance and dependability objectives required by regulations and customers. Thus, we propose a dependability benchmark to support NFV providers at making informed decisions about which virtualization, management, and application-level solutions can achieve the best dependability. We define in detail the use cases, measures, and faults to be injected. Moreover, we present a benchmarking case study on two alternative, production-grade virtualization solutions, namely VMware ESXi/vSphere (hypervisor-based) and Linux/Docker (container-based), on which we deploy an NFV-oriented IMS system. Despite the promise of higher performance and manageability, our experiments suggest that the container-based configuration can be less dependable than the hypervisor-based one, and point out which faults NFV designers should address to improve dependability. Domenico Cotroneo, Luigi De Simone, Roberto Natella |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2016 | Software Aging Analysis of the Android Mobile OSabstractMobile devices are significantly complex, feature-rich, and heavily customized, thus they are prone to software reliability and performance issues. This paper considers the problem of software aging in Android mobile OS, which causes the device to gradually degrade in responsiveness, and to eventually fail. We present a methodology to identify factors (such as workloads and device configurations) and resource utilization metrics that are correlated with software aging. Moreover, we performed an empirical analysis of recent Android devices, finding that software aging actually affects them. The analysis pointed out processes and components of the Android OS affected by software aging, and metrics useful as indicators of software aging to schedule software rejuvenation actions. Domenico Cotroneo, Francesco Fucci, Antonio Ken Iannillo, Roberto Natella, Roberto Pietrantuono |
ISSRE | 4 |
| 2016 | Recovery From Software Failures Caused by MandelbugsabstractSoftware failures are still a major concern in mission- and enterprise-critical contexts, despite significant efforts spent in software testing. In fact, while software testing is effective against easily-reproducible bugs (Bohrbugs), it is considerably less suitable for dealing with bugs that lead to hard-to-reproduce failures (Mandelbugs). On the positive side, the elusive nature of Mandelbugs provides opportunities for failure recovery, which are investigated in this paper. Based on real cases of Mandelbugs in eleven Information Technology (IT) systems running in production, the paper proposes a model that describes the recovery processes in IT systems. It then presents closed-form expressions, and a numerical analysis, of the mean time to recovery, and the software (un)availability. This analysis allows the designer to compare recovery strategies, as well as to determine the parameters having a high influence on the efficacy of recovery from failures caused by Mandelbugs. Michael Grottke, Dong Seong Kim 0001, Rajesh K. Mansharamani, Manoj Nambiar 0001, Roberto Natella, Kishor S. Trivedi |
IEEE Trans. Reliab. | 5 |
| 2015 | No PAIN, No Gain? The Utility of PArallel Fault INjectionsabstractSoftware Fault Injection (SFI) is an established technique for assessing the robustness of a software under test by exposing it to faults in its operational environment. Depending on the complexity of this operational environment, the complexity of the software under test, and the number and type of faults, a thorough SFI assessment can entail (a) numerous experiments and (b) long experiment run times, which both contribute to a considerable execution time for the tests. In order to counteract this increase when dealing with complex systems, recent works propose to exploit parallel hardware to execute multiple experiments at the same time. While Parallel fault Injections (PAIN) yield higher experiment throughput, they are based on an implicit assumption of non-interference among the simultaneously executing experiments. In this paper we investigate the validity of this assumption and determine the trade-off between increased throughput and the accuracy of experimental results obtained from PAIN experiments. Stefan Winter 0001, Oliver Schwahn, Roberto Natella, Neeraj Suri, Domenico Cotroneo |
ICSE (1) | 3 |
| 2015 | MoIO: Run-time monitoring for I/O protocol violations in storage device driversabstractBugs affecting storage device drivers include the so-called protocol violation bugs, which silently corrupt data and commands exchanged with I/O devices. Protocol violations are very difficult to prevent, since testing device driver is notoriously difficult. To address them, we present a monitoring approach for device drivers (MoIO) to detect HO protocol violations at run-time. The approach infers a model of the interactions between the storage device driver, the OS kernel, and the hardware (the device driver protocol) by analyzing execution traces. The model is then used as a reference for detecting violations in production. The approach has been designed to have a low overhead and to overcome the lack of source code and protocol documentation. We show that the approach is feasible and effective by applying it on the SATA/AHCI storage device driver of the Linux kernel, and by performing fault injection and long-running tests. Domenico Cotroneo, Luigi De Simone, Francesco Fucci, Roberto Natella |
ISSRE | 4 |
| 2015 | Dependability evaluation and benchmarking of Network Function Virtualization InfrastructuresabstractNetwork Function Virtualization (NFV) is an emerging solution that aims at improving the flexibility, the efficiency and the manageability of networks, by leveraging virtualization and cloud computing technologies to run network appliances in software. However, the “softwarization” of network functions raises reliability concerns, as they will be exposed to faults in commodity hardware and software components. In this paper, we propose a methodology for the dependability evaluation and benchmarking of NFV Infrastructures (NFVIs), based on fault injection. We discuss the application of the methodology in the context of a virtualized IP Multimedia Subsystem (IMS), and the pitfalls in the design of a reliable NFVI. Domenico Cotroneo, Luigi De Simone, Antonio Ken Iannillo, Anna Lanzaro, Roberto Natella |
NetSoft | 5 |
| 2014 | An empirical study of injected versus actual interface errorsabstractThe reuse of software components is a common practice in commercial applications and increasingly appearing in safety critical systems as driven also by cost considerations. This practice puts dependability at risk, as differing operating conditions in different reuse scenarios may expose residual software faults in the components. Consequently, software fault injection techniques are used to assess how residual faults of reused software components may affect the system, and to identify appropriate counter-measures. As fault injection in components’ code suffers from a number of practical disadvantages, it is often replaced by error injection at the component interface level. However, it is still an open issue, whether such injected errors are actually representative of the effects of residual faults. To this end, we propose a method for analyzing how software faults turn into interface errors, with the ultimate aim of supporting more representative interface error injection experiments. Our analysis in the context of widely used software libraries reveals that existing interface error models are not suitable for emulating software faults, and provides useful insights for improving the representativeness of interface error injection. Anna Lanzaro, Roberto Natella, Stefan Winter 0001, Domenico Cotroneo, Neeraj Suri |
ISSTA | 2 |
| 2014 | A survey of software aging and rejuvenation studiesabstractSoftware aging is a phenomenon plaguing many long-running complex software systems, which exhibit performance degradation or an increasing failure rate. Several strategies based on the proactive rejuvenation of the software state have been proposed to counteract software aging and prevent failures. This survey article provides an overview of studies on Software Aging and Rejuvenation (SAR) that have appeared in major journals and conference proceedings, with respect to the statistical approaches that have been used to forecast software aging phenomena and to plan rejuvenation, the kind of systems and aging effects that have been studied, and the techniques that have been proposed to rejuvenate complex software systems. The analysis is useful to identify key results from SAR research, and it is leveraged in this article to highlight trends and open issues. Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2013 | Analysis and Prediction of Mandelbugs in an Industrial Software SystemabstractMandelbugs are faults that are triggered by complex conditions, such as interaction with hardware and other software, and timing or ordering of events. These faults are considerably difficult to detect with traditional testing techniques, since it can be challenging to control their complex triggering conditions in a testing environment. Therefore, it is necessary to adopt specific verification and/or fault-tolerance strategies for dealing with them in a cost-effective way. In this paper, we investigate how to predict the location of Mandelbugs in complex software systems, in order to focus V&V activities and fault tolerance mechanisms in those modules where Mandelbugs are most likely present. In the context of an industrial complex software system, we empirically analyze Mandelbugs, and investigate an approach for Mandelbug prediction based on a set of novel software complexity metrics. Results show that Mandelbugs account for a noticeable share of faults, and that the proposed approach can predict Mandelbug-prone modules with greater accuracy than the sole adoption of traditional software metrics. Gabriella Carrozza, Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ICST | 3 |
| 2013 | Fault triggers in open-source software: An experience reportabstractWith software systems becoming increasingly large and complex, many difficulties in coping with software bugs arise for developers. Despite good development practices, thorough testing, and proper maintenance policies, a non-negligible number of bugs remain in the released software. Understanding the type of residual bugs is fundamental for adopting proper countermeasures in current and future software releases. Depending on the fault triggering conditions that lead to a failure, developers can introduce fault-tolerance mechanisms and plan verification and validation strategies. In this paper, we analyze bugs in four large open-source software systems during their lifecycle, based on the concept of fault triggers. We first investigate how the type of system affects the bug type proportions, and their evolution over years. Then, an analysis of bug subtypes is performed, so as to better understand their nature, followed by a comparison with respect to attributes such as their average time to fix and severity. Domenico Cotroneo, Michael Grottke, Roberto Natella, Roberto Pietrantuono, Kishor S. Trivedi |
ISSRE | 3 |
| 2013 | SABRINE: State-based robustness testing of operating systemsabstractThe assessment of operating systems robustness with respect to unexpected or anomalous events is a fundamental requirement for mission-critical systems. Robustness can be tested by deliberately exposing the system to erroneous events during its execution, and then analyzing the OS behavior to evaluate its ability to gracefully handle these events. Since OSs are complex and stateful systems, robustness testing needs to account for the timing of erroneous events, in order to evaluate the robust behavior of the OS under different states. This paper presents SABRINE (StAte-Based Robustness testIng of operatiNg systEms), an approach for state-aware robustness testing of OSs. SABRINE automatically extracts state models from execution traces, and generates a set of test cases that cover different OS states. We evaluate the approach on a Linux-based Real-Time Operating System adopted in the avionic domain. Experimental results show that SABRINE can automatically identify relevant OS states, and find robustness vulnerabilities while keeping low the number of test cases. Domenico Cotroneo, Domenico Di Leo, Francesco Fucci, Roberto Natella |
ASE | 4 |
| 2013 | State-Driven Testing of Distributed Systems
Domenico Cotroneo, Roberto Natella, Stefano Russo 0001, Fabio Scippacercola |
OPODIS | 2 |
| 2013 | Predicting aging-related bugs using software complexity metrics
Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono |
Perform. Evaluation | 2 |
| 2013 | On Fault Representativeness of Software Fault InjectionabstractThe injection of software faults in software components to assess the impact of these faults on other components or on the system as a whole, allowing the evaluation of fault tolerance, is relatively new compared to decades of research on hardware fault injection. This paper presents an extensive experimental study (more than 3.8 million individual experiments in three real systems) to evaluate the representativeness of faults injected by a state-of-the-art approach (G-SWFIT). Results show that a significant share (up to 72 percent) of injected faults cannot be considered representative of residual software faults as they are consistently detected by regression tests, and that the representativeness of injected faults is affected by the fault location within the system, resulting in different distributions of representative/nonrepresentative faults across files and functions. Therefore, we propose a new approach to refine the faultload by removing faults that are not representative of residual software faults. This filtering is essential to assure meaningful results and to reduce the cost (in terms of number of faults) of software fault injection campaigns in complex software. The proposed approach is based on classification algorithms, is fully automatic, and can be used for improving fault representativeness of existing software fault injection approaches. Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira |
IEEE Trans. Software Eng. | 1 |
| 2011 | A Case Study on State-Based Robustness Testing of an Operating System for the Avionic Domain
Domenico Cotroneo, Domenico Di Leo, Roberto Natella, Roberto Pietrantuono |
SAFECOMP | 3 |
| 2010 | Assessing and improving the effectiveness of logs for the analysis of software faultsabstractEvent logs are the primary source of data to characterize the dependability behavior of a computing system during the operational phase. However, they are inadequate to provide evidence of software faults, which are nowadays among the main causes of system outages. This paper proposes an approach based on software fault injection to assess the effectiveness of logs to keep track of software faults triggered in the field. Injection results are used to provide guidelines to improve the ability of logging mechanisms to report the effects of software faults. The benefits of the approach are shown by means of experimental results on three widely used software systems. Marcello Cinque, Domenico Cotroneo, Roberto Natella, Antonio Pecchia |
DSN | 3 |
| 2010 | Representativeness analysis of injected software faults in complex softwareabstractDespite of the existence of several techniques for emulating software faults, there are still open issues regarding representativeness of the faults being injected. An important aspect, not considered by existing techniques, is the non-trivial activation condition (trigger) of real faults, which causes them to elude testing and remain hidden until operation. In this paper, we investigate how the representativeness of injected software faults can be improved regarding the representativeness of triggers, by proposing a set of generic criteria to select representative faults from afaultload. We used the G-SWFIT technique to inject software faults in a DBMS, resulting in over 40 thousands faults and 2 million runs of a real test suite. We analyzed faults with respect to their triggers, and concluded that a non-negligible share (15%) would not realistically elude testing. Our proposed criteria decreased the percentage of non-elusive faults in the faultload, improving its representativeness. Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira |
DSN | 1 |
| 2010 | Software Aging Analysis of the Linux Operating SystemabstractSoftware systems running continuously for a long time tend to show degrading performance and an increasing failure occurrence rate, due to error conditions that accrue over time and eventually lead the system to failure. This phenomenon is usually referred to as Software Aging. Several long-running mission and safety critical applications have been reported to experience catastrophic aging-related failures. Software aging sources (i.e., aging-related bugs) may be hidden in several layers of a complex software system, ranging from the Operating System (OS) to the user application level. This paper presents a software aging analysis at the Operating System level, investigating software aging sources inside the Linux kernel. Linux is increasingly being employed in critical scenarios; this analysis intends to shed light on its behaviour from the aging perspective. The study is based on an experimental campaign designed to investigate the kernel internal behaviour over long running executions. By means of a kernel tracing tool specifically developed for this study, we collected relevant parameters of several kernel subsystems. Statistical analysis of collected data allowed us to confirm the presence of aging sources in Linux and to relate the observed aging dynamics to the monitored subsystems behaviour. The analysis output allowed us to infer potential sources of aging in the kernel subsystems. Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 2 |
| 2010 | Memory leak analysis of mission-critical middleware
Gabriella Carrozza, Domenico Cotroneo, Roberto Natella, Antonio Pecchia, Stefano Russo 0001 |
J. Syst. Softw. | 3 |
| 2009 | Assessment and Improvement of Hang Detection in the Linux Operating SystemabstractWe propose a fault injection framework to assess hang detection facilities within the Linux operating system (OS). The novelty of the framework consists in the adoption of a more representative fault load than existing ones, and in the effectiveness in terms of number of hang failures produced; representativeness is supported by a field data study on the Linux OS. Using the proposed fault injection framework, along with realistic workloads, we find that the Linux OS is unable to detect hangs in several cases. We experience a relative coverage of 75%. To improve detection facilities, we propose a simple yet effective hang detector, which periodically tests OS liveness, as perceived by applications, by means of I/O system calls; it is shown that this approach can improve relative coverage up to 94%. The hang detector can be deployed on any Linux system, with an acceptable overhead. Domenico Cotroneo, Roberto Natella, Stefano Russo 0001 |
SRDS | 2 |