EDBT 2026 Demo / reviewers in the wild / expert
Domenico Cotroneo
dblp:c/DomenicoCotroneo
· DBLP profile ↗
121ranked-venue papers
60as first author
34since 2021 · last 2026
0000-0001-7103-592XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 55 · 30 first-author · 20 since 2021Security and privacy · 29 · 13 first-author · 7 since 2021Systems, architecture and hardware · 27 · 13 first-author · 4 since 2021Computer networks · 7 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Makes Software Bugs Escape Testing? Evidence from a Large-Scale Empirical Study
Domenico Cotroneo, Giuseppe De Rosa, Cristina Improta, Benedetta Gaia Varriale |
DSN | 1 |
| 2026 | Elevating Cyber Threat Intelligence against disinformation campaigns with LLM-based concept extraction and the FakeCTI datasetabstractThe swift spread of fake news and disinformation campaigns poses a significant threat to public trust, political stability, and cybersecurity. Traditional Cyber Threat Intelligence (CTI) approaches, which rely on low-level indicators such as domain names and social media handles, are easily evaded by adversaries who frequently modify their online infrastructure. To address these limitations, we introduce a novel CTI framework that focuses on high-level, semantic indicators derived from recurrent narratives and relationships of disinformation campaigns. Our approach extracts structured CTI indicators from unstructured disinformation content, capturing key entities and their contextual dependencies within fake news using Large Language Models (LLMs). We further introduce FakeCTI, the first dataset that systematically links fake news to disinformation campaigns and threat actors. To evaluate the effectiveness of our CTI framework, we analyze multiple fake news attribution techniques, spanning from traditional Natural Language Processing (NLP) to fine-tuned LLMs. This work shifts the focus from low-level artifacts to persistent conceptual structures, establishing a scalable and adaptive approach to tracking and countering disinformation campaigns. Domenico Cotroneo, Roberto Natella, Vittorio Orbinato |
J. Syst. Softw. | 1 |
| 2026 | Machine learning for software aging detection: A systematic mapping study
Rafael José Moura, Maria Gizele Nascimento, Fumio Machida, Domenico Cotroneo, Ermeson Carneiro de Andrade |
J. Syst. Softw. | 4 |
| 2026 | Aging-Related Bug Prediction Based on Multi-View Graph Feature Learning and Graph-TransformerabstractSoftware aging, characterized by an increasing failure rate or performance degradation in long-running software systems, poses significant risks, including substantial financial losses and potential threats to human lives. This phenomenon is primarily driven by the accumulation of runtime errors, commonly referred to as aging-related bugs (ARBs). Aging-related bug prediction (ARBP) has been proposed to facilitate the detection and remediation of ARBs prior to software release. However, ARBP’s effectiveness heavily depends on the quality of dataset features used. Previous research has largely relied on a standard set of manually designed metrics, often overlooking that these metrics may fail to distinguish between code segments with different semantics, even when they exhibit identical metric values. While some studies have attempted to develop models that learn semantic features from source code, they typically focus on token-level or graph-level features, neglecting a comprehensive exploration of ARB characteristics within the source code. Specifically, there is insufficient discussion on whether deep semantic features can adequately capture the essential traits that trigger aging phenomena. In this paper, we propose a novel multi-view graph feature learning framework based on Graph-Transformer, which integrates newly proposed ARB features extracted from Abstract Syntax Trees with Code Property Graphs for feature learning. Our approach effectively captures hierarchical structures and variable dependencies, facilitating the identification of complex interactions that contribute to ARBs. Additionally, we implement sub-graph sampling and class imbalance strategies to enhance model performance. Experimental results across three datasets demonstrate that our method surpasses state-of-the-art approaches, a code property graph-based feature extraction method (specifically SGT), achieving precision improvements of 8.2% on Linux, 15.4% on MySQL, and 2.5% on NetBSD, thereby establishing a new benchmark for ARB prediction. Jianwen Xiang, Roberto Natella, Roberto Pietrantuono, Domenico Cotroneo |
IEEE Trans. Software Eng. | 8 |
| 2025 | Human-Written vs. AI-Generated Code: A Large-Scale Study of Defects, Vulnerabilities, and ComplexityabstractAs AI code assistants become increasingly integrated into software development workflows, understanding how their code compares to human-written programs is critical for ensuring reliability, maintainability, and security. In this paper, we present a large-scale comparison of code authored by human developers and three state-of-the-art LLMs, i.e., ChatGPT, DeepSeek-Coder, and Qwen-Coder, on multiple dimensions of software quality: code defects, security vulnerabilities, and structural complexity. Our evaluation spans over 500k code samples in two widely used languages, Python and Java, classifying defects via Orthogonal Defect Classification and security vulnerabilities using the Common Weakness Enumeration. We find that AI-generated code is generally simpler and more repetitive, yet more prone to unused constructs and hardcoded debugging, while humanwritten code exhibits greater structural complexity and a higher concentration of maintainability issues. Notably, AI-generated code also contains more high-risk security vulnerabilities. These findings highlight the distinct defect profiles of AI-and humanauthored code and underscore the need for specialized quality assurance practices in AI-assisted programming. Domenico Cotroneo, Cristina Improta, Pietro Liguori |
ISSRE | 1 |
| 2025 | Open-FARI: An Open-source testbed for Federated Anomaly detection in the Railway Industrial Internet of ThingsabstractThe paper presents Open-FARI, an open-source testbed for evaluating federated learning algorithms for anomaly detection in the railway Industrial Internet of Things domain. Open-FARI uses synthetic data generation modules trained from real train sensor data to generate realistic sensor data of a fleet of trains. Generated data encompass normal and anomalous data, enabling the evaluation of federated learning algorithms for anomaly detection. The paper addresses the lack of testbeds and datasets tailored to the railway domain, which represents an obstacle to research on Machine Learning-driven solutions in this domain. Alessandra Rizzardi, Raffaele Della Corte, Jesús Fernando Cevallos Moreno, Simona De Vivo, Vittorio Orbinato, Sabrina Sicari, Domenico Cotroneo, Alberto Coen-Porisini |
IWCMC | 7 |
| 2025 | Quality In, Quality Out: Investigating Training Data's Role in AI Code GenerationabstractDeep Learning (DL)-based code generators have seen significant advancements in recent years. Tools such as GitHub Copilot are used by thousands of developers with the main promise of a boost in productivity. However, researchers have recently questioned their impact on code quality showing, for example, that code generated by DL-based tools may be affected by security vulnerabilities. Since DL models are trained on large code corpora, one may conjecture that low-quality code they output is the result of low-quality code they have seen during training. However, there is very little empirical evidence documenting this phenomenon. Indeed, most of previous work look at the frequency with which commercial code generators (e.g., Copilot, ChatGPT) recommend low-quality code without the possibility of relating this to their (publicly unavailable) training set. In this paper, we investigate the extent to which low-quality code instances seen during training affect the quality of the code generated at inference time. We start by fine-tuning a pre-trained DL model on a large-scale dataset ($>4.4 \mathrm{M}$functions) being representative of those usually adopted in the training of code generators. We show that 4.98 % of functions in this dataset exhibit one or more quality issues related to security, maintainability, coding practices, etc. We use the fine-tuned model to generate 551 k Python functions, showing that 5.85 % of them are affected by at least one quality issue. We then remove from the training set the low-quality functions, and use the cleaned dataset to fine-tune a second model which has been used to generate the same 551k Python functions. We show that the model trained on the cleaned dataset exhibits similar performance in terms of functional correctness as compared to the original model (i.e., the one trained on the whole dataset) while, however, generating a statistically significant lower number of low-quality functions (2.16 %). Our study empirically documents the importance of high-quality training data for code generators. Cristina Improta, Rosalia Tufano, Pietro Liguori, Domenico Cotroneo, Gabriele Bavota |
ICPC | 4 |
| 2025 | Enhancing robustness of AI offensive code generators via data augmentation
Cristina Improta, Pietro Liguori, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
Empir. Softw. Eng. | 5 |
| 2025 | DeVAIC: A tool for security assessment of AI-generated code
Domenico Cotroneo, Roberta De Luca, Pietro Liguori |
Inf. Softw. Technol. | 1 |
| 2025 | COSMOS: A Fault Injection Framework to Assess Hardware-Assisted HypervisorsabstractHardware-assisted virtualization represents a pillar technology for large-scale clusters and cloud-based applications. Hardware faults are still frequent as technology advances, potentially resulting in serious reliability concerns. This paper introduces COSMOS, a fault injection framework tailored for testing hardware-assisted hypervisors. By exploiting nested virtualization, COSMOS does not require instrumentation of the target and enables the assessment of multiple hypervisors. We performed an extensive fault injection campaign to assess popular hardware-assisted hypervisors like KVM, Xen, and Jailhouse. The results show a non-negligible percentage of non–fail-stop behaviors, with notable differences in hypervisors’ ability to log failures and prevent fault propagation with a timely recovery. Marcello Cinque, Domenico Cotroneo, Giuseppe De Rosa, Luigi De Simone, Giorgio Farina |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Enhancing AI-based Generation of Software Exploits with Contextual InformationabstractThis practical experience report explores Neural Machine Translation (NMT) models’ capability to generate offensive security code from natural language (NL) descriptions, highlighting the significance of contextual understanding and its impact on model performance. Our study employs a dataset comprising real shellcodes to evaluate the models across various scenarios, including missing information, necessary context, and unnecessary context. The experiments are designed to assess the models’ resilience against incomplete descriptions, their proficiency in leveraging context for enhanced accuracy, and their ability to discern irrelevant information. The findings reveal that the introduction of contextual data significantly improves performance. However, the benefits of additional context diminish beyond a certain point, indicating an optimal level of contextual information for model training. Moreover, the models demonstrate an ability to filter out unnecessary context, maintaining high levels of accuracy in the generation of offensive security code. This study paves the way for future research on optimizing context use in AI-driven code generation, particularly for applications requiring a high degree of technical precision such as the generation of offensive code. Pietro Liguori, Cristina Improta, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
ISSRE | 5 |
| 2024 | Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning AttacksabstractAI-based code generators have become pivotal in assisting developers in writing software starting from natural language (NL). However, they are trained on large amounts of data, often collected from unsanitized online sources (e.g., GitHub, HuggingFace). As a consequence, AI models become an easy target for data poisoning, i.e., an attack that injects malicious samples into the training data to generate vulnerable code. Domenico Cotroneo, Cristina Improta, Pietro Liguori, Roberto Natella |
ICPC | 1 |
| 2024 | RaiIRED: a Node-RED-Based Framework for Modeling Train Control Management SystemsabstractThe modeling and simulation of Internet of Things (IoT) and Industrial IoT (IIoT) systems allow practitioners to obtain valuable insights into the system's behavior before their actual deployment in the field. Early designing permits the analysis of the interactions among the involved entities, evaluating the effects of modifications, and understanding the impact of failures on the system. In particular, this is exacerbated in the context of IoT/IIoT, which is characterized by multiple and heterogeneous subsystems, different processing levels, and communication protocols. In such a direction, recent innovations in IT devices have enabled the rail industry to gather information from Train Control and Monitoring Systems (TCMS) to check conditions constantly and prevent issues, thus improving relia-bility and safety and, in some cases, leading to cost-saving by optimizing maintenance resources. In such a scenario, this paper presents RailRED, a framework for simulating and prototyping a TCMS based on the Node-RED tool. In RaiIRED, the main TCMS subsystems are modeled using Node-RED flows, while the subsystem interconnections are performed through a low footprint and encrypted gateway based on the MQTT protocol. The proposal can also generate diagnostic data that mimic the behavior of a real-world TCMS. RailRED communication latency and its ability to generate diagnostic data have been analyzed, with the latter evaluated by using clusters of diagnostic events collected from a real-world TCMS running on a high-speed train. Alessandra Rizzardi, Raffaele Della Corte, Jesús Fernando Cevallos Moreno, Vittorio Orbinato, Simona De Vivo, Sabrina Sicari, Domenico Cotroneo, Alberto Coen-Porisini |
WiMob | 7 |
| 2024 | DRACO: Distributed Resource-aware Admission Control for large-scale, multi-tier systemsabstractModern distributed systems are designed to manage overload conditions, by throttling the traffic in excess that cannot be served through overload control techniques. However, the adoption of large-scale NoSQL datastores make systems vulnerable to unbalanced overloads, where specific datastore nodes are overloaded because of hot-spot resources and hogs. In this paper, we propose DRACO, a novel overload control solution that is aware of data dependencies between the application and the datastore tiers. DRACO performs selective admission control of application requests, by only dropping the ones that map to resources on overloaded datastore nodes, while achieving high resource utilization on non-overloaded datastore nodes. We evaluate DRACO on two case studies with high availability and performance requirements, a virtualized IP Multimedia Subsystem and a distributed fileserver. Results show that the solution can achieve high performance and resource utilization even under extreme overload conditions, up to 100x the engineered capacity. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
J. Parallel Distributed Comput. | 1 |
| 2024 | Automating the correctness assessment of AI-generated code for security contextsabstractEvaluating the correctness of code generated by AI is a challenging open problem. In this paper, we propose a fully automated method, named ACCA, to evaluate the correctness of AI-generated code for security purposes. The method uses symbolic execution to assess whether the AI-generated code behaves as a reference implementation. We use ACCA to assess four state-of-the-art models trained to generate security-oriented assembly code and compare the results of the evaluation with different baseline solutions, including output similarity metrics, widely used in the field, and the well-known ChatGPT, the AI-powered language model developed by OpenAI. Our experiments show that our method outperforms the baseline solutions and assesses the correctness of the AI-generated code similar to the human-based evaluation, which is considered the ground truth for the assessment in the field. Moreover, ACCA has a very strong correlation with the human evaluation (Pearson’s correlation coefficient r=0.84 on average). Finally, since it is a full y automated solution that does not require any human intervention, the proposed method performs the assessment of every code snippet in ∼0.17 s on average, which is definitely lower than the average time required by human analysts to manually inspect the code, based on our experience. Domenico Cotroneo, Alessio Foggia, Cristina Improta, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 1 |
| 2024 | SGT: Aging-related bug prediction via semantic feature learning based on graph-transformer
Jianwen Xiang, Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono |
J. Syst. Softw. | 5 |
| 2024 | Laccolith: Hypervisor-Based Adversary Emulation With Anti-DetectionabstractAdvanced Persistent Threats (APTs) represent the most threatening form of attack nowadays since they can stay undetected for a long time. Adversary emulation is a proactive approach for preparing against these attacks. However, adversary emulation tools lack the anti-detection abilities of APTs. We introduce Laccolith, a hypervisor-based solution for adversary emulation with anti-detection to fill this gap. We also present an experimental study to compare Laccolith with MITRE CALDERA, a state-of-the-art solution for adversary emulation, against five popular anti-virus products. We found that CALDERA cannot evade detection, limiting the realism of emulated attacks, even when combined with a state-of-the-art anti-detection framework. Our experiments show that Laccolith can hide its activities from all the tested anti-virus products, thus making it suitable for realistic emulations. Vittorio Orbinato, Marco Carlo Feliciano, Domenico Cotroneo, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | IRIS: a Record and Replay Framework to Enable Hardware-assisted Virtualization FuzzingabstractNowadays, industries are looking into virtualization as an effective means to build safe applications, thanks to the isolation it can provide among virtual machines (VMs) running on the same hardware. In this context, a fundamental issue is understanding to what extent the isolation is guaranteed, despite possible (or induced) problems in the virtualization mechanisms. Uncovering such isolation issues is still an open challenge, especially for hardware-assisted virtualization, since the search space should include all the possible VM states (and the linked hypervisor state), which is prohibitive. In this paper, we propose IRIS, a framework to record (learn) sequences of inputs (i.e., VM seeds) from the real guest execution (e.g., OS boot), replay them as-is to reach valid and complex VM states, and finally use them as valid seed to be mutated for enabling fuzzing solutions for hardware-assisted hypervisors. We demonstrate the accuracy and efficiency of IRIS in automatically reproducing valid VM behaviors, with no need to execute guest workloads. We also provide a proof-of-concept fuzzer, based on the proposed architecture, showing its potential on the Xen hypervisor. Carmine Cesarano 0002, Marcello Cinque, Domenico Cotroneo, Luigi De Simone, Giorgio Farina |
DSN | 3 |
| 2023 | IFCM: An improved Fuzzy C-means clustering method to handle Class Overlap on Aging-related Software Bug PredictionabstractSoftware aging refers to a problem of performance decay in long-running software systems. This phenomenon is primarily attributed to the accumulation of run-time errors, commonly known as aging-related bugs (ARBs). Detecting ARBs through Aging-related Bug Prediction (ARBP) is crucial in ensuring system reliability. The effectiveness of ARBP heavily relies on the quality of datasets. However, ARB datasets often suffer from class overlap, where instances from different classes exhibit similar feature values. Class overlap poses a significant challenge as it compromises the quality of training data and subsequently impacts ARBP accuracy. To address this issue, we propose an improved Fuzzy C-means clustering method named IFCM, designed to mitigate class overlap in ARBP tasks. IFCM can identify whether an instance occurs overlap, and identify the overlap degree of this instance through the predefined parameters. We evaluate our proposed method on two public datasets Linux and MySQL and one self-collected dataset NetBSD using five different classifiers with five performance metrics (AUC, F1, Balance, PD, PF). Comparison with four existing methods (No clean, NCL, IKMCCA, ROCT) demonstrates that IFCM is effective in alleviating class overlap in ARBP. For Instance, IFCM achieves promising results in terms of AUC blue (which are 0.762, 0.757, and 0.642) and Balance (which are 0.709, 0.736, and 0.595) at the dataset level. Shuo Feng 0003, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang, Roberto Pietrantuono, Roberto Natella, Domenico Cotroneo |
ISSRE | 8 |
| 2023 | Who evaluates the evaluators? On automatic metrics for assessing AI-based offensive code generatorsabstractAI-based code generators are an emerging solution for automatically writing programs starting from descriptions in natural language, by using deep neural networks (Neural Machine Translation, NMT). In particular, code generators have been used for ethical hacking and offensive security testing by generating proof-of-concept attacks. Unfortunately, the evaluation of code generators still faces several issues. The current practice uses output similarity metrics, i.e., automatic metrics that compute the textual similarity of generated code with ground-truth references. However, it is not clear what metric to use, and which metric is most suitable for specific contexts. This work analyzes a large set of output similarity metrics on offensive code generators. We apply the metrics on two state-of-the-art NMT models using two datasets containing offensive assembly and Python code with their descriptions in the English language. We compare the estimates from the automatic metrics with human evaluation and provide practical insights into their strengths and limitations. Pietro Liguori, Cristina Improta, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
Expert Syst. Appl. | 5 |
| 2023 | Run-time failure detection via non-intrusive event analysis in a large-scale cloud computing platformabstractCloud computing systems fail in complex and unforeseen ways due to unexpected combinations of events and interactions among hardware and software components. These failures are especially problematic when they are silent, i.e., not accompanied by any explicit failure notification, hindering the timely detection and recovery. In this work, we propose an approach to run-time failure detection tailored for monitoring multi-tenant and concurrent cloud computing systems. The approach uses a non-intrusive form of event tracing, without manual changes to the system’s internals to propagate session identifiers (IDs), and builds a set of lightweight monitoring rules from fault-free executions. We evaluated the effectiveness of the approach in detecting failures in the context of the OpenStack cloud computing platform, a complex and “off-the-shelf” distributed system, by executing a campaign of fault injection experiments in a multi-tenant scenario. Our experiments show that the approach detects the failure with an F1 score (0.85) and accuracy (0.77) higher than the ones provided by the OpenStack failure logging mechanisms (0.53 and 0.50) and two non-session-aware run-time verification approaches (both lower than 0.15). Moreover, the approach significantly decreases the average time to detect failures at run-time (∼114 seconds) compared to the OpenStack logging mechanisms. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 1 |
| 2023 | Guest editorial: special issue on emerging challenges in software certification and verification
Luigi De Simone, Nuno Laranjeiro, Domenico Cotroneo |
Softw. Qual. J. | 3 |
| 2023 | A Comparative Analysis of Software Aging in Image Classifiers on Cloud and EdgeabstractImage classifiers for recognizing real-world objects are widely used in the Internet of Things (IoT) and Cyber-Physical Systems(CPSs). A classifier is trained offline by machine learning algorithms with training data sets, and then it is deployed on a cloud or an edge computing system for online label predictions. As the classifier's performance depends on the underlying software infrastructure, it may degrade over time due to software faults causing software aging. In this paper, we address this issue and experimentally investigate software aging observed in an image classification system that continuously runs on cloud and edge computing environments. We apply several statistical techniques to analyze degradation trends in the systems under stress tests. Our statistical trend analysis confirms the degradation trends in the throughput as well as the available memory resources both in the cloud and the edge environments. Contrary to our expectation, the edge computing environment under test had much less impact on the performance degradation than our cloud environment when the workload is high, although the latter one has four times larger allocated memory resources. We also show that the observed performance degradation trends are associated with the memory usage of specific processes by performing correlation analysis. Ermeson Carneiro de Andrade, Roberto Pietrantuono, Fumio Machida, Domenico Cotroneo |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | Automatic Mapping of Unstructured Cyber Threat Intelligence: An Experimental Study: (Practical Experience Report)abstractProactive approaches to security, such as adversary emulation, leverage information about threat actors and their techniques (Cyber Threat Intelligence, CTI). However, most CTI still comes in unstructured forms (i.e., natural language), such as incident reports and leaked documents. To support proactive security efforts, we present an experimental study on the automatic classification of unstructured CTI into attack techniques using machine learning (ML). We contribute with two new datasets for CTI analysis, and we evaluate several ML models, including both traditional and deep learning-based ones. We present several lessons learned about how ML can perform at this task, which classifiers perform best and under which conditions, which are the main causes of classification errors, and the challenges ahead for CTI analysis. Vittorio Orbinato, Mariarosaria Barbaraci, Roberto Natella, Domenico Cotroneo |
ISSRE | 4 |
| 2022 | An Empirical Study on Software Aging of Long-Running Object Detection AlgorithmsabstractEfficient and effective object detection is a key problem in Computer Vision. Numerous object detection algorithms have been developed, whose aim is to achieve two conflicting goals, namely accuracy and efficiency, while being executed in real-time with high robustness. Many of these algorithms must run for an extended period of time, i.e., in video surveillance or in self-driving cars – a working condition that make them subject to the risk of software aging.In this work, we focus on evaluating several object detection algorithms to understand if and to what extent they are affected by software aging. A measurement-based aging approach was adopted, with a series of long-running tests and subsequent data analysis. The results report significant trends of performance degradation, sometimes leading to aging-related failures, as well as memory consumption trends, which turned out to be the main issue across all the experiments. Roberto Pietrantuono, Domenico Cotroneo, Ermeson Carneiro de Andrade, Fumio Machida |
QRS | 2 |
| 2022 | Can we generate shellcodes via natural language? An empirical studyabstractAbstract Writing software exploits is an important practice for offensive security analysts to investigate and prevent attacks. In particular, shellcodes are especially time-consuming and a technical challenge, as they are written in assembly language. In this work, we address the task of automatically generating shellcodes, starting purely from descriptions in natural language, by proposing an approach based on Neural Machine Translation (NMT). We then present an empirical study using a novel dataset ( Shellcode_IA32 ), which consists of 3200 assembly code snippets of real Linux/x86 shellcodes from public databases, annotated using natural language. Moreover, we propose novel metrics to evaluate the accuracy of NMT at generating shellcodes. The empirical analysis shows that NMT can generate assembly code snippets from the natural language with high accuracy and that in many cases can generate entire shellcodes with no errors. Pietro Liguori, Erfan Al-Hossami, Domenico Cotroneo, Roberto Natella, Bojan Cukic, Samira Shaikh |
Autom. Softw. Eng. | 3 |
| 2022 | Virtualizing mixed-criticality systems: A survey on industrial trends and issues
Marcello Cinque, Domenico Cotroneo, Luigi De Simone, Stefano Rosiello |
Future Gener. Comput. Syst. | 2 |
| 2022 | ThorFI: a Novel Approach for Network Fault Injection as a Service
Domenico Cotroneo, Luigi De Simone, Roberto Natella |
J. Netw. Comput. Appl. | 1 |
| 2022 | Software micro-rejuvenation for Android mobile systems
Domenico Cotroneo, Luigi De Simone, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 1 |
| 2022 | Fault Injection Analytics: A Novel Approach to Discover Failure Modes in Cloud-Computing SystemsabstractCloud computing systems fail in complex and unexpected ways due to unexpected combinations of events and interactions between hardware and software components. Fault injection is an effective means to bring out these failures in a controlled environment. However, fault injection experiments produce massive amounts of data, and manually analyzing these data is inefficient and error-prone, as the analyst can miss severe failure modes that are yet unknown. This article introduces a new paradigm (fault injection analytics) that applies unsupervised machine learning on execution traces of the injected system, to ease the discovery and interpretation of failure modes. We evaluated the proposed approach in the context of fault injection experiments on the OpenStack cloud computing platform, where we show that the approach can accurately identify failure modes with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | EVIL: Exploiting Software via Natural LanguageabstractWriting exploits for security assessment is a challenging task. The writer needs to master programming and obfuscation techniques to develop a successful exploit. To make the task easier, we propose an approach (EVIL) to automatically generate exploits in assembly/Python language from descriptions in natural language. The approach leverages Neural Machine Translation (NMT) techniques and a dataset that we developed for this work. We present an extensive experimental study to evaluate the feasibility of EVIL, using both automatic and manual analysis, and both at generating individual statements and entire exploits. The generated code achieved high accuracy in terms of syntactic and semantic correctness. Pietro Liguori, Erfan Al-Hossami, Vittorio Orbinato, Roberto Natella, Samira Shaikh, Domenico Cotroneo, Bojan Cukic |
ISSRE | 6 |
| 2021 | Timing covert channel analysis of the VxWorks MILS embedded hypervisor under the common criteria security certification
Domenico Cotroneo, Luigi De Simone, Roberto Natella |
Comput. Secur. | 1 |
| 2021 | Enhancing the analysis of software failures in cloud computing systems with deep learning
Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 1 |
| 2021 | Dependability Assessment of the Android OS Through Fault InjectionabstractThe reliability of mobile devices is a challenge for vendors, since the mobile software stack has significantly grown in complexity. In this article, we study how to assess the impact of faults on the quality of user experience in the Android mobile OS through fault injection. We first address the problem of identifying a realistic fault model for the Android OS, by providing to developers a set of lightweight and systematic guidelines for fault modeling. Then, we present an extensible fault injection tool (AndroFIT) to apply such fault model on actual, commercial Android devices. Finally, we present a large fault injection experimentation on three Android products from major vendors, and point out several reliability issues and opportunities for improving the Android OS. Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella, Stefano Rosiello |
IEEE Trans. Reliab. | 1 |
| 2020 | ProFIPy: Programmable Software Fault Injection as-a-ServiceabstractIn this paper, we present a new fault injection tool (ProFIPy) for Python software. The tool is designed to be programmable, in order to enable users to specify their software fault model, using a domain-specific language (DSL) for fault injection. Moreover, to achieve better usability, ProFIPy is provided as software-as-a-service and supports the user through the configuration of the faultload and workload, failure data analysis, and full automation of the experiments using container- based virtualization and parallelization. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
DSN | 1 |
| 2020 | Dependability Evaluation of Middleware Technology for Large-scale Distributed CachingabstractDistributed caching systems (e.g., Memcached) are widely used by service providers to satisfy accesses by millions of concurrent clients. Given their large-scale, modern distributed systems rely on a middleware layer to manage caching nodes, to make applications easier to develop, and to apply load balancing and replication strategies. In this work, we performed a dependability evaluation of three popular middleware platforms, namely Twemproxy by Twitter, Mcrouter by Facebook, and Dynomite by Netflix, to assess availability and performance under faults, including failures of Memcached nodes and congestion due to unbalanced workloads and network link bandwidth bottlenecks. We point out the different availability and performance trade-offs achieved by the three platforms, and scenarios in which few faulty components cause cascading failures of the whole distributed system. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
ISSRE | 1 |
| 2020 | A comprehensive study on software aging across android versions and vendors
Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella, Roberto Pietrantuono |
Empir. Softw. Eng. | 1 |
| 2020 | Analyzing the Effects of Bugs on Software InterfacesabstractCritical systems that integrate software components (e.g., from third-parties) need to address the risk of residual software defects in these components. Software fault injection is an experimental solution to gauge such risk. Many error models have been proposed for emulating faulty components, such as by injecting error codes and exceptions, or by corrupting data with bit-flips, boundary values, and random values. Even if these error models have been able to find breaches in fragile systems, it is unclear whether these errors are in fact representative of software faults. To pursue this open question, we propose a methodology to analyze how software faults in C/C++ software components turn into errors at components' interfaces (interface error propagation), and present an experimental analysis on what, where, and when to inject interface errors. The results point out that the traditional error models, as used so far, do not accurately emulate software faults, but that richer interface errors need to be injected, by: injecting both fail-stop behaviors and data corruptions; targeting larger amounts of corrupted data structures; emulating silent data corruptions not signaled by the component; combining bit-flips, boundary values, and data perturbations. Roberto Natella, Stefan Winter 0001, Domenico Cotroneo, Neeraj Suri |
IEEE Trans. Software Eng. | 3 |
| 2019 | Analyzing the Context of Bug-Fixing Changes in the OpenStack Cloud Computing PlatformabstractMany research areas in software engineering, such as mutation testing, automatic repair, fault localization, and fault injection, rely on empirical knowledge about recurring bug-fixing code changes. Previous studies in this field focus on what has been changed due to bug-fixes, such as in terms of code edit actions. However, such studies did not consider where the bug-fix change was made (i.e., the context of the change), but knowing about the context can potentially narrow the search space for many software engineering techniques (e.g., by focusing mutation only on specific parts of the software). Furthermore, most previous work on bug-fixing changes focused on C and Java projects, but there is little empirical evidence about Python software. Therefore, in this paper we perform a thorough empirical analysis of bug-fixing changes in three OpenStack projects, focusing on both the what and the where of the changes. We observed that all the recurring change patterns are not oblivious with respect to the surrounding code, but tend to occur in specific code contexts. Domenico Cotroneo, Luigi De Simone, Antonio Ken Iannillo, Roberto Natella, Stefano Rosiello, Nematollah Bidokhti |
ISSRE | 1 |
| 2019 | Enhancing Failure Propagation Analysis in Cloud Computing SystemsabstractIn order to plan for failure recovery, the designers of cloud systems need to understand how their system can potentially fail. Unfortunately, analyzing the failure behavior of such systems can be very difficult and time-consuming, due to the large volume of events, non-determinism, and reuse of third-party components. To address these issues, we propose a novel approach that joins fault injection with anomaly detection to identify the symptoms of failures. We evaluated the proposed approach in the context of the OpenStack cloud computing platform. We show that our model can significantly improve the accuracy of failure analysis in terms of false positives and negatives, with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ISSRE | 1 |
| 2019 | How bad can a bug get? an empirical analysis of software failures in the OpenStack cloud computing platformabstractCloud management systems provide abstractions and APIs for programmatically configuring cloud infrastructures. Unfortunately, residual software bugs in these systems can potentially lead to high-severity failures, such as prolonged outages and data losses. In this paper, we investigate the impact of failures in the context widespread OpenStack cloud management system, by performing fault injection and by analyzing the impact of the resulting failures in terms of fail-stop behavior, failure detection through logging, and failure propagation across components. The analysis points out that most of the failures are not timely detected and notified; moreover, many of these failures can silently propagate over time and through components of the cloud management system, which call for more thorough run-time checks and fault containment. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ESEC/SIGSOFT FSE | 1 |
| 2019 | Privacy Preserving Intrusion Detection Via Homomorphic EncryptionabstractIn the recent years, we are assisting to an undiminished, and unlikely to stop number of cyber threats, that have increased the organizations/companies interest about security concerns. Further, the rising costs of an efficient IT security staff and environment is posing a significant challenge. These have created a new fast growing trend named Managed Security Services (MSS). Often customers turn to MSS providers to alleviate the pressures they face daily related to information security. One of the most critical aspect, related to the outsourcing of security issues, is privacy. Security monitoring and in general security services require access to as much data as possible, in order to provide an effective and reliable service. It is the well known conflict between privacy and security, a particularly evident problem in security monitoring solutions. This paper analyzes a scenario of MSS in order to provide a privacy preserving solution that allows the security monitoring without violating the privacy requirements. The basic idea relies on the usage of the Homomorphic Encryption technology. Encrypting data using homomorphic schemes, cloud computing and MSS providers can perform different computations on encrypted data without ever having access to their decryption. This solution keeps data confidential and secured, not only during exchange and storage, but also during processing. We provide an ad-hoc Intrusion Detection System architecture for privacy preserving security monitoring, considering as counter threats Code Injection attacks on homomorphically encrypted fields. Luigi Sgaglione, Luigi Coppolino, Salvatore D'Antonio, Giovanni Mazzeo, Luigi Romano, Domenico Cotroneo, Andrea Scognamiglio |
WETICE | 6 |
| 2019 | Evolutionary Fuzzing of Android OS Vendor System Services
Domenico Cotroneo, Antonio Ken Iannillo, Roberto Natella |
Empir. Softw. Eng. | 1 |
| 2019 | A framework for on-line timing error detection in software systems
Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia |
Future Gener. Comput. Syst. | 2 |
| 2019 | Overload control for virtual network functions under CPU contention
Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
Future Gener. Comput. Syst. | 1 |
| 2019 | Empirical Analysis and Validation of Security Alerts Filtering TechniquesabstractSystem administrators cope with security incidents through a variety of monitors, such as intrusion detection systems, event logs, security information and event management systems. Monitors generate large volumes of alerts that overwhelm the operations team and make forensics time-consuming. Filtering is a consolidated technique to reduce the amount of alerts. In spite of the number of filtering proposals, few studies have addressed the validation of filtering results in real production datasets. This paper analyzes a number of state-of-the-art filtering techniques that are used to address security datasets. We use 14 months of alerts generated in a SaaS Cloud. Our analysis aims to measure and compare the reduction of the alerts volume obtained by the filters. The analysis highlights pros and cons of each filter and provides insights into the practical implications of filtering as affected by the characteristics of a dataset. We complement the analysis with a method to validate the output of a filter in absence of ground truth, i.e., the knowledge of the incidents occurred in the system at the time the alerts were generated. The analysis addresses blacklist, conceptual clustering and bytes techniques, and our filtering proposal based on term weighting. Domenico Cotroneo, Andrea Paudice, Antonio Pecchia |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Faultprog: Testing the Accuracy of Binary-Level Software Fault InjectionabstractOff-The-Shelf (OTS) software components are the cornerstone of modern systems, including safety-critical ones. However, the dependability of OTS components is uncertain due to the lack of source code, design artifacts and test cases, since only their binary code is supplied. Fault injection in components' binary code is a solution to understand the risks posed by buggy OTS components. In this paper, we consider the problem of the accurate mutation of binary code for fault injection purposes. Fault injection emulates bugs in high-level programming constructs (assignments, expressions, function calls, ...) by mutating their translation in binary code. However, the semantic gap between the source code and its binary translation often leads to inaccurate mutations. We propose Faultprog, a systematic approach for testing the accuracy of binary mutation tools. Faultprog automatically generates synthetic programs using a stochastic grammar, and mutates both their binary code with the tool under test, and their source code as reference for comparisons. Moreover, we present a case study on a commercial binary mutation tool, where Faultprog was adopted to identify code patterns and compiler optimizations that affect its mutation accuracy. Domenico Cotroneo, Anna Lanzaro, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Guest Editors' Introduction: Special Issue on Data-Driven Dependability and SecurityabstractThe six papers in this special issue aim to concentrate novel contributions addressing dependability and security of computer systems through data analysis, and to publish consolidated research results focusing on data-driven methodologies, measurements from production systems, and analysis of large datasets. This information provides valuable contributions related to log-based measurements, operating systems dependability, and attack detection. Domenico Cotroneo, Karthik Pattabiraman, Antonio Pecchia |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Run-Time Detection of Protocol Bugs in Storage I/O Device DriversabstractProtocol violation bugs in storage device drivers are a critical threat for data integrity, since these bugs can silently corrupt the commands and data flowing between the OS and storage devices. Due to their nature, these bugs are notoriously difficult to find by traditional testing. In this paper, we propose a run-time monitoring approach for storage device drivers, in order to detect I/O protocol violations that would otherwise silently escalate in corruptions of users' data. The monitoring approach detects violations of I/O protocols by automatically learning a reference model from failure-free execution traces. The approach focuses on selected portions of the storage controller interface, in order to achieve a good tradeoff in terms of low performance overhead and high coverage and accuracy of failure detection. We assess these properties on three real-world storage device drivers from the Linux kernel, through fault injection and stress tests. Moreover, we show that the monitoring approach only requires few minutes of training workload, and that it is robust to differences between the operational and the training workloads. Domenico Cotroneo, Luigi De Simone, Roberto Natella |
IEEE Trans. Reliab. | 1 |
| 2017 | A Fault Correlation Approach to Detect Performance Anomalies in Virtual Network Function ChainsabstractNetwork Function Virtualization is an emerging paradigm to allow the creation, at software level, of complex network services by composing simpler ones. However, this paradigm shift exposes network services to faults and bottlenecks in the complex software virtualization infrastructure they rely on. Thus, NFV services require effective anomaly detection systems to detect the occurrence of network problems. The paper proposes a novel approach to ease the adoption of anomaly detection in production NFV services, by avoiding the need to train a model or to calibrate a threshold. The approach infers the service health status by collecting metrics from multiple elements in the NFV service chain, and by analyzing their (lack of) correlation over the time. We validate this approach on an NFV-oriented Interactive Multimedia System, to detect problems affecting the quality of service, such as the overload, component crashes, avalanche restarts and physical resource contention. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
ISSRE | 1 |
| 2017 | Chizpurfle: A Gray-Box Android Fuzzer for Vendor Service CustomizationsabstractAndroid has become the most popular mobile OS, as it enables device manufacturers to introduce customizations to compete with value-added services. However, customizations make the OS less dependable and secure, since they can introduce software flaws. Such flaws can be found by using fuzzing, a popular testing technique among security researchers.This paper presents Chizpurfle, a novel "gray-box" fuzzing tool for vendor-specific Android services. Testing these services is challenging for existing tools, since vendors do not provide source code and the services cannot be run on a device emulator. Chizpurfle has been designed to run on an unmodified Android OS on an actual device. The tool automatically discovers, fuzzes, and profiles proprietary services. This work evaluates the applicability and performance of Chizpurfle on the Samsung Galaxy S6 Edge, and discusses software bugs found in privileged vendor services. Antonio Ken Iannillo, Roberto Natella, Domenico Cotroneo, Cristina Nita-Rotaru |
ISSRE | 3 |
| 2017 | Secure crisis information sharing through an interoperability framework among first responders: The SECTOR practical experienceabstractThe increasing occurrence of large scale disasters calls out for a collaborative approach to crisis management, where multiple and heterogeneous organizations of first responders are deployed within the damaged area and must interact with each others in order to cooperate in the damage assessment and recovery actions. Such an approach requires a suitable communication platform to allow these organizations to exchange crisis information among their members, despite their heterogeneity. Current research is investigating such point and several solutions have been proposed; however, there are other key requirements that such a platform needs to address in order to be successfully used in practical cases. Among these requirements, security plays a key role. This paper introduces the issue of confidential and private communications for platforms supporting collaborative crisis management, and identifies a possible solution developed within the context of the EU-funded project named SECTOR. Marcello Cinque, Domenico Cotroneo, Christian Esposito 0001, Mario Fiorentino |
WiMob | 2 |
| 2017 | Debugging-workflow-aware software reliability growth analysisabstractSummary Software reliability growth models support the prediction/assessment of product quality, release time, and testing/debugging cost. Several software reliability growth model extensions take into account the bug correction process. However, their estimates may be significantly inaccurate when debugging fails to fully fit modelling assumptions. This paper proposes debugging‐workflow‐aware software reliability growth method (DWA‐SRGM), a method for reliability growth analysis leveraging the debugging data usually managed by companies in bug tracking systems. On the basis of a characterization of the debugging workflow within the software project under consideration (in terms of bug features and treatment phases), DWA‐SRGM pinpoints the factors impacting the estimates and to spot bottlenecks, thus supporting process improvement decisions. Two industrial case studies are presented, a customer relationship management system and an enterprise resource planning system, whose defects span a period of about 17 and 13 months, respectively. DWA‐SRGM revealed effective to obtain more realistic estimates and to capitalize on the awareness of critical factors for improving debugging. Marcello Cinque, Domenico Cotroneo, Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001 |
Softw. Test. Verification Reliab. | 2 |
| 2017 | NFV-Throttle: An Overload Control Framework for Network Function VirtualizationabstractNetwork function virtualization (NFV) aims to provide high-performance network services through cloud computing and virtualization technologies. However, network overloads represent a major challenge. While elastic cloud computing can partially address overloads by scaling on-demand, this mechanism is not quick enough to meet the strict high-availability requirements of “carrier-grade” telecom services. Thus, in this paper we propose a novel overload control framework (NFV-Throttle) to protect NFV services from failures due to an excess of traffic in the short term, by filtering the incoming traffic toward virtual network functions (VNFs) to make the best use of the available capacity, and to preserve the QoS of traffic flows admitted in the network. Moreover, the framework has been designed to fit the service models of NFV, including VNFaaS and NFVIaaS. We present an extensive experimental evaluation on the NFV-oriented Clearwater IMS, showing that the solution is robust and able to sustain severe overload conditions with a very small performance overhead. Domenico Cotroneo, Roberto Natella, Stefano Rosiello |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2017 | NFV-Bench: A Dependability Benchmark for Network Function Virtualization SystemsabstractNetwork function virtualization (NFV) envisions the use of cloud computing and virtualization technology to reduce costs and innovate network services. However, this paradigm shift poses the question whether NFV will be able to fulfill the strict performance and dependability objectives required by regulations and customers. Thus, we propose a dependability benchmark to support NFV providers at making informed decisions about which virtualization, management, and application-level solutions can achieve the best dependability. We define in detail the use cases, measures, and faults to be injected. Moreover, we present a benchmarking case study on two alternative, production-grade virtualization solutions, namely VMware ESXi/vSphere (hypervisor-based) and Linux/Docker (container-based), on which we deploy an NFV-oriented IMS system. Despite the promise of higher performance and manageability, our experiments suggest that the container-based configuration can be less dependable than the hypervisor-based one, and point out which faults NFV designers should address to improve dependability. Domenico Cotroneo, Luigi De Simone, Roberto Natella |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2016 | Software Aging Analysis of the Android Mobile OSabstractMobile devices are significantly complex, feature-rich, and heavily customized, thus they are prone to software reliability and performance issues. This paper considers the problem of software aging in Android mobile OS, which causes the device to gradually degrade in responsiveness, and to eventually fail. We present a methodology to identify factors (such as workloads and device configurations) and resource utilization metrics that are correlated with software aging. Moreover, we performed an empirical analysis of recent Android devices, finding that software aging actually affects them. The analysis pointed out processes and components of the Android OS affected by software aging, and metrics useful as indicators of software aging to schedule software rejuvenation actions. Domenico Cotroneo, Francesco Fucci, Antonio Ken Iannillo, Roberto Natella, Roberto Pietrantuono |
ISSRE | 1 |
| 2016 | Automated root cause identification of security alerts: Evaluation in a SaaS Cloud
Domenico Cotroneo, Andrea Paudice, Antonio Pecchia |
Future Gener. Comput. Syst. | 1 |
| 2016 | How do bugs surface? A comprehensive study on the characteristics of software bugs manifestation
Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001, Kishor S. Trivedi |
J. Syst. Softw. | 1 |
| 2016 | To Cloudify or Not to Cloudify: The Question for a Scientific Data CenterabstractThe idea of turning data centers executing scientific batch jobs into private clouds is as attractive as troubling. Cloud platforms may help both in limiting power consumption and in implementing fault tolerance strategies. However, there is also the fear that performance may worsen, and that the electricity required for longer job duration and fault tolerance implementation may overcome the saved one. In this paper, we present the consumability analysis for assessing the impact of cloud and fault tolerance tunings on scientific processing systems. The analysis considers performance, consumption, and dependability aspects, jointly. The aim is to pinpoint if, for a given system, there is a setting where consumption and job failure rate decrease, while performance is not affected. Applied to the scientific data center at our University, the analysis allowed us to find the proper selection of virtual machines' configuration, consolidation strategy, and fault tolerance tuning. Marcello Cinque, Domenico Cotroneo, Flavio Frattini, Stefano Russo 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2016 | Characterizing Direct Monitoring Techniques in Software SystemsabstractMonitoring is a consolidated practice to characterize the dependability behavior of a software system. A variety of techniques, such as event logging and operating system probes, are currently used to generate monitoring data for troubleshooting and failure analysis. In spite of the importance of monitoring, whose role can be essential in critical software systems, there is a lack of studies addressing the assessment and the comparison of the techniques aiming to monitor the occurrence of failures during operations. This paper proposes a method to characterize the monitoring techniques implemented in a software system. The method is based on a fault injection approach and allows measuring 1) precision and recall of a monitoring technique and 2) the dissimilarity of the data it generates upon failures. The method has been used in two critical software systems implementing event logging, assertion checking, and source code instrumentation techniques. We analyzed a total of 3 844 failures. With respect to our data, we observed that the effectiveness of a technique is strongly affected by the system and type of failure, and that the combination of different techniques is potentially beneficial to increase the overall failure reporting ability. More important, our analysis revealed a number of practical implications to be taken into account when developing a monitoring technique. Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia |
IEEE Trans. Reliab. | 2 |
| 2016 | RELAI Testing: A Technique to Assess and Improve Software ReliabilityabstractTesting software to assess or improve reliability presents several practical challenges. Conventional operational testing is a fundamental strategy that simulates the real usage of the system in order to expose failures with the highest occurrence probability. However, practitioners find it unsuitable for assessing/achieving very high reliability levels; also, they do not see the adoption of a “real” usage profile estimate as a sensible idea, being it a source of non-quantifiable uncertainty. Oppositely, debug testing aims to expose as many failures as possible, but regardless of their impact on runtime reliability. These strategies are used either to assess or to improve reliability, but cannot improve and assess reliability in the same testing session. This article proposes Reliability Assessment and Improvement (RELAI) testing, a new technique thought to improve the delivered reliability by an adaptive testing scheme, while providing, at the same time, a continuous assessment of reliability attained through testing and fault removal. The technique also quantifies the impact of a partial knowledge of the operational profile. RELAI is positively evaluated on four software applications compared, in separate experiments, with techniques conceived either for reliability improvement or for reliability assessment, demonstrating substantial improvements in both cases. Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2015 | Impact of Malfunction on the Energy Efficiency of Batch Processing SystemsabstractEnergy efficiency of large processing systems is usually assessed as the relation between a performance and a power consumption metric, neglecting malfunction. Execution failures have a tangible cost in terms of wasted energy, however. They are often managed through fault tolerance mechanisms, which in turn consume electricity. We introduce the consumability attribute for batch processing systems, encompassing performance, consumption, and dependability aspects altogether. We propose a metric for its quantification and a methodology for its analysis. Using a real 500-node batch system as a case study, we show that consumability is representative of both efficiency and effectiveness, and we show the usefulness of the proposed metric and the suitability of the proposed methodology. Marcello Cinque, Domenico Cotroneo, Flavio Frattini, Stefano Russo 0001 |
DSN | 2 |
| 2015 | Industry Practices and Event Logging: Assessment of a Critical Software Development ProcessabstractPractitioners widely recognize the importance of event logging for a variety of tasks, such as accounting, system measurements and troubleshooting. Nevertheless, in spite of the importance of the tasks based on the logs collected under real workload conditions, event logging lacks systematic design and implementation practices. The implementation of the logging mechanism strongly relies on the human expertise. This paper proposes a measurement study of event logging practices in a critical industrial domain. We assess a software development process at Selex ES, a leading Finmeccanica company in electronic and information solutions for critical systems. Our study combines source code analysis, inspection of around 2.3 millions log entries, and direct feedback from the development team to gain process-wide insights ranging from programming practices, logging objectives and issues impacting log analysis. The findings of our study were extremely valuable to prioritize event logging reengineering tasks at Selex ES. Antonio Pecchia, Marcello Cinque, Gabriella Carrozza, Domenico Cotroneo |
ICSE (2) | 4 |
| 2015 | No PAIN, No Gain? The Utility of PArallel Fault INjectionsabstractSoftware Fault Injection (SFI) is an established technique for assessing the robustness of a software under test by exposing it to faults in its operational environment. Depending on the complexity of this operational environment, the complexity of the software under test, and the number and type of faults, a thorough SFI assessment can entail (a) numerous experiments and (b) long experiment run times, which both contribute to a considerable execution time for the tests. In order to counteract this increase when dealing with complex systems, recent works propose to exploit parallel hardware to execute multiple experiments at the same time. While Parallel fault Injections (PAIN) yield higher experiment throughput, they are based on an implicit assumption of non-interference among the simultaneously executing experiments. In this paper we investigate the validity of this assumption and determine the trade-off between increased throughput and the accuracy of experimental results obtained from PAIN experiments. Stefan Winter 0001, Oliver Schwahn, Roberto Natella, Neeraj Suri, Domenico Cotroneo |
ICSE (1) | 5 |
| 2015 | MoIO: Run-time monitoring for I/O protocol violations in storage device driversabstractBugs affecting storage device drivers include the so-called protocol violation bugs, which silently corrupt data and commands exchanged with I/O devices. Protocol violations are very difficult to prevent, since testing device driver is notoriously difficult. To address them, we present a monitoring approach for device drivers (MoIO) to detect HO protocol violations at run-time. The approach infers a model of the interactions between the storage device driver, the OS kernel, and the hardware (the device driver protocol) by analyzing execution traces. The model is then used as a reference for detecting violations in production. The approach has been designed to have a low overhead and to overcome the lack of source code and protocol documentation. We show that the approach is feasible and effective by applying it on the SATA/AHCI storage device driver of the Linux kernel, and by performing fault injection and long-running tests. Domenico Cotroneo, Luigi De Simone, Francesco Fucci, Roberto Natella |
ISSRE | 1 |
| 2015 | Dependability evaluation and benchmarking of Network Function Virtualization InfrastructuresabstractNetwork Function Virtualization (NFV) is an emerging solution that aims at improving the flexibility, the efficiency and the manageability of networks, by leveraging virtualization and cloud computing technologies to run network appliances in software. However, the “softwarization” of network functions raises reliability concerns, as they will be exposed to faults in commodity hardware and software components. In this paper, we propose a methodology for the dependability evaluation and benchmarking of NFV Infrastructures (NFVIs), based on fault injection. We discuss the application of the methodology in the context of a virtualized IP Multimedia Subsystem (IMS), and the pitfalls in the design of a reliable NFVI. Domenico Cotroneo, Luigi De Simone, Antonio Ken Iannillo, Anna Lanzaro, Roberto Natella |
NetSoft | 1 |
| 2014 | What Logs Should You Look at When an Application Fails? Insights from an Industrial Case StudyabstractEvent logs are the first place where to find useful information about application failures. Event logs are available at different system levels, such as application, middleware and operating system. In this paper we analyze the failure reporting capability of event logs collected at different levels of an industrial system in the Air Traffic Control (ATC) domain. The study is based on a data set of 3,159 failures induced in the system by means of software fault injection. Results indicate that the reporting ability of event logs collected at a given level is strongly affected by the type of failure observed at runtime. For example, even if operating system logs catch almost all application crashes, they are strongly ineffective in face of silent and erratic failures in the considered system. Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia |
DSN | 2 |
| 2014 | Assessing Direct Monitoring Techniques to Analyze Failures of Critical Industrial SystemsabstractThe analysis of monitoring data is extremely valuable for critical computer systems. It allows to gain insights into the failure behavior of a given system under real workload conditions, which is crucial to assure service continuity and downtime reduction. This paper proposes an experimental evaluation of different direct monitoring techniques, namely event logs, assertions, and source code instrumentation, that are widely used in the context of critical industrial systems. We inject 12,733 software faults in a real-world air traffic control (ATC) middleware system with the aim of analyzing the ability of mentioned techniques to produce information in case of failures. Experimental results indicate that each technique is able to cover a limited number of failure manifestations. Moreover, we observe that the quality of collected data to support failure diagnosis tasks strongly varies across the techniques considered in this study. Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia |
ISSRE | 2 |
| 2014 | An empirical study of injected versus actual interface errorsabstractThe reuse of software components is a common practice in commercial applications and increasingly appearing in safety critical systems as driven also by cost considerations. This practice puts dependability at risk, as differing operating conditions in different reuse scenarios may expose residual software faults in the components. Consequently, software fault injection techniques are used to assess how residual faults of reused software components may affect the system, and to identify appropriate counter-measures. As fault injection in components’ code suffers from a number of practical disadvantages, it is often replaced by error injection at the component interface level. However, it is still an open issue, whether such injected errors are actually representative of the effects of residual faults. To this end, we propose a method for analyzing how software faults turn into interface errors, with the ultimate aim of supporting more representative interface error injection experiments. Our analysis in the context of widely used software libraries reveals that existing interface error models are not suitable for emulating software faults, and provides useful insights for improving the representativeness of interface error injection. Anna Lanzaro, Roberto Natella, Stefan Winter 0001, Domenico Cotroneo, Neeraj Suri |
ISSTA | 4 |
| 2014 | A survey of software aging and rejuvenation studiesabstractSoftware aging is a phenomenon plaguing many long-running complex software systems, which exhibit performance degradation or an increasing failure rate. Several strategies based on the proactive rejuvenation of the software state have been proposed to counteract software aging and prevent failures. This survey article provides an overview of studies on Software Aging and Rejuvenation (SAR) that have appeared in major journals and conference proceedings, with respect to the statistical approaches that have been used to forecast software aging phenomena and to plan rejuvenation, the kind of systems and aging effects that have been studied, and the techniques that have been proposed to rejuvenate complex software systems. The analysis is useful to identify key results from SAR research, and it is leveraged in this article to highlight trends and open issues. Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2013 | Towards secure monitoring and control systems: Diversify!abstractCyber attacks have become surprisingly sophisticated over the past fifteen years. While early infections mostly targeted individual machines, recent threats leverage the widespread network connectivity to develop complex and highly coordinated attacks involving several distributed nodes [1]. Attackers are currently targeting very diverse domains, e.g., e-commerce systems, corporate networks, datacenter facilities and industrial systems, to achieve a variety of objectives, which range from credentials compromise to sabotage of physical devices, by means of smarter and smarter worms and rootkits. Stuxnet is a recent worm that well emphasizes the strong technical advances achieved by the attackers' community. It was discovered in July 2010 and firstly affected Iranian nuclear plants [2]. Stuxnet compromises the regular behavior of the supervisory control and data acquisition (SCADA) system by reprogramming the code of programmable logic controllers (PLC). Once compromised, PLCs can progressively destroy a device (e.g., components of a centrifuge, such as the case of the Iranian plant) by sending malicious control signals. Stuxnet combines a relevant number of challenging features: it exploits zero-days vulnerabilities of the Windows OS to affect the nodes connected to the PLC; it propagates either locally (e.g., by means of USB sticks) or remotely (e.g., via shared folders or the print spooler vulnerability); it is able to modify its behavior during the progression of the attack, and communicates with a remote command and control server. More importantly, Stuxnet can remain undetected for many months [3] because it is able to fool the SCADA system by emulating regular monitoring signals. Domenico Cotroneo, Antonio Pecchia, Stefano Russo 0001 |
DSN | 1 |
| 2013 | A learning-based method for combining testing techniquesabstractThis work presents a method to combine testing techniques adaptively during the testing process. It intends to mitigate the sources of uncertainty of software testing processes, by learning from past experience and, at the same time, adapting the technique selection to the current testing session. The method is based on machine learning strategies. It uses offline strategies to take historical information into account about the techniques performance collected in past testing sessions; then, online strategies are used to adapt the selection of test cases to the data observed as the testing proceeds. Experimental results show that techniques performance can be accurately characterized from features of the past testing sessions, by means of machine learning algorithms, and that integrating this result into the online algorithm allows improving the fault detection effectiveness with respect to single testing techniques, as well as to their random combination. Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 1 |
| 2013 | Analysis and Prediction of Mandelbugs in an Industrial Software SystemabstractMandelbugs are faults that are triggered by complex conditions, such as interaction with hardware and other software, and timing or ordering of events. These faults are considerably difficult to detect with traditional testing techniques, since it can be challenging to control their complex triggering conditions in a testing environment. Therefore, it is necessary to adopt specific verification and/or fault-tolerance strategies for dealing with them in a cost-effective way. In this paper, we investigate how to predict the location of Mandelbugs in complex software systems, in order to focus V&V activities and fault tolerance mechanisms in those modules where Mandelbugs are most likely present. In the context of an industrial complex software system, we empirically analyze Mandelbugs, and investigate an approach for Mandelbug prediction based on a set of novel software complexity metrics. Results show that Mandelbugs account for a noticeable share of faults, and that the proposed approach can predict Mandelbug-prone modules with greater accuracy than the sole adoption of traditional software metrics. Gabriella Carrozza, Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ICST | 2 |
| 2013 | Fault triggers in open-source software: An experience reportabstractWith software systems becoming increasingly large and complex, many difficulties in coping with software bugs arise for developers. Despite good development practices, thorough testing, and proper maintenance policies, a non-negligible number of bugs remain in the released software. Understanding the type of residual bugs is fundamental for adopting proper countermeasures in current and future software releases. Depending on the fault triggering conditions that lead to a failure, developers can introduce fault-tolerance mechanisms and plan verification and validation strategies. In this paper, we analyze bugs in four large open-source software systems during their lifecycle, based on the concept of fault triggers. We first investigate how the type of system affects the bug type proportions, and their evolution over years. Then, an analysis of bug subtypes is performed, so as to better understand their nature, followed by a comparison with respect to attributes such as their average time to fix and severity. Domenico Cotroneo, Michael Grottke, Roberto Natella, Roberto Pietrantuono, Kishor S. Trivedi |
ISSRE | 1 |
| 2013 | SABRINE: State-based robustness testing of operating systemsabstractThe assessment of operating systems robustness with respect to unexpected or anomalous events is a fundamental requirement for mission-critical systems. Robustness can be tested by deliberately exposing the system to erroneous events during its execution, and then analyzing the OS behavior to evaluate its ability to gracefully handle these events. Since OSs are complex and stateful systems, robustness testing needs to account for the timing of erroneous events, in order to evaluate the robust behavior of the OS under different states. This paper presents SABRINE (StAte-Based Robustness testIng of operatiNg systEms), an approach for state-aware robustness testing of OSs. SABRINE automatically extracts state models from execution traces, and generates a set of test cases that cover different OS states. We evaluate the approach on a Linux-based Real-Time Operating System adopted in the avionic domain. Experimental results show that SABRINE can automatically identify relevant OS states, and find robustness vulnerabilities while keeping low the number of test cases. Domenico Cotroneo, Domenico Di Leo, Francesco Fucci, Roberto Natella |
ASE | 1 |
| 2013 | State-Driven Testing of Distributed Systems
Domenico Cotroneo, Roberto Natella, Stefano Russo 0001, Fabio Scippacercola |
OPODIS | 1 |
| 2013 | On reliability in publish/subscribe services
Christian Esposito 0001, Domenico Cotroneo, Stefano Russo 0001 |
Comput. Networks | 2 |
| 2013 | Testing techniques selection based on ODC fault types and software metrics
Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 1 |
| 2013 | Predicting aging-related bugs using software complexity metrics
Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono |
Perform. Evaluation | 1 |
| 2013 | A measurement-based ageing analysis of the JVMabstractSUMMARY In this work, a software ageing analysis of Java‐based software systems is conducted. The JVM is the core layer in Java‐based systems, and its dependability greatly affects the overall system quality. Starting from an experimental campaign on a real‐world test bed, this work isolates the contribution of the JVM to the overall ageing trend, and identifies, through statistical methods, which workload parameters are the most relevant to ageing dynamics. Results revealed the presence of several ageing dynamics in the JVM, including (i) a throughput loss trend mainly dependent on the execution unit, (ii) a slow memory depletion drift due to the just‐in‐time‐compiler activity and (iii) a fast memory depletion drift caused by dynamics inside the garbage collector. The outlined procedure and obtained results are useful in order to (i) identify the presence of ageing phenomena, (ii) perform online ageing detection and time‐to‐exhaustion prediction and (iii) define optimal rejuvenation techniques. Copyright © 2011 John Wiley & Sons, Ltd. Domenico Cotroneo, Salvatore Orlando 0002, Roberto Pietrantuono, Stefano Russo 0001 |
Softw. Test. Verification Reliab. | 1 |
| 2013 | Combining Operational and Debug Testing for Improving ReliabilityabstractThis paper addresses the challenge of reliability-driven testing, i.e., of testing software systems with the specific objective of increasing its operational reliability. We first examined the most relevant approach oriented toward this goal, namely operational testing. The main issues that in the past hindered its wide-scale adoption and practical application are first discussed, followed by the analysis of its performance under different conditions and configurations. Then, a new approach conceived to overcome the limits of operational testing in delivering high reliability is proposed. The two testing strategies are evaluated probabilistically, and by simulation. Results report on the performance of operational testing when several involved parameters are taken into account, and on the effectiveness of the new proposed approach in achieving better reliability. At a higher level, the findings of the paper also suggest that a different view of the testing for reliability improvement concept may help to devise new testing approaches for high-reliability, demanding systems. Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
IEEE Trans. Reliab. | 1 |
| 2013 | Event Logs for the Analysis of Software Failures: A Rule-Based ApproachabstractEvent logs have been widely used over the last three decades to analyze the failure behavior of a variety of systems. Nevertheless, the implementation of the logging mechanism lacks a systematic approach and collected logs are often inaccurate at reporting software failures: This is a threat to the validity of log-based failure analysis. This paper analyzes the limitations of current logging mechanisms and proposes a rule-based approach to make logs effective to analyze software failures. The approach leverages artifacts produced at system design time and puts forth a set of rules to formalize the placement of the logging instructions within the source code. The validity of the approach, with respect to traditional logging mechanisms, is shown by means of around 12,500 software fault injection experiments into real-world systems. Marcello Cinque, Domenico Cotroneo, Antonio Pecchia |
IEEE Trans. Software Eng. | 2 |
| 2013 | On Fault Representativeness of Software Fault InjectionabstractThe injection of software faults in software components to assess the impact of these faults on other components or on the system as a whole, allowing the evaluation of fault tolerance, is relatively new compared to decades of research on hardware fault injection. This paper presents an extensive experimental study (more than 3.8 million individual experiments in three real systems) to evaluate the representativeness of faults injected by a state-of-the-art approach (G-SWFIT). Results show that a significant share (up to 72 percent) of injected faults cannot be considered representative of residual software faults as they are consistently detected by regression tests, and that the representativeness of injected faults is affected by the fault location within the system, resulting in different distributions of representative/nonrepresentative faults across files and functions. Therefore, we propose a new approach to refine the faultload by removing faults that are not representative of residual software faults. This filtering is essential to assure meaningful results and to reduce the cost (in terms of number of faults) of software fault injection campaigns in complex software. The proposed approach is based on classification algorithms, is fully automatic, and can be used for improving fault representativeness of existing software fault injection approaches. Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira |
IEEE Trans. Software Eng. | 2 |
| 2012 | Assessing time coalescence techniques for the analysis of supercomputer logsabstractThis paper presents a novel approach to assess time coalescence techniques. These techniques are widely used to reconstruct the failure process of a system and to estimate dependability measurements from its event logs. The approach is based on the use of automatically generated logs, accompanied by the exact knowledge of the ground truth on the failure process. The assessment is conducted by comparing the presumed failure process, reconstructed via coalescence, with the ground truth. We focus on supercomputer logs, due to increasing importance of automatic event log analysis for these systems. Experimental results show how the approach allows to compare different time coalescence techniques and to identify their weaknesses with respect to given system settings. In addition, results revealed an interesting correlation between errors caused by the coalescence and errors in the estimation of dependability measurements. Catello Di Martino, Marcello Cinque, Domenico Cotroneo |
DSN | 3 |
| 2012 | On the Aging Effects Due to Concurrency Bugs: A Case Study on MySQLabstractThis study investigates software aging effects caused by the activation of concurrency bugs in a wellknown database management system (DBMS), namely MySQL. Experiments with different workloads are performed in order to reproduce the most likely conditions for concurrency bugs activation. Besides the typical aging effects observed in many operational systems (i.e., a gradual degradation over time), results highlight that both available resources and DBMS performance (e.g. service rate, service time, and connection latency) can decrease with time in a hard-to-predict way. We observed that, due to the activation of concurrency bug, the DBMS enters a degraded state in which: i) the estimation of Time-To-Failure (TTF) by means of memory depletion trend analysis is highly inaccurate, and ii) the failure rate does not depend on the instantaneous and/or mean accumulated work. Results suggest that, in such cases, finer-grained indicators and/or different techniques need to be taken into account for properly preventing failures. Antonio Bovenzi, Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 2 |
| 2012 | Automated Generation of Performance and Dependability Models for the Assessment of Wireless Sensor NetworksabstractWireless Sensor Networks (WSNs) are widely recognized as a promising solution to build next-generation monitoring systems. Their industrial uptake is however still compromised by the low level of trust on their performance and dependability. Whereas analytical models represent a valid mean to assess nonfunctional properties via simulation, their wide use is still limited by the complexity and dynamicity of WSNs, which lead to unaffordable modeling costs. To reduce this gap between research achievements and industrial development, this paper presents a framework for the assessment of WSNs based on the automated generation of analytical models. The framework hides modeling details, and it allows designers to focus on simulation results to drive their design choices. Models are generated starting from a high-level specification of the system and by a preliminary characterization of its fault-free behavior, using behavioral simulators. The benefits of the framework are shown in the context of two case studies, based on the wireless monitoring of civil structures. Catello Di Martino, Marcello Cinque, Domenico Cotroneo |
IEEE Trans. Computers | 3 |
| 2011 | 5th international workshop on adaptive and dependable mobile ubiquitous systems ADAMUS 2011abstractThe vision of mobile and ubiquitous systems is becoming a reality thanks to the recent advances in wireless communication and device miniaturization. However, the widespread industrial uptake of these systems is still compromised by the highly error-prone and heterogeneous mobile provisioning environments, which induce several impairments to normal operation. Thus, how to improve the dependability of these systems is still an open issue. Domenico Cotroneo, Vincenzo De Florio |
DSN | 1 |
| 2011 | Improving Log-based Field Failure Data Analysis of multi-node computing systemsabstractLog-based Field Failure Data Analysis (FFDA) is a widely-adopted methodology to assess dependability properties of an operational system. A key step in FFDA is filtering out entries that are not useful and redundant error entries from the log. The latter is challenging: a fault, once triggered, can generate multiple errors that propagate within the system. Grouping the error entries related to the same fault manifestation is crucial to obtain realistic measurements. This paper deals with the issues of the tuple heuristic, used to group the error entries in the log, in multi-node computing systems. We demonstrate that the tuple heuristic can group entries incorrectly; thus, an improved heuristic that adopts statistical indicators is proposed. We assess the impact of inaccurate grouping on dependability measurements by comparing the results obtained with both the heuristics. The analysis encompasses the log of the Mercury cluster at the National Center for Supercomputing Applications. Antonio Pecchia, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 2 |
| 2011 | Workload Characterization for Software Aging AnalysisabstractThe phenomenon of software aging is increasingly recognized as a relevant problem of long-running systems. Numerous experiments have been carried out in the last decade to empirically analyze software aging. Such experiments, besides highlighting the relevance of the phenomenon, have shown that aging is tightly related to the applied workload. However, due to the differences among the experimented applications and among the experimental conditions, results of past studies are not comparable to each other. This prevent from drawing general conclusions (e.g., about the aging-workload relationship), and from comparing systems from the aging perspective. In this paper, we propose a procedure to carry out aging experiments in different applications for: i) assessing aging trend of the individual systems, as well as assessing differences among them (i.e., obtaining comparable results), ii) inferring workload-aging relationships from experiments performed on different applications, by highlighting the most relevant workload parameters. The procedure is applied, through a set of long-running experiments, to three real-scale software applications, namely Apache Web Server, James Mail Server, and CARDAMOM, a middleware for the development of air traffic control (ATC) systems. Antonio Bovenzi, Domenico Cotroneo, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 2 |
| 2011 | A Case Study on State-Based Robustness Testing of an Operating System for the Avionic Domain
Domenico Cotroneo, Domenico Di Leo, Roberto Natella, Roberto Pietrantuono |
SAFECOMP | 1 |
| 2011 | Identifying Compromised Users in Shared Computing Infrastructures: A Data-Driven Bayesian Network ApproachabstractThe growing demand for processing and storage capabilities has led to the deployment of high-performance computing infrastructures. Users log into the computing infrastructure remotely, by providing their credentials (e.g., username and password), through the public network and using well-established authentication protocols, e.g., SSH. However, user credentials can be stolen and an attacker (using a stolen credential) can masquerade as the legitimate user and penetrate the system as an insider. This paper deals with security incidents initiated by using stolen credentials and occurred during the last three years at the National Center for Supercomputing Applications (NCSA) at the University of Illinois. We analyze the key characteristics of the security data produced by the monitoring tools during the incidents and use a Bayesian network approach to correlate (i) data provided by different security tools (e.g., IDS and Net Flows) and (ii) information related to the users' profiles to identify compromised users, i.e., the users whose credentials have been stolen. The technique is validated with the real incident data. The experimental results demonstrate that the proposed approach is effective in detecting compromised users, while allows eliminating around 80% of false positives (i.e., not compromised user being declared compromised). Antonio Pecchia, Aashish Sharma, Zbigniew T. Kalbarczyk, Domenico Cotroneo, Ravishankar K. Iyer |
SRDS | 4 |
| 2010 | Assessing and improving the effectiveness of logs for the analysis of software faultsabstractEvent logs are the primary source of data to characterize the dependability behavior of a computing system during the operational phase. However, they are inadequate to provide evidence of software faults, which are nowadays among the main causes of system outages. This paper proposes an approach based on software fault injection to assess the effectiveness of logs to keep track of software faults triggered in the field. Injection results are used to provide guidelines to improve the ability of logging mechanisms to report the effects of software faults. The benefits of the approach are shown by means of experimental results on three widely used software systems. Marcello Cinque, Domenico Cotroneo, Roberto Natella, Antonio Pecchia |
DSN | 2 |
| 2010 | Representativeness analysis of injected software faults in complex softwareabstractDespite of the existence of several techniques for emulating software faults, there are still open issues regarding representativeness of the faults being injected. An important aspect, not considered by existing techniques, is the non-trivial activation condition (trigger) of real faults, which causes them to elude testing and remain hidden until operation. In this paper, we investigate how the representativeness of injected software faults can be improved regarding the representativeness of triggers, by proposing a set of generic criteria to select representative faults from afaultload. We used the G-SWFIT technique to inject software faults in a DBMS, resulting in over 40 thousands faults and 2 million runs of a real test suite. We analyzed faults with respect to their triggers, and concluded that a non-negligible share (15%) would not realistically elude testing. Our proposed criteria decreased the percentage of non-elusive faults in the faultload, improving its representativeness. Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira |
DSN | 2 |
| 2010 | An effective approach for injecting faults in wireless sensor network operating systemsabstractThis paper presents an effective approach for injecting faults/errors in WSN nodes operating systems. The approach is based on the injection of faults at the assembly level. Results show that depending on the concurrency model and on the memory management, the operating systems react to injected errors differently, indicating that fault containment strategies and hang-checking assertions should be implemented to avoid spreading and activations of errors. Marcello Cinque, Domenico Cotroneo, Catello Di Martino, Alessandro Testa |
ISCC | 2 |
| 2010 | Reliable Event Dissemination over Wide-Area Networks without Severe Performance FluctuationsabstractPublish/subscribe middleware is being increasingly used to devise large-scale critical systems. Although several reliable publish/subscribe solutions have been proposed, none of them properly address the problem of assuring message dissemination even if network omissions happen without breaking any temporal constraints. In order to fill this gap, we have investigated how to guarantee a resilient and timely event dissemination despite of message losses. The contribution of this paper is on proposing a FEC approach, where encoding functionality is placed at the root and on a subset of interior nodes in the multicast tree, combined to a gossiping algorithm. Simulation-based experiments demonstrate that the proposed approach allows all the interested subscribers to receive all the published messages and the adopted resiliency mean does not affect the timeliness of the multicast protocol. Christian Esposito 0001, Domenico Cotroneo, Stefano Russo 0001 |
ISORC | 2 |
| 2010 | Software Aging Analysis of the Linux Operating SystemabstractSoftware systems running continuously for a long time tend to show degrading performance and an increasing failure occurrence rate, due to error conditions that accrue over time and eventually lead the system to failure. This phenomenon is usually referred to as Software Aging. Several long-running mission and safety critical applications have been reported to experience catastrophic aging-related failures. Software aging sources (i.e., aging-related bugs) may be hidden in several layers of a complex software system, ranging from the Operating System (OS) to the user application level. This paper presents a software aging analysis at the Operating System level, investigating software aging sources inside the Linux kernel. Linux is increasingly being employed in critical scenarios; this analysis intends to shed light on its behaviour from the aging perspective. The study is based on an experimental campaign designed to investigate the kernel internal behaviour over long running executions. By means of a kernel tracing tool specifically developed for this study, we collected relevant parameters of several kernel subsystems. Statistical analysis of collected data allowed us to confirm the presence of aging sources in Linux and to relate the observed aging dynamics to the monitored subsystems behaviour. The analysis output allowed us to infer potential sources of aging in the kernel subsystems. Domenico Cotroneo, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
ISSRE | 1 |
| 2010 | Memory leak analysis of mission-critical middleware
Gabriella Carrozza, Domenico Cotroneo, Roberto Natella, Antonio Pecchia, Stefano Russo 0001 |
J. Syst. Softw. | 2 |
| 2009 | AVR-INJECT: A tool for injecting faults in Wireless Sensor NodesabstractAs the incidence of faults in real Wireless Sensor Networks (WSNs) increases, fault injection is starting to be adopted to verify and validate their design choices. Following this recent trend, this paper presents a tool, named AVR-INJECT, designed to automate the fault injection, and analysis of results, on WSN nodes. The tool emulates the injection of hardware faults, such as bit flips, acting via software at the assembly level. This allows to attain simplicity, while preserving the low level of abstraction needed to inject such faults. The potential of the tool is shown by using it to perform a large number of fault injection experiments, which allow to study the reaction to faults of real WSN software. Marcello Cinque, Domenico Cotroneo, Catello Di Martino, Stefano Russo 0001, Alessandro Testa |
IPDPS | 2 |
| 2009 | Assessment and Improvement of Hang Detection in the Linux Operating SystemabstractWe propose a fault injection framework to assess hang detection facilities within the Linux operating system (OS). The novelty of the framework consists in the adoption of a more representative fault load than existing ones, and in the effectiveness in terms of number of hang failures produced; representativeness is supported by a field data study on the Linux OS. Using the proposed fault injection framework, along with realistic workloads, we find that the Linux OS is unable to detect hangs in several cases. We experience a relative coverage of 75%. To improve detection facilities, we propose a simple yet effective hang detector, which periodically tests OS liveness, as perceived by applications, by means of I/O system calls; it is shown that this approach can improve relative coverage up to 94%. The hang detector can be deployed on any Linux system, with an acceptable overhead. Domenico Cotroneo, Roberto Natella, Stefano Russo 0001 |
SRDS | 1 |
| 2009 | Calibrating RSS-Based Indoor Positioning SystemsabstractLocation estimation based on received signal strength (RSS) is the prevalent method in indoor positioning. For RSS-based methods a massive collection of training RSS samples is needed to calibrate the positioning system and to achieve a high positioning quality. The quality of these methods is directly related to the placement of the wireless sensors in the workspace and the radio map used to compute the user location. Traditionally deploying the reference points and building the radio map require human intervention and are extremely time-consuming. In this paper we aim to reduce these manual calibration efforts. We propose an automatic approach both to build a radio map in the given environment and to assess the best system calibration that fits the required positioning quality. The approach has been tested on the most used radio frequency-based technologies, i.e., IEEE 802.11 and Bluetooth. Christian Esposito 0001, Domenico Cotroneo, Massimo Ficco |
WiMob | 2 |
| 2009 | Self-adaptive handoff management for mobile streaming continuityabstractSelf-adaptive management and quality adaptation of multimedia services are open challenges in the heterogeneous wireless Internet, where different wireless access points potentially enable anywhere anytime Internet connectivity. One of the most challenging issues is to guarantee streaming continuity with maximum quality, despite possible handoffs at multimedia provisioning time. To enable handoff management to self-adapt to specific application requirements with minimum resource consumption, this paper offers three main contributions. First, it proposes a simple way to specify handoff-related service-level objectives that are focused on quality metrics and tolerable delay. Second, it presents how to automatically derive from these objectives a set of parameters to guide system-level configuration about handoff strategies and dynamic buffer tuning. Third, it describes the design and implementation of a novel handoff management infrastructure for maximizing streaming quality while minimizing resource consumption. Our infrastructure exploits i) experimentally evaluated tuning diagrams for resource management and ii) handoff prediction/awareness. The reported results show the effectiveness of our approach, which permits to achieve the desired quality-delay tradeoff in common Internet deployment environments, even in presence of vertical handoffs. Paolo Bellavista, Marcello Cinque, Domenico Cotroneo, Luca Foschini 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2008 | Dependability Evaluation and Modeling of the Bluetooth Data Communication ChannelabstractThis work presents a measurement-based dependability evaluation of the Bluetooth data communication channel, i.e., the Baseband layer. The main contribution is the definition of the Baseband's error/recovery model according to the Markov chains formalism. The model is derived by analyzing field data, which are collected via a commercial air sniffer deployed over real- world Bluetooth piconets. The model is parametric and actual values for its parameters are estimated by analyzing the field data. The paper also proposes the evaluation of dependability statistics (e.g., the error and failure times distributions, and the availability estimate), and the study of the failing behavior of the Bluetooth communication channel under Wi-Fi interferences. Gabriella Carrozza, Marcello Cinque, Domenico Cotroneo, Stefano Russo 0001 |
PDP | 3 |
| 2008 | Securing services in nomadic computing environments
Domenico Cotroneo, Cristiano di Flora, Almerindo Graziano, Stefano Russo 0001 |
Inf. Softw. Technol. | 1 |
| 2007 | How Do Mobile Phones Fail? A Failure Data Analysis of Symbian OS Smart PhonesabstractWhile the new generation of hand-held devices, e.g., smart phones, support a rich set of applications, growing complexity of the hardware and runtime environment makes the devices susceptible to accidental errors and malicious attacks. Despite these concerns, very few studies have looked into the dependability of mobile phones. This paper presents measurement-based failure characterization of mobile phones. The analysis starts with a high level failure characterization of mobile phones based on data from publicly available web forums, where users post information on their experiences in using hand-held devices. This initial analysis is then used to guide the development of a failure data logger for collecting failure-related information on SymbianOS-based smart phones. Failure data is collected from 25 phones (in Italy and USA) over the period of 14 months. Key findings indicate that: (i) the majority of kernel exceptions are due to memory access violation errors (56%) and heap management problems (18%), and (ii) on average users experience a failure (freeze or self shutdown) every 11 days. While the study provide valuable insight into the failure sensitivity of smart-phones, more data and further analysis are needed before generalizing the results. Marcello Cinque, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 2 |
| 2007 | Modeling and Assessing the Dependability ofWireless Sensor NetworksabstractThis paper proposes a flexible framework for dependability modeling and assessing of Wireless Sensor Networks (WSNs). The framework takes into account network related aspects (topology, routing, network traffic) as well as hardware/software characteristics of nodes (type of sensors, running applications, power consumption). It is composed of two basic elements: i) a parametric Stochastic Activity Networks (SAN) failure model, reproducing WSN failure behavior as inferred from a detailed Failure Mode Effect Analysis (FMEA), and ii) an external library reproducing network behavior on behalf of the SAN model. This library specializes the SAN model by feeding it with quantitative parameters obtained by simulation or by experimental campaigns; it is also in charge of updating the network state in response to failure events during the simulation (e.g., routing tree updated due to node failures). The framework is thus suited to evaluate the dependability of several WSNs, with different topologies, routing algorithms, hardware/software platforms, without requiring any changes to its structure. The use of the external library makes the model simpler, decoupling the network behavior from the failure behavior. Simulation experiments are discussed that provide a quantitative evaluation of WSN dependability for a sample scenario: results show how the proposed framework supports WSN developers to find proper cost-reliability trade-offs for the system being deployed. Marcello Cinque, Domenico Cotroneo, Catello Di Martino, Stefano Russo 0001 |
SRDS | 2 |
| 2007 | Characterizing Aging Phenomena of the Java Virtual MachineabstractIn this work we investigate software aging phenomena inside the Java Virtual Machine (JVM). Starting from an experimental campaign on real world testbeds, this work isolates the contribution of the JVM to the overall aging trend, and identifies, through statistical methods, which workload parameters are more relevant to aging dynamics. Experimental results show that the Sun Hotpost JVM experiences software aging phenomena. A consistent memory depletion trend (up to 50 KB/min) has been observed during periods of low garbage collector activity; the Just-In-Time compiler is also responsible for a lighter, but not negligible, memory depletion trend; finally, a consistent throughput loss (up to 24 KB/min) has been observed. Domenico Cotroneo, Salvatore Orlando 0002, Stefano Russo 0001 |
SRDS | 1 |
| 2007 | The Esperanto Broker: a communication platform for nomadic computing systemsabstractAbstract There is an increasing demand for middleware for nomadic computing applications. Owing to the inherent characteristics of such environments, these platforms have to address two fundamental issues: (i) device disconnections and the limitations of wireless networks may force users to experience short periods of service unavailability; and (ii) the complexity to design and develop next‐generation mobile computing applications. This paper proposes the Esperanto Broker (EB), a communication platform that addresses mobility issues via an integrated approach, i.e. at data‐link, network, and middleware levels. Decoupling interactions are achieved via a tuple‐space underlying infrastructure. To support developers with advanced services, the EB enhances the distributed objects computing model providing the abstraction for the communication paradigms standardized by the W3C. Esperanto applications can be modeled as sets of objects that are distributed over mobile devices, which communicate via remote method invocations (RMIs). RMIs natively implement pull and push models, in both one‐to‐one and one‐to‐many multiplicity. The paper focuses on the EB design issues, essential aspects of the implementation, and performance evaluations of the implemented prototype. Copyright © 2006 John Wiley & Sons, Ltd. Domenico Cotroneo, Armando Migliaccio, Stefano Russo 0001 |
Softw. Pract. Exp. | 1 |
| 2006 | Collecting and Analyzing Failure Data of Bluetooth Personal Area NetworksabstractThis work presents a failure data analysis campaign on Bluetooth personal area networks (PANs) conducted on two kind of heterogeneous testbeds (working for more than one year). The obtained results reveal how failures distribution is characterized and suggest how to improve the dependability of Bluetooth PANs. Specifically, we define the failure model and we then identify the most effective recovery actions and masking strategies that can be adopted for each failure. We then integrate the discovered recovery actions and masking strategies in our testbeds, improving the availability and the reliability of 3.64% (up to 36.6%) and 202% (referred to the mean time to failure), respectively Marcello Cinque, Domenico Cotroneo, Stefano Russo 0001 |
DSN | 2 |
| 2006 | Failure classification and analysis of the Java Virtual MachineabstractThis paper presents a failure analysis of the Java Virtual Machine providing useful insights into the nature of reported failures and to improve the understanding of its dependability aspects. Failure data is extracted from publicly available bug databases, where developers and users of Java applications usually submit failures/bugs. Presented results clearly indicate that much more efforts have still to be done in order to improve the dependability of the JVM. In particular, the conducted analysis revealed that i) builtin error detection mechanism are characterized by a low coverage; ii) the JVM does not achieve the same levels of dependability across different platforms iii) developers have to pursue a tradeoff between performance and reliability. Finally, code fragments reproducing failures submitted in bug database are injected into Java Applications. Preliminary results show that often these faults could be removed changing the environment of the JVM. Domenico Cotroneo, Salvatore Orlando 0002, Stefano Russo 0001 |
ICDCS | 1 |
| 2006 | Automated Logging of Mobile Phones Failures DataabstractThe increasing complexity of mobile phones directly affects their reliability, while the user tolerance for failures becomes to decrease, especially when the phone is used for business- or mission-critical applications. Despite these concerns, there is still little understanding on how and why these devices fail and no techniques have been defined to gather useful information about failures manifestation from the phone. This paper presents the design of a logger application to collect failure-related information from mobile phones. Preliminary failure data collected from real-world mobile phones confirm the proposed logger is a useful instrument to gain knowledge about mobile phone failure's dynamics and causes. Paolo Ascione, Marcello Cinque, Domenico Cotroneo |
ISORC | 3 |
| 2005 | ESPERANTO: a middleware platform to achieve interoperability in nomadic computing domainsabstractSummary form only given. The most challenging issues in nomadic computing environments arise from the combination of heterogeneity, dynamism, context-awareness, and mobility. Driven by these issues, this paper presents a new middleware infrastructure, named ESPERANTO, to support the integration of diverse nomadic computing domains. This middleware aims to glue the emerging heterogeneous nomadic computing technologies and service oriented architectures. Marcello Cinque, Domenico Cotroneo, Cristiano di Flora, Armando Migliaccio, Stefano Russo 0001 |
AICCSA | 2 |
| 2005 | A Communication Broker for Nomadic Computing Systems
Domenico Cotroneo, Armando Migliaccio, Stefano Russo 0001 |
HPCC | 1 |
| 2005 | CSAR-2: A Case Study of Parallel File System Dependability Analysis
Domenico Cotroneo, Generoso Paolillo, Stefano Russo 0001, Mario Lauria |
HPCC | 1 |
| 2005 | An Automated Distributed Infrastructure for Collecting Bluetooth Field Failure DataabstractThe widespread use of mobile and wireless computing platforms is leading to a growing interest on dependability issues. Several research studies have been conducted on dependability of mobile environments, but none of them attempted to identify system bottlenecks and to quantify dependability measures. This paper proposes a distributed automated infrastructure for monitoring and collecting spontaneous failures of the Bluetooth infrastructure, which is nowadays more and more recognized as an enabler for mobile systems. Information sources for failure data are presented, and preliminary experimental results are discussed. Marcello Cinque, Fabio Cornevilli, Domenico Cotroneo, Stefano Russo 0001 |
ISORC | 3 |
| 2004 | Effective Fault Treatment for Improving the Dependability of COTS and Legacy-Based ApplicationsabstractThis paper proposes a novel methodology and an architectural framework for handling multiple classes of faults (namely, hardware-induced software errors in the application, process and/or host crashes or hangs, and errors in the persistent system stable storage) in a COTS and legacy-based application. The basic idea is to use an evidence-accruing fault tolerance manager to choose and carry out one of multiple fault recovery strategies, depending upon the perceived severity of the fault. The methodology and the framework have been applied to a case study system consisting of a legacy system, which makes use of a COTS DBMS for persistent storage facilities. A thorough performability analysis has also been conducted via combined use of direct measurements and analytical modeling. Experimental results demonstrate that effective fault treatment, consisting of careful diagnosis and damage assessment, plays a key role in leveraging the dependability of COTS and legacy-based applications. Andrea Bondavalli, Silvano Chiaradonna, Domenico Cotroneo, Luigi Romano |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2003 | Modeling and Detecting Failures in Next-generation Distributed Multimedia ApplicationsabstractIn this paper we investigate dependability issues of next-generation distributed multimedia applications. Examples of such applications are autonomous vehicle control, tele-medicine, and audio/video control. For these applications the quality of the delivered multimedia data is a critical factor. According to the ITU-T (working group SG 12), the quality of a multimedia service as perceived by end-users is defined by three parameters: delay, delay variation, and information loss. It is paramount to formalize the concept of a failure from the user's perspective. This paper defines the correctness of a multimedia service as a function of temporal distributions of the user-related parameters. It proposes a strategy for modeling and detecting failures of the considered applications. In particular, the detection process is based on error filtering functions. We show that the combination of threshold-based mechanisms is suitable for implementing an efficient detection strategy. We also evaluate the effectiveness of the proposed mechanism both by simulations and by experiments performed on a prototype. Such a prototype is tested with respect to a case study application, consisting of distributed remote-control based on RTP/RTCP standard streaming protocols. Domenico Cotroneo, Cristiano di Flora, Generoso Paolillo, Stefano Russo 0001 |
SRDS | 1 |
| 2003 | An architecture for security-oriented perfective maintenance of legacy software
Domenico Cotroneo, Antonino Mazzeo, Luigi Romano, Stefano Russo 0001 |
Inf. Softw. Technol. | 1 |
| 2003 | An Enhanced Service Oriented Architecture for Developing Web-based Applications
Domenico Cotroneo, Cristiano di Flora, Stefano Russo 0001 |
J. Web Eng. | 1 |
| 2002 | An active security protocol against DoS attacksabstractDenial of service (DoS) attacks represent, in today's Internet, one of the most complex issues to address. A session is under a DoS attack if it cannot achieve its intended throughput due to the misbehavior of other sessions. Many research studies dealt with DoS, proposing models and/or architectures mostly based on an attack prevention approach. Prevention techniques lead to different models, each suitable for a single type of misbehavior, but do not guarantee the protection of a system from a more general DoS attack. We analyze the fundamental requirements to be satisfied in order to protect hosts and routers from any form of distributed DoS (DDoS). Then we propose a network signaling protocol, named active security protocol(ASP), which satisfies most of the defined requirements. ASP provides an active protection from a DDoS attack, being able to adapt its defense strategies to the current type of violation. Protocol specification and design are performed using an object oriented methodology: we used Unified Modeling Language (UML) as a software description language. Domenico Cotroneo, L. Peluso, Simon Pietro Romano, Giorgio Ventre |
ISCC | 1 |
| 2002 | Implementation of Threshold-based Diagnostic Mechanisms for COTS-Based ApplicationsabstractThis work investigates feasibility issues that must be addressed when threshold-based mechanisms are to be used for diagnostic purposes in COTS-based distributed systems. Threshold based mechanisms have typically been used for such purposes in embedded systems. A variety of solutions exist, with different characteristics of completeness, accuracy, and induced overhead. We first discuss the challenges related to applying such mechanisms to COTS-based distributed applications. We then identify alternative strategies for diagnosis, which use run-time data on COTS component service failures to trigger alarms to reconfiguration and fault treatment mechanisms. We implement those strategies in a system prototype, which is based on a substantial application, i.e. a real world (as opposed to a toy) application. We discuss the relationships between the sensitivity of the quality of service (QoS) provided by the diagnostic mechanisms and the accuracy of the available failure data. Our considerations and preliminary experiments on the prototype suggest that a careful evaluation of tradeoffs must be conducted, in order to achieve the best compromise between accuracy and cost, which depends on application characteristics, and service deployment requirements. Luigi Romano, Andrea Bondavalli, Silvano Chiaradonna, Domenico Cotroneo |
SRDS | 4 |
| 2002 | Building a dependable system from a legacy application with CORBA
Domenico Cotroneo, Nicola Mazzocca, Luigi Romano, Stefano Russo 0001 |
J. Syst. Archit. | 1 |