Antonio Pecchia

dblp:62/8357 · DBLP profile ↗
← Back
48ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0003-2869-8423ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 17 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 13 · 3 first-author · 4 since 2021Systems, architecture and hardware · 9 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Similarity Is Not Enough: Issues with Adversarial Perturbations of Traffic Features against Intrusion Detection Systems
Marta Catillo, Antonio Pecchia, Umberto Villano
ICISSP (1)2
2025 USB-IDS-TC: A Flow-Based Intrusion Detection Dataset of DoS Attacks in Different Network Scenarios
Marta Catillo, Antonio Pecchia, Umberto Villano
ICISSP (1)2
2025 Topic Modeling for Graph-Based Analysis of Fake News Diffusion
abstract
Fake news diffusion is a primary driver of misinformation. Analyzing deliberately false and misleading content is tough because social media platforms make it incredibly easy to create and spread huge amounts of information quickly. The intricate dynamics of fake news propagation demand the availability of ready-to-use frameworks for its analysis. This paper explores the automatic topic identification component of SPREADSHOT, a graph-based method designed to analyze fake news dissemination by examining two key factors: spreaders and topics. When it comes to news content, fake news frequently revolves around rapidly evolving topics due to its strong connection to current events. Consequently, topic modeling has gained significant traction for analyzing news articles. In our analysis, we explore two distinct topic modeling techniques: Latent Dirichlet Allocation (LDA) and BERTopic. While both offer valuable insights, we carefully justify which of these two techniques is best suited for integration into the SPREADSHOT framework for topic modeling.
Pasquale Avella, Carmela Bernardo, Marta Catillo, Antonio Pecchia, Francesco Vasca, Umberto Villano
WETICE4
2024 Towards realistic problem-space adversarial attacks against machine learning in network intrusion detection
abstract
Current trends in network intrusion detection systems (NIDS) capitalize on the extraction of features from network traffic and the use of up-to-date machine and deep learning techniques to infer a detection model; in consequence, NIDS can be vulnerable to adversarial attacks. Differently from the plethora of contributions that apply (and misuse) feature-level attacks envisioned in application domains far from NIDS, this paper proposes a novel approach to adversarial attacks, which consists in a realistic problem-space perturbation of the network traffic. The perturbation is achieved through a traffic control utility. Experiments are based on normal and Denial of Service traffic in both legitimate and adversarial conditions, and the application of four popular techniques to learn the NIDS models. The results highlight the transferability of the adversarial examples generated by the proposed problem-space attack as well as the effectiveness at inducing traffic misclassifications across the NIDS models obtained.
Marta Catillo, Antonio Pecchia, Antonio Repola, Umberto Villano
ARES2
2024 DEFEDGE: Threat-Driven Security Testing and Proactive Defense Identification for Edge-Cloud Systems
Valentina Casola, Marta Catillo, Alessandra De Benedictis, Felice Moretta, Antonio Pecchia, Massimiliano Rak, Umberto Villano
AINA (5)5
2024 Exploring the effect of training-time randomness on the performance of deep neural networks for intrusion detection
Marta Catillo, Antonio Pecchia, Umberto Villano
Soft Comput.2
2024 Successful intrusion detection with a single deep autoencoder: theory and practice
Marta Catillo, Antonio Pecchia, Umberto Villano
Softw. Qual. J.2
2023 A Case Study with CICIDS2017 on the Robustness of Machine Learning against Adversarial Attacks in Intrusion Detection
abstract
Intrusion detection systems (IDS) play a key role to assure security properties of modern computer networks. IDS are often based on machine and deep learning techniques; as such, IDS are vulnerable to various forms of adversarial attacks. This paper presents an initial case study on the robustness of machine learning for network intrusion detection against adversarial attacks. Experiments are based on a recent fix of the widely-used CICIDS2017 benchmark dataset, two well-known machine learning techniques for intrusion detection (i.e., deep autoencoders and decision trees), and the virtual adversarial method (VAM) to generate the adversarial examples. Based on the data and experiments at hand, the results provide many interesting findings on the robustness of the IDS models assessed. The autoencoder-based IDS is more robust to evasion rather than overstimulation. On the contrary, the decision tree is vulnerable to evasion; moreover, changes to the learning parameters can strongly affect the robustness of the decision tree against the VAM attack.
Marta Catillo, Andrea Del Vecchio, Antonio Pecchia, Umberto Villano
ARES3
2023 Traditional vs Federated Learning with Deep Autoencoders: a Study in IoT Intrusion Detection
abstract
Security of Internet of Things (IoT) devices and networks is a primary concern. Many intrusion detection systems (IDS) proposals in the IoT leverage machine and deep learning algorithms to learn models that can be used to discriminate normal behaviors from intrusions. Due to the dynamicity and scale of modern IoT networks, it is hard to learn and maintain one separate IDS model per device; on the other hand, the Cloud-Edge-IoT architecture allows learning a single IDS model (instead of many separate models). This paper compares two paradigms, i.e., traditional and federated, to learn a single IDS model atop the traffic of different IoT devices. The former assumes the availability of an all-in-one training dataset at a unique learning node; the latter aggregates the outcomes of independent learning procedures executed on individual training datasets hosted by different nodes. The experiments are done with a well-established public benchmark of nine IoT devices and the use of deep autoencoders. In the experiment and dataset at hand, federated learning lead to an increase of the false positive rate of six devices compared to the traditional scenario. Such an increase was balanced by a narrower variability of the false positive rate across all the devices and a mitigation of potential overfitting.
Marta Catillo, Antonio Pecchia, Umberto Villano
CloudCom2
2023 CPS-GUARD: Intrusion detection for cyber-physical systems and IoT devices using outlier-aware deep autoencoders
abstract
Detecting attacks to Cyber-Physical Systems (CPSs) is of utmost importance, due to their increasingly frequent use in many critical assets. Intrusion detection in CPSs and other domains, such as the Internet of Things, is often addressed through machine and deep learning. However, many existing proposals tend to favor the application of complex detection models over the usability in real-world operations. This paper presents CPS-GUARD, a novel intrusion detection approach based on a single semi-supervised autoencoder and a technique to set the threshold used to discriminate normal operations from attacks. The technique is outlier-aware, in that it relies on outlier detection to mitigate inherent imperfections of the training data. CPS-GUARD is evaluated by means of direct experiments with normal and intrusion data points pertaining to individual sensing devices, an HTTP server and four full-fledged systems, including CPSs. Experiments are based on a wide spectrum of attacks available in six state-of-the-art datasets. The intrusion detection results of CPS-GUARD are within 0.949-1.000 recall, 0.961-0.999 precision and 0.006-0.027 false positive rate depending on the specific system. The results are competitive with other existing intrusion detection methods. The evaluation is complemented by a comparative study on alternative threshold selection and outlier detection techniques.
Marta Catillo, Antonio Pecchia, Umberto Villano
Comput. Secur.2
2022 Botnet Detection in the Internet of Things through All-in-one Deep Autoencoding
abstract
In the past years Internet of Things (IoT) has received increasing attention by academia and industry due to the potential use in several human activities; however, IoT devices are vulnerable to various types of attacks. Many existing intrusion detection proposals in the IoT leverage complex machine learning architectures, which may provide one separate model per device or per attack. These solutions are not suited to the dynamicity and scale of modern IoT environments. This paper proposes an initial analysis of the problem in the context of deep autoencoders and the detection of botnet attacks. Our findings, obtained by means of the N-BaIoT dataset, indicate that it is relatively easy to achieve impressive detection results by training-testing separate and minimal deep autoenconders on the top of the data individual IoT devices. More important, our all-in-one deep autoencoding proposal, which consists in training a single model with the benign traffic collected from different IoT devices, allows to preserve the overall detection performance obtained through separate autoencoders. The all-in-one model can pave the way for more scalable intrusion detection solutions in the context of IoT.
Marta Catillo, Antonio Pecchia, Umberto Villano
ARES2
2022 AutoLog: Anomaly detection by deep autoencoding of system logs
Marta Catillo, Antonio Pecchia, Umberto Villano
Expert Syst. Appl.2
2022 No more DoS? An empirical study on defense techniques for web server Denial of Service mitigation
Marta Catillo, Antonio Pecchia, Umberto Villano
J. Netw. Comput. Appl.2
2022 Micro2vec: Anomaly detection in microservices systems by mining numeric representations of computer logs
abstract
This paper describes a study on log mining in the domain of microservices technologies. We focus on the detection of anomalies from logs, i.e., events requiring deeper inspection by analysts. Log mining is challenging in microservices systems due to the high number of heterogeneous logs. We present Micro2vec, a novel approach to mine numeric representations of computer logs without making assumptions on the format of underlying data and requiring no application knowledge; representations computed by Micro2vec are suited for anomaly detection. To cope with the lack of publicly-available datasets of labeled logs from production systems, we validate our approach by means of a mixture of direct measurements from logs, one-class classification experiments and generation of log variants. The study has been conducted in the context of a Clearwater IP Multimedia Subsystem setup consisting of microservices deployed in Docker containers, and on a real-world critical information system from the Air Traffic Control domain, which implements a communication model typically used with microservices.
Marcello Cinque, Raffaele Della Corte, Antonio Pecchia
J. Netw. Comput. Appl.3
2022 Transferability of machine learning models learned from public intrusion detection datasets: the CICIDS2017 case study
Marta Catillo, Andrea Del Vecchio, Antonio Pecchia, Umberto Villano
Softw. Qual. J.3
2022 Microservices Monitoring with Event Logs and Black Box Execution Tracing
abstract
Monitoring is a core practice in any software system. Trends in microservices systems exacerbate the role of monitoring and pose novel challenges to data sources being used for monitoring, such as event logs. Current deployments create a distinct log per microservice; moreover, composing microservices by different vendors exacerbates format and semantic heterogeneity of logs. Understanding and traversing the logs from different microservices demands for substantial cognitive work by human experts. This paper proposes a novel approach to accompany microservices logs with black box tracing to help practitioners in making informed decisions for troubleshooting. Our approach is based on the passive tracing of request-response messages of the REpresentational State Transfer (REST) communication model. Differently from many existing tools for microservices, our tracing is application transparent and non-intrusive. We present an implementation called MetroFunnel and conduct an assessment in the context of two case studies: a Clearwater IP Multimedia Subsystem (IMS) setup consisting of Docker microservices and a Kubernetes orchestrator deployment hosting tens of microservices. MetroFunnel allows making useful attributions in traversing the logs; more important, it reduces the size of collected monitoring data at negligible performance overhead with respect to traditional logs.
Marcello Cinque, Raffaele Della Corte, Antonio Pecchia
IEEE Trans. Serv. Comput.3
2021 On the Quality of Network Flow Records for IDS Evaluation: A Collaborative Filtering Approach
Marta Catillo, Andrea Del Vecchio, Antonio Pecchia, Umberto Villano
ICTSS3
2021 Microservices Monitoring with Event Logs and Black Box Execution Tracing
abstract
Monitoring is a core practice in any software system, and entails gathering a variety of data sources that pertain the execution of a given system. Trends in microservices systems exacerbate the role of monitoring. Microservices put forth reduced size, independency, flexibility and modularity principles, which well cope with ever-changing business environments. However, as real-world applications are decomposed, they can easily reach hundreds of microservices. This inherent complexity determines an increasing difficulty in debugging, monitoring and forensics, and poses novel challenges to monitoring data sources, such as event logs.
Marcello Cinque, Raffaele Della Corte, Antonio Pecchia
SERVICES3
2021 Demystifying the role of public intrusion datasets: A replication study of DoS network traffic data
Marta Catillo, Antonio Pecchia, Massimiliano Rak, Umberto Villano
Comput. Secur.2
2020 A case study on the representativeness of public DoS network traffic data for cybersecurity research
abstract
The availability of ready-to-use public security datasets is fostering measurement-driven research by a wide community of academics and practitioners. Recent trends in this area put forth a substantial body of literature on anomaly and attack detection on the top of public labelled datasets. Much of this literature blindly reuses existing datasets by overlooking the cybersecurity facets of the network traffic therein, in terms of its real impact on service availability and performance of operations.
Marta Catillo, Antonio Pecchia, Massimiliano Rak, Umberto Villano
ARES2
2020 Measurement-Based Analysis of a DoS Defense Module for an Open Source Web Server
Marta Catillo, Antonio Pecchia, Umberto Villano
ICTSS2
2020 An empirical analysis of error propagation in critical software systems
Marcello Cinque, Raffaele Della Corte, Antonio Pecchia
Empir. Softw. Eng.3
2020 Contextual filtering and prioritization of computer application logs for security situational awareness
Marcello Cinque, Raffaele Della Corte, Antonio Pecchia
Future Gener. Comput. Syst.3
2020 Discovering process models for the analysis of application failures under uncertainty of event logs
Antonio Pecchia, Ingo Weber, Marcello Cinque, Yu Ma 0001
Knowl. Based Syst.1
2020 Security Log Analysis in Critical Industrial Systems Exploiting Game Theoretic Feature Selection and Evidence Combination
abstract
Critical industrial systems have become profitable targets for cyber-attackers. Practitioners and administrators rely on a variety of data sources to develop security situation awareness at runtime. In spite of the advances in security information and event management products and services for handling heterogeneous data sources, analysis of proprietary logs generated by industrial systems keeps posing many challenges due to the lack of standard practices, formats, and threat models. This article addresses log analysis to detect anomalies, such as failures and misuse, in a critical industrial system. We conduct our study with a real-life system by a top leading industry provider in the air traffic control domain. The system emits massive volumes of highly-unstructured proprietary textual logs at runtime. We propose to extract quantitative metrics from logs and to detect anomalies by means of game theoretic feature selection and evidence combination. Experiments indicate that the proposed approach achieves high precision and recall at small tuning efforts.
Marcello Cinque, Christian Esposito 0001, Antonio Pecchia
IEEE Trans. Ind. Informatics3
2020 Assessing Invariant Mining Techniques for Cloud-Based Utility Computing Systems
abstract
Likely system invariants model properties that hold in operating conditions of a computing system. Invariants may be mined offline from training datasets, or inferred during execution. Scientific work has shown that invariants' mining techniques support several activities, including capacity planning and detection of failures, anomalies and violations of Service Level Agreements. However their practical application by operation engineers is still a challenge. We aim to fill this gap through an empirical analysis of three major techniques for mining invariants in cloud-based utility computing systems: clustering, association rules, and decision list. The experiments use independent datasets from real-world systems: a Google cluster, whose traces are publicly available, and a Software-as-a-Service platform used by various companies worldwide. We assess the techniques in two invariants' applications, namely executions characterization and anomaly detection, using the metrics of coverage, recall and precision. A sensitivity analysis is performed. Experimental results allow inferring practical usage implications, showing that relatively few invariants characterize the majority of operating conditions, that precision and recall may drop significantly when trying to achieve a large coverage, and that techniques exhibit similar precision, though the supervised one a higher recall. Finally, we propose a general heuristic for selecting likely invariants from a dataset.
Antonio Pecchia, Stefano Russo 0001, Santonu Sarkar
IEEE Trans. Serv. Comput.1
2019 RT-CASEs: Container-Based Virtualization for Temporally Separated Mixed-Criticality Task Sets
abstract
Real-time containers are a promising solution to reduce latencies in time-sensitive cloud systems. Recent efforts are emerging to extend their usage in industrial edge systems with mixed-criticality constraints. In these contexts, isolation becomes a major concern: a disturbance (such as timing faults or unexpected overloads) affecting a container must not impact the behavior of other containers deployed on the same hardware. In this paper, we propose a novel architectural solution to achieve isolation in real-time containers, based on real-time co-kernels, hierarchical scheduling, and time-division networking. The architecture has been implemented on Linux patched with the Xenomai co-kernel, extended with a new hierarchical scheduling policy, named SCHED_DS, and integrating the RTNet stack. Experimental results are promising in terms of overhead and latency compared to other Linux-based solutions. More importantly, the isolation of containers is guaranteed even in presence of severe co-located disturbances, such as faulty tasks (elapsing more time than declared) or high CPU, network, or I/O stress on the same machine.
Marcello Cinque, Raffaele Della Corte, Antonio Eliso, Antonio Pecchia
ECRTS4
2019 A framework for on-line timing error detection in software systems
Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia
Future Gener. Comput. Syst.4
2019 Empirical Analysis and Validation of Security Alerts Filtering Techniques
abstract
System administrators cope with security incidents through a variety of monitors, such as intrusion detection systems, event logs, security information and event management systems. Monitors generate large volumes of alerts that overwhelm the operations team and make forensics time-consuming. Filtering is a consolidated technique to reduce the amount of alerts. In spite of the number of filtering proposals, few studies have addressed the validation of filtering results in real production datasets. This paper analyzes a number of state-of-the-art filtering techniques that are used to address security datasets. We use 14 months of alerts generated in a SaaS Cloud. Our analysis aims to measure and compare the reduction of the alerts volume obtained by the filters. The analysis highlights pros and cons of each filter and provides insights into the practical implications of filtering as affected by the characteristics of a dataset. We complement the analysis with a method to validate the output of a filter in absence of ground truth, i.e., the knowledge of the incidents occurred in the system at the time the alerts were generated. The analysis addresses blacklist, conceptual clustering and bytes techniques, and our filtering proposal based on term weighting.
Domenico Cotroneo, Andrea Paudice, Antonio Pecchia
IEEE Trans. Dependable Secur. Comput.3
2018 Guest Editors' Introduction: Special Issue on Data-Driven Dependability and Security
abstract
The six papers in this special issue aim to concentrate novel contributions addressing dependability and security of computer systems through data analysis, and to publish consolidated research results focusing on data-driven methodologies, measurements from production systems, and analysis of large datasets. This information provides valuable contributions related to log-based measurements, operating systems dependability, and attack detection.
Domenico Cotroneo, Karthik Pattabiraman, Antonio Pecchia
IEEE Trans. Dependable Secur. Comput.3
2018 Multiobjective Testing Resource Allocation Under Uncertainty
abstract
Testing resource allocation is the problem of planning the assignment of resources to testing activities of software components so as to achieve a target goal under given constraints. Existing methods build on software reliability growth models (SRGMs), aiming at maximizing reliability given time/cost constraints, or at minimizing cost given quality/time constraints. We formulate it as a multiobjective debug-aware and robust optimization problem under uncertainty of data, advancing the state-of-the-art in the following ways. Multiobjective optimization produces a set of solutions, allowing to evaluate alternative tradeoffs among reliability, cost, and release time. Debug awareness relaxes the traditional assumptions of SRGMs-in particular the very unrealistic immediate repair of detected faults-and incorporates the bug assignment activity. Robustness provides solutions valid in spite of a degree of uncertainty on input parameters. We show results with a real-world case study.
Roberto Pietrantuono, Pasqualina Potena, Antonio Pecchia, Daniel Rodríguez-García, Stefano Russo 0001, Luis Fernández-Sanz
IEEE Trans. Evol. Comput.3
2017 Entropy-Based Security Analytics: Measurements from a Critical Information System
abstract
Critical information systems strongly rely on event logging techniques to collect data, such as housekeeping/error events, execution traces and dumps of variables, into unstructured text logs. Event logs are the primary source to gain actionable intelligence from production systems. In spite of the recognized importance, system/application logs remain quite underutilized in security analytics when compared to conventional and structured data sources, such as audit traces, network flows and intrusion detection logs. This paper proposes a method to measure the occurrence of interesting activity (i.e., entries that should be followed up by analysts) within textual and heterogeneous runtime log streams. We use an entropy-based approach, which makes no assumptions on the structure of underlying log entries. Measurements have been done in a real-world Air Traffic Control information system through a data analytics framework. Experiments suggest that our entropy-based method represents a valuable complement to security analytics solutions.
Marcello Cinque, Raffaele Della Corte, Antonio Pecchia
DSN3
2017 On the injection of hardware faults in virtualized multicore systems
Marcello Cinque, Antonio Pecchia
J. Parallel Distributed Comput.2
2017 Debugging-workflow-aware software reliability growth analysis
abstract
Summary Software reliability growth models support the prediction/assessment of product quality, release time, and testing/debugging cost. Several software reliability growth model extensions take into account the bug correction process. However, their estimates may be significantly inaccurate when debugging fails to fully fit modelling assumptions. This paper proposes debugging‐workflow‐aware software reliability growth method (DWA‐SRGM), a method for reliability growth analysis leveraging the debugging data usually managed by companies in bug tracking systems. On the basis of a characterization of the debugging workflow within the software project under consideration (in terms of bug features and treatment phases), DWA‐SRGM pinpoints the factors impacting the estimates and to spot bottlenecks, thus supporting process improvement decisions. Two industrial case studies are presented, a customer relationship management system and an enterprise resource planning system, whose defects span a period of about 17 and 13 months, respectively. DWA‐SRGM revealed effective to obtain more realistic estimates and to capitalize on the awareness of critical factors for improving debugging.
Marcello Cinque, Domenico Cotroneo, Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001
Softw. Test. Verification Reliab.3
2016 Automated root cause identification of security alerts: Evaluation in a SaaS Cloud
Domenico Cotroneo, Andrea Paudice, Antonio Pecchia
Future Gener. Comput. Syst.3
2016 Characterizing Direct Monitoring Techniques in Software Systems
abstract
Monitoring is a consolidated practice to characterize the dependability behavior of a software system. A variety of techniques, such as event logging and operating system probes, are currently used to generate monitoring data for troubleshooting and failure analysis. In spite of the importance of monitoring, whose role can be essential in critical software systems, there is a lack of studies addressing the assessment and the comparison of the techniques aiming to monitor the occurrence of failures during operations. This paper proposes a method to characterize the monitoring techniques implemented in a software system. The method is based on a fault injection approach and allows measuring 1) precision and recall of a monitoring technique and 2) the dissimilarity of the data it generates upon failures. The method has been used in two critical software systems implementing event logging, assertion checking, and source code instrumentation techniques. We analyzed a total of 3 844 failures. With respect to our data, we observed that the effectiveness of a technique is strongly affected by the system and type of failure, and that the combination of different techniques is potentially beneficial to increase the overall failure reporting ability. More important, our analysis revealed a number of practical implications to be taken into account when developing a monitoring technique.
Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia
IEEE Trans. Reliab.4
2015 Industry Practices and Event Logging: Assessment of a Critical Software Development Process
abstract
Practitioners widely recognize the importance of event logging for a variety of tasks, such as accounting, system measurements and troubleshooting. Nevertheless, in spite of the importance of the tasks based on the logs collected under real workload conditions, event logging lacks systematic design and implementation practices. The implementation of the logging mechanism strongly relies on the human expertise. This paper proposes a measurement study of event logging practices in a critical industrial domain. We assess a software development process at Selex ES, a leading Finmeccanica company in electronic and information solutions for critical systems. Our study combines source code analysis, inspection of around 2.3 millions log entries, and direct feedback from the development team to gain process-wide insights ranging from programming practices, logging objectives and issues impacting log analysis. The findings of our study were extremely valuable to prioritize event logging reengineering tasks at Selex ES.
Antonio Pecchia, Marcello Cinque, Gabriella Carrozza, Domenico Cotroneo
ICSE (2)1
2014 What Logs Should You Look at When an Application Fails? Insights from an Industrial Case Study
abstract
Event logs are the first place where to find useful information about application failures. Event logs are available at different system levels, such as application, middleware and operating system. In this paper we analyze the failure reporting capability of event logs collected at different levels of an industrial system in the Air Traffic Control (ATC) domain. The study is based on a data set of 3,159 failures induced in the system by means of software fault injection. Results indicate that the reporting ability of event logs collected at a given level is strongly affected by the type of failure observed at runtime. For example, even if operating system logs catch almost all application crashes, they are strongly ineffective in face of silent and erratic failures in the considered system.
Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia
DSN4
2014 On the Impact of Debugging on Software Reliability Growth Analysis: A Case Study
Marcello Cinque, Claudio Gaiani, Daniele De Stradis, Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001
ICCSA (5)4
2014 Assessing Direct Monitoring Techniques to Analyze Failures of Critical Industrial Systems
abstract
The analysis of monitoring data is extremely valuable for critical computer systems. It allows to gain insights into the failure behavior of a given system under real workload conditions, which is crucial to assure service continuity and downtime reduction. This paper proposes an experimental evaluation of different direct monitoring techniques, namely event logs, assertions, and source code instrumentation, that are widely used in the context of critical industrial systems. We inject 12,733 software faults in a real-world air traffic control (ATC) middleware system with the aim of analyzing the ability of mentioned techniques to produce information in case of failures. Experimental results indicate that each technique is able to cover a limited number of failure manifestations. Moreover, we observe that the quality of collected data to support failure diagnosis tasks strongly varies across the techniques considered in this study.
Marcello Cinque, Domenico Cotroneo, Raffaele Della Corte, Antonio Pecchia
ISSRE4
2013 Towards secure monitoring and control systems: Diversify!
abstract
Cyber attacks have become surprisingly sophisticated over the past fifteen years. While early infections mostly targeted individual machines, recent threats leverage the widespread network connectivity to develop complex and highly coordinated attacks involving several distributed nodes [1]. Attackers are currently targeting very diverse domains, e.g., e-commerce systems, corporate networks, datacenter facilities and industrial systems, to achieve a variety of objectives, which range from credentials compromise to sabotage of physical devices, by means of smarter and smarter worms and rootkits. Stuxnet is a recent worm that well emphasizes the strong technical advances achieved by the attackers' community. It was discovered in July 2010 and firstly affected Iranian nuclear plants [2]. Stuxnet compromises the regular behavior of the supervisory control and data acquisition (SCADA) system by reprogramming the code of programmable logic controllers (PLC). Once compromised, PLCs can progressively destroy a device (e.g., components of a centrifuge, such as the case of the Iranian plant) by sending malicious control signals. Stuxnet combines a relevant number of challenging features: it exploits zero-days vulnerabilities of the Windows OS to affect the nodes connected to the PLC; it propagates either locally (e.g., by means of USB sticks) or remotely (e.g., via shared folders or the print spooler vulnerability); it is able to modify its behavior during the progression of the attack, and communicates with a remote command and control server. More importantly, Stuxnet can remain undetected for many months [3] because it is able to fool the SCADA system by emulating regular monitoring signals.
Domenico Cotroneo, Antonio Pecchia, Stefano Russo 0001
DSN2
2013 Event Logs for the Analysis of Software Failures: A Rule-Based Approach
abstract
Event logs have been widely used over the last three decades to analyze the failure behavior of a variety of systems. Nevertheless, the implementation of the logging mechanism lacks a systematic approach and collected logs are often inaccurate at reporting software failures: This is a threat to the validity of log-based failure analysis. This paper analyzes the limitations of current logging mechanisms and proposes a rule-based approach to make logs effective to analyze software failures. The approach leverages artifacts produced at system design time and puts forth a set of rules to formalize the placement of the logging instructions within the source code. The validity of the approach, with respect to traditional logging mechanisms, is shown by means of around 12,500 software fault injection experiments into real-world systems.
Marcello Cinque, Domenico Cotroneo, Antonio Pecchia
IEEE Trans. Software Eng.3
2012 Detection of Software Failures through Event Logs: An Experimental Study
abstract
Software faults are recognized to be among the main responsible for system failures in many application domains. Event logs play a key role to support the analysis of failures occurring under real workload conditions. Nevertheless, field experience suggests that event logs may be inaccurate at reporting software failures or they fail to provide accurate support for understanding their causes. This paper analyzes the factors that determine accurate detection of software failures through event logs. The study is based on a data set of 17,387 experiments where failures have been induced by means of software fault injection into three systems. Analysis reveals that the reporting ability of logs collected during the experiments, is not influenced by the type of fault that is activated at runtime. More importantly, analysis demonstrates that, despite the considered systems adopt very similar detection mechanisms, the ability of logs at reporting a given type of failure changes significantly across the systems. A closer inspection of collected logs reveals that characteristics, such as system architecture, placement of the logging instructions and specific supports provided by the execution environment, significantly increase accuracy of logs at runtime.
Antonio Pecchia, Stefano Russo 0001
ISSRE1
2011 Improving Log-based Field Failure Data Analysis of multi-node computing systems
abstract
Log-based Field Failure Data Analysis (FFDA) is a widely-adopted methodology to assess dependability properties of an operational system. A key step in FFDA is filtering out entries that are not useful and redundant error entries from the log. The latter is challenging: a fault, once triggered, can generate multiple errors that propagate within the system. Grouping the error entries related to the same fault manifestation is crucial to obtain realistic measurements. This paper deals with the issues of the tuple heuristic, used to group the error entries in the log, in multi-node computing systems. We demonstrate that the tuple heuristic can group entries incorrectly; thus, an improved heuristic that adopts statistical indicators is proposed. We assess the impact of inaccurate grouping on dependability measurements by comparing the results obtained with both the heuristics. The analysis encompasses the log of the Mercury cluster at the National Center for Supercomputing Applications.
Antonio Pecchia, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN1
2011 Criticality-Driven Component Integration in Complex Software Systems
Antonio Pecchia, Roberto Pietrantuono, Stefano Russo 0001
SAFECOMP1
2011 Identifying Compromised Users in Shared Computing Infrastructures: A Data-Driven Bayesian Network Approach
abstract
The growing demand for processing and storage capabilities has led to the deployment of high-performance computing infrastructures. Users log into the computing infrastructure remotely, by providing their credentials (e.g., username and password), through the public network and using well-established authentication protocols, e.g., SSH. However, user credentials can be stolen and an attacker (using a stolen credential) can masquerade as the legitimate user and penetrate the system as an insider. This paper deals with security incidents initiated by using stolen credentials and occurred during the last three years at the National Center for Supercomputing Applications (NCSA) at the University of Illinois. We analyze the key characteristics of the security data produced by the monitoring tools during the incidents and use a Bayesian network approach to correlate (i) data provided by different security tools (e.g., IDS and Net Flows) and (ii) information related to the users' profiles to identify compromised users, i.e., the users whose credentials have been stolen. The technique is validated with the real incident data. The experimental results demonstrate that the proposed approach is effective in detecting compromised users, while allows eliminating around 80% of false positives (i.e., not compromised user being declared compromised).
Antonio Pecchia, Aashish Sharma, Zbigniew T. Kalbarczyk, Domenico Cotroneo, Ravishankar K. Iyer
SRDS1
2010 Assessing and improving the effectiveness of logs for the analysis of software faults
abstract
Event logs are the primary source of data to characterize the dependability behavior of a computing system during the operational phase. However, they are inadequate to provide evidence of software faults, which are nowadays among the main causes of system outages. This paper proposes an approach based on software fault injection to assess the effectiveness of logs to keep track of software faults triggered in the field. Injection results are used to provide guidelines to improve the ability of logging mechanisms to report the effects of software faults. The benefits of the approach are shown by means of experimental results on three widely used software systems.
Marcello Cinque, Domenico Cotroneo, Roberto Natella, Antonio Pecchia
DSN4
2010 Memory leak analysis of mission-critical middleware
Gabriella Carrozza, Domenico Cotroneo, Roberto Natella, Antonio Pecchia, Stefano Russo 0001
J. Syst. Softw.4