Tommaso Zoppi

dblp:25/9388 · DBLP profile ↗
← Back
25ranked-venue papers
19as first author
15since 2021 · last 2026
0000-0001-9820-6047ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 9 first-author · 6 since 2021Software engineering, systems software and programming languages · 10 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Fail-Controlled Classifiers: A Swiss-Army Knife Toward Trustworthy Systems
abstract
ABSTRACT Background Modern critical systems often require to take decisions and classify data and scenarios autonomously without having detrimental effects on people, infrastructures or the environment, ensuring desired dependability attributes. Researchers typically strive to craft classifiers with perfect accuracy, which should be always correct and as such never threaten the encompassing system. Unfortunately, this is a very unrealistic goal, as classification tasks are typically complex and may encounter a wide variety of unexpected operating conditions and unknown inputs. Methods Classifiers should be considered as building blocks that interact with other components that help rejecting those predictions that are suspected to be misclassifications, triggering system‐level mitigation strategies instead. Fail‐Controlled Classifiers (FCCs) are software components that can either correctly classify, misclassify, or reject outputs: ideally, they would reject all and only outputs that correspond to misclassifications. Nine different FCCs are presented: Self‐Checking Classifiers (SCC), Watchdog Timers (WT), Input Processor (IP), Output processor (OP), Safety Wrapper (SW), Recovery Blocks (RB), weighted and non‐weighted Voting (VT, WVT) and Stacking (STK). Results These 9 FCCs are instantiated in experiments with tabular and image classifiers, showing their potential in rejecting most misclassifications and paving the ways for trustworthy decisions to be deployed in critical systems. If the system can tolerate more omissions, the IP FCC is a good choice. On the other hand, if achieving the highest accuracy is the priority, RB FCC performs better. Conclusions Findings show that FCCs do not primarily aim at improving correct classifications, but allow for transforming many misclassifications into rejections, which may be easily handled by the encompassing system and paving the way for trustworthy decisions to be deployed in critical systems.
Fahad Ahmed KhoKhar, Tommaso Zoppi, Andrea Ceccarelli, Leonardo Montecchi, Andrea Bondavalli
Softw. Pract. Exp.2
2025 Orchestrating Fail-Safe, Black-Box Models Within Federated Learning Scenarios
abstract
Federated Learning (FL) stands out in the realm of collaborative learning that ensures data privacy for each client involved in the federation. Despite showing astounding potential, it inherently brings challenges that are often difficult to address and make it hardly applicable in industrial scenarios. First, all clients must agree on the same algorithm to be trained locally to enable global averaging, limiting autonomy. Second, clients are required to disclose insights of their models (i.e., weights or gradients) to create the global model upon averaging and aggregating techniques, threatening privacy. Third, applications of FL in critical systems where each component has to be trusted are not solid enough at best. This paper designs Trustworthy, Black-Box FL (TBB-FL) software architectures that allow clients to locally train any algorithm they want to solve a specific task. Only executables of local models are sent to the server and herein treated as black-boxes to create the global model as an adjudication of clients' opinions. Moreover, local and global models are self-checking, software components that quantify confidence in a prediction to suspect prediction errors, ultimately rejecting outputs that cannot be trusted. We validate this approach through classification experiments on both image and tabular datasets using the Flower framework, comparing TBB-FL against traditional FL and against individual local models. TBB-FL heavily reduces misclassifications compared to traditional FL, with minimal accuracy drop, and has better classification performance than local models alone.
Fahad Ahmed KhoKhar, Tommaso Zoppi, Jamal Hussain Shah
PRDC2
2025 A Strategy for Predicting the Performance of Supervised and Unsupervised Tabular Data Classifiers
abstract
Abstract Machine Learning algorithms that perform classification are increasingly been adopted in Information and Communication Technology (ICT) systems and infrastructures due to their capability to profile their expected behavior and detect anomalies due to ongoing errors or intrusions. Deploying a classifier for a given system requires conducting comparison and sensitivity analyses that are time-consuming, require domain expertise, and may even not achieve satisfactory classification performance, resulting in a waste of money and time for practitioners and stakeholders. This paper predicts the expected performance of classifiers without needing to select, craft, exercise, or compare them, requiring minimal expertise and machinery. Should classification performance be predicted worse than expectations, the users could focus on improving data quality and monitoring systems instead of wasting time in exercising classifiers, saving key time and money. The prediction strategy uses scores of feature rankers, which are processed by regressors to predict metrics such as Matthews Correlation Coefficient (MCC) and Area Under ROC-Curve (AUC) for quantifying classification performance. We validate our prediction strategy through a massive experimental analysis using up to 12 feature rankers that process features from 23 public datasets, creating additional variants in the process and exercising supervised and unsupervised classifiers. Our findings show that it is possible to predict the value of performance metrics for supervised or unsupervised classifiers with a mean average error (MAE) of residuals lower than 0.1 for many classification tasks. The predictors are publicly available in a Python library whose usage is straightforward and does not require domain-specific skill or expertise.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
Data Sci. Eng.1
2024 Deploying a Generic Threat Model for Detecting Anomalies in a Power Grid Digital Twin
abstract
Monitoring power grid infrastructures typically generates a massive amount of power consumption data related to different components or communication channels. This is typically employed for power optimization, but does not suffice for conducting other key tasks for guaranteeing desirable properties as reliability, safety and security. In these cases, the grid should be monitored for detecting anomalies due to security threats, component failures, environmental damages, or other hazards. This is the case of the Grid Data’s Digital Twin industrial scenario, which provides an up-to-date grid image that combines actual measurement data and a time-series-based grid model that closely approximates reality. To tackle this, this paper analyzes the state of the art of existing threat models for smart grids, proposing a generic and comprehensive threat and anomaly model that is then used to craft power consumption anomaly detectors for the case study above. This work was conducted by members from academia and industrial partners to show how to deploy power consumption anomaly detectors in the wild, showing a methodology that is generic enough to be applied also by other stakeholders.
Tommaso Zoppi, Irene Bicchierai, Francesco Brancati, Andrea Bondavalli, Hans-Peter Schwefel
PRDC1
2024 Fail-Controlled Classifiers: Do they Know when they don't Know?
abstract
Domain experts are desperately looking to solve decision-making problems by designing and training Machine Learning algorithms that can perform classification with the highest possible accuracy. No matter how hard they try, classifiers will always be prone to misclassifications due to a variety of reasons that make the decision boundary unclear. This complicates the integration of classifiers into critical systems, where misclassifications could directly impact people, infrastructures, or the environment. The paper proposes to consider a classifier as a structural part of the system instead of an individual component to be tested in isolation and included in the system afterward. This allows for omitting those predictions that are suspected to be misclassifications, triggering system-level mitigation strategies. The resulting fail-controlled classifiers (FCCs) are software components that can correctly classify, misclassify, or omit outputs: ideally, they would omit all and only outputs that correspond to misclassifications. After presenting the theoretical foundations of FCCs, the paper proposes metrics to quantify their performance, 5 software architectures for FCCs, and an experimental analysis involving tabular data and image classifiers. Overall, this paper advocates the need for a system and software design in which ML classifiers are not separate components, but should rather be considered building blocks that interact with other components for improved performance.
Tommaso Zoppi, Fahad Ahmed KhoKhar, Andrea Ceccarelli, Leonardo Montecchi, Andrea Bondavalli
PRDC1
2024 Anomaly-based error and intrusion detection in tabular data: No DNN outperforms tree-based classifiers
abstract
Recent years have seen a growing involvement of researchers and practitioners in crafting Deep Neural Networks (DNNs) that seem to outperform existing machine learning approaches for solving classification problems as anomaly-based error and intrusion detection. Undoubtedly, classifiers may be very diverse among themselves, and choosing one or another is typically due to the specific task and target system. Designing and training the optimal tabular data classifier requires extensive experimentation, sensitivity analyses, big datasets, and domain-specific knowledge that may not be available at will or considered a non-strategical asset by many companies and stakeholders. This paper compares, using a total of 23 public datasets: i) traditional (tree-based, statistical) supervised classifiers, ii) DNNs that are specifically designed for classifying tabular data, iii) DNNs for image classification that are applied to tabular data after converting data points into images, alone and as ensembles. Experimental results and related discussions show clear advantages in adopting tree-based classifiers for anomaly-based error and intrusion detection in tabular data as they outperform their competitors, including DNNs. Then, individual classifiers are compared against ensembles using different combinations of the classifiers considered in this study as base-learners, providing a unified final response through many meta-learning strategies. Results show that there is no benefit in building ensembles instead of using a tree-based classifier as Random Forests, eXtreme Gradient Boosting or Extra Trees. The paper concludes that anomaly-based error and intrusion detectors for critical systems should use the old (but gold) tree-based classifiers, which are also easier to fine-tune, and understand; plus, they require less time and resources to learn their model.
Tommaso Zoppi, Stefano Gazzini, Andrea Ceccarelli
Future Gener. Comput. Syst.1
2023 Ensembling Uncertainty Measures to Improve Safety of Black-Box Classifiers
abstract
Machine Learning (ML) algorithms that perform classification may predict the wrong class, experiencing misclassifications. It is well-known that misclassifications may have cascading effects on the encompassing system, possibly resulting in critical failures. This paper proposes SPROUT, a Safety wraPper thROugh ensembles of UncertainTy measures, which suspects misclassifications by computing uncertainty measures on the inputs and outputs of a black-box classifier. If a misclassification is detected, SPROUT blocks the propagation of the output of the classifier to the encompassing system. The resulting impact on safety is that SPROUT transforms erratic outputs (misclassifications) into data omission failures, which can be easily managed at the system level. SPROUT has a broad range of applications as it fits binary and multi-class classification, comprising image and tabular datasets. We experimentally show that SPROUT always identifies a huge fraction of the misclassifications of supervised classifiers, and it is able to detect all misclassifications in specific cases. SPROUT implementation contains pre-trained wrappers, it is publicly available and ready to be deployed with minimal effort.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
ECAI1
2023 Intrusion detection without attack knowledge: generating out-of-distribution tabular data
abstract
Anomaly-based intrusion detectors are machine learners trained to distinguish between normal and anomalous data. The normal data is generally easy to collect when building the train set; instead, collecting anomalous data requires historical data or penetration testing campaigns. Unfortunately, the first is most often unavailable or unusable, and the latter is usually expensive and unfeasible, as it requires hacking the target system. It turns out that the possibility of training an intrusion detector without attack knowledge, i.e., without anomalies, is attractive. This paper reviews strategies to train anomaly detectors in the absence of anomalies, from shallow machine learning to deep learning and computer vision approaches, and applies such strategies to the domain of intrusion detection. We experimentally show that training an intrusion detector without attack knowledge is effective when normal and attack data distributions are distinguishable. Detection performance severely drops in the case of complex (but more realistic) datasets, making all the existing solutions inadequate for real applications. However, the recent advancements of out-of-distribution research in deep learning and computer vision show interesting prospective results.
Andrea Ceccarelli, Tommaso Zoppi
ISSRE2
2023 Anomaly Detectors for Self-Aware Edge and IoT Devices
abstract
With the growing processing power of computing systems and the increasing availability of massive datasets, machine learning algorithms have led to major breakthroughs in many different areas. This applies also to resource-constrained IoT and edge devices, which will often benefit from relatively small – but smart – local anomaly detection tasks that aim at protecting the device, or the information they convey from sensors towards a central node. This provides the device with fault detection capabilities that are typically required when engineering dependable devices, services or systems. This paper overviews a pitfall-free process to provide small devices with anomaly detection capabilities, to make them self-aware of their health condition, and possibly take appropriate countermeasures. Our methodology applies to a wide range of Linux-based devices: we show an application to a specific ARANCINO device, which has already been successfully used in many smart cities and sensing applications. We craft anomaly detectors that are very effective in detecting most of the anomalies. Additionally, we comment on the beneficial impact of time-series analysis, which could help improve detection performance even further, allowing to equip any small device with responsive and accurate anomaly detection machinery.
Tommaso Zoppi, Giovanni Merlino, Andrea Ceccarelli, Antonio Puliafito, Andrea Bondavalli
QRS1
2023 Which algorithm can detect unknown attacks? Comparison of supervised, unsupervised and meta-learning algorithms for intrusion detection
abstract
There is an astounding growth in the adoption of machine learners (MLs) to craft intrusion detection systems (IDSs). These IDSs model the behavior of a target system during a training phase, making them able to detect attacks at runtime. Particularly, they can detect known attacks, whose information is available during training, at the cost of a very small number of false alarms, i.e., the detector suspects attacks but no attack is actually threatening the system. However, the attacks experienced at runtime will likely differ from those learned during training and thus will be unknown to the IDS. Consequently, the ability to detect unknown attacks becomes a relevant distinguishing factor for an IDS. This study aims to evaluate and quantify such ability by exercising multiple ML algorithms for IDSs. We apply 47 supervised, unsupervised, deep learning, and meta-learning algorithms in an experimental campaign embracing 11 attack datasets, and with a methodology that simulates the occurrence of unknown attacks. Detecting unknown attacks is not trivial: however, we show how unsupervised meta-learning algorithms have better detection capabilities of unknowns and may even outperform classification performance of other ML algorithms when dealing with unknown attacks.
Tommaso Zoppi, Andrea Ceccarelli, Tommaso Puccetti, Andrea Bondavalli
Comput. Secur.1
2023 Safe Maintenance of Railways using COTS Mobile Devices: The Remote Worker Dashboard
abstract
The railway domain is regulated by rigorous safety standards to ensure that specific safety goals are met. Often, safety-critical systems rely on custom hardware-software components that are built from scratch to achieve specific functional and non-functional requirements. Instead, the (partial) usage of Commercial Off-The-Shelf (COTS) components is very attractive as it potentially allows reducing cost and time to market. Unfortunately, COTS components do not individually offer enough guarantees in terms of safety and security to be used in critical systems as they are. In such a context, RFI (Rete Ferroviaria Italiana), a major player in Europe for railway infrastructure management, aims at equipping track-side workers with COTS devices to remotely and safely interact with the existing interlocking system, drastically improving the performance of maintenance operations. This paper describes the first effort to update existing (embedded) railway systems to a more recent cyber-physical system paradigm. Our Remote Worker Dashboard (RWD) pairs the existing safe interlocking machinery alongside COTS mobile components, making cyber and physical components cooperate to provide the user with responsive, safe, and secure service. Specifically, the RWD is a SIL4 cyber-physical system to support maintenance of actuators and railways in which COTS mobile devices are safely used by track-side workers. The concept, development, implementation, verification, and validation activities to build the RWD were carried out in compliance with the applicable CENELEC standards required by certification bodies to declare compliance with specific guidelines.
Tommaso Zoppi, Innocenzo Mungiello, Andrea Ceccarelli, Alberto Cirillo, Lorenzo Sarti, Lorenzo Esposito, Giuseppe Scaglione, Sergio Repetto, Andrea Bondavalli
ACM Trans. Cyber Phys. Syst.1
2021 Detecting Intrusions by Voting Diverse Machine Learners: Is It Really Worth?
abstract
Recent years have seen an astounding growth in the adoption of Machine Learning algorithms to classify data gathered through monitoring activities. Those algorithms can effectively classify data as system indicators, network packets, and logs according to a model they infer during training. This way, they provide sophisticated means to conduct intrusion detection by suspecting anomalies due to attacks in the value of those features. Additionally, Meta-Learners as Bagging and Boosting build ensembles of homogeneous classifiers that are known to improve classification performance with positive impact on intrusion detection. On the other hand, it is not yet clear if ensembles of heterogeneous or diverse classifiers can build better intrusion detectors. To such extent, we first recap on n-version programming, k-out-of-m (k-o-o-m) systems and the role of diversity. Then, we present k-o-o-m systems of classifiers for intrusion detection, expanding on meta-learning and diversity measures to be applied to classifiers. This paves the way for an experimental campaign which exercises supervised and unsupervised classifiers as well as k-o-o-m voting ensembles. After presenting and discussing results, we conclude that voting ensembles of diverse classifiers does not improve intrusion detection. Therefore, while voting has been acknowledged since decades as a staple to manage n-version programming for reliable systems engineering, it is not as effective as a meta-learner to improve classification performance of intrusion detectors.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
PRDC1
2021 Prepare for trouble and make it double! Supervised - Unsupervised stacking for anomaly-based intrusion detection
Tommaso Zoppi, Andrea Ceccarelli
J. Netw. Comput. Appl.1
2021 Meta-Learning to Improve Unsupervised Intrusion Detection in Cyber-Physical Systems
abstract
Artificial Intelligence (AI)- based classifiers rely on Machine Learning (ML) algorithms to provide functionalities that system architects are often willing to integrate into critical Cyber-Physical Systems (CPSs) . However, such algorithms may misclassify observations, with potential detrimental effects on the system itself or on the health of people and of the environment. In addition, CPSs may be subject to threats that were not previously known, motivating the need for building Intrusion Detectors (IDs) that can effectively deal with zero-day attacks. Different studies were directed to compare misclassifications of various algorithms to identify the most suitable one for a given system. Unfortunately, even the most suitable algorithm may still show an unsatisfactory number of misclassifications when system requirements are strict. A possible solution may rely on the adoption of meta-learners, which build ensembles of base-learners to reduce misclassifications and that are widely used for supervised learning. Meta-learners have the potential to reduce misclassifications with respect to non-meta learners: however, misleading base-learners may let the meta-learner leaning towards misclassifications and therefore their behavior needs to be carefully assessed through empirical evaluation. To such extent, in this paper we investigate, expand, empirically evaluate, and discuss meta-learning approaches that rely on ensembles of unsupervised algorithms to detect (zero-day) intrusions in CPSs. Our experimental comparison is conducted by means of public datasets belonging to network intrusion detection and biometric authentication systems, which are common IDSs for CPSs. Overall, we selected 21 datasets, 15 unsupervised algorithms and 9 different meta-learning approaches. Results allow discussing the applicability and suitability of meta-learning for unsupervised anomaly detection, comparing metric scores achieved by base algorithms and meta-learners. Analyses and discussion end up showing how the adoption of meta-learners significantly reduces misclassifications when detecting (zero-day) intrusions in CPSs.
Tommaso Zoppi, Mohamad Gharib, Muhammad Atif 0001, Andrea Bondavalli
ACM Trans. Cyber Phys. Syst.1
2021 MADneSs: A Multi-Layer Anomaly Detection Framework for Complex Dynamic Systems
abstract
Anomaly detection can infer the presence of errors without observing the target services, but detecting variations in the observable parts of the system on which the services reside. This is a promising technique in complex software-intensive systems, because either instrumenting the services' internals is exceedingly time-consuming, or encapsulation makes them not accessible. Unfortunately, in such systems anomaly detection is often ineffective due to their dynamicity, which implies changes in the services or their expected workload. Here we present our approach to enhance the efficacy of anomaly detection in complex, dynamic software-intensive systems. After discussing the related challenges, we present MADneSs, an anomaly detection framework tailored for the above systems that includes an adaptive multi-layer monitoring module. Monitored data are then processed by the anomaly detector, which adapts its parameters depending on the current system behavior. An anomaly alert is provided if the analysis conducted by the anomaly detector identify unexpected trends in the data. MADneSs is evaluated through an experimental campaign on two service-oriented architectures; software faults are injected in the application layer, and detected through monitoring of underlying system layers. Lastly, we quantitatively and qualitatively discuss our results with respect to state-of-the-art solutions, highlighting the key contributions of MADneSs.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
IEEE Trans. Dependable Secur. Comput.1
2020 On the educated selection of unsupervised algorithms via attacks and anomaly classes
abstract
Anomaly detection aims at finding patterns in data that do not conform to the expected behavior. It is largely adopted in intrusion detection systems, relying on unsupervised algorithms that have the potential to detect zero-day attacks; however, efficacy of algorithms varies depending on the observed system and the attacks. Selecting the algorithm that maximizes detection capability is a challenging task with no master key. This paper tackles the challenge above by devising and applying a methodology to identify relations between attack families, anomaly classes and algorithms. The implication is that an unknown attack belonging to a specific attack family is most likely to get observed by unsupervised algorithms that are particularly effective on such attack family. This paves the way to rules for the selection of algorithms based on the identification of attack families. The paper proposes and applies a methodology based on analytical and experimental investigations supported by a tool to i) identify which anomaly classes are most likely raised by the different attack families, ii) study suitability of anomaly detection algorithms to detect anomaly classes, iii) combine previous results to relate anomaly detection algorithms and attack families, and iv) define guidelines to select unsupervised algorithms for intrusion detection.
Tommaso Zoppi, Andrea Ceccarelli, Lorenzo Salani, Andrea Bondavalli
J. Inf. Secur. Appl.1
2019 Evaluation of Anomaly Detection Algorithms Made Easy with RELOAD
abstract
Anomaly detection aims at identifying patterns in data that do not conform to the expected behavior. Despite anomaly detection has been arising as one of the most powerful techniques to suspect attacks or failures, dedicated support for the experimental evaluation is actually scarce. In fact, existing frameworks are mostly intended for the broad purposes of data mining and machine learning. Intuitive tools tailored for evaluating anomaly detection algorithms for failure and attack detection with an intuitive support to sliding windows are currently missing. This paper presents RELOAD, a flexible and intuitive tool for the Rapid EvaLuation Of Anomaly Detection algorithms. RELOAD is able to automatically i) fetch data from an existing data set, ii) identify the most informative features of the data set, iii) run anomaly detection algorithms, including those based on sliding windows, iv) apply multiple strategies to features and decide on anomalies, and v) provide conclusive results following an extensive set of metrics, along with plots of algorithms scores. Finally, RELOAD includes a simple GUI to set up the experiments and examine results. After describing the structure of the tool and detailing inputs and outputs of RELOAD, we exercise RELOAD to analyze an intrusion detection dataset available on a public platform, showing its setup, metric scores and plots.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
ISSRE1
2019 An Initial Investigation on Sliding Windows for Anomaly-Based Intrusion Detection
abstract
The growing systems complexity calls for dedicated monitoring and data analysis strategies aiming to detect faults, attacks and errors before they escalate into failures. Distributed and heterogeneous systems are more likely to expose vulnerabilities that attackers may target to get unauthorized access to a system, make it unavailable or steal sensitive data. As countermeasure, traditionally techniques for attacks and intrusion detection are based on signature recognition and requires knowledge on the attacks pattern: therefore, they are not well-suited to detect zero-days attacks. A viable alternative is anomaly detection, where deviation from the expected behavior are suspected as attacks. However, anomaly detection is generally not applicable in systems where the expected behavior changes through time. In this paper we explore anomaly detection strategies based on sliding windows, which are intended for evolving and dynamic systems as IoT, in which system configuration and behavior may change continuously. We first describe the context and the key features of sliding windows, and then we proceed detailing their possible drawbacks. Discussion is substantiated by quantitative analyses directed to evaluate detection capabilities. The experimental campaign is based on state-of-the-art algorithms and datasets, and results have been made publicly available.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
SERVICES1
2019 Threat Analysis in Systems-of-Systems: An Emergence-Oriented Approach
abstract
Cyber-physical Systems of Systems (SoSs) are large-scale systems made of independent and autonomous cyber-physical Constituent Systems (CSs) which may interoperate to achieve high-level goals also with the intervention of humans. Providing security in such SoSs means, among other features, forecasting and anticipating evolving SoS functionalities, ultimately identifying possible detrimental phenomena that may result from the interactions of CSs and humans. Such phenomena, usually called emergent phenomena , are often complex and difficult to capture: the first appearance of an emergent phenomenon in a cyber-physical SoS is often a surprise to the observers. Adequate support to understand emergent phenomena will assist in reducing both the likelihood of design or operational flaws, and the time needed to analyze the relations amongst the CSs, which always has a key economic significance. This article presents a threat analysis methodology and a supporting tool aimed at (i) identifying (emerging) threats in evolving SoSs, (ii) reducing the cognitive load required to understand an SoS and the relations among CSs, and (iii) facilitating SoS risk management by proposing mitigation strategies for SoS administrators. The proposed methodology, as well as the tool, is empirically validated on Smart Grid case studies by submitting questionnaires to a user base composed of 3 stakeholders and 18 BSc and MSc students.
Andrea Ceccarelli, Tommaso Zoppi, Alexandr Vasenev, Marco Mori, Dan Ionita, Lorena Montoya, Andrea Bondavalli
ACM Trans. Cyber Phys. Syst.2
2018 On Algorithms Selection for Unsupervised Anomaly Detection
abstract
Anomaly detection, which aims at identifying unexpected trends and data patterns, has widely been used to build error detectors, failure predictors or intrusion detectors. Internal faults or malicious attacks have a different impact on the behavior of the system. They usually manifest as different observable deviations from the expected behavior, which may be identified by anomaly detection algorithms. Our study aims at investigating the suitability of unsupervised algorithms and their families in detecting either point, contextual or collective anomalies. To provide a complete picture, we consider both sliding and non-sliding window algorithms which operate in unsupervised mode. Along with qualitative analyses of each algorithm and family, we conduct an experimental campaign in which we run each algorithm on three state-of-the-art datasets in which we inject either point, contextual or collective anomalies. Results show that non-sliding algorithms are capable to detect point and collective anomalies, while they cannot effectively deal with contextual ones. Instead, sliding window algorithms require shorter periods of training and naturally build a local context, which allow them to effectively deal with contextual anomalies. Such observations are summarized to support the choice of the correct algorithm depending on the investigated class(es) of anomaly.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
PRDC1
2018 Labelling relevant events to support the crisis management operator
abstract
Abstract Thanks to the large availability of portable devices and the growing interest in the Internet of Things, during crises, social networks, or alerts sent through mobile devices or sensor networks are available and can be matched each other to perform situational analysis. However, the inclusion of multiple heterogeneous sources in situational analyses leads to 2 main issues: (1) a source could deliver (voluntarily or erroneously) wrong data damaging the integrity and the correctness of the analysis, and (2) a significant amount of heterogeneous data need to be processed. As a consequence, the crisis management operator faces a large amount of potentially unreliable data. In this paper, we present a relevance labelling strategy to process information gathered from heterogeneous data streams to select the most relevant events. These are presented to the crisis management operator with the highest priority. Our strategy is evaluated using events collected by the Secure! crisis management system, considering 3 real crisis scenarios happened in Italy in 2015. Results show that our strategy is able to correctly identify sets of relevant events, supporting the activities of the crisis management operator.
Tommaso Zoppi, Andrea Ceccarelli, Francesco Lo Piccolo, Paolo Lollini, Gabriele Giunta, Vito Morreale, Andrea Bondavalli
J. Softw. Evol. Process.1
2016 Context-Awareness to Improve Anomaly Detection in Dynamic Service Oriented Architectures
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
SAFECOMP1
2016 Challenging Anomaly Detection in Complex Dynamic Systems
abstract
Software infrastructures are becoming more and more complex, making performance and dependability monitoring in wide and dynamic contexts such as Distributed Systems, Systems of Systems (SoS) and Cloud environments an unachievable goal. Consequently, it is very difficult to know how all the specific parts, services and modules of these systems behave. This negatively impacts our ability in detecting anomalies, because the boundaries between normal and anomalous behaviors are not always known. The paper describes the context and the targeted problem highlighting the research directions that the student will follow in the next years. In particular, after introducing the relevance of this work with respect to the academic and the industrial state of the art, we carefully define the problem and summarize the main challenges that arise according to such problem definition.
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli
SRDS1
2015 A Multi-layer Anomaly Detector for Dynamic Service-Based Systems
Andrea Ceccarelli, Tommaso Zoppi, Massimiliano Leone Itria, Andrea Bondavalli
SAFECOMP2
2014 A Testbed for Evaluating Anomaly Detection Monitors through Fault Injection
abstract
Amongst the features of Service Oriented Architectures (SOAs), their flexibility, dynamicity, and scalability make them particularly attractive for adoption in the ICT infrastructure of organizations. Such features come at the cost of improved difficulty in monitoring the SOA for error detection: i) faults may manifest themselves differently due to services and SOA evolution, and ii) interactions between a service and its monitors may need reconfiguration at each service update. This calls for monitoring solutions that operate at different layers than the application layer (services layer). In this paper we present our ongoing work towards the definition of a monitoring framework for SOAs and services, which relies on anomaly detection performed at the Application Server (AS) and the Operating System (OS) layers to identify events whose manifestation or effect is not adequately described a-priori. Specifically the paper introduces the key concepts of our work and presents the case study built to exercise and set-up our monitor. The case study uses Life ray as application layer and it includes fault injection and data collection instruments to perform extended testing campaigns.
Andrea Ceccarelli, Tommaso Zoppi, Andrea Bondavalli, Fabio Duchi, Giuseppe Vella
ISORC2