VLDB 2026 Research / reviewers in the wild / expert
Andrea Bondavalli
dblp:31/4416
· DBLP profile ↗
99ranked-venue papers
26as first author
19since 2021 · last 2026
0000-0001-7366-6530ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 40 · 9 first-author · 6 since 2021Software engineering, systems software and programming languages · 21 · 2 first-author · 10 since 2021Systems, architecture and hardware · 17 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 5 since 2021Computer networks · 8 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fail-Controlled Classifiers: A Swiss-Army Knife Toward Trustworthy SystemsabstractABSTRACT Background Modern critical systems often require to take decisions and classify data and scenarios autonomously without having detrimental effects on people, infrastructures or the environment, ensuring desired dependability attributes. Researchers typically strive to craft classifiers with perfect accuracy, which should be always correct and as such never threaten the encompassing system. Unfortunately, this is a very unrealistic goal, as classification tasks are typically complex and may encounter a wide variety of unexpected operating conditions and unknown inputs. Methods Classifiers should be considered as building blocks that interact with other components that help rejecting those predictions that are suspected to be misclassifications, triggering system‐level mitigation strategies instead. Fail‐Controlled Classifiers (FCCs) are software components that can either correctly classify, misclassify, or reject outputs: ideally, they would reject all and only outputs that correspond to misclassifications. Nine different FCCs are presented: Self‐Checking Classifiers (SCC), Watchdog Timers (WT), Input Processor (IP), Output processor (OP), Safety Wrapper (SW), Recovery Blocks (RB), weighted and non‐weighted Voting (VT, WVT) and Stacking (STK). Results These 9 FCCs are instantiated in experiments with tabular and image classifiers, showing their potential in rejecting most misclassifications and paving the ways for trustworthy decisions to be deployed in critical systems. If the system can tolerate more omissions, the IP FCC is a good choice. On the other hand, if achieving the highest accuracy is the priority, RB FCC performs better. Conclusions Findings show that FCCs do not primarily aim at improving correct classifications, but allow for transforming many misclassifications into rejections, which may be easily handled by the encompassing system and paving the way for trustworthy decisions to be deployed in critical systems. Fahad Ahmed KhoKhar, Tommaso Zoppi, Andrea Ceccarelli, Leonardo Montecchi, Andrea Bondavalli |
Softw. Pract. Exp. | 5 |
| 2025 | Quantitative Comparison of System Architectures for an SAE Level 4 Highway PilotabstractThe development of SAE Level 4 autonomous driving systems requires fault-tolerant architectures to ensure proper operation and robustness under demanding operational conditions. These architectures must gracefully handle hardware and software faults to maintain system functionality for a minimum operational timeframe, ensuring a fail-operational capability. The selection of an appropriate system architecture is crucial for managing the complexity of autonomous driving systems while achieving effective and efficient fault tolerance. This paper presents a preliminary quantitative comparison of different system architectures—symmetric and asymmet-ric—for a reference use case of an SAE Level 4 Highway Pilot. Specifically, we evaluate a Triple Modular Redundancy symmetric architecture, a Channel-Wise Doer/Checker/Fallback asymmetric architecture, and a Layer-Wise Doer/Checker/-Fallback asymmetric architecture. The comparison leverages Stochastic Activity Networks (SAN) modeled in the Mobius tool to assess key dependability metrics, particularly safety and reliability. Our preliminary results provide initial insights into the differences between symmetry and asymmetry in faulttolerant designs, highlighting their potential impact on safetycritical autonomous driving systems. This work contributes to a deeper understanding of architectural trade-offs in the design of dependable autonomous systems and provides guidance to system designers in making informed architectural decisions while strengthening safety argumentation. Manuel Drago, Andrea Bondavalli, Paolo Lollini, Georg Niedrist, Moritz Antlanger |
QRS | 2 |
| 2025 | Analyzing the 2015 Ukraine Power Grid Cyber-Attack: A Quantitative Assessment of Adversary Behavior and Impact*abstractThe security of critical infrastructures, such as power grids, water treatment facilities, transportation networks, financial systems, and communication networks, is essential for social stability. These systems deliver vital services, but are increasingly reliant on digital control mechanisms, making them vulnerable to cyber threats. A successful cyber-attack on any of these infrastructures could lead to widespread disruptions, significant financial losses, and in severe cases, risks to public safety. An effective cyber-security risk assessment process requires structured methodologies that identify vulnerabilities and anticipate adversarial behavior.Traditional risk assessment approaches rely on static and qualitative analyses that focus on known vulnerabilities and configurations, but lack dynamic attack simulation. In contrast, formal modeling and simulation-based techniques provide a quantitative framework to analyze possible attack paths and their likelihood of success. Among these formal methods, the ADVISE (ADversary VIew Security Evaluation) formalism offers a structured approach to assess cyber threats from the perspective of an adversary.This paper explores the application of the formal security evaluation framework, ADVISE, to model and analyze the 2015 Ukraine Power Grid cyber-attack. It specifically highlights the impact and the importance of the execution timing of the attacks, the adversary capabilities, and the effects of countermeasures throughout the progression of cyber-attacks. This framework simulates attack dynamics and quantifies the security risks associated with the Ukrainian Power Grid, thereby complementing the qualitative analyses conducted in previous studies. Marzieh Kordi, Syed Muhammad Fasih Ali, Paolo Lollini, Andrea Bondavalli |
SMC | 4 |
| 2025 | A Strategy for Predicting the Performance of Supervised and Unsupervised Tabular Data ClassifiersabstractAbstract Machine Learning algorithms that perform classification are increasingly been adopted in Information and Communication Technology (ICT) systems and infrastructures due to their capability to profile their expected behavior and detect anomalies due to ongoing errors or intrusions. Deploying a classifier for a given system requires conducting comparison and sensitivity analyses that are time-consuming, require domain expertise, and may even not achieve satisfactory classification performance, resulting in a waste of money and time for practitioners and stakeholders. This paper predicts the expected performance of classifiers without needing to select, craft, exercise, or compare them, requiring minimal expertise and machinery. Should classification performance be predicted worse than expectations, the users could focus on improving data quality and monitoring systems instead of wasting time in exercising classifiers, saving key time and money. The prediction strategy uses scores of feature rankers, which are processed by regressors to predict metrics such as Matthews Correlation Coefficient (MCC) and Area Under ROC-Curve (AUC) for quantifying classification performance. We validate our prediction strategy through a massive experimental analysis using up to 12 feature rankers that process features from 23 public datasets, creating additional variants in the process and exercising supervised and unsupervised classifiers. Our findings show that it is possible to predict the value of performance metrics for supervised or unsupervised classifiers with a mean average error (MAE) of residuals lower than 0.1 for many classification tasks. The predictors are publicly available in a Python library whose usage is straightforward and does not require domain-specific skill or expertise. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
Data Sci. Eng. | 3 |
| 2025 | A Systematic Literature Review on Application of Agile Software Development Process Models for the Development of Safety-Critical Systems in Multiple DomainsabstractThis paper presents a literature review on using agile for safety‐critical systems (SCSs). We have systematically selected and evaluated relevant literature to find out major areas of concern for adapting agile in the development of SCSs. In the paper, we have listed the most used Agile process models and reasons for their suitability for SCS, then we have outlined phases of the software development life cycle (SDLC) where changes are required to make an agile process suitable for the development of SCSs. Thirdly, we have elaborated on problems and other important aspects according to specific domains where agile is used for SCS. This paper serves as an insight into the latest trends and problems regarding the use of Agile process models to develop SCSs. H. Maria Maqsood, Joelma Choma, Eduardo Guerra 0001, Andrea Bondavalli |
IET Softw. | 4 |
| 2024 | Deploying a Generic Threat Model for Detecting Anomalies in a Power Grid Digital TwinabstractMonitoring power grid infrastructures typically generates a massive amount of power consumption data related to different components or communication channels. This is typically employed for power optimization, but does not suffice for conducting other key tasks for guaranteeing desirable properties as reliability, safety and security. In these cases, the grid should be monitored for detecting anomalies due to security threats, component failures, environmental damages, or other hazards. This is the case of the Grid Data’s Digital Twin industrial scenario, which provides an up-to-date grid image that combines actual measurement data and a time-series-based grid model that closely approximates reality. To tackle this, this paper analyzes the state of the art of existing threat models for smart grids, proposing a generic and comprehensive threat and anomaly model that is then used to craft power consumption anomaly detectors for the case study above. This work was conducted by members from academia and industrial partners to show how to deploy power consumption anomaly detectors in the wild, showing a methodology that is generic enough to be applied also by other stakeholders. Tommaso Zoppi, Irene Bicchierai, Francesco Brancati, Andrea Bondavalli, Hans-Peter Schwefel |
PRDC | 4 |
| 2024 | Fail-Controlled Classifiers: Do they Know when they don't Know?abstractDomain experts are desperately looking to solve decision-making problems by designing and training Machine Learning algorithms that can perform classification with the highest possible accuracy. No matter how hard they try, classifiers will always be prone to misclassifications due to a variety of reasons that make the decision boundary unclear. This complicates the integration of classifiers into critical systems, where misclassifications could directly impact people, infrastructures, or the environment. The paper proposes to consider a classifier as a structural part of the system instead of an individual component to be tested in isolation and included in the system afterward. This allows for omitting those predictions that are suspected to be misclassifications, triggering system-level mitigation strategies. The resulting fail-controlled classifiers (FCCs) are software components that can correctly classify, misclassify, or omit outputs: ideally, they would omit all and only outputs that correspond to misclassifications. After presenting the theoretical foundations of FCCs, the paper proposes metrics to quantify their performance, 5 software architectures for FCCs, and an experimental analysis involving tabular data and image classifiers. Overall, this paper advocates the need for a system and software design in which ML classifiers are not separate components, but should rather be considered building blocks that interact with other components for improved performance. Tommaso Zoppi, Fahad Ahmed KhoKhar, Andrea Ceccarelli, Leonardo Montecchi, Andrea Bondavalli |
PRDC | 5 |
| 2023 | Ensembling Uncertainty Measures to Improve Safety of Black-Box ClassifiersabstractMachine Learning (ML) algorithms that perform classification may predict the wrong class, experiencing misclassifications. It is well-known that misclassifications may have cascading effects on the encompassing system, possibly resulting in critical failures. This paper proposes SPROUT, a Safety wraPper thROugh ensembles of UncertainTy measures, which suspects misclassifications by computing uncertainty measures on the inputs and outputs of a black-box classifier. If a misclassification is detected, SPROUT blocks the propagation of the output of the classifier to the encompassing system. The resulting impact on safety is that SPROUT transforms erratic outputs (misclassifications) into data omission failures, which can be easily managed at the system level. SPROUT has a broad range of applications as it fits binary and multi-class classification, comprising image and tabular datasets. We experimentally show that SPROUT always identifies a huge fraction of the misclassifications of supervised classifiers, and it is able to detect all misclassifications in specific cases. SPROUT implementation contains pre-trained wrappers, it is publicly available and ready to be deployed with minimal effort. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
ECAI | 3 |
| 2023 | Modeling of GPGPU architectures for performance analysis of CUDA programsabstractGraphics Processing Units (GPUs), originally developed for computer graphics, are now commonly used to accelerate parallel applications. Given that GPUs are designed to be as efficient as possible, evaluating their performance is crucial. This problem has been tackled in the last years by researchers that started to propose solutions such as analytical models and digital simulators, which are, however, often complex to use and/or to adapt to the needs of the user. Thanks to its high flexibility, model-based analysis is widely used to evaluate systems’ properties, including performance. Researchers started working on developing GPU models that can represent both their architecture and the software in execution, but they often use strong assumptions that undermine their usability. In this work we develop a Stochastic Activity Network model to evaluate the performance of CUDA applications running on NVIDIA GPUs. The model takes as input a representation of the program’s instruction, parsed from the CUDA SASS assembly file, and a list of parameters to offer configurability to the user. We tune our model to match the architecture of two different NVIDIA GPUs and simulate the execution of a CUDA program. We then compare the results with those obtained from the execution of the program over the real GPUs. Francesco Terrosi, Francesco Mariotti, Paolo Lollini, Andrea Bondavalli |
QRS | 4 |
| 2023 | Anomaly Detectors for Self-Aware Edge and IoT DevicesabstractWith the growing processing power of computing systems and the increasing availability of massive datasets, machine learning algorithms have led to major breakthroughs in many different areas. This applies also to resource-constrained IoT and edge devices, which will often benefit from relatively small – but smart – local anomaly detection tasks that aim at protecting the device, or the information they convey from sensors towards a central node. This provides the device with fault detection capabilities that are typically required when engineering dependable devices, services or systems. This paper overviews a pitfall-free process to provide small devices with anomaly detection capabilities, to make them self-aware of their health condition, and possibly take appropriate countermeasures. Our methodology applies to a wide range of Linux-based devices: we show an application to a specific ARANCINO device, which has already been successfully used in many smart cities and sensing applications. We craft anomaly detectors that are very effective in detecting most of the anomalies. Additionally, we comment on the beneficial impact of time-series analysis, which could help improve detection performance even further, allowing to equip any small device with responsive and accurate anomaly detection machinery. Tommaso Zoppi, Giovanni Merlino, Andrea Ceccarelli, Antonio Puliafito, Andrea Bondavalli |
QRS | 5 |
| 2023 | Which algorithm can detect unknown attacks? Comparison of supervised, unsupervised and meta-learning algorithms for intrusion detectionabstractThere is an astounding growth in the adoption of machine learners (MLs) to craft intrusion detection systems (IDSs). These IDSs model the behavior of a target system during a training phase, making them able to detect attacks at runtime. Particularly, they can detect known attacks, whose information is available during training, at the cost of a very small number of false alarms, i.e., the detector suspects attacks but no attack is actually threatening the system. However, the attacks experienced at runtime will likely differ from those learned during training and thus will be unknown to the IDS. Consequently, the ability to detect unknown attacks becomes a relevant distinguishing factor for an IDS. This study aims to evaluate and quantify such ability by exercising multiple ML algorithms for IDSs. We apply 47 supervised, unsupervised, deep learning, and meta-learning algorithms in an experimental campaign embracing 11 attack datasets, and with a methodology that simulates the occurrence of unknown attacks. Detecting unknown attacks is not trivial: however, we show how unsupervised meta-learning algorithms have better detection capabilities of unknowns and may even outperform classification performance of other ML algorithms when dealing with unknown attacks. Tommaso Zoppi, Andrea Ceccarelli, Tommaso Puccetti, Andrea Bondavalli |
Comput. Secur. | 4 |
| 2023 | Safe Maintenance of Railways using COTS Mobile Devices: The Remote Worker DashboardabstractThe railway domain is regulated by rigorous safety standards to ensure that specific safety goals are met. Often, safety-critical systems rely on custom hardware-software components that are built from scratch to achieve specific functional and non-functional requirements. Instead, the (partial) usage of Commercial Off-The-Shelf (COTS) components is very attractive as it potentially allows reducing cost and time to market. Unfortunately, COTS components do not individually offer enough guarantees in terms of safety and security to be used in critical systems as they are. In such a context, RFI (Rete Ferroviaria Italiana), a major player in Europe for railway infrastructure management, aims at equipping track-side workers with COTS devices to remotely and safely interact with the existing interlocking system, drastically improving the performance of maintenance operations. This paper describes the first effort to update existing (embedded) railway systems to a more recent cyber-physical system paradigm. Our Remote Worker Dashboard (RWD) pairs the existing safe interlocking machinery alongside COTS mobile components, making cyber and physical components cooperate to provide the user with responsive, safe, and secure service. Specifically, the RWD is a SIL4 cyber-physical system to support maintenance of actuators and railways in which COTS mobile devices are safely used by track-side workers. The concept, development, implementation, verification, and validation activities to build the RWD were carried out in compliance with the applicable CENELEC standards required by certification bodies to declare compliance with specific guidelines. Tommaso Zoppi, Innocenzo Mungiello, Andrea Ceccarelli, Alberto Cirillo, Lorenzo Sarti, Lorenzo Esposito, Giuseppe Scaglione, Sergio Repetto, Andrea Bondavalli |
ACM Trans. Cyber Phys. Syst. | 9 |
| 2022 | Failure modes and failure mitigation in GPGPUs: a reference model and its applicationabstractGeneral Purpose GPUs (GPGPUs) are highly susceptible to both transient and permanent faults. This is a serious concern for their safe and reliable usage in many domains, from autonomous driving to High Performance Computing. The research and industrial community responded fiercely to this issue, by analyzing failures impact and devising failure mitigation strategies. This led to the definition of several failure modes and mitigation approaches. Unfortunately, these are often based on different foundations, and it is not easy to position them in a consistent view. This work elaborates a GPGPU failures model, identifying relations between the GPGPU failure modes and components, and then it analyzes mitigations proposed in the literature. By proposing a unified view on failures and mitigations, the resulting model i) positions each research on the subject, ii) easily identifies the current gaps, and iii) sets the basis for further research on GPGPU failures. Francesco Terrosi, Andrea Ceccarelli, Andrea Bondavalli |
COMPSAC | 3 |
| 2022 | Impact of Machine Learning on Safety Monitors
Francesco Terrosi, Lorenzo Strigini, Andrea Bondavalli |
SAFECOMP | 3 |
| 2022 | A cyber-physical-social approach for engineering Functional Safety Requirements for automotive systems
Mohamad Gharib, Andrea Ceccarelli, Paolo Lollini, Andrea Bondavalli |
J. Syst. Softw. | 4 |
| 2022 | Stochastic Activity Networks Templates: Supporting Variability in Performability ModelsabstractModel-based evaluation is extensively used to estimate the performance and reliability of dependable systems. Traditionally, these systems were small and self-contained, and the main challenge for model-based evaluation has been the efficiency of the solution process. Recently, the problem of specifying and maintaining complex models has increasingly gained attention, as modern systems are characterized by many components and complex interactions. Components share similarities, but at the same time, also exhibit variations in their behavior due to different configurations or roles in the system. From the modeling perspective, variations lead to replicating and altering a small set of base models multiple times. Variability is taken into account only informally, by defining a sample model and explaining its possible variations. In this article, we address the problem of including variability in performability models, focusing on stochastic activity networks (SANs). We introduce the formal definition of stochastic activity networks templates (SAN-T), a formalism based on SANs with the addition of variability aspects. Differently from other approaches, parameters can also affect the structure of the model, like the number of cases of activities. We apply the SAN-T formalism to the modeling of the backbone network of an environmental monitoring infrastructure. In particular, we show how existing SAN models from the literature can be generalized using the newly introduced formalism. Leonardo Montecchi, Paolo Lollini, Andrea Bondavalli |
IEEE Trans. Reliab. | 3 |
| 2021 | Detecting Intrusions by Voting Diverse Machine Learners: Is It Really Worth?abstractRecent years have seen an astounding growth in the adoption of Machine Learning algorithms to classify data gathered through monitoring activities. Those algorithms can effectively classify data as system indicators, network packets, and logs according to a model they infer during training. This way, they provide sophisticated means to conduct intrusion detection by suspecting anomalies due to attacks in the value of those features. Additionally, Meta-Learners as Bagging and Boosting build ensembles of homogeneous classifiers that are known to improve classification performance with positive impact on intrusion detection. On the other hand, it is not yet clear if ensembles of heterogeneous or diverse classifiers can build better intrusion detectors. To such extent, we first recap on n-version programming, k-out-of-m (k-o-o-m) systems and the role of diversity. Then, we present k-o-o-m systems of classifiers for intrusion detection, expanding on meta-learning and diversity measures to be applied to classifiers. This paves the way for an experimental campaign which exercises supervised and unsupervised classifiers as well as k-o-o-m voting ensembles. After presenting and discussing results, we conclude that voting ensembles of diverse classifiers does not improve intrusion detection. Therefore, while voting has been acknowledged since decades as a staple to manage n-version programming for reliable systems engineering, it is not as effective as a meta-learner to improve classification performance of intrusion detectors. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
PRDC | 3 |
| 2021 | Meta-Learning to Improve Unsupervised Intrusion Detection in Cyber-Physical SystemsabstractArtificial Intelligence (AI)- based classifiers rely on Machine Learning (ML) algorithms to provide functionalities that system architects are often willing to integrate into critical Cyber-Physical Systems (CPSs) . However, such algorithms may misclassify observations, with potential detrimental effects on the system itself or on the health of people and of the environment. In addition, CPSs may be subject to threats that were not previously known, motivating the need for building Intrusion Detectors (IDs) that can effectively deal with zero-day attacks. Different studies were directed to compare misclassifications of various algorithms to identify the most suitable one for a given system. Unfortunately, even the most suitable algorithm may still show an unsatisfactory number of misclassifications when system requirements are strict. A possible solution may rely on the adoption of meta-learners, which build ensembles of base-learners to reduce misclassifications and that are widely used for supervised learning. Meta-learners have the potential to reduce misclassifications with respect to non-meta learners: however, misleading base-learners may let the meta-learner leaning towards misclassifications and therefore their behavior needs to be carefully assessed through empirical evaluation. To such extent, in this paper we investigate, expand, empirically evaluate, and discuss meta-learning approaches that rely on ensembles of unsupervised algorithms to detect (zero-day) intrusions in CPSs. Our experimental comparison is conducted by means of public datasets belonging to network intrusion detection and biometric authentication systems, which are common IDSs for CPSs. Overall, we selected 21 datasets, 15 unsupervised algorithms and 9 different meta-learning approaches. Results allow discussing the applicability and suitability of meta-learning for unsupervised anomaly detection, comparing metric scores achieved by base algorithms and meta-learners. Analyses and discussion end up showing how the adoption of meta-learners significantly reduces misclassifications when detecting (zero-day) intrusions in CPSs. Tommaso Zoppi, Mohamad Gharib, Muhammad Atif 0001, Andrea Bondavalli |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2021 | MADneSs: A Multi-Layer Anomaly Detection Framework for Complex Dynamic SystemsabstractAnomaly detection can infer the presence of errors without observing the target services, but detecting variations in the observable parts of the system on which the services reside. This is a promising technique in complex software-intensive systems, because either instrumenting the services' internals is exceedingly time-consuming, or encapsulation makes them not accessible. Unfortunately, in such systems anomaly detection is often ineffective due to their dynamicity, which implies changes in the services or their expected workload. Here we present our approach to enhance the efficacy of anomaly detection in complex, dynamic software-intensive systems. After discussing the related challenges, we present MADneSs, an anomaly detection framework tailored for the above systems that includes an adaptive multi-layer monitoring module. Monitored data are then processed by the anomaly detector, which adapts its parameters depending on the current system behavior. An anomaly alert is provided if the analysis conducted by the anomaly detector identify unexpected trends in the data. MADneSs is evaluated through an experimental campaign on two service-oriented architectures; software faults are injected in the application layer, and detected through monitoring of underlying system layers. Lastly, we quantitatively and qualitatively discuss our results with respect to state-of-the-art solutions, highlighting the key contributions of MADneSs. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2020 | Agility of Security Practices and Agile Process Models: An Evaluation of Cost for Incorporating Security in Agile Process Models
H. Maria Maqsood, Andrea Bondavalli |
ENASE | 2 |
| 2020 | On the educated selection of unsupervised algorithms via attacks and anomaly classesabstractAnomaly detection aims at finding patterns in data that do not conform to the expected behavior. It is largely adopted in intrusion detection systems, relying on unsupervised algorithms that have the potential to detect zero-day attacks; however, efficacy of algorithms varies depending on the observed system and the attacks. Selecting the algorithm that maximizes detection capability is a challenging task with no master key. This paper tackles the challenge above by devising and applying a methodology to identify relations between attack families, anomaly classes and algorithms. The implication is that an unknown attack belonging to a specific attack family is most likely to get observed by unsupervised algorithms that are particularly effective on such attack family. This paves the way to rules for the selection of algorithms based on the identification of attack families. The paper proposes and applies a methodology based on analytical and experimental investigations supported by a tool to i) identify which anomaly classes are most likely raised by the different attack families, ii) study suitability of anomaly detection algorithms to detect anomaly classes, iii) combine previous results to relate anomaly detection algorithms and attack families, and iv) define guidelines to select unsupervised algorithms for intrusion detection. Tommaso Zoppi, Andrea Ceccarelli, Lorenzo Salani, Andrea Bondavalli |
J. Inf. Secur. Appl. | 4 |
| 2020 | A Template-Based Methodology for the Specification and Automated Composition of Performability ModelsabstractDependability and performance analysis of modern systems is facing great challenges: their scale is growing, they are becoming massively distributed, interconnected, and evolving. Such complexity makes model-based assessment a difficult and time-consuming task. For the evaluation of large systems, reusable submodels are typically adopted as an effective way to address the complexity and to improve the maintainability of models. When using state-based models, a common approach is to define libraries of generic submodels, and then compose concrete instances by state sharing, following predefined “patterns” that depend on the class of systems being modeled. However, such composition patterns are rarely formalized, or not even documented at all. In this paper, we address this problem using a model-driven approach, which combines a language to specify reusable submodels and composition patterns, and an automated composition algorithm. Clearly defining libraries of reusable submodels, together with patterns for their composition, allows complex models to be automatically assembled, based on a high-level description of the scenario to be evaluated. This paper provides a solution to this problem focusing on: formally defining the concept of model templates, defining a specification language for model templates, defining an automated instantiation and composition algorithm, and applying the approach to a case study of a large-scale distributed system. Leonardo Montecchi, Paolo Lollini, Andrea Bondavalli |
IEEE Trans. Reliab. | 3 |
| 2019 | Evaluation of Anomaly Detection Algorithms Made Easy with RELOADabstractAnomaly detection aims at identifying patterns in data that do not conform to the expected behavior. Despite anomaly detection has been arising as one of the most powerful techniques to suspect attacks or failures, dedicated support for the experimental evaluation is actually scarce. In fact, existing frameworks are mostly intended for the broad purposes of data mining and machine learning. Intuitive tools tailored for evaluating anomaly detection algorithms for failure and attack detection with an intuitive support to sliding windows are currently missing. This paper presents RELOAD, a flexible and intuitive tool for the Rapid EvaLuation Of Anomaly Detection algorithms. RELOAD is able to automatically i) fetch data from an existing data set, ii) identify the most informative features of the data set, iii) run anomaly detection algorithms, including those based on sliding windows, iv) apply multiple strategies to features and decide on anomalies, and v) provide conclusive results following an extensive set of metrics, along with plots of algorithms scores. Finally, RELOAD includes a simple GUI to set up the experiments and examine results. After describing the structure of the tool and detailing inputs and outputs of RELOAD, we exercise RELOAD to analyze an intrusion detection dataset available on a public platform, showing its setup, metric scores and plots. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
ISSRE | 3 |
| 2019 | An Initial Investigation on Sliding Windows for Anomaly-Based Intrusion DetectionabstractThe growing systems complexity calls for dedicated monitoring and data analysis strategies aiming to detect faults, attacks and errors before they escalate into failures. Distributed and heterogeneous systems are more likely to expose vulnerabilities that attackers may target to get unauthorized access to a system, make it unavailable or steal sensitive data. As countermeasure, traditionally techniques for attacks and intrusion detection are based on signature recognition and requires knowledge on the attacks pattern: therefore, they are not well-suited to detect zero-days attacks. A viable alternative is anomaly detection, where deviation from the expected behavior are suspected as attacks. However, anomaly detection is generally not applicable in systems where the expected behavior changes through time. In this paper we explore anomaly detection strategies based on sliding windows, which are intended for evolving and dynamic systems as IoT, in which system configuration and behavior may change continuously. We first describe the context and the key features of sliding windows, and then we proceed detailing their possible drawbacks. Discussion is substantiated by quantitative analyses directed to evaluate detection capabilities. The experimental campaign is based on state-of-the-art algorithms and datasets, and results have been made publicly available. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
SERVICES | 3 |
| 2019 | Threat Analysis in Systems-of-Systems: An Emergence-Oriented ApproachabstractCyber-physical Systems of Systems (SoSs) are large-scale systems made of independent and autonomous cyber-physical Constituent Systems (CSs) which may interoperate to achieve high-level goals also with the intervention of humans. Providing security in such SoSs means, among other features, forecasting and anticipating evolving SoS functionalities, ultimately identifying possible detrimental phenomena that may result from the interactions of CSs and humans. Such phenomena, usually called emergent phenomena , are often complex and difficult to capture: the first appearance of an emergent phenomenon in a cyber-physical SoS is often a surprise to the observers. Adequate support to understand emergent phenomena will assist in reducing both the likelihood of design or operational flaws, and the time needed to analyze the relations amongst the CSs, which always has a key economic significance. This article presents a threat analysis methodology and a supporting tool aimed at (i) identifying (emerging) threats in evolving SoSs, (ii) reducing the cognitive load required to understand an SoS and the relations among CSs, and (iii) facilitating SoS risk management by proposing mitigation strategies for SoS administrators. The proposed methodology, as well as the tool, is empirically validated on Smart Grid case studies by submitting questionnaires to a user base composed of 3 stakeholders and 18 BSc and MSc students. Andrea Ceccarelli, Tommaso Zoppi, Alexandr Vasenev, Marco Mori, Dan Ionita, Lorena Montoya, Andrea Bondavalli |
ACM Trans. Cyber Phys. Syst. | 7 |
| 2018 | On Algorithms Selection for Unsupervised Anomaly DetectionabstractAnomaly detection, which aims at identifying unexpected trends and data patterns, has widely been used to build error detectors, failure predictors or intrusion detectors. Internal faults or malicious attacks have a different impact on the behavior of the system. They usually manifest as different observable deviations from the expected behavior, which may be identified by anomaly detection algorithms. Our study aims at investigating the suitability of unsupervised algorithms and their families in detecting either point, contextual or collective anomalies. To provide a complete picture, we consider both sliding and non-sliding window algorithms which operate in unsupervised mode. Along with qualitative analyses of each algorithm and family, we conduct an experimental campaign in which we run each algorithm on three state-of-the-art datasets in which we inject either point, contextual or collective anomalies. Results show that non-sliding algorithms are capable to detect point and collective anomalies, while they cannot effectively deal with contextual ones. Instead, sliding window algorithms require shorter periods of training and naturally build a local context, which allow them to effectively deal with contextual anomalies. Such observations are summarized to support the choice of the correct algorithm depending on the investigated class(es) of anomaly. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
PRDC | 3 |
| 2018 | A Requirements-Driven Methodology for the Proper Selection and Configuration of BlockchainsabstractIn recent years, the interest in blockchain has grown exponentially, and nowadays it is foreseen as a technology with the potential to revolutionize the way data is maintained and transferred around the globe. The reason of this excitement is ascribable to the ability of enabling new forms of transactions and interactions between mistrusting and decentralized entities. Indeed, it has attracted interests and huge investments from enterprises, and it is predictable that in a near future many industries will adopt it. However, it is not a panacea and in some cases may even become useless or not convenient. Moreover, even when it can really constitute an added value, selecting the proper blockchain and configuring it may not be trivial. Trying to go beyond the hype and to address this problem, this paper proposes a methodology addressing: i) whether, given a specific problem requirements, the blockchain is a proper solution for it ii) in such a case which is the blockchain category more suitable, and finally iii) guiding the designer throughout its configuration. Mirko Staderini, Enrico Schiavone, Andrea Bondavalli |
SRDS | 3 |
| 2018 | Systems-of-systems modeling using a comprehensive viewpoint-based SysML profileabstractAbstract In recent years, more and more efforts have been devoted in supporting the design of systems‐of‐systems (SoS). Designing such systems is a multidisciplinary problem which involves considering emergent phenomena, assuring the achievement of dependability/security requirements, guaranteeing system responsiveness, and supporting dynamicity/evolution and multicriticality of provided services. A first step towards a viable design approach is to provide a conceptual model of SoS which captures SoS concepts, and their interrelationships aiming at enhancing the understandability of SoS to stakeholders and providing the basis for further automated analysis. In this context, the AMADEOS European project is bringing together researchers and practitioners to provide the support to design SoS starting from the definition of a domain specific ontology serving as a vocabulary for SoS. Our contribution consists in the modeling of the key SoS concepts and relationships defined in AMADEOS adopting a systems modeling language visual modeling language. We propose a systems modeling language profile for SoS, and we show its applicability in a Smart Grid scenario. We show how to use the profile in a model‐driven engineering process to support different types of analyses, and we discuss how to integrate the profile in a user‐friendly model‐driven engineering tool for SoS rapid modeling, validation, code‐generation, and simulation. Marco Mori, Andrea Ceccarelli, Paolo Lollini, Bernhard Frömel, Francesco Brancati, Andrea Bondavalli |
J. Softw. Evol. Process. | 6 |
| 2018 | Labelling relevant events to support the crisis management operatorabstractAbstract Thanks to the large availability of portable devices and the growing interest in the Internet of Things, during crises, social networks, or alerts sent through mobile devices or sensor networks are available and can be matched each other to perform situational analysis. However, the inclusion of multiple heterogeneous sources in situational analyses leads to 2 main issues: (1) a source could deliver (voluntarily or erroneously) wrong data damaging the integrity and the correctness of the analysis, and (2) a significant amount of heterogeneous data need to be processed. As a consequence, the crisis management operator faces a large amount of potentially unreliable data. In this paper, we present a relevance labelling strategy to process information gathered from heterogeneous data streams to select the most relevant events. These are presented to the crisis management operator with the highest priority. Our strategy is evaluated using events collected by the Secure! crisis management system, considering 3 real crisis scenarios happened in Italy in 2015. Results show that our strategy is able to correctly identify sets of relevant events, supporting the activities of the crisis management operator. Tommaso Zoppi, Andrea Ceccarelli, Francesco Lo Piccolo, Paolo Lollini, Gabriele Giunta, Vito Morreale, Andrea Bondavalli |
J. Softw. Evol. Process. | 7 |
| 2018 | PrivAPP: An integrated approach for the design of privacy-aware applicationsabstractSummary Nowadays, personal information is collected, stored, and managed through web applications and services. Companies are interested in keeping such information private due to regulation laws and privacy concerns of customers. Furthermore, the reputation of a company can be dependent on privacy protection, ie, the more a company protects the privacy of its customers, the more credibility it gets. This paper proposes an integrated approach that relies on models and design tools to help in the analysis, design, and development of web applications and services with privacy concerns. Using the approach, these applications can be developed consistently with their privacy policies to enforce them, protecting personal information from different sources of privacy violation. The approach is composed of a conceptual model, a reference architecture, and a Unified Modified Language Profile, ie, an extension of the Unified Modified Language for including privacy protection. The idea is to systematize the privacy concepts in the scope of web applications and services, organizing the privacy domain knowledge and providing features and functionalities that must be addressed to protect the privacy of the users in the design and development of web applications. Validation has been performed by analyzing the ability of the approach to model privacy policies from real web applications and by applying it to a simple application example of an online bookstore. Results show that privacy protection can be implemented in a model‐based approach, bringing values for the stakeholders and being an important contribution toward improving the process of designing web applications in the privacy domain. Tânia Basso, Leonardo Montecchi, Regina Lúcia de Oliveira Moraes, Mário Jino, Andrea Bondavalli |
Softw. Pract. Exp. | 5 |
| 2017 | Continuous Biometric Verification for Non-Repudiation of Remote ServicesabstractAs our society massively relies on ICT, security services are becoming essential to protect users and entities involved. Amongst such services, non-repudiation provides evidences of actions, protects against their denial, and helps solving disputes between parties. For example, it prevents denial of past behaviors as having sent or received messages. Noteworthy, if the information flow is continuous, evidences should be produced for the entirety of the flow and not only at specific points. Further, non-repudiation should be guaranteed by mechanisms that do not reduce the usability of the system or application. To meet these challenges, in this paper, we propose two solutions for non-repudiation of remote services based on multi-biometric continuous authentication. We present an application scenario that discusses how users and service providers are protected with such solutions. We also discuss the technological readiness of biometrics for non-repudiation services: the outcome is that, under specific assumptions, it is actually ready. Enrico Schiavone, Andrea Ceccarelli, Andrea Bondavalli |
ARES | 3 |
| 2017 | Dealing with Functional Safety Requirements for Automotive Systems: A Cyber-Physical-Social Approach
Mohamad Gharib, Paolo Lollini, Andrea Ceccarelli, Andrea Bondavalli |
CRITIS | 4 |
| 2017 | Identification of critical situations via Event Processing and Event Trust Analysis
Massimiliano Leone Itria, Melinda Kocsis-Magyar, Andrea Ceccarelli, Paolo Lollini, Gabriele Giunta, Andrea Bondavalli |
Knowl. Inf. Syst. | 6 |
| 2016 | A Model-Based Approach to Support Safety-Related Decisions in the Petroleum DomainabstractAccidents on petroleum installations can have huge consequences, to mitigate the risk, a number of safety barriers are devised. Faults and unexpected events may cause barriers to temporarily deviate from their nominal state. For safety reasons, a work permit process is in place: decision makers accept or reject work permits based on the current state of barriers. However, this is difficult to estimate, as it depends on a multitude of physical, technical and human factors. Information obtained from different sources needs to be aggregated by humans, typically within a limited amount of time. In this paper we propose an approach to provide an automated decision support to the work permit system, which consists in the evaluation of quantitative measures of the risk associated with the execution of work. The approach relies on state-based stochastic models, which can be automatically composed based on the work permit to be examined. Leonardo Montecchi, Atle Refsdal, Paolo Lollini, Andrea Bondavalli |
DSN | 4 |
| 2016 | On the Dependability for Dynamic Software Product Lines: A Comparative Systematic Mapping StudyabstractSoftware Product Lines (SPLs) are techniques where several artefacts are reused (domain), and some are customised (variation points). An SPL can bind variation points statically (compilation time) or dynamically (runtime). Dynamic Software Product Lines (DSPLs) use dynamic binding to adapt to the environment or requirements changes. DSPLs are commonly used to build dependable systems, defined as systems with the ability to avoid more frequent or severe service failures than the acceptable. The main dependability attributes are availability, confidentiality, integrity, reliability, maintainability, and safety. To better understand this context, a Systematic Mapping Study (SMS) was applied searching proposals that include dependability attributes in DSPLs. Our results suggest that few solutions handle dependability in DSPL context. We selected nine primary studies in this regard. We performed a comparative study of the results, analysing other dimensions, and facets, aiming for a better understanding of this research area. Jane Dirce A. Sandim Eleuterio, Felipe Nunes Gaia, Andrea Bondavalli, Paolo Lollini, Genaína Nunes Rodrigues, Cecília M. F. Rubira |
SEAA | 3 |
| 2016 | Context-Awareness to Improve Anomaly Detection in Dynamic Service Oriented Architectures
Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
SAFECOMP | 3 |
| 2016 | Continuous Authentication and Non-repudiation for the Security of Critical SystemsabstractUser authentication is a key service, especially for systems that can be considered critical for the data stored and the functionalities offered. In those cases, traditional authentication mechanisms can be inadequate to face intrusions: they usually verify user's identity only at login, and even repeating this step, frequently asking for passwords or PIN would reduce system's usability. Biometric continuous authentication, instead, is emerging as viable alternative approach that can guarantee accurate and transparent verification for the entire session: the traits can be repeatedly acquired avoiding disturbing the user's activity. Another security service that these systems may need is nonrepudiation, which protect against the denial of having used the system or executed some commands with it. The paper focuses on biometric continuous authentication and nonrepudiation, and it briefly presents a preliminary solution based on a specific case study. This work presents the current research direction of the author and describes some challenges that the student aims to address in the next years. Enrico Schiavone, Andrea Ceccarelli, Andrea Bondavalli |
SRDS | 3 |
| 2016 | Challenging Anomaly Detection in Complex Dynamic SystemsabstractSoftware infrastructures are becoming more and more complex, making performance and dependability monitoring in wide and dynamic contexts such as Distributed Systems, Systems of Systems (SoS) and Cloud environments an unachievable goal. Consequently, it is very difficult to know how all the specific parts, services and modules of these systems behave. This negatively impacts our ability in detecting anomalies, because the boundaries between normal and anomalous behaviors are not always known. The paper describes the context and the targeted problem highlighting the research directions that the student will follow in the next years. In particular, after introducing the relevance of this work with respect to the academic and the industrial state of the art, we carefully define the problem and summarize the main challenges that arise according to such problem definition. Tommaso Zoppi, Andrea Ceccarelli, Andrea Bondavalli |
SRDS | 3 |
| 2016 | Modeling QoE in Dependable Tele-Immersive Applications: A Case Study of World OperaabstractWith the advent of recent technological advances, more demanding tele-immersive applications have started to emerge. In the World Opera application, artists from different opera houses across the globe can participate in a single united performance, and interact almost as if they were co-located. One of the main design challenges in this application domain is to assess to what extent the inevitable failures of some of the numerous and complex hardware, software, and network components affect the quality of experience for the user. This challenge cannot be addressed by traditional system-centric methods for dependability evaluation, which do not take personalized user perspective into account when considering meaningful and acceptable degradation of services. In this paper, we propose a novel method to assess the quality of experience in presence of failures, based on a new metric called perceived reliability. The method takes the human perspective into account and allows considering factors such as human perception of video and audio, characteristics of the audience, as well as performance elements and artistic content. This method can help system designers and engineers compare architectural variants and determine the dependability budget. We show the feasibility of our method by applying it to a World Opera performance. To this end, we construct a SAN-based model and run simulations in the Möbius framework. The obtained results provide useful guidelines for system engineers towards improving the quality of experience of World Opera performances despite the presence of failures. Narasimha Raghavan, Leonardo Montecchi, Nicola Nostro, Roman Vitenberg, Hein Meling, Andrea Bondavalli |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2015 | A Multi-layer Anomaly Detector for Dynamic Service-Based Systems
Andrea Ceccarelli, Tommaso Zoppi, Massimiliano Leone Itria, Andrea Bondavalli |
SAFECOMP | 4 |
| 2015 | An OS-level Framework for Anomaly Detection in Complex Software SystemsabstractRevealing anomalies at the operating system (OS) level to support online diagnosis activities of complex software systems is a promising approach when traditional detection mechanisms (e.g., based on event logs, probes and heartbeats) are inadequate or cannot be applied. In this paper we propose a configurable detection framework to reveal anomalies in the OS behavior, related to system misbehaviors. The detector is based on online statistical analyses techniques, and it is designed for systems that operate under variable and non-stationary conditions. The framework is evaluated to detect the activation of software faults in a complex distributed system for Air Traffic Management (ATM). Results of experiments with two different OSs, namely Linux Red Hat EL5 and Windows Server 2008, show that the detector is effective for mission-critical systems. The framework can be configured to select the monitored indicators so as to tune the level of intrusivity. A sensitivity analysis of the detector parameters is carried out to show their impact on the performance and to give to practitioners guidelines for its field tuning. Antonio Bovenzi, Francesco Brancati, Stefano Russo 0001, Andrea Bondavalli |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2015 | Continuous and Transparent User Identity Verification for Secure Internet ServicesabstractSession management in distributed Internet services is traditionally based on username and password, explicit logouts and mechanisms of user session expiration using classic timeouts. Emerging biometric solutions allow substituting username and password with biometric data during session establishment, but in such an approach still a single verification is deemed sufficient, and the identity of a user is considered immutable during the entire session. Additionally, the length of the session timeout may impact on the usability of the service and consequent client satisfaction. This paper explores promising alternatives offered by applying biometrics in the management of sessions. A secure protocol is defined for perpetual authentication through continuous user verification. The protocol determines adaptive timeouts based on the quality, frequency and type of biometric data transparently acquired from the user. The functional behavior of the protocol is illustrated through Matlab simulations, while model-based quantitative analysis is carried out to assess the ability of the protocol to contrast security attacks exercised by different kinds of attackers. Finally, the current prototype for PCs and Android smartphones is discussed. Andrea Ceccarelli, Leonardo Montecchi, Francesco Brancati, Paolo Lollini, Angelo Marguglio, Andrea Bondavalli |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2014 | A Testbed for Evaluating Anomaly Detection Monitors through Fault InjectionabstractAmongst the features of Service Oriented Architectures (SOAs), their flexibility, dynamicity, and scalability make them particularly attractive for adoption in the ICT infrastructure of organizations. Such features come at the cost of improved difficulty in monitoring the SOA for error detection: i) faults may manifest themselves differently due to services and SOA evolution, and ii) interactions between a service and its monitors may need reconfiguration at each service update. This calls for monitoring solutions that operate at different layers than the application layer (services layer). In this paper we present our ongoing work towards the definition of a monitoring framework for SOAs and services, which relies on anomaly detection performed at the Application Server (AS) and the Operating System (OS) layers to identify events whose manifestation or effect is not adequately described a-priori. Specifically the paper introduces the key concepts of our work and presents the case study built to exercise and set-up our monitor. The case study uses Life ray as application layer and it includes fault injection and data collection instruments to perform extended testing campaigns. Andrea Ceccarelli, Tommaso Zoppi, Andrea Bondavalli, Fabio Duchi, Giuseppe Vella |
ISORC | 3 |
| 2013 | Meeting the challenges in the design and evaluation of a trackside real-time safety-critical systemabstractHighly distributed, autonomous and self-powered systems operating in harsh, outdoors environments face several threats in terms of dependability, timeliness and security, due to the challenging operating conditions determined by the environment. Despite such difficulties, there is an increasing demand to deploy these systems to support critical services, thus calling for severe timeliness, safety, and security requirements. Several challenges need to be faced and overcome. First, the designed architecture must be able to cope with the environmental challenges and satisfy dependability, timeliness and security requirements. Second, the assessment of the system must be carried on despite potentially incomplete field-data, and complex cascading effects that small modifications in system properties and operating conditions may have on the targeted metrics. In this paper we present our experience from the EU-funded project ALARP (A railway automatic track warning system based on distributed personal mobile terminals), which aims to build and validate a distributed, real-time, safety-critical system that detects trains approaching a railway worksite and notifies their arrivals to railway trackside workers. The paper describes the challenges we faced, and the solutions we adopted, when architecting and evaluating the ALARP system. Leonardo Montecchi, Andrea Ceccarelli, Paolo Lollini, Andrea Bondavalli |
ISORC | 4 |
| 2012 | Model-based analysis of a protocol for reliable communication in railway worksitesabstractIn this paper we perform a model-based analysis of the Timed Reliable Communication (TRC) protocol, which is being used within the EU funded ALARP project for railway worksite com-munication. TRC is a group communication protocol based on IEEE 802.11 networks, targeting safety-critical applications with limited bandwidth requirements. The paper contains an in-depth analysis of the performance and reliability characteristics of the protocol using a Stochastic Activity Networks model. The results are first compared with available experimental measurements for the sake of model validation. The validated model is then used for a thorough analysis of a set of key metrics under different envi-ronment and network conditions. The obtained results allow: i) to assess that the protocol allows to satisfy the ALARP targeted performance and reliability requirements, and ii) to evaluate the existing tradeoffs and help in choosing parameter values for the final implementation. Leonardo Montecchi, Paolo Lollini, Boris Malinowsky, Jesper Grønbæk, Andrea Bondavalli |
MSWiM | 5 |
| 2012 | Improving Security of Internet Services through Continuous and Transparent User Identity VerificationabstractSession management in distributed Internet services is traditionally based on username and password, and explicit logouts and timeouts that expire due to idle activity of the user. Emerging biometric solutions allow substituting username and password with biometric data, but still a single verification is deemed sufficient, and the identity of a user is considered immutable during the entire session. Additionally, the length of the timeout may impact on the usability of the service and consequent client satisfaction. This paper explores promising alternatives offered by biometrics for the management of sessions. A secure protocol is defined for perpetual authentication through continuous user verification. The protocol determines adaptive timeouts selected on the basis of the quality, frequency and type of biometric data acquired transparently from the user. Protocol behavior is shown through simulations. Andrea Ceccarelli, Andrea Bondavalli, Francesco Brancati, Ernesto La Mattina |
SRDS | 2 |
| 2012 | Adaptare: Supporting automatic and dependable adaptation in dynamic environmentsabstractDistributed protocols executing in uncertain environments, like the Internet or ambient computing systems, should dynamically adapt to environment changes in order to preserve Quality of Service (QoS). In earlier work, it was shown that QoS adaptation should be dependable, if correctness of protocol properties is to be maintained. More recently, some ideas concerning specific strategies and methodologies for improving QoS adaptation have been proposed. In this article we describe Adaptare , a complete framework for dependable QoS adaptation. We assume that during its lifetime, a system alternates periods where its temporal behavior is well characterized, with transition periods during which a variation of the environment conditions occurs. Our method is based on the following: if the environment is generically characterized in analytical terms, and we can detect the alternation of these stable and transient phases, we can improve the effectiveness and dependability of QoS adaptation. To prove our point we provide detailed evaluation results of the proposed solutions. Our evaluation is based on synthetic data flows generated from probabilistic distributions, as well as on real data traces collected in various Internet-based environments. We compare our solution with other approaches and we show that Adaptare, albeit more complex, is very effective, allowing protocols to adapt to the available resources in a dependable way. Monica Dixit, António Casimiro, Paolo Lollini, Andrea Bondavalli, Paulo Veríssimo |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2011 | Towards a MDE Transformation Workflow for Dependability AnalysisabstractIn the last ten years, Model Driven Engineering (MDE) approaches have been extensively used for the analysis of extra-functional properties of complex systems, like safety, dependability, security, predictability, quality of service. To this purpose, engineering languages (like UML and AADL) have been extended with additional features to model the required non-functional attributes, and transformations have been used to automatically generate the analysis models to be solved by appropriate analysis tools. In most of the available works, however, the transformations are not inte grated into a more general development process, aimed to support both domain-specific design analysis and verification of extra-functional properties. In this paper we explore this research direction presenting a transformation work flow for dependability analysis that is part of an industrial-quality infrastructure for the specification, analysis and verification of extra-functional properties, currently under development within the ARTEMIS-JU CHESS project. Specifically, the paper provides the following major contributions: i) definition of the required transformation steps to automatically assess the system dependability properties starting from the CHESS Modeling Language, ii) definition of a new Intermediate Dependability Model (IDM) acting as a bridge between the CHESS Modeling Language and the low-level analysis models, iii) definition of transformations from the CHESS Modeling Language to IDM models. Leonardo Montecchi, Paolo Lollini, Andrea Bondavalli |
ICECCS | 3 |
| 2011 | A Statistical Anomaly-Based Algorithm for On-line Fault Detection in Complex Software Critical Systems
Antonio Bovenzi, Francesco Brancati, Stefano Russo 0001, Andrea Bondavalli |
SAFECOMP | 4 |
| 2011 | The HIDENETS Holistic Approach for the Analysis of Large Critical Mobile SystemsabstractDealing with large, critical mobile systems and infrastructures where ongoing changes and resilience are paramount leads to very complex and difficult challenges for system evaluation. These challenges call for approaches that are able to integrate several evaluation methods for the quantitative assessment of QoS indicators which have been applied so far only to a limited extent. In this paper, we propose the holistic evaluation framework developed during the recently concluded FP6-HIDENETS project. It is based on abstraction and decomposition, and it exploits the interactions among different evaluation techniques including analytical, simulative, and experimental measurement approaches, to manage system complexity. The feasibility of the holistic approach for the analysis of a complete end-to-end scenario is first illustrated presenting two examples where mobility simulation is used in combination with stochastic analytical modeling, and then through the development and implementation of an evaluation workflow integrating several tools and model transformation steps. Andrea Bondavalli, Ossama Hamouda, Mohamed Kaâniche, Paolo Lollini, István Majzik, Hans-Peter Schwefel |
IEEE Trans. Mob. Comput. | 1 |
| 2010 | Improving Robustness of Network Fault Diagnosis to Uncertainty in ObservationsabstractPerforming decentralized network fault diagnosis based on network traffic is challenging. Besides inherent stochastic behaviour of observations, measurements may be subject to errors degrading diagnosis timeliness and accuracy. In this paper we present a novel approach in which we aim to mitigate issues of measurement errors by quantifying uncertainty. The uncertainty information is applied in the diagnostic component to improve its robustness. Three diagnosis components have been proposed based on the Hidden Markov Model formalism: (H0) representing a classical approach, (H1) a static compensation of (H0) to uncertainties and (H2) dynamically adapting diagnosis to uncertainty information. From uncertainty injection scenarios of added measurement noise we demonstrate how using uncertainty information can provide a structured approach of improving diagnosis. Jesper Grønbæk, Hans-Peter Schwefel, Andrea Ceccarelli, Andrea Bondavalli |
NCA | 4 |
| 2010 | Experimental Validation of a Synchronization Uncertainty-Aware Software ClockabstractA software clock capable of self-evaluating its synchronization uncertainty is experimentally validated for a specific implementation on a node synchronized through NTP. The validation methodology takes advantage of an external node equipped with a GPS-synchronized clock acting as a reference, which is connected to the node hosting the system under test through a fast Ethernet connection. Experiments are carried out for different values of the software clock parameters and different types of workload, and address the possible occurrence of faults in the system under test and in the NTP synchronization mechanism. The validation methodology is designed to be as less intrusive as possible and to grant a resolution of the order of few hundreds of microseconds. The experimental results show very good performance of R&SAClock, and their analysis gives precious hints for further improvements. Andrea Bondavalli, Francesco Brancati, Andrea Ceccarelli, Michele Vadursi |
SRDS | 1 |
| 2010 | Practical Aspects in Analyzing and Sharing the Results of Experimental EvaluationabstractDependability evaluation techniques such as the ones based on testing, or on the analysis of field data on computer faults, are a fundamental process in assessing complex and critical systems. Recently a new approach has been proposed consisting in collecting the row data produced in the experimental evaluation and store it in a multidimensional data structure. This paper reports the work in progress activities of the entire process of collecting, storing and analyzing the experimental data in order to perform a sound experimental evaluation. This is done through describing the various steps on a running example. Francesco Brancati, Andrea Bondavalli |
SRDS | 2 |
| 2009 | Trustworthy Evaluation of a Safe Driver Machine Interface through Software-Implemented Fault InjectionabstractExperimental evaluation is aimed at providing useful insights and results that constitute a confident representation of the system under evaluation. Although guidelines and good practices exist and are often applied, the uncertainty of results and the quality of the measuring system is rarely discussed. To complement such guidelines and good practices in experimental evaluation, metrology principles can contribute in improving experimental evaluation activities by assessing the measuring systems and the results achieved. In this paper we present the experimental evaluation by software-implemented fault injection of a safe train-borne driver machine interface (DMI), to evaluate its behavior in presence of faults. The measuring system built for the purpose and the results obtained on the assessment of the DMI are scrutinized along basic principles of metrology and good practices of fault injection. Trustfulness in results has been estimated satisfactory and the experimental campaign has shown that the safety mechanisms of the DMI correctly identify the faults injected and that a proper reaction is executed. Andrea Ceccarelli, Andrea Bondavalli, Danilo Iovino |
PRDC | 2 |
| 2009 | A Decomposition-Based Modeling Framework for Complex SystemsabstractStochastic model-based approaches are widely used for performability evaluation of complex software/hardware systems. Many techniques have been developed to mitigate the complexity of the associated models, but most of them are domain-specific, and they support the analysis of a limited class of systems. This paper provides a contribution in the definition of a general modeling framework that adopts three different types of decomposition techniques to deal with model complexity. Paolo Lollini, Andrea Bondavalli, Felicita Di Giandomenico |
IEEE Trans. Reliab. | 2 |
| 2008 | International Workshop on Resilience Assessment and Dependability Benchmarking (RADB 2008)abstractThis workshop summary gives a brief overview of the workshop on ldquoResilience Assessment and Dependability Benchmarkingrdquo held in conjunction with the 38th IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2008). The workshop aims at the presentation and exchange of ideas from the world wide research community and fostering discussions in order to give answers to the need for improving trustworthiness and understand the current risks inherent to computer systems and infrastructures. In particular the workshop aims at addressing key research challenges related to effective and accurate methods for measuring, assessing and benchmarking dependability and resilience. Andrea Bondavalli, István Majzik, Aad P. A. van Moorsel |
DSN | 1 |
| 2008 | Assuring Resilient Time SynchronizationabstractIn many distributed and pervasive systems the clocks of nodes are required to be synchronized to a unique global time. Due to unpredictable system and environment characteristics, the distance of a local clock from global time is a variable factor very hard to predict. Systems usually adopt measures to guarantee an upper bound on such distance from global time that are very often quite far from typical execution scenarios and thus are of practical little use. As a consequence, while in many circumstances reliable information on the actual distance from global time would improve system behaviour, unfortunately such information is usually not available. In this paper we propose the Reliable and Self-Aware Clock (R&SAClock), a low-intrusive software service that is able to compute a conservative estimation of distance from an external global time. R&SAClock acts as a new clock that couples information gained from synchronization mechanisms with information collected from the local clock to provide both current time and a self-adaptive reliable estimation of distance from global time. This paper describes the R&SAClock as a system component: we define its main functions, services and time-related mechanisms. Finally details of an implementation of the R&SAClock for the NTP synchronization mechanism and Linux OS are shown. Andrea Bondavalli, Andrea Ceccarelli, Lorenzo Falai |
SRDS | 1 |
| 2007 | Foundations of Measurement Theory Applied to the Evaluation of Dependability AttributesabstractIncreasing interest is being paid to quantitative evaluation based on measurements of dependability attributes and metrics of computer systems and infrastructures. Despite measurands are generally sensibly identified, different approaches make it difficult to compare different results. Moreover, measurement tools are seldom recognized for what they are: measuring instruments. In this paper, many measurement tools, present in the literature, are critically evaluated at the light of metrology concepts and rules. With no claim of being exhaustive, the paper (i) investigates if and how deeply such tools have been validated in accordance to measurement theory, and (ii) tries to evaluate (if possible) their measurement properties. The intention is to take advantage of knowledge available in a recognized discipline such as metrology and to propose criteria and indicators taken from such discipline to improve the quality of measurements performed in evaluation of dependability attributes. Andrea Bondavalli, Andrea Ceccarelli, Lorenzo Falai, Michele Vadursi |
DSN | 1 |
| 2007 | Towards Making NekoStat a Proper Measurement Tool for the Validation of Distributed SystemsabstractNekoStat is a Java framework and tool developed for qualitative and quantitative evaluation of dependability attributes of distributed algorithms. In this paper, NekoStat is analyzed along the lines of metrology. First the relevant metrological properties that a tool such as NekoStat should possess are introduced. The lack of a rigorous metrological characterization of the accuracy of collected measures is noticed as there is no estimation of how biased the collected data can be. To solve this, a new component, called OffsetDetector, is introduced and described. OffsetDetector allows to estimate the uncertainty of collected data and enables NekoStat to be aware of the accuracy level of the local clocks during distributed executions. The collected time measurements can thus be distinguished depending on the synchronization quality at the instant they were collected. In this way, the trustworthiness in the results is widely enhanced as shown through a case study illustrated in the paper Andrea Bondavalli, Andrea Ceccarelli, Lorenzo Falai, Michele Vadursi |
ISADS | 1 |
| 2007 | A Self-Aware Clock for Pervasive Computing SystemsabstractThe paper addresses the challenges and opportunities of instrumenting pervasive computing systems with a logical clock, aware of the quality of synchronization with respect to a time reference. Pervasive computing systems are: i) mobile; ii) dynamic; and iii) composed of a large number of distributed components; in systems with these characteristics the availability of a "smart" clock that is: i) capable to use different mechanisms for the synchronization with the global distributed time reference; and ii) aware of the current quality of synchronization with such time reference, can be very useful in order to build dependable middleware services and applications Andrea Bondavalli, Andrea Ceccarelli, Lorenzo Falai |
PDP | 1 |
| 2007 | Online Diagnosis and Recovery: On the Choice and Impact of Tuning ParametersabstractA sequenced process of Fault Detection followed by the erroneous node's Isolation and system Reconfiguration (node exclusion or recovery), that is, the FDIR process, characterizes the sustained operations of a fault-tolerant system. For distributed systems utilizing message passing, a number of diagnostic (and associated FDIR) approaches, including our prior algorithms, exist in literature and practice. Invariably, the focus is on proving the completeness and correctness (all and only the faulty nodes are isolated) for the chosen fault model, without explicitly segregating permanent from transient faulty nodes. To capture diagnostic issues related to the persistence of errors (transient, intermittent, and permanent), we advocate the integration of count-and-threshold mechanisms into the FDIR framework. Targeting pragmatic system issues, we develop an adaptive online FDIR framework that handles a continuum of fault models and diagnostic protocols and comprehensively characterizes the role of various probabilistic parameters that, due to the count-and-threshold approach, influence the correctness and completeness of diagnosis and system reliability such as the fault detection frequency. The FDIR framework has been implemented on two prototypes for automotive and aerospace applications. The tuning of the protocol parameters at design time allows a significant improvement with respect to prior design choices. Marco Serafini, Andrea Bondavalli, Neeraj Suri |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2006 | Hidden Markov Models as a Support for Diagnosis: Formalization of the Problem and Synthesis of the SolutionabstractIn modern information infrastructures, diagnosis must be able to assess the status or the extent of the damage of individual components. Traditional one-shot diagnosis is not adequate, but streams of data on component behavior need to be collected and filtered over time as done by some existing heuristics. This paper proposes instead a general framework and a formalism to model such over-time diagnosis scenarios, and to find appropriate solutions. As such, it is very beneficial to system designers to support design choices. Taking advantage of the characteristics of the hidden Markov models formalism, widely used in pattern recognition, the paper proposes a formalization of the diagnosis process, addressing the complete chain constituted by monitored component, deviation detection and state diagnosis. Hidden Markov models are well suited to represent problems where the internal state of a certain entity is not known and can only be inferred from external observations of what this entity emits. Such over-time diagnosis is a first class representative of this category of problems. The accuracy of diagnosis carried out through the proposed formalization is then discussed, as well as how to concretely use it to perform state diagnosis and allow direct comparison of alternative solutions Alessandro Daidone, Felicita Di Giandomenico, Andrea Bondavalli, Silvano Chiaradonna |
SRDS | 3 |
| 2006 | Guest Editorial for the Special Issue on the 2005 IEEE/IFIP Conference on Dependable Systems and Networks, including the Dependable Computing and Communications and Performance and Dependability SymposiaabstractNo abstract available. Jean Arlat, Andrea Bondavalli, Boudewijn R. Haverkort, Paulo Veríssimo |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2005 | Experimental Evaluation of the QoS of Failure Detectors on Wide Area NetworkabstractThis paper describes an experiment performed on wide area network to assess and fairly compare the quality of service provided by a large family of failure detectors. Failure detectors are a popular middleware mechanism used for improving the dependability of distributed systems and applications. Their QoS greatly influences the QoS that upper layers may provide. It is thus of uttermost importance to equip a system with an appropriate failure detector and to properly tune its parameters for the most desirable QoS to be provided. The paper first analyzes the QoS indicators and the structure of push-style failure detectors and then introduces the choices for estimators and safety margins used to build several (30) failure detectors. The experimental setup designed and implemented to allow a fair comparison of QoS of the several alternatives in a real representative experimental setting is then described. Finally the results obtained through the experiments and their interpretation are provided. Lorenzo Falai, Andrea Bondavalli |
DSN | 2 |
| 2004 | Congestion analysis during outage, congestion treatment and outage recovery for simple GPRS networksabstractThis paper deals with congestion analysis of a simple GPRS network composed by two cells partially overlapping. In particular, we consider that one of the two cells is affected by an outage and we analyze the effectiveness of applying a class of congestion treatment techniques that ultimately results in a switching of users from the congested cell to the other one. For this purpose, we introduce a modelling technique to support a proper calibration of the parameters involved in a reconfiguration action, in order to successfully treat the congestion phenomenon. The effectiveness of a reconfiguration action is evaluated in terms of indicators that represent the quality of service (QoS) perceived by the users in the congested and adjacent cells. Paolo Lollini, Andrea Bondavalli, Felicita Di Giandomenico, Stefano Porcarelli |
ISCC | 2 |
| 2004 | A Freshness Detection Mechanism for Railway ApplicationsabstractRailway control systems are based on on-board and trackside subsystems for signaling purposes. Several factors demand for new design and implementation solutions for such railway control systems. These factors are related to the design of interoperable railway networks in Europe, the introduction of new technologies and equipment, and the competition in the market of railway products. The safety of such new design and implementation solutions should still be proved in accordance with the CENELEC recommendations. SFDA, safe message freshness detection algorithm among trackside subsystems, is deeply described. SFDA is included in a new message passing safety protocol stack and it allows the detection of "old" messages and the meeting of real time and safety requirements of trackside railway systems. It is demonstrated that the SFDA can detect all the old messages. Moreover a preliminary analysis of its availability characteristics to check whether it is suitable for railways systems is performed through simulation. Andrea Bondavalli, Enrico De Giudici, Stefano Porcarelli, Salvatore Sabina, Fabrizio Zanini |
PRDC | 1 |
| 2004 | Modeling and Analysis of a Scheduled Maintenance System: a DSPN ApproachabstractThis paper describes a way of managing the modeling and analysis of Scheduled Maintenance Systems (SMSs) within an analytically tractable context. We chose a significant case study having a variety of interesting features like a heavily redundant architecture and a test and maintenance policy whose execution is made on-line without halting the system. We applied a methodology we previously developed based on the Deterministic Stochastic Petri Net (DSPN) approach, where the underlying stochastic process is Markov regenerative (MRGP) solved in our setting using an efficient analytical solution method. This methodology is implemented by the DEEM tool specifically developed for modeling and evaluating the dependability of Phased Mission Systems (PMSs). We test our methodology with such a case study to check whether it can master real and complex SMS problems and to compare its efficacy with traditional approaches (fault trees). The paper also investigates the problem of the optimal tuning of a maintenance program, giving a useful decision support tool for evaluating the system performance from the early design stage. Andrea Bondavalli, Roberto Filippini |
Comput. J. | 1 |
| 2004 | Effective Fault Treatment for Improving the Dependability of COTS and Legacy-Based ApplicationsabstractThis paper proposes a novel methodology and an architectural framework for handling multiple classes of faults (namely, hardware-induced software errors in the application, process and/or host crashes or hangs, and errors in the persistent system stable storage) in a COTS and legacy-based application. The basic idea is to use an evidence-accruing fault tolerance manager to choose and carry out one of multiple fault recovery strategies, depending upon the perceived severity of the fault. The methodology and the framework have been applied to a case study system consisting of a legacy system, which makes use of a COTS DBMS for persistent storage facilities. A thorough performability analysis has also been conducted via combined use of direct measurements and analytical modeling. Experimental results demonstrate that effective fault treatment, consisting of careful diagnosis and damage assessment, plays a key role in leveraging the dependability of COTS and legacy-based applications. Andrea Bondavalli, Silvano Chiaradonna, Domenico Cotroneo, Luigi Romano |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2004 | Dependability modeling and evaluation of multiple-phased systems using DEEMabstractMultiple-Phased Systems (MPS), i.e., systems whose operational life can be partitioned in a set of disjoint periods, called "phases", include several classes of systems such as Phased Mission Systems and Scheduled Maintenance Systems. Because of their deployment in critical applications, the dependability modeling and analysis of Multiple-Phased Systems is a task of primary relevance. The phased behavior makes the analysis of Multiple-Phased Systems extremely complex. This paper describes the modeling methodology and the solution procedure implemented in DEEM, a dependability modeling and evaluation tool specifically tailored for Multiple Phased Systems. It also describes its use for the solution of representative MPS problems. DEEM relies upon Deterministic and Stochastic Petri Nets as the modeling formalism, and on Markov Regenerative Processes for the model solution. When compared to existing general-purpose tools based on similar formalisms, DEEM offers advantages on both the modeling side (sub-models neatly model the phase-dependent behaviors of MPS), and on the evaluation side (a specialized algorithm allows a considerable reduction of the solution cost and time). Thus, DEEM is able to deal with all the scenarios of MPS which have been analytically treated in the literature, at a cost which is comparable with that of the cheapest ones, completely solving the issues posed by the phased-behavior of MPS. Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico, Ivan Mura |
IEEE Trans. Reliab. | 1 |
| 2003 | Guest Editorial: Special Issue on Reliable Distributed SystemsabstractESIGNERS of distributed systems are concerned with developing architectures, networking, software, algorithms, and applications. While research in this direction addresses some of the fundamental issues in distributed computing, topics related to modeling and simulation of multiple processor systems, real-time operation, reliability, fault tolerance, information assurance, performance measurements, and evaluation are also critical for the successful functioning of distributed systems. The purpose of this special issue is to serve researchers, designers, and implementers of distributed systems, with emphasis on system properties such as reliability, availability, and performability. In addition to conceptual advancement, this issue is intended to recognize the efforts that are aimed toward experimentation, testbeds, development, exploratory or emerging applications, and measurements from operational systems. The theme of this special issue was made to coincide with the 19th IEEE Symposium on Reliable Distributed Systems held at Nuernberg, Germany, 2000, but the topics and submissions were not restricted to the proceedings of this symposium. We received a total of 55 submissions, of which we selected 11 regular papers. Every submission was sent to at least five referees. We received a total of 201 reviews back from 141 referees. Several papers that are not included in this special issue have been forwarded for consideration in the regular issues of this transaction. Although we wanted to have a good mix of current, successful efforts, innovative ideas on reliable designs and open problems—both conceptual and experimental, the space limitation in the special issue and the type of submissions we received may have precluded some key topics of reliable distributed systems. Yet, we believe that the special issue encompasses major properties of a reliable distributed system. The selected papers are classified into four groups: Shambhu J. Upadhyaya, Andrea Bondavalli |
IEEE Trans. Computers | 2 |
| 2003 | Service-Level Availability Estimation of GPRSabstractThe General Packet Radio Service (GPRS) extends the Global System Mobile Communication (GSM) by introducing a packet-switched transmission service. This paper analyzes the GPRS behavior under critical conditions. In particular, we focus on outages, which significantly impact the GPRS dependability. In fact, during outage periods, the cumulative number of users trying to access the service grows proportionally over time. When the system resumes its operations, the overload caused by accumulated users determines a higher probability of collisions on resources assignment and, therefore, a degradation of the overall QoS. This paper adopts a stochastic activity network modeling approach for evaluating the dependability of a GPRS network under outage conditions. The major contribution of this study lies in the novel perspective the dependability study is framed in. Starting from a quite classical availability analysis, the network dependability figures are incorporated into a very detailed service model that is used to analyze the overload effect GPRS has to face after outages, gaining deep insights on its impact on user's perceived QoS. The result of this modeling is an enhanced availability analysis, which takes into account not only the bare estimation of unavailability periods, but also the important congestion phenomenon following outages that contribute to service degradation for a certain period of time after operations resume. Stefano Porcarelli, Felicita Di Giandomenico, Andrea Bondavalli, Massimo Barbera, Ivan Mura |
IEEE Trans. Mob. Comput. | 3 |
| 2002 | Performance Analysis of a Consensus Algorithm Combining Stochastic Activity Networks and MeasurementsabstractProtocols which solve agreement problems are essential building blocks for fault tolerant distributed applications. While many protocols have been published, little has been done to analyze their performance. This paper represents a starting point for such studies, by focusing on the consensus problem, a problem related to most other agreement problems. The paper analyzes the latency of a consensus algorithm designed for the asynchronous model with failure detectors, by combining experiments on a cluster of PCs and simulation using stochastic activity networks. We evaluated the latency in runs (1) with no failures nor failure suspicions, (2) with failures but no wrong suspicions and (3) with no failures but with (wrong) failure suspicions. We validated the adequacy and the usability of the stochastic activity network model by comparing experimental results with those obtained from the model. This has led us to identify limitations of the model and the measurements, and suggests new directions for evaluating the performance of agreement protocols. Andrea Coccoli, Péter Urbán, Andrea Bondavalli |
DSN | 3 |
| 2002 | Analyzing quality of service of GPRS network systems from a user's perspectiveabstractWith reference to the General Packet Radio Service (GPRS), an extension of the Global System for Mobile Communication (GSM) addressing packet-oriented traffic, this paper contributes to the analysis of the service accomplishment level perceived by GPRS users. The proposed modeling approach builds separately the GPRS and user models; the focus is on the GPRS random access procedure on the one side, and different classes of user behavior on the other side. The overall model is composed of the basic submodels. Quantitative analysis, performed using a simulation approach, is carried out, showing the impact of users' characteristics and network load on identified indicators expressing the QoS as perceived by users. Stefano Porcarelli, Felicita Di Giandomenico, Andrea Bondavalli |
ISCC | 3 |
| 2002 | Implementation of Threshold-based Diagnostic Mechanisms for COTS-Based ApplicationsabstractThis work investigates feasibility issues that must be addressed when threshold-based mechanisms are to be used for diagnostic purposes in COTS-based distributed systems. Threshold based mechanisms have typically been used for such purposes in embedded systems. A variety of solutions exist, with different characteristics of completeness, accuracy, and induced overhead. We first discuss the challenges related to applying such mechanisms to COTS-based distributed applications. We then identify alternative strategies for diagnosis, which use run-time data on COTS component service failures to trigger alarms to reconfiguration and fault treatment mechanisms. We implement those strategies in a system prototype, which is based on a substantial application, i.e. a real world (as opposed to a toy) application. We discuss the relationships between the sensitivity of the quality of service (QoS) provided by the diagnostic mechanisms and the accuracy of the available failure data. Our considerations and preliminary experiments on the prototype suggest that a careful evaluation of tradeoffs must be conducted, in order to achieve the best compromise between accuracy and cost, which depends on application characteristics, and service deployment requirements. Luigi Romano, Andrea Bondavalli, Silvano Chiaradonna, Domenico Cotroneo |
SRDS | 2 |
| 2002 | An adaptive approach to achieving hardware and software fault tolerance in a distributed computing environment
Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico |
J. Syst. Archit. | 1 |
| 2001 | Analysis of the Effects of Outages on the Quality of Service of GPRS Network SystemsabstractThe General Packet Radio Service (GPRS) extends the Global System Mobile Communications (GSM) by addressing packet-oriented traffic. Availability is the most important dependability requirement for such communication systems as GPRS. Focusing on the contention phase, where users compete for channel reservation, this paper analyses the GPRS with the objective to understand its behaviour under critical conditions, as determined by periods of outages, which significantly impact on the resulting dependability. In fact, during outages (service unavailability), users trying to access the service accumulate, leading to an overload of the system. When the system resumes its operations, the accumulated users determine a higher probability of collisions on resources assignment (and therefore a degradation of the QoS perceived by the users). Our analysis, performed using a simulation approach, allowed us to gain insights on the impact of outages on the QoS and of the overload that GPRS systems have to face after outages. F. Tataranni, Stefano Porcarelli, Felicita Di Giandomenico, Andrea Bondavalli |
DSN | 4 |
| 2001 | Analysis and Estimation of the Quality of Service of Group Communication ProtocolsabstractQoS (defined as a proper set of quantitative characteristics) analysis is a necessary step for the early verification and validation of an appropriate design, and for taking design decisions about the most rewarding choice, in relation to user requirements. We describe an analytical approach for the evaluation of the QoS offered by a family of group communication protocols in a wireless environment, and use experimental data to feed our models. Specific indicators have been defined and evaluated, which capture the main characteristics of the protocols and of the environment, focusing our attention on performance and dependability attributes. The defined models account for the correlation among successive packet transmissions due to fading and user mobility. The main purpose of our analysis is to provide a fast, cost effective, and formally sound way to further analyze and understand the protocol behavior and its environment. Andrea Coccoli, Andrea Bondavalli, Felicita Di Giandomenico |
ISORC | 2 |
| 2001 | Tuning of Database Audits to Improve Scheduled Maintenance in Communication Systems
Stefano Porcarelli, Felicita Di Giandomenico, Amine Chohra, Andrea Bondavalli |
SAFECOMP | 4 |
| 2001 | Evaluation of Fault-Tolerant Multiprocessor Systems for High Assurance ApplicationsabstractIn designing high assurance systems, the dependability goals are achieved through the adoption of several fault-tolerance techniques. Unfortunately, their combined effect on the system cannot be, in the general case, derived by straightforward composition of the stand-alone component's analysis, because of mutual dependence of their controlling parameters. In this paper the assessment of overall system dependability induced by such integrated fault-tolerance organization is carried out through a stochastic simulation approach. To this purpose, a few fault-tolerant multiprocessor architectures, based on the integrated usage of standard error-processing structures with a recently-proposed diagnostic mechanism, called $\alpha$-count, are selected and evaluated. The diagnostic mechanism gets its input (error signals) from the error-processing mechanism, whose behaviour is in turn influenced by the rapidity and correctness with which $\alpha$-count identifies permanently/intermittently faulty processors. The choice of the basic fault-tolerance mechanisms to adopt, as well as the reference-system architecture, has been driven by the characteristics of the envisaged target applications: mainly, stringent dependability requirements, to be traded with adequate levels of performance and cost. The analysis has focused on performability, which is an appropriate measure to evaluate whether a certain design is ‘better’ than another under dependability and performance point of view. Fabrizio Grandoni 0002, Silvano Chiaradonna, Felicita Di Giandomenico, Andrea Bondavalli |
Comput. J. | 4 |
| 2001 | Markov Regenerative Stochastic Petri Nets to Model and Evaluate Phased Mission Systems DependabilityabstractThis study deals with model-based dependability transient analysis of phased mission systems. A review of the studies in the literature showed that several aspects of multiphased systems pose challenging problems to the dependability evaluation methods and tools. To attack the weak points of the state-of-the-art we propose a modeling methodology that exploits the power of the class of Markov regenerative stochastic Petri net models. By exploiting the techniques available in the literature for the analysis of the Markov Regenerative Processes, we obtain an analytical solution technique with a low computational complexity, basically dominated by the cost of the separate analysis of the system inside each phase. Last, the existence of analytical solutions allows us to derive the sensitivity functions of the dependability measures, thus providing the dependability engineer with additional means for the study of phased mission systems. Ivan Mura, Andrea Bondavalli |
IEEE Trans. Computers | 2 |
| 2000 | DEEM: A Tool for the Dependability Modeling and Evaluation of Multiple Phased SystemsabstractMultiple-phased systems, whose operational life can be partitioned into a set of disjoint periods called "phases", include several classes of systems, such as phased mission systems and scheduled maintenance systems. Because of their deployment in critical applications, the dependability modeling and analysis of multiple-phased systems is a task of primary relevance. However, the phased behavior makes the analysis of multiple-phased systems extremely complex. This paper is centered on the description and application of DEEM, a dependability modeling and evaluation tool for multiple-phased systems. DEEM supports a powerful and efficient methodology for the analytical dependability modeling and evaluation of multiple-phased systems, based on deterministic and stochastic Petri nets and on Markov regenerative processes. Andrea Bondavalli, Ivan Mura, Silvano Chiaradonna, Roberto Filippini, S. Poli, F. Sandrini |
DSN | 1 |
| 2000 | A Position on Design, Methods, and Tools for Object-Oriented Real-Time ComputingabstractReal-time is a major characteristic of many systems, increasingly employed today in disparate sectors of our society. To specify and program systems, exhibiting real-time properties, a number of computing paradigms have been adopted, both explicitly defined to specify real-time behaviours and imported from other application areas with the addition of mechanisms to deal with real-time. The paper discusses object oriented real-time system design and tools. Andrea Bondavalli, Felicita Di Giandomenico |
ISORC | 1 |
| 2000 | Design, Methods, and Tools for ORC
Edgar Nett, Andrea Bondavalli, Bruce Powel Douglass, Carlos Eduardo Pereira, Douglas C. Schmidt, Bran Selic, Kelvin D. Nilsen |
ISORC | 2 |
| 2000 | Scheduling Solutions for Supporting Dependable Real-Time ApplicationsabstractThis paper deals with tolerance to timing faults in time-constrained systems. TAFT (Time Aware Fault-Tolerant) is a recently devised approach which applies tolerance to timing violations. According to TAFT, a task is structured in a pair, to guarantee that deadlines are met (although possibly offering a degraded service) without requiring the knowledge of task attributes difficult to estimate in practice. Wide margin of actions is left by the TAFT approach in scheduling the task pairs, leading to disparate performances; up to now, poor attention has been devoted to analyse this aspect. The goal of this work is to investigate on the most appropriate scheduling policies to adopt in a system structured in the TAFT fashion, in accordance with system conditions and application requirements. To this end, all experimental evaluation will be conducted based on a variety of scheduling policies, to derive useful indications for the system designer about the most rewarding policies to apply. F. Sandrini, Felicita Di Giandomenico, Andrea Bondavalli, Edgar Nett |
ISORC | 3 |
| 2000 | The meaning and role of value in scheduling flexible real-time systems
Alan Burns 0001, Divya Prasad, Andrea Bondavalli, Felicita Di Giandomenico, Krithi Ramamritham, John A. Stankovic, Lorenzo Strigini |
J. Syst. Archit. | 3 |
| 2000 | Threshold-Based Mechanisms to Discriminate Transient from Intermittent FaultsabstractThis paper presents a class of count-and-threshold mechanisms, collectively named /spl alpha/-count, which are able to discriminate between transient faults and intermittent faults in computing systems. For many years, commercial systems have been using transient fault discrimination via threshold-based techniques. We aim to contribute to the utility of count-and-threshold schemes, by exploring their effects on the system. We adopt a mathematically defined structure, which is simple enough to analyze by standard tools. /spl alpha/-count is equipped with internal parameters that can be tuned to suit environmental variables (such as transient fault rate, intermittent fault occurrence patterns). We carried out an extensive behavior analysis for two versions of the count-and-threshold scheme, assuming, first, exponentially distributed fault occurrencies and, then, more realistic fault patterns. Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico, Fabrizio Grandoni 0002 |
IEEE Trans. Computers | 1 |
| 1999 | Automated Dependability Analysis of UML DesignsabstractThe paper deals with the automatic dependability analysis of systems designed using UML. An automatic transformation is defined for the generation of models to capture systems dependability attributes, like reliability. The transformation concentrates on structural UML views, available early in the design, to operate at different levels of refinement, and tries to capture only the information relevant for dependability to limit the size (state space) of the models. Due to the modular construction, these models can be refined later as more detailed, relevant information becomes available. Moreover a careful selection of those critical parts to be detailed allows one to avoid explosion of the size. An implementation of the transformation is in progress and will be integrated in the toolsets available for the ESPRIT LTR HIDE project. Andrea Bondavalli, Ivan Mura, István Majzik |
ISORC | 1 |
| 1999 | An Optimal Value-Based Admission Policy and its Reflective Use in Real-Time Systems
Andrea Bondavalli, Felicita Di Giandomenico, Ivan Mura |
Real Time Syst. | 1 |
| 1999 | A Contribution to the Evaluation of the Reliability of Iterative-Execution SoftwareabstractThis paper deals with the reliability of software executed iteratively, as for example in process control applications. The probability of mission survival is evaluated taking account of two characteristics of iterative software: (a) system failure, defined in terms of the behaviour of the software over successive iterations, because the controlled system can usually tolerate short bursts of errors; (b) the probabilistic correlation between successive executions of the software, which is to be expected for various reasons. The paper presents models accounting for these characteristics and evaluates their effects. The interesting case of fault-tolerant software is considered as well. Using the example of a ‘pair-and-spare’ type fault-tolerant scheme, the relationships between different aspects of failure behaviour that are covered by the models developed here, and those used elsewhere for fault-tolerant software, are shown. Copyright © 1999 John Wiley & Sons, Ltd. Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico, Lorenzo Strigini |
Softw. Test. Verification Reliab. | 1 |
| 1999 | GUARDS: A Generic Upgradable Architecture for Real-Time Dependable SystemsabstractThe development and validation of fault-tolerant computers for critical real-time applications are currently both costly and time consuming. Often, the underlying technology is out-of-date by the time the computers are ready for deployment. Obsolescence can become a chronic problem when the systems in which they are embedded have lifetimes of several decades. This paper gives an overview of the work carried out in a project that is tackling the issues of cost and rapid obsolescence by defining a generic fault-tolerant computer architecture based essentially on commercial off-the-shelf (COTS) components (both processor hardware boards and real-time operating systems). The architecture uses a limited number of specific, but generic, hardware and software components to implement an architecture that can be configured along three dimensions: redundant channels, redundant lanes, and integrity levels. The two dimensions of physical redundancy allow the definition of a wide variety of instances with different fault tolerance strategies. The integrity level dimension allows application components of different levels of criticality to coexist in the same instance. The paper describes the main concepts of the architecture, the supporting environments for development and validation, and the prototypes currently being implemented. David Powell, Jean Arlat, Ljerka Beus-Dukic, Andrea Bondavalli, P. Coppola, Alessandro Fantechi, Eric Jenn, Christophe Rabéjac, Andy J. Wellings |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 1998 | State Restoration in a COTS-Based N-Modular ArchitectureabstractMechanisms for restoring the state of a channel in an N-modular redundant architecture are necessary to prevent redundancy attrition due to transient faults and to allow failed channels to be brought back on line after repair. This paper considers software-implemented mechanisms for state restoration (SR) in a generic fault-tolerant architecture in which both the underlying hardware and operating system are commercial off-the-shelf (COTS) components. State restoration involves copying the values of state variables from the active channel(s) across to the joining channel. Concurrent updating of state variables by application tasks is considered. Two state restoration schemes are considered: Running SR and Recursive SR. In the former, each state variable is copied exactly once while concurrent updates are written through to the joining channel. In the latter state variables are copied once and then recopied recursively until no concurrent updates are detected. Andrea Bondavalli, Felicita Di Giandomenico, Fabrizio Grandoni 0002, David Powell, Christophe Rabéjac |
ISORC | 1 |
| 1995 | Dependability of Iterative Software: A Model for Evaluating the Effects of Input Correlation
Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico, S. La Torre |
SAFECOMP | 1 |
| 1994 | Efficient Fault Tolerance: An Approach to Deal with Transient Faults in Multiprocessor ArchitecturesabstractDynamic error processing approaches are an important mechanism to increase the reliability in a multiprocessor system, while making efficient use of the available resources. To this end, dynamic error processing must be integrated with a fault treatment approach aiming at optimising resource utilisation. In this paper we propose a diagnosis approach that, accounting for transient faults, tries to remove units very cautiously and to balance between two conflicting requirements. The first is to avoid the removal of units that have experienced transient faults and can be still useful for the system and the other is to avoid to keep failed units whose usage may lead to a premature failure of the system. The proposed fault treatment approach is integrated with a mechanism for dynamic error processing in a complete fault tolerance strategy. Reliability analyses based on the Markov approach and an efficiency evaluation performed by simulation are carried out. Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico |
ICPADS | 1 |
| 1993 | Functional paradigm for designing dependable large-scale parallel computing systemsabstractThe authors propose the use of a functional language and of a dataflow computing model for the design of large-scale parallel computing systems for which dependability, in its aspects of reliability, timeliness, parallelism, and distributedness, is the requirement of main concern. The design methodology is sufficiently flexible to allow for verification and validation of such systems from the very first steps of the design. The advantages of such an approach reside in the potential for parallelism it admits, in being characterized by referential transparency, and in the property of composability. The design description language is extended for dealing with dependability issues, with the creation of a library of fault tolerance schemes which can be used for modular insertion of redundancy. A set of tools is being defined as part of a design development environment allowing the designer to proceed in successive interactive steps, each of which can be validated.> Andrea Bondavalli, Luca Simoncini |
ISADS | 1 |
| 1993 | Data Flow Control Systems: an Example of Safety Validation
Cinzia Bernardeschi, Luca Simoncini, Andrea Bondavalli |
SAFECOMP | 3 |
| 1992 | Dataflow-Like Languages for Real-Time Systems: Issues of Computational Models and NotationsabstractThe use of dataflow-like models for the in-the-large design of real-time applications is discussed. In these models, modules can only communicate by (asynchronously) receiving messages when activated and transmitting result messages when terminating. This rather restrictive computational model allows the description of typical, cyclic control programs, with predictable, well-verifiable behavior. In particular, important timing properties can be dealt with in the in-the-large design. The case for the use of dataflow-like models is outlined, and the choice of appropriate notations, which implies a tradeoff between predictability of behavior and expressive power, and the potential for an advanced design support environment are discussed.> Andrea Bondavalli, Lorenzo Strigini, Luca Simoncini |
SRDS | 1 |
| 1992 | Destination Stripping Dual Ring: A New Protocol for MANs
Andrea Bondavalli, Lorenzo Strigini, Matteo Sereno |
Comput. Networks ISDN Syst. | 1 |
| 1991 | DSDR: A Fair and Efficient Access Protocol for Ring-Topology MANsabstractA new media access control (MAC) protocol for ring-topology metropolitan network (MANs) is defined. The destination-stripping dual ring (DSDR) is defined for slotted-medium dual-ring network. The main feature is the destination-stripping capability, which yields a very high throughput through immediate reuse of the slots. It is shown that the protocol can be tailored for different compromises between guaranteed throughput and total throughput. On each of the two rings, guaranteed throughput, with bounded access delay, can be close to twice the medium capacity, and total throughput, under realistic hypotheses, can be close to four times the medium capacity. The maximum interval between accesses by the same station can be as low as N/2 slot times, with N stations on the network. Multicast can be supported with the same simplicity and the same cost as in other LANs and MANs.> Andrea Bondavalli, Lorenzo Strigini |
INFOCOM | 1 |
| 1989 | MAC Protocols for High-Speed MANs: Performance Comparisons for a Family of Fasnet-Based Protocols
Andrea Bondavalli, Marco Conti, Enrico Gregori, Luciano Lenzini, Lorenzo Strigini |
Comput. Networks ISDN Syst. | 1 |