VLDB 2026 Research / reviewers in the wild / expert
Luigi De Simone
dblp:156/2317
· DBLP profile ↗
31ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-6008-2656ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 3 first-author · 7 since 2021Systems, architecture and hardware · 6 · 5 since 2021Computer networks · 5 · 2 first-author · 4 since 2021Security and privacy · 5 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PREEMPT-FaaS: Taming Orchestration Times in Latency-Sensitive Serverless EnvironmentsabstractThe orchestration of application instances is critical for the efficient management of cloud computing platforms. Specifically, the serverless paradigm automates container spawning and de-spawning based on actual load, mitigating inefficiencies, such as over- and under-provisioning, that might compromise Service Level Objectives (SLOs). This dynamic behavior introduces significant challenges concerning initialization and termination latencies, which are exacerbated when enforcing real-time requirements in mixed-criticality systems. The existing literature already addresses key issues, such as reducing cold-start times and assuring real-time performance to deployed instances. However, container orchestration times remain an overlooked factor that can severely affect instance startup times, especially when the orchestrator is subject to intense workloads. In this paper, we present PREEMPT-FaaS, an orchestration controller that, unlike commonly adopted controllers, adopts a fixed-priority preemptive scheduling of requests to guarantee reduced orchestration times to high-priority and highly critical instances. We implemented PREEMPT-FaaS as a Rust custom controller for Kubernetes (K8s), along with a patch for Knative, a popular serverless platform built upon K8s. We perform an extensive experimental campaign of PREEMPT-FaaS, including the serving of AI workloads, such as, recurring neural networks and video analytics, showing up to ∼6× reduction of orchestration times under high load and improving end-to-end cold-start times of critical instances, with a consequent reduction of service-level latencies (up to ∼2 s reduction under stress at the 95th percentile). Marcello Cinque, Luigi De Simone, Raffaele Della Corte, Stefano Toscano |
ECRTS | 2 |
| 2026 | SLO-aware Prioritization of Orchestration Times for Containerized ServicesabstractIn this article, we present a timing analysis of orchestration times for containerized services, revealing the inability of current container orchestrators to fully prioritize services under concurrent requests. The analysis identifies the sources of orchestration delays that impact services to be prioritized potentially violating their Service Level Objectives (SLOs). Based on the findings of the timing analysis, we highlight three alternative SLO-aware orchestration system designs aimed at preventing and/or mitigating delays for high-priority services. We provide principles and guidelines that must drive the implementation of these designs. We then introduce Ulysses , a Kubernetes -based prototype embodying the simplest of the three designs. Ulysses modifies the core Kubernetes control plane components to manage events synchronously and with fixed priority. Through experiments conducted with both synthetic workloads and a containerized cloud-native 5G core network, we demonstrate that Ulysses ensures stable orchestration times for high-priority services, with a reduction of up to 78% under high orchestration load. Marco Barletta, Marcello Cinque, Luigi De Simone |
ACM Trans. Internet Techn. | 3 |
| 2025 | PREEMPT-K8S: Pod Prioritization for Mixed-Criticality Edge-Cloud ServicesabstractIn this paper, we design and implement a fixedpriority and fully preemptable controller for Kubernetes. The controller is designed to manage mixed-criticality services, handling orchestration requests and cluster events with a priority level that matches each service’s criticality. The controller aims at providing predictable and schedulable orchestration times according to services’ priority, even in the presence of interfering orchestration events. Experimental results show a reduction of up to 99% of the time spent in the control plane to handle highpriority requests. Stefano Toscano, Luigi De Simone, Marco Barletta, Marcello Cinque |
DSD | 2 |
| 2025 | Performability Management of 5G Service Chains with Rejuvenation: The Open5GS Use CaseabstractThis paper presents a stochastic framework for managing the performability (performance and availability) of 5G-based service function chains (SFCs). By integrating an$M / G / m$queueing model for latency estimation and Stochastic Reward Networks (SRNs) for availability assessment, we evaluate the impact of software rejuvenation on 5 G network performability. The final goal is to derive the optimal 5G setting that meets both performance (e.g., delay threshold) and availability (e.g., the “five nines”). Our testbed, based on Open5GS, validates the model and provides insights into optimal 5G settings that balance performance, availability, and resource utilization. Luigi De Simone, Mario Di Mauro, Maurizio Longo, Roberto Natella, Fabio Postiglione |
NetSoft | 1 |
| 2025 | COSMOS: A Fault Injection Framework to Assess Hardware-Assisted HypervisorsabstractHardware-assisted virtualization represents a pillar technology for large-scale clusters and cloud-based applications. Hardware faults are still frequent as technology advances, potentially resulting in serious reliability concerns. This paper introduces COSMOS, a fault injection framework tailored for testing hardware-assisted hypervisors. By exploiting nested virtualization, COSMOS does not require instrumentation of the target and enables the assessment of multiple hypervisors. We performed an extensive fault injection campaign to assess popular hardware-assisted hypervisors like KVM, Xen, and Jailhouse. The results show a non-negligible percentage of non–fail-stop behaviors, with notable differences in hypervisors’ ability to log failures and prevent fault propagation with a timely recovery. Marcello Cinque, Domenico Cotroneo, Giuseppe De Rosa, Luigi De Simone, Giorgio Farina |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Temporal isolation assessment in virtualized safety-critical mixed-criticality systems: A case study on Xen hypervisorabstractToday, we are witnessing the increasing use of the cloud and virtualization technologies, which are a prominent way for the industry to develop mixed-criticality systems (MCSs) and reduce SWaP-C factors (size, weight, power, and cost) by flexibly consolidating multiple critical and non-critical software on the same System-on-a-Chip (SoC). Unfortunately, using virtualization leads to several issues in assessing isolation aspects, especially temporal behaviors, which must be evaluated due to safety-related standards (e.g., EN50128 in the railway domain). This study proposes a systematic approach for verifying temporal isolation properties in virtualized MCSs to characterize and mitigate timing failures, which is a fundamental aspect of dependability. In particular, as proof of the effectiveness of our proposal, we exploited the real-time flavor of Xen hypervisor used to deploy a virtualized 2 out of 2-based MCS scenario provided in the framework of an academic-industrial partnership, in the context of the railway domain. The results point out that virtualization overhead must be carefully tuned in a real industrial scenario according to the several features provided by a specific hypervisor solution. Further, we identify a set of directions toward employing virtualization in industry in the context of ARM-based mixed-criticality systems. Marcello Cinque, Luigi De Simone, Daniele Ottaviano |
J. Syst. Softw. | 2 |
| 2024 | Criticality-aware Monitoring and Orchestration for Containerized Industry 4.0 EnvironmentsabstractThe evolution of industrial environments makes the reconfigurability and flexibility key requirements to rapidly adapt to changeable market needs. Computing paradigms like Edge/Fog computing are able to provide the required flexibility and scalability while guaranteeing low latencies and response times. Orchestration systems play a key role in these environments, enforcing automatic management of resources and workloads’ lifecycle, and drastically reducing the need for manual interventions. However, they do not currently meet industrial non-functional requirements, such as real-timeliness, determinism, reliability, and support for mixed-criticality workloads. In this article, we present k4.0s, an orchestration system for Industry 4.0 (I4.0) environments, which enables the support for real-time and mixed-criticality workloads. We highlight through experiments the need for novel monitoring approaches and propose a workflow for selecting monitoring metrics, which depends on both workload requirements and hosting node guarantees. We introduce new abstractions for the components of a cluster in order to enable criticality-aware monitoring and orchestration of real-time industrial workloads. Finally, we design an orchestration system architecture that reflects the proposed model, introducing new components and prototyping a Kubernetes-based implementation, taking the first steps towards a fully I4.0-enabled orchestration system. Marco Barletta, Marcello Cinque, Luigi De Simone, Raffaele Della Corte |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Performance and Availability Challenges in Designing Resilient 5G ArchitecturesabstractThis work proposes a stochastic characterization of resilient 5G architectures, where attributes such as performance and availability play a crucial role. As regards performance, we focus on the delay associated with the Packet Data Unit session establishment, a 5G procedure recognized as critical for its impact on the Quality of Service and Experience of end-users. To formally characterize this aspect, we employ the non-product-form queueing networks framework where: i) main nodes of a 5G architecture have been realistically modeled as G/G/m queues which do not admit analytical solutions; ii) the decomposition method useful to catch subtle quantities involved in the chain of 5G interconnected nodes has been conveniently customized. The results of performance characterization constitute the input of the availability modeling, where we design a hierarchical scheme to characterize the probabilistic failure/repair behavior of 5G nodes combining two formalisms: i) the Reliability Block Diagrams, useful to capture the high-level interconnections between nodes; ii) the Stochastic Reward Networks to model the internal structure of each node. The final result is an optimal resilient 5G setting that fulfills both a performance constraint (e.g., a temporal threshold) and an availability constraint (e.g., the so-called five nines) at the minimum cost, namely, with the smallest number of redundant elements. The theoretical part is complemented by an empirical assessment carried out through Open5GS, a 5G testbed that we have deployed to realistically estimate main performance and availability metrics. Luigi De Simone, Mario Di Mauro, Roberto Natella, Fabio Postiglione |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | IRIS: a Record and Replay Framework to Enable Hardware-assisted Virtualization FuzzingabstractNowadays, industries are looking into virtualization as an effective means to build safe applications, thanks to the isolation it can provide among virtual machines (VMs) running on the same hardware. In this context, a fundamental issue is understanding to what extent the isolation is guaranteed, despite possible (or induced) problems in the virtualization mechanisms. Uncovering such isolation issues is still an open challenge, especially for hardware-assisted virtualization, since the search space should include all the possible VM states (and the linked hypervisor state), which is prohibitive. In this paper, we propose IRIS, a framework to record (learn) sequences of inputs (i.e., VM seeds) from the real guest execution (e.g., OS boot), replay them as-is to reach valid and complex VM states, and finally use them as valid seed to be mutated for enabling fuzzing solutions for hardware-assisted hypervisors. We demonstrate the accuracy and efficiency of IRIS in automatically reproducing valid VM behaviors, with no need to execute guest workloads. We also provide a proof-of-concept fuzzer, based on the proposed architecture, showing its potential on the Xen hypervisor. Carmine Cesarano 0002, Marcello Cinque, Domenico Cotroneo, Luigi De Simone, Giorgio Farina |
DSN | 4 |
| 2023 | LSTM-based failure prediction for railway rolling stock equipmentabstractIn the railway domain, rolling stock maintenance affects service operation time and efficiency. Minimizing train unavailability is essential for reducing capital loss and operational costs. To this aim, prediction of failures of rolling stock equipment is crucial to proactively trigger proper maintenance activities. Indeed, predictive maintenance is a golden example of the digital transformation within Industry 4.0, which affects several engineering processes in the railway domain. Nowadays, it may leverage artificial intelligence and machine learning algorithms to forecast failures and schedule the optimal time for maintenance actions. Generally, rail systems deteriorate gradually over time or fail directly, leading to data that vary extremely slowly. Indeed, ML approaches for predictive maintenance should consider this type of data to accurately predict and forecast failures. This paper proposes a methodology based on Long Short-Term Memory deep learning algorithms for predictive maintenance of railway rolling stock equipment. The methodology allows us to properly learn long-term dependencies for gradually changing data, and both predicting and forecasting failures of rail equipment. In the framework of an academic-industrial partnership, the methodology is experimented on a train traction converter cooling system, demonstrating its applicability and benefits. The results show that it outperforms state-of-the-art methods, reaching a failure prediction and forecasting accuracy over 99%, with a false alarm rate of ∼0.4% and a mean absolute error in the order of 10−4, respectively. Luigi De Simone, Enzo Caputo, Marcello Cinque, Antonio Galli, Vincenzo Moscato, Stefano Russo 0001, Guido Cesaro, Vincenzo Criscuolo, Giuseppe Giannini |
Expert Syst. Appl. | 1 |
| 2023 | Run-time failure detection via non-intrusive event analysis in a large-scale cloud computing platformabstractCloud computing systems fail in complex and unforeseen ways due to unexpected combinations of events and interactions among hardware and software components. These failures are especially problematic when they are silent, i.e., not accompanied by any explicit failure notification, hindering the timely detection and recovery. In this work, we propose an approach to run-time failure detection tailored for monitoring multi-tenant and concurrent cloud computing systems. The approach uses a non-intrusive form of event tracing, without manual changes to the system’s internals to propagate session identifiers (IDs), and builds a set of lightweight monitoring rules from fault-free executions. We evaluated the effectiveness of the approach in detecting failures in the context of the OpenStack cloud computing platform, a complex and “off-the-shelf” distributed system, by executing a campaign of fault injection experiments in a multi-tenant scenario. Our experiments show that the approach detects the failure with an F1 score (0.85) and accuracy (0.77) higher than the ones provided by the OpenStack failure logging mechanisms (0.53 and 0.50) and two non-session-aware run-time verification approaches (both lower than 0.15). Moreover, the approach significantly decreases the average time to detect failures at run-time (∼114 seconds) compared to the OpenStack logging mechanisms. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 2 |
| 2023 | Evaluating virtualization for fog monitoring of real-time applications in mixed-criticality systemsabstractAbstract Technological advances in embedded systems and the advent of fog computing led to improved quality of service of applications of cyber-physical systems. In fact, the deployment of such applications on powerful and heterogeneous embedded systems, such as multiprocessors system-on-chips (MPSoCs), allows them to meet latency requirements and real-time operation. Highly relevant to the industry and our reference case-study, the challenging field of nuclear fusion deploys the aforementioned applications, involving high-frequency control with hard real-time and safety constraints. The use of fog computing and MPSoCs is promising to achieve safety, low latency, and timeliness of such control. Indeed, on one hand, applications designed according to fog computing distribute computation across hierarchically organized and geographically distributed edge devices, enabling timely anomaly detection during high-frequency sampling of time series, and, on the other hand, MPSoCs allow leveraging fog computing and integrating monitoring by deploying tasks on a flexible platform suited for mixed-criticality software, leading to so-called mixed criticality systems (MCSs). However, the integration of such software on the same MPSoC opens challenges related to predictability and reliability guarantees, as tasks interfering with each other when accessing the same shared MPSoC resources may introduce non-deterministic latency, possibly leading to failures on account of deadline overruns. Addressing the design, deployment, and evaluation of MCSs on MPSoCs, we propose a model-based system development process that facilitates the integration of real-time and monitoring software on the same platform by means of a formal notation for modeling the design and deployment of MPSoCs. The proposed notation allows developers to leverage embedded hypervisors for monitoring real-time applications and guaranteeing predictability by isolation of hardware resources. Providing evidence of the feasibility of our system development process and evaluating the industry-relevant class of nuclear fusion applications, we experiment with a safety-critical case-study in the context of the ITER nuclear fusion reactor. Our experimentation involves the design and evaluation of several prototypes deployed as MCSs on a virtualized MPSoC, showing that deployment choices linked to the monitor placement and virtualization configurations (e.g., resource allocation, partitioning, and scheduling policies) can significantly impact the predictability of MCSs in terms of Worst-Case Execution Times and other related metrics. Marcello Cinque, Luigi De Simone, Nicola Mazzocca, Daniele Ottaviano, Francesco Vitale |
Real Time Syst. | 2 |
| 2023 | Guest editorial: special issue on emerging challenges in software certification and verification
Luigi De Simone, Nuno Laranjeiro, Domenico Cotroneo |
Softw. Qual. J. | 1 |
| 2023 | Multi-Provider IMS Infrastructure With Controlled Redundancy: A Performability EvaluationabstractIn modern telecommunication networks, services are provided through Service Function Chains (SFC), where network resources are implemented by leveraging virtualization and containerization technologies. In particular, the possibility of easily adding or removing network resources has prompted service providers to redefine some concepts including performance and availability. In line with this new trend, we propose a performability study of a multi-provider containerized IP Multimedia Subsystem (cIMS), an SFC-like infrastructure used in the core part of 4G/5G networks to handle multimedia sessions. On the one hand, performance issues are tackled by modeling each cIMS node in terms of a G/G/m queueing system to derive the Call Setup Delay (CSD), a performance metric related to the user-end experience in multimedia communications. On the other hand, availability issues are addressed through the Multi-State System (MSS) formalism, to take into account different performance rates of the system. Then, we devise an algorithm called PE-MUGF (Performability Evaluation through Multidimensional Universal Generating Function) to identify the minimum-redundancy cIMS configuration which meets given performance and availability targets at the same time. Finally, an extensive experimental analysis based on Clearwater, a containerized IMS testbed, allows us to estimate most of system parameters whose robustness is evaluated through a sensitivity analysis. Luigi De Simone, Mario Di Mauro, Maurizio Longo, Roberto Natella, Fabio Postiglione |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | A Latency-Driven Availability Assessment for Multi-Tenant Service ChainsabstractNowadays, most telecommunication services adhere to the Service Function Chain (SFC) paradigm, where network functions are implemented via software. In particular, container virtualization is becoming a popular approach to deploy network functions and to enable resource slicing among several tenants. The resulting infrastructure is a complex system composed by a huge amount of containers implementing different SFC functionalities, along with different tenants sharing the same chain. The complexity of such a scenario lead us to evaluate two critical metrics: the steady-state availability (the probability that a system is functioning in long runs) and the latency (the time between a service request and the pertinent response). Consequently, we propose a latency-driven availability assessment for multi-tenant service chains implemented via Containerized Network Functions (CNFs). We adopt a multi-state system to model single CNFs and the queueing formalism to characterize the service latency. To efficiently compute the availability, we develop a modified version of the Multidimensional Universal Generating Function (MUGF) technique. Finally, we solve an optimization problem to minimize the SFC cost under an availability constraint. As a relevant example of SFC, we consider a containerized version of IP Multimedia Subsystem, whose parameters have been estimated through fault injection techniques and load tests. Luigi De Simone, Mario Di Mauro, Roberto Natella, Fabio Postiglione |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | Performability Assessment of Containerized Multi-Tenant IMS through Multidimensional UGFabstractWe advance a performability assessment of a multi- tenant containerized IP Multimedia Subsystem (cIMS), i.e.: one and the same infrastructure is shared among different providers (or tenants). Specifically, we: i) model each cIMS node (a.k.a. Containerized Network Function - CNF) through the Multi-State System (MSS) formalism to capture the dimensionality of the multi-tenant arrangement, and characterize each tenant through queueing theory attributes to catch latency-dependent performance aspects; ii) afford an availability analysis of cIMS by means of an extended version of the Universal Generating Function (UGF) technique, dubbed Multidimensional UGF (MUGF); iii) solve an optimization problem to retrieve the cIMS deployment minimizing costs while guaranteeing high availability requirements. The whole assessment is supported by an experiment based on the containerized IMS platform Clearwater which we deploy to derive some realistic system parameters by means of fault injection techniques. Luigi De Simone, Mario Di Mauro, Maurizio Longo, Roberto Natella, Fabio Postiglione |
CNSM | 1 |
| 2022 | Achieving Isolation in Mixed-Criticality Industrial Edge Systems with Real-Time ContainersabstractReal-time containers are a promising solution to reduce latencies in time-sensitive cloud systems. Recent efforts are emerging to extend their usage in industrial edge systems with mixed-criticality constraints. In these contexts, isolation becomes a major concern: a disturbance (such as timing faults or unexpected overloads) affecting a container must not impact the behavior of other containers deployed on the same hardware. In this paper, we propose a novel architectural solution to achieve isolation in real-time containers, based on real-time co-kernels, hierarchical scheduling, and time-division networking. The architecture has been implemented on Linux patched with the Xenomai co-kernel, extended with a new hierarchical scheduling policy, named SCHED_DS, and integrating the RTNet stack. Experimental results are promising in terms of overhead and latency compared to other Linux-based solutions. More importantly, the isolation of containers is guaranteed even in presence of severe co-located disturbances, such as faulty tasks (elapsing more time than declared) or high CPU, network, or I/O stress on the same machine. Marco Barletta, Marcello Cinque, Luigi De Simone, Raffaele Della Corte |
ECRTS | 3 |
| 2022 | Virtualizing mixed-criticality systems: A survey on industrial trends and issues
Marcello Cinque, Domenico Cotroneo, Luigi De Simone, Stefano Rosiello |
Future Gener. Comput. Syst. | 3 |
| 2022 | ThorFI: a Novel Approach for Network Fault Injection as a Service
Domenico Cotroneo, Luigi De Simone, Roberto Natella |
J. Netw. Comput. Appl. | 2 |
| 2022 | Software micro-rejuvenation for Android mobile systems
Domenico Cotroneo, Luigi De Simone, Roberto Natella, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 2 |
| 2022 | Fault Injection Analytics: A Novel Approach to Discover Failure Modes in Cloud-Computing SystemsabstractCloud computing systems fail in complex and unexpected ways due to unexpected combinations of events and interactions between hardware and software components. Fault injection is an effective means to bring out these failures in a controlled environment. However, fault injection experiments produce massive amounts of data, and manually analyzing these data is inefficient and error-prone, as the analyst can miss severe failure modes that are yet unknown. This article introduces a new paradigm (fault injection analytics) that applies unsupervised machine learning on execution traces of the injected system, to ease the discovery and interpretation of failure modes. We evaluated the proposed approach in the context of fault injection experiments on the OpenStack cloud computing platform, where we show that the approach can accurately identify failure modes with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Timing covert channel analysis of the VxWorks MILS embedded hypervisor under the common criteria security certification
Domenico Cotroneo, Luigi De Simone, Roberto Natella |
Comput. Secur. | 2 |
| 2021 | Enhancing the analysis of software failures in cloud computing systems with deep learning
Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 2 |
| 2020 | ProFIPy: Programmable Software Fault Injection as-a-ServiceabstractIn this paper, we present a new fault injection tool (ProFIPy) for Python software. The tool is designed to be programmable, in order to enable users to specify their software fault model, using a domain-specific language (DSL) for fault injection. Moreover, to achieve better usability, ProFIPy is provided as software-as-a-service and supports the user through the configuration of the faultload and workload, failure data analysis, and full automation of the experiments using container- based virtualization and parallelization. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
DSN | 2 |
| 2019 | Analyzing the Context of Bug-Fixing Changes in the OpenStack Cloud Computing PlatformabstractMany research areas in software engineering, such as mutation testing, automatic repair, fault localization, and fault injection, rely on empirical knowledge about recurring bug-fixing code changes. Previous studies in this field focus on what has been changed due to bug-fixes, such as in terms of code edit actions. However, such studies did not consider where the bug-fix change was made (i.e., the context of the change), but knowing about the context can potentially narrow the search space for many software engineering techniques (e.g., by focusing mutation only on specific parts of the software). Furthermore, most previous work on bug-fixing changes focused on C and Java projects, but there is little empirical evidence about Python software. Therefore, in this paper we perform a thorough empirical analysis of bug-fixing changes in three OpenStack projects, focusing on both the what and the where of the changes. We observed that all the recurring change patterns are not oblivious with respect to the surrounding code, but tend to occur in specific code contexts. Domenico Cotroneo, Luigi De Simone, Antonio Ken Iannillo, Roberto Natella, Stefano Rosiello, Nematollah Bidokhti |
ISSRE | 2 |
| 2019 | Enhancing Failure Propagation Analysis in Cloud Computing SystemsabstractIn order to plan for failure recovery, the designers of cloud systems need to understand how their system can potentially fail. Unfortunately, analyzing the failure behavior of such systems can be very difficult and time-consuming, due to the large volume of events, non-determinism, and reuse of third-party components. To address these issues, we propose a novel approach that joins fault injection with anomaly detection to identify the symptoms of failures. We evaluated the proposed approach in the context of the OpenStack cloud computing platform. We show that our model can significantly improve the accuracy of failure analysis in terms of false positives and negatives, with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ISSRE | 2 |
| 2019 | How bad can a bug get? an empirical analysis of software failures in the OpenStack cloud computing platformabstractCloud management systems provide abstractions and APIs for programmatically configuring cloud infrastructures. Unfortunately, residual software bugs in these systems can potentially lead to high-severity failures, such as prolonged outages and data losses. In this paper, we investigate the impact of failures in the context widespread OpenStack cloud management system, by performing fault injection and by analyzing the impact of the resulting failures in terms of fail-stop behavior, failure detection through logging, and failure propagation across components. The analysis points out that most of the failures are not timely detected and notified; moreover, many of these failures can silently propagate over time and through components of the cloud management system, which call for more thorough run-time checks and fault containment. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ESEC/SIGSOFT FSE | 2 |
| 2018 | Run-Time Detection of Protocol Bugs in Storage I/O Device DriversabstractProtocol violation bugs in storage device drivers are a critical threat for data integrity, since these bugs can silently corrupt the commands and data flowing between the OS and storage devices. Due to their nature, these bugs are notoriously difficult to find by traditional testing. In this paper, we propose a run-time monitoring approach for storage device drivers, in order to detect I/O protocol violations that would otherwise silently escalate in corruptions of users' data. The monitoring approach detects violations of I/O protocols by automatically learning a reference model from failure-free execution traces. The approach focuses on selected portions of the storage controller interface, in order to achieve a good tradeoff in terms of low performance overhead and high coverage and accuracy of failure detection. We assess these properties on three real-world storage device drivers from the Linux kernel, through fault injection and stress tests. Moreover, we show that the monitoring approach only requires few minutes of training workload, and that it is robust to differences between the operational and the training workloads. Domenico Cotroneo, Luigi De Simone, Roberto Natella |
IEEE Trans. Reliab. | 2 |
| 2017 | NFV-Bench: A Dependability Benchmark for Network Function Virtualization SystemsabstractNetwork function virtualization (NFV) envisions the use of cloud computing and virtualization technology to reduce costs and innovate network services. However, this paradigm shift poses the question whether NFV will be able to fulfill the strict performance and dependability objectives required by regulations and customers. Thus, we propose a dependability benchmark to support NFV providers at making informed decisions about which virtualization, management, and application-level solutions can achieve the best dependability. We define in detail the use cases, measures, and faults to be injected. Moreover, we present a benchmarking case study on two alternative, production-grade virtualization solutions, namely VMware ESXi/vSphere (hypervisor-based) and Linux/Docker (container-based), on which we deploy an NFV-oriented IMS system. Despite the promise of higher performance and manageability, our experiments suggest that the container-based configuration can be less dependable than the hypervisor-based one, and point out which faults NFV designers should address to improve dependability. Domenico Cotroneo, Luigi De Simone, Roberto Natella |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2015 | MoIO: Run-time monitoring for I/O protocol violations in storage device driversabstractBugs affecting storage device drivers include the so-called protocol violation bugs, which silently corrupt data and commands exchanged with I/O devices. Protocol violations are very difficult to prevent, since testing device driver is notoriously difficult. To address them, we present a monitoring approach for device drivers (MoIO) to detect HO protocol violations at run-time. The approach infers a model of the interactions between the storage device driver, the OS kernel, and the hardware (the device driver protocol) by analyzing execution traces. The model is then used as a reference for detecting violations in production. The approach has been designed to have a low overhead and to overcome the lack of source code and protocol documentation. We show that the approach is feasible and effective by applying it on the SATA/AHCI storage device driver of the Linux kernel, and by performing fault injection and long-running tests. Domenico Cotroneo, Luigi De Simone, Francesco Fucci, Roberto Natella |
ISSRE | 2 |
| 2015 | Dependability evaluation and benchmarking of Network Function Virtualization InfrastructuresabstractNetwork Function Virtualization (NFV) is an emerging solution that aims at improving the flexibility, the efficiency and the manageability of networks, by leveraging virtualization and cloud computing technologies to run network appliances in software. However, the “softwarization” of network functions raises reliability concerns, as they will be exposed to faults in commodity hardware and software components. In this paper, we propose a methodology for the dependability evaluation and benchmarking of NFV Infrastructures (NFVIs), based on fault injection. We discuss the application of the methodology in the context of a virtualized IP Multimedia Subsystem (IMS), and the pitfalls in the design of a reliable NFVI. Domenico Cotroneo, Luigi De Simone, Antonio Ken Iannillo, Anna Lanzaro, Roberto Natella |
NetSoft | 2 |