EDBT 2026 Demo / reviewers in the wild / expert
Filippo Poltronieri
dblp:229/2346
· DBLP profile ↗
27ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0003-2860-4204ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the Impact of Scheduling Strategies in Kubernetes with the KubeTwin Platform
José Santos 0001, Davide Borsatti, Walter Cerroni, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
NetSoft | 5 |
| 2026 | KubeTwin 2.0: Demonstrating the Impact of Scheduling Strategies in Kubernetes
José Santos 0001, Davide Borsatti, Walter Cerroni, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
NetSoft | 5 |
| 2025 | Beyond TimeGraph: A Comparative Analysis of Temporal Generators for Evolving Network GraphsabstractThe growing adoption of Artificial Intelligence (AI) in network and service management demands extensive, diverse, and high-fidelity datasets for training and evaluation. However, collecting real-world network data at scale often faces significant challenges, including privacy concerns, operational constraints, and the rarity of certain events or conditions. Generative AI offers a promising solution by synthesizing realistic data that mirrors complex network dynamics and user behavior.In many application domains-such as mobile connectivity, cybersecurity, and disaster recovery-realism is not only defined by accurate replication of structural features (e.g., connectivity graphs), but also by the ability to model how these features evolve over time. Capturing these temporal dynamics is critical to ensure that AI models trained on synthetic data can generalize effectively to real-world scenarios. One effective approach to this challenge is to transform raw graph data into a compact latent representation, which can then be processed by a temporal generative model. This two-stage framework enables the learning of both structural and temporal characteristics of the underlying system, offering a more comprehensive generative pipeline.Building on previous work that employed Time-series Generative Adversarial Networks (TimeGAN) for this purpose, this paper explores an alternative temporal generative model: DoppelGANger. By integrating DoppelGANger into the graph generation pipeline, we aim to assess whether it can more accurately capture the dynamics of evolving graph structures. Furthermore, we introduce a more rigorous and detailed evaluation of the generated data by comparing decoded synthetic graph sequences against their real-world counterparts using distribution-aware and graph-structural metrics. These metrics provide a clearer picture of the quality and fidelity of the generated data, highlighting key differences between the TimeGAN and DoppelGANger approaches. Edoardo Di Caro, Nicolas Belletti, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
CNSM | 3 |
| 2025 | Investigating Neurosymbolic AI for Intent-based Service ManagementabstractThe increasing complexity and increasing demands of IT applications, especially in federated multi-cluster environments, pose significant challenges for service orchestration. To address these, Zero-Touch Service Management (ZSM) and intent-based management paradigms are gaining traction, allowing users to specify high-level goals rather than low-level configurations. However, current intent-driven approaches often rely on rigid Domain Specific Languages (DSLs) or graphic user interfaces, limiting expressiveness and usability. In this work, we propose a neurosymbolic intent-based platform that leverages Large Language Models (LLMs) for natural language intent ingestion and Answer Set Programming (ASP), a declarative programming paradigm used for solving complex combinatorial problems. The system translates natural language descriptions of microservice requirements into structured policies, enabling explainable service-to-cluster matching across federated Kubernetes environments. We validate our approach through experiments that evaluate both the syntactic correctness and efficiency of various LLMs in intent translation, as well as the computational time of the symbolic placement algorithm. Lorenzo Colombi, Sara Cavicchi, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Pál Varga |
CNSM | 3 |
| 2025 | FedEdge-Learn: a Semi-Supervised Federated Learning Framework for Industry 5.0abstractRecent advancements in Machine Learning (ML) and MLOps for Industry 5.0 have significantly boosted productivity in manufacturing by enabling predictive maintenance and optimizing industrial workflows. However, implementing ML applications in real-world industrial environments presents several challenges, including limited access to labeled data, stringent privacy requirements, and the decentralized nature of industrial data. An effective solution for distributed learning with unlabeled data is essential to address these issues. In this paper, we introduce FedEdge-Learn, a novel Federated Learning framework tailored for Industry 5.0 applications. It focuses on unsupervised K-means clustering enhanced by globally shared data. Our approach safeguards data privacy other than accelerating the onboarding of new machines by leveraging the globally trained model. We validate our framework using both public datasets and real-world industrial data, demonstrating its effectiveness in real-world scenarios. The results show how, with our framework, the K-means algorithm is effective in federated settings, without a significant performance decrease compared to the centralized case. Lorenzo Colombi, Edoardo Di Caro, Simon Dahdal, Filippo Poltronieri, Filippo Tabanelli, Mauro Tortonesi, Cesare Stefanelli, Massimiliano Vignoli |
ISCC | 4 |
| 2025 | TimeGraph: Synthetic Generation of Graph Sequences for Realistic Mobile Connectivity ModelsabstractSoftwarized networking solutions are a key enabler for effective and efficient communications in natural disaster recovery scenarios. However, the design development of reliable and robust softwarization solutions in this context is hampered by the scarcity of reference datasets which accurately capture the real-world behavior – and variability – of those environments. This paper presents a method for synthetic generation of sequences of graphs using state-of-the-art Graph Neural Networks (GNNs) and Time-series Generative Adversarial Networks (TimeGAN). By leveraging available real-world data, the proposed approach generates synthetic datasets that closely replicate the features and connectivity patterns found in actual scenarios. These synthetic datasets not only support the training of AI models but also enable testing and evaluation of solutions across different but similar scenarios. Preliminary results using the Anglova scenario show that our solution accurately captures spatio-temporal behaviour in disrupted networks, making it a powerful tool for developing and validating systems in fields where access to real-world data is limited, enhancing their generalizability and reliability. Edoardo Di Caro, Matteo Brina, Nicolas Belletti, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
NetSoft | 4 |
| 2025 | Chaos Engineering Based Kubernetes Pod Rescheduling Through Deep Sets and Reinforcement LearningabstractKubernetes (K8S) is a widely used orchestration solution that helps manage complex IT applications by providing mechanisms for autoscaling, health checking, cluster formation, and replication, which are essential to deploy and manage the multitude of connected microservices. However, they may suffer in case of unexpected faults which can severely change the underlying computing infrastructure and lead to service outages, highlighting the need for resilient solutions capable of mitigating the adverse effects of faults. To address this, the TELKA sched-uler integrates Chaos Engineering (CE), Reinforcement Learning (RL), and Digital Twin (DT) to reallocate K8S pods evicted due to unexpected faults. While TELKA showed promising results in reallocating evicted pods, its preliminary implementations suffered from scalability issues, as the RL agent could only effectively operate on scenarios with the same number of nodes seen during training. To overcome this limitation, this paper improves TELKA by incorporating a neural network architecture called Deep Sets (DS), which can generalize the operation of TELKA on different numbers of nodes. Experimental results not only demonstrate the validity of the improved TELKA but also show how it can be used to identify good operating conditions. Mattia Zaccarini, Filippo Poltronieri, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 2 |
| 2025 | Hybridized Hot Restart via Reinforcement Learning for Microservice OrchestrationabstractThe Compute Continuum (CC) represents a set of computing resources residing from remote cloud datacenters to dedicated hardware at the edge of the network. Modern applications based on the composition of several microservices can strongly benefit from tools that realize optimal deployment based on the characteristics of each microservice, the availability of computing resources across the CC, the end-users latency, and pricing perspectives. To realize such goals, there is the need for sophisticated solutions capable of efficiently exploring a large space of potential configurations. In this regard, Reinforcement Learning (RL) and Computational Intelligence (CI) techniques represent valuable approaches. However, one of the critical challenges remains the adaption of these deployments to highly dynamic scenarios such as the CC. When the availability of the computing resources changes, there is the need to re-optimize the deployment efficiently. This calls for solutions with reactive or proactive capabilities to deal with the dynamicity of these ecosystems. The work considers a CC-inspired scenario by exploring hybridization techniques that combine CI and RL to devise a hot restart approach for metaheuristics and assess a new deployment solution efficiently. Preliminary results show the soundness of the proposed hybridization methods in handling severe system changes. Mattia Zaccarini, Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 2 |
| 2025 | HephaestusForge: Optimal microservice deployment across the Compute Continuum via Reinforcement LearningabstractWith the advent of containerization technologies, microservices have revolutionized application deployment by converting old monolithic software into a group of loosely coupled containers, aiming to offer greater flexibility and improve operational efficiency. This transition made applications more complex, consisting of tens to hundreds of microservices. Designing effective orchestration mechanisms remains a crucial challenge, especially for emerging distributed cloud paradigms such as the Compute Continuum (CC). Orchestration across multiple clusters is still not extensively explored in the literature since most works consider single-cluster scenarios. In the CC scenario, the orchestrator must decide the optimal locations for each microservice, deciding whether instances are deployed altogether or placed across different clusters, significantly increasing orchestration complexity. This paper addresses orchestration in a containerized CC environment by studying a Reinforcement Learning (RL) approach for efficient microservice deployment in Kubernetes (K8s) clusters, a widely adopted container orchestration platform. This work demonstrates the effectiveness of RL in achieving near-optimal deployment schemes under dynamic conditions, where network latency and resource capacity fluctuate. We extensively evaluate a multi-objective reward function that aims to minimize overall latency, reduce deployment costs, and promote fair distribution of microservice instances, and we compare it against typical heuristic-based approaches. The results from an implemented OpenAI Gym framework, named as HephaestusForge, show that RL algorithms achieve minimal rejection rates (as low as 0.002%, 90x less than the baseline Karmada scheduler). Cost-aware strategies result in lower deployment costs (2.5 units), and latency-aware functions achieve lower latency (268–290 ms), improving by 1.5x and 1.3x, respectively, over the best-performing baselines. HephaestusForge is available in a public open-source repository, allowing researchers to validate their own placement algorithms. This study also highlights the adaptability of the DeepSets (DS) neural network in optimizing microservice placement across diverse multi-cluster setups without retraining. The DS neural network can handle inputs and outputs as arbitrarily sized sets, enabling the RL algorithm to learn a policy not bound to a fixed number of clusters. José Santos 0001, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Nicola Di Cicco, Filip De Turck |
Future Gener. Comput. Syst. | 3 |
| 2025 | RoamML distributed continual learning: Adaptive and flexible data-driven response for disaster recovery operationsabstractIn the aftermath of natural disasters, Human Assistance & Disaster Recovery (HADR) operations have to deal with disrupted communication networks and constrained resources. Such harsh conditions make high-communication-overhead ML approaches — either centralized or distributed — impractical, thus hindering the adoption of AI solutions to implement a critical function for HADR operations: building accurate and up-to-date situational awareness. To address this issue we developed Roaming Machine Learning (RoamML), a novel Distributed Continual Learning Framework designed for HADR operations and based on the premise that moving an ML model is more efficient and robust than either large dataset transfers or frequent model parameter updates. RoamML deploys a mobile AI agent that incrementally train models across network nodes containing yet unprocessed data; at each stop, the agent initiate a local training phase to update its internal ML model parameters. To prioritize the processing of strategically valuable data, RoamML Agents follow a navigation system based upon the concept of Data Gravity, leveraging Multi-Criteria Decision Making techniques to simultaneously consider many objectives for Agent routing optimization, including model learning efficiency and network resource utilization, while seamlessly blending subjective insights from expert judgments with objective metrics derived from quantifiable data to determine each next hop. We conducted extensive experiments to evaluate RoamML, demonstrating the framework’s efficiency to train ML models under highly dynamic, resource-constrained environments. RoamML achieves similar performance to centralized ML training under ideal network conditions and outperforms it in a more realistic scenario with reduced network resources, ultimately saving up to 75% in bandwidth utilization across all experiments. • Human Assistance & Disaster Recovery (HADR) requires accurate situational awareness. • Most distributed ML approaches assume a stable network and are unsuited for HADR. • Approaches based on distributed AI agents and continual learning are more resilient. • Data Gravity represents a solid foundational concept for agent routing optimization. • Data Gravity and MCDM allow to prioritize the processing of critical datasets. Simon Dahdal, Sara Cavicchi, Alessandro Gilli, Filippo Poltronieri, Mauro Tortonesi, Niranjan Suri, Cesare Stefanelli |
J. Netw. Comput. Appl. | 4 |
| 2024 | Multi-Objective Scheduling and Resource Allocation of Kubernetes Replicas Across the Compute ContinuumabstractOrchestrating microservice applications deployed on a federation of globally distributed Kubernetes clusters is a challenging and multifaceted optimization problem. It is not only computationally hard, but also requires balancing a delicate trade-off between competing performance metrics, such as latency, deployment cost, and service interruption frequency. Classical approaches in the literature merge multiple objectives into a single one via, e.g., linear combinations. However, in practice, it is complex to express a priori a quantitative preference between heterogeneous objectives, let alone with simple linear combinations. This paper adopts a more comprehensive approach leveraging proper Multi-Objective Optimization (MOO), with the goal of producing multiple solutions from the Pareto Front (PF). Therefore, the orchestrator can inspect a posteriori all possible "optimal" trade-offs and decide on the strategy that best fits their operating requirements. To solve the MOO problem, this paper adopts state-of-the-art Multi-Objective Evolutionary Algorithms and shows their effectiveness in solving the MOO problem. Illustrative results highlight the practical benefits of a MOO formulation, providing several tens of nondominated solutions and evenly covering the objectives’ space. Nicola Di Cicco, Filippo Poltronieri, José Santos 0001, Mattia Zaccarini, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
CNSM | 2 |
| 2024 | RoamML Platform: Enabling Distributed Continual Learning for Disaster Relief OperationsabstractMachine learning offers a promising avenue for improving the efficiency and effectiveness of decision-making in disaster recovery and relief efforts. These operations face significant hurdles due to the large volumes of data, intermittent connectivity, and infrastructure limitations. In this paper, we present the RoamML Platform, a sophisticated modular implementation of the RoamML framework, designed specifically to address these challenges and enable efficient distributed machine learning. We advocate for a foundational principle that "the transmission of the ML model itself is usually more efficient than the costly transfer of large datasets", leading to a more adaptable training regime. The platform orchestrates the activities of the RoamML model along with its related metadata, collectively referred to as the "RoamML Agent", while faithfully observing the Data Gravity principle to guarantee thorough model training. We extensively validated the platform through a simulated disaster recovery scenario employing the Mininet-WiFi emulator. Our results highlight the benefits of integrating the RoamML framework, including enhanced ML performance and significant bandwidth savings. Simon Dahdal, Alessandro Gilli, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Niranjan Suri |
ISCC | 3 |
| 2024 | TELKA: Twin-Enhanced Learning for Kubernetes ApplicationsabstractChaos engineering is the discipline of injecting computing and network faults, such as increased network latency and unavailability of computing nodes, into an IT system to help developers in identifying problems that could arise in a production environment and tackle them. Several tools have emerged to ease the application of chaos engineering to complex IT systems, leveraging microservice and container-based applications deployed on Kubernetes. However, applying of such tools requires several phases to be put into practice, from defining a steady state to establishing an effective response plan if something goes wrong. To ease the application of chaos engineering in improving the resilience of Kubernetes applications, this work presents a smart scheduler for Kubernetes called TELKA: a Twin-Enhanced Learning for Kubernetes Applications, which combines chaos engineering, Digital Twin (DT), and Reinforcement Learning (RL) methodologies to mitigate the effects of computing and network faults. Instead of interacting directly with the physical Kubernetes application, TELKA learns by interacting with a digital twin, thus reducing the learning time and the operation costs related to the application of chaos engineering. Experiment results compare TELKA with other approaches to show its effectiveness in mitigating the adverse effects of injected faults. Mattia Zaccarini, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
ISCC | 7 |
| 2024 | Efficient Microservice Deployment in Kubernetes Multi-Clusters through Reinforcement LearningabstractMicroservices have revolutionized application deployment in popular cloud platforms, offering flexible scheduling of loosely-coupled containers and improving operational efficiency. However, this transition made applications more complex, consisting of tens to hundreds of microservices. Efficient orchestration remains an enormous challenge, especially with emerging paradigms such as Fog Computing and novel use cases as autonomous vehicles. Also, multi-cluster scenarios are still not vastly explored today since most literature focuses mainly on a single-cluster setup. The scheduling problem becomes significantly more challenging since the orchestrator needs to find optimal locations for each microservice while deciding whether instances are deployed altogether or placed into different clusters. This paper studies the multi-cluster orchestration challenge by proposing a Reinforcement Learning (RL)-based approach for efficient microservice deployment in Kubernetes (K8s), a widely adopted container orchestration platform. The study demonstrates the effectiveness of RL agents in achieving near-optimal allocation schemes, emphasizing latency reduction and deployment cost minimization. Additionally, the work highlights the versatility of the DeepSets neural network in optimizing microservice placement across diverse multi-cluster setups without retraining. Results show that DeepSets algorithms optimize the placement of microservices in a multi-cluster setup 32 times higher than its trained scenario. José Santos 0001, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Sleianelli, Nicola Di Cicco, Filip De Turck |
NOMS | 3 |
| 2024 | The Evolution of Kubernetes Management: Introducing the KubeTwin FrameworkabstractThe current trend in the management of complex and highly distributed microservices architecture is to leverage container-based orchestration tools such as Kubernetes. Kubernetes provides several capabilities that help service providers to deal with essential aspects of service provisioning such as service replication, scalability, and cluster federation. However, finding a suitable configuration for a Kubernetes application and optimizing its deployment in complex scenarios like Compute Continuum ones can be a very time-consuming and challenging operation. To simplify this process, this research project introduces the KubeTwin framework, a Digital Twin-based solution to evaluate the behavior and the impact of Kubernetes applications in a shorter computation time. Due to an accurate virtual representation of a Kubernetes application, KubeTwin provides a plethora of functionalities that can help providers to assess their ecosystems by employing different mechanisms that span from performance optimization to chaos engineering. Results shown by the first experimental model evaluations demonstrate the strength of this project and encourage for future major efforts. Mattia Zaccarini, Mauro Tortonesi, Filippo Poltronieri |
NOMS | 3 |
| 2024 | Chaos Engineering for Resilience Assessment of Digital TwinsabstractWithin the Industry 4.0 vision, digital twins (DTs) have gained great attention as a promising approach to improve remote monitoring and control by means of virtual representations of physical objects. However, while DTs are becoming more and more sophisticated and even adopted for mission-critical applications, their resilience assessment has not received the required consideration yet. This article originally proposes chaos engineering to assess and improve the resilience of DTs by testing multiple aspects of industrial environments in a coordinated, automated, and replicable manner. First, the article discusses why and how chaos engineering is promising to improve the resilience of DTs. Then, it identifies and introduces a set of chaos engineering profiles specifically designed to take into account the many aspects an industrial environment is composed of. Finally, it shows the feasibility of assessing the resilience of a proof-of-concept DT through a testbed based on widely-adopted, open-source tools. Mattia Fogli, Carlo Giannelli, Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | KubeTwin: A Digital Twin Framework for Kubernetes Deployments at ScaleabstractKubernetes is a well-known orchestration and management solution for complex and large-scale service architectures in the Cloud Continuum. While it provides very valuable functions from the operation perspective, the high number of control loops it implements significantly enlarges the already wide space of configuration parameters and policies to consider for management purposes. We argue that optimizing complex Kubernetes deployments considering a multi-cloud and edge computing environment would significantly benefit from a Digital Twin approach, enabling an accurate virtual representation of a Kubernetes application to optimize its deployment and management policies. Towards that goal, this work illustrates the design of KubeTwin, a framework to implement Digital Twins of Kubernetes deployments. Furthermore, we present a validation of KubeTwin in a Multi-access Edge Computing (MEC) scenario, which shows its soundness in reenacting realistic Digital Twins of complex and highly distributed Kubernetes deployments. We believe that KubeTwin can provide useful guidance to the research community working in this field. Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2023 | Characterization of Microservice Response Time in Kubernetes: A Mixture Density Network ApproachabstractThe use of microservice-based applications is becoming more prominent also in the telecommunication field. The current 5G core network, for instance, is already built around the concept of a “Service Based Architecture”, and it is foreseeable that 6G will push even further this concept to enable more flexible and pervasive deployments. However, the increasing complexity of future networks calls for sophisticated platforms that could help network providers with their deployments design. In this framework, a central research trend is the development of digital twins of the physical infrastructures. These digital representations should closely mimic the behavior of the managed system, allowing the operators to test new configurations, analyze what-if scenarios, or train their reinforcement learning algorithms in safe environments. Considering that Kubernetes is becoming the de-facto standard platform for container orchestration and microservice-based application lifecycle management, the implementation of a Kubernetes digital twin requires an accurate characterization of the microservice response time, possibly leveraging suitable Machine Learning techniques trained with measurement data collected in the field. In this paper we introduce a new methodology, based on Mixture Density Networks, to accurately estimate the statistical distribution of the response time of microservice-based applications. We show the improvement in performance with respect to simulation-based inference procedures proposed in literature. Lorenzo Manca, Davide Borsatti, Filippo Poltronieri, Mattia Zaccarini, Domenico Scotece, Gianluca Davoli, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Walter Cerroni |
CNSM | 3 |
| 2023 | Modeling Digital Twins of Kubernetes-Based ApplicationsabstractKubernetes provides several functions that can help service providers to deal with the management of complex container-based applications. However, most of these functions need a time-consuming and costly customization process to address service-specific requirements. The adoption of Digital Twin (DT) solutions can ease the configuration process by enabling the evaluation of multiple configurations and custom policies by means of simulation-based what-if scenario analysis. To facilitate this process, this paper proposes KubeTwin, a framework to enable the definition and evaluation of DTs of Kubernetes applications. Specifically, this work presents an innovative simulation-based inference approach to define accurate DT models for a Kubernetes environment. We experimentally validate the proposed solution by implementing a DT model of an image recognition application that we tested under different conditions to verify the accuracy of the DT model. The soundness of these results demonstrates the validity of the KubeTwin approach and calls for further investigation. Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini |
ISCC | 5 |
| 2023 | Enabling civil-military collaboration for disaster relief operations in smart city environments
Lorenzo Campioni, Filippo Poltronieri, Cesare Stefanelli, Niranjan Suri, Mauro Tortonesi, Konrad S. Wrona |
Future Gener. Comput. Syst. | 2 |
| 2022 | Water 4.0: enabling Smart Water and Environmental Data MeteringabstractSmart metering represents an interesting field, where IoT can bring huge benefits to collect and analyze water and environmental data. However, it presents well-known issues such as the presence of high heterogeneity at the hardware, software, and network layers. This is even amplified by the vendors’ tendency to adopt proprietary solutions and different low-power wireless communication protocols for IoT sensors and metering devices. Realizing an interoperable platform for the collection and analysis of smart Water and environmental data metering thus becomes a complex task, that needs to address multiple requirements at several levels, starting from the collection of data on the field (IoT sensors, smart meters) and its processing on Cloud Computing platforms. The Water 4.0 project, which involves a collaboration between universities and private companies, including Dipietro Group, aims at addressing these challenges. This paper presents the comprehensive smart environmental data metering solution that we realized within Water 4.0, that enables data collection, data-processing, and Over-The-Air (OTA) for IoT devices, with the ultimate goal of reducing water losses and monitoring water quality. Nicola Caldognetto, Luca Pasquali Evangelisti, Filippo Poltronieri, Michele Russo, Cesare Stefanelli, Sara Tenani, Sara Toboli, Mauro Tortonesi |
NOMS | 3 |
| 2022 | Value-of-Information Middleware Solutions for Fog and Edge ComputingabstractFog and Edge Computing aim to deliver low-latency, immersive, and powerful services by processing information close to both devices and users. This is well suited for IoT applications in Smart City, where IoT gateways, Cloudlets, Base Stations, and other computational nodes can process (part of) the data generated by the multitude of IoT sensors directly at the edge of the network. However, the implementation of Fog and Edge Computing is challenging because it requires to deal with a (limited number of) constrained devices, dynamic services’ requirements, and heterogeneous network conditions. Differently from the Cloud, where computational resources are supposed to be unlimited, Fog and Edge services should be capable to adapt to scarce and constrained resources and deal with the deluge of IoT data. To facilitate the adoption of Fog and Edge Computing this work proposes middleware solutions that leverage Value-of-Information (VoI) as interesting criterion to select only the most valuable piece of information for processing and dissemination and to scale computational workload in an automated fashion. Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 1 |
| 2022 | A Chaos Engineering Approach for Improving the Resiliency of IT Services ConfigurationsabstractTesting the resiliency of complex IT services deployed in hybrid Cloud scenarios is a challenging task that requires expensive and possibly destructive operations. An interesting approach lies in Chaos Engineering, a set of practices to test the resiliency of software systems running in a production environment. However, Chaos Engineering is an expensive practice that requires the setup of complicated operations that further increase the complexity of management operations. To reduce this complexity, Chaos Engineering can benefit from the adoption of non-destructive approaches such as the definition of realistic digital twins. A digital twin is a virtual replica of a real-system on which experimenting with management configurations. This paper embraces this research avenue by extending our previous efforts to integrate Chaos Engineering techniques into an IT services management framework called ChaosTwin. ChaosTwin leverages novel methodologies and tools capable of identifying and promptly react to unexpected failures. Finally, to implement autonomous fault management, ChaosTwin defines scaling and migration policies that can quickly explore for more resilient placements of software components in case of system failures. We believe that ChaosTwin can provide useful guidance to service providers in finding cost-effective service configurations capable of minimizing the negative effects of unpredictable events. Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
NOMS | 1 |
| 2022 | BDMaaS+: Business-Driven and Simulation-Based Optimization of IT Services in the Hybrid CloudabstractThe maturity of heterogeneous and hybrid public Cloud environments enables service providers to deploy there their complex IT services trusting these large and complex infrastructures. At the same time, evaluating the impact of changes at service configuration before and at the runtime is still a very challenging and difficult task. Moreover, a comprehensive performance evaluation of IT service configurations should not be limited just to costs for IT resource acquisition, but also include risk related elements such as Service Level Agreement (SLA) violation penalties and other intangibles. To support IT service providers in this difficult task, we developed Business-Driven Management as a Service Plus (BDMaaS+), a novel decision support tool that can evaluate IT service configuration through simulation with realistic service and network models. By allowing service providers to define expanded operational parameters, BDMaaS+ also enables what-if scenario analysis, thereby opening interesting possibilities at the planning level. Experimental results, collected from our thorough evaluations, demonstrate how a service provider can leverage BDMaaS+ to explore the potential of high-level business SLA changes and data center additions. Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2021 | ChaosTwin: A Chaos Engineering and Digital Twin Approach for the Design of Resilient IT ServicesabstractChaos Engineering represents an interesting software engineering methodology to improve the resilience of a complex IT system operating in a live production environment by injecting simulated faults, observing the system reaction, and devising mitigating solutions. However, Chaos Engineering is an expensive practice with a high setup and operation overhead and it often focuses on the evaluation of the system behavior from a relatively narrow technical perspective instead of a more comprehensive business level one. To enlarge the audience of Chaos Engineering there is the need for novel solutions that can give service providers the tools to deal with the deployment and testing of complex IT services. To fill this gap, this paper presents ChaosTwin, a novel solution exploring an innovative approach to apply Chaos Engineering to a digital-twin, i.e., a virtual representation of a physical object or a system. By creating realistic digital twin of an IT service, injecting faults on the digital twin and evaluating how different service configuration and fault management strategies would perform from a business level perspective, ChaosTwin provides useful guidance to service providers in finding cost-effective service configurations that can minimize the negative effects of unpredictable events. Experimental results, collected from the evaluation of a realistic case study, demonstrate how ChaosTwin is capable of minimizing both the associated costs and the effects of injected Chaos faults. Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
CNSM | 1 |
| 2021 | Reinforcement Learning for value-based Placement of Fog Services
Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Niranjan Suri |
IM | 1 |
| 2020 | Value of Information based Optimal Service Fabric Management for Fog ComputingabstractService fabric management in Fog Computing is a challenging task, which has to deal with a complex and resource scarce environment. We argue that approaches leveraging Value-of-Information (VoI) concepts and tools are particularly interesting to support the realization of that objective. This paper describes innovative methodologies and reference models for the service fabric management for Fog Computing applications. First, we formalize the VoI concept and discuss its adoption in Fog Computing environments. Then, we propose a formal model that aims at maximizing the allocation of Fog services from a value-based perspective. To overcome the complexity of this model, we present two possible approaches (simulation-based optimization and a model approximation) and we compare them by adopting Evolutionary Algorithms (EAs) as optimization techniques. Experimental results prove the validity of both models in finding resource allocation solutions capable of minimizing network latency and maximizing the utility for the end-users of Fog Computing services. Finally, we show how the results of the approximated model can be adopted as a first approximated approach for resource management of Fog Computing services. Filippo Poltronieri, Mauro Tortonesi, Alessandro Morelli, Cesare Stefanelli, Niranjan Suri |
NOMS | 1 |