VLDB 2026 Research / reviewers in the wild / expert
Mattia Zaccarini
dblp:339/6353
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0003-6679-5125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the Impact of Scheduling Strategies in Kubernetes with the KubeTwin Platform
José Santos 0001, Davide Borsatti, Walter Cerroni, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
NetSoft | 4 |
| 2026 | KubeTwin 2.0: Demonstrating the Impact of Scheduling Strategies in Kubernetes
José Santos 0001, Davide Borsatti, Walter Cerroni, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
NetSoft | 4 |
| 2025 | Chaos Engineering Based Kubernetes Pod Rescheduling Through Deep Sets and Reinforcement LearningabstractKubernetes (K8S) is a widely used orchestration solution that helps manage complex IT applications by providing mechanisms for autoscaling, health checking, cluster formation, and replication, which are essential to deploy and manage the multitude of connected microservices. However, they may suffer in case of unexpected faults which can severely change the underlying computing infrastructure and lead to service outages, highlighting the need for resilient solutions capable of mitigating the adverse effects of faults. To address this, the TELKA sched-uler integrates Chaos Engineering (CE), Reinforcement Learning (RL), and Digital Twin (DT) to reallocate K8S pods evicted due to unexpected faults. While TELKA showed promising results in reallocating evicted pods, its preliminary implementations suffered from scalability issues, as the RL agent could only effectively operate on scenarios with the same number of nodes seen during training. To overcome this limitation, this paper improves TELKA by incorporating a neural network architecture called Deep Sets (DS), which can generalize the operation of TELKA on different numbers of nodes. Experimental results not only demonstrate the validity of the improved TELKA but also show how it can be used to identify good operating conditions. Mattia Zaccarini, Filippo Poltronieri, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 1 |
| 2025 | Hybridized Hot Restart via Reinforcement Learning for Microservice OrchestrationabstractThe Compute Continuum (CC) represents a set of computing resources residing from remote cloud datacenters to dedicated hardware at the edge of the network. Modern applications based on the composition of several microservices can strongly benefit from tools that realize optimal deployment based on the characteristics of each microservice, the availability of computing resources across the CC, the end-users latency, and pricing perspectives. To realize such goals, there is the need for sophisticated solutions capable of efficiently exploring a large space of potential configurations. In this regard, Reinforcement Learning (RL) and Computational Intelligence (CI) techniques represent valuable approaches. However, one of the critical challenges remains the adaption of these deployments to highly dynamic scenarios such as the CC. When the availability of the computing resources changes, there is the need to re-optimize the deployment efficiently. This calls for solutions with reactive or proactive capabilities to deal with the dynamicity of these ecosystems. The work considers a CC-inspired scenario by exploring hybridization techniques that combine CI and RL to devise a hot restart approach for metaheuristics and assess a new deployment solution efficiently. Preliminary results show the soundness of the proposed hybridization methods in handling severe system changes. Mattia Zaccarini, Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 1 |
| 2025 | HephaestusForge: Optimal microservice deployment across the Compute Continuum via Reinforcement LearningabstractWith the advent of containerization technologies, microservices have revolutionized application deployment by converting old monolithic software into a group of loosely coupled containers, aiming to offer greater flexibility and improve operational efficiency. This transition made applications more complex, consisting of tens to hundreds of microservices. Designing effective orchestration mechanisms remains a crucial challenge, especially for emerging distributed cloud paradigms such as the Compute Continuum (CC). Orchestration across multiple clusters is still not extensively explored in the literature since most works consider single-cluster scenarios. In the CC scenario, the orchestrator must decide the optimal locations for each microservice, deciding whether instances are deployed altogether or placed across different clusters, significantly increasing orchestration complexity. This paper addresses orchestration in a containerized CC environment by studying a Reinforcement Learning (RL) approach for efficient microservice deployment in Kubernetes (K8s) clusters, a widely adopted container orchestration platform. This work demonstrates the effectiveness of RL in achieving near-optimal deployment schemes under dynamic conditions, where network latency and resource capacity fluctuate. We extensively evaluate a multi-objective reward function that aims to minimize overall latency, reduce deployment costs, and promote fair distribution of microservice instances, and we compare it against typical heuristic-based approaches. The results from an implemented OpenAI Gym framework, named as HephaestusForge, show that RL algorithms achieve minimal rejection rates (as low as 0.002%, 90x less than the baseline Karmada scheduler). Cost-aware strategies result in lower deployment costs (2.5 units), and latency-aware functions achieve lower latency (268–290 ms), improving by 1.5x and 1.3x, respectively, over the best-performing baselines. HephaestusForge is available in a public open-source repository, allowing researchers to validate their own placement algorithms. This study also highlights the adaptability of the DeepSets (DS) neural network in optimizing microservice placement across diverse multi-cluster setups without retraining. The DS neural network can handle inputs and outputs as arbitrarily sized sets, enabling the RL algorithm to learn a policy not bound to a fixed number of clusters. José Santos 0001, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Nicola Di Cicco, Filip De Turck |
Future Gener. Comput. Syst. | 2 |
| 2024 | Multi-Objective Scheduling and Resource Allocation of Kubernetes Replicas Across the Compute ContinuumabstractOrchestrating microservice applications deployed on a federation of globally distributed Kubernetes clusters is a challenging and multifaceted optimization problem. It is not only computationally hard, but also requires balancing a delicate trade-off between competing performance metrics, such as latency, deployment cost, and service interruption frequency. Classical approaches in the literature merge multiple objectives into a single one via, e.g., linear combinations. However, in practice, it is complex to express a priori a quantitative preference between heterogeneous objectives, let alone with simple linear combinations. This paper adopts a more comprehensive approach leveraging proper Multi-Objective Optimization (MOO), with the goal of producing multiple solutions from the Pareto Front (PF). Therefore, the orchestrator can inspect a posteriori all possible "optimal" trade-offs and decide on the strategy that best fits their operating requirements. To solve the MOO problem, this paper adopts state-of-the-art Multi-Objective Evolutionary Algorithms and shows their effectiveness in solving the MOO problem. Illustrative results highlight the practical benefits of a MOO formulation, providing several tens of nondominated solutions and evenly covering the objectives’ space. Nicola Di Cicco, Filippo Poltronieri, José Santos 0001, Mattia Zaccarini, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
CNSM | 4 |
| 2024 | TELKA: Twin-Enhanced Learning for Kubernetes ApplicationsabstractChaos engineering is the discipline of injecting computing and network faults, such as increased network latency and unavailability of computing nodes, into an IT system to help developers in identifying problems that could arise in a production environment and tackle them. Several tools have emerged to ease the application of chaos engineering to complex IT systems, leveraging microservice and container-based applications deployed on Kubernetes. However, applying of such tools requires several phases to be put into practice, from defining a steady state to establishing an effective response plan if something goes wrong. To ease the application of chaos engineering in improving the resilience of Kubernetes applications, this work presents a smart scheduler for Kubernetes called TELKA: a Twin-Enhanced Learning for Kubernetes Applications, which combines chaos engineering, Digital Twin (DT), and Reinforcement Learning (RL) methodologies to mitigate the effects of computing and network faults. Instead of interacting directly with the physical Kubernetes application, TELKA learns by interacting with a digital twin, thus reducing the learning time and the operation costs related to the application of chaos engineering. Experiment results compare TELKA with other approaches to show its effectiveness in mitigating the adverse effects of injected faults. Mattia Zaccarini, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
ISCC | 1 |
| 2024 | Efficient Microservice Deployment in Kubernetes Multi-Clusters through Reinforcement LearningabstractMicroservices have revolutionized application deployment in popular cloud platforms, offering flexible scheduling of loosely-coupled containers and improving operational efficiency. However, this transition made applications more complex, consisting of tens to hundreds of microservices. Efficient orchestration remains an enormous challenge, especially with emerging paradigms such as Fog Computing and novel use cases as autonomous vehicles. Also, multi-cluster scenarios are still not vastly explored today since most literature focuses mainly on a single-cluster setup. The scheduling problem becomes significantly more challenging since the orchestrator needs to find optimal locations for each microservice while deciding whether instances are deployed altogether or placed into different clusters. This paper studies the multi-cluster orchestration challenge by proposing a Reinforcement Learning (RL)-based approach for efficient microservice deployment in Kubernetes (K8s), a widely adopted container orchestration platform. The study demonstrates the effectiveness of RL agents in achieving near-optimal allocation schemes, emphasizing latency reduction and deployment cost minimization. Additionally, the work highlights the versatility of the DeepSets neural network in optimizing microservice placement across diverse multi-cluster setups without retraining. Results show that DeepSets algorithms optimize the placement of microservices in a multi-cluster setup 32 times higher than its trained scenario. José Santos 0001, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Sleianelli, Nicola Di Cicco, Filip De Turck |
NOMS | 2 |
| 2024 | The Evolution of Kubernetes Management: Introducing the KubeTwin FrameworkabstractThe current trend in the management of complex and highly distributed microservices architecture is to leverage container-based orchestration tools such as Kubernetes. Kubernetes provides several capabilities that help service providers to deal with essential aspects of service provisioning such as service replication, scalability, and cluster federation. However, finding a suitable configuration for a Kubernetes application and optimizing its deployment in complex scenarios like Compute Continuum ones can be a very time-consuming and challenging operation. To simplify this process, this research project introduces the KubeTwin framework, a Digital Twin-based solution to evaluate the behavior and the impact of Kubernetes applications in a shorter computation time. Due to an accurate virtual representation of a Kubernetes application, KubeTwin provides a plethora of functionalities that can help providers to assess their ecosystems by employing different mechanisms that span from performance optimization to chaos engineering. Results shown by the first experimental model evaluations demonstrate the strength of this project and encourage for future major efforts. Mattia Zaccarini, Mauro Tortonesi, Filippo Poltronieri |
NOMS | 1 |
| 2024 | KubeTwin: A Digital Twin Framework for Kubernetes Deployments at ScaleabstractKubernetes is a well-known orchestration and management solution for complex and large-scale service architectures in the Cloud Continuum. While it provides very valuable functions from the operation perspective, the high number of control loops it implements significantly enlarges the already wide space of configuration parameters and policies to consider for management purposes. We argue that optimizing complex Kubernetes deployments considering a multi-cloud and edge computing environment would significantly benefit from a Digital Twin approach, enabling an accurate virtual representation of a Kubernetes application to optimize its deployment and management policies. Towards that goal, this work illustrates the design of KubeTwin, a framework to implement Digital Twins of Kubernetes deployments. Furthermore, we present a validation of KubeTwin in a Multi-access Edge Computing (MEC) scenario, which shows its soundness in reenacting realistic Digital Twins of complex and highly distributed Kubernetes deployments. We believe that KubeTwin can provide useful guidance to the research community working in this field. Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini |
IEEE Trans. Netw. Serv. Manag. | 11 |
| 2023 | Characterization of Microservice Response Time in Kubernetes: A Mixture Density Network ApproachabstractThe use of microservice-based applications is becoming more prominent also in the telecommunication field. The current 5G core network, for instance, is already built around the concept of a “Service Based Architecture”, and it is foreseeable that 6G will push even further this concept to enable more flexible and pervasive deployments. However, the increasing complexity of future networks calls for sophisticated platforms that could help network providers with their deployments design. In this framework, a central research trend is the development of digital twins of the physical infrastructures. These digital representations should closely mimic the behavior of the managed system, allowing the operators to test new configurations, analyze what-if scenarios, or train their reinforcement learning algorithms in safe environments. Considering that Kubernetes is becoming the de-facto standard platform for container orchestration and microservice-based application lifecycle management, the implementation of a Kubernetes digital twin requires an accurate characterization of the microservice response time, possibly leveraging suitable Machine Learning techniques trained with measurement data collected in the field. In this paper we introduce a new methodology, based on Mixture Density Networks, to accurately estimate the statistical distribution of the response time of microservice-based applications. We show the improvement in performance with respect to simulation-based inference procedures proposed in literature. Lorenzo Manca, Davide Borsatti, Filippo Poltronieri, Mattia Zaccarini, Domenico Scotece, Gianluca Davoli, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Walter Cerroni |
CNSM | 4 |
| 2023 | Modeling Digital Twins of Kubernetes-Based ApplicationsabstractKubernetes provides several functions that can help service providers to deal with the management of complex container-based applications. However, most of these functions need a time-consuming and costly customization process to address service-specific requirements. The adoption of Digital Twin (DT) solutions can ease the configuration process by enabling the evaluation of multiple configurations and custom policies by means of simulation-based what-if scenario analysis. To facilitate this process, this paper proposes KubeTwin, a framework to enable the definition and evaluation of DTs of Kubernetes applications. Specifically, this work presents an innovative simulation-based inference approach to define accurate DT models for a Kubernetes environment. We experimentally validate the proposed solution by implementing a DT model of an image recognition application that we tested under different conditions to verify the accuracy of the DT model. The soundness of these results demonstrates the validity of the KubeTwin approach and calls for further investigation. Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini |
ISCC | 10 |