VLDB 2026 Research / reviewers in the wild / expert
Mauro Tortonesi
dblp:65/1897
· DBLP profile ↗
64ranked-venue papers
5as first author
33since 2021 · last 2026
0000-0002-7417-4455ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 2 first-author · 9 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the Impact of Scheduling Strategies in Kubernetes with the KubeTwin Platform
José Santos 0001, Davide Borsatti, Walter Cerroni, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
NetSoft | 6 |
| 2026 | KubeTwin 2.0: Demonstrating the Impact of Scheduling Strategies in Kubernetes
José Santos 0001, Davide Borsatti, Walter Cerroni, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
NetSoft | 6 |
| 2026 | Embedding Models for Multivariate Time Series Anomaly Detection in Industry 5.0abstractAbstract Industrial processes often involve the generation—and the analysis—of multivariate time series data, which poses several challenges from the anomaly detection perspective. In addition to the need to detect previously unseen anomalies, the high dimensionality of industrial datasets introduces the complexity of simultaneously analyzing multiple features and their interactions. Finally, industrial datasets are typically highly imbalanced, with minimal information on anomalous processes. To address these issues, we propose a novel anomaly detection framework that introduces two embedding models, based on Time2Vec and Discrete Wavelet Transforms, leveraging their capabilities to represent multivariate time series as vectors while capturing and preserving temporal dependencies and combining them with several classifiers to enhance the overall performance of anomaly detection. We tested our solution using a publicly available benchmark dataset and a real industrial use case, particularly data collected from a Bonfiglioli gear manufacturing plant. The results demonstrate that, unlike traditional reconstruction-based autoencoders, which often struggle with sporadic noise, our embedding-based solutions maintain high performance across various noise conditions. Lorenzo Colombi, Michela Vespa, Nicolas Belletti, Matteo Brina, Simon Dahdal, Filippo Tabanelli, Francesco Resca, Elena Bellodi, Mauro Tortonesi, Cesare Stefanelli, Massimiliano Vignoli |
Data Sci. Eng. | 9 |
| 2026 | Smart and Sustainable Ice Cream Making Through Edge Machine LearningabstractThe manufacturing process of frozen dairy desserts, such as ice cream and gelato, is very sensitive to human errors in ingredient preparation: even minor variations in the ingredient mix can lead to quality issues and material waste. To become more sustainable, next generation ice cream making machines need to implement intelligent and adaptive processes that are both efficient and forgiving of human mistakes in mixture preparations. Toward that goal, we developed Hard-O-Tronic AI-driven (HOT-AI), a novel edge AI solution specifically designed for Carpigiani’s ice cream making machines. Leveraging the innovative multimilestone classification methodology, HOT-AI performs inference at multiple stages—or milestones—during the ice cream making process, with increasing accuracy over time. This enables HOT-AI to take corrective actions by adapting the preparation process accordingly, thus improving batch-to-batch uniformity, minimizing ingredient waste, and enhancing production efficiency, cost-effectiveness, and sustainability. HOT-AI has been successfully validated under real production conditions, and its large-scale implementation is planned across Carpigiani Group machines. Filippo Tabanelli, Simon Dahdal, Nicolas Belletti, Elena Bellodi, Franck Ngatcha, Roberto Lazzarini, Cesare Stefanelli, Mauro Tortonesi |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | Beyond TimeGraph: A Comparative Analysis of Temporal Generators for Evolving Network GraphsabstractThe growing adoption of Artificial Intelligence (AI) in network and service management demands extensive, diverse, and high-fidelity datasets for training and evaluation. However, collecting real-world network data at scale often faces significant challenges, including privacy concerns, operational constraints, and the rarity of certain events or conditions. Generative AI offers a promising solution by synthesizing realistic data that mirrors complex network dynamics and user behavior.In many application domains-such as mobile connectivity, cybersecurity, and disaster recovery-realism is not only defined by accurate replication of structural features (e.g., connectivity graphs), but also by the ability to model how these features evolve over time. Capturing these temporal dynamics is critical to ensure that AI models trained on synthetic data can generalize effectively to real-world scenarios. One effective approach to this challenge is to transform raw graph data into a compact latent representation, which can then be processed by a temporal generative model. This two-stage framework enables the learning of both structural and temporal characteristics of the underlying system, offering a more comprehensive generative pipeline.Building on previous work that employed Time-series Generative Adversarial Networks (TimeGAN) for this purpose, this paper explores an alternative temporal generative model: DoppelGANger. By integrating DoppelGANger into the graph generation pipeline, we aim to assess whether it can more accurately capture the dynamics of evolving graph structures. Furthermore, we introduce a more rigorous and detailed evaluation of the generated data by comparing decoded synthetic graph sequences against their real-world counterparts using distribution-aware and graph-structural metrics. These metrics provide a clearer picture of the quality and fidelity of the generated data, highlighting key differences between the TimeGAN and DoppelGANger approaches. Edoardo Di Caro, Nicolas Belletti, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
CNSM | 4 |
| 2025 | Machine Learning for Performance Optimization in Resource-Constrained EnvironmentsabstractThe rapid growth of digital technology has significantly increased reliance on computer systems across many different fields. This widespread adoption has also led to the deployment of devices in challenging scenarios characterized by limited resources, such as constrained computational capabilities or unreliable communication. These environments require alternative, innovative solutions, often specifically tailored to individual use cases. However, if operational conditions change at runtime, causing systems to function in unforeseen scenarios, there is a considerable risk of performance degradation or inefficient use of already limited and valuable resources. To address these complexities, Machine Learning (ML) methods offer significant potential by enabling system to autonomously adapt to evolving conditions. Techniques including predictive analytics and self-optimization allow systems to intelligently adjust at runtime, dynamically optimizing resource usage and performance. However, the effectiveness of ML-driven optimization is significantly limited by the scarcity of realistic data, particularly in scenarios involving limited resources. To mitigate this challenge, generative artificial intelligence emerges as a pivotal tool, facilitating data augmentation, enhancing the effectiveness of the development and testing processes of software solutions. Edoardo Di Caro, Mauro Tortonesi |
CNSM | 2 |
| 2025 | Investigating Neurosymbolic AI for Intent-based Service ManagementabstractThe increasing complexity and increasing demands of IT applications, especially in federated multi-cluster environments, pose significant challenges for service orchestration. To address these, Zero-Touch Service Management (ZSM) and intent-based management paradigms are gaining traction, allowing users to specify high-level goals rather than low-level configurations. However, current intent-driven approaches often rely on rigid Domain Specific Languages (DSLs) or graphic user interfaces, limiting expressiveness and usability. In this work, we propose a neurosymbolic intent-based platform that leverages Large Language Models (LLMs) for natural language intent ingestion and Answer Set Programming (ASP), a declarative programming paradigm used for solving complex combinatorial problems. The system translates natural language descriptions of microservice requirements into structured policies, enabling explainable service-to-cluster matching across federated Kubernetes environments. We validate our approach through experiments that evaluate both the syntactic correctness and efficiency of various LLMs in intent translation, as well as the computational time of the symbolic placement algorithm. Lorenzo Colombi, Sara Cavicchi, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Pál Varga |
CNSM | 4 |
| 2025 | Management of Safety-Critical AI Services in the Compute ContinuumabstractThe increasing complexity and criticality of AIdriven services across the compute continuum, spanning from edge devices to cloud datacenters, necessitates resilient, intelligent, and explainable management strategies. This work addresses the challenges of deploying and orchestrating safetycritical AI services in dynamic and resource-constrained environments, such as Industry 5.0 and Human Assistance and Disaster Recovery (HADR). We present a suite of complementary solutions, including a cloud-native MLOps platform tailored for SMEs, a semi-supervised federated learning framework (FedEdge-Learn), and novel semantic communication mechanisms that optimize data transmission using LLM-driven embeddings. Furthermore, we introduce an intent-based Zerotouch Service Management (ZSM) architecture, leveraging neurosymbolic AI and collaborative intelligence to automate orchestration, model fine-tuning, and policy reasoning across federated Kubernetes clusters. These efforts pave the way for trustworthy, adaptive, and efficient AI service lifecycle management in environments characterized by disconnections, privacy constraints, and operational unpredictability. Future work focuses on extending the neuro-symbolic approach to support additional tasks, including dynamic node selection, optimized placement across the compute continuum with the goal of improving resilience and interpretability in distributed, resource-constrained environments like Industry 5.0 and HADR, while addressing challenges such as intermittent connectivity and evolving operational conditions. Lorenzo Colombi, Mauro Tortonesi |
CNSM | 2 |
| 2025 | Multi-Cluster MLOps Platform for Industry 5.0abstractMachine Learning (ML) is becoming increasingly vital in various Big Data applications within Industry 5.0, such as predictive maintenance, zero-defect manufacturing, and process or supply chain optimization. However, the dynamic and high-stakes nature of manufacturing necessitates continuous monitoring, periodic reevaluation, and potential retraining of ML models to ensure they remain accurate and aligned with the evolving operational context. This paper presents a Machine Learning Operations (MLOps) platform tailored for Industry 5.0 applications, taking advantage of the whole Compute Continuum, from the edge to the cloud, designed to be scalable and resource-efficient. Built on Kubernetes and the Kubeflow framework, the platform enables seamless management of the entire ML lifecycle, from model development to deployment, across multi-cluster environments. Using Bonfiglioli’s real-world anomaly detection use case, the platform demonstrated its ability to support robust, low-latency ML inference services under varying workloads, even on resource-constrained edge devices. Experimental results confirm the platform’s efficiency, scalability, and practical applicability in addressing zero-defect and zero-waste manufacturing requirements. Lorenzo Colombi, Ion Boleac, Matteo Brina, Simon Dahdal, Mauro Tortonesi, Massimiliano Vignoli, Cesare Stefanelli |
ISCC | 5 |
| 2025 | FedEdge-Learn: a Semi-Supervised Federated Learning Framework for Industry 5.0abstractRecent advancements in Machine Learning (ML) and MLOps for Industry 5.0 have significantly boosted productivity in manufacturing by enabling predictive maintenance and optimizing industrial workflows. However, implementing ML applications in real-world industrial environments presents several challenges, including limited access to labeled data, stringent privacy requirements, and the decentralized nature of industrial data. An effective solution for distributed learning with unlabeled data is essential to address these issues. In this paper, we introduce FedEdge-Learn, a novel Federated Learning framework tailored for Industry 5.0 applications. It focuses on unsupervised K-means clustering enhanced by globally shared data. Our approach safeguards data privacy other than accelerating the onboarding of new machines by leveraging the globally trained model. We validate our framework using both public datasets and real-world industrial data, demonstrating its effectiveness in real-world scenarios. The results show how, with our framework, the K-means algorithm is effective in federated settings, without a significant performance decrease compared to the centralized case. Lorenzo Colombi, Edoardo Di Caro, Simon Dahdal, Filippo Poltronieri, Filippo Tabanelli, Mauro Tortonesi, Cesare Stefanelli, Massimiliano Vignoli |
ISCC | 6 |
| 2025 | TimeGraph: Synthetic Generation of Graph Sequences for Realistic Mobile Connectivity ModelsabstractSoftwarized networking solutions are a key enabler for effective and efficient communications in natural disaster recovery scenarios. However, the design development of reliable and robust softwarization solutions in this context is hampered by the scarcity of reference datasets which accurately capture the real-world behavior – and variability – of those environments. This paper presents a method for synthetic generation of sequences of graphs using state-of-the-art Graph Neural Networks (GNNs) and Time-series Generative Adversarial Networks (TimeGAN). By leveraging available real-world data, the proposed approach generates synthetic datasets that closely replicate the features and connectivity patterns found in actual scenarios. These synthetic datasets not only support the training of AI models but also enable testing and evaluation of solutions across different but similar scenarios. Preliminary results using the Anglova scenario show that our solution accurately captures spatio-temporal behaviour in disrupted networks, making it a powerful tool for developing and validating systems in fields where access to real-world data is limited, enhancing their generalizability and reliability. Edoardo Di Caro, Matteo Brina, Nicolas Belletti, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
NetSoft | 5 |
| 2025 | Chaos Engineering Based Kubernetes Pod Rescheduling Through Deep Sets and Reinforcement LearningabstractKubernetes (K8S) is a widely used orchestration solution that helps manage complex IT applications by providing mechanisms for autoscaling, health checking, cluster formation, and replication, which are essential to deploy and manage the multitude of connected microservices. However, they may suffer in case of unexpected faults which can severely change the underlying computing infrastructure and lead to service outages, highlighting the need for resilient solutions capable of mitigating the adverse effects of faults. To address this, the TELKA sched-uler integrates Chaos Engineering (CE), Reinforcement Learning (RL), and Digital Twin (DT) to reallocate K8S pods evicted due to unexpected faults. While TELKA showed promising results in reallocating evicted pods, its preliminary implementations suffered from scalability issues, as the RL agent could only effectively operate on scenarios with the same number of nodes seen during training. To overcome this limitation, this paper improves TELKA by incorporating a neural network architecture called Deep Sets (DS), which can generalize the operation of TELKA on different numbers of nodes. Experimental results not only demonstrate the validity of the improved TELKA but also show how it can be used to identify good operating conditions. Mattia Zaccarini, Filippo Poltronieri, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 10 |
| 2025 | Hybridized Hot Restart via Reinforcement Learning for Microservice OrchestrationabstractThe Compute Continuum (CC) represents a set of computing resources residing from remote cloud datacenters to dedicated hardware at the edge of the network. Modern applications based on the composition of several microservices can strongly benefit from tools that realize optimal deployment based on the characteristics of each microservice, the availability of computing resources across the CC, the end-users latency, and pricing perspectives. To realize such goals, there is the need for sophisticated solutions capable of efficiently exploring a large space of potential configurations. In this regard, Reinforcement Learning (RL) and Computational Intelligence (CI) techniques represent valuable approaches. However, one of the critical challenges remains the adaption of these deployments to highly dynamic scenarios such as the CC. When the availability of the computing resources changes, there is the need to re-optimize the deployment efficiently. This calls for solutions with reactive or proactive capabilities to deal with the dynamicity of these ecosystems. The work considers a CC-inspired scenario by exploring hybridization techniques that combine CI and RL to devise a hot restart approach for metaheuristics and assess a new deployment solution efficiently. Preliminary results show the soundness of the proposed hybridization methods in handling severe system changes. Mattia Zaccarini, Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 4 |
| 2025 | HephaestusForge: Optimal microservice deployment across the Compute Continuum via Reinforcement LearningabstractWith the advent of containerization technologies, microservices have revolutionized application deployment by converting old monolithic software into a group of loosely coupled containers, aiming to offer greater flexibility and improve operational efficiency. This transition made applications more complex, consisting of tens to hundreds of microservices. Designing effective orchestration mechanisms remains a crucial challenge, especially for emerging distributed cloud paradigms such as the Compute Continuum (CC). Orchestration across multiple clusters is still not extensively explored in the literature since most works consider single-cluster scenarios. In the CC scenario, the orchestrator must decide the optimal locations for each microservice, deciding whether instances are deployed altogether or placed across different clusters, significantly increasing orchestration complexity. This paper addresses orchestration in a containerized CC environment by studying a Reinforcement Learning (RL) approach for efficient microservice deployment in Kubernetes (K8s) clusters, a widely adopted container orchestration platform. This work demonstrates the effectiveness of RL in achieving near-optimal deployment schemes under dynamic conditions, where network latency and resource capacity fluctuate. We extensively evaluate a multi-objective reward function that aims to minimize overall latency, reduce deployment costs, and promote fair distribution of microservice instances, and we compare it against typical heuristic-based approaches. The results from an implemented OpenAI Gym framework, named as HephaestusForge, show that RL algorithms achieve minimal rejection rates (as low as 0.002%, 90x less than the baseline Karmada scheduler). Cost-aware strategies result in lower deployment costs (2.5 units), and latency-aware functions achieve lower latency (268–290 ms), improving by 1.5x and 1.3x, respectively, over the best-performing baselines. HephaestusForge is available in a public open-source repository, allowing researchers to validate their own placement algorithms. This study also highlights the adaptability of the DeepSets (DS) neural network in optimizing microservice placement across diverse multi-cluster setups without retraining. The DS neural network can handle inputs and outputs as arbitrarily sized sets, enabling the RL algorithm to learn a policy not bound to a fixed number of clusters. José Santos 0001, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Nicola Di Cicco, Filip De Turck |
Future Gener. Comput. Syst. | 4 |
| 2025 | RoamML distributed continual learning: Adaptive and flexible data-driven response for disaster recovery operationsabstractIn the aftermath of natural disasters, Human Assistance & Disaster Recovery (HADR) operations have to deal with disrupted communication networks and constrained resources. Such harsh conditions make high-communication-overhead ML approaches — either centralized or distributed — impractical, thus hindering the adoption of AI solutions to implement a critical function for HADR operations: building accurate and up-to-date situational awareness. To address this issue we developed Roaming Machine Learning (RoamML), a novel Distributed Continual Learning Framework designed for HADR operations and based on the premise that moving an ML model is more efficient and robust than either large dataset transfers or frequent model parameter updates. RoamML deploys a mobile AI agent that incrementally train models across network nodes containing yet unprocessed data; at each stop, the agent initiate a local training phase to update its internal ML model parameters. To prioritize the processing of strategically valuable data, RoamML Agents follow a navigation system based upon the concept of Data Gravity, leveraging Multi-Criteria Decision Making techniques to simultaneously consider many objectives for Agent routing optimization, including model learning efficiency and network resource utilization, while seamlessly blending subjective insights from expert judgments with objective metrics derived from quantifiable data to determine each next hop. We conducted extensive experiments to evaluate RoamML, demonstrating the framework’s efficiency to train ML models under highly dynamic, resource-constrained environments. RoamML achieves similar performance to centralized ML training under ideal network conditions and outperforms it in a more realistic scenario with reduced network resources, ultimately saving up to 75% in bandwidth utilization across all experiments. • Human Assistance & Disaster Recovery (HADR) requires accurate situational awareness. • Most distributed ML approaches assume a stable network and are unsuited for HADR. • Approaches based on distributed AI agents and continual learning are more resilient. • Data Gravity represents a solid foundational concept for agent routing optimization. • Data Gravity and MCDM allow to prioritize the processing of critical datasets. Simon Dahdal, Sara Cavicchi, Alessandro Gilli, Filippo Poltronieri, Mauro Tortonesi, Niranjan Suri, Cesare Stefanelli |
J. Netw. Comput. Appl. | 5 |
| 2024 | Multi-Objective Scheduling and Resource Allocation of Kubernetes Replicas Across the Compute ContinuumabstractOrchestrating microservice applications deployed on a federation of globally distributed Kubernetes clusters is a challenging and multifaceted optimization problem. It is not only computationally hard, but also requires balancing a delicate trade-off between competing performance metrics, such as latency, deployment cost, and service interruption frequency. Classical approaches in the literature merge multiple objectives into a single one via, e.g., linear combinations. However, in practice, it is complex to express a priori a quantitative preference between heterogeneous objectives, let alone with simple linear combinations. This paper adopts a more comprehensive approach leveraging proper Multi-Objective Optimization (MOO), with the goal of producing multiple solutions from the Pareto Front (PF). Therefore, the orchestrator can inspect a posteriori all possible "optimal" trade-offs and decide on the strategy that best fits their operating requirements. To solve the MOO problem, this paper adopts state-of-the-art Multi-Objective Evolutionary Algorithms and shows their effectiveness in solving the MOO problem. Illustrative results highlight the practical benefits of a MOO formulation, providing several tens of nondominated solutions and evenly covering the objectives’ space. Nicola Di Cicco, Filippo Poltronieri, José Santos 0001, Mattia Zaccarini, Mauro Tortonesi, Cesare Stefanelli, Filip De Turck |
CNSM | 5 |
| 2024 | RoamML Platform: Enabling Distributed Continual Learning for Disaster Relief OperationsabstractMachine learning offers a promising avenue for improving the efficiency and effectiveness of decision-making in disaster recovery and relief efforts. These operations face significant hurdles due to the large volumes of data, intermittent connectivity, and infrastructure limitations. In this paper, we present the RoamML Platform, a sophisticated modular implementation of the RoamML framework, designed specifically to address these challenges and enable efficient distributed machine learning. We advocate for a foundational principle that "the transmission of the ML model itself is usually more efficient than the costly transfer of large datasets", leading to a more adaptable training regime. The platform orchestrates the activities of the RoamML model along with its related metadata, collectively referred to as the "RoamML Agent", while faithfully observing the Data Gravity principle to guarantee thorough model training. We extensively validated the platform through a simulated disaster recovery scenario employing the Mininet-WiFi emulator. Our results highlight the benefits of integrating the RoamML framework, including enhanced ML performance and significant bandwidth savings. Simon Dahdal, Alessandro Gilli, Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Niranjan Suri |
ISCC | 4 |
| 2024 | TELKA: Twin-Enhanced Learning for Kubernetes ApplicationsabstractChaos engineering is the discipline of injecting computing and network faults, such as increased network latency and unavailability of computing nodes, into an IT system to help developers in identifying problems that could arise in a production environment and tackle them. Several tools have emerged to ease the application of chaos engineering to complex IT systems, leveraging microservice and container-based applications deployed on Kubernetes. However, applying of such tools requires several phases to be put into practice, from defining a steady state to establishing an effective response plan if something goes wrong. To ease the application of chaos engineering in improving the resilience of Kubernetes applications, this work presents a smart scheduler for Kubernetes called TELKA: a Twin-Enhanced Learning for Kubernetes Applications, which combines chaos engineering, Digital Twin (DT), and Reinforcement Learning (RL) methodologies to mitigate the effects of computing and network faults. Instead of interacting directly with the physical Kubernetes application, TELKA learns by interacting with a digital twin, thus reducing the learning time and the operation costs related to the application of chaos engineering. Experiment results compare TELKA with other approaches to show its effectiveness in mitigating the adverse effects of injected faults. Mattia Zaccarini, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
ISCC | 11 |
| 2024 | Efficient Microservice Deployment in Kubernetes Multi-Clusters through Reinforcement LearningabstractMicroservices have revolutionized application deployment in popular cloud platforms, offering flexible scheduling of loosely-coupled containers and improving operational efficiency. However, this transition made applications more complex, consisting of tens to hundreds of microservices. Efficient orchestration remains an enormous challenge, especially with emerging paradigms such as Fog Computing and novel use cases as autonomous vehicles. Also, multi-cluster scenarios are still not vastly explored today since most literature focuses mainly on a single-cluster setup. The scheduling problem becomes significantly more challenging since the orchestrator needs to find optimal locations for each microservice while deciding whether instances are deployed altogether or placed into different clusters. This paper studies the multi-cluster orchestration challenge by proposing a Reinforcement Learning (RL)-based approach for efficient microservice deployment in Kubernetes (K8s), a widely adopted container orchestration platform. The study demonstrates the effectiveness of RL agents in achieving near-optimal allocation schemes, emphasizing latency reduction and deployment cost minimization. Additionally, the work highlights the versatility of the DeepSets neural network in optimizing microservice placement across diverse multi-cluster setups without retraining. Results show that DeepSets algorithms optimize the placement of microservices in a multi-cluster setup 32 times higher than its trained scenario. José Santos 0001, Mattia Zaccarini, Filippo Poltronieri, Mauro Tortonesi, Cesare Sleianelli, Nicola Di Cicco, Filip De Turck |
NOMS | 4 |
| 2024 | A Machine Learning Operations Platform for Streamlined Model Serving in Industry 5.0abstractMachine Learning (ML) plays an increasingly important role in many Big Data applications in Industry 5.0: predictive maintenance, zero defect manufacturing, process and/or supply chain optimization, etc. However, the dynamic and high-stakes nature of the manufacturing environment requires ML models to be maintained through continuous monitoring, periodical reevaluation, and possible retraining to ensure they remain accurate and relevant to the actual context. In addition, to match the desired performance (as well as security and safety) requirements ML models need to be executed in different locations along the edge-to-Cloud continuum (and possibly migrated in case of need), on dedicated serving runtimes that suit the specific needs of the use case. To address these issues, we realized an MLOps platform that is capable of managing ML models through their entire lifecycle and enabling their deployment in different ML serving runtimes. More specifically, the initial experimental evaluation presented in the paper focuses on Bento Yatai and TorchServe serving runtimes. It demonstrates that our platform is capable of effectively running ML models on both runtimes and provides a comparative evaluation at both the quantitative and qualitative levels. Lorenzo Colombi, Alessandro Gilli, Simon Dahdal, Ion Boleac, Mauro Tortonesi, Cesare Stefanelli, Massimiliano Vignoli |
NOMS | 5 |
| 2024 | Enabling Big Data and Machine Learning Applications in High-Stakes EnvironmentsabstractHigh-Stakes Environments, such as industrial settings and natural disaster relief operations, are pivotal areas where decision-making speed and accuracy critically influence safety, reliability, and success. Within these critical contexts, the management of Data and Machine Learning (ML) models lifecycles is essential. Converting immense volumes of data into actionable insights, thereby enhancing the effectiveness of decision-making processes and ensuring more reliable outcomes. Data management ensures the meticulous organization of extensive datasets, facilitating their use for deriving meaningful information. ML enhances this process by employing algorithms capable of learning from data, thereby automating intricate decision-making tasks. These technologies enable real-time analysis that adjusts to evolving conditions, significantly automating complex decision-making. Nevertheless, implementing ML necessitates capabilities for real-time data acquisition, processing, storage, and adaptive training of ML models to meet immediate operational demands, thereby generating actionable knowledge for swift, efficient, and effective decisions. This paper discusses the advancements in my PhD journey, detailing the research methodology, challenges, and opportunities to improve the data-driven service management in these demanding environments. Simon Dahdal, Mauro Tortonesi |
NOMS | 2 |
| 2024 | The Evolution of Kubernetes Management: Introducing the KubeTwin FrameworkabstractThe current trend in the management of complex and highly distributed microservices architecture is to leverage container-based orchestration tools such as Kubernetes. Kubernetes provides several capabilities that help service providers to deal with essential aspects of service provisioning such as service replication, scalability, and cluster federation. However, finding a suitable configuration for a Kubernetes application and optimizing its deployment in complex scenarios like Compute Continuum ones can be a very time-consuming and challenging operation. To simplify this process, this research project introduces the KubeTwin framework, a Digital Twin-based solution to evaluate the behavior and the impact of Kubernetes applications in a shorter computation time. Due to an accurate virtual representation of a Kubernetes application, KubeTwin provides a plethora of functionalities that can help providers to assess their ecosystems by employing different mechanisms that span from performance optimization to chaos engineering. Results shown by the first experimental model evaluations demonstrate the strength of this project and encourage for future major efforts. Mattia Zaccarini, Mauro Tortonesi, Filippo Poltronieri |
NOMS | 2 |
| 2024 | Chaos Engineering for Resilience Assessment of Digital TwinsabstractWithin the Industry 4.0 vision, digital twins (DTs) have gained great attention as a promising approach to improve remote monitoring and control by means of virtual representations of physical objects. However, while DTs are becoming more and more sophisticated and even adopted for mission-critical applications, their resilience assessment has not received the required consideration yet. This article originally proposes chaos engineering to assess and improve the resilience of DTs by testing multiple aspects of industrial environments in a coordinated, automated, and replicable manner. First, the article discusses why and how chaos engineering is promising to improve the resilience of DTs. Then, it identifies and introduces a set of chaos engineering profiles specifically designed to take into account the many aspects an industrial environment is composed of. Finally, it shows the feasibility of assessing the resilience of a proof-of-concept DT through a testbed based on widely-adopted, open-source tools. Mattia Fogli, Carlo Giannelli, Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | KubeTwin: A Digital Twin Framework for Kubernetes Deployments at ScaleabstractKubernetes is a well-known orchestration and management solution for complex and large-scale service architectures in the Cloud Continuum. While it provides very valuable functions from the operation perspective, the high number of control loops it implements significantly enlarges the already wide space of configuration parameters and policies to consider for management purposes. We argue that optimizing complex Kubernetes deployments considering a multi-cloud and edge computing environment would significantly benefit from a Digital Twin approach, enabling an accurate virtual representation of a Kubernetes application to optimize its deployment and management policies. Towards that goal, this work illustrates the design of KubeTwin, a framework to implement Digital Twins of Kubernetes deployments. Furthermore, we present a validation of KubeTwin in a Multi-access Edge Computing (MEC) scenario, which shows its soundness in reenacting realistic Digital Twins of complex and highly distributed Kubernetes deployments. We believe that KubeTwin can provide useful guidance to the research community working in this field. Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini |
IEEE Trans. Netw. Serv. Manag. | 10 |
| 2023 | Characterization of Microservice Response Time in Kubernetes: A Mixture Density Network ApproachabstractThe use of microservice-based applications is becoming more prominent also in the telecommunication field. The current 5G core network, for instance, is already built around the concept of a “Service Based Architecture”, and it is foreseeable that 6G will push even further this concept to enable more flexible and pervasive deployments. However, the increasing complexity of future networks calls for sophisticated platforms that could help network providers with their deployments design. In this framework, a central research trend is the development of digital twins of the physical infrastructures. These digital representations should closely mimic the behavior of the managed system, allowing the operators to test new configurations, analyze what-if scenarios, or train their reinforcement learning algorithms in safe environments. Considering that Kubernetes is becoming the de-facto standard platform for container orchestration and microservice-based application lifecycle management, the implementation of a Kubernetes digital twin requires an accurate characterization of the microservice response time, possibly leveraging suitable Machine Learning techniques trained with measurement data collected in the field. In this paper we introduce a new methodology, based on Mixture Density Networks, to accurately estimate the statistical distribution of the response time of microservice-based applications. We show the improvement in performance with respect to simulation-based inference procedures proposed in literature. Lorenzo Manca, Davide Borsatti, Filippo Poltronieri, Mattia Zaccarini, Domenico Scotece, Gianluca Davoli, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Walter Cerroni |
CNSM | 11 |
| 2023 | Modeling Digital Twins of Kubernetes-Based ApplicationsabstractKubernetes provides several functions that can help service providers to deal with the management of complex container-based applications. However, most of these functions need a time-consuming and costly customization process to address service-specific requirements. The adoption of Digital Twin (DT) solutions can ease the configuration process by enabling the evaluation of multiple configurations and custom policies by means of simulation-based what-if scenario analysis. To facilitate this process, this paper proposes KubeTwin, a framework to enable the definition and evaluation of DTs of Kubernetes applications. Specifically, this work presents an innovative simulation-based inference approach to define accurate DT models for a Kubernetes environment. We experimentally validate the proposed solution by implementing a DT model of an image recognition application that we tested under different conditions to verify the accuracy of the DT model. The soundness of these results demonstrates the validity of the KubeTwin approach and calls for further investigation. Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini |
ISCC | 9 |
| 2023 | Enabling civil-military collaboration for disaster relief operations in smart city environments
Lorenzo Campioni, Filippo Poltronieri, Cesare Stefanelli, Niranjan Suri, Mauro Tortonesi, Konrad S. Wrona |
Future Gener. Comput. Syst. | 5 |
| 2022 | Water 4.0: enabling Smart Water and Environmental Data MeteringabstractSmart metering represents an interesting field, where IoT can bring huge benefits to collect and analyze water and environmental data. However, it presents well-known issues such as the presence of high heterogeneity at the hardware, software, and network layers. This is even amplified by the vendors’ tendency to adopt proprietary solutions and different low-power wireless communication protocols for IoT sensors and metering devices. Realizing an interoperable platform for the collection and analysis of smart Water and environmental data metering thus becomes a complex task, that needs to address multiple requirements at several levels, starting from the collection of data on the field (IoT sensors, smart meters) and its processing on Cloud Computing platforms. The Water 4.0 project, which involves a collaboration between universities and private companies, including Dipietro Group, aims at addressing these challenges. This paper presents the comprehensive smart environmental data metering solution that we realized within Water 4.0, that enables data collection, data-processing, and Over-The-Air (OTA) for IoT devices, with the ultimate goal of reducing water losses and monitoring water quality. Nicola Caldognetto, Luca Pasquali Evangelisti, Filippo Poltronieri, Michele Russo, Cesare Stefanelli, Sara Tenani, Sara Toboli, Mauro Tortonesi |
NOMS | 8 |
| 2022 | Value-of-Information Middleware Solutions for Fog and Edge ComputingabstractFog and Edge Computing aim to deliver low-latency, immersive, and powerful services by processing information close to both devices and users. This is well suited for IoT applications in Smart City, where IoT gateways, Cloudlets, Base Stations, and other computational nodes can process (part of) the data generated by the multitude of IoT sensors directly at the edge of the network. However, the implementation of Fog and Edge Computing is challenging because it requires to deal with a (limited number of) constrained devices, dynamic services’ requirements, and heterogeneous network conditions. Differently from the Cloud, where computational resources are supposed to be unlimited, Fog and Edge services should be capable to adapt to scarce and constrained resources and deal with the deluge of IoT data. To facilitate the adoption of Fog and Edge Computing this work proposes middleware solutions that leverage Value-of-Information (VoI) as interesting criterion to select only the most valuable piece of information for processing and dissemination and to scale computational workload in an automated fashion. Filippo Poltronieri, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 3 |
| 2022 | A Chaos Engineering Approach for Improving the Resiliency of IT Services ConfigurationsabstractTesting the resiliency of complex IT services deployed in hybrid Cloud scenarios is a challenging task that requires expensive and possibly destructive operations. An interesting approach lies in Chaos Engineering, a set of practices to test the resiliency of software systems running in a production environment. However, Chaos Engineering is an expensive practice that requires the setup of complicated operations that further increase the complexity of management operations. To reduce this complexity, Chaos Engineering can benefit from the adoption of non-destructive approaches such as the definition of realistic digital twins. A digital twin is a virtual replica of a real-system on which experimenting with management configurations. This paper embraces this research avenue by extending our previous efforts to integrate Chaos Engineering techniques into an IT services management framework called ChaosTwin. ChaosTwin leverages novel methodologies and tools capable of identifying and promptly react to unexpected failures. Finally, to implement autonomous fault management, ChaosTwin defines scaling and migration policies that can quickly explore for more resilient placements of software components in case of system failures. We believe that ChaosTwin can provide useful guidance to service providers in finding cost-effective service configurations capable of minimizing the negative effects of unpredictable events. Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
NOMS | 2 |
| 2022 | BDMaaS+: Business-Driven and Simulation-Based Optimization of IT Services in the Hybrid CloudabstractThe maturity of heterogeneous and hybrid public Cloud environments enables service providers to deploy there their complex IT services trusting these large and complex infrastructures. At the same time, evaluating the impact of changes at service configuration before and at the runtime is still a very challenging and difficult task. Moreover, a comprehensive performance evaluation of IT service configurations should not be limited just to costs for IT resource acquisition, but also include risk related elements such as Service Level Agreement (SLA) violation penalties and other intangibles. To support IT service providers in this difficult task, we developed Business-Driven Management as a Service Plus (BDMaaS+), a novel decision support tool that can evaluate IT service configuration through simulation with realistic service and network models. By allowing service providers to define expanded operational parameters, BDMaaS+ also enables what-if scenario analysis, thereby opening interesting possibilities at the planning level. Experimental results, collected from our thorough evaluations, demonstrate how a service provider can leverage BDMaaS+ to explore the potential of high-level business SLA changes and data center additions. Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2021 | ChaosTwin: A Chaos Engineering and Digital Twin Approach for the Design of Resilient IT ServicesabstractChaos Engineering represents an interesting software engineering methodology to improve the resilience of a complex IT system operating in a live production environment by injecting simulated faults, observing the system reaction, and devising mitigating solutions. However, Chaos Engineering is an expensive practice with a high setup and operation overhead and it often focuses on the evaluation of the system behavior from a relatively narrow technical perspective instead of a more comprehensive business level one. To enlarge the audience of Chaos Engineering there is the need for novel solutions that can give service providers the tools to deal with the deployment and testing of complex IT services. To fill this gap, this paper presents ChaosTwin, a novel solution exploring an innovative approach to apply Chaos Engineering to a digital-twin, i.e., a virtual representation of a physical object or a system. By creating realistic digital twin of an IT service, injecting faults on the digital twin and evaluating how different service configuration and fault management strategies would perform from a business level perspective, ChaosTwin provides useful guidance to service providers in finding cost-effective service configurations that can minimize the negative effects of unpredictable events. Experimental results, collected from the evaluation of a realistic case study, demonstrate how ChaosTwin is capable of minimizing both the associated costs and the effects of injected Chaos faults. Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli |
CNSM | 2 |
| 2021 | Reinforcement Learning for value-based Placement of Fog Services
Filippo Poltronieri, Mauro Tortonesi, Cesare Stefanelli, Niranjan Suri |
IM | 2 |
| 2020 | A Federated Platform to Support IoT Discovery in Smart Cities and HADR ScenariosabstractSmart Cities are among the most dynamic and rapidly evolving modern environments, driven by the development of new technologies and the fast growth of the Internet of Things (IoT), which enable the acquisition and processing of very large amounts of data.However, accessing IoT assets is proving to be a challenge, as neither formal nor de facto standards to discover connected Things have emerged.Services that provide discovery and access capability for IoT resources are in the rise, but they often adopt service-specific interfaces and authorization mechanisms that hinder the development and maintainability of IoT applications.Low flexibility and interoperability become especially problematic during emergency situations, when responders might need to access resources that normally would not be allowed to access.To address these issues, this paper describes MARGOT, a distributed edge computing platform that supports domain-aware and secure discovery of IoT resources in Smart Cities.Experimental results obtained using MARGOT in an emulated network environment show that our platform can effectively reduce discovery latency and bandwidth consumption under the considered use cases and network conditions. Alessandro Morelli, Lorenzo Campioni, Niccolò Fontana, Niranjan Suri, Mauro Tortonesi |
FedCSIS | 5 |
| 2020 | Value of Information based Optimal Service Fabric Management for Fog ComputingabstractService fabric management in Fog Computing is a challenging task, which has to deal with a complex and resource scarce environment. We argue that approaches leveraging Value-of-Information (VoI) concepts and tools are particularly interesting to support the realization of that objective. This paper describes innovative methodologies and reference models for the service fabric management for Fog Computing applications. First, we formalize the VoI concept and discuss its adoption in Fog Computing environments. Then, we propose a formal model that aims at maximizing the allocation of Fog services from a value-based perspective. To overcome the complexity of this model, we present two possible approaches (simulation-based optimization and a model approximation) and we compare them by adopting Evolutionary Algorithms (EAs) as optimization techniques. Experimental results prove the validity of both models in finding resource allocation solutions capable of minimizing network latency and maximizing the utility for the end-users of Fog Computing services. Finally, we show how the results of the approximated model can be adopted as a first approximated approach for resource management of Fog Computing services. Filippo Poltronieri, Mauro Tortonesi, Alessandro Morelli, Cesare Stefanelli, Niranjan Suri |
NOMS | 2 |
| 2019 | What-if Scenario Analysis for IT Services in Hybrid Cloud Environments with BDMaaS+
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi |
IM | 5 |
| 2019 | Taming the IoT data deluge: An innovative information-centric service model for fog computing applications
Mauro Tortonesi, Marco Govoni, Alessandro Morelli, Giulio Riberto, Cesare Stefanelli, Niranjan Suri |
Future Gener. Comput. Syst. | 1 |
| 2019 | Smart Appliances and RAMI 4.0: Management and Servitization of Ice Cream MachinesabstractThe widespread adoption of information and communication technologies (ICT) is profoundly changing manufacturing. Several Internet-of-Things (IoT) and industry 4.0 solutions deployed in production environments have pushed for standardization efforts, most notably reference architecture model industrie 4.0 (RAMI 4.0), typically focusing on smart factory environments. However, the ICT evolution is also enabling novel smart appliance scenarios, where relatively cheap machines, connected and integrated, are deployed outside the typical industrial environment with a wide range of stakeholders involved. The paper reports about a real world use case composed of more than 12000 ice cream machines connected worldwide and shows how, anticipating the state of the art, the underlying design of the ICT platform presents many interesting similarities with RAMI 4.0. The integration of appliances in a smart value chain enables to develop novel services for different stakeholders, ranging from ice cream manufacturer and maintenance technicians to ice cream shop owners and final consumers. The important synergies with RAMI 4.0 and the extensive on-the-field validation make the proposed solution a compelling reference application, from which to draw useful and generally applicable guidelines for the development of future Industry 4.0 smart appliance platforms. Antonio Corradi, Luca Foschini 0001, Carlo Giannelli, Roberto Lazzarini, Cesare Stefanelli, Mauro Tortonesi, Giovanni Virgilli |
IEEE Trans. Ind. Informatics | 6 |
| 2018 | Service Placement for Hybrid Clouds Environments based on Realistic Network Measurements
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi |
CNSM | 5 |
| 2018 | Business-Driven Service Placement for Highly Dynamic and Distributed Cloud SystemsabstractThe emergence of large-scale Cloud computing environments characterized by dynamic resource pricing schemes enables valuable cost saving opportunities for service providers that could dynamically decide to change the placement of their IT service components in order to reduce their bills. However, that requires new management solutions to dynamically reconfigure IT service components placement, in order to respond to pricing changes and to control and guarantee the high-level business objectives defined by service providers. This paper proposes a novel approach based on Genetic Algorithm (GA) optimization techniques for adaptive business-driven IT service component reconfiguration. Our proposal allows to evaluate the performance of complex IT services deployments over large-scale Cloud systems in a wide range of alternative configurations, by granting prompt transitions to more convenient placements as business values and costs change dynamically. We deeply assessed our framework in a realistic scenario that consists of 2-tier service architectures with real-world pricing schemes. Collected results show the effectiveness and quantify the overhead of our solution. The results also demonstrate the suitability of business-driven IT management techniques for service components placement and reconfiguration in highly dynamic and distributed Cloud systems. Mauro Tortonesi, Luca Foschini 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2017 | Information-Centric Networking in next-generation communications scenarios
Alessandro Morelli, Mauro Tortonesi, Cesare Stefanelli, Niranjan Suri |
J. Netw. Comput. Appl. | 2 |
| 2016 | Data-driven cloud-based IT services performance forecastingabstractModern Cloud computing environments are rapidly evolving, leading to a growing adoption of dynamic pricing for virtual resources and of speedier deployment tools and to the emergence of hybrid Cloud scenarios. These trends suggest the opportunity to investigate a new generation of Cloud-based IT services, capable of adapting to changes in their operating conditions and deployment environment by dynamically realigning their configuration. This calls for new and more sophisticated management tools, that are capable of atomatically evaluating the performance of alternative configurations for Cloud-based IT services and of identifying the one that aligns better to the objectives defined by the business management. Genady Grabarnik, Mauro Tortonesi, Larisa Shwartz |
IEEE BigData | 2 |
| 2016 | SPF: An SDN-based middleware solution to mitigate the IoT information explosionabstractManaging the extremely large volume of information generated by Internet-of-Things (IoT) devices, estimated to be in excess of 400 ZB per year by 2018, is going to be an increasingly relevant issue. Most of the approaches to IoT information management proposed so far, based on the collection of IoT-generated raw data for storage and processing in the Cloud, place a significant burden on both communications and computational resources, and introduce significant latency. IoT applications would instead benefit from new paradigms to enable definition and deployment of dynamic IoT services and facilitate their use of computational resources at the edge of the network for data analysis purposes, and from smart dissemination solutions to deliver the processed information to consumers. This paper presents SPF (as in “Sieve, Process, and Forward”), an SDN solution which extends the reference ONF architecture replacing the Data Plane with an Information Processing and Dissemination Plane. By leveraging programmable information processors deployed at the Internet/IoT edge and disruption tolerant information dissemination solutions, SPF allows to define and manage IoT applications and services and represents a promising architecture for future urban computing applications. Mauro Tortonesi, James Michaelis, Alessandro Morelli, Niranjan Suri, Michael A. Baker |
ISCC | 1 |
| 2016 | Software-defined and value-based information processing and dissemination in IoT applicationsabstractIn the near term, a multitude of IoT applications are expected, each taking advantage of heterogeneous device collections ranging from environmental sensors to smartphones. However, approaches taken in many IoT systems - based on the paradigm of Cloud computing - face challenges of both high latency and network utilization. A clear demand now exists for new paradigms to facilitate IoT application usage of computational resources at the edge of the network for data analysis purposes, as well as smart dissemination solutions to deliver information to consumers. This paper presents SPF (Sieve, Process, and Forward), a Software Defined Networking (SDN) solution for creating and managing IoT applications and services. By leveraging programmable information processors deployed at the Internet/IoT edge, the SDN approach introduced by SPF represents a promising architecture for future urban computing applications. Mauro Tortonesi, James Michaelis, Niranjan Suri, Michael A. Baker |
NOMS | 1 |
| 2015 | Business-driven configuration of IT services in public and hybrid clouds based on performance forecastingabstractModern Cloud computing environments are rapidly evolving, leading to a growing adoption of dynamic pricing for virtual resources and of speedier deployment tools, and to the emergence of hybrid Cloud scenarios. These trends suggest the opportunity to investigate a new generation of Cloud-based IT services, capable of adapting to changes in their operating conditions and deployment environment by dynamically realigning their configuration. This calls for new and more sophisticated management tools, that are capable of evaluating the performance of alternative configurations for Cloud-based IT services and of identifying the one that aligns better to the objectives defined by the business management. This paper presents an optimization tool for Cloud-based IT services, based on queuing theoretic analysis of service workflows and ILP optimization. Genady Grabarnik, Mauro Tortonesi, Larisa Shwartz |
IM | 2 |
| 2015 | Exploring continuous optimization solutions for business-driven IT management problemsabstractBusiness-driven IT management practices often involve the performance optimization of a system according to business criteria. The increased attention to dynamical aspects of the system behavior, the preference for simulative approaches rather than analytical ones, and the increased level of complexity posed by business-driven performance evaluation significantly complicate the optimization of BDIM systems and demand a radical rethinking of methodologies and tools. This raises the opportunity to devise and implement common methodology and tools that could be used for a large class of different BDIM optimization problems. This paper proposes a generic framework for the dynamic and adaptive optimization of BDIM systems, introduces the Open Source ruby-mhl metaheuristics library, and provides an experimental evaluation in the context of a realistic case study. Mauro Tortonesi |
IM | 1 |
| 2014 | Business-driven optimization of component placement for complex services in federated CloudsabstractWith the advent of connected services ecosystems, new generations of services and systems are being conceived, responding to the ever growing demands of the market place. In parallel, the effective adoption of the Cloud computing paradigm is becoming an essential enabler for business enterprises. With such importance placed on the services ecosystem, the design and management of services becomes a key issue both for the providers and the users. One of the main challenges for service providers lies in the complexity of the services, comprising of a multiplicity of technologies and competing and cooperating providers, which is difficult to address through current technology-centric service design approaches, in particular for the deployment infrastructure. The work described in this paper lays a foundation for business driven service design by proposing a business goals driven model of resource allocation in the Cloud. We define a goal/loss/processing cost function for resource allocation that we optimize, while taking into account the dynamic and varying nature of requests load. Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi |
NOMS | 3 |
| 2013 | Synthetic incident generation in the reenactment of IT support organization behavior
Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
IM | 3 |
| 2013 | Adaptive and business-driven service placement in federated Cloud computing environments
Luca Foschini 0001, Mauro Tortonesi |
IM | 2 |
| 2013 | Large-scale E-maintenance: A new frontier for management?
Roberto Lazzarini, Cesare Stefanelli, Mauro Tortonesi |
IM | 3 |
| 2012 | Predicting peer interactions for opportunistic information dissemination protocolsabstractTactical edge networks provide one of the most challenging environments for communications, which significantly complicates the development of efficient and robust information dissemination solutions. In our previous work, we found that exploiting highly mobile nodes, such as Unmanned Air Vehicles, with cyclic mobility patterns, as message ferries can significantly improve the performance of information dissemination solutions. However, our experience demonstrated that robust forecasting mechanisms are essential in order to withstand frequent changes in the mobility patterns of message ferrying nodes. This paper presents an extension of the adaptive node presence forecasting component developed for DisService, a Peer-to-peer information dissemination system, that provides estimates of tolerance and accuracy of node mobility forecasts. We tested the extended forecasting mechanism in a simulation environment and found that it can lead to significant improvements in the timeliness and reliability of information dissemination. Marco Marchini, Mauro Tortonesi, Giacomo Benincasa, Niranjan Suri, Cesare Stefanelli |
ISCC | 2 |
| 2012 | Modeling IT support organizations using multiple-priority queuesabstractAs IT services grow more and more complicated, and their management becomes increasingly challenging, IT support organizations assume an essential role to ensure the delivery of Service Level Objectives. What-if scenario analysis represents a very effective tool for the performance optimization of IT support organizations, as it enables an iterative and customized performance optimization process. The problem of accurately modeling IT support organizations requires the development of sophisticated models as well as dedicated parameter inference techniques and tools. This paper presents a multiple-priority queuing model suited for the reenactment of IT support groups developed from the analysis of empirical evidence, as well as a powerful method to infer the model parameters. We applied our model to reenact the behavior of a real life IT support group with our Symian simulator. The results demonstrate that multiple-priority queuing models can reproduce real life IT support groups with a high degree of accuracy. Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 3 |
| 2012 | Potential benefits and challenges of closed-loop optimization processes for IT support organizationsabstractIT services are getting increasingly complicated, and require IT support organization to manage them. IT support organization are in charge of the incident management process and represent mission-critical structures whose performance needs to be frequently assessed and optimized. State-of-the-art research in the performance optimization of IT support organization proposes user-driven performance assessment and optimization processes based on what-if scenario analysis tools that implement sophisticated IT support organization models. This manuscript instead represents a preliminary study of a different kind of optimization processes, of the closed-loop type, that try to autonomously identify optimal IT support organization configurations according to inputs provided by the user. This paper discusses the development challenges in realizing decision support tools for closed-loop optimization processes and presents a prototype system. The preliminary evaluation of our tool demonstrates that closed-loop processes might be impractical as reference tools but can effectively complement and extend human-driven ones. Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 3 |
| 2012 | A cloud-based solution for the performance improvement of IT support organizationsabstractIT support organizations are in charge of restoring normal operations after IT service disruptions, among other tasks. Such organizations can be complex systems, with a large network of interacting support groups subject to complex management policies. The performance assessment and optimization of IT support organizations is an extremely challenging task that requires considering organization-specific structure, behavior, and business-level objectives. This paper presents Symian-Web, a decision support tool that enables IT managers to both assess and improve IT support organization performance by using what-if scenario analysis. Symian-Web features advanced information visualization concepts and metaphors, allowing for precise and timely assessment of IT support organization performance, and facilitating their redesign. Symian-Web is realized as a cloud computing-based Web application, thus enabling the tool to take advantage of the on-demand computational capabilities provided by the cloud. Claudio Bartolini, Cesare Stefanelli, Davide Targa, Mauro Tortonesi |
NOMS | 4 |
| 2011 | A web-based what-if scenario analysis tool for performance improvement of IT support organizations
Claudio Bartolini, Cesare Stefanelli, Davide Targa, Mauro Tortonesi |
CNSM | 4 |
| 2011 | Teorema: An e-maintenance platform for ice cream machinesabstractIn modern manufacturing, the integration of ICT in the maintenance process, led to the development of e-maintenance, that automates management operations. E-maintenance, that initially interested only large plant machinery, is now becoming affordable for mass-produced equipment, thanks to the recent advances in ICT. This paper presents Teorema, an innovative e-maintenance solution for Carpigiani ice cream machines, which provides several services: remote monitoring of machines, automatic notification of malfunctions, diagnostics and prognostics functions, remote assistance interventions, and automated reporting of production data. Teorema is already in production and it is significantly improving the after-sale service to Carpigiani customers. Roberto Lazzarini, Giovanni Virgilli, Cesare Stefanelli, Mauro Tortonesi |
ETFA | 4 |
| 2011 | On decision making in business-driven IT managementabstractBusiness-driven IT management (BDIM) is a recent research effort to drive IT management decisions from a business perspective by considering business indicators such as profit, cost, and customer experience. BDIM studies complicated decision making processes, dealing with the relationship between the IT function and the business value it generates. The present paper aims at stimulating the discussion on decision making theory and practice within the BDIM research community, in order to develop a better understanding of the theoretical background and consequently improve current tools and practices. To this end, the paper analyzes the most challenging aspects of decision making in BDIM and proposes a few discussion topics that could be of interest for future research studies. Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
Integrated Network Management | 3 |
| 2010 | Modeling IT support organizations from transactional logsabstractThere is great interest in building an accurate theoretical model of IT support organizations, for several purposes such as optimal workforce allocation and what-if scenario analysis. However, the complexity of real-life IT support organizations makes it extremely hard to model their organizational, structural and behavioral processes. While the adoption of stationary stochastic processes to model incident arrivals and of first-come-first-served GI/G/N queues to model support groups permits to reproduce with good enough fidelity the organization-wide behavior, this approach does not always accurately capture the internal dynamics of the organization. This paper presents an experimental analysis of transaction logs from a real-life IT support organization, provided to us by the Outsourcing Services Division of HP. The statistical analysis of transactional logs allows us to make some interesting considerations that can be used to build a more accurate model of the organization, as well as to gain useful experience in the modeling process. Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
NOMS | 3 |
| 2010 | SYMIAN: Analysis and performance improvement of the IT incident management processabstractIncident Management is the process through which IT support organizations manage to restore normal service operation after a service disruption. The complexity of real-life enterprise-class IT support organizations makes it extremely hard to understand the impact of organizational, structural and behavioral components on the performance of the currently adopted incident management strategy and, consequently, which actions could improve it. This paper presents SYMIAN, a decision support tool for the performance improvement of the incident management function in IT support organizations. SYMIAN simulates the effect of corrective measures before their actual implementation, enabling time, effort, and cost saving. To this end, SYMIAN models the IT support organization as an open queuing network, thereby enabling the evaluation of both the system-wide dynamics as well as the behavior of the individual organization components and their interactions. Experimental results show the SYMIAN effectiveness in the performance analysis and tuning of the incident management process for real-life IT support organizations. Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2009 | Business-impact analysis and simulation of critical incidents in IT service managementabstractService disruptions can have a considerable impact on business operations of IT support organizations, thus calling for the implementation of efficient incident management and service restoration processes. The evaluation and improvement of incident management strategies currently in place, in order to minimize the business-impact of major service disruptions, is a very arduous task which goes beyond the optimization with respect to IT-level metrics. This paper presents HANNIBAL, a decision support tool for the business impact analysis and improvement of the incident management process. HANNIBAL evaluates possible strategies for an IT support organization to deal with major service disruptions. HANNIBAL then selects the strategy with the best alignment to the business objectives. Experimental results collected from the HANNIBAL application to a realistic case study show that business impact-driven optimization outperforms traditional performance-driven optimization. Claudio Bartolini, Cesare Stefanelli, Mauro Tortonesi |
Integrated Network Management | 3 |
| 2008 | Session mobility in the mockets communication middlewareabstractTaking advantage of the benefits of modern networking, a growing number of users are exhibiting mobile behavior. As they roam between different network localities, they access the Internet and the Web exploiting both wired and wireless communications and using several heterogeneous devices. Mobile users want to access their subscribed services anywhere, anytime, and want to preserve their currently opened service sessions as they roam between different network localities or switch between different devices. Mobile userspsila requirements call for novel middlewares to provide support for mobility on top of the traditional Internet infrastructure. In this context, we have developed Mockets, a communication middleware specifically designed to address the challenges of wireless networks and mobile computing. In particular, Mockets supports session mobility in terms of seamless handover for preservation of end-to-end connectivity in spite of node mobility, automatic detection and exploitation of best available connectivity, and migration of service session endpoints from one node to another. Cesare Stefanelli, Mauro Tortonesi, Erika Benvegnu, Niranjan Suri |
ISCC | 2 |
| 2008 | QoS management middleware solutions for Bluetooth audio distribution
Paolo Bellavista, Cesare Stefanelli, Mauro Tortonesi |
Pervasive Mob. Comput. | 3 |
| 2004 | The ubiQoS Middleware for Audio Streaming to Bluetooth DevicesabstractThe full and seamless integration of wireless devices with traditional fixed networks is more and more important to foster the mobile and ubiquitous access to the Internet. In particular, the heterogeneity and resource limitations of wireless devices motivate novel support infrastructures that can facilitate the wired-wireless integration and can provide service tailoring depending on client characteristics. The paper presents an application-level portable middleware, called ubiQoS, for QoS-enabled audio streaming to Bluetooth clients. ubiQoS exploits support proxies for QoS tailoring and for managing the QoS over the last segment of the audio distribution path towards the clients, by using different types of Bluetooth links. Proxies execute at the wired-wireless network edges and can even migrate to follow the device movements, where and when needed. The reported experimental results show the feasibility of the application-level approach in the challenging case of QoS-enabled audio streaming to resource-limited Bluetooth devices. Paolo Bellavista, Cesare Stefanelli, Mauro Tortonesi |
MobiQuitous | 3 |
| 2004 | Middleware-Level QoS Differentiation in the Wireless Internet: The UbiQoS Solution for Audio Streaming over BluetoothabstractThe ultimate goal of mobile and ubiquitous Internet accessibility is not only the seamless integration of wireless devices with traditional fixed networks but also the dynamic differentiation of quality of service (QoS) levels depending on client characteristics. In this context, the paper presents the provisioning of audio streaming with different QoS levels in the application-level ubiQoS middleware. In particular, it focuses on how ubiQoS manages the QoS over the last segment of the audio distribution path towards Bluetooth clients by allocating different types of Bluetooth communication channels (unicast connection-oriented or broadcast connectionless) depending on the differentiated QoS requirements of different user classes. To this purpose, we have developed a library that extends the JSR82 standard with the support of active slave broadcast, thus simplifying the Java-based management of Bluetooth communications. The reported experimental results show the feasibility of our application-level middleware approach in the challenging case of audio streaming with differentiated QoS to resource-limited Bluetooth devices. Paolo Bellavista, Cesare Stefanelli, Mauro Tortonesi |
QSHINE | 3 |