Remo Andreoli

dblp:293/6232 · DBLP profile ↗
← Back
12ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0002-3268-4289ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 DistWalk: A Distributed Workload Emulator
abstract
This paper introduces DistWalk, a flexible, distributed, scalable, and open-source toolkit designed to emulate compute, network, and storage workloads across a networked infrastructure, and measure the resulting end-to-end latency. DistWalk provides fine-grained control over the workload behavior, which consists of a graph-like sequence of operations spanning multiple servers, It supports several communication protocols and traffic patterns, and enables the customization of several factors, such as the duration and parallelism of computeintensive operations, and the I/O data access and synchronization mode, among others. The proposed toolkit may be used to experiment with a variety of deployment models, from bare-metal to virtualized or containerized environments, e.g., using Cloud/Edge infrastructures, OpenStack, Kubernetes, or other orchestrators, allowing for experimental comparisons of the achievable latency across a wide range of system-level configurations.
Remo Andreoli, Tommaso Cucinotta
CCGrid1
2025 Demo: Emulating Distributed Workloads with DistWalk
abstract
This demo showcases DistWalk, an open-source distributed workload emulator designed to study the end-to-end latency implications of Linux-based systems. DistWalk is capable of deploying sequences of compute, network, and storage operations arranged within graph-like topologies, to be carried out across multiple servers. It supports a variety of communication protocols and traffic patterns, and enables the customization of several factors, such as the duration and parallelism of compute-intensive operations, the network security and connection handling strategy, and the I/O data access and synchronization mode, among others. DistWalk can be used to experiment with a variety of deployment models for distributed workloads, from bare-metal to virtualized or containerized environments, e.g., using Cloud/Edge infrastructures, OpenStack, Kubernetes, or other orchestrators. This allows Cloud/Edge researchers and developers to perform experimental comparisons of the latency achievable by distributed workload patterns across a wide range of system-level configurations.
Remo Andreoli, Tommaso Burlon, Antonio Napolitano 0003, Tommaso Cucinotta
IC2E1
2025 CloudSim 7G: An Integrated Toolkit for Modeling and Simulation of Future Generation Cloud Computing Environments
abstract
ABSTRACT Background Cloud Computing has established itself as an efficient and cost‐effective paradigm for the execution of web‐based applications, and scientific workloads, that need elasticity and on‐demand scalability capabilities. However, the evaluation of novel resource provisioning and management techniques is a major challenge due to the complexity of large‐scale data centers. Therefore, Cloud simulators are an essential tool for academic and industrial researchers, to investigate the effectiveness of novel algorithms and mechanisms in large‐scale scenarios. Aim This paper proposes CloudSim 7G, the seventh generation of CloudSim, which features a re‐engineered and generalized internal architecture to facilitate the integration of multiple CloudSim extensions within the same simulated environment. Methods As part of the new design, we introduced a set of standardized interfaces to abstract common functionalities and carried out extensive refactoring and refinement of the codebase. Results The result is a substantial reduction in lines of code with no loss in functionality, significant improvements in run‐time performance and memory efficiency (up to 25\% less heap memory allocated), as well as increased flexibility, ease‐of‐use, and extensibility of the framework. Conclusion These improvements benefit not only CloudSim developers but also researchers and practitioners using the framework for modeling and simulating next‐generation Cloud Computing environments.
Remo Andreoli, Tommaso Cucinotta, Rajkumar Buyya
Softw. Pract. Exp.1
2025 A Multi-Domain Survey on Time-Criticality in Cloud Computing
abstract
Conventional cloud services and infrastructures are mainly designed to maximize utilization of resources and provide best-effort Quality-of-Service levels. However, many emerging use cases in both public and private cloud computing scenarios are time-critical in nature. For example, automated vehicles, smart cities, and automated factories, are all application domains characterized by the need for highly reliable and consistent low-latency services. The incorporation of predictable execution properties in cloud solutions is essential to meet these requirements. This paper provides an overview of the current research landscape in cloud computing, summarizing the key aspects to enable support of time-critical applications. The paper explores various levels of the typical cloud software stack: machine virtualization and containers, resource management and orchestration, fault tolerance, serverless computing, data storage and management, and communications.
Remo Andreoli, Raquel Mini, P. Skarin, Harald Gustafsson, J. Harmatos, Luca Abeni, Tommaso Cucinotta
IEEE Trans. Serv. Comput.1
2025 RTilience: Fault-Tolerant Time-Critical Kubernetes
abstract
This paper tackles the problem of optimal configuration and deployment of fault-tolerant time-critical service chains with arbitrary DAG-alike topologies. We propose RTilience, designed according to a scalable cloud microservice paradigm, and prototyped on top of the well-known Kubernetes cloud orchestrator. It features real-time reservation scheduling of containers to guarantee temporal isolation of time-critical tasks, leading to fine-grained control of compute latencies, while allowing for sharing physical CPUs among containers. A distributed routing library, ReqRoute, is configured with a timeout and primary and secondary routes, enabling autonomous and decentralized handling of failing requests. The routes are configured by a centralized controller that performs admission control, resource management of microservice instances, task placement, and fault detection and recovery, extending the features available in Kubernetes. Admission control is based on a theoretical framework enclosing a worst-case performance model for the experienced end-to-end response-time under various fault handling options, and an optimization framework that computes the optimum resource allocation for admitted services. Extensive experimentation of the proposed solution has been performed with synthetic examples, and an autonomous transport robot use-case, verifying that end-to-end deadlines are effectively respected, even in presence of high fault rates of individual microservice instances, according to the theoretical expectations. RTilience is made available as open-source software, released under a MIT license.
Harald Gustafsson, Fredrik Svensson, Raquel Mini, Luca Abeni, Remo Andreoli, Tommaso Cucinotta
IEEE Trans. Serv. Comput.5
2024 A Logic Programming Approach to VM Placement
abstract
Placing virtual machines so to minimize the number of used physical hosts is an utterly important problem in cloud computing and next-generation virtualized networks.This article proposes a declarative reasoning methodology, and its open-source prototype, including four heuristic strategies to tackle this problem.Our proposal is extensively assessed over real data from an industrial case study and compared to state-of-the-art approaches, both in terms of execution times and solution optimality.As a result, our declarative approach determines placements that are only 6% far from optimal, outperforming a state-of-the-art genetic algorithm in terms of execution times, and a first-fit search for optimality of found placements.Last, its pipelining with a mathematical programming solution improves execution times of the latter by one order of magnitude on average, compared to using a genetic algorithm as a primer.
Remo Andreoli, Stefano Forti 0002, Luigi Pannocchi, Tommaso Cucinotta, Antonio Brogi
CLOSER1
2023 Design-Time Analysis of Time-Critical and Fault-Tolerance Constraints in Cloud Services
abstract
This work presents a model for designing and deploying time-critical, cloud-native applications under fault conditions. Our model considers the interactions and interferences among service components, as well as the possible occurrence of faults. Given a set of to-be-deployed applications with precise temporal constraints and a predefined configuration of the service components, we devised an optimizer to verify at design time if the cloud services guarantee compliance with the timing constraints while minimizing the resources needed to achieve fault tolerance.
Remo Andreoli, Harald Gustafsson, Luca Abeni, Raquel Mini, Tommaso Cucinotta
CLOUD1
2023 Inducing Huge Tail Latency on a MongoDB deployment
abstract
The NoSQL paradigm has emerged as the leading design choice for cloud providers offering highly scalable storage services. Contrary to traditional relational databases, NoSQL architectures are capable of ingesting the ever-growing volume of nowadays’ data-driven applications characterized by low-latency and high-throughput requirements. However, it is difficult to build an ultra-scalable, high-performance storage engine that can sustain an arbitrary number of concurrent clients. A common technique to increase throughput is minimizing the OS overhead, quantified as the number of context switches, through busy waiting (or "spinning"). While this simple synchronization mechanism proves to be beneficial in the high-performance computing community, it requires special care to avoid wasting resources and counter-intuitive behaviors. In this paper, we address an instance of "unsafe" busy waiting in WiredTiger, the underlying storage engine of MongoDB, which leads to a consistent, excessive increase of tail latency in high contention scenarios.
Remo Andreoli, Tommaso Cucinotta
IC2E1
2023 Towards a Holistic Cloud System with End-to-End Performance Guarantees
abstract
Computing technologies are undergoing a relentless evolution from both the hardware and software sides, incorporating new mechanisms for low-latency networking, virtualization, operating systems, hardware acceleration, smart services orchestration, serverless computing, hybrid private-public Cloud solutions and others. Therefore, Cloud infrastructures are becoming increasingly attractive for deploying a wider and wider range of applications, including those with more and more stringent timing constraints, like the emerging use case of deploying time-critical applications. However, despite the availability of a number of public Cloud offerings, and of products (or open-source suites) for deploying in-house private Cloud infrastructures, still there are no solutions readily available for managing time-critical software components with predictable end-to-end timing requirements in the range of hundreds or even tens of milliseconds. The goal of this discussion is to present the multi-domain challenges associated with orchestrating a holistic Cloud system with end-to-end guarantees, which is the subject of my current PhD investigations.
Remo Andreoli, Tommaso Cucinotta
IC2E1
2023 Fault Tolerance in Real-Time Cloud Computing
abstract
This paper presents the Fault-Tolerant Real-Time Cloud (FTRTC) project that aims to design cloud computing infrastructures capable of hosting highly reliable and real-time applications. These applications are characterized by strict timing and reliability constraints, as well as critical failure scenarios. For instance, such requirements are commonly found in the context of Industry 4.0. We present a formalization of the problem of designing real-time cloud applications supporting an adjustable level of fault tolerance throughout their distributed execution in a cloud infrastructure. The contributions presented in this paper indicate important research directions when building cloud infrastructures able to supporting ultra-reliable real-time applications.
Luca Abeni, Remo Andreoli, Harald Gustafsson, Raquel Mini, Tommaso Cucinotta
ISORC2
2023 Priority-Driven Differentiated Performance for NoSQL Database-as-a-Service
abstract
Designing data stores for native Cloud Computing services brings a number of challenges, especially if the Cloud Provider wants to offer database services capable of controlling the response time for specific customers. These requests may come from heterogeneous data-driven applications with conflicting responsiveness requirements. For instance, a batch processing workload does not require the same level of responsiveness as a time-sensitive one. Their coexistence may interfere with the responsiveness of the time-sensitive workload, such as online video gaming, virtual reality, and cloud-based machine learning. This paper presents a modification to the popular MongoDB NoSQL database to enable differentiated per-user/request performance on a priority basis by leveraging CPU scheduling and synchronization mechanisms available within the Operating System. This is achieved with minimally invasive changes to the source code and without affecting the performance and behavior of the database when the new feature is not in use. The proposed extension has been integrated with the access-control model of MongoDB for secure and controlled access to the new capability. Extensive experimentation with realistic workloads demonstrates how the proposed solution is able to reduce the response times for high-priority users/requests, with respect to lower-priority ones, in scenarios with mixed-priority clients accessing the data store.
Remo Andreoli, Tommaso Cucinotta, Daniel Bristot de Oliveira
IEEE Trans. Cloud Comput.1
2021 RT-MongoDB: A NoSQL Database with Differentiated Performance
abstract
The advent of Cloud Computing and Big Data brought several changes and innovations in the landscape of database management systems. Nowadays, a cloud-friendly storage system is required to reliably support data that is in continuous motion and of previously unthinkable magnitude, while guaranteeing high availability and optimal performance to thousands of clients. In particular, NoSQL database services are taking momentum as a key technology thanks to their relaxed requirements with respect to their relational counterparts, that are not designed to scale massively on distributed systems. Most research papers on performance of cloud storage systems propose solutions that aim to achieve the highest possible throughput, while neglecting the problem of controlling the response latency for specific users or queries. The latter research topic is particularly important for distributed real-time applications, where task completion is bounded by precise timing constraints. In this paper, the popular MongoDB NoSQL database software is modified introducing a per-client/request prioritization mechanism within the request processing engine, allowing for a better control of the temporal interference among competing requests with different priorities. Extensive experimentation with synthetic stress workloads demonstrates that the proposed solution is able to assure differentiated per-client/request performance in a shared MongoDB instance. Namely, requests with higher priorities achieve reduced and significantly more stable response times, with respect to lower priorities ones. This constitutes a basic but fundamental brick in providing assured performance to distributed real-time applications making use of NoSQL database services.
Remo Andreoli, Tommaso Cucinotta, Dino Pedreschi
CLOSER1