VLDB 2026 Research / reviewers in the wild / expert
Remo Andreoli
dblp:293/6232
· DBLP profile ↗
12ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0002-3268-4289ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DistWalk: A Distributed Workload EmulatorabstractThis paper introduces DistWalk, a flexible, distributed, scalable, and open-source toolkit designed to emulate compute, network, and storage workloads across a networked infrastructure, and measure the resulting end-to-end latency. DistWalk provides fine-grained control over the workload behavior, which consists of a graph-like sequence of operations spanning multiple servers, It supports several communication protocols and traffic patterns, and enables the customization of several factors, such as the duration and parallelism of computeintensive operations, and the I/O data access and synchronization mode, among others. The proposed toolkit may be used to experiment with a variety of deployment models, from bare-metal to virtualized or containerized environments, e.g., using Cloud/Edge infrastructures, OpenStack, Kubernetes, or other orchestrators, allowing for experimental comparisons of the achievable latency across a wide range of system-level configurations. Remo Andreoli, Tommaso Cucinotta |
CCGrid | 1 |
| 2025 | Demo: Emulating Distributed Workloads with DistWalkabstractThis demo showcases DistWalk, an open-source distributed workload emulator designed to study the end-to-end latency implications of Linux-based systems. DistWalk is capable of deploying sequences of compute, network, and storage operations arranged within graph-like topologies, to be carried out across multiple servers. It supports a variety of communication protocols and traffic patterns, and enables the customization of several factors, such as the duration and parallelism of compute-intensive operations, the network security and connection handling strategy, and the I/O data access and synchronization mode, among others. DistWalk can be used to experiment with a variety of deployment models for distributed workloads, from bare-metal to virtualized or containerized environments, e.g., using Cloud/Edge infrastructures, OpenStack, Kubernetes, or other orchestrators. This allows Cloud/Edge researchers and developers to perform experimental comparisons of the latency achievable by distributed workload patterns across a wide range of system-level configurations. Remo Andreoli, Tommaso Burlon, Antonio Napolitano 0003, Tommaso Cucinotta |
IC2E | 1 |
| 2025 | CloudSim 7G: An Integrated Toolkit for Modeling and Simulation of Future Generation Cloud Computing EnvironmentsabstractABSTRACT Background Cloud Computing has established itself as an efficient and cost‐effective paradigm for the execution of web‐based applications, and scientific workloads, that need elasticity and on‐demand scalability capabilities. However, the evaluation of novel resource provisioning and management techniques is a major challenge due to the complexity of large‐scale data centers. Therefore, Cloud simulators are an essential tool for academic and industrial researchers, to investigate the effectiveness of novel algorithms and mechanisms in large‐scale scenarios. Aim This paper proposes CloudSim 7G, the seventh generation of CloudSim, which features a re‐engineered and generalized internal architecture to facilitate the integration of multiple CloudSim extensions within the same simulated environment. Methods As part of the new design, we introduced a set of standardized interfaces to abstract common functionalities and carried out extensive refactoring and refinement of the codebase. Results The result is a substantial reduction in lines of code with no loss in functionality, significant improvements in run‐time performance and memory efficiency (up to 25\% less heap memory allocated), as well as increased flexibility, ease‐of‐use, and extensibility of the framework. Conclusion These improvements benefit not only CloudSim developers but also researchers and practitioners using the framework for modeling and simulating next‐generation Cloud Computing environments. Remo Andreoli, Tommaso Cucinotta, Rajkumar Buyya |
Softw. Pract. Exp. | 1 |
| 2025 | A Multi-Domain Survey on Time-Criticality in Cloud ComputingabstractConventional cloud services and infrastructures are mainly designed to maximize utilization of resources and provide best-effort Quality-of-Service levels. However, many emerging use cases in both public and private cloud computing scenarios are time-critical in nature. For example, automated vehicles, smart cities, and automated factories, are all application domains characterized by the need for highly reliable and consistent low-latency services. The incorporation of predictable execution properties in cloud solutions is essential to meet these requirements. This paper provides an overview of the current research landscape in cloud computing, summarizing the key aspects to enable support of time-critical applications. The paper explores various levels of the typical cloud software stack: machine virtualization and containers, resource management and orchestration, fault tolerance, serverless computing, data storage and management, and communications. Remo Andreoli, Raquel Mini, P. Skarin, Harald Gustafsson, J. Harmatos, Luca Abeni, Tommaso Cucinotta |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | RTilience: Fault-Tolerant Time-Critical KubernetesabstractThis paper tackles the problem of optimal configuration and deployment of fault-tolerant time-critical service chains with arbitrary DAG-alike topologies. We propose RTilience, designed according to a scalable cloud microservice paradigm, and prototyped on top of the well-known Kubernetes cloud orchestrator. It features real-time reservation scheduling of containers to guarantee temporal isolation of time-critical tasks, leading to fine-grained control of compute latencies, while allowing for sharing physical CPUs among containers. A distributed routing library, ReqRoute, is configured with a timeout and primary and secondary routes, enabling autonomous and decentralized handling of failing requests. The routes are configured by a centralized controller that performs admission control, resource management of microservice instances, task placement, and fault detection and recovery, extending the features available in Kubernetes. Admission control is based on a theoretical framework enclosing a worst-case performance model for the experienced end-to-end response-time under various fault handling options, and an optimization framework that computes the optimum resource allocation for admitted services. Extensive experimentation of the proposed solution has been performed with synthetic examples, and an autonomous transport robot use-case, verifying that end-to-end deadlines are effectively respected, even in presence of high fault rates of individual microservice instances, according to the theoretical expectations. RTilience is made available as open-source software, released under a MIT license. Harald Gustafsson, Fredrik Svensson, Raquel Mini, Luca Abeni, Remo Andreoli, Tommaso Cucinotta |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | A Logic Programming Approach to VM PlacementabstractPlacing virtual machines so to minimize the number of used physical hosts is an utterly important problem in cloud computing and next-generation virtualized networks.This article proposes a declarative reasoning methodology, and its open-source prototype, including four heuristic strategies to tackle this problem.Our proposal is extensively assessed over real data from an industrial case study and compared to state-of-the-art approaches, both in terms of execution times and solution optimality.As a result, our declarative approach determines placements that are only 6% far from optimal, outperforming a state-of-the-art genetic algorithm in terms of execution times, and a first-fit search for optimality of found placements.Last, its pipelining with a mathematical programming solution improves execution times of the latter by one order of magnitude on average, compared to using a genetic algorithm as a primer. Remo Andreoli, Stefano Forti 0002, Luigi Pannocchi, Tommaso Cucinotta, Antonio Brogi |
CLOSER | 1 |
| 2023 | Design-Time Analysis of Time-Critical and Fault-Tolerance Constraints in Cloud ServicesabstractThis work presents a model for designing and deploying time-critical, cloud-native applications under fault conditions. Our model considers the interactions and interferences among service components, as well as the possible occurrence of faults. Given a set of to-be-deployed applications with precise temporal constraints and a predefined configuration of the service components, we devised an optimizer to verify at design time if the cloud services guarantee compliance with the timing constraints while minimizing the resources needed to achieve fault tolerance. Remo Andreoli, Harald Gustafsson, Luca Abeni, Raquel Mini, Tommaso Cucinotta |
CLOUD | 1 |
| 2023 | Inducing Huge Tail Latency on a MongoDB deploymentabstractThe NoSQL paradigm has emerged as the leading design choice for cloud providers offering highly scalable storage services. Contrary to traditional relational databases, NoSQL architectures are capable of ingesting the ever-growing volume of nowadays’ data-driven applications characterized by low-latency and high-throughput requirements. However, it is difficult to build an ultra-scalable, high-performance storage engine that can sustain an arbitrary number of concurrent clients. A common technique to increase throughput is minimizing the OS overhead, quantified as the number of context switches, through busy waiting (or "spinning"). While this simple synchronization mechanism proves to be beneficial in the high-performance computing community, it requires special care to avoid wasting resources and counter-intuitive behaviors. In this paper, we address an instance of "unsafe" busy waiting in WiredTiger, the underlying storage engine of MongoDB, which leads to a consistent, excessive increase of tail latency in high contention scenarios. Remo Andreoli, Tommaso Cucinotta |
IC2E | 1 |
| 2023 | Towards a Holistic Cloud System with End-to-End Performance GuaranteesabstractComputing technologies are undergoing a relentless evolution from both the hardware and software sides, incorporating new mechanisms for low-latency networking, virtualization, operating systems, hardware acceleration, smart services orchestration, serverless computing, hybrid private-public Cloud solutions and others. Therefore, Cloud infrastructures are becoming increasingly attractive for deploying a wider and wider range of applications, including those with more and more stringent timing constraints, like the emerging use case of deploying time-critical applications. However, despite the availability of a number of public Cloud offerings, and of products (or open-source suites) for deploying in-house private Cloud infrastructures, still there are no solutions readily available for managing time-critical software components with predictable end-to-end timing requirements in the range of hundreds or even tens of milliseconds. The goal of this discussion is to present the multi-domain challenges associated with orchestrating a holistic Cloud system with end-to-end guarantees, which is the subject of my current PhD investigations. Remo Andreoli, Tommaso Cucinotta |
IC2E | 1 |
| 2023 | Fault Tolerance in Real-Time Cloud ComputingabstractThis paper presents the Fault-Tolerant Real-Time Cloud (FTRTC) project that aims to design cloud computing infrastructures capable of hosting highly reliable and real-time applications. These applications are characterized by strict timing and reliability constraints, as well as critical failure scenarios. For instance, such requirements are commonly found in the context of Industry 4.0. We present a formalization of the problem of designing real-time cloud applications supporting an adjustable level of fault tolerance throughout their distributed execution in a cloud infrastructure. The contributions presented in this paper indicate important research directions when building cloud infrastructures able to supporting ultra-reliable real-time applications. Luca Abeni, Remo Andreoli, Harald Gustafsson, Raquel Mini, Tommaso Cucinotta |
ISORC | 2 |
| 2023 | Priority-Driven Differentiated Performance for NoSQL Database-as-a-ServiceabstractDesigning data stores for native Cloud Computing services brings a number of challenges, especially if the Cloud Provider wants to offer database services capable of controlling the response time for specific customers. These requests may come from heterogeneous data-driven applications with conflicting responsiveness requirements. For instance, a batch processing workload does not require the same level of responsiveness as a time-sensitive one. Their coexistence may interfere with the responsiveness of the time-sensitive workload, such as online video gaming, virtual reality, and cloud-based machine learning. This paper presents a modification to the popular MongoDB NoSQL database to enable differentiated per-user/request performance on a priority basis by leveraging CPU scheduling and synchronization mechanisms available within the Operating System. This is achieved with minimally invasive changes to the source code and without affecting the performance and behavior of the database when the new feature is not in use. The proposed extension has been integrated with the access-control model of MongoDB for secure and controlled access to the new capability. Extensive experimentation with realistic workloads demonstrates how the proposed solution is able to reduce the response times for high-priority users/requests, with respect to lower-priority ones, in scenarios with mixed-priority clients accessing the data store. Remo Andreoli, Tommaso Cucinotta, Daniel Bristot de Oliveira |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | RT-MongoDB: A NoSQL Database with Differentiated PerformanceabstractThe advent of Cloud Computing and Big Data brought several changes and innovations in the landscape of database management systems. Nowadays, a cloud-friendly storage system is required to reliably support data that is in continuous motion and of previously unthinkable magnitude, while guaranteeing high availability and optimal performance to thousands of clients. In particular, NoSQL database services are taking momentum as a key technology thanks to their relaxed requirements with respect to their relational counterparts, that are not designed to scale massively on distributed systems. Most research papers on performance of cloud storage systems propose solutions that aim to achieve the highest possible throughput, while neglecting the problem of controlling the response latency for specific users or queries. The latter research topic is particularly important for distributed real-time applications, where task completion is bounded by precise timing constraints. In this paper, the popular MongoDB NoSQL database software is modified introducing a per-client/request prioritization mechanism within the request processing engine, allowing for a better control of the temporal interference among competing requests with different priorities. Extensive experimentation with synthetic stress workloads demonstrates that the proposed solution is able to assure differentiated per-client/request performance in a shared MongoDB instance. Namely, requests with higher priorities achieve reduced and significantly more stable response times, with respect to lower priorities ones. This constitutes a basic but fundamental brick in providing assured performance to distributed real-time applications making use of NoSQL database services. Remo Andreoli, Tommaso Cucinotta, Dino Pedreschi |
CLOSER | 1 |