VLDB 2026 Research / reviewers in the wild / expert
Roberto Rocco
dblp:291/4564
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-0223-2900ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications
Marco Cipollini, Simone Rizzo, Sergio Iserte, Paolo Viviani 0001, Giacomo Vitali, Matteo Barbieri, Gabriella Bettonte, Elisabetta Boella, Fulvio Ganz, Roberto Rocco, Orazio Spina, Antonio J. Peña, Petter Sandås, Iacopo Colonnelli, Alberto Scionti, Chiara Vercellino, Emanuele Dri, Jonathan Frassineti, Sara Marzella, Andrea Muratori, Daniele Ottaviani, Olivier Terzo, Bartolomeo Montrucchio, Daniele Gregori |
Future Gener. Comput. Syst. | 10 |
| 2025 | Efficient parameter tuning for a structure-based virtual screening HPC applicationabstractVirtual screening applications are highly parameterized to optimize the balance between quality and execution performance. While output quality is critical, the entire screening process must be completed within a reasonable time. In fact, a slight reduction in output accuracy may be acceptable when dealing with large datasets. Finding the optimal quality-throughput trade-off depends on the specific HPC system used and should be re-evaluated with each new deployment or significant code update. This paper presents two parallel autotuning techniques for constrained optimization in distributed High-Performance Computing (HPC) environments. These techniques extend sequential Bayesian Optimization (BO) with two parallel asynchronous approaches, and they integrate predictions from Machine Learning (ML) models to help comply with constraints. Our target application is LiGen, a real-world virtual screening software for drug discovery. The proposed methods address two relevant challenges: efficient exploration of the parameter space and performance measurement using domain-specific metrics and procedures. We conduct an experimental campaign comparing the two methods with a popular state-of-the-art autotuner. Results show that our methods find configurations that are, on average, up to 35–42% better than the ones found by the autotuner and the default expert-picked LiGen configuration. • We propose two parallel algorithms for black-box constrained optimization. • We integrate Bayesian Optimization and Machine Learning for constraint estimation. • We use a meta-scheduler to hinge on the resources available in HPC settings. • We evaluate the benefits of the proposed approach using a relevant case study. Bruno Guindani, Davide Gadioli, Roberto Rocco, Danilo Ardagna, Gianluca Palermo |
J. Parallel Distributed Comput. | 3 |
| 2025 | To repair or not to repair: Assessing fault resilience in MPI stencil applications
Roberto Rocco, Elisabetta Boella, Daniele Gregori, Gianluca Palermo |
J. Parallel Distributed Comput. | 1 |
| 2024 | A System Development Kit for Big Data Applications on FPGA-based Clusters: The EVEREST ApproachabstractModern big data workflows are characterized by computationally intensive kernels. The simulated results are often combined with knowledge extracted from AI models to ultimately support decision-making. These energy-hungry workflows are increasingly executed in data centers with energy-efficient hard-ware accelerators since FPG As are well-suited for this task due to their inherent parallelism. We present the H2020 project EVEREST, which has developed a system development kit (SDK) to simplify the creation of FPGA-accelerated kernels and manage the execution at runtime through a virtualization environment. This paper describes the main components of the EVEREST SDK and the benefits that can be achieved in our use cases. Christian Pilato, Subhadeep Banik, Jakub Beránek, Fabien Brocheton, Jerónimo Castrillón, Riccardo Cevasco, Radim Cmar, Serena Curzel, Fabrizio Ferrandi, Karl F. A. Friebel, Antonella Galizia, Matteo Grasso, Paulo Silva 0002, Jan Martinovic, Gianluca Palermo, Michele Paolino, Andrea Parodi, Antonio Parodi, Fabio Pintus, Raphael Polig, David Poulet, Francesco Regazzoni 0001, Burkhard Ringlein, Roberto Rocco, Katerina Slaninová, Tom Slooff, Stephanie Soldavini, Felix Suchert, Mattia Tibaldi, Beat Weiss, Christoph Hagleitner |
DATE | 24 |
| 2024 | Extending the Legio Resilience Framework to Handle Critical Process Failures in MPIabstractThe presence of faults in distributed executions can compromise the production of results without proper fault management techniques. The current de-facto standard for inter-process communication, MPI, lacks these features, precluding its effectiveness at a massive scale. Previous efforts produced all-in-one frameworks for fault management, but most of these works leverage checkpoint and restart functionalities, impacting the performance and scalability of the executions. Unlike those, the Legio framework adopts a graceful degradation solution, sac-rificing result accuracy for faster recovery and lower overhead. Still, it cannot handle all the faults: there may be some critical processes whose failure irremediably compromises the result of the computation. With this work, we extend the Legio framework to support critical process faults, combining the benefits of checkpoint and restart solutions with a graceful degradation approach. The experimental campaign shows that our extension does not introduce significant overheads in the executions while correctly managing the failure of critical processes. Roberto Rocco, Luca Repetti, Elisabetta Boella, Daniele Gregori, Gianluca Palermo |
PDP | 1 |
| 2024 | Integrating Bayesian Optimization and Machine Learning for the Optimal Configuration of Cloud SystemsabstractBayesian Optimization (BO) is an efficient method for finding optimal cloud configurations for several types of applications. On the other hand, Machine Learning (ML) can provide helpful knowledge about the application at hand thanks to its predicting capabilities. This work proposes a general approach based on BO, which integrates elements from ML techniques in multiple ways, to find an optimal configuration of recurring jobs running in public and private cloud environments, possibly subject to black-box constraints, e.g., application execution time or accuracy. We test our approach by considering several use cases, including edge computing, scientific computing, and Big Data applications. Results show that our solution outperforms other state-of-the-art black-box techniques, including classical autotuning and BO- and ML-based algorithms, reducing the number of unfeasible executions and corresponding costs up to 2–4 times. Bruno Guindani, Danilo Ardagna, Alessandra Guglielmi, Roberto Rocco, Gianluca Palermo |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | The Legio Fault Resilience Framework: Design and RationaleabstractThe increasing size of HPC clusters makes fault management mandatory. The current MPI standard does not specify the behaviour after the incurrence of a fault, precluding any possible solution. In this work, we present Legio, a framework leveraging the ULFM extension functionalities to introduce fault resilience properties in MPI applications. Roberto Rocco, Gianluca Palermo |
CF | 1 |
| 2023 | Fault Awareness in the MPI 4.0 Session ModelabstractMPI version 4.0 introduces new functionalities like the Session model but still lacks fault management mechanisms. Past efforts produced tools and MPI standard extensions to manage fault presence, including User Level Fault Mitigation (ULFM). These measures are effective against faults but do not fully support the new additions to the standard. In this paper, we combine the fault management possibilities of ULFM with the new Session model functionality introduced in version 4.0 of the standard. We focus on the communicator creation procedure, highlighting criticalities and proposing a method to circumvent them. The experimental campaign shows that the proposed solution does not significantly affect execution times and scalability while better managing the arise of faults. Roberto Rocco, Gianluca Palermo, Daniele Gregori |
CF | 1 |
| 2023 | Fault-Aware Group-Collective Communication Creation and Repair in MPI
Roberto Rocco, Gianluca Palermo |
Euro-Par | 1 |
| 2022 | Legio: fault resiliency for embarrassingly parallel MPI applicationsabstractAbstract Due to the increasing size of HPC machines, dealing with faults is becoming mandatory due to their high frequency. Natively, MPI cannot handle faults and it stops the execution prematurely when it finds one. With the introduction of ULFM, it is possible to continue the execution, but it requires complex integration with the application. In this paper we propose Legio, a framework that introduces fault resiliency in embarrassingly parallel MPI applications. Legio exposes its features to the application transparently, removing any integration difficulty. After a fault, the execution continues only with the non-failed processes. We also propose a hierarchical alternative, which features lower repair costs on large communicators. We evaluated our solutions on the Marconi100 cluster at CINECA with benchmarks and real-world applications, showing that the overhead introduced by the library is negligible and it does not limit the scalability properties of MPI. Roberto Rocco, Davide Gadioli, Gianluca Palermo |
J. Supercomput. | 1 |