Iacopo Colonnelli

dblp:237/8287 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-9290-2017ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-author · 14 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A comprehensive performance evaluation of TEEs for confidential DNA alignment
Lorenzo Brescia, Iacopo Colonnelli, Robert Birke, Valerio Schiavoni, Pascal Felber, Marco Aldinucci
Future Gener. Comput. Syst.2
2026 Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications
Marco Cipollini, Simone Rizzo, Sergio Iserte, Paolo Viviani 0001, Giacomo Vitali, Matteo Barbieri, Gabriella Bettonte, Elisabetta Boella, Fulvio Ganz, Roberto Rocco, Orazio Spina, Antonio J. Peña, Petter Sandås, Iacopo Colonnelli, Alberto Scionti, Chiara Vercellino, Emanuele Dri, Jonathan Frassineti, Sara Marzella, Andrea Muratori, Daniele Ottaviani, Olivier Terzo, Bartolomeo Montrucchio, Daniele Gregori
Future Gener. Comput. Syst.14
2026 A formal framework for fault tolerance in hybrid scientific workflows
abstract
In large-scale distributed systems, failures are routine events whose occurrences increase with the number of computational tasks and execution locations. The advantage of representing an application as a workflow is the possibility of exploiting Workflow Management System (WMS) features such as portability, scalability, and, crucially, reliability. Among these, reliability is essential for ensuring robust execution in dynamic and failure-prone environments. In recent years, the emergence of hybrid workflows has posed new and intriguing challenges by increasing the possibility of distributing computations involving heterogeneous and independent environments. Consequently, the number of possible points of failure during the execution increased, creating a need for sophisticated fault tolerance mechanisms capable of addressing the specific requirements of hybrid systems. This work introduces a formal framework for a fault tolerance mechanism in hybrid workflows, enabling failure recovery through a rollback approach. The framework is rigorously defined by adapting and extending an existing workflow semantics tailored for hybrid execution. Our method leverages provenance data from workflow execution up to the point of failure, and creates a recovery workflow that spans multiple infrastructures. The rollback approach provides a robust and reliable strategy to ensure resilience against step failures and potential data loss. We then implement this mechanism in the StreamFlow WMS, and evaluate it using two case studies: the 1000 Genomes workflow and a synthetic workflow featuring iterative patterns. Experiments showcase the conceptual validity of our approach and assess the overhead introduced by the mechanism, including data availability checks.
Alberto Mulone, Doriana Medic, Iacopo Colonnelli, Marco Aldinucci
Future Gener. Comput. Syst.3
2026 Dynamic transparent streaming in file-based workflows with CAPIO
Marco Edoardo Santimaria, Iacopo Colonnelli, Barbara Cantalupo, Massimo Torquati, Doriana Medic, Nicola Tuccari, Eva Sciacca, Marco Aldinucci
Future Gener. Comput. Syst.2
2026 A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva
Future Gener. Comput. Syst.7
2025 Towards RISC-V-based HPC: The Italian Pathfinding Activities in the DARE-SGA1 Project
abstract
The European Union’s efforts towards technological sovereignty in High-Performance Computing are driving research and development of RISC-V-based supercomputers. The DARE SGA1 project, in particular, aims to develop chips designed and owned by Europeans. This paper introduces the Italian contribution to DARE SGA1 regarding pathfinding activities toward future RISC-V-based accelerator designs, reliability improvements, system software, and AI and Quantum Chemistry applications.
Giovanni Agosta, Marco Aldinucci, Andrea Bartolini, Laura Bellentani, Andrea Biagioni, Daniele Cesarini, Carlotta Chiarini, Iacopo Colonnelli, Pietro Delugas, Lev Denisov, Ottorino Frezza, Marco Grangetto, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Andrea Maslov, Mauro Olivieri, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Davide Rossi 0001, Sergio Saponara, Antonio Sciarappa, Francesco Simula, Matteo Sonza Reorda, Massimo Torquati, Piero Vicini
DSD8
2025 BookedSlurm: meeting user needs for advanced resource reservations in Slurm
abstract
Modern scientific discovery is frequently backed by largescale scientific experiments, which cannot prescind from the unparalleled computational capabilities provided by highperformance computing systems. However, large data centers usually prioritize system efficiency over user accessibility, posing challenges for researchers without advanced computer science expertise. This work introduces BookedSlurm, a secure and userfocused extension of the Slurm workload manager, aiming to democratize HPC access across interdisciplinary research domains. BookedSlurm enables a partially decentralized regulation of finegrained advanced resource reservations through a novel creditbased framework, ensuring fair and predictable access to computing resources. Its modular architecture leverages dedicated microservices to manage reservations, credit handling, and accounting. These components are exposed through a secure REST API and an intuitive webbased dashboard, enhancing system usability for novice and expert users. While the dashboard simplifies interactions for nonspecialists, advanced users can directly access agentlevel APIs for more complex and automated operations. The effectiveness of BookedSlurm is validated on a realworld bioinformatics use case from the SUSMIRRI.IT project, showcasing its ability to enhance usability, optimize job scheduling, and streamline execution workflows.
Sandro Gepiro Contaldo, Lorenzo Bosio, Janneth Estefania Hoyos Rea, Elisa Li Perottino, Sergio Rabellino, Marco Aldinucci, Marco Beccuti, Iacopo Colonnelli
eScience8
2025 The Cloud-HPC infrastructure for Hazard Mapping and vulnerability Monitoring (HaMMon)
abstract
The HaMMon project is the outcome of an industrial partnership that includes many Italian research institutions and private companies. It is led by UnipolSai and Leitha, and funded by the ICSC, the Italian National Research Center for High Performance Computing, Big Data and Quantum Computing.The ambition of HaMMon is to build a flexible and scalable platform to analyze the hydrogeological and atmospheric balance of the Italian territory. The project aims to expand the current knowledge in hazard mapping, monitoring, and forecasting from an industrial perspective by leveraging innovative technologies and the interdisciplinary activities carried out by the ICSC.In this work, we present the cloud-HPC infrastructure deployed in the High-Performance Computing for Artificial Intelligence (HPC4AI) green data center of the University of Turin which supports the testing and development of HaMMon’s applications and services. We describe the current activities and preliminary results related to the integration of Photogrammetry techniques, Data Visualization and Artificial Intelligence technologies, applied on aerial images, to assess extreme natural events and evaluate their impact on risk-exposed assets.
Mauro Imbrosciano, Eva Sciacca, Fabio Vitello, Leonardo Pelonero, Francesco Franchina, Ugo Becciani, Iacopo Colonnelli, Doriana Medic
PDP7
2025 High Performance Visualization for Astrophysics and Cosmology
abstract
Modern Astrophysics and Cosmology (A&C) projects produce immense data volumes, necessitating advanced software tools for data access, storage, and analysis. Visualization Interface for the Virtual Observatory (VisIVO) is one such tool enabling multi-dimensional data analysis and knowledge discovery across complex astrophysical datasets. Leveraging containerization and virtualization, VisIVO has been deployed on various distributed computing platforms. Additionally, Blender, an open-source 3D suite, provides robust tools for rendering and processing volumetric data, making it suitable for visualizing complex datasets. At the SPACE Center of Excellence these tools are being adapted for high-performance visualization of cosmological simulations performed with GADGET and ChaNGa on pre-exascale systems. However, implementing high-performance visualization on diverse HPC platforms presents several challenges, including hardware and software compatibility, data management, scalability, performance portability, and efficient resource allocation. This paper outlines strategies to integrate VisIVO with workflow frameworks and streaming platforms to address these challenges. Workflow frameworks enhance portability, scheduling, and reproducibility of visualization workflows on pre-exascale systems used in A&C simulations. We also discuss the use of streaming platforms to enable concurrent (i.e. in-situ) analysis and visualization of simulations, reducing the need to store full simulation data by leveraging distributed databases that stream the output data in real time. Lastly, we present an adaptation of Blender to handle large-scale particle-based astrophysical data, offering high-quality visualization with interactive exploration capabilities.
Nicola Tuccari, Eva Sciacca, Fabio Vitello, Iacopo Colonnelli, Yolanda Becerra 0001, Enric Sosa Cintero, Guillermo Marin, Milan Jaros, Lubomir Riha, Petr Strakos, Sebastian Trujillo-Gomez, Emiliano Tramontana, Robert Wissing
PDP4
2025 Reproducibility Report for SC25 Paper ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
abstract
This reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage by Shen et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibilty Committee.
Iacopo Colonnelli
SC1
2025 Reproducibility Report for SC25 Paper Bridging the Gap Between Binary and Source Based Package Management in Spack
abstract
This reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper Bridging the Gap Between Binary and Source Based Package Management in Spack by Gouwar et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibilty Committee.
Iacopo Colonnelli
SC1
2024 Introducing SWIRL: An Intermediate Representation Language for Scientific Workflows
abstract
Abstract In the ever-evolving landscape of scientific computing, properly supporting the modularity and complexity of modern scientific applications requires new approaches to workflow execution, like seamless interoperability between different workflow systems, distributed-by-design workflow models, and automatic optimisation of data movements. In order to address this need, this article introduces SWIRL, an intermediate representation language for scientific workflows. In contrast with other product-agnostic workflow languages, SWIRL is not designed for human interaction but to serve as a low-level compilation target for distributed workflow execution plans. The main advantages of SWIRL semantics are low-level primitives based on the send/receive programming model and a formal framework ensuring the consistency of the semantics and the specification of translating workflow models represented by Directed Acyclic Graphs (DAGs) into SWIRL workflow descriptions. Additionally, SWIRL offers rewriting rules designed to optimise execution traces, accompanied by corresponding equivalence. An open-source SWIRL compiler toolchain has been developed using the ANTLR Python3 bindings.
Iacopo Colonnelli, Doriana Medic, Alberto Mulone, Viviana Bono, Luca Padovani, Marco Aldinucci
FM (1)1
2024 Benchmarking Parallelization Models through Karmarkar's Interior-point method
abstract
Optimization problems are one of the main focus of scientific research. Their computational-intensive nature makes them prone to be parallelized with consistent improvements in performance. This paper sheds light on different parallel models for accelerating Karmarkar's Interior-point method. To do so, we assess parallelization strategies for individual operations within Karmarkar's algorithm using OpenMP, GPU acceleration with CUDA, and the recent Parallel Standard C++ Linear Algebra library (PSTL) executing both GPU and CPU. Our different implementations yield interesting benchmark results that show the optimal approach for parallelizing interior point algorithms for general Linear Programming (LP) problems. In addition, we propose a more theoretical perspective of the parallelization of this algorithm, with a detailed study of our OpenMP implemen-tation, showing the limits of optimizing the single operations.
Marco Edoardo Santimaria, Samuele Fonio, Giulio Malenza, Iacopo Colonnelli, Marco Aldinucci
PDP4
2023 Experimenting with Emerging RISC-V Systems for Decentralised Machine Learning
abstract
Decentralised Machine Learning (DML) enables collaborative machine learning without centralised input data. Federated Learning (FL) and Edge Inference are examples of DML. While tools for DML (especially FL) are starting to flourish, many are not flexible and portable enough to experiment with novel processors (e.g., RISC-V), non-fully connected network topologies, and asynchronous collaboration schemes. We overcome these limitations via a domain-specific language allowing us to map DML schemes to an underlying middleware, i.e. the FastFlow parallel programming library. We experiment with it by generating different working DML schemes on x86-64 and ARM platforms and an emerging RISC-V one. We characterise the performance and energy efficiency of the presented schemes and systems. As a byproduct, we introduce a RISC-V porting of the PyTorch framework, the first publicly available to our knowledge.
Gianluca Mittone, Nicolò Tonci, Robert Birke, Iacopo Colonnelli, Doriana Medic, Andrea Bartolini, Roberto Esposito, Emanuele Parisi, Francesco Beneventi, Mirko Polato, Massimo Torquati, Luca Benini, Marco Aldinucci
CF4
2023 Model-Agnostic Federated Learning
Gianluca Mittone, Walter Riviera, Iacopo Colonnelli, Robert Birke, Marco Aldinucci
Euro-Par3
2023 CAPIO: a Middleware for Transparent I/O Streaming in Data- Intensive Workflows
abstract
With the increasing amount of digital data available for analysis and simulation, the class of I/O-intensive HPC workflows is fated to quickly expand, further exacerbating the performance gap between computing, memory, and storage technologies. This paper introduces CAPIO (Cross-Application Programmable I/O), a middleware capable of injecting I/O streaming capabilities into file-based workflows, improving the computation- I/O overlap without the need to change the application code. The contribution is twofold: 1) at design time, a new I/O coordination language allows users to annotate workflow data dependencies with synchronization semantics; 2) at run time, a user-space middleware automatically and transparently to the user turns a workflow batch execution into a streaming execution according to the semantics expressed in the configuration file. CAPIO has been tested on synthetic benchmarks simulating typical workflow I/O patterns and two real-world workflows. Experiments show that CAPIO reduces the execution time by 10% to 66% for data-intensive workflows that use the file system as a communication medium.
Alberto Riccardo Martinelli, Massimo Torquati, Marco Aldinucci, Iacopo Colonnelli, Barbara Cantalupo
HiPC4
2023 Pooling critical datasets with Federated Learning
abstract
Federated Learning (FL) is becoming popular in different industrial sectors where data access is critical for security, privacy and the economic value of data itself. Unlike traditional machine learning, where all the data must be globally gathered for analysis, FL makes it possible to extract knowledge from data distributed across different organizations that can be coupled with different Machine Learning paradigms. In this work, we replicate, using Federated Learning, the analysis of a pooled dataset (with AdaBoost) that has been used to define the PRAISE score, which is today among the most accurate scores to evaluate the risk of a second acute myocardial infarction. We show that thanks to the extended-OpenFL framework, which implements AdaBoost.F, we can train a federated PRAISE model that exhibits comparable accuracy and recall as the centralised model. We achieved F1 and F2 scores which are consistently comparable to the PRAISE score study of a 16-parties federation but within an order of magnitude less time.
Yasir Arfat, Gianluca Mittone, Iacopo Colonnelli, Fabrizio D'Ascenzo, Roberto Esposito, Marco Aldinucci
PDP3
2022 Distributed workflows with Jupyter
abstract
The designers of a new coordination interface enacting complex workflows have to tackle a dichotomy: choosing a language-independent or language-dependent approach. Language-independent approaches decouple workflow models from the host code’s business logic and advocate portability. Language-dependent approaches foster flexibility and performance by adopting the same host language for business and coordination code. Jupyter Notebooks, with their capability to describe both imperative and declarative code in a unique format, allow taking the best of the two approaches, maintaining a clear separation between application and coordination layers but still providing a unified interface to both aspects. We advocate the Jupyter Notebooks’ potential to express complex distributed workflows, identifying the general requirements for a Jupyter-based Workflow Management System (WMS) and introducing a proof-of-concept portable implementation working on hybrid Cloud-HPC infrastructures. As a byproduct, we extended the vanilla IPython kernel with workflow-based parallel and distributed execution capabilities. The proposed Jupyter-workflow (Jw) system is evaluated on common scenarios for High Performance Computing (HPC) and Cloud, showing its potential in lowering the barriers between prototypical Notebooks and production-ready implementations.
Iacopo Colonnelli, Marco Aldinucci, Barbara Cantalupo, Luca Padovani, Sergio Rabellino, Concetto Spampinato, Roberto Morelli, Rosario Di Carlo, Nicolò Magini, Carlo Cavazzoni
Future Gener. Comput. Syst.1
2021 TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascale
abstract
To achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research.
Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola
DSD22
2021 Practical parallelization of scientific applications with OpenMP, OpenACC and MPI
Marco Aldinucci, Valentina Cesare, Iacopo Colonnelli, Alberto Riccardo Martinelli, Gianluca Mittone, Barbara Cantalupo, Carlo Cavazzoni, Maurizio Drocco
J. Parallel Distributed Comput.3
2020 Practical Parallelization of Scientific Applications
abstract
This work aims at distilling a systematic methodology to modernize existing sequential scientific codes with a limited re-designing effort, turning an old codebase into modern code, i.e., parallel and robust code. We propose an automatable methodology to parallelize scientific applications designed with a purely sequential programming mindset, thus possibly using global variables, aliasing, random number generators, and stateful functions. We demonstrate the methodology by way of an astrophysical application, where we model at the same time the kinematic profiles of 30 disk galaxies with a Monte Carlo Markov Chain (MCMC), which is sequential by definition. The parallel code exhibits a 12 times speedup on a 48-core platform.
Valentina Cesare, Iacopo Colonnelli, Marco Aldinucci
PDP2
2019 Accelerating Spectral Graph Analysis Through Wavefronts of Linear Algebra Operations
abstract
The wavefront pattern captures the unfolding of a parallel computation in which data elements are laid out as a logical multidimensional grid and the dependency graph favours a diagonal sweep across the grid. In the emerging area of spectral graph analysis, the computing often consists in a wavefront running over a tiled matrix, involving expensive linear algebra kernels. While these applications might benefit from parallel heterogeneous platforms (multi-core with GPUs), programming wavefront applications directly with high-performance linear algebra libraries yields code that is complex to write and optimize for the specific application. We advocate a methodology based on two abstractions (linear algebra and parallel pattern-based run-time), that allows to develop portable, self-configuring, and easy-to-profile code on hybrid platforms.
Maurizio Drocco, Paolo Viviani 0001, Iacopo Colonnelli, Marco Aldinucci, Marco Grangetto
PDP3
2019 Deep Learning at Scale
abstract
This work presents a novel approach to distributed training of deep neural networks (DNNs) that aims to overcome the issues related to mainstream approaches to data parallel training. Established techniques for data parallel training are discussed from both a parallel computing and deep learning perspective, then a different approach is presented that is meant to allow DNN training to scale while retaining good convergence properties. Moreover, an experimental implementation is presented as well as some preliminary results.
Paolo Viviani 0001, Maurizio Drocco, Daniele Baccega, Iacopo Colonnelli, Marco Aldinucci
PDP4