VLDB 2026 Research / reviewers in the wild / expert
Marco Aldinucci
dblp:75/5249
· DBLP profile ↗
90ranked-venue papers
31as first author
41since 2021 · last 2026
0000-0001-8788-0829ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 18 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QSplit: A Workflow-Oriented Hybrid Quantum-Classical Optimization Framework
Mario Bifulco, Francesco Medina, Doriana Medic, Luca Roversi, Marco Aldinucci |
Euro-Par (2) | 5 |
| 2026 | Accelerating Sharded Data Parallelism at Scale with Federated Learning
Gianluca Mittone, Marco Aldinucci |
Euro-Par (2) | 2 |
| 2026 | A comprehensive performance evaluation of TEEs for confidential DNA alignment
Lorenzo Brescia, Iacopo Colonnelli, Robert Birke, Valerio Schiavoni, Pascal Felber, Marco Aldinucci |
Future Gener. Comput. Syst. | 6 |
| 2026 | Inference performance of large language models on a 64-core RISC-V CPU with silicon-enabled vectors
Adriano Marques Garcia, Giulio Malenza, Robert Birke, Marco Aldinucci |
Future Gener. Comput. Syst. | 4 |
| 2026 | A formal framework for fault tolerance in hybrid scientific workflowsabstractIn large-scale distributed systems, failures are routine events whose occurrences increase with the number of computational tasks and execution locations. The advantage of representing an application as a workflow is the possibility of exploiting Workflow Management System (WMS) features such as portability, scalability, and, crucially, reliability. Among these, reliability is essential for ensuring robust execution in dynamic and failure-prone environments. In recent years, the emergence of hybrid workflows has posed new and intriguing challenges by increasing the possibility of distributing computations involving heterogeneous and independent environments. Consequently, the number of possible points of failure during the execution increased, creating a need for sophisticated fault tolerance mechanisms capable of addressing the specific requirements of hybrid systems. This work introduces a formal framework for a fault tolerance mechanism in hybrid workflows, enabling failure recovery through a rollback approach. The framework is rigorously defined by adapting and extending an existing workflow semantics tailored for hybrid execution. Our method leverages provenance data from workflow execution up to the point of failure, and creates a recovery workflow that spans multiple infrastructures. The rollback approach provides a robust and reliable strategy to ensure resilience against step failures and potential data loss. We then implement this mechanism in the StreamFlow WMS, and evaluate it using two case studies: the 1000 Genomes workflow and a synthetic workflow featuring iterative patterns. Experiments showcase the conceptual validity of our approach and assess the overhead introduced by the mechanism, including data availability checks. Alberto Mulone, Doriana Medic, Iacopo Colonnelli, Marco Aldinucci |
Future Gener. Comput. Syst. | 4 |
| 2026 | Dynamic transparent streaming in file-based workflows with CAPIO
Marco Edoardo Santimaria, Iacopo Colonnelli, Barbara Cantalupo, Massimo Torquati, Doriana Medic, Nicola Tuccari, Eva Sciacca, Marco Aldinucci |
Future Gener. Comput. Syst. | 8 |
| 2025 | Hyperbolic Prototypical Entailment Cones for Image ClassificationabstractNon-Euclidean geometries have garnered significant research interest, particularly in their application to Deep Learning. Utilizing specific manifolds as embedding spaces has been shown to enhance neural network representational capabilities by aligning these spaces with the data’s latent structure. In this paper, we focus on hyperbolic manifolds and introduce a novel framework, Hyperbolic Prototypical Entailment Cones (HPEC). The core innovation of HPEC lies in utilizing angular relationships, rather than traditional distance metrics, to more effectively capture the similarity between data representations and their corresponding prototypes. This is achieved by leveraging hyperbolic entailment cones, a mathematical construct particularly suited for embedding hierarchical structures in the Poincare’ Ball, along with a novel Backclip mechanism. Our experimental results demonstrate that this approach significantly enhances performance in high-dimensional embedding spaces. To substantiate these findings, we evaluate HPEC on four diverse datasets across various embedding dimensions, consistently surpassing state-of-the-art methods in Prototype Learning. Samuele Fonio, Roberto Esposito, Marco Aldinucci |
AISTATS | 3 |
| 2025 | Towards RISC-V-based HPC: The Italian Pathfinding Activities in the DARE-SGA1 ProjectabstractThe European Union’s efforts towards technological sovereignty in High-Performance Computing are driving research and development of RISC-V-based supercomputers. The DARE SGA1 project, in particular, aims to develop chips designed and owned by Europeans. This paper introduces the Italian contribution to DARE SGA1 regarding pathfinding activities toward future RISC-V-based accelerator designs, reliability improvements, system software, and AI and Quantum Chemistry applications. Giovanni Agosta, Marco Aldinucci, Andrea Bartolini, Laura Bellentani, Andrea Biagioni, Daniele Cesarini, Carlotta Chiarini, Iacopo Colonnelli, Pietro Delugas, Lev Denisov, Ottorino Frezza, Marco Grangetto, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Andrea Maslov, Mauro Olivieri, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Davide Rossi 0001, Sergio Saponara, Antonio Sciarappa, Francesco Simula, Matteo Sonza Reorda, Massimo Torquati, Piero Vicini |
DSD | 2 |
| 2025 | BookedSlurm: meeting user needs for advanced resource reservations in SlurmabstractModern scientific discovery is frequently backed by largescale scientific experiments, which cannot prescind from the unparalleled computational capabilities provided by highperformance computing systems. However, large data centers usually prioritize system efficiency over user accessibility, posing challenges for researchers without advanced computer science expertise. This work introduces BookedSlurm, a secure and userfocused extension of the Slurm workload manager, aiming to democratize HPC access across interdisciplinary research domains. BookedSlurm enables a partially decentralized regulation of finegrained advanced resource reservations through a novel creditbased framework, ensuring fair and predictable access to computing resources. Its modular architecture leverages dedicated microservices to manage reservations, credit handling, and accounting. These components are exposed through a secure REST API and an intuitive webbased dashboard, enhancing system usability for novice and expert users. While the dashboard simplifies interactions for nonspecialists, advanced users can directly access agentlevel APIs for more complex and automated operations. The effectiveness of BookedSlurm is validated on a realworld bioinformatics use case from the SUSMIRRI.IT project, showcasing its ability to enhance usability, optimize job scheduling, and streamline execution workflows. Sandro Gepiro Contaldo, Lorenzo Bosio, Janneth Estefania Hoyos Rea, Elisa Li Perottino, Sergio Rabellino, Marco Aldinucci, Marco Beccuti, Iacopo Colonnelli |
eScience | 6 |
| 2025 | Fed2RC: Federated Rocket Kernels and Ridge Classifier for Time Series ClassificationabstractTime series classification is a pivotal task in modern machine learning, with widespread applications in fields such as healthcare, finance, and cybersecurity. While deep learning methods dominate recent developments, their resource demands and privacy limitations hinder deployment on low-power and decentralized environments. To address these challenges, we introduce Fed2RC, a fully federated and gradient-free approach that integrates the efficiency of Rocket-based feature extraction with the robustness of ridge regression in a privacy-preserving setting. Fed2RC builds upon two key ideas: (i) federated selection and aggregation of high-performing random convolution kernels, and (ii) incremental and communication-efficient updates of ridge classifier parameters using closed-form solutions. Additionally, we propose a novel federated protocol for selecting the global ridge regularization parameter λ, and show how to improve the communication efficiency by matrix factorization techniques. Extensive experiments on the UCR benchmark demonstrate that Fed2RC achieves state-of-the-art results with a fraction of the computation and communication costs. Code to reproduce the experiments can be found at: https://github.com/CasellaJr/Fed2RC. Bruno Casella, Samuele Fonio, Lorenzo Sciandra, Claudio Gallicchio, Marco Aldinucci, Mirko Polato, Roberto Esposito |
ECAI | 5 |
| 2025 | Covariances computation in the Gaia AVU-GSR Parallel Solver with I/O techniques: a performance study as a function of writing cycle lengthabstractThe solver module of the Astrometric Verification Unit-Global Sphere Reconstruction (AVU-GSR) pipeline aims to find the astrometric parameters of $\sim 10^{8}$ stars in the Milky Way, the attitude and instrumental settings of the Gaia satellite, and the parametrized post Newtonian parameter $\gamma$ with a resolution of 1 0 - 100 micro – arc seconds. To perform this task, the code, which runs in production on Leonardo CINECA infrastructure, solves a system of linear equations with the iterative LSQR algorithm, where the coefficient matrix is large (10-50 TB) and sparse and the iterations stop when least square convergence is reached. The solver was ported to GPU with CUDA, obtaining a $\sim 14 x$ acceleration factor over an original version CPU-parallelized with OpenMP. This work concentrates on a code section dedicated to covariances calculation, representing an important scientific task for Gaia mission, since the problems unknowns present strong correlations. Given the number of unknowns at mission end, the variances-covariances matrix is expected to occupy $\sim 1$ EB, which represents a substantial “Big Data” issue. To compute a subset of the total covariances, we defined an I/Obased pipeline made of two jobs. The first job, the LSQR, writes the files every $i \operatorname{tnCov} C P$ iterations, and the second job reads them and calculates the corresponding covariances. The two jobs can be launched either in sequence or concurrently. Previous studies demonstrated that the covariances calculation does not significantly slowdown the AVU-GSR production up to $\sim 3 \times 10^{7}$ covariances. Here we investigate the performance of the covariances pipeline as a function of $i t n \operatorname{Cov} C P$. The results show that writing smaller files more frequently or writing larger files less frequently does not affect the global performance of the solver, whose speed only depends on the number of covariances to calculate and of system unknowns. Valentina Cesare, Ugo Becciani, Alberto Vecchiato, Mario Gilberto Lattanzi, Marco Aldinucci, Beatrice Bucciarelli |
PDP | 5 |
| 2025 | Sustainable-HPC: toward Digital Twin for active management of self-cooled data centers with Renewable Energy Sources and waste heat recoveryabstractHigh-performance computing (HPC) data centers are key for advancing research and industry, but their high energy consumption and environmental impact pose critical sustainability challenges. Few HPC centers currently use renewable energy sources, particularly hydrogen-based storage, or integrate waste heat reuse. Recently, Digital Twins (DTs) are proving crucial for advancing sustainable HPC centers, enabling the integration of advanced monitoring and energy optimization with renewable energy sources. They simulate operations, predict maintenance needs, balance workloads, and support renewable integration for enhanced efficiency. The paper first reviews current practices, focusing on the potential of DTs to improve HPC sustainability by dynamically managing resources and optimizing systems, including cooling, power distribution, and load balancing. Then, it illustrates the University of Turin’s Sustainable HPC4AI (S-HPC4AI) project, which aims to develop a low impact HPC center to support Artificial Intelligence (AI) research across diverse scientific fields. The facility will feature renewable energy sources, advanced self-cooling, and waste heat recovery, setting a benchmark for energy-efficient, low-carbon HPC infrastructure. Central to the project is the integration of hydrogen and solar energy, with photovoltaic systems providing clean power and hydrogen fuel cells serving as reliable backup sources, reducing reliance on fossil fuels. Waste heat can be transformed into a productive resource for a local automated phenotyping system, reducing energy consumption and environmental impact on the overall. Furthermore, a DT will be developed, integrating BIM with sensors data about performance, resource usage and operating conditions, enabling real-time monitoring and predictive analytics for managing power, cooling, and energy resources. The potential and challenges of the S-HPC4AI model are discussed, suggesting possible solutions. Silvia Meschini, Lavinia Chiara Tagliabue, Giuseppe Martino Di Giuda, Marco Aldinucci, Paola Gasbarri, Daniele Accardo |
PDP | 4 |
| 2025 | HPC4AI@UNITO: A Use Case For Datacenter Digital TwinabstractThe HPC4AI datacenter hosted at the Computer Science Department of the University of Turin was born to address the exponentially increasing computing needs of cross-disciplinary research on AI. To address the needs of modern AI, HPC4AI rethinks the traditional usage of Cloud and HPC systems where the Cloud provides a modern interface for HPC, and HPC serves as an accelerator for the cloud. To date, it has supported 40+ research projects across a broad range of domains, from astronomy and medicine to human sciences. Furthermore, it acts as an R&D platform to study, develop, and test new datacenter technologies. As such, it hosts a zoo of exotic computing platforms and the first prototype of two-phase evaporative server cooling. This work describes operating and managing HPC4AI with its challenges and lessons learned, with an analysis of key opportunities for digital twins. Viviana Vaccaro, Robert Birke, Lavinia Chiara Tagliabue, Sergio Rabellino, Marco Aldinucci |
PDP | 5 |
| 2025 | Sustainable Data Centers: Advancing Energy Efficiency and Resource OptimizationabstractAs global demand for digital services rises, the role of data centers has become increasingly pivotal, making it essential to adopt measures to curb energy consumption and carbon emissions. Data centers rank among the highest energy consumers and are therefore significant contributors to greenhouse gas emissions within the ICT sector. This context has driven the development of innovative solutions aimed at minimizing the environmental impact of data centers while maintaining high standards of performance, security, and reliability. This research presents an in-depth analysis of the current state of the sector. It offers a detailed understanding of the strategies that can be implemented to optimize performance, focusing on energy efficiency, sustainability, and technological innovation. It aims to highlight how the adoption of innovative strategies can significantly reduce energy consumption and CO 2 emissions, while optimizing resource utilization. The study also highlights how these strategies can be integrated in a holistic approach to design and management, demonstrating the interdependence between different sustainability practices. Methodologically, this work combines an accurate analysis of existing literature and established industry practices. Energy efficiency strategies are examined, such as server consumption reduction and the adoption of more efficient cooling systems, with a focus on advanced technologies. Special attention is also given to emerging IT technologies, such as hyper-converged infrastructures (HCI) and artificial intelligence (AI), evaluating their potential to improve the efficiency and scalability of nextgeneration data centers, outlining a pathway toward a sustainable digital future that prioritizes responsible resource usage and environmental management. Viviana Vaccaro, Lavinia Chiara Tagliabue, Marco Aldinucci |
PDP | 3 |
| 2025 | Performance Portability Assessment in GaiaabstractModern scientific experiments produce ever-increasing amounts of data, soon requiring ExaFLOPs computing capacities for analysis. Reaching such performance requires purpose-built supercomputers with$O(10^{3})$nodes, each hosting multicore CPUs and multiple GPUs, and applications designed to exploit this hardware optimally. Given that each supercomputer is generally a one-off project, the need for computing frameworks portable across diverse CPU and GPU architectures without performance losses is increasingly compelling. We investigate the performance portability (ȹ) of a real-world application: the solver module of the AVU–GSR pipeline for the ESA Gaia mission. This code finds the astrometric parameters of$\sim$$10^{8}$stars in the Milky Way using the LSQR iterative algorithm. LSQR is widely used to solve linear systems of equations across a wide range of high-performance computing applications, elevating the study beyond its astrophysical relevance. The code is memory-bound, with six main compute kernels implementing sparse matrix-by-vector products. We optimize the previous CUDA implementation and port the code to further six GPU-acceleration frameworks: C++ PSTL, SYCL, OpenMP, HIP, KOKKOS, and OpenACC. We evaluate each framework's performance portability across multiple GPUs (NVIDIA and AMD) and problem sizes in terms of application and architectural efficiency. Architectural efficiency is estimated through the roofline model of the six most computationally expensive GPU kernels. Our results show that C++ library-based (C++ PSTL and KOKKOS), pragma-based (OpenMP and OpenACC), and language-specific (CUDA, HIP, and SYCL) frameworks achieve increasingly better performance portability across the supported platforms with larger problem sizes providing better ȹ scores due to higher GPU occupancies. Giulio Malenza, Valentina Cesare, Marco Edoardo Santimaria, Robert Birke, Alberto Vecchiato, Ugo Becciani, Marco Aldinucci |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | The TEXTAROSSA Project: Cool all the Way Down to the HardwareabstractThe TEXTAROSSA project aims to bridge the technology gaps that exascale computing systems will face in the near future in order to overcome their performance and energy efficiency challenges. This project provides solutions for improved energy efficiency and thermal control, seamless integration of heterogeneous accelerators in HPC multi-node platforms, and new arithmetic methods. Challenges are tacked through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models, and tools derived from European research. Antonio Filgueras, Giovanni Agosta, Marco Aldinucci, Carlos Álvarez 0001, Pasqua D'Ambra, Massimo Bernaschi, Andrea Biagioni, Daniele Cattaneo 0002, Alessandro Celestini, Massimo Celino, Carlotta Chiarini, Francesca Lo Cicero, Paolo Cretaro, William Fornaciari, Ottorino Frezza, Andrea Galimberti, Francesco Giacomini, Juan Miguel De Haro Ruiz, Francesco Iannone, Daniel Jaschke, Daniel Jiménez-González, Michal Kulczewski, Alberto Leva, Alessandro Lonardo, Michele Martinelli, Xavier Martorell, Simone Montangero, Lucas Morais, Ariel Oleksiak, Paolo Palazzari, Luca Pontisso, Federico Reghenzani, Cristian Rossi, Sergio Saponara, Carlo Saverio Lodi, Francesco Simula, Federico Terraneo, Piero Vicini, Miquel Vidal, Davide Zoni, Giuseppe Zummo |
DSD | 3 |
| 2024 | Federated Learning in a Semi-Supervised Environment for Earth Observation DataabstractWe propose FedRec, a federated learning workflow taking advantage of unlabelled data in a semi-supervised environment to assist in the training of a supervised aggregated model.In our proposed method, an encoder architecture extracting features from unlabelled data is aggregated with the feature extractor of a classification model via weight averaging.The fully connected layers of the supervised models are also averaged in a federated fashion.We show the effectiveness of our approach by comparing it with the state-of-the-art federated algorithm, an isolated and a centralised baseline, on novel cloud detection datasets.Our code is available at https://github.com/CasellaJr/FedRec.* This work has been partly supported by the Spoke "FutureHPC & BigData" of the ICSC -Centro Nazionale di Ricerca in "High Performance Computing, Big Data and Quantum Computing", funded by European Union - Bruno Casella, Alessio Barbaro Chisari, Marco Aldinucci, Sebastiano Battiato, Mario Valerio Giuffrida |
ESANN | 3 |
| 2024 | Federated Time Series Classification with ROCKET featuresabstractThis paper proposes FROCKS, a federated time series classification method using ROCKET features.Our approach dynamically adapts the models' features by selecting and exchanging the bestperforming ROCKET kernels from a federation of clients.Specifically, the server gathers the best-performing kernels of the clients together with the associated model parameters, and it performs a weighted average if a kernel is best-performing for more than one client.We compare the proposed method with state-of-the-art approaches on the UCR archive binary classification datasets and show superior performance on most datasets. Bruno Casella, Matthias Jakobs, Marco Aldinucci, Sebastian Buschjäger |
ESANN | 3 |
| 2024 | Introducing SWIRL: An Intermediate Representation Language for Scientific WorkflowsabstractAbstract In the ever-evolving landscape of scientific computing, properly supporting the modularity and complexity of modern scientific applications requires new approaches to workflow execution, like seamless interoperability between different workflow systems, distributed-by-design workflow models, and automatic optimisation of data movements. In order to address this need, this article introduces SWIRL, an intermediate representation language for scientific workflows. In contrast with other product-agnostic workflow languages, SWIRL is not designed for human interaction but to serve as a low-level compilation target for distributed workflow execution plans. The main advantages of SWIRL semantics are low-level primitives based on the send/receive programming model and a formal framework ensuring the consistency of the semantics and the specification of translating workflow models represented by Directed Acyclic Graphs (DAGs) into SWIRL workflow descriptions. Additionally, SWIRL offers rewriting rules designed to optimise execution traces, accompanied by corresponding equivalence. An open-source SWIRL compiler toolchain has been developed using the ANTLR Python3 bindings. Iacopo Colonnelli, Doriana Medic, Alberto Mulone, Viviana Bono, Luca Padovani, Marco Aldinucci |
FM (1) | 6 |
| 2024 | DALLMi: Domain Adaption for LLM-Based Multi-label Classifier
Miruna Betianu, Abele Malan, Marco Aldinucci, Robert Birke, Lydia Y. Chen |
PAKDD (3) | 3 |
| 2024 | Benchmarking Parallelization Models through Karmarkar's Interior-point methodabstractOptimization problems are one of the main focus of scientific research. Their computational-intensive nature makes them prone to be parallelized with consistent improvements in performance. This paper sheds light on different parallel models for accelerating Karmarkar's Interior-point method. To do so, we assess parallelization strategies for individual operations within Karmarkar's algorithm using OpenMP, GPU acceleration with CUDA, and the recent Parallel Standard C++ Linear Algebra library (PSTL) executing both GPU and CPU. Our different implementations yield interesting benchmark results that show the optimal approach for parallelizing interior point algorithms for general Linear Programming (LP) problems. In addition, we propose a more theoretical perspective of the parallelization of this algorithm, with a detailed study of our OpenMP implemen-tation, showing the limits of optimizing the single operations. Marco Edoardo Santimaria, Samuele Fonio, Giulio Malenza, Iacopo Colonnelli, Marco Aldinucci |
PDP | 5 |
| 2024 | FedER: Federated Learning through Experience Replay and privacy-preserving data synthesisabstractIn the medical field, multi-center collaborations are often sought to yield more generalizable findings by leveraging the heterogeneity of patient and clinical data. However, recent privacy regulations hinder the possibility to share data, and consequently, to come up with machine learning-based solutions that support diagnosis and prognosis. Federated learning (FL) aims at sidestepping this limitation by bringing AI-based solutions to data owners and only sharing local AI models, or parts thereof, that need then to be aggregated. However, most of the existing federated learning solutions are still at their infancy and show several shortcomings, from the lack of a reliable and effective aggregation scheme able to retain the knowledge learned locally to weak privacy preservation as real data may be reconstructed from model updates. Furthermore, the majority of these approaches, especially those dealing with medical data, relies on a centralized distributed learning strategy that poses robustness, scalability and trust issues. In this paper we present a federated learning strategy, FedER, that, exploiting experience replay and generative adversarial concepts, effectively integrates features from local nodes, providing models able to generalize across multiple datasets while maintaining privacy. FedER is tested on two tasks — tuberculosis and melanoma classification — using multiple datasets in order to simulate realistic non-i.i.d. medical data scenarios. Results show that our approach achieves performance comparable to standard (non-federated) learning and significantly outperforms state-of-the-art federated methods. Remarkably, we also observe that FedER enables any node model to be used as a global federation model. Indeed, the experience replay strategy with privacy-preserving synthetic data allows all node models to converge to reach the same optimum without the need of a single shared model. Code is available at https://github.com/perceivelab/FedER. Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Bruno Casella, Marco Aldinucci, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 5 |
| 2024 | Analyzing FOSS license usage in publicly available software at scale via the SWH-analytics frameworkabstractAbstract The Software Heritage (SWH) dataset represents an invaluable source of open-source code as it aims to collect, preserve, and share all publicly available software in source code form ever produced by humankind. Although designed to archive deduplicated small files thanks to the use of a Merkle tree as the underlying data structure, querying the SWH dataset presents challenges due to the nature of these structures, which organize content based on hash values rather than any locality principle. The magnitude of the repository, coupled with the resource-intensive nature of the download process, highlights the need for specialized infrastructure and computational resources to effectively handle and study the extensive dataset housed within SWH. Currently, there is a lack of infrastructures specifically tailored for running analytics on the SWH dataset, leaving users to handle these issues manually. To address these challenges, we implemented the SWH-Analytics (SWHA) framework, a development environment that transparently runs custom analytic applications on publicly available software data preserved over time by SWH. Specifically, this work shows how SWHA can be effectively exploited to study usage patterns of free and open-source software licenses, highlighting the need to improve license literacy among developers. Alessia Antelmi, Massimo Torquati, Giacomo Corridori, Daniele Gregori, Francesco Polzella, Gianmarco Spinatelli, Marco Aldinucci |
J. Supercomput. | 7 |
| 2024 | Toward HPC application portability via C++ PSTL: the Gaia AVU-GSR code assessmentabstractAbstract The computing capacity needed to process the data generated in modern scientific experiments is approaching ExaFLOPs. Currently, achieving such performances is only feasible through GPU-accelerated supercomputers. Different languages were developed to program GPUs at different levels of abstraction. Typically, the more abstract the languages, the more portable they are across different GPUs. However, the less abstract and co-designed with the hardware, the more room for code optimization and, eventually, the more performance. In the HPC context, portability and performance are a fairly traditional dichotomy. The current C++ Parallel Standard Template Library (PSTL) has the potential to go beyond this dichotomy. In this work, we analyze the main performance benefits and limitations of PSTL using as a use-case the Gaia Astrometric Verification Unit-Global Sphere Reconstruction parallel solver developed by the European Space Agency Gaia mission. The code aims to find the astrometric parameters of $$\sim10^8$$ ∼ 10 8 stars in the Milky Way by iteratively solving a linear system of equations with the LSQR algorithm, originally GPU-ported with the CUDA language. We show that the performance obtained with the PSTL version, which is intrinsically more portable than CUDA, is comparable to the CUDA one on NVIDIA GPU architecture. Giulio Malenza, Valentina Cesare, Marco Aldinucci, Ugo Becciani, Alberto Vecchiato |
J. Supercomput. | 3 |
| 2024 | A Convolutional-Transformer Model for FFR and iFR Assessment From Coronary AngiographyabstractThe quantification of stenosis severity from X-ray catheter angiography is a challenging task. Indeed, this requires to fully understand the lesion's geometry by analyzing dynamics of the contrast material, only relying on visual observation by clinicians. To support decision making for cardiac intervention, we propose a hybrid CNN-Transformer model for the assessment of angiography-based non-invasive fractional flow-reserve (FFR) and instantaneous wave-free ratio (iFR) of intermediate coronary stenosis. Our approach predicts whether a coronary artery stenosis is hemodynamically significant and provides direct FFR and iFR estimates. This is achieved through a combination of regression and classification branches that forces the model to focus on the cut-off region of FFR (around 0.8 FFR value), which is highly critical for decision-making. We also propose a spatio-temporal factorization mechanisms that redesigns the transformer's self-attention mechanism to capture both local spatial and temporal interactions between vessel geometry, blood flow dynamics, and lesion morphology. The proposed method achieves state-of-the-art performance on a dataset of 778 exams from 389 patients. Unlike existing methods, our approach employs a single angiography view and does not require knowledge of the key frame; supervision at training time is provided by a classification loss (based on a threshold of the FFR/iFR values) and a regression loss for direct estimation. Finally, the analysis of model interpretability and calibration shows that, in spite of the complexity of angiographic imaging data, our method can robustly identify the location of the stenosis and correlate prediction uncertainty to the provided output scores. Raffaele Mineo, Federica Proietto Salanitri, Giovanni Bellitto, Isaak Kavasidis, Ovidio De Filippo, M. Millesimo, Gaetano Maria de Ferrari, Marco Aldinucci, Daniela Giordano, Simone Palazzo, Fabrizio D'Ascenzo, Concetto Spampinato |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Adaptive multi-tier intelligent data manager for ExascaleabstractThe main objective of the ADMIRE project1 is the creation of an active I/O stack that dynamically adjusts computation and storage requirements through intelligent global coordination, the elasticity of computation and I/O, and the scheduling of storage resources along all levels of the storage hierarchy, while offering quality-of-service (QoS), energy efficiency, and resilience for accessing extremely large data sets in very heterogeneous computing and storage environments. We have developed a framework prototype that is able to dynamically adjust computation and storage requirements through intelligent global coordination, separated control, and data paths, the malleability of computation and I/O, the scheduling of storage resources along all levels of the storage hierarchy, and scalable monitoring techniques. The leading idea in ADMIRE is to co-design applications with ad-hoc storage systems that can be deployed with the application and adapt their computing and I/O behaviour on runtime, using malleability techniques, to increase the performance of applications and the throughput of the applications. Jesús Carretero 0001, Francisco Javier García Blas, Marco Aldinucci, Jean-Baptiste Besnard, Jean-Thomas Acquaviva, André Brinkmann, Marc-Andre Vef, Emmanuel Jeannot, Alberto Miranda, Ramon Nou, Morris Riedel, Massimo Torquati, Felix Wolf 0001 |
CF | 3 |
| 2023 | Experimenting with Emerging RISC-V Systems for Decentralised Machine LearningabstractDecentralised Machine Learning (DML) enables collaborative machine learning without centralised input data. Federated Learning (FL) and Edge Inference are examples of DML. While tools for DML (especially FL) are starting to flourish, many are not flexible and portable enough to experiment with novel processors (e.g., RISC-V), non-fully connected network topologies, and asynchronous collaboration schemes. We overcome these limitations via a domain-specific language allowing us to map DML schemes to an underlying middleware, i.e. the FastFlow parallel programming library. We experiment with it by generating different working DML schemes on x86-64 and ARM platforms and an emerging RISC-V one. We characterise the performance and energy efficiency of the presented schemes and systems. As a byproduct, we introduce a RISC-V porting of the PyTorch framework, the first publicly available to our knowledge. Gianluca Mittone, Nicolò Tonci, Robert Birke, Iacopo Colonnelli, Doriana Medic, Andrea Bartolini, Roberto Esposito, Emanuele Parisi, Francesco Beneventi, Mirko Polato, Massimo Torquati, Luca Benini, Marco Aldinucci |
CF | 13 |
| 2023 | A Proposal for a Continuum-aware Programming Model: From Workflows to Services Autonomously Interacting in the Compute ContinuumabstractThis paper proposes a continuum-aware programming model enabling the execution of application workflows across the compute continuum: cloud, fog and edge resources. It simplifies the management of heterogeneous nodes while alleviating the burden of programmers and unleashing innovation. This model optimizes the continuum through advanced development experiences by transforming workflows into autonomous service collaborations. It reduces complexity in positioning/interconnecting services across the continuum. A meta-model introduces high-level workflow descriptions as service networks with defined contracts and quality of service, thus enabling the deployment/management of workflows as first-class entities. It also provides automation based on policies, monitoring and heuristics. Tailored mechanisms orchestrate/manage services across the continuum, optimizing performance, cost, data protection and sustainability while managing risks. This model facilitates incremental development with visibility of design impacts and seamless evolution of applications and infrastructures. In this work, we explore this new computing paradigm showing how it can trigger the development of a new generation of tools to support the compute continuum progress. Marco Aldinucci, Robert Birke, Antonio Brogi, Emanuele Carlini 0001, Massimo Coppola, Marco Danelutto, Patrizio Dazzi, Luca Ferrucci, Stefano Forti 0002, Hanna Kavalionak, Gabriele Mencagli, Matteo Mordacchini, Marcelo Pasin, Federica Paganelli, Massimo Torquati |
COMPSAC | 1 |
| 2023 | Towards formal model for location aware workflowsabstractDesigning complex applications and executing them on large-scale topologies of heterogeneous architectures is becoming increasingly crucial in many scientific domains. As a result, diverse workflow modelling paradigms are developed, most of them with no formalisation provided. In these circumstances, comparing two different models or switching from one system to the other becomes a hard nut to crack.This paper investigates the capability of process algebra to model a location aware workflow system. Distributed π-calculus is considered as the base of the formal model due to its ability to describe the communicating components that change their structure as an outcome of the communication. Later, it is discussed how the base model could be extended or modified to capture different features of location aware workflow system.The intention of this paper is to highlight the fact that due to its flexibility, π-calculus, could be a good candidate to represent the behavioural perspective of the workflow system. Doriana Medic, Marco Aldinucci |
COMPSAC | 2 |
| 2023 | Porting the Variant Calling Pipeline for NGS data in cloud-HPC environmentabstractIn recent years we have understood the importance of analyzing and sequencing human genetic variation. A relevant aspect that emerged from the Covid-19 pandemic was the need to obtain results very quickly; this involved using High-Performance Computing (HPC) environments to execute the Next Generation Sequencing (NGS) pipeline. However, HPC is not always the most suitable environment for the entire execution of a pipeline, especially when it involves many heterogeneous tools. The ability to execute parts of the pipeline on different environments can lead to higher performance but also cheaper executions. This work shows the design and optimization process that led us to a state-of-the-art Variant Calling hybrid workflow based on the StreamFlow Workflow Management System (WfMS). We also compare StreamFlow with Snakemake, an established WfMS targeting HPC facilities, observing comparable performance on single environments and satisfactory improvements with a hybrid cloud-HPC configuration. Alberto Mulone, Sherine Awad, Davide Chiarugi, Marco Aldinucci |
COMPSAC | 4 |
| 2023 | Hercules: Scalable and Network Portable In-Memory Ad-Hoc File System for Data-Centric and High-Performance Applications
Francisco Javier García Blas, Genaro Sanchez-Gallegos, Cosmin Petre, Alberto Riccardo Martinelli, Marco Aldinucci, Jesús Carretero 0001 |
Euro-Par | 5 |
| 2023 | Model-Agnostic Federated Learning
Gianluca Mittone, Walter Riviera, Iacopo Colonnelli, Robert Birke, Marco Aldinucci |
Euro-Par | 5 |
| 2023 | CAPIO: a Middleware for Transparent I/O Streaming in Data- Intensive WorkflowsabstractWith the increasing amount of digital data available for analysis and simulation, the class of I/O-intensive HPC workflows is fated to quickly expand, further exacerbating the performance gap between computing, memory, and storage technologies. This paper introduces CAPIO (Cross-Application Programmable I/O), a middleware capable of injecting I/O streaming capabilities into file-based workflows, improving the computation- I/O overlap without the need to change the application code. The contribution is twofold: 1) at design time, a new I/O coordination language allows users to annotate workflow data dependencies with synchronization semantics; 2) at run time, a user-space middleware automatically and transparently to the user turns a workflow batch execution into a streaming execution according to the semantics expressed in the configuration file. CAPIO has been tested on synthetic benchmarks simulating typical workflow I/O patterns and two real-world workflows. Experiments show that CAPIO reduces the execution time by 10% to 66% for data-intensive workflows that use the file system as a communication medium. Alberto Riccardo Martinelli, Massimo Torquati, Marco Aldinucci, Iacopo Colonnelli, Barbara Cantalupo |
HiPC | 3 |
| 2023 | Pooling critical datasets with Federated LearningabstractFederated Learning (FL) is becoming popular in different industrial sectors where data access is critical for security, privacy and the economic value of data itself. Unlike traditional machine learning, where all the data must be globally gathered for analysis, FL makes it possible to extract knowledge from data distributed across different organizations that can be coupled with different Machine Learning paradigms. In this work, we replicate, using Federated Learning, the analysis of a pooled dataset (with AdaBoost) that has been used to define the PRAISE score, which is today among the most accurate scores to evaluate the risk of a second acute myocardial infarction. We show that thanks to the extended-OpenFL framework, which implements AdaBoost.F, we can train a federated PRAISE model that exhibits comparable accuracy and recall as the centralised model. We achieved F1 and F2 scores which are consistently comparable to the PRAISE score study of a 16-parties federation but within an order of magnitude less time. Yasir Arfat, Gianluca Mittone, Iacopo Colonnelli, Fabrizio D'Ascenzo, Roberto Esposito, Marco Aldinucci |
PDP | 6 |
| 2022 | Boosting the Federation: Cross-Silo Federated Learning without Gradient DescentabstractFederated Learning has been proposed to develop better AI systems without compromising the privacy of final users and the legitimate interests of private companies. Initially deployed by Google to predict text input on mobile devices, FL has been deployed in many other industries. Since its introduction, Federated Learning mainly exploited the inner working of neural networks and other gradient descent-based algorithms by either exchanging the weights of the model or the gradients computed during learning. While this approach has been very successful, it rules out applying FL in contexts where other models are preferred, e.g., easier to interpret or known to work better. This paper proposes FL algorithms that build federated models without relying on gradient descent-based methods. Specifically, we leverage distributed versions of the AdaBoost algorithm to acquire strong federated models. In contrast with previous approaches, our proposal does not put any constraint on the client-side learning models. We perform a large set of experiments on ten UCI datasets, comparing the algorithms in six non-iidness settings. Mirko Polato, Roberto Esposito, Marco Aldinucci |
IJCNN | 3 |
| 2022 | Distributed workflows with JupyterabstractThe designers of a new coordination interface enacting complex workflows have to tackle a dichotomy: choosing a language-independent or language-dependent approach. Language-independent approaches decouple workflow models from the host code’s business logic and advocate portability. Language-dependent approaches foster flexibility and performance by adopting the same host language for business and coordination code. Jupyter Notebooks, with their capability to describe both imperative and declarative code in a unique format, allow taking the best of the two approaches, maintaining a clear separation between application and coordination layers but still providing a unified interface to both aspects. We advocate the Jupyter Notebooks’ potential to express complex distributed workflows, identifying the general requirements for a Jupyter-based Workflow Management System (WMS) and introducing a proof-of-concept portable implementation working on hybrid Cloud-HPC infrastructures. As a byproduct, we extended the vanilla IPython kernel with workflow-based parallel and distributed execution capabilities. The proposed Jupyter-workflow (Jw) system is evaluated on common scenarios for High Performance Computing (HPC) and Cloud, showing its potential in lowering the barriers between prototypical Notebooks and production-ready implementations. Iacopo Colonnelli, Marco Aldinucci, Barbara Cantalupo, Luca Padovani, Sergio Rabellino, Concetto Spampinato, Roberto Morelli, Rosario Di Carlo, Nicolò Magini, Carlo Cavazzoni |
Future Gener. Comput. Syst. | 2 |
| 2021 | The Italian research on HPC key technologies across EuroHPCabstractHigh-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects. Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati |
CF | 1 |
| 2021 | TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascaleabstractTo achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research. Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola |
DSD | 19 |
| 2021 | An explainable AI system for automated COVID-19 assessment and lesion categorization from CT-scans
Matteo Pennisi, Isaak Kavasidis, Concetto Spampinato, Vincenzo Schininà, Simone Palazzo, Federica Proietto Salanitri, Giovanni Bellitto, Francesco Rundo, Marco Aldinucci, Massimo Cristofaro, Paolo Campioni, Elisa Pianura, Federica Di Stefano 0002, Ada Petrone, Fabrizio Albarello, Giuseppe Ippolito, Salvatore Cuzzocrea, Sabrina Conoci |
Artif. Intell. Medicine | 9 |
| 2021 | Advantages of using graph databases to explore chromatin conformation capture experimentsabstractBACKGROUND: High-throughput sequencing Chromosome Conformation Capture (Hi-C) allows the study of DNA interactions and 3D chromosome folding at the genome-wide scale. Usually, these data are represented as matrices describing the binary contacts among the different chromosome regions. On the other hand, a graph-based representation can be advantageous to describe the complex topology achieved by the DNA in the nucleus of eukaryotic cells. METHODS: Here we discuss the use of a graph database for storing and analysing data achieved by performing Hi-C experiments. The main issue is the size of the produced data and, working with a graph-based representation, the consequent necessity of adequately managing a large number of edges (contacts) connecting nodes (genes), which represents the sources of information. For this, currently available graph visualisation tools and libraries fall short with Hi-C data. The use of graph databases, instead, supports both the analysis and the visualisation of the spatial pattern present in Hi-C data, in particular for comparing different experiments or for re-mapping omics data in a space-aware context efficiently. In particular, the possibility of describing graphs through statistical indicators and, even more, the capability of correlating them through statistical distributions allows highlighting similarities and differences among different Hi-C experiments, in different cell conditions or different cell types. RESULTS: These concepts have been implemented in NeoHiC, an open-source and user-friendly web application for the progressive visualisation and analysis of Hi-C networks based on the use of the Neo4j graph database (version 3.5). CONCLUSION: With the accumulation of more experiments, the tool will provide invaluable support to compare neighbours of genes across experiments and conditions, helping in highlighting changes in functional domains and identifying new co-organised genomic compartments. Daniele D'Agostino, Pietro Liò, Marco Aldinucci, Ivan Merelli |
BMC Bioinform. | 3 |
| 2021 | Practical parallelization of scientific applications with OpenMP, OpenACC and MPI
Marco Aldinucci, Valentina Cesare, Iacopo Colonnelli, Alberto Riccardo Martinelli, Gianluca Mittone, Barbara Cantalupo, Carlo Cavazzoni, Maurizio Drocco |
J. Parallel Distributed Comput. | 1 |
| 2020 | Practical Parallelization of Scientific ApplicationsabstractThis work aims at distilling a systematic methodology to modernize existing sequential scientific codes with a limited re-designing effort, turning an old codebase into modern code, i.e., parallel and robust code. We propose an automatable methodology to parallelize scientific applications designed with a purely sequential programming mindset, thus possibly using global variables, aliasing, random number generators, and stateful functions. We demonstrate the methodology by way of an astrophysical application, where we model at the same time the kinematic profiles of 30 disk galaxies with a Monte Carlo Markov Chain (MCMC), which is sequential by definition. The parallel code exhibits a 12 times speedup on a 48-core platform. Valentina Cesare, Iacopo Colonnelli, Marco Aldinucci |
PDP | 3 |
| 2020 | Enforcing Deadlines for Skeleton-based Parallel ProgrammingabstractHigh throughput applications with real-time guarantees are increasingly relevant. For these applications, parallelism must be exposed to meet deadlines. Directed Acyclic Graphs (DAGs) are a popular and very general application model that can capture any possible interaction among threads. However, we argue that by constraining the application structure to a set of composable “skeletons”, at the price of losing some generality w.r.t. DAGs, the following advantages are gained: (i) a finer model of the application enables tighter analysis, (ii) specialised scheduling policies are applicable, (iii) programming is simplified, (iv) specialised implementation techniques can be exploited transparently, and (v) the program can be automatically tuned to minimise resource usage while still meeting its hard deadlines. As a first step towards a set of real-time skeletons we conduct a case study with the job farm skeleton and the hard real-time XMOS xCore-200 microcontroller. We present an analytical framework for job farms that reduces the number of required cores by scheduling jobs in batches, while ensuring that deadlines are still met. Our experimental results demonstrate that batching reduces the minimum sustainable period by up to 22%, leading to a reduced number of required cores. The framework chooses the best parameters in 83% of cases and never selects parameters that cause deadline misses. Finally, we show that the overheads introduced by the skeleton abstraction layer are negligible. Paul Metzger, Murray Cole, Christian Fensch, Marco Aldinucci, Enrico Bini |
RTAS | 4 |
| 2020 | Data stream processing in HPC systems: New frameworks and architectures for high-frequency streaming
Marco Aldinucci, Valeria Cardellini, Gabriele Mencagli, Massimo Torquati |
Parallel Comput. | 1 |
| 2020 | Programming languages for data-Intensive HPC applications: A systematic mapping study
Vasco Amaral 0001, Beatriz Norberto, Miguel Goulão, Marco Aldinucci, Siegfried Benkner, Andrea Bracciali, Paulo Carreira 0001, Edgars Celms, Luís Correia 0001, Clemens Grelck, Helen D. Karatza, Christoph W. Kessler, Peter Kilpatrick, Hugo F. M. C. Martiniano, Ilias Mavridis, Sabri Pllana, Ana Respício, José Simão, Luís Veiga, Ari Visa |
Parallel Comput. | 4 |
| 2020 | Challenging the abstraction penalty in parallel patterns libraries
José Daniel García, David del Rio Astorga, Marco Aldinucci, Fabio Tordini, Marco Danelutto, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 3 |
| 2019 | Accelerating Spectral Graph Analysis Through Wavefronts of Linear Algebra OperationsabstractThe wavefront pattern captures the unfolding of a parallel computation in which data elements are laid out as a logical multidimensional grid and the dependency graph favours a diagonal sweep across the grid. In the emerging area of spectral graph analysis, the computing often consists in a wavefront running over a tiled matrix, involving expensive linear algebra kernels. While these applications might benefit from parallel heterogeneous platforms (multi-core with GPUs), programming wavefront applications directly with high-performance linear algebra libraries yields code that is complex to write and optimize for the specific application. We advocate a methodology based on two abstractions (linear algebra and parallel pattern-based run-time), that allows to develop portable, self-configuring, and easy-to-profile code on hybrid platforms. Maurizio Drocco, Paolo Viviani 0001, Iacopo Colonnelli, Marco Aldinucci, Marco Grangetto |
PDP | 4 |
| 2019 | Deep Learning at ScaleabstractThis work presents a novel approach to distributed training of deep neural networks (DNNs) that aims to overcome the issues related to mainstream approaches to data parallel training. Established techniques for data parallel training are discussed from both a parallel computing and deep learning perspective, then a different approach is presented that is meant to allow DNN training to scale while retaining good convergence properties. Moreover, an experimental implementation is presented as well as some preliminary results. Paolo Viviani 0001, Maurizio Drocco, Daniele Baccega, Iacopo Colonnelli, Marco Aldinucci |
PDP | 5 |
| 2019 | Power-aware pipelining with automatic concurrency controlabstractSummary Continuous streaming computations are usually composed of different modules, exchanging data through shared message queues. The selection of the algorithm used to access such queues (ie, theconcurrency control) is a critical aspect both for performance and power consumption. In this paper, we describe the design of automatic concurrency control algorithm for implementing power‐efficient communications on shared‐memory multicores. The algorithm automatically switches between nonblocking and blocking concurrency protocols, getting the best from the two worlds, ie, obtaining the same throughput offered by the nonblocking implementation and the same power efficiency of the blocking concurrency protocol. We demonstrate the effectiveness of our approach using two micro‐benchmarks and two real streaming applications. Massimo Torquati, Daniele De Sensi, Gabriele Mencagli, Marco Aldinucci, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Exploiting Docker containers over Grid computing for a comprehensive study of chromatin conformation in different cell types
Ivan Merelli, Federico Fornari, Fabio Tordini, Daniele D'Agostino, Marco Aldinucci, Daniele Cesini |
J. Parallel Distributed Comput. | 5 |
| 2019 | On dynamic memory allocation in sliding-window parallel patterns for streaming analytics
Massimo Torquati, Gabriele Mencagli, Maurizio Drocco, Marco Aldinucci, Tiziano De Matteis, Marco Danelutto |
J. Supercomput. | 4 |
| 2018 | HPC4AI: an AI-on-demand federated platform endeavourabstractIn April 2018, under the auspices of the POR-FESR 2014-2020 program of Italian Piedmont Region, the Turin's Centre on High-Performance Computing for Artificial Intelligence (HPC4AI) was funded with a capital investment of 4.5M€ and it began its deployment. HPC4AI aims to facilitate scientific research and engineering in the areas of Artificial Intelligence and Big Data Analytics. HPC4AI will specifically focus on methods for the on-demand provisioning of AI and BDA Cloud services to the regional and national industrial community, which includes the large regional ecosystem of Small-Medium Enterprises (SMEs) active in many different sectors such as automotive, aerospace, mechatronics, manufacturing, health and agrifood. Marco Aldinucci, Sergio Rabellino, Marco Pironti, Filippo Spiga, Paolo Viviani 0001, Maurizio Drocco, Marco Guerzoni, Guido Boella, Marco Mellia, Paolo Margara, Idilio Drago, Roberto Marturano, Guido Marchetto, Elio Piccolo, Stefano Bagnasco, Stefano Lusso, Sara Vallero, Giuseppe Attardi, Alex Barchiesi, Alberto Colla, Fulvio Galeazzi |
CF | 1 |
| 2018 | Scaling Dense Linear Algebra on Multicore and Beyond: A SurveyabstractThe present trend in big-data analytics is to exploit algorithms with (sub-)linear time complexity, in this sense it is usually worth to investigate if the available techniques can be approximated to reach an affordable complexity. However, there are still problems in data science and engineering that involve algorithms with higher time complexity, like matrix inversion or Singular Value Decomposition (SVD). This work presents the results of a survey that reviews a number of tools meant to perform dense linear algebra at “Big Data” scale: namely, the proposed approach aims first to define a feasibility boundary for the problem size of shared-memory matrix factorizations, then to understand whether it is convenient to employ specific tools meant to scale out such dense linear algebra tasks on distributed platforms. The survey will eventually discuss the presented tools from the point of view of domain experts (data scientist, engineers), hence focusing on the trade-off between usability and performance. Paolo Viviani 0001, Maurizio Drocco, Marco Aldinucci |
PDP | 3 |
| 2018 | PiCo: High-performance data analytics pipelines in modern C++
Claudia Misale, Maurizio Drocco, Guy Tremblay, Alberto Riccardo Martinelli, Marco Aldinucci |
Future Gener. Comput. Syst. | 5 |
| 2018 | Harnessing sliding-window execution semantics for parallel stream processing
Gabriele Mencagli, Massimo Torquati, Fabio Lucattini, Salvatore Cuomo, Marco Aldinucci |
J. Parallel Distributed Comput. | 5 |
| 2018 | A parallel pattern for iterative stencil + reduce
Marco Aldinucci, Marco Danelutto, Maurizio Drocco, Peter Kilpatrick, Claudia Misale, Guilherme Peretti Pezzi, Massimo Torquati |
J. Supercomput. | 1 |
| 2017 | Deep learning for automated skeletal bone age assessment in X-ray images
Concetto Spampinato, Simone Palazzo, Daniela Giordano, Marco Aldinucci, Rosalia Leonardi |
Medical Image Anal. | 4 |
| 2016 | A Cluster-as-Accelerator Approach for SPMD-Free Data ParallelismabstractIn this paper we present a novel approach for functional-style programming of distributed-memory clusters, targeting data-centric applications. The programming model proposed is purely sequential, SPMD-free and based on high-level functional features introduced since C++11 specification. Additionally, we propose a novel cluster-as-accelerator design principle. In this scheme, cluster nodes act as general interpreters of user-defined functional tasks over node-local portions of distributed data structures. We envision coupling a simple yet powerful programming model with a lightweight, locality-aware distributed runtime as a promising step along the road towards high-performance data analytics, in particular under the perspective of the upcoming exascale era. We implemented the proposed approach in SkeDaTo, a prototyping C++ library of data-parallel skeletons exploiting cluster-as-accelerator at the bottom layer of the runtime software stack. Maurizio Drocco, Claudia Misale, Marco Aldinucci |
PDP | 3 |
| 2016 | RPL: A Domain-Specific Language for Designing and Implementing Parallel C++ ApplicationsabstractParallelising sequential applications is usually a very hard job, due to many different ways in which an application can be parallelised and a large number of programming models (each with its own advantages and disadvantages) that can be used. In this paper, we describe a method to semi-automatically generate and evaluate different parallelisations of the same application, allowing programmers to find the best parallelisation without significant manual reengineering of the code. We describe a novel, high-level domain-specific language, Refactoring Pattern Language (RPL), that is used to represent the parallel structure of an application and to capture its extra-functional properties (such as service time). We then describe a set of RPL rewrite rules that can be used to generate alternative, but semantically equivalent, parallel structures (parallelisations) of the same application. We also describe the RPL Shell that can be used to evaluate these parallelisations, in terms of the desired extra-functional properties. Finally, we describe a set of C++ refactorings, targeting OpenMP, Intel TBB and FastFlow parallel programming models, that semi-automatically apply the desired parallelisation to the application's source code, therefore giving a parallel version of the code. We demonstrate how the RPL and the refactoring rules can be used to derive efficient parallelisations of two realistic C++ use cases (Image Convolution and Ant Colony Optimisation). Vladimir Janjic, Christopher Brown 0002, Kenneth MacKenzie, Kevin Hammond, Marco Danelutto, Marco Aldinucci, José Daniel García |
PDP | 6 |
| 2016 | PWHATSHAP: efficient haplotyping for future generation sequencingabstractBACKGROUND: Haplotype phasing is an important problem in the analysis of genomics information. Given a set of DNA fragments of an individual, it consists of determining which one of the possible alleles (alternative forms of a gene) each fragment comes from. Haplotype information is relevant to gene regulation, epigenetics, genome-wide association studies, evolutionary and population studies, and the study of mutations. Haplotyping is currently addressed as an optimisation problem aiming at solutions that minimise, for instance, error correction costs, where costs are a measure of the confidence in the accuracy of the information acquired from DNA sequencing. Solutions have typically an exponential computational complexity. WHATSHAP is a recent optimal approach which moves computational complexity from DNA fragment length to fragment overlap, i.e., coverage, and is hence of particular interest when considering sequencing technology's current trends that are producing longer fragments. RESULTS: Given the potential relevance of efficient haplotyping in several analysis pipelines, we have designed and engineered PWHATSHAP, a parallel, high-performance version of WHATSHAP. PWHATSHAP is embedded in a toolkit developed in Python and supports genomics datasets in standard file formats. Building on WHATSHAP, PWHATSHAP exhibits the same complexity exploring a number of possible solutions which is exponential in the coverage of the dataset. The parallel implementation on multi-core architectures allows for a relevant reduction of the execution time for haplotyping, while the provided results enjoy the same high accuracy as that provided by WHATSHAP, which increases with coverage. CONCLUSIONS: Due to its structure and management of the large datasets, the parallelisation of WHATSHAP posed demanding technical challenges, which have been addressed exploiting a high-level parallel programming framework. The result, PWHATSHAP, is a freely available toolkit that improves the efficiency of the analysis of genomics information. Andrea Bracciali, Marco Aldinucci, Murray Patterson, Tobias Marschall, Nadia Pisanti, Ivan Merelli, Massimo Torquati |
BMC Bioinform. | 2 |
| 2015 | Memory-Optimised Parallel Processing of Hi-C DataabstractThis paper presents the optimisation efforts on the creation of a graph-based mapping representation of gene adjacency. The method is based on the Hi-C process, starting from Next Generation Sequencing data, and it analyses a huge amount of static data in order to produce maps for one or more genes. Straightforward parallelisation of this scheme does not yield acceptable performance on multicore architectures since the scalability is rather limited due to the memory bound nature of the problem. This work focuses on the memory optimisations that can be applied to the graph construction algorithm and its (complex) data structures to derive a cache-oblivious algorithm and eventually to improve the memory bandwidth utilisation. We used as running example not, a tool for annotation and statistic analysis of Hi-C data that creates a gene-centric neighborhood graph. The proposed approach, which is exemplified for Hi-C, addresses several common issue in the parallelisation of memory bound algorithms for multicore. Results show that the proposed approach is able to increase the parallel speedup from 7x to 22x (on a 32-core platform). Finally, the proposed C++ implementation outperforms the first R Nu Chart prototype, by which it was not possible to complete the graph generation because of strong memory-saturation problems. Maurizio Drocco, Claudia Misale, Guilherme Peretti Pezzi, Fabio Tordini, Marco Aldinucci |
PDP | 5 |
| 2015 | Parallel Exploration of the Nuclear Chromosome Conformation with NuChart-IIabstractHigh-throughput molecular biology techniques are widely used to identify physical interactions between genetic elements located throughout the human genome. Chromosome Conformation Capture (3C) and other related techniques allow to investigate the spatial organisation of chromosomes in the cell's natural state. Recent results have shown that there is a large correlation between co-localization and co-regulation of genes, but these important information are hampered by the lack of biologists-friendly analysis and visualisation software. In this work we introduce NuChart-II, a tool for Hi-C data analysis that provides a gene-centric view of the chromosomal neighbourhood in a graph-based manner. NuChart-II is an efficient and highly optimized C++ re-implementation of a previous prototype package developed in R. Representing Hi-C data using a graph-based approach overcomes the common view relying on genomic coordinates and permits the use of graph analysis techniques to explore the spatial conformation of a gene neighbourhood. Fabio Tordini, Maurizio Drocco, Claudia Misale, Luciano Milanesi, Pietro Liò, Ivan Merelli, Marco Aldinucci |
PDP | 7 |
| 2014 | Parallel stochastic systems biology in the cloudabstractThe stochastic modelling of biological systems, coupled with Monte Carlo simulation of models, is an increasingly popular technique in bioinformatics. The simulation-analysis workflow may result computationally expensive reducing the interactivity required in the model tuning. In this work, we advocate the high-level software design as a vehicle for building efficient and portable parallel simulators for the cloud. In particular, the Calculus of Wrapped Components (CWC) simulator for systems biology, which is designed according to the FastFlow pattern-based approach, is presented and discussed. Thanks to the FastFlow framework, the CWC simulator is designed as a high-level workflow that can simulate CWC models, merge simulation results and statistically analyse them in a single parallel workflow in the cloud. To improve interactivity, successive phases are pipelined in such a way that the workflow begins to output a stream of analysis results immediately after simulation is started. Performance and effectiveness of the CWC simulator are validated on the Amazon Elastic Compute Cloud. Marco Aldinucci, Massimo Torquati, Concetto Spampinato, Maurizio Drocco, Claudia Misale, Cristina Calcagno, Mario Coppo |
Briefings Bioinform. | 1 |
| 2014 | Decision tree building on multi-core using FastFlowabstractSUMMARY The whole computer hardware industry embraced the multi‐core. The extreme optimisation of sequential algorithms is then no longer sufficient to squeeze the real machine power, which can be only exploited via thread‐level parallelism. Decision tree algorithms exhibit natural concurrency that makes them suitable to be parallelised. This paper presents an in‐depth study of the parallelisation of an implementation of the C4.5 algorithm for multi‐core architectures. We characterise elapsed time lower bounds for the forms of parallelisations adopted and achieve close to optimal performance. Our implementation is based on the FastFlow parallel programming environment, and it requires minimal changes to the original sequential code. Copyright © 2013 John Wiley & Sons, Ltd. Marco Aldinucci, Salvatore Ruggieri, Massimo Torquati |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Discovering biological knowledge by integrating high-throughput data and scientific literature on the cloudabstractSUMMARY In this paper, we present a bioinformatics knowledge discovery tool for extracting and validating associations between biological entities. By mining specialized scientific literature, the tool not only generates biological hypotheses in the form of associations between genes, proteins, miRNA and diseases but also validates the plausibility of such associations against high‐throughput biological data (e.g. microarray) and annotated databases (e.g. Gene Ontology). Both the knowledge discovery system and its validation are carried out by exploiting the advantages and the potentialities of the Cloud, which allowed us to derive and check the validity of thousands of biological associations in a reasonable amount of time. The system was tested on a dataset containing more than 1000 gene–disease associations achieving an average recall of about 71%, outperforming existing approaches. The results also showed that porting a data‐intensive application in an Infrastructure as a Service cloud environment boosts significantly the application's efficiency. Copyright © 2013 John Wiley & Sons, Ltd. Concetto Spampinato, Isaak Kavasidis, Marco Aldinucci, Carmelo Pino, Daniela Giordano, Alberto Faro |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Parallel Stochastic Simulators in System Biology: The Evolution of the SpeciesabstractThe stochastic simulation of biological systems is an increasingly popular technique in Bioinformatics. It is often an enlightening technique, especially for multi-stable systems which dynamics can be hardly captured with ordinary differential equations. To be effective, stochastic simulations should be supported by powerful statistical analysis tools. The simulation-analysis workflow may however result in being computationally expensive, thus compromising the interactivity required in model tuning. In this work we advocate the high-level design of simulators for stochastic systems as a vehicle for building efficient and portable parallel simulators. In particular, the Calculus of Wrapped Components (CWC) simulator, which is designed according to the FastFlow's pattern-based approach, is presented and discussed in this work. FastFlow has been extended to support also clusters of multi-cores with minimal coding effort, assessing the portability of the approach. Marco Aldinucci, Maurizio Drocco, Fabio Tordini, Mario Coppo, Massimo Torquati |
PDP | 1 |
| 2012 | An Efficient Unbounded Lock-Free Queue for Multi-core Systems
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati |
Euro-Par | 1 |
| 2012 | Parallel Patterns + Macro Data Flow for Multi-core ProgrammingabstractData flow techniques have been around since the early '70s when they were used in compilers for sequential languages. Shortly after their introduction they were also considered as a possible model for parallel computing, although the impact here was limited. Recently, however, data flow has been identified as a candidate for efficient implementation of various programming models on multi-core architectures. In most cases, however, the burden of determining data flow ``macro'' instructions is left to the programmer, while the compiler/run time system manages only the efficient scheduling of these instructions. We discuss a structured parallel programming approach supporting automatic compilation of programs to macro data flow and we show experimental results demonstrating the feasibility of the approach and the efficiency of the resulting ``object'' code on different classes of state-of-the-art multi-core architectures. The experimental results use different base mechanisms to implement the macro data flow run time support, from plain pthreads with condition variables to more modern and effective lock- and fence-free parallel frameworks. Experimental results comparing efficiency of the proposed approach with those achieved using other, more classical, parallel frameworks are also presented. Marco Aldinucci, L. Anardu, Marco Danelutto, Massimo Torquati, Peter Kilpatrick |
PDP | 1 |
| 2011 | Accelerating Code on Multi-cores with FastFlow
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati |
Euro-Par (2) | 1 |
| 2011 | On Designing Multicore-Aware Simulators for Biological SystemsabstractThe stochastic simulation of biological systems is an increasingly popular technique in bioinformatics. It often is an enlightening technique, which may however result in being computational expensive. We discuss the main opportunities to speed it up on multi-core platforms, which pose new challenges for parallelisation techniques. These opportunities are developed in two general families of solutions involving both the single simulation and a bulk of independent simulations (either replicas of derived from parameter sweep). Proposed solutions are tested on the parallelisation of the CWC simulator (Calculus of Wrapped Compartments) that is carried out according to proposed solutions by way of the Fast Flow programming framework making possible fast development and efficient execution on multi-cores. Marco Aldinucci, Mario Coppo, Ferruccio Damiani, Maurizio Drocco, Massimo Torquati, Angelo Troina |
PDP | 1 |
| 2010 | Efficient Smith-Waterman on Multi-core with FastFlowabstractShared memory multiprocessors have returned to popularity thanks to rapid spreading of commodity multi-core architectures. However, little attention has been paid to supporting effective streaming applications on these architectures. In this paper we describe FastFlow, a low-level programming framework based on lock-free queues explicitly designed to support high-level languages for streaming applications. We compare FastFlow with state-of-the-art programming frameworks such as Cilk, OpenMP, and Intel TBB. We experimentally demonstrate that FastFlow is always more efficient than them on a given real world application: the speedup of FastFlow over other solutions may be substantial for fine grain tasks, for example +35% over OpenMP, +226% over Cilk, +96% over TBB for the alignment of protein P01111 against UniProt DB using the Smith-Waterman algorithm. Marco Aldinucci, Massimiliano Meneghin, Massimo Torquati |
PDP | 1 |
| 2010 | Porting Decision Tree Algorithms to Multicore Using FastFlow
Marco Aldinucci, Salvatore Ruggieri, Massimo Torquati |
ECML/PKDD (1) | 1 |
| 2009 | Stkm on Sca: A Unified Framework with Components, Workflows and Algorithmic Skeletons
Marco Aldinucci, Hinde-Lilia Bouziane, Marco Danelutto, Christian Pérez |
Euro-Par | 1 |
| 2009 | Autonomic management of non-functional concerns in distributed & parallel application programmingabstractAn approach to the management of non-functional concerns in massively parallel and/or distributed architectures that marries parallel programming patterns with autonomic computing is presented. The necessity and suitability of the adoption of autonomic techniques are evidenced. Issues arising in the implementation of autonomic managers taking care of multiple concerns and of coordination among hierarchies of such autonomic managers are discussed. Experimental results are presented that demonstrate the feasibility of the approach. Marco Aldinucci, Marco Danelutto, Peter Kilpatrick |
IPDPS | 1 |
| 2009 | Towards Hierarchical Management of Autonomic Components: A Case StudyabstractWe address the issue of autonomic management in hierarchical component-based distributed systems. The long term aim is to provide a modeling framework for autonomic management in which QoS goals can be defined, plans for system adaptation described and proofs of achievement of goals by (sequences of) adaptations furnished. Here we present an early step on this path. We restrict our focus to skeleton-based systems in order to exploit their well-defined structure. The autonomic cycle is described using the Orc system orchestration language while the plans are presented as structural modifications together with associated costs and benefits. A case study is presented to illustrate the interaction of managers to maintain QoS goals for throughput under varying conditions of resource availability. Marco Aldinucci, Marco Danelutto, Peter Kilpatrick |
PDP | 1 |
| 2008 | Behavioural Skeletons in GCM: Autonomic Management of Grid ComponentsabstractAutonomic management can be used to improve the QoS provided by parallel/distributed applications. We discuss behavioural skeletons introduced in earlier work: rather than relying on programmer ability to design "from scratch" efficient autonomic policies, we encapsulate general autonomic controller features into algorithmic skeletons. Then we leave to the programmer the duty of specifying the parameters needed to specialise the skeletons to the needs of the particular application at hand. This results in the programmer having the ability to fast prototype and tune distributed/parallel applications with non-trivial autonomic management capabilities. We discuss how behavioural skeletons have been implemented in the framework of GCM (the grid component model developed within the CoreGRID NoE and currently being implemented within the GridCOMP STREP project). We present results evaluating the overhead introduced by autonomic management activities as well as the overall behaviour of the skeletons. We also present results achieved with a long running application subject to autonomic management and dynamically adapting to changing features of the target architecture. Overall the results demonstrate both the feasibility of implementing autonomic control via behavioural skeletons and the effectiveness of our sample behavioural skeletons in managing the "functional replication" pattern(s). Marco Aldinucci, Sonia Campa, Marco Danelutto, Marco Vanneschi, Peter Kilpatrick, Patrizio Dazzi, Domenico Laforenza, Nicola Tonellotto |
PDP | 1 |
| 2008 | The VirtuaLinux Storage Abstraction Layer for Ef?cient Virtual ClusteringabstractVirtuaLinux is a meta-distribution that enables a standard Linux distribution to support robust physical and virtualized clusters. VirtuaLinux helps in avoiding the "single point of failure" effect by means of a combination of architectural strategies, including the transparent support for disk-less and master-less cluster configuration. VirtuaLinux supports the creation and management of Virtual Clusters in seamless way: VirtuaLinux Virtual Cluster Manager enables the system administrator to create, save, restore Xen- based Virtual Clusters, and to map and dynamically re-map them onto the nodes of the physical cluster. In this paper we introduce and discuss VirtuaLinux virtualization architecture, features, and tools, and in particular, the novel disk abstraction layer, which permits the fast and space-efficient creation of Virtual Clusters. Marco Aldinucci, Massimo Torquati, Marco Vanneschi, Pierfrancesco Zuccato |
PDP | 1 |
| 2008 | Securing skeletal systems with limited performance penalty: The muskel
Marco Aldinucci, Marco Danelutto |
J. Syst. Archit. | 1 |
| 2007 | Management in Distributed Systems: A Semi-formal Approach
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick |
Euro-Par | 1 |
| 2007 | The cost of security in skeletal systemsabstractSkeletal systems exploit algorithmical skeletons technology to provide the user very high level, efficient parallel programming environments. They have been recently demonstrated to be suitable for highly distributed architectures, such as workstation clusters, networks and grids. However, when using skeletal system for grid programming care must be taken to secure data and code transfers across non-dedicated, non-secure network links. In this work we take into account the cost of security introduction in muskel, a Java based skeletal system exploiting macro data flow implementation technology. We consider the adoption of mechanisms that allow securing all the communications taking place between remote, unreliable nodes and we evaluate the cost of such mechanisms. In particular, we consider the implications on the computational grains needed to scale secure and insecure skeletal computations. Marco Aldinucci, Marco Danelutto |
PDP | 1 |
| 2007 | Skeleton-based parallel programming: Functional and parallel semantics in a single shot
Marco Aldinucci, Marco Danelutto |
Comput. Lang. Syst. Struct. | 1 |
| 2006 | Autonomic QoS in ASSIST Grid-Aware ComponentsabstractCurrent grid-aware applications are developed on existing software infrastructures, such as Globus, by developers who are experts on grid software implementation. Although many useful applications have been produced this way, this approach may hardly support the additional complexity to quality of service (QoS) control in real application. We describe the ASSIST programming environment, the prototype of parallel programming environment currently under development at our group, as a suitable basis to capture all the desired features for QoS control for the grid. Grid applications, built as compositions of ASSIST components, are supported by an innovative grid abstract machine, which includes essential abstractions of standard middleware services and a hierarchical application manager, which may be considered as an early prototype of autonomic manager. Marco Aldinucci, Marco Danelutto, Marco Vanneschi |
PDP | 1 |
| 2006 | Algorithmic skeletons meeting grids
Marco Danelutto, Marco Aldinucci |
Parallel Comput. | 2 |
| 2005 | Dynamic Reconfiguration of Grid-Aware Applications in ASSIST
Marco Aldinucci, Alessandro Petrocelli, Edoardo Pistoletti, Massimo Torquati, Marco Vanneschi, Luca Veraldi, Corrado Zoccolo |
Euro-Par | 1 |
| 2004 | Targeting Heterogeneous Architectures in ASSIST: Experimental Results
Marco Aldinucci, Sonia Campa, Massimo Coppola, Silvia Magini, Paolo Pesciullesi, Laura Potiti, Roberto Ravazzolo, Massimo Torquati, Corrado Zoccolo |
Euro-Par | 1 |
| 2004 | Accelerating Apache Farms Through Ad-HOC Distributed Scalable Object Repository
Marco Aldinucci, Massimo Torquati |
Euro-Par | 1 |
| 2004 | Topic 8: Parallel Computer Architecture and Instruction-Level Parallelism
Kemal Ebcioglu, Wolfgang Karl, André Seznec, Marco Aldinucci |
Euro-Par | 4 |
| 2003 | ASSIST Demo: A High Level, High Performance Portable, Structured Parallel Programming Environment at Work
Marco Aldinucci, Sonia Campa, Pierpaolo Ciullo, Massimo Coppola, Marco Danelutto, Paolo Pesciullesi, Roberto Ravazzolo, Massimo Torquati, Marco Vanneschi, Corrado Zoccolo |
Euro-Par | 1 |
| 2003 | The Implementation of ASSIST, an Environment for Parallel and Distributed Programming
Marco Aldinucci, Sonia Campa, Pierpaolo Ciullo, Massimo Coppola, Silvia Magini, Paolo Pesciullesi, Laura Potiti, Roberto Ravazzolo, Massimo Torquati, Marco Vanneschi, Corrado Zoccolo |
Euro-Par | 1 |
| 2003 | An advanced environment supporting structured parallel programming in Java
Marco Aldinucci, Marco Danelutto, P. Teti |
Future Gener. Comput. Syst. | 1 |