Guido Walter Di Donato

dblp:249/3374 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-2026-9755ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-GPU Greedy Scheduling Through a Polyglot Runtime
abstract
Multi-GPU architectures are increasingly being deployed in cloud data centers, but using GPUs efficiently from high-level programming languages remains a challenge.Moreover, exploiting the full capabilities of multi-GPU systems is an arduous task due to the complex interconnection topology between available accelerators and the variety of inter-GPU communication patterns exhibited by different workloads.This work introduces a novel scheduler for multi-task GPU computations that provides transparent asynchronous execution on multi-GPU systems without requiring prior information about the program dependencies or the underlying system architecture.It integrates with the polyglot GraalVM ecosystem and is therefore available for multiple high-level languages, providing a general framework that can significantly lower the barriers to entry to multi-GPU acceleration.We validate our work on representative workloads to investigate scalability and inter-GPU communication.Experimental results show how our scheduler automatically achieves 80-90% peak performance against hand-optimized CUDA host code on Volta and Ampere multi-GPU systems.
Ian Di Dio Lavore, Guido Walter Di Donato, Alberto Parravicini, Francesco Sgherzi, Daniele Bonetta, Marco D. Santambrogio
CF2
2025 Bridging Research and Entrepreneurship: An Innovative Educational and Experiential Approach
abstract
NECSTLab at Politecnico di Milano is a pioneering research laboratory that integrates cutting-edge academic research with entrepreneurial ventures. Its mission is to bridge the gap between research and real-world applications, encouraging students and researchers to consider the societal impact of their innovations. NECSTLab promotes an interdisciplinary approach, combining technical expertise with business acumen to drive innovation. The lab's strategy includes fostering an entrepreneurial mindset through courses that equip students with the tools to transform research into viable products or services. This paper introduces an innovative approach to address the challenges of transitioning academic research into market commercialization. By focusing on early-stage collaboration between researchers and industry, and incorporating market analysis, prototype development, and business model validation, this process supports the commercialization of research. The proposed pipeline has been implemented and validated within the NECSTLab environment, demonstrating its efficacy in fostering entrepreneurship and translating academic research into successful commercial ventures. This is exemplified by case studies showcasing the pipeline's effectiveness in transforming innovative research into market-ready solutions while fostering entrepreneurial initiatives and driving impactful innovation.
Susanna Bardini, Mirko Coggi, Laura Ginestretti, Guido Walter Di Donato, Marco D. Santambrogio
EDUCON4
2025 GpJSON: High-performance JSON Data Processing on GPUs
abstract
The JavaScript Object Notation (JSON) format is ubiquitous, and countless applications depend on it to store and exchange high volumes of data. Despite its great popularity, JSON is nevertheless a very inefficient data format: decoding and querying JSON data is often a major bottleneck for many data-intensive applications. In this paper, we explore how Graphics Processing Units (GPUs) can be used to parallelize both JSON de-serialization and querying. We show how JSON parsing can be implemented on GPUs by means of parallel structural index construction, and we describe how JSON data can then be queried in situ using a lightweight query engine designed to run on GPUs. We present the design and implementation of GpJSON, a GPU-based JSON data processing library. The library can be used from high-level languages such as JavaScript or Python, and features bindings for the GraalVM language runtime. Our evaluation on real-world datasets shows that, on a single NVIDIA Ampere A100, GpJSON achieves at least 2.9x speedup on end-to-end performance (de-serialization plus querying) over state-of-the-art parallel JSON parsers and query engines, and 6-8 x over NVIDIA RAPIDS.
Sacheendra Talluri, Guido Walter Di Donato, Luca Danelutti, Koen Vlaswinkel, Marco Arnaboldi, Arnaud Delamare, Marco D. Santambrogio, Daniele Bonetta
Proc. VLDB Endow.2
2025 A Novel Methodology for a Comprehensive Analysis of Genomic Sequence-to-Graph Alignment Tools
abstract
Genome graphs have proved to be a more compact and efficient way of representing genetic inter- and intra-individual variability. Although they overcome the traditional sequence-based genome references in many use cases, analyzing genome graphs introduces new computational challenges. The workhorse of graph-based genome analysis is the sequence-to-graph alignment process, which consists of finding the path in the graph that better represents a query sequence. This search is highly computationally intensive, and different solutions have been proposed to solve it efficiently, either by adapting sequence-to-sequence strategies or exploiting novel graph-specific algorithms. However, comparing sequence-to-graph alignment tools is quite challenging because of the complexity and relative novelty of this task, and the resulting lack of standardization. Therefore, here we propose a methodology for a comprehensive and structured comparison of such tools. First, we define a set of KPIs for the qualitative analysis of an aligner's usability, accuracy, and performance. Then, we introduce the first open-sourcehttps://github.com/Mirkocoggi/GGBSbenchmark suite for the quantitative analysis of multiple sequence-to-graph aligners. We test the proposed methodology on state-of-the-art tools, proving how it easily provides valuable insights about the compared aligners. Finally, we conclude the paper by drawing some guidelines to drive the improvement of this promising research field.
Mirko Coggi, Guido Walter Di Donato, Marco D. Santambrogio
IEEE Trans. Comput. Biol. Bioinform.2
2023 On the Genome Sequence Alignment FPGA Acceleration via KSW2z
abstract
Pairwise sequence alignment is a fundamental step for many genomics and molecular biology applications. Given the quadratic time complexity of alignment algorithms, the community demands innovative, fast, and efficient techniques to perform this task. Furthermore, general-purpose architectures lack the necessary performance to address the computational load of these algorithms. In this context, we present the first open-source FPGA implementation of the popular KSW2z algorithm employed by minimap2. Our design also implements the$Z- \mathbf{drop}$heuristic and banded alignment as the original software to further reduce the processing time if needed. The proposed multi-core accelerator achieves up to$\mathbf{7.70}\times$improvement in speedup and$\mathbf{20.07}\times$in energy efficiency compared to the multi-threaded software implementation run on a Xeon Platinum 8167M processor.
Alberto Zeni, Guido Walter Di Donato, Alessia Della Valle, Filippo Carloni, Marco D. Santambrogio
ISCAS2
2022 On the Automation of Radiomics-Based Identification and Characterization of NSCLC
abstract
Proper detection and accurate characterization of Non-Small Cell Lung Cancer (NSCLC) are an open challenge in the imaging field. Biomedical imaging is fundamental in lung cancer assessment and offers the possibility of calculating predictive biomarkers impacting patients' management. Within this context, radiomics, which consists of extracting quantitative features from digital images, shows encouraging results for clinical applications, but the sub-optimal standardization of the procedure and the lack of definitive results are still a concern in the field. For these reasons, this work proposes the design and development of LuCIFEx, a fully-automated pipeline for non-invasive in-vivo characterization of NSCLC, aiming to speed up the analysis process and enable an early diagnosis of the tumor.LuCIFEx pipeline relies on routinely acquired [18F]FDG-PET/CT images for the automatic segmentation of the cancer lesion, allowing the computation of accurate radiomic features, then employed for cancer characterization through Machine Learning algorithms. The proposed multi-stage segmentation process can identify the lesion with a mean accuracy of 94.2±5.0%. Finally, the proposed data analysis pipeline demonstrates the potential of PET/CT features for the automatic recognition of lung metastases and NSCLC histological subtypes, while highlighting the main current limitations of the radiomic approach.
Eleonora D'Arnese, Guido Walter Di Donato, Emanuele Del Sozzo, Martina Sollini, Donatella Sciuto, Marco D. Santambrogio
IEEE J. Biomed. Health Informatics2
2021 The Importance of Being X-Drop: High Performance Genome Alignment on Reconfigurable Hardware
abstract
Pairwise sequence alignment accounts for the majority of key genome analysis applications' runtime. Because of the quadratic time complexity of exact alignment algorithms, the community is moving away from exact algorithms in favor of heuristics that only compute high-quality results. However, the state of the art lacks hardware-accelerated versions of these heuristic algorithms as the vast majority of the available solutions still rely on implementing exact alignment algorithms. Moreover, hardware-based implementations lack high-level APIs that can simplify their integration in commonly used genomic pipelines, hindering their applicability in real-world scenarios. In this context, we present the first high-performance FPGA implementation of the popular X-drop heuristic alignment algorithm and provide an easy-to-use API for its integration. On a Xilinx Alveo U280, our FPGA design achieves up to 5× speed-up over SeqAn, the state-of-the-art software version of the algorithm, running on two Intel Xeon processors using 80 CPU threads. Moreover, our design is also 3.45× faster than ksw2, a state-of-the-art vectorized alignment algorithm that performs a similar heuristic to the one employed in the X-drop algorithm. Finally, our implementation also outperforms LOGAN, a recently published GPU implementation of X-drop running on an Nvidia Tesla V100, by a factor of 1.5×.
Alberto Zeni, Guido Walter Di Donato, Lorenzo Di Tucci, Marco Rabozzi, Marco D. Santambrogio
FCCM2