Ian Di Dio Lavore

dblp:274/4994 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0009-1572-3221ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GASP: GPU-Accelerated Shortest Path for Graph Analytics
abstract
GPUs have shown significant potential for accelerating database analytical queries, but leveraging them for graph analytics queries involving weighted shortest-path computations remains challenging. This stems from the need to efficiently support diverse query patterns that require multi-source, multi-target computations, tracking all shortest paths, ranking paths based on cost, as well as ease of integration with columnar data processing systems.
Ian Di Dio Lavore, Rathijit Sen, Yuanyuan Tian 0001
DaMoN1
2026 Locality-Aware Distributed Allocators for High-Performance Global Data Structures
Ian Di Dio Lavore, Beatrice Branchini, Vito Giovanni Castellana, Marco D. Santambrogio
IPDPS1
2025 Multi-GPU Greedy Scheduling Through a Polyglot Runtime
abstract
Multi-GPU architectures are increasingly being deployed in cloud data centers, but using GPUs efficiently from high-level programming languages remains a challenge.Moreover, exploiting the full capabilities of multi-GPU systems is an arduous task due to the complex interconnection topology between available accelerators and the variety of inter-GPU communication patterns exhibited by different workloads.This work introduces a novel scheduler for multi-task GPU computations that provides transparent asynchronous execution on multi-GPU systems without requiring prior information about the program dependencies or the underlying system architecture.It integrates with the polyglot GraalVM ecosystem and is therefore available for multiple high-level languages, providing a general framework that can significantly lower the barriers to entry to multi-GPU acceleration.We validate our work on representative workloads to investigate scalability and inter-GPU communication.Experimental results show how our scheduler automatically achieves 80-90% peak performance against hand-optimized CUDA host code on Volta and Ampere multi-GPU systems.
Ian Di Dio Lavore, Guido Walter Di Donato, Alberto Parravicini, Francesco Sgherzi, Daniele Bonetta, Marco D. Santambrogio
CF1
2025 Harnessing GPU Acceleration for Exact DNA Sequence Matching via the KMP Algorithm
abstract
Identifying recurrent patterns and mutations in the DNA is essential for helping clinicians formulate faster diagnoses and develop personalized treatments. Here, exact matching represents a core procedure. However, due to its computational intensity, it embodies a bottleneck in many genome analysis pipelines. In this context, the computational efficiency of the chosen algorithm used for exact matching combined with the GPUs’ computing power is crucial to speeding up the process. This paper introduces a high-performance multi-GPU solution for exact DNA sequence matching based on the Knuth-Morris-Pratt (KMP) algorithm, designed to identify all possible occurrences of patterns within a reference DNA. Our solution offers a multi-pattern search that exploits multi-buffering to maximize data reuse and overlap data transfers with computation. Experimental results show that our approach for the KMP algorithm attains 5.83× over a state-of-the-art software for genome analysis. Also, our solution run on an AMD MI210 attains up to 1.59 over the best-performing FPGA solution in the literature. ×
Beatrice Branchini, Pierluigi Negro, Ian Di Dio Lavore, Marco D. Santambrogio
ISCAS3
2025 On the Characterization of GraphML Frameworks: The Case of Semi-Supervised Node Classification
abstract
In recent years, the application of Machine Learning techniques on graphs has produced a considerable interest, leading to the development of many Graph Machine Learning (GraphML) frameworks. However, the proper framework has to be selected depending on the application, requiring time and resources. To solve this issue, this work characterizes three GraphML frameworks: PyTorch Geometric (PyG), Deep Graph Library (DGL), and Stellargraph on four different GPU architectures on the task of Semi-Supervised Node Classification. We compare both the training and inference time and the accuracy and loss curves for each configuration under identical model setups. Results show that PyTorch-based frameworks are faster than those using TensorFlow. Furthermore, we evaluate how DGL has a steeper convergence and can outperform PyG in the presented case study, while the training time per epoch of PyG is faster. Additionally, our evaluation highlights how the frameworks only sometimes fully exploit newer generations of server-grade GPUs. This study demonstrates how selecting the most suitable GraphML framework is a multifaced problem that can directly impact the performance for the end-user.
Alessandro La Conca, Leonardo De Grandis, Ian Di Dio Lavore, Beatrice Branchini, Marco D. Santambrogio
ISCAS3
2025 On the Effectiveness of Unified Memory in Multi-GPU Collective Communication
abstract
Modern supercomputers are becoming increasingly dense with accelerators. Industry leaders offer multi-GPU architectures with high interconnection bandwidth between the devices to match the requirements of modern workloads. While those technologies advance, it is up to the programmer to successfully exploit them. Recognizing this burden, multiple abstractions have been built. We focus on the NVIDIA Collective Communication Library (NCCL) and Unified Memory (UM). The former provides MPI-like directives integrated within the GPU runtime, allowing lower latencies and increasing the bandwidth over previous approaches. The latter simplifies the programming paradigm, offering a unified virtual address space. Moreover, it enables memory oversubscription, drastically reducing the efforts towards handling larger problems without completely restructuring the codebase. This work provides the first joint analysis of NCCL and UM from single-node multi-GPU architectures to a production supercomputer. We explore all the available collective communication directives concerning their power requirements and overall throughput. Moreover, we study the effects of various hyperparameters, e.g., message sizes, oversubscription level, and memory advice, on the overall obtainable performance. Our findings showcase how using UM brings negligible increased energy consumption; moreover, in distributed settings, other restricting factors, such as network bottlenecks, surpass the overhead introduced by UM’s page-eviction mechanisms.
Riccardo Strina, Ian Di Dio Lavore, Marco D. Santambrogio, Michael E. Papka, Zhiling Lan
PDP2
2024 Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems
abstract
The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globally-consistent meta-data. In this paper, we propose a novel data-structure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.
Vito Giovanni Castellana, Burcu O. Mutlu, Ian Di Dio Lavore, Jesun Sahariar Firoz, Katherine E. Wolf, Marco Minutoli, John Feo
IEEE Big Data3