Paolo Sylos Labini

dblp:243/7280 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-7950-4396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2024 Scaling Expected Force: Efficient Identification of Key Nodes in Network-Based Epidemic Models
abstract
Structural centrality measures are often used to approximate or predict dynamical influence in a network. The recently proposed Expected Force of Infection (ExF) measures the entropy of all potential transmission paths starting at a node, effectively characterizing a node's role in epidemic diffusion processes. However, this promising metric has seen limited adoption mainly due to an inefficient formulation and the lack of an open-source implementation. In this paper, we present a novel cluster-centric, parallel algorithm enhancing ExF's efficiency and scalability. Compared to the simple parallel version of the original formulation of the ExF our efficient, open-source GPU implementation enables key nodes detection at previously intractable scales, with speed-ups of up to 300 x on networks with up to 44 million edges. Leveraging on our algorithm, we compare the ExF with other well-known centrality metrics, upon six real and synthetic contact networks. The ExF emerges as the best of the considered metrics in a few, important tasks: it predicts the likelihood of a global epidemic and its diffusion speed, based on the centrality of the seed node; and it predicts how many other infections will occur as a consequence, in some sense, of a specific node having caught the disease.
Paolo Sylos Labini, Andrej Jurco, Matteo Ceccarello, Stefano Guarino, Enrico Mastrostefano, Flavio Vella
PDP1
2024 High Performance Unstructured SpMM Computation Using Tensor Cores
abstract
High-performance sparse matrix-matrix (SpMM) multiplication is paramount for science and industry, as the ever-increasing sizes of data prohibit using dense data structures. Yet, existing hardware, such as Tensor Cores (TC), is ill-suited for SpMM, as it imposes strict constraints on data structures that cannot be met by unstructured sparsity found in many applications. To address this, we introduce (S)parse (Ma)trix Matrix (T)ensor Core-accelerated (SMaT): a novel SpMM library that utilizes TCs for unstructured sparse matrices. Our block-sparse library leverages the low-level CUDA MMA (matrix-matrix-accumulate) API, maximizing the performance offered by modern GPUs. Algorithmic optimizations such as sparse matrix permutation, further improve performance by minimizing the number of non-zero blocks. The evaluation on NVIDIA A100 shows that SMaT outperforms SotA libraries (DASP, cuSPARSE, and Magicube) by up to 125x (on average 2.6x). SMaT can be used to accelerate many workloads in scientific computing, large model training, inference, and others.
Patrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta, Flavio Vella, Torsten Hoefler
SC3
2023 High-Performance and Programmable Attentional Graph Neural Networks with Global Tensor Formulations
abstract
Graph attention models (A-GNNs), a type of Graph Neural Networks (GNNs), have been shown to be more powerful than simpler convolutional GNNs (C-GNNs). However, A-GNNs are more complex to program and difficult to scale. To address this, we develop a novel mathematical formulation, based on tensors that group all the feature vectors, targeting both training and inference of A-GNNs. The formulation enables straightforward adoption of communication-minimizing routines, it fosters optimizations such as vectorization, and it enables seamless integration with established linear algebra DSLs or libraries such as GraphBLAS. Our implementation uses a data redistribution scheme explicitly developed for sparse-dense tensor operations used heavily in GNNs, and fusing optimizations that further minimize memory usage and communication cost. We ensure theoretical asymptotic reductions in communicated data compared to the established message-passing GNN paradigm. Finally, we provide excellent scalability and speedups of even 4--5x over modern libraries such as Deep Graph Library.
Maciej Besta, Pawel Renc, Robert Gerstenberger, Paolo Sylos Labini, Alexandros Nikolaos Ziogas, Tiancheng Chen, Lukas Gianinazzi, Florian Scheidl, Kalman Szenes, Armon Carigiet, Patrick Iff, Grzegorz Kwasniewski, Raghavendra Kanakagiri, Chio Ge, Sammy Jaeger, Jaroslaw Was, Flavio Vella, Torsten Hoefler
SC4
2022 ProbGraph: High-Performance and High-Accuracy Graph Mining with Probabilistic Set Representations
abstract
Important graph mining problems such as Clustering are computationally demanding. To significantly accelerate these problems, we propose ProbGraph: a graph representation that enables simple and fast approximate parallel graph mining with strong theoretical guarantees on work, depth, and result accuracy. The key idea is to represent sets of vertices using probabilistic set representations such as Bloom filters. These representations are much faster to process than the original vertex sets thanks to vectorizability and small size. We use these representations as building blocks in important parallel graph mining algorithms such as Clique Counting or Clustering. When enhanced with ProbGraph, these algorithms significantly outperform tuned parallel exact baselines (up to nearly 50 x on 32 cores) while ensuring accuracy of more than 90% for many input graph datasets. Our novel bounds and algorithms based on probabilistic set representations with desirable statistical properties are of separate interest for the data analytics community. Proofs of theorems & more results: http://arxiv.org/abs/2208.11469
Maciej Besta, Cesare Miglioli, Paolo Sylos Labini, Jakub Tetek, Patrick Iff, Raghavendra Kanakagiri, Saleh Ashkboos, Kacper Janda, Michal Podstawski, Grzegorz Kwasniewski, Niels Gleinig, Flavio Vella, Onur Mutlu, Torsten Hoefler
SC3
2021 Analysis of SARS-CoV-2 protein interactome map
abstract
By calculating the centrality measures of the nodes of the SARS-CoV-2 protein interactome network, we have identified the viral proteins of potential greatest interest for further experimental investigation to understand the mechanisms by which SARS-CoV-2 attacks cells and to identify possible therapeutic targets. The proteins identified in this study including NSP13, NSP7, ORF3a, ORF8a, and ORF8b, were found to be involved in crucial processes of the viral life cycle, and some of them are currently suspected to be antiviral targets. These results thus demonstrate the importance - and the predictive power- of the in silico analysis of the viral interactome to guide and support experimental investigation, which could otherwise be too complex and time-consuming to carry out in clinical and experimental research, given the size and interaction density of the viral protein network and the current still partial knowledge of this new virus.
Paola Lecca, Bruno Carpentieri, Paolo Sylos Labini, Flavio Vella, Emidio Troiani, Attilio Cavezzi
BIBM3
2021 On the Anatomy of Predictive Models for Accelerating GPU Convolution Kernels and Beyond
abstract
Efficient HPC libraries often expose multiple tunable parameters, algorithmic implementations, or a combination of them, to provide optimized routines. The optimal parameters and algorithmic choices may depend on input properties such as the shapes of the matrices involved in the operation. Traditionally, these parameters are manually tuned or set by auto-tuners. In emerging applications such as deep learning, this approach is not effective across the wide range of inputs and architectures used in practice. In this work, we analyze different machine learning techniques and predictive models to accelerate the convolution operator and GEMM. Moreover, we address the problem of dataset generation, and we study the performance, accuracy, and generalization ability of the models. Our insights allow us to improve the performance of computationally expensive deep learning primitives on high-end GPUs as well as low-power embedded GPU architectures on three different libraries. Experimental results show significant improvement in the target applications from 50% up to 300% compared to auto-tuned and high-optimized vendor-based heuristics by using simple decision tree- and MLP-based models.
Paolo Sylos Labini, Marco Cianfriglia, Damiano Perri, Osvaldo Gervasi, Grigori Fursin, Anton Lokhmotov, Cedric Nugteren, Bruno Carpentieri, Fabiana Zollo, Flavio Vella
ACM Trans. Archit. Code Optim.1
2019 Towards a Learning-Based Performance Modeling for Accelerating Deep Neural Networks
Damiano Perri, Paolo Sylos Labini, Osvaldo Gervasi, Sergio Tasso, Flavio Vella
ICCSA (1)2