EDBT 2026 Demo / reviewers in the wild / expert
Alessandro Celestini
dblp:125/2357
· DBLP profile ↗
12ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-8493-4550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the energy efficiency of sparse matrix computations on multi-GPU clustersabstractWe investigate the energy efficiency of a library designed for parallel computations with sparse matrices. The library leverages high-performance, energy-efficient Graphics Processing Unit (GPU) accelerators to enable large-scale scientific applications. Our primary development objective was to maximize parallel performance and scalability in solving sparse linear systems whose dimensions far exceed the memory capacity of a single node. To this end, we devised methods that expose a high degree of parallelism while optimizing algorithmic implementations for efficient multi-GPU usage. Previous work has already demonstrated the library’s performance efficiency on large-scale systems comprising thousands of NVIDIA GPUs, achieving improvements over state-of-the-art solutions. In this paper, we extend those results by providing energy profiles that address the growing sustainability requirements of modern HPC platforms. We present our methodology and tools for accurate runtime energy measurements of the library’s core components and discuss the findings. Our results confirm that optimizing GPU computations and minimizing data movement across memory and computing nodes reduces both time-to-solution and energy consumption. Moreover, we show that the library delivers substantial advantages over comparable software frameworks on standard benchmarks. Massimo Bernaschi, Alessandro Celestini, Giorgio Richelli, Pasqua D'Ambra |
Future Gener. Comput. Syst. | 2 |
| 2025 | Communication-reduced Conjugate Gradient Variants for GPU-accelerated ClustersabstractLinear solvers are key components in any software platform for scientific and engineering computing. The solution of large and sparse linear systems lies at the core of physics-driven numerical simulations relying on partial differential equations (PDEs) and often represents a significant bottleneck in data-driven procedures, such as scientific machine learning. In this paper, we present an efficient implementation of the preconditioned s-step Conjugate Gradient (CG) method, originally proposed by Chronopoulos and Gear in 1989, for large clusters of Nvidia GPU-accelerated computing nodes. The method, often referred to as communication-reduced or communication-avoiding CG, reduces global synchronizations and data communication steps compared to the standard approach, enhancing strong and weak scalability on parallel computers. Our main contribution is the design of a parallel solver that fully exploits the aggregation of low-granularity operations inherent to the s-step CG method to leverage the high throughput of GPU accelerators. Additionally, it applies overlap between data communication and computation in the multi-GPU sparse matrix-vector product. Experiments on classic benchmark datasets, derived from the discretization of the Poisson PDE, demonstrate the potential of the method. Massimo Bernaschi, Mauro Carrozzo, Alessandro Celestini, Giacomo Piperno, Pasqua D'Ambra |
PDP | 3 |
| 2025 | Multi GPU Sparse Matrix by Sparse Matrix MultiplicationabstractABSTRACT The paper focuses on the improvement of the existing nsparse Nagasaka et al. algorithm and its extension to the multi‐GPU setting for the application of real engineering problems. In this work, we propose a distributed multi‐GPU framework for SpGEMM that is designed specifically for the nsparse like algorithms. The results show ∼2 times speed‐up for nsparse and close to ideal scalability of the multi‐GPU extension with the number of GPUs. Finally, we test the proposed algorithm in the AMG setting by computing the double SpGEMM product. Artem Mavliutov, Giovanni Isotton, Carlo Janna, Alessandro Celestini, Massimo Bernaschi |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | The TEXTAROSSA Project: Cool all the Way Down to the HardwareabstractThe TEXTAROSSA project aims to bridge the technology gaps that exascale computing systems will face in the near future in order to overcome their performance and energy efficiency challenges. This project provides solutions for improved energy efficiency and thermal control, seamless integration of heterogeneous accelerators in HPC multi-node platforms, and new arithmetic methods. Challenges are tacked through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models, and tools derived from European research. Antonio Filgueras, Giovanni Agosta, Marco Aldinucci, Carlos Álvarez 0001, Pasqua D'Ambra, Massimo Bernaschi, Andrea Biagioni, Daniele Cattaneo 0002, Alessandro Celestini, Massimo Celino, Carlotta Chiarini, Francesca Lo Cicero, Paolo Cretaro, William Fornaciari, Ottorino Frezza, Andrea Galimberti, Francesco Giacomini, Juan Miguel De Haro Ruiz, Francesco Iannone, Daniel Jaschke, Daniel Jiménez-González, Michal Kulczewski, Alberto Leva, Alessandro Lonardo, Michele Martinelli, Xavier Martorell, Simone Montangero, Lucas Morais, Ariel Oleksiak, Paolo Palazzari, Luca Pontisso, Federico Reghenzani, Cristian Rossi, Sergio Saponara, Carlo Saverio Lodi, Francesco Simula, Federico Terraneo, Piero Vicini, Miquel Vidal, Davide Zoni, Giuseppe Zummo |
DSD | 9 |
| 2023 | A Multi-GPU Aggregation-Based AMG Preconditioner for Iterative Linear SolversabstractWe present and release in open source format a sparse linear solver which efficiently exploits heterogeneous parallel computers. The solver can be easily integrated into scientific applications that need to solve large and sparse linear systems on modern parallel computers made of hybrid nodes hosting Nvidia Graphics Processing Unit (GPU) accelerators. The work extends previous efforts of some of the authors in the exploitation of a single GPU accelerator and proposes an implementation, based on the hybrid MPI-CUDA software environment, of a Krylov-type linear solver relying on an efficient Algebraic MultiGrid (AMG) preconditioner already available in theBootCMatchGlibrary. Our design for the hybrid implementation has been driven by the best practices for minimizing data communication overhead when multiple GPUs are employed, yet preserving the efficiency of the GPU kernels. Strong and weak scalability results of the new version of the library on well-known benchmark test cases are discussed. Comparisons with the Nvidia AmgX solution show a speedup, in the solve phase, up to 2.0x. Massimo Bernaschi, Alessandro Celestini, Flavio Vella, Pasqua D'Ambra |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Onion under Microscope: An in-depth analysis of the Tor WebabstractAbstract Tor is an open source software that allows accessing various kinds of resources, known as hidden services, while guaranteeing sender and receiver anonymity. Tor relies on a free, worldwide, overlay network, managed by volunteers, that works according to the principles of onion routing in which messages are encapsulated in layers of encryption, analogous to layers of an onion. The Tor Web is the set of web resources that exist on the Tor network, and Tor websites are part of the so-called dark web. Recent research works have evaluated Tor security, its evolution over time, and its thematic organization. Nevertheless, limited information is available about the structure of the graph defined by the network of Tor websites, not to be mistaken with the network of nodes that supports the onion routing. The limited number of entry points that can be used to crawl the network, makes the study of this graph far from being simple. In the present paper we analyze two graph representations of the Tor Web and the relationship between contents and structural features, considering three crawling datasets collected over a five-month time frame. Among other findings, we show that Tor consists of a tiny strongly connected component, in which link directories play a central role, and of a multitude of services that can (only) be reached from there. From this viewpoint, the graph appears inefficient. Nevertheless, if we only consider mutual connections, a more efficient subgraph emerges, that is, probably, the backbone of social interactions in Tor. Massimo Bernaschi, Alessandro Celestini, Marco Cianfriglia, Stefano Guarino, Flavio Lombardi, Enrico Mastrostefano |
World Wide Web | 2 |
| 2019 | Critical nodes reveal peculiar features of human essential genes and protein interactomeabstractNetwork-based ranking methods (e.g., centrality analysis) have found extensive use in systems biology and network medicine for the prediction of essential proteins, for the prioritization of drug targets candidates in the treatment of several pathologies and in biomarker discovery, and for human disease genes identification. We here studied the connectivity of the human protein-protein interaction network (i.e., the interactome) to find the nodes whose removal has the heaviest impact on the network, i.e., maximizes its fragmentation. Such nodes are known as Critical Nodes (CNs). Specifically, we implemented a Critical Node Heuristic (CNH) and compared its performance against other four heuristics based on well known centrality measures. To better understand the structure of the interactome, the CNs' role played in the network, and the different heuristics' capabilities to grasp biologically relevant nodes, we compared the sets of nodes identified as CNs by each heuristic with two experimentally validated sets of essential genes, i.e., the genes whose removal impact on a given organism's ability to survive. Our results show that classical centrality measures (i.e., closeness centrality, degree) found more essential genes with respect to CNH on the current version of the human interactome, however the removal of such nodes does not have the greatest impact on interactome connectivity, while, interestingly, the genes identified by CNH show peculiar characteristics both from the topological and the biological point of view. Finally, even if a relevant fraction of essential genes is found via the classical centrality measures, the same measures seem to fail in identifying the whole set of essential genes, suggesting once again that some of them are not central in the network, that there may be biases in the current interaction data, and that different, combined graph theoretical and other techniques should be applied for their discovery. Alessandro Celestini, Marco Cianfriglia, Enrico Mastrostefano, Alessandro Palma, Filippo Castiglione, Paolo Tieri |
BIBM | 1 |
| 2019 | Analysing the tor web with high performance graph algorithmsabstractThe exploration and analysis of Web graphs has flourished in the recent past, producing a large number of relevant and interesting research results. However, the unique characteristics of the Tor network demand for specific algorithms to explore and analyze it. Tor is an anonymity network that allows offering and accessing various Internet resources while guaranteeing a high degree of provider and user anonymity. So far the attention of the research community has focused on assessing the security of the Tor infrastructure. Most research work on the Tor network aimed at discovering protocol vulnerabilities to de-anonymize users and services, while little or no information is available about the topology of the Tor Web graph or the relationship between pages' content and topological structure. With our work we aim at addressing such lack of information. We describe the topology of the Tor Web graph measuring both global and local properties by means of well-known metrics that require due to the size of the network, high performance algorithms. We consider three different snapshots obtained by extensively crawling Tor three times over a 5 months time frame. Finally we present a correlation analysis of pages' semantics and topology, discussing novel insights about the Tor Web organization and its content. Our findings show that the Tor graph presents some of the characteristics of social and surface web graphs, along with a few unique peculiarities. Massimo Bernaschi, Alessandro Celestini, Stefano Guarino, Flavio Lombardi, Enrico Mastrostefano |
CF | 2 |
| 2019 | Spiders like Onions: on the Network of Tor Hidden ServicesabstractTor hidden services allow offering and accessing various Internet resources while guaranteeing a high degree of provider and user anonymity. So far, most research work on the Tor network aimed at discovering protocol vulnerabilities to de-anonymize users and services. Other work aimed at estimating the number of available hidden services and classifying them. Something that still remains largely unknown is the structure of the graph defined by the network of Tor services. In this paper, we describe the topology of the Tor graph (aggregated at the hidden service level) measuring both global and local properties by means of well-known metrics. We consider three different snapshots obtained by extensively crawling Tor three times over a 5 months time frame. We separately study these three graphs and their shared “stable” core. In doing so, other than assessing the renowned volatility of Tor hidden services, we make it possible to distinguish time dependent and structural aspects of the Tor graph. Our findings show that, among other things, the graph of Tor hidden services presents some of the characteristics of social and surface web graphs, along with a few unique peculiarities, such as a very high percentage of nodes having no outbound links. Massimo Bernaschi, Alessandro Celestini, Stefano Guarino, Flavio Lombardi, Enrico Mastrostefano |
WWW | 2 |
| 2017 | Exploring and Analyzing the Tor Hidden Services GraphabstractThe exploration and analysis of Web graphs has flourished in the recent past, producing a large number of relevant and interesting research results. However, the unique characteristics of the Tor network limit the applicability of standard techniques and demand for specific algorithms to explore and analyze it. The attention of the research community has focused on assessing the security of the Tor infrastructure (i.e., its ability to actually provide the intended level of anonymity) and on discussing what Tor is currently being used for. Since there are no foolproof techniques for automatically discovering Tor hidden services, little or no information is available about the topology of the Tor Web graph. Even less is known on the relationship between content similarity and topological structure. The present article aims at addressing such lack of information. Among its contributions: a study on automatic Tor Web exploration/data collection approaches; the adoption of novel representative metrics for evaluating Tor data; a novel in-depth analysis of the hidden services graph; a rich correlation analysis of hidden services’ semantics and topology. Finally, a broad interesting set of novel insights/considerations over the Tor Web organization and content are provided. Massimo Bernaschi, Alessandro Celestini, Stefano Guarino, Flavio Lombardi |
ACM Trans. Web | 2 |
| 2014 | Reputation-Based Composition of Social Web ServicesabstractSocial Web Services (SWSs) constitute a novel paradigm of service-oriented computing, where Web services, just like humans, sign up in social networks that guarantee, e.g., better service discovery for users and faster replacement in case of service failures. In past work, composition of SWSs was mainly supported by specialised social networks of competitor services and cooperating ones. In this work, we continue this line of research, by proposing a novel SWSs composition procedure driven by the SWSs reputation. Making use of a well-known formal language and associated tools, we specify the composition steps and we prove that such reputation-driven approach assures better results in terms of the overall quality of service of the compositions, with respect to randomly selecting SWSs. Alessandro Celestini, Gianpiero Costantino, Rocco De Nicola, Zakaria Maamar, Fabio Martinelli, Marinella Petrocchi, Francesco Tiezzi 0001 |
AINA | 1 |
| 2013 | Asymptotic Risk Analysis for Trust and Reputation Systems
Michele Boreale, Alessandro Celestini |
SOFSEM | 2 |