EDBT 2026 Demo / reviewers in the wild / expert
Vincenzo Bonnici
dblp:65/8509
· DBLP profile ↗
22ranked-venue papers
14as first author
13since 2021 · last 2025
0000-0002-1637-7545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 11 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Graph Information for Spatially Informed Patient Data Analysis with GISTabstractPatient data such as tissue samples analyzed through spatial transcriptomics have transformed our ability to study cellular subpopulations within their native microenvironments, providing unprecedented insights into tissue architecture and cellular interactions. However, accurately identifying spatial domains remains a computational challenge. Despite notable progress, no gold standard currently exists for spatial domain identification, and significant opportunities remain for further improvement.In this study, we introduce Graph Information for Spatial Transcriptomics (GIST), a Graph Neural Network (GNN)-based framework that integrates gene expression data with spatial coordinates to construct a biologically meaningful graph representation of tissue architecture. By explicitly modeling spatial dependencies and leveraging contrastive learning to optimize node embeddings, GIST substantially improves spatial domain identification. It outperforms existing methods on key clustering metrics such as the Adjusted Rand Index (ARI) demonstrating its effectiveness in capturing the true structure of spatial transcriptomic data.Furthermore, we introduce the Silhouette Spatial Score (SSS)—an extension of the traditional Silhouette Score that incorporates spatial neighborhood information. SSS enables more accurate evaluation of both transcriptomic similarity and spatial contiguity within identified domains. GIST outperforms existing methods in SSS, highlighting its ability to identify domains that are not only transcriptomically meaningful but also spatially contiguous. Gospel Ozioma Nnadi, Vincenzo Bonnici, Simone Avesani, Eva Viesi, Rosalba Giugno |
CIBCB | 2 |
| 2025 | MultiGraphMatch: A Subgraph Matching Algorithm for MultigraphsabstractSubgraph matching is the problem of finding all the occurrences of a small graph, called the query, in a larger graph, called the target. Although the problem has been widely studied in simple graphs, few solutions have been proposed for multigraphs, in which two nodes can be connected by multiple edges, each denoting a possibly different type of relationship. In our new algorithm MultiGraphMatch (MGM), nodes and edges can be associated with labels and multiple properties. MGM introduces a novel data structure called bit matrix to efficiently index both the query and the target and filter the set of target edges that are matchable with each query edge. In addition, the algorithm proposes a new technique for ordering the processing of query edges based on the cardinalities of the sets of matchable edges. Using the CYPHER query definition language, MGM can perform queries with logical conditions on node and edge labels. We compare MGM with SuMGra and graph database systems Memgraph and Neo4J, showing comparable or better performance in all queries on a wide variety of synthetic and real-world graphs. Giovanni Micale, Antonio Di Maria, Roberto Grasso, Vincenzo Bonnici, Alfredo Ferro, Dennis E. Shasha, Rosalba Giugno, Alfredo Pulvirenti |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | ArcMatch: high-performance subgraph matching for labeled graphs by exploiting edge domainsabstractAbstract Consider a large labeled graph (network), denoted the target. Subgraph matching is the problem of finding all instances of a small subgraph, denoted the query, in the target graph. Unlike the majority of existing methods that are restricted to graphs with labels solely on vertices, our proposed approach, named can effectively handle graphs with labels on both vertices and edges. ntroduces an efficient new vertex/edge domain data structure filtering procedure to speed up subgraph queries. The procedure, called path-based reduction, filters initial domains by scanning them for paths up to a specified length that appear in the query graph. Additionally, ncorporates existing techniques like variable ordering and parent selection, as well as adapting the core search process, to take advantage of the information within edge domains. Experiments in real scenarios such as protein–protein interaction graphs, co-authorship networks, and email networks, show that s faster than state-of-the-art systems varying the number of distinct vertex labels over the whole target graph and query sizes. Vincenzo Bonnici, Roberto Grasso, Giovanni Micale, Antonio Di Maria, Dennis E. Shasha, Alfredo Pulvirenti, Rosalba Giugno |
Data Min. Knowl. Discov. | 1 |
| 2023 | MODIMO: Workshop on Multi-Omics Data Integration for Modelling Biological SystemsabstractMulti-omics analysis aims at extracting previously uncovered biological knowledge by integrating information across multiple single-omic sources. Past approaches have focused on the simultaneous analysis of a small number of omic data sets. Current challenges face the problem of integrating multiple omic sources into a unified complex model, or of combining already available tools for two-by-two omics analyses and merging their outcomes. By doing so and leveraging integrated system-level knowledge, multi-omic approaches ought to enable the development of better qualitative and quantitative models for descriptive and predictive analyses. To move this area forward, new statistical and algorithmic frameworks are needed, for example for generalizing classical graph theory results to heterogeneous networks, and applying them to diverse problems such as drug repurposing or understanding the immune response to infections. Thus, in short, this workshop aims at investigating novel methodologies for providing crucial insights into multi-omics data management, integration, and analysis to enable biological discoveries. The workshop will be sponsored by the InfoLife CINI National Laboratory (https://www.consorzio-cini.it/index.php/en/ ). Simone Avesani, Vincenzo Bonnici, Simone Pernice, Marco Beccuti, Rosalba Giugno |
CIKM | 2 |
| 2023 | BIOCHAIN: towards a platform for securely sharing microbiological dataabstractThere is a need to persuade public and private entities to share their currently unexposed bio-data banks by preserving ownership and secrecy. The reason is to make available results that can be obtained by massively exploiting the content of such data by modern machine learning approaches. Digital catalogues of data collections are being provided. However, they are not developed to protect private content that may be shared according to privileges assigned by the owners. Here, we present BIOCHAIN, a data-sharing module which will be the basis for a computational platform aimed at performing federated data analysis. The platform is intended to be used by a consortium of private and public institutions in the field of microbiology. BIOCHAIN makes use of blockchain technology to guarantee fairness among entities of the consortium by allowing them to securely share their data. Vincenzo Bonnici, Vincenzo Arceri, Alessio Diana, Flavio Bertini 0001, Eleonora Iotti, Alessia Levante, Valentina Bernini, Erasmo Neviani, Alessandro Dal Palù |
IDEAS | 1 |
| 2023 | PanDelos-frags: A methodology for discovering pangenomic content of incomplete microbial assembliesabstractPangenomics was originally defined as the problem of comparing the composition of genes into gene families within a set of bacterial isolates belonging to the same species. The problem requires the calculation of sequence homology among such genes. When combined with metagenomics, namely for human microbiome composition analysis, gene-oriented pangenome detection becomes a promising method to decipher ecosystem functions and population-level evolution. Established computational tools are able to investigate the genetic content of isolates for which a complete genomic sequence is available. However, there is a plethora of incomplete genomes that are available on public resources, which only a few tools may analyze. Incomplete means that the process for reconstructing their genomic sequence is not complete, and only fragments of their sequence are currently available. However, the information contained in these fragments may play an essential role in the analyses. Here, we present PanDelos-frags, a computational tool which exploits and extends previous results in analyzing complete genomes. It provides a new methodology for inferring missing genetic information and thus for managing incomplete genomes. PanDelos-frags outperforms state-of-the-art approaches in reconstructing gene families in synthetic benchmarks and in a real use case of metagenomics. PanDelos-frags is publicly available at https://github.com/InfOmics/PanDelos-frags. Vincenzo Bonnici, Claudia Mengoni, Manuel Mangoni, Giuditta Franco, Rosalba Giugno |
J. Biomed. Informatics | 1 |
| 2022 | PANPROVA: pangenomic prokaryotic evolution of full assembliesabstractMOTIVATION: Computational tools for pangenomic analysis have gained increasing interest over the past two decades in various applications such as evolutionary studies and vaccine development. Synthetic benchmarks are essential for the systematic evaluation of their performance. Currently, benchmarking tools represent a genome as a set of genetic sequences and fail to simulate the complete information of the genomes, which is essential for evaluating pangenomic detection between fragmented genomes. RESULTS: We present PANPROVA, a benchmark tool to simulate prokaryotic pangenomic evolution by evolving the complete genomic sequence of an ancestral isolate. In this way, the possibility of operating in the preassembly phase is enabled. Gene set variations, sequence variation and horizontal acquisition from a pool of external genomes are the evolutionary features of the tool. AVAILABILITY AND IMPLEMENTATION: PANPROVA is publicly available at https://github.com/InfOmics/PANPROVA. The manuscript explicitelly refers to the github repository. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vincenzo Bonnici, Rosalba Giugno |
Bioinform. | 1 |
| 2021 | MODIMO: Workshop on Multi-Omics Data Integration for Modelling Biological SystemsabstractMulti-omics analysis aims at extracting previously uncovered biological knowledge by integrating information across multiple single-omic sources. Past approaches have focused on the simultaneous analysis of a small number of omic data sets. Current challenges face the problem of integrating multiple omic sources into a unified complex model, or of combining already available tools for two-by-two omics analyses and merging their outcomes. By doing so and leveraging integrated system-level knowledge, multi-omic approaches ought to enable the development of better qualitative and quantitative models for descriptive and predictive analyses. To move this area forward, new statistical and algorithmic frameworks are needed, for example for generalizing classical graph theory results to heterogeneous networks and applying them to diverse problems such as drug repurposing or understanding the immune response to infections. Thus, in short, this workshop aims at investigating novel methodologies for providing crucial insights into multi-omics data management, integration, and analysis in order to enable biological discoveries. Marco Beccuti, Vincenzo Bonnici, Rosalba Giugno |
CIKM | 2 |
| 2021 | TEDAR: Temporal dynamic signal detection of adverse reactions
Antonino Aparo, Pietro Sala, Vincenzo Bonnici, Rosalba Giugno |
Artif. Intell. Medicine | 3 |
| 2021 | Challenges in gene-oriented approaches for pangenome content discoveryabstractGiven a group of genomes, represented as the sets of genes that belong to them, the discovery of the pangenomic content is based on the search of genetic homology among the genes for clustering them into families. Thus, pangenomic analyses investigate the membership of the families to the given genomes. This approach is referred to as the gene-oriented approach in contrast to other definitions of the problem that takes into account different genomic features. In the past years, several tools have been developed to discover and analyse pangenomic contents. Because of the hardness of the problem, each tool applies a different strategy for discovering the pangenomic content. This results in a differentiation of the performance of each tool that depends on the composition of the input genomes. This review reports the main analysis instruments provided by the current state of the art tools for the discovery of pangenomic contents. Moreover, unlike previous works, the presented study compares pangenomic tools from a methodological perspective, analysing the causes that lead a given methodology to outperform other tools. The analysis is performed by taking into account different bacterial populations, which are synthetically generated by changing evolutionary parameters. The benchmarks used to compare the pangenomic tools, in addition to the computational pipeline developed for this purpose, are available at https://github.com/InfOmics/pangenes-review. Contact: V. Bonnici, R. Giugno Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. Vincenzo Bonnici, Emiliano Maresi, Rosalba Giugno |
Briefings Bioinform. | 1 |
| 2021 | GRAPES-DD: exploiting decision diagrams for index-driven search in biological graph databasesabstractBACKGROUND: Graphs are mathematical structures widely used for expressing relationships among elements when representing biomedical and biological information. On top of these representations, several analyses are performed. A common task is the search of one substructure within one graph, called target. The problem is referred to as one-to-one subgraph search, and it is known to be NP-complete. Heuristics and indexing techniques can be applied to facilitate the search. Indexing techniques are also exploited in the context of searching in a collection of target graphs, referred to as one-to-many subgraph problem. Filter-and-verification methods that use indexing approaches provide a fast pruning of target graphs or parts of them that do not contain the query. The expensive verification phase is then performed only on the subset of promising targets. Indexing strategies extract graph features at a sufficient granularity level for performing a powerful filtering step. Features are memorized in data structures allowing an efficient access. Indexing size, querying time and filtering power are key points for the development of efficient subgraph searching solutions. RESULTS: An existing approach, GRAPES, has been shown to have good performance in terms of speed-up for both one-to-one and one-to-many cases. However, it suffers in the size of the built index. For this reason, we propose GRAPES-DD, a modified version of GRAPES in which the indexing structure has been replaced with a Decision Diagram. Decision Diagrams are a broad class of data structures widely used to encode and manipulate functions efficiently. Experiments on biomedical structures and synthetic graphs have confirmed our expectation showing that GRAPES-DD has substantially reduced the memory utilization compared to GRAPES without worsening the searching time. CONCLUSION: The use of Decision Diagrams for searching in biochemical and biological graphs is completely new and potentially promising thanks to their ability to encode compactly sets by exploiting their structure and regularity, and to manipulate entire sets of elements at once, instead of exploring each single element explicitly. Search strategies based on Decision Diagram makes the indexing for biochemical graphs, and not only, more affordable allowing us to potentially deal with huge and ever growing collections of biochemical and biological structures. Nicola Licheri, Vincenzo Bonnici, Marco Beccuti, Rosalba Giugno |
BMC Bioinform. | 2 |
| 2021 | GRAFIMO: Variant and haplotype aware motif scanning on pangenome graphsabstractTranscription factors (TFs) are proteins that promote or reduce the expression of genes by binding short genomic DNA sequences known as transcription factor binding sites (TFBS). While several tools have been developed to scan for potential occurrences of TFBS in linear DNA sequences or reference genomes, no tool exists to find them in pangenome variation graphs (VGs). VGs are sequence-labelled graphs that can efficiently encode collections of genomes and their variants in a single, compact data structure. Because VGs can losslessly compress large pangenomes, TFBS scanning in VGs can efficiently capture how genomic variation affects the potential binding landscape of TFs in a population of individuals. Here we present GRAFIMO (GRAph-based Finding of Individual Motif Occurrences), a command-line tool for the scanning of known TF DNA motifs represented as Position Weight Matrices (PWMs) in VGs. GRAFIMO extends the standard PWM scanning procedure by considering variations and alternative haplotypes encoded in a VG. Using GRAFIMO on a VG based on individuals from the 1000 Genomes project we recover several potential binding sites that are enhanced, weakened or missed when scanning only the reference genome, and which could constitute individual-specific binding events. GRAFIMO is available as an open-source tool, under the MIT license, at https://github.com/pinellolab/GRAFIMO and https://github.com/InfOmics/GRAFIMO. Manuel Tognon, Vincenzo Bonnici, Erik Garrison, Rosalba Giugno, Luca Pinello |
PLoS Comput. Biol. | 2 |
| 2021 | Spectral concepts in genome informational analysis
Vincenzo Bonnici, Giuditta Franco, Vincenzo Manca |
Theor. Comput. Sci. | 1 |
| 2019 | LErNet: characterization of lncRNAs via context-aware network expansion and enrichment analysisabstractLong non-coding RNAs (lncRNAs) have recently acquired a boost of interest for their implication in several biological conditions. However, many of these elements are not yet characterized. LErNet is a method to in silico define and predict the roles of IncRNAs. The core of the approach is a network expansion algorithm which enriches the genomic context of IncRNAs. The context is built by integrating the genes encoding proteins that are found next to the non-coding elements both at genomic and system level. The pipeline is particularly useful in situations where the functions of discovered IncRNAs are not yet known. The results show both the outperformance of LErNet compared to enrichment approaches in literature and its robustness in case of partially missing context information. LErNet is provided as an R package. It is available at https://github.com/InfOmics/LErNet. Vincenzo Bonnici, Simone Caligola, Giulia Fiorini, Luca Giudice, Rosalba Giugno |
CIBCB | 1 |
| 2019 | Parallel Searching on Biological NetworksabstractSoftware applications for biological networks analysis rely on graphs to model the structure interactions. A great part of them requires searching for subgraphs in a target graph or in collections of graphs. Even though very efficient algorithms have been defined to solve such a subgraph isomorphisms problem, the complexity of current real biological networks make their sequential execution time prohibitive. On the other hand, parallel architectures, from multi-core to manycore, have become pervasive to deal with the problem of the data size. Nevertheless, the sequential nature of the graph searching algorithms makes their implementation for parallel architectures very challenging. This paper presents three different parallel solutions for the graph searching problem. The first two target the exact search for multi-core CPUs and manycore GPUs, respectively. The third one targets the approximate search for GPUs, which handles node, edge, and node label mismatches. The paper shows how different techniques have been developed in all the solutions to reduce the search space complexity. The paper shows the performance of the proposed solutions on representative biological networks containing antiviral chemical compounds and protein interactions networks. Nicola Bombieri, Vincenzo Bonnici, Rosalba Giugno |
PDP | 2 |
| 2018 | An Efficient Implementation of a Subgraph Isomorphism Algorithm for GPUs
Vincenzo Bonnici, Rosalba Giugno, Nicola Bombieri |
BIBM | 1 |
| 2018 | cuRnet: an R package for graph traversing on GPUabstractBACKGROUND: R has become the de-facto reference analysis environment in Bioinformatics. Plenty of tools are available as packages that extend the R functionality, and many of them target the analysis of biological networks. Several algorithms for graphs, which are the most adopted mathematical representation of networks, are well-known examples of applications that require high-performance computing, and for which classic sequential implementations are becoming inappropriate. In this context, parallel approaches targeting GPU architectures are becoming pervasive to deal with the execution time constraints. Although R packages for parallel execution on GPUs are already available, none of them provides graph algorithms. RESULTS: This work presents cuRnet, a R package that provides a parallel implementation for GPUs of the breath-first search (BFS), the single-source shortest paths (SSSP), and the strongly connected components (SCC) algorithms. The package allows offloading computing intensive applications to GPU devices for massively parallel computation and to speed up the runtime up to one order of magnitude with respect to the standard sequential computations on CPU. We have tested cuRnet on a benchmark of large protein interaction networks and for the interpretation of high-throughput omics data thought network analysis. CONCLUSIONS: cuRnet is a R package to speed up graph traversal and analysis through parallel computation on GPUs. We show the efficiency of cuRnet applied both to biological network analysis, which requires basic graph algorithms, and to complex existing procedures built upon such algorithms. Vincenzo Bonnici, Federico Busato, Stefano Aldegheri, Murodzhon Akhmedov, Luciano Cascione, Alberto Arribas Carmena, Francesco Bertoni, Nicola Bombieri, Ivo Kwee, Rosalba Giugno |
BMC Bioinform. | 1 |
| 2018 | Correction to: cuRnet: an R package for graph traversing on GPUabstractAfter publication of this supplement article [1], it was brought to our attention that reference 10 and reference 12 in the article are incorrect. Vincenzo Bonnici, Federico Busato, Stefano Aldegheri, Murodzhon Akhmedov, Luciano Cascione, Alberto Arribas Carmena, Francesco Bertoni, Nicola Bombieri, Ivo Kwee, Rosalba Giugno |
BMC Bioinform. | 1 |
| 2018 | Arena-Idb: a platform to build human non-coding RNA interaction networksabstractBACKGROUND: High throughput technologies have provided the scientific community an unprecedented opportunity for large-scale analysis of genomes. Non-coding RNAs (ncRNAs), for a long time believed to be non-functional, are emerging as one of the most important and large family of gene regulators and key elements for genome maintenance. Functional studies have been able to assign to ncRNAs a wide spectrum of functions in primary biological processes, and for this reason they are assuming a growing importance as a potential new family of cancer therapeutic targets. Nevertheless, the number of functionally characterized ncRNAs is still too poor if compared to the number of new discovered ncRNAs. Thus platforms able to merge information from available resources addressing data integration issues are necessary and still insufficient to elucidate ncRNAs biological roles. RESULTS: In this paper, we describe a platform called Arena-Idb for the retrieval of comprehensive and non-redundant annotated ncRNAs interactions. Arena-Idb provides a framework for network reconstruction of ncRNA heterogeneous interactions (i.e., with other type of molecules) and relationships with human diseases which guide the integration of data, extracted from different sources, via mapping of entities and minimization of ambiguity. CONCLUSIONS: Arena-Idb provides a schema and a visualization system to integrate ncRNA interactions that assists in discovering ncRNA functions through the extraction of heterogeneous interaction networks. The Arena-Idb is available at http://arenaidb.ba.itb.cnr.it. Vincenzo Bonnici, Giorgio De Caro, Giorgio Constantino, Sabino Liuni, Domenica D'Elia, Nicola Bombieri, Flavio Licciulli, Rosalba Giugno |
BMC Bioinform. | 1 |
| 2017 | On the Variable Ordering in Subgraph Isomorphism AlgorithmsabstractGraphs are mathematical structures to model several biological data. Applications to analyze them require to apply solutions for the subgraph isomorphism problem, which is NP-complete. Here, we investigate the existing strategies to reduce the subgraph isomorphism algorithm running time with emphasis on the importance of the order with which the graph vertices are taken into account during the search, called variable ordering, and its incidence on the total running time of the algorithms. We focus on two recent solutions, which are based on an effective variable ordering strategy. We discuss their comparison both with the variable ordering strategies reviewed in the paper and the other algorithms present in the ICPR2014 contest on graph matching algorithms for pattern search in biological databases. Vincenzo Bonnici, Rosalba Giugno |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2016 | APPAGATO: an APproximate PArallel and stochastic GrAph querying TOol for biological networksabstractMOTIVATION: Biological network querying is a problem requiring a considerable computational effort to be solved. Given a target and a query network, it aims to find occurrences of the query in the target by considering topological and node similarities (i.e. mismatches between nodes, edges, or node labels). Querying tools that deal with similarities are crucial in biological network analysis because they provide meaningful results also in case of noisy data. In addition, as the size of available networks increases steadily, existing algorithms and tools are becoming unsuitable. This is rising new challenges for the design of more efficient and accurate solutions. RESULTS: This paper presents APPAGATO, a stochastic and parallel algorithm to find approximate occurrences of a query network in biological networks. APPAGATO handles node, edge and node label mismatches. Thanks to its randomic and parallel nature, it applies to large networks and, compared with existing tools, it provides higher performance as well as statistically significant more accurate results. Tests have been performed on protein-protein interaction networks annotated with synthetic and real gene ontology terms. Case studies have been done by querying protein complexes among different species and tissues. AVAILABILITY AND IMPLEMENTATION: APPAGATO has been developed on top of CUDA-C ++ Toolkit 7.0 framework. The software is available online http://profs.sci.univr.it/∼bombieri/APPAGATO CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vincenzo Bonnici, Federico Busato, Giovanni Micale, Nicola Bombieri, Alfredo Pulvirenti, Rosalba Giugno |
Bioinform. | 1 |
| 2013 | A subgraph isomorphism algorithm and its application to biochemical dataabstractBACKGROUND: Graphs can represent biological networks at the molecular, protein, or species level. An important query is to find all matches of a pattern graph to a target graph. Accomplishing this is inherently difficult (NP-complete) and the efficiency of heuristic algorithms for the problem may depend upon the input graphs. The common aim of existing algorithms is to eliminate unsuccessful mappings as early as and as inexpensively as possible. RESULTS: We propose a new subgraph isomorphism algorithm which applies a search strategy to significantly reduce the search space without using any complex pruning rules or domain reduction procedures. We compare our method with the most recent and efficient subgraph isomorphism algorithms (VFlib, LAD, and our C++ implementation of FocusSearch which was originally distributed in Modula2) on synthetic, molecules, and interaction networks data. We show a significant reduction in the running time of our approach compared with these other excellent methods and show that our algorithm scales well as memory demands increase. CONCLUSIONS: Subgraph isomorphism algorithms are intensively used by biochemical tools. Our analysis gives a comprehensive comparison of different software approaches to subgraph isomorphism highlighting their weaknesses and strengths. This will help researchers make a rational choice among methods depending on their application. We also distribute an open-source package including our system and our own C++ implementation of FocusSearch together with all the used datasets (http://ferrolab.dmi.unict.it/ri.html). In future work, our findings may be extended to approximate subgraph isomorphism algorithms. Vincenzo Bonnici, Rosalba Giugno, Alfredo Pulvirenti, Dennis E. Shasha, Alfredo Ferro |
BMC Bioinform. | 1 |