Rosalba Giugno

dblp:02/5792 · DBLP profile ↗
← Back
61ranked-venue papers
2as first author
21since 2021 · last 2025
0000-0001-9843-7638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 11 · 4 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Leveraging Graph Information for Spatially Informed Patient Data Analysis with GIST
abstract
Patient data such as tissue samples analyzed through spatial transcriptomics have transformed our ability to study cellular subpopulations within their native microenvironments, providing unprecedented insights into tissue architecture and cellular interactions. However, accurately identifying spatial domains remains a computational challenge. Despite notable progress, no gold standard currently exists for spatial domain identification, and significant opportunities remain for further improvement.In this study, we introduce Graph Information for Spatial Transcriptomics (GIST), a Graph Neural Network (GNN)-based framework that integrates gene expression data with spatial coordinates to construct a biologically meaningful graph representation of tissue architecture. By explicitly modeling spatial dependencies and leveraging contrastive learning to optimize node embeddings, GIST substantially improves spatial domain identification. It outperforms existing methods on key clustering metrics such as the Adjusted Rand Index (ARI) demonstrating its effectiveness in capturing the true structure of spatial transcriptomic data.Furthermore, we introduce the Silhouette Spatial Score (SSS)—an extension of the traditional Silhouette Score that incorporates spatial neighborhood information. SSS enables more accurate evaluation of both transcriptomic similarity and spatial contiguity within identified domains. GIST outperforms existing methods in SSS, highlighting its ability to identify domains that are not only transcriptomically meaningful but also spatially contiguous.
Gospel Ozioma Nnadi, Vincenzo Bonnici, Simone Avesani, Eva Viesi, Rosalba Giugno
CIBCB5
2025 Benchmarking transcription factor binding site prediction models: a comparative analysis on synthetic and biological data
abstract
Transcription factors (TFs) are essential regulatory proteins controlling the cellular transcriptional states by binding to specific DNA sequences known as transcription factor binding sites (TFBSs) or motifs. Accurate TFBS identification is crucial for unraveling regulatory mechanisms driving cellular dynamics. Over the years, various computational approaches have been developed to model TFBSs, with position weight matrices (PWMs) being one of the most widely adopted methods. PWMs provide a probabilistic framework by representing nucleotide frequencies at every position within the binding site. While effective and interpretable, PWMs face significant limitations, such as their inability to capture positional dependencies or model complex interactions. To address these, advanced methods, like support vector machine (SVM)-based, and deep learning (DL)-based models, have been introduced. Leveraging human ChIP-seq data from ENCODE, we systematically benchmarked the predictive performance of PWM, SVM-, and DL-based models across different scenarios. We evaluate the impact of key factors such as training dataset size, sequence length, and kernel functions (for SVMs) on models' performance. Additionally, we explore the impact of synthetic versus real biological background data during model training. Our analysis highlights strengths and limitations of each approach under different conditions, providing practical guidance for selecting and tailoring models to specific biological datasets. To complement our analysis, we present a comprehensive database of pretrained SVM models for TFBS detection, trained on human ChIP-seq data from diverse cell lines and tissues. This resource aims to facilitate broader adoption of SVM-based methods in TFBS prediction and enhance their practical utility in regulatory genomics research.
Manuel Tognon, Alisa Kumbara, Andrea Betti, Lorenzo Ruggeri, Rosalba Giugno
Briefings Bioinform.5
2025 An efficient solution for GPUs to the ST-connectivity problem on dynamic graphs
abstract
ST-connectivity poses a decision problem, determining whether, for vertices s and t within a graph, t is reachable from s . The challenge arises in the context of dynamic real-world graphs that undergo rapid evolution over time. In these scenarios, repeatedly solving the s-t connectivity problem from the beginning after each graph modification becomes impractical. Although parallel solutions, especially designed for GPUs, have been introduced to tackle the size complexity of static graphs, none have specifically addressed the concern of work efficiency in dynamic graphs. We propose an efficient solution for GPUs to the st-connectivity problem that can handle concurrent processing of batches of graph updates. We use batch information strategically to reduce the overall workload needed for updating the connectivity result. We provide experimental results based on standard datasets and with graphs of different characteristics and batch sizes to evaluate the proposed solutions efficiency. • The article presents an STCON solution for GPU architectures. • The solution targets dynamic graphs. • It shows the speedup and improvements wrt the static solution. • The results have been conducted on two large and standard datasets.
Leonardo Fraccaroli, Federico Busato, Rosalba Giugno, Nicola Bombieri
Pattern Recognit. Lett.3
2025 MultiGraphMatch: A Subgraph Matching Algorithm for Multigraphs
abstract
Subgraph matching is the problem of finding all the occurrences of a small graph, called the query, in a larger graph, called the target. Although the problem has been widely studied in simple graphs, few solutions have been proposed for multigraphs, in which two nodes can be connected by multiple edges, each denoting a possibly different type of relationship. In our new algorithm MultiGraphMatch (MGM), nodes and edges can be associated with labels and multiple properties. MGM introduces a novel data structure called bit matrix to efficiently index both the query and the target and filter the set of target edges that are matchable with each query edge. In addition, the algorithm proposes a new technique for ordering the processing of query edges based on the cardinalities of the sets of matchable edges. Using the CYPHER query definition language, MGM can perform queries with logical conditions on node and edge labels. We compare MGM with SuMGra and graph database systems Memgraph and Neo4J, showing comparable or better performance in all queries on a wide variety of synthetic and real-world graphs.
Giovanni Micale, Antonio Di Maria, Roberto Grasso, Vincenzo Bonnici, Alfredo Ferro, Dennis E. Shasha, Rosalba Giugno, Alfredo Pulvirenti
ACM Trans. Knowl. Discov. Data7
2024 GPU-Accelerated BFS for Dynamic Networks
Filippo Ziche, Nicola Bombieri, Federico Busato, Rosalba Giugno
Euro-Par (3)4
2024 GateMeClass: Gate Mining and Classification of cytometry data
abstract
MOTIVATION: Cytometry comprises powerful techniques for analyzing the cell heterogeneity of a biological sample by examining the expression of protein markers. These technologies impact especially the field of oncoimmunology, where cell identification is essential to analyze the tumor microenvironment. Several classification tools have been developed for the annotation of cytometry datasets, which include supervised tools that require a training set as a reference (i.e. reference-based) and semisupervised tools based on the manual definition of a marker table. The latter is closer to the traditional annotation of cytometry data based on manual gating. However, they require the manual definition of a marker table that cannot be extracted automatically in a reference-based fashion. Therefore, we are lacking methods that allow both classification approaches while maintaining the high biological interpretability given by the marker table. RESULTS: We present a new tool called GateMeClass (Gate Mining and Classification) which overcomes the limitation of the current methods of classification of cytometry data allowing both semisupervised and supervised annotation based on a marker table that can be defined manually or extracted from an external annotated dataset. We measured the accuracy of GateMeClass for annotating three well-established benchmark mass cytometry datasets and one flow cytometry dataset. The performance of GateMeClass is comparable to reference-based methods and marker table-based techniques, offering greater flexibility and rapid execution times. AVAILABILITY AND IMPLEMENTATION: GateMeClass is implemented in R language and is publicly available at https://github.com/simo1c/GateMeClass.
Simone Caligola, Luca Giacobazzi, Stefania Canè, Antonio Vella, Annalisa Adamo, Stefano Ugel, Rosalba Giugno, Vincenzo Bronte
Bioinform.7
2024 ArcMatch: high-performance subgraph matching for labeled graphs by exploiting edge domains
abstract
Abstract Consider a large labeled graph (network), denoted the target. Subgraph matching is the problem of finding all instances of a small subgraph, denoted the query, in the target graph. Unlike the majority of existing methods that are restricted to graphs with labels solely on vertices, our proposed approach, named can effectively handle graphs with labels on both vertices and edges. ntroduces an efficient new vertex/edge domain data structure filtering procedure to speed up subgraph queries. The procedure, called path-based reduction, filters initial domains by scanning them for paths up to a specified length that appear in the query graph. Additionally, ncorporates existing techniques like variable ordering and parent selection, as well as adapting the core search process, to take advantage of the information within edge domains. Experiments in real scenarios such as protein–protein interaction graphs, co-authorship networks, and email networks, show that s faster than state-of-the-art systems varying the number of distinct vertex labels over the whole target graph and query sizes.
Vincenzo Bonnici, Roberto Grasso, Giovanni Micale, Antonio Di Maria, Dennis E. Shasha, Alfredo Pulvirenti, Rosalba Giugno
Data Min. Knowl. Discov.7
2024 From translational bioinformatics computational methodologies to personalized medicine
Barbara Di Camillo, Rosalba Giugno
J. Biomed. Informatics2
2024 Identifying the joint signature of brain atrophy and gene variant scores in Alzheimer's Disease
abstract
The joint modeling of genetic data and brain imaging information allows for determining the pathophysiological pathways of neurodegenerative diseases such as Alzheimer's disease (AD). This task has typically been approached using mass-univariate methods that rely on a complete set of Single Nucleotide Polymorphisms (SNPs) to assess their association with selected image-derived phenotypes (IDPs). However, such methods are prone to multiple comparisons bias and, most importantly, fail to account for potential cross-feature interactions, resulting in insufficient detection of significant associations. Ways to overcome these limitations while reducing the number of traits aim at conveying genetic information at the gene level and capturing the integrated genetic effects of a set of genetic variants, rather than looking at each SNP individually. Their associations with brain IDPs are still largely unexplored in the current literature, though they can uncover new potential genetic determinants for brain modulations in the AD continuum. In this work, we explored an explainable multivariate model to analyze the genetic basis of the grey matter modulations, relying on the AD Neuroimaging Initiative (ADNI) phase 3 dataset. Cortical thicknesses and subcortical volumes derived from T1-weighted Magnetic Resonance were considered to describe the imaging phenotypes. At the same time the genetic counterpart was represented by gene variant scores extracted by the Sequence Kernel Association Test (SKAT) filtering model. Moreover, transcriptomic analysis was carried on to assess the expression of the resulting genes in the main brain structures as a form of validation. Results highlighted meaningful genotype-phenotype interactionsas defined by three latent components showing a significant difference in the projection scores between patients and controls. Among the significant associations, the model highlighted EPHX1 and BCAS1 gene variant scores involved in neurodegenerative and myelination processes, hence relevant for AD. In particular, the first was associated with decreased subcortical volumes and the second with decreasedtemporal lobe thickness. Noteworthy, BCAS1 is particularly expressed in the dentate gyrus. Overall, the proposed approach allowed capturing genotype-phenotype interactions in a restricted study cohort that was confirmed by transcriptomic analysis, offering insights into the underlying mechanisms of neurodegeneration in AD in line with previous findings and suggesting new potential disease biomarkers.
Federica Cruciani, Antonino Aparo, Lorenza Brusini, Carlo Combi, Silvia Francesca Storti, Rosalba Giugno, Gloria Menegaz, Ilaria Boscolo Galazzo
J. Biomed. Informatics6
2023 MODIMO: Workshop on Multi-Omics Data Integration for Modelling Biological Systems
abstract
Multi-omics analysis aims at extracting previously uncovered biological knowledge by integrating information across multiple single-omic sources. Past approaches have focused on the simultaneous analysis of a small number of omic data sets. Current challenges face the problem of integrating multiple omic sources into a unified complex model, or of combining already available tools for two-by-two omics analyses and merging their outcomes. By doing so and leveraging integrated system-level knowledge, multi-omic approaches ought to enable the development of better qualitative and quantitative models for descriptive and predictive analyses. To move this area forward, new statistical and algorithmic frameworks are needed, for example for generalizing classical graph theory results to heterogeneous networks, and applying them to diverse problems such as drug repurposing or understanding the immune response to infections. Thus, in short, this workshop aims at investigating novel methodologies for providing crucial insights into multi-omics data management, integration, and analysis to enable biological discoveries. The workshop will be sponsored by the InfoLife CINI National Laboratory (https://www.consorzio-cini.it/index.php/en/ ).
Simone Avesani, Vincenzo Bonnici, Simone Pernice, Marco Beccuti, Rosalba Giugno
CIKM5
2023 A survey on algorithms to characterize transcription factor binding sites
abstract
Transcription factors (TFs) are key regulatory proteins that control the transcriptional rate of cells by binding short DNA sequences called transcription factor binding sites (TFBS) or motifs. Identifying and characterizing TFBS is fundamental to understanding the regulatory mechanisms governing the transcriptional state of cells. During the last decades, several experimental methods have been developed to recover DNA sequences containing TFBS. In parallel, computational methods have been proposed to discover and identify TFBS motifs based on these DNA sequences. This is one of the most widely investigated problems in bioinformatics and is referred to as the motif discovery problem. In this manuscript, we review classical and novel experimental and computational methods developed to discover and characterize TFBS motifs in DNA sequences, highlighting their advantages and drawbacks. We also discuss open challenges and future perspectives that could fill the remaining gaps in the field.
Manuel Tognon, Rosalba Giugno, Luca Pinello
Briefings Bioinform.2
2023 PanDelos-frags: A methodology for discovering pangenomic content of incomplete microbial assemblies
abstract
Pangenomics was originally defined as the problem of comparing the composition of genes into gene families within a set of bacterial isolates belonging to the same species. The problem requires the calculation of sequence homology among such genes. When combined with metagenomics, namely for human microbiome composition analysis, gene-oriented pangenome detection becomes a promising method to decipher ecosystem functions and population-level evolution. Established computational tools are able to investigate the genetic content of isolates for which a complete genomic sequence is available. However, there is a plethora of incomplete genomes that are available on public resources, which only a few tools may analyze. Incomplete means that the process for reconstructing their genomic sequence is not complete, and only fragments of their sequence are currently available. However, the information contained in these fragments may play an essential role in the analyses. Here, we present PanDelos-frags, a computational tool which exploits and extends previous results in analyzing complete genomes. It provides a new methodology for inferring missing genetic information and thus for managing incomplete genomes. PanDelos-frags outperforms state-of-the-art approaches in reconstructing gene families in synthetic benchmarks and in a real use case of metagenomics. PanDelos-frags is publicly available at https://github.com/InfOmics/PanDelos-frags.
Vincenzo Bonnici, Claudia Mengoni, Manuel Mangoni, Giuditta Franco, Rosalba Giugno
J. Biomed. Informatics5
2022 PANPROVA: pangenomic prokaryotic evolution of full assemblies
abstract
MOTIVATION: Computational tools for pangenomic analysis have gained increasing interest over the past two decades in various applications such as evolutionary studies and vaccine development. Synthetic benchmarks are essential for the systematic evaluation of their performance. Currently, benchmarking tools represent a genome as a set of genetic sequences and fail to simulate the complete information of the genomes, which is essential for evaluating pangenomic detection between fragmented genomes. RESULTS: We present PANPROVA, a benchmark tool to simulate prokaryotic pangenomic evolution by evolving the complete genomic sequence of an ancestral isolate. In this way, the possibility of operating in the preassembly phase is enabled. Gene set variations, sequence variation and horizontal acquisition from a pool of external genomes are the evolutionary features of the tool. AVAILABILITY AND IMPLEMENTATION: PANPROVA is publicly available at https://github.com/InfOmics/PANPROVA. The manuscript explicitelly refers to the github repository. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vincenzo Bonnici, Rosalba Giugno
Bioinform.2
2022 From translational bioinformatics computational methodologies to personalized medicine
Barbara Di Camillo, Rosalba Giugno
J. Biomed. Informatics2
2022 Ten quick tips for biomarker discovery and validation analyses using machine learning
abstract
IntroductionAU : Pleaseconfirmthatallheadinglevelsarerepresentedcorrectly
Ramón Díaz-Uriarte, Elisa Gómez de Lope, Rosalba Giugno, Holger Fröhlich, Petr V. Nazarov, Isabel A. Nepomuceno-Chamorro, Armin Rauschenberger, Enrico Glaab
PLoS Comput. Biol.3
2021 MODIMO: Workshop on Multi-Omics Data Integration for Modelling Biological Systems
abstract
Multi-omics analysis aims at extracting previously uncovered biological knowledge by integrating information across multiple single-omic sources. Past approaches have focused on the simultaneous analysis of a small number of omic data sets. Current challenges face the problem of integrating multiple omic sources into a unified complex model, or of combining already available tools for two-by-two omics analyses and merging their outcomes. By doing so and leveraging integrated system-level knowledge, multi-omic approaches ought to enable the development of better qualitative and quantitative models for descriptive and predictive analyses. To move this area forward, new statistical and algorithmic frameworks are needed, for example for generalizing classical graph theory results to heterogeneous networks and applying them to diverse problems such as drug repurposing or understanding the immune response to infections. Thus, in short, this workshop aims at investigating novel methodologies for providing crucial insights into multi-omics data management, integration, and analysis in order to enable biological discoveries.
Marco Beccuti, Vincenzo Bonnici, Rosalba Giugno
CIKM3
2021 TEDAR: Temporal dynamic signal detection of adverse reactions
Antonino Aparo, Pietro Sala, Vincenzo Bonnici, Rosalba Giugno
Artif. Intell. Medicine4
2021 Challenges in gene-oriented approaches for pangenome content discovery
abstract
Given a group of genomes, represented as the sets of genes that belong to them, the discovery of the pangenomic content is based on the search of genetic homology among the genes for clustering them into families. Thus, pangenomic analyses investigate the membership of the families to the given genomes. This approach is referred to as the gene-oriented approach in contrast to other definitions of the problem that takes into account different genomic features. In the past years, several tools have been developed to discover and analyse pangenomic contents. Because of the hardness of the problem, each tool applies a different strategy for discovering the pangenomic content. This results in a differentiation of the performance of each tool that depends on the composition of the input genomes. This review reports the main analysis instruments provided by the current state of the art tools for the discovery of pangenomic contents. Moreover, unlike previous works, the presented study compares pangenomic tools from a methodological perspective, analysing the causes that lead a given methodology to outperform other tools. The analysis is performed by taking into account different bacterial populations, which are synthetically generated by changing evolutionary parameters. The benchmarks used to compare the pangenomic tools, in addition to the computational pipeline developed for this purpose, are available at https://github.com/InfOmics/pangenes-review. Contact: V. Bonnici, R. Giugno Supplementary information: Supplementary data are available at Briefings in Bioinformatics online.
Vincenzo Bonnici, Emiliano Maresi, Rosalba Giugno
Briefings Bioinform.3
2021 GRAPES-DD: exploiting decision diagrams for index-driven search in biological graph databases
abstract
BACKGROUND: Graphs are mathematical structures widely used for expressing relationships among elements when representing biomedical and biological information. On top of these representations, several analyses are performed. A common task is the search of one substructure within one graph, called target. The problem is referred to as one-to-one subgraph search, and it is known to be NP-complete. Heuristics and indexing techniques can be applied to facilitate the search. Indexing techniques are also exploited in the context of searching in a collection of target graphs, referred to as one-to-many subgraph problem. Filter-and-verification methods that use indexing approaches provide a fast pruning of target graphs or parts of them that do not contain the query. The expensive verification phase is then performed only on the subset of promising targets. Indexing strategies extract graph features at a sufficient granularity level for performing a powerful filtering step. Features are memorized in data structures allowing an efficient access. Indexing size, querying time and filtering power are key points for the development of efficient subgraph searching solutions. RESULTS: An existing approach, GRAPES, has been shown to have good performance in terms of speed-up for both one-to-one and one-to-many cases. However, it suffers in the size of the built index. For this reason, we propose GRAPES-DD, a modified version of GRAPES in which the indexing structure has been replaced with a Decision Diagram. Decision Diagrams are a broad class of data structures widely used to encode and manipulate functions efficiently. Experiments on biomedical structures and synthetic graphs have confirmed our expectation showing that GRAPES-DD has substantially reduced the memory utilization compared to GRAPES without worsening the searching time. CONCLUSION: The use of Decision Diagrams for searching in biochemical and biological graphs is completely new and potentially promising thanks to their ability to encode compactly sets by exploiting their structure and regularity, and to manipulate entire sets of elements at once, instead of exploring each single element explicitly. Search strategies based on Decision Diagram makes the indexing for biochemical graphs, and not only, more affordable allowing us to potentially deal with huge and ever growing collections of biochemical and biological structures.
Nicola Licheri, Vincenzo Bonnici, Marco Beccuti, Rosalba Giugno
BMC Bioinform.4
2021 GRAFIMO: Variant and haplotype aware motif scanning on pangenome graphs
abstract
Transcription factors (TFs) are proteins that promote or reduce the expression of genes by binding short genomic DNA sequences known as transcription factor binding sites (TFBS). While several tools have been developed to scan for potential occurrences of TFBS in linear DNA sequences or reference genomes, no tool exists to find them in pangenome variation graphs (VGs). VGs are sequence-labelled graphs that can efficiently encode collections of genomes and their variants in a single, compact data structure. Because VGs can losslessly compress large pangenomes, TFBS scanning in VGs can efficiently capture how genomic variation affects the potential binding landscape of TFs in a population of individuals. Here we present GRAFIMO (GRAph-based Finding of Individual Motif Occurrences), a command-line tool for the scanning of known TF DNA motifs represented as Position Weight Matrices (PWMs) in VGs. GRAFIMO extends the standard PWM scanning procedure by considering variations and alternative haplotypes encoded in a VG. Using GRAFIMO on a VG based on individuals from the 1000 Genomes project we recover several potential binding sites that are enhanced, weakened or missed when scanning only the reference genome, and which could constitute individual-specific binding events. GRAFIMO is available as an open-source tool, under the MIT license, at https://github.com/pinellolab/GRAFIMO and https://github.com/InfOmics/GRAFIMO.
Manuel Tognon, Vincenzo Bonnici, Erik Garrison, Rosalba Giugno, Luca Pinello
PLoS Comput. Biol.4
2021 SystemC Implementation of Stochastic Petri Nets for Simulation and Parameterization of Biological Networks
Nicola Bombieri, Silvia Scaffeo, Antonio Mastrandrea, Simone Caligola, Tommaso Carlucci, Franco Fummi, Carlo Laudanna, Gabriela Constantin, Rosalba Giugno
ACM Trans. Embed. Comput. Syst.9
2020 CRISPRitz: rapid, high-throughput and variant-aware in silico off-target site identification for CRISPR genome editing
abstract
MOTIVATION: Clustered regularly interspaced short palindromic repeats (CRISPR) technologies allow for facile genomic modification in a site-specific manner. A key step in this process is the in silico design of single guide RNAs to efficiently and specifically target a site of interest. To this end, it is necessary to enumerate all potential off-target sites within a given genome that could be inadvertently altered by nuclease-mediated cleavage. Currently available software for this task is limited by computational efficiency, variant support or annotation, and assessment of the functional impact of potential off-target effects. RESULTS: To overcome these limitations, we have developed CRISPRitz, a suite of software tools to support the design and analysis of CRISPR/CRISPR-associated (Cas) experiments. Using efficient data structures combined with parallel computation, we offer a rapid, reliable, and exhaustive search mechanism to enumerate a comprehensive list of putative off-target sites. As proof-of-principle, we performed a head-to-head comparison with other available tools on several datasets. This analysis highlighted the unique features and superior computational performance of CRISPRitz including support for genomic searching with DNA/RNA bulges and mismatches of arbitrary size as specified by the user as well as consideration of genetic variants (variant-aware). In addition, graphical reports are offered for coding and non-coding regions that annotate the potential impact of putative off-target sites that lie within regions of functional genomic annotation (e.g. insulator and chromatin accessible sites from the ENCyclopedia Of DNA Elements [ENCODE] project). AVAILABILITY AND IMPLEMENTATION: The software is freely available at: https://github.com/pinellolab/CRISPRitzhttps://github.com/InfOmics/CRISPRitz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Samuele Cancellieri, Matthew C. Canver, Nicola Bombieri, Rosalba Giugno, Luca Pinello
Bioinform.4
2019 LErNet: characterization of lncRNAs via context-aware network expansion and enrichment analysis
abstract
Long non-coding RNAs (lncRNAs) have recently acquired a boost of interest for their implication in several biological conditions. However, many of these elements are not yet characterized. LErNet is a method to in silico define and predict the roles of IncRNAs. The core of the approach is a network expansion algorithm which enriches the genomic context of IncRNAs. The context is built by integrating the genes encoding proteins that are found next to the non-coding elements both at genomic and system level. The pipeline is particularly useful in situations where the functions of discovered IncRNAs are not yet known. The results show both the outperformance of LErNet compared to enrichment approaches in literature and its robustness in case of partially missing context information. LErNet is provided as an R package. It is available at https://github.com/InfOmics/LErNet.
Vincenzo Bonnici, Simone Caligola, Giulia Fiorini, Luca Giudice, Rosalba Giugno
CIBCB5
2019 Automatic Parameterization of the Purine Metabolism Pathway through Discrete Event-based Simulation
abstract
Stochastic Petri Nets (SPN)are recognized as one of the standard formalisms to model metabolic networks. They allow incorporating randomness in the model and taking into account possible fluctuations and noise due to molecule interactions in the environment. Even though some frameworks have been proposed to implement and simulate SPN (e.g., Snoopy, Monalisa), they do not allow for automatic model parameterization, which is a crucial task to identify the network configurations that lead the model to satisfy certain biological properties. We present a framework to synthesize the SPN model of a metabolic network into executable code that can be simulated through a discrete event-based simulator. The framework allows the user to formally define the network properties to be observed and to automatically extrapolate, through Assertion-based Verification (ABV), the parameter configurations that lead the network to satisfy such properties. We applied the framework to model the purine metabolism and to reproduce the metabolomics data obtained from naive lymphocytes and autoreactive T cells implicated in the induction of experimental autoimmune disorders. We show system parameterization extrapolated by the framework to reproduce the experimental results and to simulate the model under different conditions.
Simone Caligola, Tommaso Carlucci, Franco Fummi, Carlo Laudanna, Gabriela Constantin, Nicola Bombieri, Rosalba Giugno
CIBCB7
2019 Efficient Simulation and Parametrization of Stochastic Petri Nets in SystemC: A Case study from Systems Biology
abstract
Stochastic Petri nets (SPN) are a form of Petri net where the transitions fire after a probabilistic and randomly determined delay. They are adopted in a wide range of applications thanks to their capability of incorporating randomness in the models and taking into account possible fluctuations and environmental noise. In Systems Biology, they are becoming a reference formalism to model metabolic networks, in which the noise due to molecule interactions in the environment plays a crucial role. Some frameworks have been proposed to implement and dynamically simulate SPN. Nevertheless, they do not allow for automatic model parametrization, which is a crucial task to identify the network configurations that lead the model to satisfy temporal properties of the model. This paper presents a framework that synthesizes the SPN models into SystemC code. The framework allows the user to formally define the network properties to be observed and to automatically extrapolate, through Assertion-based Verification (ABV), the parameter configurations that lead the network to satisfy such properties. We applied the framework to implement and simulate a complex biological network, i.e., the purine metabolism, with the aim of reproducing the metabolomics data obtained in-vitro from naive lymphocytes and autoreactive T cells implicated in the induction of experimental autoimmune disorders.
Simone Caligola, Tommaso Carlucci, Franco Fummi, Carlo Laudanna, Gabriela Constantin, Nicola Bombieri, Rosalba Giugno
FDL7
2019 Parallel Searching on Biological Networks
abstract
Software applications for biological networks analysis rely on graphs to model the structure interactions. A great part of them requires searching for subgraphs in a target graph or in collections of graphs. Even though very efficient algorithms have been defined to solve such a subgraph isomorphisms problem, the complexity of current real biological networks make their sequential execution time prohibitive. On the other hand, parallel architectures, from multi-core to manycore, have become pervasive to deal with the problem of the data size. Nevertheless, the sequential nature of the graph searching algorithms makes their implementation for parallel architectures very challenging. This paper presents three different parallel solutions for the graph searching problem. The first two target the exact search for multi-core CPUs and manycore GPUs, respectively. The third one targets the approximate search for GPUs, which handles node, edge, and node label mismatches. The paper shows how different techniques have been developed in all the solutions to reduce the search space complexity. The paper shows the performance of the proposed solutions on representative biological networks containing antiviral chemical compounds and protein interactions networks.
Nicola Bombieri, Vincenzo Bonnici, Rosalba Giugno
PDP3
2019 The 2017 Network Tools and Applications in Biology (NETTAB) workshop: aims, topics and outcomes
abstract
The 17th International NETTAB workshop was held in Palermo, Italy, on October 16-18, 2017. The special topic for the meeting was "Methods, tools and platforms for Personalised Medicine in the Big Data Era", but the traditional topics of the meeting series were also included in the event. About 40 scientific contributions were presented, including four keynote lectures, five guest lectures, and many oral communications and posters. Also, three tutorials were organised before and after the workshop. Full papers from some of the best works presented in Palermo were submitted for this Supplement of BMC Bioinformatics. Here, we provide an overview of meeting aims and scope. We also shortly introduce selected papers that have been accepted for publication in this Supplement, for a complete presentation of the outcomes of the meeting.
Paolo Romano 0001, Arnaud Céol, Andreas Dräger, Antonino Fiannaca, Rosalba Giugno, Massimo La Rosa, Luciano Milanesi, Ulrich Pfeffer, Riccardo Rizzo, Soo-Yong Shin, Junfeng Xia, Alfonso Urso
BMC Bioinform.5
2018 An Efficient Implementation of a Subgraph Isomorphism Algorithm for GPUs
Vincenzo Bonnici, Rosalba Giugno, Nicola Bombieri
BIBM2
2018 cuRnet: an R package for graph traversing on GPU
abstract
BACKGROUND: R has become the de-facto reference analysis environment in Bioinformatics. Plenty of tools are available as packages that extend the R functionality, and many of them target the analysis of biological networks. Several algorithms for graphs, which are the most adopted mathematical representation of networks, are well-known examples of applications that require high-performance computing, and for which classic sequential implementations are becoming inappropriate. In this context, parallel approaches targeting GPU architectures are becoming pervasive to deal with the execution time constraints. Although R packages for parallel execution on GPUs are already available, none of them provides graph algorithms. RESULTS: This work presents cuRnet, a R package that provides a parallel implementation for GPUs of the breath-first search (BFS), the single-source shortest paths (SSSP), and the strongly connected components (SCC) algorithms. The package allows offloading computing intensive applications to GPU devices for massively parallel computation and to speed up the runtime up to one order of magnitude with respect to the standard sequential computations on CPU. We have tested cuRnet on a benchmark of large protein interaction networks and for the interpretation of high-throughput omics data thought network analysis. CONCLUSIONS: cuRnet is a R package to speed up graph traversal and analysis through parallel computation on GPUs. We show the efficiency of cuRnet applied both to biological network analysis, which requires basic graph algorithms, and to complex existing procedures built upon such algorithms.
Vincenzo Bonnici, Federico Busato, Stefano Aldegheri, Murodzhon Akhmedov, Luciano Cascione, Alberto Arribas Carmena, Francesco Bertoni, Nicola Bombieri, Ivo Kwee, Rosalba Giugno
BMC Bioinform.10
2018 Correction to: cuRnet: an R package for graph traversing on GPU
abstract
After publication of this supplement article [1], it was brought to our attention that reference 10 and reference 12 in the article are incorrect.
Vincenzo Bonnici, Federico Busato, Stefano Aldegheri, Murodzhon Akhmedov, Luciano Cascione, Alberto Arribas Carmena, Francesco Bertoni, Nicola Bombieri, Ivo Kwee, Rosalba Giugno
BMC Bioinform.10
2018 Arena-Idb: a platform to build human non-coding RNA interaction networks
abstract
BACKGROUND: High throughput technologies have provided the scientific community an unprecedented opportunity for large-scale analysis of genomes. Non-coding RNAs (ncRNAs), for a long time believed to be non-functional, are emerging as one of the most important and large family of gene regulators and key elements for genome maintenance. Functional studies have been able to assign to ncRNAs a wide spectrum of functions in primary biological processes, and for this reason they are assuming a growing importance as a potential new family of cancer therapeutic targets. Nevertheless, the number of functionally characterized ncRNAs is still too poor if compared to the number of new discovered ncRNAs. Thus platforms able to merge information from available resources addressing data integration issues are necessary and still insufficient to elucidate ncRNAs biological roles. RESULTS: In this paper, we describe a platform called Arena-Idb for the retrieval of comprehensive and non-redundant annotated ncRNAs interactions. Arena-Idb provides a framework for network reconstruction of ncRNA heterogeneous interactions (i.e., with other type of molecules) and relationships with human diseases which guide the integration of data, extracted from different sources, via mapping of entities and minimization of ambiguity. CONCLUSIONS: Arena-Idb provides a schema and a visualization system to integrate ncRNA interactions that assists in discovering ncRNA functions through the extraction of heterogeneous interaction networks. The Arena-Idb is available at http://arenaidb.ba.itb.cnr.it.
Vincenzo Bonnici, Giorgio De Caro, Giorgio Constantino, Sabino Liuni, Domenica D'Elia, Nicola Bombieri, Flavio Licciulli, Rosalba Giugno
BMC Bioinform.8
2018 INBIA: a boosting methodology for proteomic network inference
abstract
BACKGROUND: The analysis of tissue-specific protein interaction networks and their functional enrichment in pathological and normal tissues provides insights on the etiology of diseases. The Pan-cancer proteomic project, in The Cancer Genome Atlas, collects protein expressions in human cancers and it is a reference resource for the functional study of cancers. However, established protocols to infer interaction networks from protein expressions are still missing. RESULTS: We have developed a methodology called Inference Network Based on iRefIndex Analysis (INBIA) to accurately correlate proteomic inferred relations to protein-protein interaction (PPI) networks. INBIA makes use of 14 network inference methods on protein expressions related to 16 cancer types. It uses as reference model the iRefIndex human PPI network. Predictions are validated through non-interacting and tissue specific PPI networks resources. The first, Negatome, takes into account likely non-interacting proteins by combining both structure properties and literature mining. The latter, TissueNet and GIANT, report experimentally verified PPIs in more than 50 human tissues. The reliability of the proposed methodology is assessed by comparing INBIA with PERA, a tool which infers protein interaction networks from Pathway Commons, by both functional and topological analysis. CONCLUSION: Results show that INBIA is a valuable approach to predict proteomic interactions in pathological conditions starting from the current knowledge of human protein interactions.
Davide S. Sardina, Giovanni Micale, Alfredo Ferro, Alfredo Pulvirenti, Rosalba Giugno
BMC Bioinform.5
2018 Fast analytical methods for finding significant labeled graph motifs
Giovanni Micale, Rosalba Giugno, Alfredo Ferro, Misael Mongiovì, Dennis E. Shasha, Alfredo Pulvirenti
Data Min. Knowl. Discov.2
2017 A novel computational method for inferring competing endogenous interactions
abstract
Posttranscriptional cross talk and communication between genes mediated by microRNA response element (MREs) yield large regulatory competing endogenous RNA (ceRNA) networks. Their inference may improve the understanding of pathologies and shed new light on biological mechanisms. A variety of RNA: messenger RNA, transcribed pseudogenes, noncoding RNA, circular RNA and proteins related to RNA-induced silencing complex complex interacting with RNA transfer and ribosomal RNA have been experimentally proved to be ceRNAs. We retrace the ceRNA hypothesis of posttranscriptional regulation from its original formulation [Salmena L, Poliseno L, Tay Y, et al. Cell 2011;146:353-8] to the most recent experimental and computational validations. We experimentally analyze the methods in literature [Li J-H, Liu S, Zhou H, et al. Nucleic Acids Res 2013;42:D92-7; Sumazin P, Yang X, Chiu H-S, et al. Cell 2011;147:370-81; Sarver AL, Subramanian S. Bioinformation 2012;8:731-3] comparing them with a general machine learning approach, called ceRNA predIction Algorithm, evaluating the performance in predicting novel MRE-based ceRNAs.
Davide S. Sardina, Salvatore Alaimo, Alfredo Ferro, Alfredo Pulvirenti, Rosalba Giugno
Briefings Bioinform.5
2017 On the Variable Ordering in Subgraph Isomorphism Algorithms
abstract
Graphs are mathematical structures to model several biological data. Applications to analyze them require to apply solutions for the subgraph isomorphism problem, which is NP-complete. Here, we investigate the existing strategies to reduce the subgraph isomorphism algorithm running time with emphasis on the importance of the order with which the graph vertices are taken into account during the search, called variable ordering, and its incidence on the total running time of the algorithms. We focus on two recent solutions, which are based on an effective variable ordering strategy. We discuss their comparison both with the variable ordering strategies reviewed in the paper and the other algorithms present in the ICPR2014 contest on graph matching algorithms for pattern search in biological databases.
Vincenzo Bonnici, Rosalba Giugno
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 APPAGATO: an APproximate PArallel and stochastic GrAph querying TOol for biological networks
abstract
MOTIVATION: Biological network querying is a problem requiring a considerable computational effort to be solved. Given a target and a query network, it aims to find occurrences of the query in the target by considering topological and node similarities (i.e. mismatches between nodes, edges, or node labels). Querying tools that deal with similarities are crucial in biological network analysis because they provide meaningful results also in case of noisy data. In addition, as the size of available networks increases steadily, existing algorithms and tools are becoming unsuitable. This is rising new challenges for the design of more efficient and accurate solutions. RESULTS: This paper presents APPAGATO, a stochastic and parallel algorithm to find approximate occurrences of a query network in biological networks. APPAGATO handles node, edge and node label mismatches. Thanks to its randomic and parallel nature, it applies to large networks and, compared with existing tools, it provides higher performance as well as statistically significant more accurate results. Tests have been performed on protein-protein interaction networks annotated with synthetic and real gene ontology terms. Case studies have been done by querying protein complexes among different species and tissues. AVAILABILITY AND IMPLEMENTATION: APPAGATO has been developed on top of CUDA-C ++ Toolkit 7.0 framework. The software is available online http://profs.sci.univr.it/∼bombieri/APPAGATO CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vincenzo Bonnici, Federico Busato, Giovanni Micale, Nicola Bombieri, Alfredo Pulvirenti, Rosalba Giugno
Bioinform.6
2015 A SystemC Platform for Signal Transduction Modelling and Simulation in Systems Biology
abstract
Signal transduction is a class of cell's biological processes, which are commonly represented as highly concurrent reactive systems. In the Systems Biology community, modelling and simulation of signal transduction require overcoming issues like discrete event-based execution of complex systems, description from building blocks through composition and encapsulation, description at different levels of granularity, methods for abstraction and refinement. This paper presents a signal transduction modelling and simulation platform based on SystemC, and shows how the platform allows handling the system complexity by modelling it at different abstraction levels. The paper reports the results obtained by applying the platform to model the intracellular signalling network controlling integrin activation mediating leukocyte recruitment from the blood into the tissues.
Rosario Distefano, Franco Fummi, Carlo Laudanna, Nicola Bombieri, Rosalba Giugno
ACM Great Lakes Symposium on VLSI5
2013 Drug-target interaction prediction through domain-tuned network-based inference
abstract
MOTIVATION: The identification of drug-target interaction (DTI) represents a costly and time-consuming step in drug discovery and design. Computational methods capable of predicting reliable DTI play an important role in the field. Recently, recommendation methods relying on network-based inference (NBI) have been proposed. However, such approaches implement naive topology-based inference and do not take into account important features within the drug-target domain. RESULTS: In this article, we present a new NBI method, called domain tuned-hybrid (DT-Hybrid), which extends a well-established recommendation technique by domain-based knowledge including drug and target similarity. DT-Hybrid has been extensively tested using the last version of an experimentally validated DTI database obtained from DrugBank. Comparison with other recently proposed NBI methods clearly shows that DT-Hybrid is capable of predicting more reliable DTIs. AVAILABILITY: DT-Hybrid has been developed in R and it is available, along with all the results on the predictions, through an R package at the following URL: http://sites.google.com/site/ehybridalgo/.
Salvatore Alaimo, Alfredo Pulvirenti, Rosalba Giugno, Alfredo Ferro
Bioinform.3
2013 A subgraph isomorphism algorithm and its application to biochemical data
abstract
BACKGROUND: Graphs can represent biological networks at the molecular, protein, or species level. An important query is to find all matches of a pattern graph to a target graph. Accomplishing this is inherently difficult (NP-complete) and the efficiency of heuristic algorithms for the problem may depend upon the input graphs. The common aim of existing algorithms is to eliminate unsuccessful mappings as early as and as inexpensively as possible. RESULTS: We propose a new subgraph isomorphism algorithm which applies a search strategy to significantly reduce the search space without using any complex pruning rules or domain reduction procedures. We compare our method with the most recent and efficient subgraph isomorphism algorithms (VFlib, LAD, and our C++ implementation of FocusSearch which was originally distributed in Modula2) on synthetic, molecules, and interaction networks data. We show a significant reduction in the running time of our approach compared with these other excellent methods and show that our algorithm scales well as memory demands increase. CONCLUSIONS: Subgraph isomorphism algorithms are intensively used by biochemical tools. Our analysis gives a comprehensive comparison of different software approaches to subgraph isomorphism highlighting their weaknesses and strengths. This will help researchers make a rational choice among methods depending on their application. We also distribute an open-source package including our system and our own C++ implementation of FocusSearch together with all the used datasets (http://ferrolab.dmi.unict.it/ri.html). In future work, our findings may be extended to approximate subgraph isomorphism algorithms.
Vincenzo Bonnici, Rosalba Giugno, Alfredo Pulvirenti, Dennis E. Shasha, Alfredo Ferro
BMC Bioinform.2
2013 VIRGO: visualization of A-to-I RNA editing sites in genomic sequences
abstract
BACKGROUND: RNA Editing is a type of post-transcriptional modification that takes place in the eukaryotes. It alters the sequence of primary RNA transcripts by deleting, inserting or modifying residues. Several forms of RNA editing have been discovered including A-to-I, C-to-U, U-to-C and G-to-A. In recent years, the application of global approaches to the study of A-to-I editing, including high throughput sequencing, has led to important advances. However, in spite of enormous efforts, the real biological mechanism underlying this phenomenon remains unknown. DESCRIPTION: In this work, we present VIRGO (http://atlas.dmi.unict.it/virgo/), a web-based tool that maps Ato-G mismatches between genomic and EST sequences as candidate A-to-I editing sites. VIRGO is built on top of a knowledge-base integrating information of genes from UCSC, EST of NCBI, SNPs, DARNED, and Next Generations Sequencing data. The tool is equipped with a user-friendly interface allowing users to analyze genomic sequences in order to identify candidate A-to-I editing sites. CONCLUSIONS: VIRGO is a powerful tool allowing a systematic identification of putative A-to-I editing sites in genomic sequences. The integration of NGS data allows the computation of p-values and adjusted p-values to measure the mapped editing sites confidence. The whole knowledge base is available for download and will be continuously updated as new NGS data becomes available.
Rosario Distefano, Giovanni Nigita, Valentina Macca, Alessandro Laganà, Rosalba Giugno, Alfredo Pulvirenti, Alfredo Ferro
BMC Bioinform.5
2013 Bioinformatics in Italy: BITS 2012, the ninth annual meeting of the Italian Society of Bioinformatics
abstract
Abstract The BITS2012 meeting, held in Catania on May 2-4, 2012, brought together almost 100 Italian researchers working in the field of Bioinformatics, as well as students in the same or related disciplines. About 90 original research works were presented either as oral communication or as posters, representing a landscape of Italian current research in bioinformatics. This preface provides a brief overview of the meeting and introduces the manuscripts that were accepted for publication in this supplement, after a strict and careful peer-review by an International board of referees.
Carmela Gissi, Paolo Romano 0001, Alfredo Ferro, Rosalba Giugno, Alfredo Pulvirenti, Angelo M. Facchiano, Manuela Helmer-Citterich
BMC Bioinform.4
2013 Enhancing density-based clustering: Parameter reduction and outlier detection
Carmelo Cassisi, Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti
Inf. Syst.3
2012 miR-EdiTar: a database of predicted A-to-I edited miRNA target sites
abstract
MOTIVATION: A-to-I RNA editing is an important mechanism that consists of the conversion of specific adenosines into inosines in RNA molecules. Its dysregulation has been associated to several human diseases including cancer. Recent work has demonstrated a role for A-to-I editing in microRNA (miRNA)-mediated gene expression regulation. In fact, edited forms of mature miRNAs can target sets of genes that differ from the targets of their unedited forms. The specific deamination of mRNAs can generate novel binding sites in addition to potentially altering existing ones. RESULTS: This work presents miR-EdiTar, a database of predicted A-to-I edited miRNA binding sites. The database contains predicted miRNA binding sites that could be affected by A-to-I editing and sites that could become miRNA binding sites as a result of A-to-I editing. AVAILABILITY: miR-EdiTar is freely available online at http://microrna.osumc.edu/mireditar. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alessandro Laganà, Alessio Paone, Dario Veneziano, Luciano Cascione, Pierluigi Gasparini, Stefania Carasi, Francesco Russo 0004, Giovanni Nigita, Valentina Macca, Rosalba Giugno, Alfredo Pulvirenti, Dennis E. Shasha, Alfredo Ferro, Carlo Maria Croce
Bioinform.10
2011 DBStrata: a system for density-based clustering and outlier detection based on stratification
abstract
Clustering is a widely used unsupervised data mining technique. In density-based clustering, a cluster is defined as a connected dense component and grows in the direction set by the density. In this paper we present a software system called DBStrata that implements the density-based clustering architecture together with several extensions able to boost the clustering performances and to efficiently identify outliers.
Marco Aliotta, Andrea Cannata, Carmelo Cassisi, Rosalba Giugno, Placido Montalto, Alfredo Pulvirenti
SISAP4
2011 Obstacles constrained group mobility models in event-driven wireless networks with movable base stations
S. Cristaldi, Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti
Ad Hoc Networks3
2011 Editorial
abstract
NETTAB workshops are a series of international events on ‘Network Tools and Applications in Biology’ that are aimed at presenting and discussing emerging Information and Communication Technologies whose adoption in support of biology appear to be of particular interest. Since 2001, many different topics were discussed including: XML standardization for data integration (2001), multi agent systems (2002), scientific workflows (2005), Web Services (2006) and the Semantic Web (2007). The NETTAB 2009 workshop that was held at the Mathematics and Computer Science Department of the University of Catania, Italy, 10–13 June 2009, was focused on ‘Collaborative Bioinformatics Research and Development’ and on ‘Tools for RNA Analysis’. The workshop included many original contributions devoted to these innovative research domains, the best of which have been carefully peer reviewed and included in this Special Issue on ‘Collaborative Bioinformatics and RNA Analysis’. There is a clear switch from previous focus themes since these are aimed at different aspects of data integration, ranging from standardization, to automation of procedures. In this case, the accent is on human collaboration and ways to implement it. While many decades ago research was mainly carried out by researchers in a single institute or laboratory, it is nowadays common that researchers from different, far away laboratories carry out experiments in collaboration. However, such a collaboration does not usually involve sharing of research platforms. Research is done independently and results are then exchanged. The advent of high-speed communication networks, together with the development of software enabling actual sharing of data, is quickly leading to new ways of conducting experiments and making research in collaboration. Romano et al. present a survey of some of the tools available for collaborative research and development in the first paper of this Issue. We review the principles and some social networking applications are able to support scientific collaboration, discuss the wiki approach to collaborative document creation and review some biological wikis. Finally, we present collaborative platforms for software development and some examples of Learning Management Systems as tools for bioinformatics. Some good examples of how this approach can effectively be implemented are provided by the following two papers. Splendiani et al. present DC-THERA Directory, an information system designed to support collaborative knowledge management in DC-THERA. This is a European ‘Network of Excellence’ (NoE) involving a large interdisciplinary research community that is focused on the development of novel immunotherapies derived from research in dendritic cell immunobiology that is producing large data sets and a great wealth of knowledge. The Directory is based on Semantic Web technologies. The authors show how these technologies, along with properly defined metadata such as reference vocabularies and ontologies can support a better organization of the knowledge and constitute a model for an effective data management solution during and beyond a project lifecycle. The dramatic increase of data being generated by high-throughput technologies are opening a new universe of problems that will certainly provide great challenges in the coming years. The bottom line of such problems is represented by data and knowledge integration. Due to the vast amount of data, its curation cannot rely on individual researchers. Accordingly, community-driven tools are becoming very popular in the life sciences. Navas-Delgado et al. describe the collaborative improvement of metabolic pathways by means of a new tool with social networking capabilities. The NETTAB workshop was also focused on tools for RNA analysis. We continue the special issue with two papers focusing on microRNAs (miRNAs) and non-coding RNAs (ncRNAs). miRNAs are a class of small, non-coding regulatory RNAs that are crucial in post-transcriptional gene silencing. They may regulate gene expression by binding the 3′-UTR site of their targets. miRNAs may then play a key role in many biological processes, such as cell proliferation, cell death and oncogenesis. Since their discovery, computational approaches have been important in understanding the biology of miRNAs. Web-based-miRNA databases now provide thousands of miRNA sequences, annotations and putative target genes. Although much research has been done, target prediction is still a challenging task and due to the high false positive rate it is a major obstacle in biological validation of miRNA/mRNA interaction. In this Special Issue, Corrada et al., after surveying target prediction tools, propose a meta-system to improve overall prediction. Their tool is based on a knowledge base that is built on top of the most reliable prediction methods implementing an integration strategy for computing ranked miRNA-target lists together with functional annotations. The RNA of an organism is mostly composed of molecules that are not translated into proteins, so called ‘non-coding RNA’. Some ncRNAs, such as transfer RNA and ribosomal RNA, have well-known functions, whereas, for the remaining, their activity is still unknown or not validated. Work aimed at understanding ncRNA suggests that they have functions that are as yet unknown. Recalling that biological activity is largely due to the three dimensional (3D) structure of molecules, tools aimed at predicting the 3D structure of RNA are of an extreme relevance in this context. In this special issue, ModeRNA, a tool for the prediction of the tertiary structure of a RNA, is included. In their paper, Rother et al. guide the readers through the preparation of the input to the evaluation of the results, and to the prediction of tRNA molecules from Escherichia Coli, as an example of its practical utilization. We conclude this special issue with a review on metagenomics data analysis and associated issues by Ribeca and Valiente. It is well known that metagenomics may also allow the sequencing and analysis of genomes of organisms that cannot grow in laboratories either because their lives depend on other organisms or because they are extinct. In the former case, DNA is taken from the environment, while, in the letter, it is extracted from fossils. Metagenomics applications include the monitoring of bacterial composition of an individual's gastrointestinal fauna as well as the detection of viruses in the environment. Next-generation sequencing is used and its relative low cost makes it affordable for many laboratories. From the computational point of view, metagenomics analysis is conducted by assembling random short reads, aligning them with full genomes and performing taxonomic classification. The major issues include the enormous amounts of resulting data, which make the analysis not trivial, the sequencing error models and the assignment of reads to the correct species.
Paolo Romano 0001, Rosalba Giugno, Alfredo Pulvirenti
Briefings Bioinform.2
2011 Tools and collaborative environments for bioinformatics research
abstract
Advanced research requires intensive interaction among a multitude of actors, often possessing different expertise and usually working at a distance from each other. The field of collaborative research aims to establish suitable models and technologies to properly support these interactions. In this article, we first present the reasons for an interest of Bioinformatics in this context by also suggesting some research domains that could benefit from collaborative research. We then review the principles and some of the most relevant applications of social networking, with a special attention to networks supporting scientific collaboration, by also highlighting some critical issues, such as identification of users and standardization of formats. We then introduce some systems for collaborative document creation, including wiki systems and tools for ontology development, and review some of the most interesting biological wikis. We also review the principles of Collaborative Development Environments for software and show some examples in Bioinformatics. Finally, we present the principles and some examples of Learning Management Systems. In conclusion, we try to devise some of the goals to be achieved in the short term for the exploitation of these technologies.
Paolo Romano 0001, Rosalba Giugno, Alfredo Pulvirenti
Briefings Bioinform.2
2010 An Efficient Duplicate Record Detection Using q-Grams Array Inverted Index
Alfredo Ferro, Rosalba Giugno, Piera Laura Puglisi, Alfredo Pulvirenti
DaWak2
2010 MySQL Data Mining: Extending MySQL to Support Data Mining Primitives (Demo)
Alfredo Ferro, Rosalba Giugno, Piera Laura Puglisi, Alfredo Pulvirenti
KES (3)2
2010 SING: Subgraph search In Non-homogeneous Graphs
abstract
BACKGROUND: Finding the subgraphs of a graph database that are isomorphic to a given query graph has practical applications in several fields, from cheminformatics to image understanding. Since subgraph isomorphism is a computationally hard problem, indexing techniques have been intensively exploited to speed up the process. Such systems filter out those graphs which cannot contain the query, and apply a subgraph isomorphism algorithm to each residual candidate graph. The applicability of such systems is limited to databases of small graphs, because their filtering power degrades on large graphs. RESULTS: In this paper, SING (Subgraph search In Non-homogeneous Graphs), a novel indexing system able to cope with large graphs, is presented. The method uses the notion of feature, which can be a small subgraph, subtree or path. Each graph in the database is annotated with the set of all its features. The key point is to make use of feature locality information. This idea is used to both improve the filtering performance and speed up the subgraph isomorphism task. CONCLUSIONS: Extensive tests on chemical compounds, biological networks and synthetic graphs show that the proposed system outperforms the most popular systems in query time over databases of medium and large graphs. Other specific tests show that the proposed system is effective for single large graphs.
Raffaele Di Natale, Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Alfredo Pulvirenti, Dennis E. Shasha
BMC Bioinform.3
2009 BitCube: A Bottom-Up Cubing Engineering
Alfredo Ferro, Rosalba Giugno, Piera Laura Puglisi, Alfredo Pulvirenti
DaWaK2
2009 Distributed randomized algorithms for low-support data mining
abstract
Data mining in distributed systems has been facilitated by using high-support association rules. Less attention has been paid to distributed low-support/high-correlation data mining. This has proved useful in several fields such as computational biology, wireless networks, web mining, security and rare events analysis in industrial plants. In this paper we present distributed versions of efficient algorithms for low-support/high-correlation data mining such as Min-Hashing, K-Min-Hashing and Locality-Sensitive-Hashing. Experimental results on real data concerning scalability, speed-up and network traffic are reported.
Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Alfredo Pulvirenti
IPDPS2
2008 GraphFind: enhancing graph searching by low support data mining techniques
abstract
BACKGROUND: Biomedical and chemical databases are large and rapidly growing in size. Graphs naturally model such kinds of data. To fully exploit the wealth of information in these graph databases, a key role is played by systems that search for all exact or approximate occurrences of a query graph. To deal efficiently with graph searching, advanced methods for indexing, representation and matching of graphs have been proposed. RESULTS: This paper presents GraphFind. The system implements efficient graph searching algorithms together with advanced filtering techniques that allow approximate search. It allows users to select candidate subgraphs rather than entire graphs. It implements an effective data storage based also on low-support data mining. CONCLUSIONS: GraphFind is compared with Frowns, GraphGrep and gIndex. Experiments show that GraphFind outperforms the compared systems on a very large collection of small graphs. The proposed low-support mining technique which applies to any searching system also allows a significant index space reduction.
Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Alfredo Pulvirenti, Dmitry Skripin, Dennis E. Shasha
BMC Bioinform.2
2007 NetMatch: a Cytoscape plugin for searching biological networks
abstract
UNLABELLED: NetMatch is a Cytoscape plugin which allows searching biological networks for subcomponents matching a given query. Queries may be approximate in the sense that certain parts of the subgraph-query may be left unspecified. To make the query creation process easy, a drawing tool is provided. Cytoscape is a bioinformatics software platform for the visualization and analysis of biological networks. AVAILABILITY: The full package, a tutorial and associated examples are available at the following web sites: http://alpha.dmi.unict.it/~ctnyu/netmatch.html, http://baderlab.org/Software/NetMatch.
Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti, Dmitry Skripin, Gary D. Bader, Dennis E. Shasha
Bioinform.2
2007 Sequence similarity is more relevant than species specificity in probabilistic backtranslation
abstract
BACKGROUND: Backtranslation is the process of decoding a sequence of amino acids into the corresponding codons. All synthetic gene design systems include a backtranslation module. The degeneracy of the genetic code makes backtranslation potentially ambiguous since most amino acids are encoded by multiple codons. The common approach to overcome this difficulty is based on imitation of codon usage within the target species. RESULTS: This paper describes EasyBack, a new parameter-free, fully-automated software for backtranslation using Hidden Markov Models. EasyBack is not based on imitation of codon usage within the target species, but instead uses a sequence-similarity criterion. The model is trained with a set of proteins with known cDNA coding sequences, constructed from the input protein by querying the NCBI databases with BLAST. Unlike existing software, the proposed method allows the quality of prediction to be estimated. When tested on a group of proteins that show different degrees of sequence conservation, EasyBack outperforms other published methods in terms of precision. CONCLUSION: The prediction quality of a protein backtranslation methis markedly increased by replacing the criterion of most used codon in the same species with a Hidden Markov Model trained with a set of most similar sequences from all species. Moreover, the proposed method allows the quality of prediction to be estimated probabilistically.
Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti, Cinzia Di Pietro, Michele Purrello, Marco Ragusa
BMC Bioinform.2
2006 Distributed antipole clustering for efficient data search and management in Euclidean and metric spaces
abstract
In this paper a simple and efficient distributed version of the introduced antipole clustering algorithm for general metric spaces is proposed. This combines ideas from the M-tree, the multi-vantage point structure and the FQ-tree to create a new structure in the "bisector tree" class, called the antipole tree. Bisection is based on the proximity to an "antipole" pair of elements generated by a suitable linear randomized tournament. The final winners (A, B) of such a tournament are far enough apart to approximate the diameter of the splitting set. A simple linear algorithm computing antipoles in Euclidean spaces with exponentially small approximation ratio is proposed. The antipole tree clustering has been shown to be very effective in important applications such as range and k-nearest neighbor searching, mobile objects clustering in centralized wireless networks with movable base stations and multiple alignment of biological sequences. In many of such applications an efficient distributed clustering algorithm is needed. In the proposed distributed versions of antipole clustering the amount of data passed from one node to another is either constant or proportional to the number of nodes in the network. The distributed antipole tree is equipped with additional information in order to perform efficient range search and dynamic clusters management. This is achieved by adding to the randomized tournaments technique, methodologies taken from established systems such as BFR and BIRCH*. Experiments show the good performance of the proposed algorithms on both real and synthetic data
Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Giuseppe Pigola, Alfredo Pulvirenti
IPDPS2
2004 Locally sensitive backtranslation based on multiple sequence alignment
abstract
Backtranslation is the process of decoding an amino acid sequence into a corresponding nucleic acid. Classical methods are based on the construction of a codon usage table by clustering and detection of the most probable codon used for each amino acid. We present a new method for backtranslation which is sensitive to the local position of the amino acid in the input sequence. The method makes use of multiple sequence alignment of the set of proteins under analysis. A local codon usage table stores for each amino acid X and for each position of X in the alignment the most used codon. We compared our method with EMBOSS using both ClustalW and AntiClustAl for multiple sequence alignment. Experiments showed that our method outperforms EMBOSS in terms of precision of backtranslation: the matching between the proteins obtained by our method and the original protein templates is clearly superior to that obtained by EMBOSS. This enforces the validity of a locally sensitive approach.
Rosalba Giugno, Alfredo Pulvirenti, Marco Ragusa, Loredana Facciola, Laura Patelmo, Valentina Di Pietro, Cinzia Di Pietro, Michele Purrello, Alfredo Ferro
CIBCB1
2003 Temporal Probabilistic Object Bases
abstract
There are numerous applications where we have to deal with temporal uncertainty associated with objects. The ability to automatically store and manipulate time, probabilities, and objects is important. We propose a data model and algebra for temporal probabilistic object bases (TPOBs), which allows us to specify the probability with which an event occurs at a given time point. In explicit TPOB-instances, the sets of time points along with their probability intervals are explicitly enumerated. In implicit TPOB-instances, sets of time points are expressed by constraints and their probability intervals by probability distribution functions. Thus, implicit object base instances are succinct representations of explicit ones; they allow for an efficient implementation of algebraic operations, while their explicit counterparts make defining algebraic operations easy. We extend the relational algebra to both explicit and implicit instances and prove that the operations on implicit instances correctly implement their counterpart on explicit instances.
Veronica Biazzo, Rosalba Giugno, Thomas Lukasiewicz, V. S. Subrahmanian
IEEE Trans. Knowl. Data Eng.2
2002 P-SHOQ(D): A Probabilistic Extension of SHOQ(D) for Probabilistic Ontologies in the Semantic Web
Rosalba Giugno, Thomas Lukasiewicz
JELIA1
2002 Algorithmics and Applications of Tree and Graph Searching
abstract
Modern search engines answer keyword-based queries extremely efficiently. The impressive speed is due to clever inverted index structures, caching, a domain-independent knowledge of strings, and thousands of machines. Several research efforts have attempted to generalize keyword search to keytree and keygraph searching, because trees and graphs have many applications in next-generation database systems. This paper surveys both algorithms and applications, giving some emphasis to our own work.
Dennis E. Shasha, Jason Tsong-Li Wang, Rosalba Giugno
PODS3
2001 Best-Match Retrieval for Structured Images
abstract
Propose a methodology for fast best-match retrieval of structured images. A triangle inequality property for the tree-distance introduced by Oflazer (1997) is proven. This property is, in turn, applied to obtain a saturation algorithm of the trie used to store the database of the collection of pictures. The new approach can be considered as a substantial optimization of Oflazer's technique and can be applied to the retrieval of homogeneous hierarchically structured objects of any kind. The new technique inscribes itself in the number of distance-based search strategies and it is of interest for the indexing and maintenance of large collections of historical and pictorial data. We demonstrate the proposed approach on an example and report data about the speed-up that it introduces in query processing. Direct comparison with an MVP-trees algorithm is also presented.
Alfredo Ferro, Giovanni Gallo, Rosalba Giugno, Alfredo Pulvirenti
IEEE Trans. Pattern Anal. Mach. Intell.3