Alfredo Pulvirenti

dblp:56/727 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-9764-0295ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 6 since 2021Databases, data management, data science and information retrieval · 9 · 3 since 2021Artificial intelligence and machine learning · 5 · 1 since 2021Systems, architecture and hardware · 3Computer networks · 1
YearPublicationVenuePosition
2026 CMiner: An Algorithm to Discover Frequent Structures in Conceptual Models
Simone Avellino, Emanuele Valore, Giovanni Micale, Antonio Di Maria, Mattia Fumagalli, Tiago Prince Sales, Alfredo Pulvirenti, Diego Calvanese
EDBT7
2025 MultiGraphMatch: A Subgraph Matching Algorithm for Multigraphs
abstract
Subgraph matching is the problem of finding all the occurrences of a small graph, called the query, in a larger graph, called the target. Although the problem has been widely studied in simple graphs, few solutions have been proposed for multigraphs, in which two nodes can be connected by multiple edges, each denoting a possibly different type of relationship. In our new algorithm MultiGraphMatch (MGM), nodes and edges can be associated with labels and multiple properties. MGM introduces a novel data structure called bit matrix to efficiently index both the query and the target and filter the set of target edges that are matchable with each query edge. In addition, the algorithm proposes a new technique for ordering the processing of query edges based on the cardinalities of the sets of matchable edges. Using the CYPHER query definition language, MGM can perform queries with logical conditions on node and edge labels. We compare MGM with SuMGra and graph database systems Memgraph and Neo4J, showing comparable or better performance in all queries on a wide variety of synthetic and real-world graphs.
Giovanni Micale, Antonio Di Maria, Roberto Grasso, Vincenzo Bonnici, Alfredo Ferro, Dennis E. Shasha, Rosalba Giugno, Alfredo Pulvirenti
ACM Trans. Knowl. Discov. Data8
2024 NetMe 2.0: a web-based platform for extracting and modeling knowledge from biomedical literature as a labeled graph
abstract
MOTIVATION: The rapid increase of bio-medical literature makes it harder and harder for scientists to keep pace with the discoveries on which they build their studies. Therefore, computational tools have become more widespread, among which network analysis plays a crucial role in several life-science contexts. Nevertheless, building correct and complete networks about some user-defined biomedical topics on top of the available literature is still challenging. RESULTS: We introduce NetMe 2.0, a web-based platform that automatically extracts relevant biomedical entities and their relations from a set of input texts-i.e. in the form of full-text or abstract of PubMed Central's papers, free texts, or PDFs uploaded by users-and models them as a BioMedical Knowledge Graph (BKG). NetMe 2.0 also implements an innovative Retrieval Augmented Generation module (Graph-RAG) that works on top of the relationships modeled by the BKG and allows the distilling of well-formed sentences that explain their content. The experimental results show that NetMe 2.0 can infer comprehensive and reliable biological networks with significant Precision-Recall metrics when compared to state-of-the-art approaches. AVAILABILITY AND IMPLEMENTATION: https://netme.click/.
Antonio Di Maria, Lorenzo Bellomo, Fabrizio Billeci, Alfio Cardillo, Salvatore Alaimo, Paolo Ferragina, Alfredo Ferro, Alfredo Pulvirenti
Bioinform.8
2024 ArcMatch: high-performance subgraph matching for labeled graphs by exploiting edge domains
abstract
Abstract Consider a large labeled graph (network), denoted the target. Subgraph matching is the problem of finding all instances of a small subgraph, denoted the query, in the target graph. Unlike the majority of existing methods that are restricted to graphs with labels solely on vertices, our proposed approach, named can effectively handle graphs with labels on both vertices and edges. ntroduces an efficient new vertex/edge domain data structure filtering procedure to speed up subgraph queries. The procedure, called path-based reduction, filters initial domains by scanning them for paths up to a specified length that appear in the query graph. Additionally, ncorporates existing techniques like variable ordering and parent selection, as well as adapting the core search process, to take advantage of the information within edge domains. Experiments in real scenarios such as protein–protein interaction graphs, co-authorship networks, and email networks, show that s faster than state-of-the-art systems varying the number of distinct vertex labels over the whole target graph and query sizes.
Vincenzo Bonnici, Roberto Grasso, Giovanni Micale, Antonio Di Maria, Dennis E. Shasha, Alfredo Pulvirenti, Rosalba Giugno
Data Min. Knowl. Discov.6
2024 MultiplexSAGE: A Multiplex Embedding Algorithm for Inter-Layer Link Prediction
abstract
Research on graph representation learning has received great attention in recent years. However, most of the studies so far have focused on the embedding of single-layer graphs. The few studies dealing with the problem of representation learning of multilayer structures rely on the strong hypothesis that the inter-layer links are known, and this limits the range of possible applications. Here we propose MultiplexSAGE, a generalization of the GraphSAGE algorithm that allows embedding multiplex networks. We show that MultiplexSAGE is capable to reconstruct both the intra-layer and the inter-layer connectivity, outperforming competing methods. Next, through a comprehensive experimental analysis, we shed light also on the performance of the embedding, both in simple and multiplex networks, showing that both the density of the graph and the randomness of the links strongly influences the quality of the embedding.
Luca Gallo, Vito Latora, Alfredo Pulvirenti
IEEE Trans. Neural Networks Learn. Syst.3
2023 MASFENON: Multi-Agent Adaptive Simulation Framework for Evolution in Networks of Networks
abstract
In this paper, we present MASFENON, a novel multi-agent network interactions simulation algorithm that allows us to consider the dynamics within and between each agent and its associated network. MASFENON can be applied in various domains. Here, we will focus on an application related to epidemics network modeling. By combining propagation, dissipation, and conservation principles with some principles inspired by chaos theory, MASFENON offers a novel approach to model infectious disease spread across communities.
Giorgio Locicero, Salvatore Alaimo, Alfredo Ferro, Alfredo Pulvirenti
BIBM4
2023 DEGGs: an R package with shiny app for the identification of differentially expressed gene-gene interactions in high-throughput sequencing data
abstract
SUMMARY: The discovery of differential gene-gene correlations across phenotypical groups can help identify the activation/deactivation of critical biological processes underlying specific conditions. The presented R package, provided with a count and design matrix, extract networks of group-specific interactions that can be interactively explored through a shiny user-friendly interface. For each gene-gene link, differential statistical significance is provided through robust linear regression with an interaction term. AVAILABILITY AND IMPLEMENTATION: DEGGs is implemented in R and available on GitHub at https://github.com/elisabettasciacca/DEGGs. The package is also under submission on Bioconductor.
Elisabetta C. Sciacca, Salvatore Alaimo, Gianmarco Silluzio, Alfredo Ferro, Vito Latora, Costantino Pitzalis, Alfredo Pulvirenti, Myles J. Lewis
Bioinform.7
2022 Virus finding tools: current solutions and limitations
abstract
MOTIVATION: The study of the Human Virome remains challenging nowadays. Viral metagenomics, through high-throughput sequencing data, is the best choice for virus discovery. The metagenomics approach is culture-independent and sequence-independent, helping search for either known or novel viruses. Though it is estimated that more than 40% of the viruses found in metagenomics analysis are not recognizable, we decided to analyze several tools to identify and discover viruses in RNA-seq samples. RESULTS: We have analyzed eight Virus Tools for the identification of viruses in RNA-seq data. These tools were compared using a synthetic dataset of 30 viruses and a real one. Our analysis shows that no tool succeeds in recognizing all the viruses in the datasets. So we can conclude that each of these tools has pros and cons, and their choice depends on the application domain. AVAILABILITY: Synthetic data used through the review and raw results of their analysis can be found at https://zenodo.org/record/6426147. FASTQ files of real data can be found in GEO (https://www.ncbi.nlm.nih.gov/gds) or ENA (https://www.ebi.ac.uk/ena/browser/home). Raw results of their analysis can be downloaded from https://zenodo.org/record/6425917.
Grete Francesca Privitera, Salvatore Alaimo, Alfredo Ferro, Alfredo Pulvirenti
Briefings Bioinform.4
2021 RNAdetector: a free user-friendly stand-alone and cloud-based system for RNA-Seq data analysis
abstract
BACKGROUND: RNA-Seq is a well-established technology extensively used for transcriptome profiling, allowing the analysis of coding and non-coding RNA molecules. However, this technology produces a vast amount of data requiring sophisticated computational approaches for their analysis than other traditional technologies such as Real-Time PCR or microarrays, strongly discouraging non-expert users. For this reason, dozens of pipelines have been deployed for the analysis of RNA-Seq data. Although interesting, these present several limitations and their usage require a technical background, which may be uncommon in small research laboratories. Therefore, the application of these technologies in such contexts is still limited and causes a clear bottleneck in knowledge advancement. RESULTS: Motivated by these considerations, we have developed RNAdetector, a new free cross-platform and user-friendly RNA-Seq data analysis software that can be used locally or in cloud environments through an easy-to-use Graphical User Interface allowing the analysis of coding and non-coding RNAs from RNA-Seq datasets of any sequenced biological species. CONCLUSIONS: RNAdetector is a new software that fills an essential gap between the needs of biomedical and research labs to process RNA-Seq data and their common lack of technical background in performing such analysis, which usually relies on outsourcing such steps to third party bioinformatics facilities or using expensive commercial software.
Alessandro La Ferlita, Salvatore Alaimo, Sebastiano Di Bella, Emanuele Martorana, Georgios I. Laliotis, Francesco Bertoni, Luciano Cascione, Philip N. Tsichlis, Alfredo Ferro, Roberta Bosotti, Alfredo Pulvirenti
BMC Bioinform.11
2021 PHENSIM: Phenotype Simulator
abstract
Despite the unprecedented growth in our understanding of cell biology, it still remains challenging to connect it to experimental data obtained with cells and tissues' physiopathological status under precise circumstances. This knowledge gap often results in difficulties in designing validation experiments, which are usually labor-intensive, expensive to perform, and hard to interpret. Here we propose PHENSIM, a computational tool using a systems biology approach to simulate how cell phenotypes are affected by the activation/inhibition of one or multiple biomolecules, and it does so by exploiting signaling pathways. Our tool's applications include predicting the outcome of drug administration, knockdown experiments, gene transduction, and exposure to exosomal cargo. Importantly, PHENSIM enables the user to make inferences on well-defined cell lines and includes pathway maps from three different model organisms. To assess our approach's reliability, we built a benchmark from transcriptomics data gathered from NCBI GEO and performed four case studies on known biological experiments. Our results show high prediction accuracy, thus highlighting the capabilities of this methodology. PHENSIM standalone Java application is available at https://github.com/alaimos/phensim, along with all data and source codes for benchmarking. A web-based user interface is accessible at https://phensim.tech/.
Salvatore Alaimo, Rosaria Valentina Rapicavoli, Gioacchino P. Marceca, Alessandro La Ferlita, Oksana B. Serebrennikova, Philip N. Tsichlis, Bud Mishra, Alfredo Pulvirenti, Alfredo Ferro
PLoS Comput. Biol.8
2020 A benchmarking of pipelines for detecting ncRNAs from RNA-Seq data
abstract
Next-Generation Sequencing (NGS) is a high-throughput technology widely applied to genome sequencing and transcriptome profiling. RNA-Seq uses NGS to reveal RNA identities and quantities in a given sample. However, it produces a huge amount of raw data that need to be preprocessed with fast and effective computational methods. RNA-Seq can look at different populations of RNAs, including ncRNAs. Indeed, in the last few years, several ncRNAs pipelines have been developed for ncRNAs analysis from RNA-Seq experiments. In this paper, we analyze eight recent pipelines (iSmaRT, iSRAP, miARma-Seq, Oasis 2, SPORTS1.0, sRNAnalyzer, sRNApipe, sRNA workbench) which allows the analysis not only of single specific classes of ncRNAs but also of more than one ncRNA classes. Our systematic performance evaluation aims at guiding users to select the appropriate pipeline for processing each ncRNA class, focusing on three key points: (i) accuracy in ncRNAs identification, (ii) accuracy in read count estimation and (iii) deployment and ease of use.
Sebastiano Di Bella, Alessandro La Ferlita, Giovanni Carapezza, Salvatore Alaimo, Antonella Isacchi, Alfredo Ferro, Alfredo Pulvirenti, Roberta Bosotti
Briefings Bioinform.7
2019 TACITuS: transcriptomic data collector, integrator, and selector on big data platform
abstract
BACKGROUND: Several large public repositories of microarray datasets and RNA-seq data are available. Two prominent examples include ArrayExpress and NCBI GEO. Unfortunately, there is no easy way to import and manipulate data from such resources, because the data is stored in large files, requiring large bandwidth to download and special purpose data manipulation tools to extract subsets relevant for the specific analysis. RESULTS: TACITuS is a web-based system that supports rapid query access to high-throughput microarray and NGS repositories. The system is equipped with modules capable of managing large files, storing them in a cloud environment and extracting subsets of data in an easy and efficient way. The system also supports the ability to import data into Galaxy for further analysis. CONCLUSIONS: TACITuS automates most of the pre-processing needed to analyze high-throughput microarray and NGS data from large publicly-available repositories. The system implements several modules to manage large files in an easy and efficient way. Furthermore, it is capable deal with Galaxy environment allowing users to analyze data through a user-friendly interface.
Salvatore Alaimo, Antonio Di Maria, Dennis E. Shasha, Alfredo Ferro, Alfredo Pulvirenti
BMC Bioinform.5
2018 INBIA: a boosting methodology for proteomic network inference
abstract
BACKGROUND: The analysis of tissue-specific protein interaction networks and their functional enrichment in pathological and normal tissues provides insights on the etiology of diseases. The Pan-cancer proteomic project, in The Cancer Genome Atlas, collects protein expressions in human cancers and it is a reference resource for the functional study of cancers. However, established protocols to infer interaction networks from protein expressions are still missing. RESULTS: We have developed a methodology called Inference Network Based on iRefIndex Analysis (INBIA) to accurately correlate proteomic inferred relations to protein-protein interaction (PPI) networks. INBIA makes use of 14 network inference methods on protein expressions related to 16 cancer types. It uses as reference model the iRefIndex human PPI network. Predictions are validated through non-interacting and tissue specific PPI networks resources. The first, Negatome, takes into account likely non-interacting proteins by combining both structure properties and literature mining. The latter, TissueNet and GIANT, report experimentally verified PPIs in more than 50 human tissues. The reliability of the proposed methodology is assessed by comparing INBIA with PERA, a tool which infers protein interaction networks from Pathway Commons, by both functional and topological analysis. CONCLUSION: Results show that INBIA is a valuable approach to predict proteomic interactions in pathological conditions starting from the current knowledge of human protein interactions.
Davide S. Sardina, Giovanni Micale, Alfredo Ferro, Alfredo Pulvirenti, Rosalba Giugno
BMC Bioinform.4
2018 Fast analytical methods for finding significant labeled graph motifs
Giovanni Micale, Rosalba Giugno, Alfredo Ferro, Misael Mongiovì, Dennis E. Shasha, Alfredo Pulvirenti
Data Min. Knowl. Discov.6
2017 A novel computational method for inferring competing endogenous interactions
abstract
Posttranscriptional cross talk and communication between genes mediated by microRNA response element (MREs) yield large regulatory competing endogenous RNA (ceRNA) networks. Their inference may improve the understanding of pathologies and shed new light on biological mechanisms. A variety of RNA: messenger RNA, transcribed pseudogenes, noncoding RNA, circular RNA and proteins related to RNA-induced silencing complex complex interacting with RNA transfer and ribosomal RNA have been experimentally proved to be ceRNAs. We retrace the ceRNA hypothesis of posttranscriptional regulation from its original formulation [Salmena L, Poliseno L, Tay Y, et al. Cell 2011;146:353-8] to the most recent experimental and computational validations. We experimentally analyze the methods in literature [Li J-H, Liu S, Zhou H, et al. Nucleic Acids Res 2013;42:D92-7; Sumazin P, Yang X, Chiu H-S, et al. Cell 2011;147:370-81; Sarver AL, Subramanian S. Bioinformation 2012;8:731-3] comparing them with a general machine learning approach, called ceRNA predIction Algorithm, evaluating the performance in predicting novel MRE-based ceRNAs.
Davide S. Sardina, Salvatore Alaimo, Alfredo Ferro, Alfredo Pulvirenti, Rosalba Giugno
Briefings Bioinform.4
2016 APPAGATO: an APproximate PArallel and stochastic GrAph querying TOol for biological networks
abstract
MOTIVATION: Biological network querying is a problem requiring a considerable computational effort to be solved. Given a target and a query network, it aims to find occurrences of the query in the target by considering topological and node similarities (i.e. mismatches between nodes, edges, or node labels). Querying tools that deal with similarities are crucial in biological network analysis because they provide meaningful results also in case of noisy data. In addition, as the size of available networks increases steadily, existing algorithms and tools are becoming unsuitable. This is rising new challenges for the design of more efficient and accurate solutions. RESULTS: This paper presents APPAGATO, a stochastic and parallel algorithm to find approximate occurrences of a query network in biological networks. APPAGATO handles node, edge and node label mismatches. Thanks to its randomic and parallel nature, it applies to large networks and, compared with existing tools, it provides higher performance as well as statistically significant more accurate results. Tests have been performed on protein-protein interaction networks annotated with synthetic and real gene ontology terms. Case studies have been done by querying protein complexes among different species and tissues. AVAILABILITY AND IMPLEMENTATION: APPAGATO has been developed on top of CUDA-C ++ Toolkit 7.0 framework. The software is available online http://profs.sci.univr.it/∼bombieri/APPAGATO CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vincenzo Bonnici, Federico Busato, Giovanni Micale, Nicola Bombieri, Alfredo Pulvirenti, Rosalba Giugno
Bioinform.5
2016 KAOS: a new automated computational method for the identification of overexpressed genes
abstract
BACKGROUND: Kinase over-expression and activation as a consequence of gene amplification or gene fusion events is a well-known mechanism of tumorigenesis. The search for novel rearrangements of kinases or other druggable genes may contribute to understanding the biology of cancerogenesis, as well as lead to the identification of new candidate targets for drug discovery. However this requires the ability to query large datasets to identify rare events occurring in very small fractions (1-3 %) of different tumor subtypes. This task is different from what is normally done by conventional tools that are able to find genes differentially expressed between two experimental conditions. RESULTS: We propose a computational method aimed at the automatic identification of genes which are selectively over-expressed in a very small fraction of samples within a specific tissue. The method does not require a healthy counterpart or a reference sample for the analysis and can be therefore applied also to transcriptional data generated from cell lines. In our implementation the tool can use gene-expression data from microarray experiments, as well as data generated by RNASeq technologies. CONCLUSIONS: The method was implemented as a publicly available, user-friendly tool called KAOS (Kinase Automatic Outliers Search). The tool enables the automatic execution of iterative searches for the identification of extreme outliers and for the graphical visualization of the results. Filters can be applied to select the most significant outliers. The performance of the tool was evaluated using a synthetic dataset and compared to state-of-the-art tools. KAOS performs particularly well in detecting genes that are overexpressed in few samples or when an extreme outlier stands out on a high variable expression background. To validate the method on real case studies, we used publicly available tumor cell line microarray data, and we were able to identify genes which are known to be overexpressed in specific samples, as well as novel ones.
Angelo Nuzzo, Giovanni Carapezza, Sebastiano Di Bella, Alfredo Pulvirenti, Antonella Isacchi, Roberta Bosotti
BMC Bioinform.4
2013 Drug-target interaction prediction through domain-tuned network-based inference
abstract
MOTIVATION: The identification of drug-target interaction (DTI) represents a costly and time-consuming step in drug discovery and design. Computational methods capable of predicting reliable DTI play an important role in the field. Recently, recommendation methods relying on network-based inference (NBI) have been proposed. However, such approaches implement naive topology-based inference and do not take into account important features within the drug-target domain. RESULTS: In this article, we present a new NBI method, called domain tuned-hybrid (DT-Hybrid), which extends a well-established recommendation technique by domain-based knowledge including drug and target similarity. DT-Hybrid has been extensively tested using the last version of an experimentally validated DTI database obtained from DrugBank. Comparison with other recently proposed NBI methods clearly shows that DT-Hybrid is capable of predicting more reliable DTIs. AVAILABILITY: DT-Hybrid has been developed in R and it is available, along with all the results on the predictions, through an R package at the following URL: http://sites.google.com/site/ehybridalgo/.
Salvatore Alaimo, Alfredo Pulvirenti, Rosalba Giugno, Alfredo Ferro
Bioinform.2
2013 A subgraph isomorphism algorithm and its application to biochemical data
abstract
BACKGROUND: Graphs can represent biological networks at the molecular, protein, or species level. An important query is to find all matches of a pattern graph to a target graph. Accomplishing this is inherently difficult (NP-complete) and the efficiency of heuristic algorithms for the problem may depend upon the input graphs. The common aim of existing algorithms is to eliminate unsuccessful mappings as early as and as inexpensively as possible. RESULTS: We propose a new subgraph isomorphism algorithm which applies a search strategy to significantly reduce the search space without using any complex pruning rules or domain reduction procedures. We compare our method with the most recent and efficient subgraph isomorphism algorithms (VFlib, LAD, and our C++ implementation of FocusSearch which was originally distributed in Modula2) on synthetic, molecules, and interaction networks data. We show a significant reduction in the running time of our approach compared with these other excellent methods and show that our algorithm scales well as memory demands increase. CONCLUSIONS: Subgraph isomorphism algorithms are intensively used by biochemical tools. Our analysis gives a comprehensive comparison of different software approaches to subgraph isomorphism highlighting their weaknesses and strengths. This will help researchers make a rational choice among methods depending on their application. We also distribute an open-source package including our system and our own C++ implementation of FocusSearch together with all the used datasets (http://ferrolab.dmi.unict.it/ri.html). In future work, our findings may be extended to approximate subgraph isomorphism algorithms.
Vincenzo Bonnici, Rosalba Giugno, Alfredo Pulvirenti, Dennis E. Shasha, Alfredo Ferro
BMC Bioinform.3
2013 VIRGO: visualization of A-to-I RNA editing sites in genomic sequences
abstract
BACKGROUND: RNA Editing is a type of post-transcriptional modification that takes place in the eukaryotes. It alters the sequence of primary RNA transcripts by deleting, inserting or modifying residues. Several forms of RNA editing have been discovered including A-to-I, C-to-U, U-to-C and G-to-A. In recent years, the application of global approaches to the study of A-to-I editing, including high throughput sequencing, has led to important advances. However, in spite of enormous efforts, the real biological mechanism underlying this phenomenon remains unknown. DESCRIPTION: In this work, we present VIRGO (http://atlas.dmi.unict.it/virgo/), a web-based tool that maps Ato-G mismatches between genomic and EST sequences as candidate A-to-I editing sites. VIRGO is built on top of a knowledge-base integrating information of genes from UCSC, EST of NCBI, SNPs, DARNED, and Next Generations Sequencing data. The tool is equipped with a user-friendly interface allowing users to analyze genomic sequences in order to identify candidate A-to-I editing sites. CONCLUSIONS: VIRGO is a powerful tool allowing a systematic identification of putative A-to-I editing sites in genomic sequences. The integration of NGS data allows the computation of p-values and adjusted p-values to measure the mapped editing sites confidence. The whole knowledge base is available for download and will be continuously updated as new NGS data becomes available.
Rosario Distefano, Giovanni Nigita, Valentina Macca, Alessandro Laganà, Rosalba Giugno, Alfredo Pulvirenti, Alfredo Ferro
BMC Bioinform.6
2013 Bioinformatics in Italy: BITS 2012, the ninth annual meeting of the Italian Society of Bioinformatics
abstract
Abstract The BITS2012 meeting, held in Catania on May 2-4, 2012, brought together almost 100 Italian researchers working in the field of Bioinformatics, as well as students in the same or related disciplines. About 90 original research works were presented either as oral communication or as posters, representing a landscape of Italian current research in bioinformatics. This preface provides a brief overview of the meeting and introduces the manuscripts that were accepted for publication in this supplement, after a strict and careful peer-review by an International board of referees.
Carmela Gissi, Paolo Romano 0001, Alfredo Ferro, Rosalba Giugno, Alfredo Pulvirenti, Angelo M. Facchiano, Manuela Helmer-Citterich
BMC Bioinform.5
2013 Enhancing density-based clustering: Parameter reduction and outlier detection
Carmelo Cassisi, Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti
Inf. Syst.5
2012 miR-EdiTar: a database of predicted A-to-I edited miRNA target sites
abstract
MOTIVATION: A-to-I RNA editing is an important mechanism that consists of the conversion of specific adenosines into inosines in RNA molecules. Its dysregulation has been associated to several human diseases including cancer. Recent work has demonstrated a role for A-to-I editing in microRNA (miRNA)-mediated gene expression regulation. In fact, edited forms of mature miRNAs can target sets of genes that differ from the targets of their unedited forms. The specific deamination of mRNAs can generate novel binding sites in addition to potentially altering existing ones. RESULTS: This work presents miR-EdiTar, a database of predicted A-to-I edited miRNA binding sites. The database contains predicted miRNA binding sites that could be affected by A-to-I editing and sites that could become miRNA binding sites as a result of A-to-I editing. AVAILABILITY: miR-EdiTar is freely available online at http://microrna.osumc.edu/mireditar. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alessandro Laganà, Alessio Paone, Dario Veneziano, Luciano Cascione, Pierluigi Gasparini, Stefania Carasi, Francesco Russo 0004, Giovanni Nigita, Valentina Macca, Rosalba Giugno, Alfredo Pulvirenti, Dennis E. Shasha, Alfredo Ferro, Carlo Maria Croce
Bioinform.11
2011 DBStrata: a system for density-based clustering and outlier detection based on stratification
abstract
Clustering is a widely used unsupervised data mining technique. In density-based clustering, a cluster is defined as a connected dense component and grows in the direction set by the density. In this paper we present a software system called DBStrata that implements the density-based clustering architecture together with several extensions able to boost the clustering performances and to efficiently identify outliers.
Marco Aliotta, Andrea Cannata, Carmelo Cassisi, Rosalba Giugno, Placido Montalto, Alfredo Pulvirenti
SISAP6
2011 Obstacles constrained group mobility models in event-driven wireless networks with movable base stations
S. Cristaldi, Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti
Ad Hoc Networks5
2011 Editorial
abstract
NETTAB workshops are a series of international events on ‘Network Tools and Applications in Biology’ that are aimed at presenting and discussing emerging Information and Communication Technologies whose adoption in support of biology appear to be of particular interest. Since 2001, many different topics were discussed including: XML standardization for data integration (2001), multi agent systems (2002), scientific workflows (2005), Web Services (2006) and the Semantic Web (2007). The NETTAB 2009 workshop that was held at the Mathematics and Computer Science Department of the University of Catania, Italy, 10–13 June 2009, was focused on ‘Collaborative Bioinformatics Research and Development’ and on ‘Tools for RNA Analysis’. The workshop included many original contributions devoted to these innovative research domains, the best of which have been carefully peer reviewed and included in this Special Issue on ‘Collaborative Bioinformatics and RNA Analysis’. There is a clear switch from previous focus themes since these are aimed at different aspects of data integration, ranging from standardization, to automation of procedures. In this case, the accent is on human collaboration and ways to implement it. While many decades ago research was mainly carried out by researchers in a single institute or laboratory, it is nowadays common that researchers from different, far away laboratories carry out experiments in collaboration. However, such a collaboration does not usually involve sharing of research platforms. Research is done independently and results are then exchanged. The advent of high-speed communication networks, together with the development of software enabling actual sharing of data, is quickly leading to new ways of conducting experiments and making research in collaboration. Romano et al. present a survey of some of the tools available for collaborative research and development in the first paper of this Issue. We review the principles and some social networking applications are able to support scientific collaboration, discuss the wiki approach to collaborative document creation and review some biological wikis. Finally, we present collaborative platforms for software development and some examples of Learning Management Systems as tools for bioinformatics. Some good examples of how this approach can effectively be implemented are provided by the following two papers. Splendiani et al. present DC-THERA Directory, an information system designed to support collaborative knowledge management in DC-THERA. This is a European ‘Network of Excellence’ (NoE) involving a large interdisciplinary research community that is focused on the development of novel immunotherapies derived from research in dendritic cell immunobiology that is producing large data sets and a great wealth of knowledge. The Directory is based on Semantic Web technologies. The authors show how these technologies, along with properly defined metadata such as reference vocabularies and ontologies can support a better organization of the knowledge and constitute a model for an effective data management solution during and beyond a project lifecycle. The dramatic increase of data being generated by high-throughput technologies are opening a new universe of problems that will certainly provide great challenges in the coming years. The bottom line of such problems is represented by data and knowledge integration. Due to the vast amount of data, its curation cannot rely on individual researchers. Accordingly, community-driven tools are becoming very popular in the life sciences. Navas-Delgado et al. describe the collaborative improvement of metabolic pathways by means of a new tool with social networking capabilities. The NETTAB workshop was also focused on tools for RNA analysis. We continue the special issue with two papers focusing on microRNAs (miRNAs) and non-coding RNAs (ncRNAs). miRNAs are a class of small, non-coding regulatory RNAs that are crucial in post-transcriptional gene silencing. They may regulate gene expression by binding the 3′-UTR site of their targets. miRNAs may then play a key role in many biological processes, such as cell proliferation, cell death and oncogenesis. Since their discovery, computational approaches have been important in understanding the biology of miRNAs. Web-based-miRNA databases now provide thousands of miRNA sequences, annotations and putative target genes. Although much research has been done, target prediction is still a challenging task and due to the high false positive rate it is a major obstacle in biological validation of miRNA/mRNA interaction. In this Special Issue, Corrada et al., after surveying target prediction tools, propose a meta-system to improve overall prediction. Their tool is based on a knowledge base that is built on top of the most reliable prediction methods implementing an integration strategy for computing ranked miRNA-target lists together with functional annotations. The RNA of an organism is mostly composed of molecules that are not translated into proteins, so called ‘non-coding RNA’. Some ncRNAs, such as transfer RNA and ribosomal RNA, have well-known functions, whereas, for the remaining, their activity is still unknown or not validated. Work aimed at understanding ncRNA suggests that they have functions that are as yet unknown. Recalling that biological activity is largely due to the three dimensional (3D) structure of molecules, tools aimed at predicting the 3D structure of RNA are of an extreme relevance in this context. In this special issue, ModeRNA, a tool for the prediction of the tertiary structure of a RNA, is included. In their paper, Rother et al. guide the readers through the preparation of the input to the evaluation of the results, and to the prediction of tRNA molecules from Escherichia Coli, as an example of its practical utilization. We conclude this special issue with a review on metagenomics data analysis and associated issues by Ribeca and Valiente. It is well known that metagenomics may also allow the sequencing and analysis of genomes of organisms that cannot grow in laboratories either because their lives depend on other organisms or because they are extinct. In the former case, DNA is taken from the environment, while, in the letter, it is extracted from fossils. Metagenomics applications include the monitoring of bacterial composition of an individual's gastrointestinal fauna as well as the detection of viruses in the environment. Next-generation sequencing is used and its relative low cost makes it affordable for many laboratories. From the computational point of view, metagenomics analysis is conducted by assembling random short reads, aligning them with full genomes and performing taxonomic classification. The major issues include the enormous amounts of resulting data, which make the analysis not trivial, the sequencing error models and the assignment of reads to the correct species.
Paolo Romano 0001, Rosalba Giugno, Alfredo Pulvirenti
Briefings Bioinform.3
2011 Tools and collaborative environments for bioinformatics research
abstract
Advanced research requires intensive interaction among a multitude of actors, often possessing different expertise and usually working at a distance from each other. The field of collaborative research aims to establish suitable models and technologies to properly support these interactions. In this article, we first present the reasons for an interest of Bioinformatics in this context by also suggesting some research domains that could benefit from collaborative research. We then review the principles and some of the most relevant applications of social networking, with a special attention to networks supporting scientific collaboration, by also highlighting some critical issues, such as identification of users and standardization of formats. We then introduce some systems for collaborative document creation, including wiki systems and tools for ontology development, and review some of the most interesting biological wikis. We also review the principles of Collaborative Development Environments for software and show some examples in Bioinformatics. Finally, we present the principles and some examples of Learning Management Systems. In conclusion, we try to devise some of the goals to be achieved in the short term for the exploitation of these technologies.
Paolo Romano 0001, Rosalba Giugno, Alfredo Pulvirenti
Briefings Bioinform.3
2010 An Efficient Duplicate Record Detection Using q-Grams Array Inverted Index
Alfredo Ferro, Rosalba Giugno, Piera Laura Puglisi, Alfredo Pulvirenti
DaWak4
2010 MySQL Data Mining: Extending MySQL to Support Data Mining Primitives (Demo)
Alfredo Ferro, Rosalba Giugno, Piera Laura Puglisi, Alfredo Pulvirenti
KES (3)4
2010 SING: Subgraph search In Non-homogeneous Graphs
abstract
BACKGROUND: Finding the subgraphs of a graph database that are isomorphic to a given query graph has practical applications in several fields, from cheminformatics to image understanding. Since subgraph isomorphism is a computationally hard problem, indexing techniques have been intensively exploited to speed up the process. Such systems filter out those graphs which cannot contain the query, and apply a subgraph isomorphism algorithm to each residual candidate graph. The applicability of such systems is limited to databases of small graphs, because their filtering power degrades on large graphs. RESULTS: In this paper, SING (Subgraph search In Non-homogeneous Graphs), a novel indexing system able to cope with large graphs, is presented. The method uses the notion of feature, which can be a small subgraph, subtree or path. Each graph in the database is annotated with the set of all its features. The key point is to make use of feature locality information. This idea is used to both improve the filtering performance and speed up the subgraph isomorphism task. CONCLUSIONS: Extensive tests on chemical compounds, biological networks and synthetic graphs show that the proposed system outperforms the most popular systems in query time over databases of medium and large graphs. Other specific tests show that the proposed system is effective for single large graphs.
Raffaele Di Natale, Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Alfredo Pulvirenti, Dennis E. Shasha
BMC Bioinform.5
2009 BitCube: A Bottom-Up Cubing Engineering
Alfredo Ferro, Rosalba Giugno, Piera Laura Puglisi, Alfredo Pulvirenti
DaWaK4
2009 Distributed randomized algorithms for low-support data mining
abstract
Data mining in distributed systems has been facilitated by using high-support association rules. Less attention has been paid to distributed low-support/high-correlation data mining. This has proved useful in several fields such as computational biology, wireless networks, web mining, security and rare events analysis in industrial plants. In this paper we present distributed versions of efficient algorithms for low-support/high-correlation data mining such as Min-Hashing, K-Min-Hashing and Locality-Sensitive-Hashing. Experimental results on real data concerning scalability, speed-up and network traffic are reported.
Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Alfredo Pulvirenti
IPDPS4
2008 GraphFind: enhancing graph searching by low support data mining techniques
abstract
BACKGROUND: Biomedical and chemical databases are large and rapidly growing in size. Graphs naturally model such kinds of data. To fully exploit the wealth of information in these graph databases, a key role is played by systems that search for all exact or approximate occurrences of a query graph. To deal efficiently with graph searching, advanced methods for indexing, representation and matching of graphs have been proposed. RESULTS: This paper presents GraphFind. The system implements efficient graph searching algorithms together with advanced filtering techniques that allow approximate search. It allows users to select candidate subgraphs rather than entire graphs. It implements an effective data storage based also on low-support data mining. CONCLUSIONS: GraphFind is compared with Frowns, GraphGrep and gIndex. Experiments show that GraphFind outperforms the compared systems on a very large collection of small graphs. The proposed low-support mining technique which applies to any searching system also allows a significant index space reduction.
Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Alfredo Pulvirenti, Dmitry Skripin, Dennis E. Shasha
BMC Bioinform.4
2007 NetMatch: a Cytoscape plugin for searching biological networks
abstract
UNLABELLED: NetMatch is a Cytoscape plugin which allows searching biological networks for subcomponents matching a given query. Queries may be approximate in the sense that certain parts of the subgraph-query may be left unspecified. To make the query creation process easy, a drawing tool is provided. Cytoscape is a bioinformatics software platform for the visualization and analysis of biological networks. AVAILABILITY: The full package, a tutorial and associated examples are available at the following web sites: http://alpha.dmi.unict.it/~ctnyu/netmatch.html, http://baderlab.org/Software/NetMatch.
Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti, Dmitry Skripin, Gary D. Bader, Dennis E. Shasha
Bioinform.4
2007 Sequence similarity is more relevant than species specificity in probabilistic backtranslation
abstract
BACKGROUND: Backtranslation is the process of decoding a sequence of amino acids into the corresponding codons. All synthetic gene design systems include a backtranslation module. The degeneracy of the genetic code makes backtranslation potentially ambiguous since most amino acids are encoded by multiple codons. The common approach to overcome this difficulty is based on imitation of codon usage within the target species. RESULTS: This paper describes EasyBack, a new parameter-free, fully-automated software for backtranslation using Hidden Markov Models. EasyBack is not based on imitation of codon usage within the target species, but instead uses a sequence-similarity criterion. The model is trained with a set of proteins with known cDNA coding sequences, constructed from the input protein by querying the NCBI databases with BLAST. Unlike existing software, the proposed method allows the quality of prediction to be estimated. When tested on a group of proteins that show different degrees of sequence conservation, EasyBack outperforms other published methods in terms of precision. CONCLUSION: The prediction quality of a protein backtranslation methis markedly increased by replacing the criterion of most used codon in the same species with a Hidden Markov Model trained with a set of most similar sequences from all species. Moreover, the proposed method allows the quality of prediction to be estimated probabilistically.
Alfredo Ferro, Rosalba Giugno, Giuseppe Pigola, Alfredo Pulvirenti, Cinzia Di Pietro, Michele Purrello, Marco Ragusa
BMC Bioinform.4
2006 Distributed antipole clustering for efficient data search and management in Euclidean and metric spaces
abstract
In this paper a simple and efficient distributed version of the introduced antipole clustering algorithm for general metric spaces is proposed. This combines ideas from the M-tree, the multi-vantage point structure and the FQ-tree to create a new structure in the "bisector tree" class, called the antipole tree. Bisection is based on the proximity to an "antipole" pair of elements generated by a suitable linear randomized tournament. The final winners (A, B) of such a tournament are far enough apart to approximate the diameter of the splitting set. A simple linear algorithm computing antipoles in Euclidean spaces with exponentially small approximation ratio is proposed. The antipole tree clustering has been shown to be very effective in important applications such as range and k-nearest neighbor searching, mobile objects clustering in centralized wireless networks with movable base stations and multiple alignment of biological sequences. In many of such applications an efficient distributed clustering algorithm is needed. In the proposed distributed versions of antipole clustering the amount of data passed from one node to another is either constant or proportional to the number of nodes in the network. The distributed antipole tree is equipped with additional information in order to perform efficient range search and dynamic clusters management. This is achieved by adding to the randomized tournaments technique, methodologies taken from established systems such as BFR and BIRCH*. Experiments show the good performance of the proposed algorithms on both real and synthetic data
Alfredo Ferro, Rosalba Giugno, Misael Mongiovì, Giuseppe Pigola, Alfredo Pulvirenti
IPDPS5
2005 Antipole Tree Indexing to Support Range Search and K-Nearest Neighbor Search in Metric Spaces
abstract
Range and k-nearest neighbor searching are core problems in pattern recognition. Given a database S of objects in a metric space M and a query object q in M, in a range searching problem the goal is to find the objects of S within some threshold distance to g, whereas in a k-nearest neighbor searching problem, the k elements of S closest to q must be produced. These problems can obviously be solved with a linear number of distance calculations, by comparing the query object against every object in the database. However, the goal is to solve such problems much faster. We combine and extend ideas from the M-tree, the multivantage point structure, and the FQ-tree to create a new structure in the "bisector tree" class, called the Antipole tree. Bisection is based on the proximity to an "Antipole" pair of elements generated by a suitable linear randomized tournament. The final winners a, b of such a tournament is far enough apart to approximate the diameter of the splitting set. If dist(a, b) is larger than the chosen cluster diameter threshold, then the cluster is split. The proposed data structure is an indexing scheme suitable for (exact and approximate) best match searching on generic metric spaces. The Antipole tree outperforms by a factor of approximately two existing structures such as list of clusters, M-trees, and others and, in many cases, it achieves better clustering properties.
Domenico Cantone, Alfredo Ferro, Alfredo Pulvirenti, Diego Reforgiato Recupero, Dennis E. Shasha
IEEE Trans. Knowl. Data Eng.3
2004 Locally sensitive backtranslation based on multiple sequence alignment
abstract
Backtranslation is the process of decoding an amino acid sequence into a corresponding nucleic acid. Classical methods are based on the construction of a codon usage table by clustering and detection of the most probable codon used for each amino acid. We present a new method for backtranslation which is sensitive to the local position of the amino acid in the input sequence. The method makes use of multiple sequence alignment of the set of proteins under analysis. A local codon usage table stores for each amino acid X and for each position of X in the alignment the most used codon. We compared our method with EMBOSS using both ClustalW and AntiClustAl for multiple sequence alignment. Experiments showed that our method outperforms EMBOSS in terms of precision of backtranslation: the matching between the proteins obtained by our method and the original protein templates is clearly superior to that obtained by EMBOSS. This enforces the validity of a locally sensitive approach.
Rosalba Giugno, Alfredo Pulvirenti, Marco Ragusa, Loredana Facciola, Laura Patelmo, Valentina Di Pietro, Cinzia Di Pietro, Michele Purrello, Alfredo Ferro
CIBCB2
2003 GENIUS: a simple and easy way to access computational and data grids
A. Andronico, Roberto Barbera, Alberto Falzone, Peter Z. Kunszt, Giuseppe Lo Re, Alfredo Pulvirenti, A. Rodolico
Future Gener. Comput. Syst.6
2001 Best-Match Retrieval for Structured Images
abstract
Propose a methodology for fast best-match retrieval of structured images. A triangle inequality property for the tree-distance introduced by Oflazer (1997) is proven. This property is, in turn, applied to obtain a saturation algorithm of the trie used to store the database of the collection of pictures. The new approach can be considered as a substantial optimization of Oflazer's technique and can be applied to the retrieval of homogeneous hierarchically structured objects of any kind. The new technique inscribes itself in the number of distance-based search strategies and it is of interest for the indexing and maintenance of large collections of historical and pictorial data. We demonstrate the proposed approach on an example and report data about the speed-up that it introduces in query processing. Direct comparison with an MVP-trees algorithm is also presented.
Alfredo Ferro, Giovanni Gallo, Rosalba Giugno, Alfredo Pulvirenti
IEEE Trans. Pattern Anal. Mach. Intell.4