Graziano Pesole

dblp:26/2581 · DBLP profile ↗
← Back
61ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-3663-0859ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 56 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 kMetaShot: a fast and reliable taxonomy classifier for metagenome-assembled genomes
abstract
The advent of high-throughput sequencing (HTS) technologies unlocked the complexity of the microbial world through the development of metagenomics, which now provides an unprecedented and comprehensive overview of its taxonomic and functional contribution in a huge variety of macro- and micro-ecosystems. In particular, shotgun metagenomics allows the reconstruction of microbial genomes, through the assembly of reads into MAGs (metagenome-assembled genomes). In fact, MAGs represent an information-rich proxy for inferring the taxonomic composition and the functional contribution of microbiomes, even if the relevant analytical approaches are not trivial and still improvable. In this regard, tools like CAMITAX and GTDBtk have implemented complex approaches, relying on marker gene identification and sequence alignments, requiring a large processing time. With the aim of deploying an effective tool for fast and reliable MAG taxonomic classification, we present here kMetaShot, a taxonomy classifier based on k-mer/minimizer counting. We benchmarked kMetaShot against CAMITAX and GTDBtk by using both in silico and real mock communities and demonstrated how, while implementing a fast and concise algorithm, it outperforms the other tools in terms of classification accuracy. Additionally, kMetaShot is an easy-to-install and easy-to-use bioinformatic tool that is also suitable for researchers with few command-line skills. It is available and documented at https://github.com/gdefazio/kMetaShot.
Giuseppe Defazio, Marco Antonio Tangaro, Graziano Pesole, Bruno Fosso
Briefings Bioinform.3
2025 REDInet: a temporal convolutional network-based classifier for A-to-I RNA editing detection harnessing million known events
abstract
A-to-I ribonucleic acid (RNA) editing detection is still a challenging task. Current bioinformatics tools rely on empirical filters and whole genome sequencing or whole exome sequencing data to remove background noise, sequencing errors, and artifacts. Sometimes they make use of cumbersome and time-consuming computational procedures. Here, we present REDInet, a temporal convolutional network-based deep learning algorithm, to profile RNA editing in human RNA sequencing (RNAseq) data. It has been trained on REDIportal RNA editing sites, the largest collection of human A-to-I changes from >8000 RNAseq data of the genotype-tissue expression project. REDInet can classify editing events with high accuracy harnessing RNAseq nucleotide frequencies of 101-base windows without the need for coupled genomic data.
Adriano Fonzino, Pietro Luca Mazzacuva, Adam Handen, Domenico Alessandro Silvestris, Annette Arnold, Riccardo Pecori, Graziano Pesole, Ernesto Picardi
Briefings Bioinform.7
2021 Next generation sequencing of SARS-CoV-2 genomes: challenges, applications and opportunities
abstract
Various next generation sequencing (NGS) based strategies have been successfully used in the recent past for tracing origins and understanding the evolution of infectious agents, investigating the spread and transmission chains of outbreaks, as well as facilitating the development of effective and rapid molecular diagnostic tests and contributing to the hunt for treatments and vaccines. The ongoing COVID-19 pandemic poses one of the greatest global threats in modern history and has already caused severe social and economic costs. The development of efficient and rapid sequencing methods to reconstruct the genomic sequence of SARS-CoV-2, the etiological agent of COVID-19, has been fundamental for the design of diagnostic molecular tests and to devise effective measures and strategies to mitigate the diffusion of the pandemic. Diverse approaches and sequencing methods can, as testified by the number of available sequences, be applied to SARS-CoV-2 genomes. However, each technology and sequencing approach has its own advantages and limitations. In the current review, we will provide a brief, but hopefully comprehensive, account of currently available platforms and methodological approaches for the sequencing of SARS-CoV-2 genomes. We also present an outline of current repositories and databases that provide access to SARS-CoV-2 genomic data and associated metadata. Finally, we offer general advice and guidelines for the appropriate sharing and deposition of SARS-CoV-2 data and metadata, and suggest that more efficient and standardized integration of current and future SARS-CoV-2-related data would greatly facilitate the struggle against this new pathogen. We hope that our 'vademecum' for the production and handling of SARS-CoV-2-related sequencing data, will contribute to this objective.
Matteo Chiara, Anna Maria D'Erchia, Carmela Gissi, Caterina Manzari, Antonio Parisi, Nicoletta Resta, Federico Zambelli, Ernesto Picardi, Giulio Pavesi, David Stephen Horner, Graziano Pesole
Briefings Bioinform.11
2021 VINYL: Variant prIoritizatioN bY survivaL analysis
abstract
MOTIVATION: Clinical applications of genome re-sequencing technologies typically generate large amounts of data that need to be carefully annotated and interpreted to identify genetic variants potentially associated with pathological conditions. In this context, accurate and reproducible methods for the functional annotation and prioritization of genetic variants are of fundamental importance. RESULTS: In this article, we present VINYL, a flexible and fully automated system for the functional annotation and prioritization of genetic variants. Extensive analyses of both real and simulated datasets suggest that VINYL can identify clinically relevant genetic variants in a more accurate manner compared to equivalent state of the art methods, allowing a more rapid and effective prioritization of genetic variants in different experimental settings. As such we believe that VINYL can establish itself as a valuable tool to assist healthcare operators and researchers in clinical genomics investigations. AVAILABILITY AND IMPLEMENTATION: VINYL is available at http://beaconlab.it/VINYL and https://github.com/matteo14c/VINYL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matteo Chiara, Pietro Mandreoli, Marco Antonio Tangaro, Anna Maria D'Erchia, Sandro Sorrentino, Cinzia Forleo, David Stephen Horner, Federico Zambelli, Graziano Pesole
Bioinform.9
2021 CorGAT: a tool for the functional annotation of SARS-CoV-2 genomes
abstract
SUMMARY: While over 200 000 genomic sequences are currently available through dedicated repositories, ad hoc methods for the functional annotation of SARS-CoV-2 genomes do not harness all currently available resources for the annotation of functionally relevant genomic sites. Here, we present CorGAT, a novel tool for the functional annotation of SARS-CoV-2 genomic variants. By comparisons with other state of the art methods we demonstrate that, by providing a more comprehensive and rich annotation, our method can facilitate the identification of evolutionary patterns in the genome of SARS-CoV-2. AVAILABILITYAND IMPLEMENTATION: Galaxy. http://corgat.cloud.ba.infn.it/galaxy; software: https://github.com/matteo14c/CorGAT/tree/Revision_V1; docker: https://hub.docker.com/r/laniakeacloud/galaxy_corgat. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matteo Chiara, Federico Zambelli, Marco Antonio Tangaro, Pietro Mandreoli, David Stephen Horner, Graziano Pesole
Bioinform.6
2021 ITSoneWB: profiling global taxonomic diversity of eukaryotic communities on Galaxy
abstract
SUMMARY: ITSoneWB (ITSone WorkBench) is a Galaxy-based bioinformatic environment where comprehensive and high-quality reference data are connected with established pipelines and new tools in an automated and easy-to-use service targeted at global taxonomic analysis of eukaryotic communities based on Internal Transcribed Spacer 1 variants high-throughput sequencing. AVAILABILITY AND IMPLEMENTATION: ITSoneWB has been deployed on the INFN-Bari ReCaS cloud facility and is freely available on the web at http://itsonewb.cloud.ba.infn.it/galaxy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco Antonio Tangaro, Giuseppe Defazio, Bruno Fosso, Flavio Licciulli, Giorgio Grillo, Giacinto Donvito, Enrico Lavezzo, Giacomo Baruzzo, Graziano Pesole, Monica Santamaria
Bioinform.9
2021 Laniakea@ReCaS: exploring the potential of customisable Galaxy on-demand instances as a cloud-based service
abstract
BACKGROUND: Improving the availability and usability of data and analytical tools is a critical precondition for further advancing modern biological and biomedical research. For instance, one of the many ramifications of the COVID-19 global pandemic has been to make even more evident the importance of having bioinformatics tools and data readily actionable by researchers through convenient access points and supported by adequate IT infrastructures. One of the most successful efforts in improving the availability and usability of bioinformatics tools and data is represented by the Galaxy workflow manager and its thriving community. In 2020 we introduced Laniakea, a software platform conceived to streamline the configuration and deployment of "on-demand" Galaxy instances over the cloud. By facilitating the set-up and configuration of Galaxy web servers, Laniakea provides researchers with a powerful and highly customisable platform for executing complex bioinformatics analyses. The system can be accessed through a dedicated and user-friendly web interface that allows the Galaxy web server's initial configuration and deployment. RESULTS: "Laniakea@ReCaS", the first instance of a Laniakea-based service, is managed by ELIXIR-IT and was officially launched in February 2020, after about one year of development and testing that involved several users. Researchers can request access to Laniakea@ReCaS through an open-ended call for use-cases. Ten project proposals have been accepted since then, totalling 18 Galaxy on-demand virtual servers that employ ~ 100 CPUs, ~ 250 GB of RAM and ~ 5 TB of storage and serve several different communities and purposes. Herein, we present eight use cases demonstrating the versatility of the platform. CONCLUSIONS: During this first year of activity, the Laniakea-based service emerged as a flexible platform that facilitated the rapid development of bioinformatics tools, the efficient delivery of training activities, and the provision of public bioinformatics services in different settings, including food safety and clinical research. Laniakea@ReCaS provides a proof of concept of how enabling access to appropriate, reliable IT resources and ready-to-use bioinformatics tools can considerably streamline researchers' work.
Marco Antonio Tangaro, Pietro Mandreoli, Matteo Chiara, Giacinto Donvito, Marica Antonacci, Antonio Parisi, Angelica Bianco, Angelo Romano, Daniela Manila Bianchi, Davide Cangelosi, Paolo Uva, Ivan Molineris, Vladimir Nosi, Raffaele A. Calogero, Luca Alessandrì, Elena Pedrini, Marina Mordenti, Emanuele Bonetti, Luca Sangiorgi, Graziano Pesole, Federico Zambelli
BMC Bioinform.20
2020 Critical assessment of bioinformatics methods for the characterization of pathological repeat expansions with single-molecule sequencing data
abstract
A number of studies have reported the successful application of single-molecule sequencing technologies to the determination of the size and sequence of pathological expanded microsatellite repeats over the last 5 years. However, different custom bioinformatics pipelines were employed in each study, preventing meaningful comparisons and somewhat limiting the reproducibility of the results. In this review, we provide a brief summary of state-of-the-art methods for the characterization of expanded repeats alleles, along with a detailed comparison of bioinformatics tools for the determination of repeat length and sequence, using both real and simulated data. Our reanalysis of publicly available human genome sequencing data suggests a modest, but statistically significant, increase of the error rate of single-molecule sequencing technologies at genomic regions containing short tandem repeats. However, we observe that all the methods herein tested, irrespective of the strategy used for the analysis of the data (either based on the alignment or assembly of the reads), show high levels of sensitivity in both the detection of expanded tandem repeats and the estimation of the expansion size, suggesting that approaches based on single-molecule sequencing technologies are highly effective for the detection and quantification of tandem repeat expansions and contractions.
Matteo Chiara, Federico Zambelli, Ernesto Picardi, David Stephen Horner, Graziano Pesole
Briefings Bioinform.5
2020 ELIXIR-IT HPC@CINECA: high performance computing resources for the bioinformatics community
abstract
BACKGROUND: The advent of Next Generation Sequencing (NGS) technologies and the concomitant reduction in sequencing costs allows unprecedented high throughput profiling of biological systems in a cost-efficient manner. Modern biological experiments are increasingly becoming both data and computationally intensive and the wealth of publicly available biological data is introducing bioinformatics into the "Big Data" era. For these reasons, the effective application of High Performance Computing (HPC) architectures is becoming progressively more recognized also by bioinformaticians. Here we describe HPC resources provisioning pilot programs dedicated to bioinformaticians, run by the Italian Node of ELIXIR (ELIXIR-IT) in collaboration with CINECA, the main Italian supercomputing center. RESULTS: Starting from April 2016, CINECA and ELIXIR-IT launched the pilot Call "ELIXIR-IT HPC@CINECA", offering streamlined access to HPC resources for bioinformatics. Resources are made available either through web front-ends to dedicated workflows developed at CINECA or by providing direct access to the High Performance Computing systems through a standard command-line interface tailored for bioinformatics data analysis. This allows to offer to the biomedical research community a production scale environment, continuously updated with the latest available versions of publicly available reference datasets and bioinformatic tools. Currently, 63 research projects have gained access to the HPC@CINECA program, for a total handout of ~ 8 Millions of CPU/hours and, for data storage, ~ 100 TB of permanent and ~ 300 TB of temporary space. CONCLUSIONS: Three years after the beginning of the ELIXIR-IT HPC@CINECA program, we can appreciate its impact over the Italian bioinformatics community and draw some considerations. Several Italian researchers who applied to the program have gained access to one of the top-ranking public scientific supercomputing facilities in Europe. Those investigators had the opportunity to sensibly reduce computational turnaround times in their research projects and to process massive amounts of data, pursuing research approaches that would have been otherwise difficult or impossible to undertake. Moreover, by taking advantage of the wealth of documentation and training material provided by CINECA, participants had the opportunity to improve their skills in the usage of HPC systems and be better positioned to apply to similar EU programs of greater scale, such as PRACE. To illustrate the effective usage and impact of the resources awarded by the program - in different research applications - we report five successful use cases, which have already published their findings in peer-reviewed journals.
Tiziana Castrignanò, Silvia Gioiosa, Tiziano Flati, Mirko Cestari, Ernesto Picardi, Matteo Chiara, Maddalena Fratelli, Stefano Amente, Marco Cirilli, Marco Antonio Tangaro, Giovanni Chillemi, Graziano Pesole, Federico Zambelli
BMC Bioinform.12
2020 HPC-REDItools: a novel HPC-aware tool for improved large scale RNA-editing analysis
abstract
BACKGROUND: RNA editing is a widespread co-/post-transcriptional mechanism that alters primary RNA sequences through the modification of specific nucleotides and it can increase both the transcriptome and proteome diversity. The automatic detection of RNA-editing from RNA-seq data is computational intensive and limited to small data sets, thus preventing a reliable genome-wide characterisation of such process. RESULTS: In this work we introduce HPC-REDItools, an upgraded tool for accurate RNA-editing events discovery from large dataset repositories. AVAILABILITY: https://github.com/BioinfoUNIBA/REDItools2 . CONCLUSIONS: HPC-REDItools is dramatically faster than the previous version, REDItools, enabling big-data analysis by means of a MPI-based implementation and scaling almost linearly with the number of available cores.
Tiziano Flati, Silvia Gioiosa, Nicola Spallanzani, Ilario Tagliaferri, Maria Angela Diroma, Graziano Pesole, Giovanni Chillemi, Ernesto Picardi, Tiziana Castrignanò
BMC Bioinform.6
2019 Elucidating the editome: bioinformatics approaches for RNA editing detection
abstract
RNA editing is a widespread co/posttranscriptional mechanism affecting primary RNAs by specific nucleotide modifications, which plays relevant roles in molecular processes including regulation of gene expression and/or the processing of noncoding RNAs. In recent years, the detection of editing sites has been improved through the availability of high-throughput RNA sequencing (RNA-Seq) technologies. Accurate bioinformatics pipelines are essential for the analysis of next-generation sequencing (NGS) data to ensure the correct identification of edited sites. Several pipelines, using various read mappers and variant callers with a wide range of adjustable parameters, are available for the detection of RNA editing events. In this review, we discuss some of the most recent and popular tools and provide guidelines for RNA-Seq data generation and analysis for the detection of RNA editing in massive transcriptome data. Using simulated and real data sets, we provide an overview of their behavior, emphasizing the fact that the RNA editing detection in NGS data sets remains a challenging task.
Maria Angela Diroma, Loredana Ciaccia, Graziano Pesole, Ernesto Picardi
Briefings Bioinform.3
2017 Unbiased Taxonomic Annotation of Metagenomic Samples
Bruno Fosso, Graziano Pesole, Francesc Rosselló, Gabriel Valiente
ISBRA2
2017 MetaShot: an accurate workflow for taxon classification of host-associated microbiome from shotgun metagenomic data
abstract
SUMMARY: Shotgun metagenomics by high-throughput sequencing may allow deep and accurate characterization of host-associated total microbiomes, including bacteria, viruses, protists and fungi. However, the analysis of such sequencing data is still extremely challenging in terms of both overall accuracy and computational efficiency, and current methodologies show substantial variability in misclassification rate and resolution at lower taxonomic ranks or are limited to specific life domains (e.g. only bacteria). We present here MetaShot, a workflow for assessing the total microbiome composition from host-associated shotgun sequence data, and show its overall optimal accuracy performance by analyzing both simulated and real datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/bfosso/MetaShot. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bruno Fosso, Monica Santamaria, Mattia D'Antonio, Domenica Lovero, Giacomo Corrado, Enrico Vizza, N. Pássaro, Anna Rosa Garbuglia, M. R. Capobianchi, Marco Crescenzi, Gabriel Valiente, Graziano Pesole
Bioinform.12
2015 MSA-PAD: DNA multiple sequence alignment framework based on PFAM accessed domain information
abstract
Abstract Summary: Here we present the MSA-PAD application, a DNA multiple sequence alignment framework that uses PFAM protein domain information to align DNA sequences encoding either single or multiple protein domains. MSA-PAD has two alignment options: gene and genome mode. Availability and Implementation: MSA-PAD is available as a web application (https://recasgateway.ba.infn.it/) and as two Taverna workflows corresponding to two alignment modes (Gene mode: http://www.myexperiment.org/workflows/4549.html; Genome Mode: http://www.myexperiment.org/workflows/4551.html). Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Bachir Balech, Saverio Vicario, Giacinto Donvito, Alfonso Monaco, Pasquale Notarangelo, Graziano Pesole
Bioinform.6
2015 BioMaS: a modular pipeline for Bioinformatic analysis of Metagenomic AmpliconS
abstract
BACKGROUND: Substantial advances in microbiology, molecular evolution and biodiversity have been carried out in recent years thanks to Metagenomics, which allows to unveil the composition and functions of mixed microbial communities in any environmental niche. If the investigation is aimed only at the microbiome taxonomic structure, a target-based metagenomic approach, here also referred as Meta-barcoding, is generally applied. This approach commonly involves the selective amplification of a species-specific genetic marker (DNA meta-barcode) in the whole taxonomic range of interest and the exploration of its taxon-related variants through High-Throughput Sequencing (HTS) technologies. The accessibility to proper computational systems for the large-scale bioinformatic analysis of HTS data represents, currently, one of the major challenges in advanced Meta-barcoding projects. RESULTS: BioMaS (Bioinformatic analysis of Metagenomic AmpliconS) is a new bioinformatic pipeline designed to support biomolecular researchers involved in taxonomic studies of environmental microbial communities by a completely automated workflow, comprehensive of all the fundamental steps, from raw sequence data upload and cleaning to final taxonomic identification, that are absolutely required in an appropriately designed Meta-barcoding HTS-based experiment. In its current version, BioMaS allows the analysis of both bacterial and fungal environments starting directly from the raw sequencing data from either Roche 454 or Illumina HTS platforms, following two alternative paths, respectively. BioMaS is implemented into a public web service available at https://recasgateway.ba.infn.it/ and is also available in Galaxy at http://galaxy.cloud.ba.infn.it:8080 (only for Illumina data). CONCLUSION: BioMaS is a friendly pipeline for Meta-barcoding HTS data analysis specifically designed for users without particular computing skills. A comparative benchmark, carried out by using a simulated dataset suitably designed to broadly represent the currently known bacterial and fungal world, showed that BioMaS outperforms QIIME and MOTHUR in terms of extent and accuracy of deep taxonomic sequence assignments.
Bruno Fosso, Monica Santamaria, Marinella Marzano, Daniel Alonso-Alemany, Gabriel Valiente, Giacinto Donvito, Alfonso Monaco, Pasquale Notarangelo, Graziano Pesole
BMC Bioinform.9
2014 MToolBox: a highly automated pipeline for heteroplasmy annotation and prioritization analysis of human mitochondrial variants in high-throughput sequencing
abstract
MOTIVATION: The increasing availability of mitochondria-targeted and off-target sequencing data in whole-exome and whole-genome sequencing studies (WXS and WGS) has risen the demand of effective pipelines to accurately measure heteroplasmy and to easily recognize the most functionally important mitochondrial variants among a huge number of candidates. To this purpose, we developed MToolBox, a highly automated pipeline to reconstruct and analyze human mitochondrial DNA from high-throughput sequencing data. RESULTS: MToolBox implements an effective computational strategy for mitochondrial genomes assembling and haplogroup assignment also including a prioritization analysis of detected variants. MToolBox provides a Variant Call Format file featuring, for the first time, allele-specific heteroplasmy and annotation files with prioritized variants. MToolBox was tested on simulated samples and applied on 1000 Genomes WXS datasets. AVAILABILITY AND IMPLEMENTATION: MToolBox package is available at https://sourceforge.net/projects/mtoolbox/.
Claudia Calabrese, Domenico Simone, Maria Angela Diroma, Mariangela Santorsola, Cristiano Guttà, Giuseppe Gasparre, Ernesto Picardi, Graziano Pesole, Marcella Attimonelli
Bioinform.8
2014 EasyCluster2: an improved tool for clustering and assembling long transcriptome reads
abstract
BACKGROUND: Expressed sequences (e.g. ESTs) are a strong source of evidence to improve gene structures and predict reliable alternative splicing events. When a genome assembly is available, ESTs are suitable to generate gene-oriented clusters through the well-established EasyCluster software. Nowadays, EST-like sequences can be massively produced using Next Generation Sequencing (NGS) technologies. In order to handle genome-scale transcriptome data, we present here EasyCluster2, a reimplementation of EasyCluster able to speed up the creation of gene-oriented clusters and facilitate downstream analyses as the assembly of full-length transcripts and the detection of splicing isoforms. RESULTS: EasyCluster2 has been developed to facilitate the genome-based clustering of EST-like sequences generated through the NGS 454 technology. Reads mapped onto the reference genome can be uploaded using the standard GFF3 file format. Alignment parsing is initially performed to produce a first collection of pseudo-clusters by grouping reads according to the overlap of their genomic coordinates on the same strand. EasyCluster2 then refines read grouping by including in each cluster only reads sharing at least one splice site and optionally performs a Smith-Waterman alignment in the region surrounding splice sites in order to correct for potential alignment errors. In addition, EasyCluster2 can include unspliced reads, which generally account for >50% of 454 datasets, and collapses overlapping clusters. Finally, EasyCluster2 can assemble full-length transcripts using a Directed-Acyclic-Graph-based strategy, simplifying the identification of alternative splicing isoforms, thanks also to the implementation of the widespread AStalavista methodology. Accuracy and performances have been tested on real as well as simulated datasets. CONCLUSIONS: EasyCluster2 represents a unique tool to cluster and assemble transcriptome reads produced with 454 technology, as well as ESTs and full-length transcripts. The clustering procedure is enhanced with the employment of genome annotations and unspliced reads. Overall, EasyCluster2 is able to perform an effective detection of splicing isoforms, since it can refine exon-exon junctions and explore alternative splicing without known reference transcripts. Results in GFF3 format can be browsed in the UCSC Genome Browser. Therefore, EasyCluster2 is a powerful tool to generate reliable clusters for gene expression studies, facilitating the analysis also to researchers not skilled in bioinformatics.
Vitoantonio Bevilacqua, Nicola Pietroleonardo, Ely Ignazio Giannino, Fabio Stroppa, Domenico Simone, Graziano Pesole, Ernesto Picardi
BMC Bioinform.6
2013 Clustering and Assembling Large Transcriptome Datasets by EasyCluster2
Vitoantonio Bevilacqua, Nicola Pietroleonardo, Ely Ignazio Giannino, Fabio Stroppa, Graziano Pesole, Ernesto Picardi
ICIC (3)5
2013 Motif discovery and transcription factor binding sites before and after the next-generation sequencing era
abstract
Motif discovery has been one of the most widely studied problems in bioinformatics ever since genomic and protein sequences have been available. In particular, its application to the de novo prediction of putative over-represented transcription factor binding sites in nucleotide sequences has been, and still is, one of the most challenging flavors of the problem. Recently, novel experimental techniques like chromatin immunoprecipitation (ChIP) have been introduced, permitting the genome-wide identification of protein-DNA interactions. ChIP, applied to transcription factors and coupled with genome tiling arrays (ChIP on Chip) or next-generation sequencing technologies (ChIP-Seq) has opened new avenues in research, as well as posed new challenges to bioinformaticians developing algorithms and methods for motif discovery.
Federico Zambelli, Graziano Pesole, Giulio Pavesi
Briefings Bioinform.2
2013 REDItools: high-throughput RNA editing detection made easy
abstract
SUMMARY: The reliable detection of RNA editing sites from massive sequencing data remains challenging and, although several methodologies have been proposed, no computational tools have been released to date. Here, we introduce REDItools a suite of python scripts to perform high-throughput investigation of RNA editing using next-generation sequencing data. AVAILABILITY AND IMPLEMENTATION: REDItools are in python programming language and freely available at http://code.google.com/p/reditools/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ernesto Picardi, Graziano Pesole
Bioinform.2
2013 NGS-Trex: Next Generation Sequencing Transcriptome profile explorer
abstract
BACKGROUND: Next-Generation Sequencing (NGS) technology has exceptionally increased the ability to sequence DNA in a massively parallel and cost-effective manner. Nevertheless, NGS data analysis requires bioinformatics skills and computational resources well beyond the possibilities of many "wet biology" laboratories. Moreover, most of projects only require few sequencing cycles and standard tools or workflows to carry out suitable analyses for the identification and annotation of genes, transcripts and splice variants found in the biological samples under investigation. These projects can take benefits from the availability of easy to use systems to automatically analyse sequences and to mine data without the preventive need of strong bioinformatics background and hardware infrastructure. RESULTS: To address this issue we developed an automatic system targeted to the analysis of NGS data obtained from large-scale transcriptome studies. This system, we named NGS-Trex (NGS Transcriptome profile explorer) is available through a simple web interface http://www.ngs-trex.org and allows the user to upload raw sequences and easily obtain an accurate characterization of the transcriptome profile after the setting of few parameters required to tune the analysis procedure. The system is also able to assess differential expression at both gene and transcript level (i.e. splicing isoforms) by comparing the expression profile of different samples.By using simple query forms the user can obtain list of genes, transcripts, splice sites ranked and filtered according to several criteria. Data can be viewed as tables, text files or through a simple genome browser which helps the visual inspection of the data. CONCLUSIONS: NGS-Trex is a simple tool for RNA-Seq data analysis mainly targeted to "wet biology" researchers with limited bioinformatics skills. It offers simple data mining tools to explore transcriptome profiles of samples investigated taking advantage of NGS technologies.
Ilenia Boria, Lara Boatti, Graziano Pesole, Flavio Mignone
BMC Bioinform.3
2013 WEP: a high-performance analysis pipeline for whole-exome data
abstract
BACKGROUND: The advent of massively parallel sequencing technologies (Next Generation Sequencing, NGS) profoundly modified the landscape of human genetics.In particular, Whole Exome Sequencing (WES) is the NGS branch that focuses on the exonic regions of the eukaryotic genomes; exomes are ideal to help us understanding high-penetrance allelic variation and its relationship to phenotype. A complete WES analysis involves several steps which need to be suitably designed and arranged into an efficient pipeline.Managing a NGS analysis pipeline and its huge amount of produced data requires non trivial IT skills and computational power. RESULTS: Our web resource WEP (Whole-Exome sequencing Pipeline web tool) performs a complete WES pipeline and provides easy access through interface to intermediate and final results. The WEP pipeline is composed of several steps:1) verification of input integrity and quality checks, read trimming and filtering; 2) gapped alignment; 3) BAM conversion, sorting and indexing; 4) duplicates removal; 5) alignment optimization around insertion/deletion (indel) positions; 6) recalibration of quality scores; 7) single nucleotide and deletion/insertion polymorphism (SNP and DIP) variant calling; 8) variant annotation; 9) result storage into custom databases to allow cross-linking and intersections, statistics and much more. In order to overcome the challenge of managing large amount of data and maximize the biological information extracted from them, our tool restricts the number of final results filtering data by customizable thresholds, facilitating the identification of functionally significant variants. Default threshold values are also provided at the analysis computation completion, tuned with the most common literature work published in recent years. CONCLUSIONS: Through our tool a user can perform the whole analysis without knowing the underlying hardware and software architecture, dealing with both paired and single end data. The interface provides an easy and intuitive access for data submission and a user-friendly web interface for annotated variant visualization.Non-IT mastered users can access through WEP to the most updated and tested WES algorithms, tuned to maximize the quality of called variants while minimizing artifacts and false positives.The web tool is available at the following web address: http://www.caspur.it/wep.
Mattia D'Antonio, Paolo D'Onorio De Meo, Daniele Paoletti, Berardino Elmi, Matteo Pallocca, Nico Sanna, Ernesto Picardi, Graziano Pesole, Tiziana Castrignanò
BMC Bioinform.8
2012 Reference databases for taxonomic assignment in metagenomics
abstract
Metagenomics is providing an unprecedented access to the environmental microbial diversity. The amplicon-based metagenomics approach involves the PCR-targeted sequencing of a genetic locus fitting different features. Namely, it must be ubiquitous in the taxonomic range of interest, variable enough to discriminate between different species but flanked by highly conserved sequences, and of suitable size to be sequenced through next-generation platforms. The internal transcribed spacers 1 and 2 (ITS1 and ITS2) of the ribosomal DNA operon and one or more hyper-variable regions of 16S ribosomal RNA gene are typically used to identify fungal and bacterial species, respectively. In this context, reliable reference databases and taxonomies are crucial to assign amplicon sequence reads to the correct phylogenetic ranks. Several resources provide consistent phylogenetic classification of publicly available 16S ribosomal DNA sequences, whereas the state of ribosomal internal transcribed spacers reference databases is notably less advanced. In this review, we aim to give an overview of existing reference resources for both types of markers, highlighting strengths and possible shortcomings of their use for metagenomics purposes. Moreover, we present a new database, ITSoneDB, of well annotated and phylogenetically classified ITS1 sequences to be used as a reference collection in metagenomic studies of environmental fungal communities. ITSoneDB is available for download and browsing at http://itsonedb.ba.itb.cnr.it/.
Monica Santamaria, Bruno Fosso, Arianna Consiglio, Giorgio De Caro, Giorgio Grillo, Flavio Licciulli, Sabino Liuni, Marinella Marzano, Daniel Alonso-Alemany, Gabriel Valiente, Graziano Pesole
Briefings Bioinform.11
2012 Editorial
abstract
These are exciting times for molecular biology. The advent of next-generation sequencing technologies has opened up new avenues in several fields. In particular, the large-scale exploration of metagenomics data gives us the extraordinary and unprecedented possibility of unravelling the taxonomic complexity of all living organisms as well as gaining a comprehensive overview of the products of billion years of evolution and selection in different environments and conditions. New genes and functions may be discovered to foster a large variety of biotechnological processes and applications. The comprehensive analysis of the genetic material extracted from environmental samples can provide an accurate and effective inventory of microbial species and their functional activities without the need for culture and at much lower cost than with previous sequencing technologies. Indeed, the explosion of metagenomics data in a diverse variety of projects poses a tremendous challenge to fill the gap between the data generation and their interpretation. Most of the bioinformatics tools in use now were conceived for the analysis of genomic data, and they are not appropriate for the analysis of metagenomics data because of the high complexity of microbial communities and the use of high-throughput sequencing technologies. Therefore, there is a strong need for better bioinformatics analysis pipelines in metagenomics. We have invited some of the most influential researchers in the field of metagenomics analysis to help us delineate the current state-of-the-art as well as to identify opportunities and challenges for the further development of the field. The identification and classification of metagenomic sequences is addressed by Alice McHardy et al. (Taxonomic Binning of Metagenome Samples Generated by Next-Generation Sequencing Technologies), by John Wooley et al. (Ultrafast Clustering Algorithms for Metagenomic Sequence Analysis), and by Sharmila Mande et al. (Classification of Metagenomic Sequences: Methods and Challenges). Reference databases for the taxonomic analysis of metagenomic sequences are dealt with by Graziano Pesole et al. (Reference databases for amplicon-based metagenomics analysis). Functional analysis of metagenomic sequences is addressed by Duccio Cavalieri et al. (Bioinformatic Approaches for Pathway Reconstruction from Metagenomics Data) and by Todd Taylor et al. (Functional Assignment of Metagenomic Data: Challenges and Applications). Opportunities and challenges for further development are discussed by Frank Glöckner et al. (Current Opportunities and Challenges in Metagenome Analysis: a Bioinformatic Perspective) and by Sarah Hunter et al. (Metagenomic Analysis: the Challenge of the Data Bonanza). Last, but not the least, biomedical applications of metagenomic analysis are further discussed by Catherine Ngom-Bru and Caroline Barretto (Gut Microbiota: Methodological Aspects to Describe Taxonomy and Functionality) and by James Brown (Data Mining the Human Gut Microbiota for Therapeutic Targets), and the human microbiome is further studied by Elhanan Borenstein (Computational Systems Biology and In Silico Modeling of the Human Microbiome).
Gabriel Valiente, Graziano Pesole
Briefings Bioinform.2
2012 PIntron: a fast method for detecting the gene structure due to alternative splicing via maximal pairings of a pattern and a text
abstract
BACKGROUND: A challenging issue in designing computational methods for predicting the gene structure into exons and introns from a cluster of transcript (EST, mRNA) sequences, is guaranteeing accuracy as well as efficiency in time and space, when large clusters of more than 20,000 ESTs and genes longer than 1 Mb are processed. Traditionally, the problem has been faced by combining different tools, not specifically designed for this task. RESULTS: We propose a fast method based on ad hoc procedures for solving the problem. Our method combines two ideas: a novel algorithm of proved small time complexity for computing spliced alignments of a transcript against a genome, and an efficient algorithm that exploits the inherent redundancy of information in a cluster of transcripts to select, among all possible factorizations of EST sequences, those allowing to infer splice site junctions that are largely confirmed by the input data. The EST alignment procedure is based on the construction of maximal embeddings, that are sequences obtained from paths of a graph structure, called embedding graph, whose vertices are the maximal pairings of a genomic sequence T and an EST P. The procedure runs in time linear in the length of P and T and in the size of the output.The method was implemented into the PIntron package. PIntron requires as input a genomic sequence or region and a set of EST and/or mRNA sequences. Besides the prediction of the full-length transcript isoforms potentially expressed by the gene, the PIntron package includes a module for the CDS annotation of the predicted transcripts. CONCLUSIONS: PIntron, the software tool implementing our methodology, is available at http://www.algolab.eu/PIntron under GNU AGPL. PIntron has been shown to outperform state-of-the-art methods, and to quickly process some critical genes. At the same time, PIntron exhibits high accuracy (sensitivity and specificity) when benchmarked with ENCODE annotations.
Yuri Pirola, Raffaella Rizzi, Ernesto Picardi, Graziano Pesole, Gianluca Della Vedova, Paola Bonizzoni
BMC Bioinform.4
2011 ExpEdit: a webserver to explore human RNA editing in RNA-Seq experiments
abstract
UNLABELLED: ExpEdit is a web application for assessing RNA editing in human at known or user-specified sites supported by transcript data obtained by RNA-Seq experiments. Mapping data (in SAM/BAM format) or directly sequence reads [in FASTQ/short read archive (SRA) format] can be provided as input to carry out a comparative analysis against a large collection of known editing sites collected in DARNED database as well as other user-provided potentially edited positions. Results are shown as dynamic tables containing University of California, Santa Cruz (UCSC) links for a quick examination of the genomic context. AVAILABILITY: ExpEdit is freely available on the web at http://www.caspur.it/ExpEdit/.
Ernesto Picardi, Mattia D'Antonio, Danilo Carrabino, Tiziana Castrignanò, Graziano Pesole
Bioinform.5
2010 New Tools for Expression Alternative Splicing Validation
Vitoantonio Bevilacqua, Ernesto Picardi, Graziano Pesole, Daniele Ranieri, Vincenzo Stola, Vito Renò
ICIC (3)3
2010 Bioinformatics approaches for genomics and post genomics applications of next-generation sequencing
abstract
Technical advances such as the development of molecular cloning, Sanger sequencing, PCR and oligonucleotide microarrays are key to our current capacity to sequence, annotate and study complete organismal genomes. Recent years have seen the development of a variety of so-called 'next-generation' sequencing platforms, with several others anticipated to become available shortly. The previously unimaginable scale and economy of these methods, coupled with their enthusiastic uptake by the scientific community and the potential for further improvements in accuracy and read length, suggest that these technologies are destined to make a huge and ongoing impact upon genomic and post-genomic biology. However, like the analysis of microarray data and the assembly and annotation of complete genome sequences from conventional sequencing data, the management and analysis of next-generation sequencing data requires (and indeed has already driven) the development of informatics tools able to assemble, map, and interpret huge quantities of relatively or extremely short nucleotide sequence data. Here we provide a broad overview of bioinformatics approaches that have been introduced for several genomics and functional genomics applications of next-generation sequencing.
David Stephen Horner, Giulio Pavesi, Tiziana Castrignanò, Paolo D'Onorio De Meo, Sabino Liuni, Michael Sammeth, Ernesto Picardi, Graziano Pesole
Briefings Bioinform.8
2009 Statistical assessment of discriminative features for protein-coding and non coding cross-species conserved sequence elements
abstract
Abstract Background The identification of protein coding elements in sets of mammalian conserved elements is one of the major challenges in the current molecular biology research. Many features have been proposed for automatically distinguishing coding and non coding conserved sequences, making so necessary a systematic statistical assessment of their differences. A comprehensive study should be composed of an association study, i.e. a comparison of the distributions of the features in the two classes, and a prediction study in which the prediction accuracies of classifiers trained on single and groups of features are analyzed, conditionally to the compared species and to the sequence lengths. Results In this paper we compared distributions of a set of comparative and non comparative features and evaluated the prediction accuracy of classifiers trained for discriminating sequence elements conserved among human, mouse and rat species. The association study showed that the analyzed features are statistically different in the two classes. In order to study the influence of the sequence lengths on the feature performances, a predictive study was performed on different data sets composed of coding and non coding alignments in equal number and equally long with an ascending average length. We found that the most discriminant feature was a comparative measure indicating the proportion of synonymous nucleotide substitutions per synonymous sites. Moreover, linear discriminant classifiers trained by using comparative features in general outperformed classifiers based on intrinsic ones. Finally, the prediction accuracy of classifiers trained on comparative features increased significantly by adding intrinsic features to the set of input variables, independently on sequence length (Kolmogorov-Smirnov P-value ≤ 0.05). Conclusion We observed distinct and consistent patterns for individual and combined use of comparative and intrinsic classifiers, both with respect to different lengths of sequences/alignments and with respect to error rates in the classification of coding and non-coding elements. In particular, we noted that comparative features tend to be more accurate in the classification of coding sequences – this is likely related to the fact that such features capture deviations from strictly neutral evolution expected as a consequence of the characteristics of the genetic code.
Teresa Maria Creanza, David Stephen Horner, Annarita D'Addabbo, Rosalia Maglietta, Flavio Mignone, Nicola Ancona, Graziano Pesole
BMC Bioinform.7
2009 EasyCluster: a fast and efficient gene-oriented clustering tool for large-scale transcriptome data
abstract
BACKGROUND: ESTs and full-length cDNAs represent an invaluable source of evidence for inferring reliable gene structures and discovering potential alternative splicing events. In newly sequenced genomes, these tasks may not be practicable owing to the lack of appropriate training sets. However, when expression data are available, they can be used to build EST clusters related to specific genomic transcribed loci. Common strategies recently employed to this end are based on sequence similarity between transcripts and can lead, in specific conditions, to inconsistent and erroneous clustering. In order to improve the cluster building and facilitate all downstream annotation analyses, we developed a simple genome-based methodology to generate gene-oriented clusters of ESTs when a genomic sequence and a pool of related expressed sequences are provided. Our procedure has been implemented in the software EasyCluster and takes into account the spliced nature of ESTs after an ad hoc genomic mapping. METHODS: EasyCluster uses the well-known GMAP program in order to perform a very quick EST-to-genome mapping in addition to the detection of reliable splice sites. Given a genomic sequence and a pool of ESTs/FL-cDNAs, EasyCluster starts building genomic and EST local databases and runs GMAP. Subsequently, it parses results creating an initial collection of pseudo-clusters by grouping ESTs according to the overlap of their genomic coordinates on the same strand. In the final step, EasyCluster refines the clustering by again running GMAP on each pseudo-cluster and groups together ESTs sharing at least one splice site. RESULTS: The higher accuracy of EasyCluster with respect to other clustering tools has been verified by means of a manually cured benchmark of human EST clusters. Additional datasets including the Unigene cluster Hs.122986 and ESTs related to the human HOXA gene family have also been used to demonstrate the better clustering capability of EasyCluster over current genome-based web service tools such as ASmodeler and BIPASS. EasyCluster has also been used to provide a first compilation of gene-oriented clusters in the Ricinus communis oilseed plant for which no Unigene clusters are yet available, as well as an evaluation of the alternative splicing in this plant species.
Ernesto Picardi, Flavio Mignone, Graziano Pesole
BMC Bioinform.3
2009 Accurate discrimination of conserved coding and non-coding regions through multiple indicators of evolutionary dynamics
abstract
BACKGROUND: The conservation of sequences between related genomes has long been recognised as an indication of functional significance and recognition of sequence homology is one of the principal approaches used in the annotation of newly sequenced genomes. In the context of recent findings that the number non-coding transcripts in higher organisms is likely to be much higher than previously imagined, discrimination between conserved coding and non-coding sequences is a topic of considerable interest. Additionally, it should be considered desirable to discriminate between coding and non-coding conserved sequences without recourse to the use of sequence similarity searches of protein databases as such approaches exclude the identification of novel conserved proteins without characterized homologs and may be influenced by the presence in databases of sequences which are erroneously annotated as coding. RESULTS: Here we present a machine learning-based approach for the discrimination of conserved coding sequences. Our method calculates various statistics related to the evolutionary dynamics of two aligned sequences. These features are considered by a Support Vector Machine which designates the alignment coding or non-coding with an associated probability score. CONCLUSION: We show that our approach is both sensitive and accurate with respect to comparable methods and illustrate several situations in which it may be applied, including the identification of conserved coding regions in genome sequences and the discrimination of coding from non-coding cDNA sequences.
Matteo Ré, Graziano Pesole, David Stephen Horner
BMC Bioinform.2
2008 GenePC and ASPIC Integrate Gene Predictions with Expressed Sequence Alignments To Predict Alternative Transcripts
Tyler S. Alioto, Roderic Guigó, Ernesto Picardi, Graziano Pesole
APBC4
2008 HT-RLS: High-Throughput Web Tool for Analysis of DNA Microarray Data Using RLS classifiers
abstract
Gene expression from DNA microarray data offers biologists and pathologists the possibility to deal with the problem of disease (e. g. cancer) diagnosis and prognosis from a quantitative point of view. Microarray data provide a snapshot of the molecular status of a sample of cells in a given tissue, returning the expression levels of thousands of genes simultaneously. Several mathematical methods from learning theory, such as Regularized Least Squares (RLS) classifiers or Support Vector Machines (SVM), have been extensively adopted to classify gene expression data. These methods can be useful to answer some relevant questions such as 1) what is the right amount of data to build an accurate classifier? 2) How many and which genes are correlated with a specific pathology? The computational analysis to statistically estimate the accuracy of the chosen models is particularly time consuming, burning several days of CPU time and without high-throughput or high- performance tools becomes practically unfeasible to obtain results in a reasonable time for biomedical community. We have implemented an independent, flexible and scalable platform, for a high-throughput large-scale microarray gene expression data analysis and classification, based on R tool for statistical computing. It integrates databases and computational intensive algorithms, based on RLS classifiers and a powerful web client for data training and graphical visualization of predicted results. Our platform provides statistically significant answers to the study of the gene expression by means of microarray data and supplying useful information to relevant questions in the diagnosis and prognosis of diseases in a reasonable time. The web resource is available free of charge for academic and non-profit institutions.
Paolo D'Onorio De Meo, Danilo Carrabino, Mattia D'Antonio, Annarita D'Addabbo, Sabino Liuni, Flavio Mignone, Graziano Pesole, Nicola Ancona
CCGRID7
2008 Correlated substitution analysis and the prediction of amino acid structural contacts
abstract
It has long been suspected that analysis of correlated amino acid substitutions should uncover pairs or clusters of sites that are spatially proximal in mature protein structures. Accordingly, methods based on different mathematical principles such as information theory, correlation coefficients and maximum likelihood have been developed to identify co-evolving amino acids from multiple sequence alignments. Sets of pairs of sites whose behaviour is identified by these methods as correlated are often significantly enriched in pairs of spatially proximal residues. However, relatively high levels of false-positive predictions typically render such methods, in isolation, of little use in the ab initio prediction of protein structure. Misleading signal (or problems with the estimation of significance levels) can be caused by phylogenetic correlations between homologous sequences and from correlation due to factors other than spatial proximity (for example, correlation of sites which are not spatially close but which are involved in common functional properties of the protein). In recent years, several workers have suggested that information from correlated substitutions should be combined with other sources of information (secondary structure, solvent accessibility, evolutionary rates) in an attempt to reduce the proportion of false-positive predictions. We review methods for the detection of correlated amino acid substitutions, compare their relative performance in contact prediction and predict future directions in the field.
David Stephen Horner, Walter Pirovano, Graziano Pesole
Briefings Bioinform.3
2008 ASPicDB: A database resource for alternative splicing analysis
abstract
MOTIVATION: Alternative splicing has recently emerged as a key mechanism responsible for the expansion of transcriptome and proteome complexity in human and other organisms. Although several online resources devoted to alternative splicing analysis are available they may suffer from limitations related both to the computational methodologies adopted and to the extent of the annotations they provide that prevent the full exploitation of the available data. Furthermore, current resources provide limited query and download facilities. RESULTS: ASPicDB is a database designed to provide access to reliable annotations of the alternative splicing pattern of human genes and to the functional annotation of predicted splicing isoforms. Splice-site detection and full-length transcript modeling have been carried out by a genome-wide application of the ASPic algorithm, based on the multiple alignments of gene-related transcripts (typically a Unigene cluster) to the genomic sequence, a strategy that greatly improves prediction accuracy compared to methods based on independent and progressive alignments. Enhanced query and download facilities for annotations and sequences allow users to select and extract specific sets of data related to genes, transcripts and introns fulfilling a combination of user-defined criteria. Several tabular and graphical views of the results are presented, providing a comprehensive assessment of the functional implication of alternative splicing in the gene set under investigation. ASPicDB, which is regularly updated on a monthly basis, also includes information on tissue-specific splicing patterns of normal and cancer cells, based on available EST sequences and their library source annotation. AVAILABILITY: www.caspur.it/ASPicDB
Tiziana Castrignanò, Mattia D'Antonio, Anna Anselmo, Danilo Carrabino, A. D'Onorio De Meo, Anna Maria D'Erchia, Flavio Licciulli, Marina Mangiulli, Flavio Mignone, Giulio Pavesi, Ernesto Picardi, Alberto Riva, Raffaella Rizzi, Paola Bonizzoni, Graziano Pesole
Bioinform.15
2008 Bioinformatics in Italy: BITS2007, the fourth annual meeting of the Italian Society of Bioinformatics
Graziano Pesole, Manuela Helmer-Citterich
BMC Bioinform.1
2007 Selection of relevant genes in cancer diagnosis based on their prediction accuracy
Rosalia Maglietta, Annarita D'Addabbo, Ada Piepoli, Francesco Perri, Sabino Liuni, Graziano Pesole, Nicola Ancona
Artif. Intell. Medicine6
2007 Statistical assessment of functional categories of genes deregulated in pathological conditions by using microarray data
abstract
MOTIVATION: A major challenge in current biomedical research is the identification of cellular processes deregulated in a given pathology through the analysis of gene expression profiles. To this end, predefined lists of genes, coding specific functions, are compared with a list of genes ordered according to their values of differential expression measured by suitable univariate statistics. RESULTS: We propose a statistically well-founded method for measuring the relevance of predefined lists of genes and for assessing their statistical significance starting from their raw expression levels as recorded on the microarray. We use prediction accuracy as a measure of relevance of the list. The rationale is that a functional category, coded through a list of genes, is perturbed in a given pathology if it is possible to correctly predict the occurrence of the disease in new subjects on the basis of the expression levels of the genes belonging to the list only. The accuracy is estimated with multiple random validation strategy and its statistical significance is assessed against a couple of null hypothesis, by using two independent permutation tests. The utility of the proposed methodology is illustrated by analyzing the relevance of Gene Ontology terms belonging to biological process category in colon and prostate cancer, by using three different microarray data sets and by comparing it with current approaches. AVAILABILITY: Source code for the algorithms is available from author upon request. SUPPLEMENTARY INFORMATION: Colon cancer data set and a complete description of experimental results are available at: ftp://bioftp:[email protected]/supp-info.htm.
Rosalia Maglietta, Ada Piepoli, Domenico Catalano, Piepoli Licciulli, Massimo Carella, Sabino Liuni, Graziano Pesole, Francesco Perri, Nicola Ancona
Bioinform.7
2007 Bioinformatics in Italy: BITS2006, the third annual meeting of the Italian Society of Bioinformatics
Rita Casadio, Manuela Helmer-Citterich, Graziano Pesole
BMC Bioinform.3
2007 WeederH: an algorithm for finding conserved regulatory motifs and regions in homologous sequences
abstract
BACKGROUND: This work addresses the problem of detecting conserved transcription factor binding sites and in general regulatory regions through the analysis of sequences from homologous genes, an approach that is becoming more and more widely used given the ever increasing amount of genomic data available. RESULTS: We present an algorithm that identifies conserved transcription factor binding sites in a given sequence by comparing it to one or more homologs, adapting a framework we previously introduced for the discovery of sites in sequences from co-regulated genes. Differently from the most commonly used methods, the approach we present does not need or compute an alignment of the sequences investigated, nor resorts to descriptors of the binding specificity of known transcription factors. The main novel idea we introduce is a relative measure of conservation, assuming that true functional elements should present a higher level of conservation with respect to the rest of the sequence surrounding them. We present tests where we applied the algorithm to the identification of conserved annotated sites in homologous promoters, as well as in distal regions like enhancers. CONCLUSION: Results of the tests show how the algorithm can provide fast and reliable predictions of conserved transcription factor binding sites regulating the transcription of a gene, with better performances than other available methods for the same task. We also show examples on how the algorithm can be successfully employed when promoter annotations of the genes investigated are missing, or when regulatory sites and regions are located far away from the genes.
Giulio Pavesi, Federico Zambelli, Graziano Pesole
BMC Bioinform.3
2007 p53FamTaG: a database resource of human p53, p63 and p73 direct target genes combining in silico prediction and microarray data
abstract
BACKGROUND: The p53 gene family consists of the three genes p53, p63 and p73, which have polyhedral non-overlapping functions in pivotal cellular processes such as DNA synthesis and repair, growth arrest, apoptosis, genome stability, angiogenesis, development and differentiation. These genes encode sequence-specific nuclear transcription factors that recognise the same responsive element (RE) in their target genes. Their inactivation or aberrant expression may determine tumour progression or developmental disease. The discovery of several protein isoforms with antagonistic roles, which are produced by the expression of different promoters and alternative splicing, widened the complexity of the scenario of the transcriptional network of the p53 family members. Therefore, the identification of the genes transactivated by p53 family members is crucial to understand the specific role for each gene in cell cycle regulation. We have combined a genome-wide computational search of p53 family REs and microarray analysis to identify new direct target genes. The huge amount of biological data produced has generated a critical need for bioinformatic tools able to manage and integrate such data and facilitate their retrieval and analysis. DESCRIPTION: We have developed the p53FamTaG database (p53 FAMily TArget Genes), a modular relational database, which contains p53 family direct target genes selected in the human genome searching for the presence of the REs and the expression profile of these target genes obtained by microarray experiments. p53FamTaG database also contains annotations of publicly available databases and links to other experimental data. The genome-wide computational search of the REs was performed using PatSearch, a pattern-matching program implemented in the DNAfan tool. These data were integrated with the microarray results we produced from the overexpression of different isoforms of p53, p63 and p73 stably transfected in isogenic cell lines, allowing the comparative study of the transcriptional activity of all the proteins in the same cellular background.p53FamTaG database is available free at http://www2.ba.itb.cnr.it/p53FamTaG/ CONCLUSION: p53FamTaG represents a unique integrated resource of human direct p53 family target genes that is extensively annotated and provides the users with an efficient query/retrieval system which displays the results of our microarray experiments and allows the export of RE sequences. The database was developed for supporting and integrating high-throughput in silico and experimental analyses and represents an important reference source of knowledge for research groups involved in the field of oncogenesis, apoptosis and cell cycle regulation.
Elisabetta Sbisà, Domenico Catalano, Giorgio Grillo, Flavio Licciulli, Antonio Turi, Sabino Liuni, Graziano Pesole, Anna De Grassi, Mariano Francesco Caratozzolo, Anna Maria D'Erchia, Beatriz Navarro, Apollonia Tullo, Cecilia Saccone, Andreas Gisel
BMC Bioinform.7
2007 A high performance grid-web service framework for the identification of 'conserved sequence tags'
Paolo D'Onorio De Meo, Danilo Carrabino, Nico Sanna, Tiziana Castrignanò, Giorgio Grillo, Flavio Licciulli, Sabino Liuni, Matteo Ré, Flavio Mignone, Graziano Pesole
Future Gener. Comput. Syst.10
2006 Classification error as a measure of gene relevance in cancer diagnosis
abstract
One of the main problems in cancer diagnosis by using DNA microarray data is selecting genes relevant for the pathology by analyzing their expression profiles in tissues in two different phenotypical conditions. The question we pose is the following: how do we measure the relevance of a single gene in a given pathology? A gene is relevant for a particular disease if it is possible to correctly predict the occurrence of the pathology in new patients on the basis of expression level of this gene only. In other words, a gene is informative for the disease if its expression levels are useful for training a classifier able to generalize, that is, able to correctly predict the status of new patients. In this paper we present a selection bias free, statistically well founded method for finding relevant genes on the basis of their classification ability. We applied the method on a colon cancer data set and produced a list of relevant genes, ranked on the basis of their prediction accuracy. We found, out of more than 6500 available genes, 54 overexpressed in normal tissue and 77 overexpressed in tumor tissue having prediction accuracy greater than 7 0 % with p-value p ≤ 0.05.
Rosalia Maglietta, Annarita D'Addabbo, Ada Piepoli, Francesco Perri, Sabino Liuni, Graziano Pesole, Nicola Ancona
IJCNN6
2006 GenoMiner: a tool for genome-wide search of coding and non-coding conserved sequence tags
abstract
GenoMiner is a software tool that searches for regions of similarity between user-submitted genome or transcript sequences and user-specified whole genome assemblies. The program then identifies conserved sequence tags (CSTs) in these homologous regions and provides a prediction of their coding or non-coding nature. The analysis is carried out through three steps: (1) definition of sequence regions homologous to the query sequence in the selected target genomes by a fast BLAT alignment; (2) identification of CSTs by a more sensitive BLAST-like alignment between the query and the homologous regions in the target genomes and (3) assessment of the coding or non-coding nature of detected CSTs through the computation of a suitable coding potential score. GenoMiner allows the user to search the query sequence against a number of vertebrate genome assemblies in a single run providing a user-friendly graphical output.
Tiziana Castrignanò, Paolo D'Onorio De Meo, Giorgio Grillo, Sabino Liuni, Flavio Mignone, Ivano Giuseppe Talamo, Graziano Pesole
Bioinform.7
2006 On the statistical assessment of classifiers using DNA microarray data
abstract
BACKGROUND: In this paper we present a method for the statistical assessment of cancer predictors which make use of gene expression profiles. The methodology is applied to a new data set of microarray gene expression data collected in Casa Sollievo della Sofferenza Hospital, Foggia--Italy. The data set is made up of normal (22) and tumor (25) specimens extracted from 25 patients affected by colon cancer. We propose to give answers to some questions which are relevant for the automatic diagnosis of cancer such as: Is the size of the available data set sufficient to build accurate classifiers? What is the statistical significance of the associated error rates? In what ways can accuracy be considered dependant on the adopted classification scheme? How many genes are correlated with the pathology and how many are sufficient for an accurate colon cancer classification? The method we propose answers these questions whilst avoiding the potential pitfalls hidden in the analysis and interpretation of microarray data. RESULTS: We estimate the generalization error, evaluated through the Leave-K-Out Cross Validation error, for three different classification schemes by varying the number of training examples and the number of the genes used. The statistical significance of the error rate is measured by using a permutation test. We provide a statistical analysis in terms of the frequencies of the genes involved in the classification. Using the whole set of genes, we found that the Weighted Voting Algorithm (WVA) classifier learns the distinction between normal and tumor specimens with 25 training examples, providing e = 21% (p = 0.045) as an error rate. This remains constant even when the number of examples increases. Moreover, Regularized Least Squares (RLS) and Support Vector Machines (SVM) classifiers can learn with only 15 training examples, with an error rate of e = 19% (p = 0.035) and e = 18% (p = 0.037) respectively. Moreover, the error rate decreases as the training set size increases, reaching its best performances with 35 training examples. In this case, RLS and SVM have error rates of e = 14% (p = 0.027) and e = 11% (p = 0.019). Concerning the number of genes, we found about 6000 genes (p < 0.05) correlated with the pathology, resulting from the signal-to-noise statistic. Moreover the performances of RLS and SVM classifiers do not change when 74% of genes is used. They progressively reduce up to e = 16% (p < 0.05) when only 2 genes are employed. The biological relevance of a set of genes determined by our statistical analysis and the major roles they play in colorectal tumorigenesis is discussed. CONCLUSIONS: The method proposed provides statistically significant answers to precise questions relevant for the diagnosis and prognosis of cancer. We found that, with as few as 15 examples, it is possible to train statistically significant classifiers for colon cancer diagnosis. As for the definition of the number of genes sufficient for a reliable classification of colon cancer, our results suggest that it depends on the accuracy required.
Nicola Ancona, Rosalia Maglietta, Ada Piepoli, Annarita D'Addabbo, R. Cotugno, M. Savino, Sabino Liuni, Massimo Carella, Graziano Pesole, Francesco Perri
BMC Bioinform.9
2005 Regularized Least Squares Cancer Classifiers from DNA microarray data
abstract
BACKGROUND: The advent of the technology of DNA microarrays constitutes an epochal change in the classification and discovery of different types of cancer because the information provided by DNA microarrays allows an approach to the problem of cancer analysis from a quantitative rather than qualitative point of view. Cancer classification requires well founded mathematical methods which are able to predict the status of new specimens with high significance levels starting from a limited number of data. In this paper we assess the performances of Regularized Least Squares (RLS) classifiers, originally proposed in regularization theory, by comparing them with Support Vector Machines (SVM), the state-of-the-art supervised learning technique for cancer classification by DNA microarray data. The performances of both approaches have been also investigated with respect to the number of selected genes and different gene selection strategies. RESULTS: We show that RLS classifiers have performances comparable to those of SVM classifiers as the Leave-One-Out (LOO) error evaluated on three different data sets shows. The main advantage of RLS machines is that for solving a classification problem they use a linear system of order equal to either the number of features or the number of training examples. Moreover, RLS machines allow to get an exact measure of the LOO error with just one training. CONCLUSION: RLS classifiers are a valuable alternative to SVM classifiers for the problem of cancer classification by gene expression data, due to their simplicity and low computational complexity. Moreover, RLS classifiers show generalization ability comparable to the ones of SVM classifiers also in the case the classification of new specimens involves very few gene expression levels.
Nicola Ancona, Rosalia Maglietta, Annarita D'Addabbo, Sabino Liuni, Graziano Pesole
BMC Bioinform.5
2005 ASPIC: a novel method to predict the exon-intron structure of a gene that is optimally compatible to a set of transcript sequences
abstract
BACKGROUND: Currently available methods to predict splice sites are mainly based on the independent and progressive alignment of transcript data (mostly ESTs) to the genomic sequence. Apart from often being computationally expensive, this approach is vulnerable to several problems--hence the need to develop novel strategies. RESULTS: We propose a method, based on a novel multiple genome-EST alignment algorithm, for the detection of splice sites. To avoid limitations of splice sites prediction (mainly, over-predictions) due to independent single EST alignments to the genomic sequence our approach performs a multiple alignment of transcript data to the genomic sequence based on the combined analysis of all available data. We recast the problem of predicting constitutive and alternative splicing as an optimization problem, where the optimal multiple transcript alignment minimizes the number of exons and hence of splice site observations. We have implemented a splice site predictor based on this algorithm in the software tool ASPIC (Alternative Splicing PredICtion). It is distinguished from other methods based on BLAST-like tools by the incorporation of entirely new ad hoc procedures for accurate and computationally efficient transcript alignment and adopts dynamic programming for the refinement of intron boundaries. ASPIC also provides the minimal set of non-mergeable transcript isoforms compatible with the detected splicing events. The ASPIC web resource is dynamically interconnected with the Ensembl and Unigene databases and also implements an upload facility. CONCLUSION: Extensive bench marking shows that ASPIC outperforms other existing methods in the detection of novel splicing isoforms and in the minimization of over-predictions. ASPIC also requires a lower computation time for processing a single gene and an EST cluster. The ASPIC web resource is available at http://aspic.algo.disco.unimib.it/aspic-devel/.
Paola Bonizzoni, Raffaella Rizzi, Graziano Pesole
BMC Bioinform.3
2005 Overview of BITS2005, the Second Annual Meeting of the Italian Bioinformatics Society
abstract
Abstract The BITS2005 Conference brought together about 200 Italian scientists working in the field of Bioinformatics, students in Biology, Computer Science and Bioinformatics on March 17–19 2005, in Milan. This Editorial provides a brief overview of the Conference topics and introduces the peer-reviewed manuscripts accepted for publication in this Supplement.
Manuela Helmer-Citterich, Rita Casadio, Alessandro Guffanti, Giancarlo Mauri, Luciano Milanesi, Graziano Pesole, Giorgio Valle, Cecilia Saccone
BMC Bioinform.6
2004 In silico representation and discovery of transcription factor binding sites
abstract
Understanding the complex mechanisms governing basic biological processes requires the characterisation of regulatory motifs modulating gene expression at transcriptional and post-transcriptional level. In particular, extent, chronology and cell-specificity of transcription are modulated by the interaction of transcription factors with their corresponding binding sites, mostly located near (or sometimes quite far away from) the transcription start site of the gene. The constantly growing amount of genomic data, complemented by other sources of information such as expression data derived from microarray experiments, has opened new opportunities to researchers in this field. Many different methods have been proposed for the identification of transcription factor binding sites in the regulatory regions of co-expressed genes: unfortunately this is a very challenging problem both from the computational and the biological viewpoint. This paper provides a survey of existing methods proposed for the problem, focusing both on the ideas underlying them and their availability to the scientific community.
Giulio Pavesi, Giancarlo Mauri, Graziano Pesole
Briefings Bioinform.3
2004 DNAfan: a software tool for automated extraction and analysis of user-defined sequence regions
abstract
SUMMARY: DNAfan (DNA Feature ANalyzer) is a tool combining sequence-filtering and pattern searching. DNAfan automatically extracts user-defined sets of sequence fragments from large sequence sets. Fragments are defined by annotated gene feature keys and co- or non-occurring patterns within the feature or close to it. A gene feature parser and a pattern-based filter tool localizes and extracts the specific subset of sequences. The selected sequence data can subsequently be retrieved for analyses or further processed with DNAfan to find the occurrence of specific patterns or structural motifs. DNAfan is a powerful tool for pattern analysis. Its filter features restricts the pattern search to a well-defined set of sequences, allowing drastic reduction in false positive hits. AVAILABILITY: http://bighost.ba.itb.cnr.it:8080/Framework.
Andreas Gisel, Maria Panetta, Giorgio Grillo, Flavio Licciulli, Sabino Liuni, Cecilia Saccone, Graziano Pesole
Bioinform.7
2004 WebVar: a resource for the rapid estimation of relative site variability from multiple sequence alignments
abstract
UNLABELLED: WebVar is an online resource that provides estimates of relative site variability from multiple alignments of homologous protein or nucleic acid sequences. WebVar provides a variety of graphic and textual representations of estimates, designed to assist in phylogenetic analysis. AVAILABILITY: The WebVar server is located at http://www.pesolelab.it/Tools/WebVar.html
Flavio Mignone, David Stephen Horner, Graziano Pesole
Bioinform.3
2004 An Algorithm for Finding Conserved Secondary Structure Motifs in Unaligned RNA Sequences
Giulio Pavesi, Giancarlo Mauri, Graziano Pesole
J. Comput. Sci. Technol.3
2003 Predicting Conserved Hairpin Motifs in Unaligned RNA Sequences
abstract
Several experiments and observations have revealed the fact that small local distinct structural features in RNA molecules are correlated with their biological function, for example in post-transcriptional regulation of gene expression. Thus, finding similar structural features in a set of RNA sequences known to play the same biological function could provide substantial information concerning which parts of the sequences are responsible for the function itself. The main difficulty lies in the fact that in nearly all the cases the structure of the molecules is unknown, has to be somehow predicted, and that sequences with little or no similarity can fold into similar structures. The algorithm we present searches for regions of the sequences that, according to base pairing rules, can fold into similar structures, where the degree of similarity can be defined by the user. Any information concerning sequence similarity in the motifs can be used either as a search constraint, or a posteriori, by post-processing the output. The search for the regions sharing structural similarity is implemented with the affix tree, a novel text-indexing structure that significantly accelerates the search for patterns having a symmetric layout, like those forming RNA hairpins. Tests based on experimentally known structures have shown that the algorithm is able to identify functional motifs in the secondary structure of non coding RNA, such as Iron Responsive Elements (IRE) in the untranslated regions of ferritin mRNA, and the domain IV stem-loop structure in SRP RNA.
Giulio Pavesi, Giancarlo Mauri, Graziano Pesole
ICTAI3
2003 A Method to Detect Gene Structure and Alternative Splice Sites by Agreeing ESTs to a Genomic Sequence
Paola Bonizzoni, Graziano Pesole, Raffaella Rizzi
WABI2
2003 The estimation of relative site variability among aligned homologous protein sequences
abstract
MOTIVATION: Maximum likelihood-based methods to estimate site by site substitution rate variability in aligned homologous protein sequences rely on the formulation of a phylogenetic tree and generally assume that the patterns of relative variability follow a pre-determined distribution. We present a phylogenetic tree-independent method to estimate the relative variability of individual sites within large datasets of homologous protein sequences. It is based upon two simple assumptions. Firstly that substitutions observed between two closely related sequences are likely, in general, to occur at the most variable sites. Secondly that non-conservative amino acid substitutions tend to occur at more variable sites. Our methodology makes no assumptions regarding the underlying pattern of relative variability between sites. RESULTS: We have compared, using data simulated under a non-gamma distributed model, the performance of this approach to that of a maximum likelihood method that assumes gamma distributed rates. At low mean rates of evolution our method inferred site by site relative substitution rates more accurately than the maximum likelihood approach in the absence of prior assumptions about the relationships between sequences. Our method does not directly account for the effects of mutational saturation, However, we have incorporated an 'ad-hoc' modification that allows the accurate estimation of relative site variability in fast evolving and saturated datasets.
David Stephen Horner, Graziano Pesole
Bioinform.2
2001 Methods for Pattern Discovery in Unaligned Biological Sequences
Giulio Pavesi, Giancarlo Mauri, Graziano Pesole
Briefings Bioinform.3
2000 The Untranslated Regions of Eukaryotic MRNAs: Structure, Function, Evolution and Bioinformatic Tools for Their Analysis
abstract
The crucial role of the non-coding portion of genomes is now widely acknowledged. In particular, mRNA untranslated regions are involved in many post-transcriptional regulatory pathways that control mRNA localisation, stability and translation efficiency. A review is given of the most recent research works on the functional characterisation of eukaryotic mRNA untranslated regions. In order to make possible a systematic and detailed sequence analysis of mRNA untranslated regions (UTRs), a non-redundant database of metazoan mRNA untranslated sequences annotated for the occurrence of specific functional elements, UTRdb, was devised. These elements, whose consensus structure has been devised on the basis of experimental assays and of comparative analyses, have been collected in the UTRsite database. A suitable pattern-matching software has been devised to search UTRsite patterns in user-submitted sequences, also assessing their statistical significance. Structural, compositional and evolutionary features of untranslated sequences of metazoan mRNAs have been investigated showing peculiar intra- and interspecific patterns.
Graziano Pesole, Giorgio Grillo, Alessandra Larizza, Sabino Liuni
Briefings Bioinform.1
2000 PatSearch: a pattern matcher software that finds functional elements in nucleotide and protein sequences and assesses their statistical significance
abstract
MOTIVATION: The identification of sequence patterns involved in gene regulation and expression is a major challenge in molecular biology. In this paper we describe a novel algorithm and the software for searching nucleotide and protein sequences for complex nucleotide patterns including potential secondary structure elements, also allowing for mismatches/mispairings below a user-fixed threshold, and assessing the statistical significance of their occurrence through a Markov chain simulation. RESULTS: The application of the proposed algorithm allowed the identification of some functional elements, such as the Iron Responsive Element, the Histone stem-loop structure and the Selenocysteine Insertion Sequence, located in the mRNA untranslated regions of post-transcriptionally regulated genes with the assessment of sensitivity and selectivity of the searching method. AVAILABILITY: A Web interface is available at: http://bigarea.area.ba.cnr.it:8000/EmbIT/Pats earch.html.
Graziano Pesole, Sabino Liuni, Mark D'Souza
Bioinform.1
1996 CLEANUP: a fast computer program for removing redundancies from nucleotide sequence databases
abstract
A key concept in comparing sequence collections is the issue of redundancy. The production of sequence collections free from redundancy is undoubtedly very useful, both in performing statistical analyses and accelerating extensive database searching on nucleotide sequences. Indeed, publicly available databases contain multiple entries of identical or almost identical sequences. Performing statistical analysis on such biased data makes the risk of assigning high significance to non-significant patterns very high. In order to carry out unbiased statistical analysis as well as more efficient database searching it is thus necessary to analyse sequence data that have been purged of redundancy. Given that a unambiguous definition of redundancy is impracticable for biological sequence data, in the present program a quantitative description of redundancy will be used, based on the measure of sequence similarity. A sequence is considered redundant if it shows a degree of similarity and overlapping with a longer sequence in the database greater than a threshold fixed by the user. In this paper we present a new algorithm based on an "approximate string matching' procedure, which is able to determine the overall degree of similarity between each pair of sequences contained in a nucleotide sequence database and to generate automatically nucleotide sequence collections free from redundancies.
Giorgio Grillo, Marcella Attimonelli, Sabino Liuni, Graziano Pesole
Comput. Appl. Biosci.4
1993 SIMD parallelization of the WORDUP algorithm for detecting statistically significant patterns in DNA sequences
abstract
The development of new techniques in sequencing nuclei acids has produced a great amount of sequence data and has led to the discovery of new relationships. In this paper, we study a method for parallelizing the algorithm WORDUP, which detects the presence of statistically significant patterns in DNA sequences. WORDUP implements an efficient method to identify the presence of statistically significant oligomers in a non-homologous group of sequences. It is based on a modified version of the Boyer-Moore algorithm, which is one of the fastest algorithms for string matching available in the literature. The aim of the parallel version of WORDUP presented here is to speed up the computational time and allow the analysis of a greater set of longer nucleotide sequences, which is usually impractical with sequential algorithms.
Sabino Liuni, N. Prunella, Graziano Pesole, Tiziana D'Orazio, Ettore Stella, Arcangelo Distante
Comput. Appl. Biosci.3
1993 FASTPAT: a fast and efficient algorithm for string searching in DNA sequences
abstract
A new string searching algorithm is presented aimed at searching for the occurrence of character patterns in longer character texts. The algorithm, specifically designed for nucleic acid sequence data, is essentially derived from the Boyer-Moore method (Comm. ACM, 20, 762-772, 1977). Both pattern and text data are compressed so that the natural 4-letter alphabet of nucleic acid sequences is considerably enlarged. The string search starts from the last character of the pattern and proceeds in large jumps through the text to be searched. The data compression and searching algorithm allows one to avoid searching for patterns not present in the text as well as to inspect, for each pattern, all text characters until the exact match with the text is found. These considerations are supported by empirical evidence and comparisons with other methods.
N. Prunella, Sabino Liuni, Marcella Attimonelli, Graziano Pesole
Comput. Appl. Biosci.4