EDBT 2026 Demo / reviewers in the wild / expert
Chiara Romualdi
dblp:74/4363
· DBLP profile ↗
22ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-4792-9047ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAMURAI: shallow analysis of copy number alterations using a reproducible and integrated bioinformatics pipelineabstractShallow whole-genome sequencing (sWGS) offers a cost-effective approach to detect copy number alterations (CNAs). However, there remains a gap for a standardized workflow specifically designed for sWGS analysis. To address this need, in this work we present SAMURAI, a bioinformatics pipeline specifically designed for analyzing CNAs from sWGS data in a standardized and reproducible manner. SAMURAI is built using established community standards, ensuring portability, scalability, and reproducibility. The pipeline features a modular design with independent blocks for data preprocessing, copy number analysis, and customized reporting. Users can select workflows tailored for either solid or liquid biopsy analysis (e.g. circulating tumor DNA), with specific tools integrated for each sample type. The final report generated by SAMURAI provides detailed results to facilitate data interpretation and potential downstream analyses. To demonstrate its robustness, SAMURAI was validated using simulated and real-world data sets. The pipeline achieved high concordance with ground truth data and maintained consistent performance across various scenarios. By promoting standardization and offering a versatile workflow, SAMURAI empowers researchers in diverse environments to reliably analyze CNAs from sWGS data. This, in turn, holds promise for advancements in precision medicine. Sara Potente, Diego Boscarino, Dino Paladin, Sergio Marchini, Luca Beltrame, Chiara Romualdi |
Briefings Bioinform. | 6 |
| 2023 | benchdamic: benchmarking of differential abundance methods for microbiome dataabstractSUMMARY: Recently, an increasing number of methodological approaches have been proposed to tackle the complexity of metagenomics and microbiome data. In this scenario, reproducibility and replicability have become two critical issues, and the development of computational frameworks for the comparative evaluations of such methods is of utmost importance. Here, we present benchdamic, a Bioconductor package to benchmark methods for the identification of differentially abundant taxa. AVAILABILITY AND IMPLEMENTATION: benchdamic is available as an open-source R package through the Bioconductor project at https://bioconductor.org/packages/benchdamic/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matteo Calgaro, Chiara Romualdi, Davide Risso, Nicola Vitulo |
Bioinform. | 2 |
| 2023 | A pan-cancer landscape of pathogenic somatic copy number variationsabstractOBJECTIVE: Copy number variations (CNVs) play crucial roles in physiological and pathological processes, including cancer. However, the functional implications of somatic CNVs in tumor progression and evolution remain unclear. This study focuses on identifying CNV alterations with high pathogenic potential that drive and sustain tumorigenesis, distinguishing them from passenger alterations that accumulate during tumor growth. Our goal is to explore the variability of CNVs across different tumor types and infer their impact on tumor cell functions. METHODS: Starting from 7352 copy number profiles across 33 different cancer types, we infer the pathogenicity of each CNV and perform both intra- and inter-tumor analyses to predict the functional impact of different genomic patterns. We evaluate the actionability of genes belonging to altered regions and we correlate the presence of pathogenic regions with genome instability patterns and patients' survival. RESULTS: Our analysis uncovered large heterogeneity among different tumors suggesting in many cases distinct genetic drivers of tumorigenesis. Recurrent genomic alterations frequently coincide with dysfunctional homologous recombination pathways and negative regulation of the immune system. In certain tumors, the number of pathogenic CNVs emerged as a prognostic biomarker, highlighting their significance in cancer progression. CONCLUSION: This study contributes to elucidate the functional impact of pathogenic CNVs in tumor progression and sheds light on their potential as prognostic markers in specific cancer types. Tommaso Becchi, Luca Beltrame, Laura Mannarino, Enrica Calura, Sergio Marchini, Chiara Romualdi |
J. Biomed. Informatics | 6 |
| 2022 | NewWave: a scalable R/Bioconductor package for the dimensionality reduction and batch effect removal of single-cell RNA-seq dataabstractSUMMARY: We present NewWave, a scalable R/Bioconductor package for the dimensionality reduction and batch effect removal of single-cell RNA sequencing data. To achieve scalability, NewWave uses mini-batch optimization and can work with out-of-memory data, enabling users to analyze datasets with millions of cells. AVAILABILITY AND IMPLEMENTATION: NewWave is implemented as an open-source R package available through the Bioconductor project at https://bioconductor.org/packages/NewWave/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Federico Agostinis, Chiara Romualdi, Gabriele Sales, Davide Risso |
Bioinform. | 2 |
| 2022 | The power of word-frequency-based alignment-free functions: a comprehensive large-scale experimental analysisabstractMOTIVATION: Alignment-free (AF) distance/similarity functions are a key tool for sequence analysis. Experimental studies on real datasets abound and, to some extent, there are also studies regarding their control of false positive rate (Type I error). However, assessment of their power, i.e. their ability to identify true similarity, has been limited to some members of the D2 family. The corresponding experimental studies have concentrated on short sequences, a scenario no longer adequate for current applications, where sequence lengths may vary considerably. Such a State of the Art is methodologically problematic, since information regarding a key feature such as power is either missing or limited. RESULTS: By concentrating on a representative set of word-frequency-based AF functions, we perform the first coherent and uniform evaluation of the power, involving also Type I error for completeness. Two alternative models of important genomic features (CIS Regulatory Modules and Horizontal Gene Transfer), a wide range of sequence lengths from a few thousand to millions, and different values of k have been used. As a result, we provide a characterization of those AF functions that is novel and informative. Indeed, we identify weak and strong points of each function considered, which may be used as a guide to choose one for analysis tasks. Remarkably, of the 15 functions that we have considered, only four stand out, with small differences between small and short sequence length scenarios. Finally, to encourage the use of our methodology for validation of future AF functions, the Big Data platform supporting it is public. AVAILABILITY AND IMPLEMENTATION: The software is available at: https://github.com/pipp8/power_statistics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Giuseppe Cattaneo, Umberto Ferraro Petrillo, Raffaele Giancarlo, Francesco Palini, Chiara Romualdi |
Bioinform. | 5 |
| 2021 | PsiNorm: a scalable normalization for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) enables transcriptome-wide gene expression measurements at single-cell resolution providing a comprehensive view of the compositions and dynamics of tissue and organism development. The evolution of scRNA-seq protocols has led to a dramatic increase of cells throughput, exacerbating many of the computational and statistical issues that previously arose for bulk sequencing. In particular, with scRNA-seq data all the analyses steps, including normalization, have become computationally intensive, both in terms of memory usage and computational time. In this perspective, new accurate methods able to scale efficiently are desirable. RESULTS: Here, we propose PsiNorm, a between-sample normalization method based on the power-law Pareto distribution parameter estimate. Here, we show that the Pareto distribution well resembles scRNA-seq data, especially those coming from platforms that use unique molecular identifiers. Motivated by this result, we implement PsiNorm, a simple and highly scalable normalization method. We benchmark PsiNorm against seven other methods in terms of cluster identification, concordance and computational resources required. We demonstrate that PsiNorm is among the top performing methods showing a good trade-off between accuracy and scalability. Moreover, PsiNorm does not need a reference, a characteristic that makes it useful in supervised classification settings, in which new out-of-sample data need to be normalized. AVAILABILITY AND IMPLEMENTATION: PsiNorm is implemented in the scone Bioconductor package and available at https://bioconductor.org/packages/scone/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matteo Borella, Graziano Martello, Davide Risso, Chiara Romualdi |
Bioinform. | 4 |
| 2019 | metaGraphite-a new layer of pathway annotation to get metabolite networksabstractMOTIVATION: Metabolomics is an emerging 'omics' science involving the characterization of metabolites and metabolism in biological systems. Few bioinformatic tools have been developed for the visualization, exploration and analysis of metabolomic data within the context of metabolic pathways: some of them became rapidly obsolete and are no longer supported, others are based on a single database. A systematic collection of existing annotations has the potential of considerably boosting the investigation and contextualization of metabolomic measurements. RESULTS: We have released a major update of our Bioconductor package graphite which explicitly tracks small molecules within pathway topologies and their interactions with proteins. The package gathers the information stored in eight major databases, oriented both at genes and at metabolites, across 14 different species. Depending on user preferences, all pathways can be retrieved as gene-only, gene metabolite or metabolite-only networks. AVAILABILITY AND IMPLEMENTATION: The new graphite version (1.24) is available on Bioconductor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Sales, Enrica Calura, Chiara Romualdi |
Bioinform. | 3 |
| 2019 | SourceSet: A graphical model approach to identify primary genes in perturbed biological pathwaysabstractTopological gene-set analysis has emerged as a powerful means for omic data interpretation. Although numerous methods for identifying dysregulated genes have been proposed, few of them aim to distinguish genes that are the real source of perturbation from those that merely respond to the signal dysregulation. Here, we propose a new method, called SourceSet, able to distinguish between the primary and the secondary dysregulation within a Gaussian graphical model context. The proposed method compares gene expression profiles in the control and in the perturbed condition and detects the differences in both the mean and the covariance parameters with a series of likelihood ratio tests. The resulting evidence is used to infer the primary and the secondary set, i.e. the genes responsible for the primary dysregulation, and the genes affected by the perturbation through network propagation. The proposed method demonstrates high specificity and sensitivity in different simulated scenarios and on several real biological case studies. In order to fit into the more traditional pathway analysis framework, SourceSet R package also extends the analysis from a single to multiple pathways and provides several graphical outputs, including Cytoscape visualization to browse the results. Elisa Salviato, Vera Djordjilovic, Monica Chiogna, Chiara Romualdi |
PLoS Comput. Biol. | 4 |
| 2017 | simPATHy: a new method for simulating data from perturbed biological PATHwaysabstractSummary: In the omic era, one of the main aims is to discover groups of functionally related genes that drive the difference between different conditions. To this end, a plethora of potentially useful multivariate statistical approaches has been proposed, but their evaluation is hindered by the absence of a gold standard. Here, we propose a method for simulating biological data – gene expression, RPKM/FPKM or protein abundances – from two conditions, namely, a reference condition and a perturbation of it. Our approach is built upon probabilistic graphical models and is thus especially suited for testing topological approaches. Availability and Implementation: The simPATHy is an R package, it is open source and freely available on CRAN. Contacts: [email protected] or [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Elisa Salviato, Vera Djordjilovic, Monica Chiogna, Chiara Romualdi |
Bioinform. | 4 |
| 2014 | timeClip: pathway analysis for time course data without replicatesabstractBACKGROUND: Time-course gene expression experiments are useful tools for exploring biological processes. In this type of experiments, gene expression changes are monitored along time. Unfortunately, replication of time series is still costly and usually long time course do not have replicates. Many approaches have been proposed to deal with this data structure, but none of them in the field of pathway analysis. Pathway analyses have acquired great relevance for helping the interpretation of gene expression data. Several methods have been proposed to this aim: from the classical enrichment to the more complex topological analysis that gains power from the topology of the pathway. None of them were devised to identify temporal variations in time course data. RESULTS: Here we present timeClip, a topology based pathway analysis specifically tailored to long time series without replicates. timeClip combines dimension reduction techniques and graph decomposition theory to explore and identify the portion of pathways that is most time-dependent. In the first step, timeClip selects the time-dependent pathways; in the second step, the most time dependent portions of these pathways are highlighted. We used timeClip on simulated data and on a benchmark dataset regarding mouse muscle regeneration model. Our approach shows good performance on different simulated settings. On the real dataset, we identify 76 time-dependent pathways, most of which known to be involved in the regeneration process. Focusing on the 'mTOR signaling pathway' we highlight the timing of key processes of the muscle regeneration: from the early pathway activation through growth factor signals to the late burst of protein production needed for the fiber regeneration. CONCLUSIONS: timeClip represents a new improvement in the field of time-dependent pathway analysis. It allows to isolate and dissect pathways characterized by time-dependent components. Furthermore, using timeClip on a mouse muscle regeneration dataset we were able to characterize the process of muscle fiber regeneration with its correct timing. Paolo G. V. Martini, Gabriele Sales, Enrica Calura, Stefano Cagnin, Monica Chiogna, Chiara Romualdi |
BMC Bioinform. | 6 |
| 2012 | graphite - a Bioconductor package to convert pathway topology to gene networkabstractBACKGROUND: Gene set analysis is moving towards considering pathway topology as a crucial feature. Pathway elements are complex entities such as protein complexes, gene family members and chemical compounds. The conversion of pathway topology to a gene/protein networks (where nodes are a simple element like a gene/protein) is a critical and challenging task that enables topology-based gene set analyses.Unfortunately, currently available R/Bioconductor packages provide pathway networks only from single databases. They do not propagate signals through chemical compounds and do not differentiate between complexes and gene families. RESULTS: Here we present graphite, a Bioconductor package addressing these issues. Pathway information from four different databases is interpreted following specific biologically-driven rules that allow the reconstruction of gene-gene networks taking into account protein complexes, gene families and sensibly removing chemical compounds from the final graphs. The resulting networks represent a uniform resource for pathway analyses. Indeed, graphite provides easy access to three recently proposed topological methods. The graphite package is available as part of the Bioconductor software suite. CONCLUSIONS: graphite is an innovative package able to gather and make easily available the contents of the four major pathway databases. In the field of topological analysis graphite acts as a provider of biological information by reducing the pathway complexity considering the biological meaning of the pathway elements. Gabriele Sales, Enrica Calura, Duccio Cavalieri, Chiara Romualdi |
BMC Bioinform. | 4 |
| 2011 | The Biological Connection Markup Language: a SBGN-compliant format for visualization, filtering and analysis of biological pathwaysabstractMOTIVATION: Many models and analysis of signaling pathways have been proposed. However, neither of them takes into account that a biological pathway is not a fixed system, but instead it depends on the organism, tissue and cell type as well as on physiological, pathological and experimental conditions. RESULTS: The Biological Connection Markup Language (BCML) is a format to describe, annotate and visualize pathways. BCML is able to store multiple information, permitting a selective view of the pathway as it exists and/or behave in specific organisms, tissues and cells. Furthermore, BCML can be automatically converted into data formats suitable for analysis and into a fully SBGN-compliant graphical representation, making it an important tool that can be used by both computational biologists and 'wet lab' scientists. AVAILABILITY AND IMPLEMENTATION: The XML schema and the BCML software suite are freely available under the LGPL for download at http://bcml.dc-atlas.net. They are implemented in Java and supported on MS Windows, Linux and OS X. Luca Beltrame, Enrica Calura, Razvan R. Popovici, Lisa Rizzetto, Damariz Rivero Guedez, Michele Donato, Chiara Romualdi, Sorin Draghici, Duccio Cavalieri |
Bioinform. | 7 |
| 2011 | parmigene - a parallel R package for mutual information estimation and gene network reconstructionabstractMOTIVATION: Inferring large transcriptional networks using mutual information has been shown to be effective in several experimental setup. Unfortunately, this approach has two main drawbacks: (i) several mutual information estimators are prone to biases and (ii) available software still has large computational costs when processing thousand of genes. RESULTS: Here, we present parmigene (PARallel Mutual Information estimation for GEne NEtwork reconstruction), an R package that tries to fill the above gaps. It implements a mutual information estimator based on k-nearest neighbor distances that is minimally biased with respect to the other methods and uses a parallel computing paradigm to reconstruct gene regulatory networks. We test parmigene on in silico and real data. We show that parmigene gives more precise results than existing softwares with strikingly less computational costs. AVAILABILITY AND IMPLEMENTATION: The parmigene package is available on the CRAN network at http://cran.r-project.org/web/packages/. CONTACT: [email protected] Gabriele Sales, Chiara Romualdi |
Bioinform. | 2 |
| 2011 | Statistical Test of Expression Pattern (STEPath): a new strategy to integrate gene expression data with genomic information in individual and meta-analysis studiesabstractBACKGROUND: In the last decades, microarray technology has spread, leading to a dramatic increase of publicly available datasets. The first statistical tools developed were focused on the identification of significant differentially expressed genes. Later, researchers moved toward the systematic integration of gene expression profiles with additional biological information, such as chromosomal location, ontological annotations or sequence features. The analysis of gene expression linked to physical location of genes on chromosomes allows the identification of transcriptionally imbalanced regions, while, Gene Set Analysis focuses on the detection of coordinated changes in transcriptional levels among sets of biologically related genes. In this field, meta-analysis offers the possibility to compare different studies, addressing the same biological question to fully exploit public gene expression datasets. RESULTS: We describe STEPath, a method that starts from gene expression profiles and integrates the analysis of imbalanced region as an a priori step before performing gene set analysis. The application of STEPath in individual studies produced gene set scores weighted by chromosomal activation. As a final step, we propose a way to compare these scores across different studies (meta-analysis) on related biological issues. One complication with meta-analysis is batch effects, which occur because molecular measurements are affected by laboratory conditions, reagent lots and personnel differences. Major problems occur when batch effects are correlated with an outcome of interest and lead to incorrect conclusions. We evaluated the power of combining chromosome mapping and gene set enrichment analysis, performing the analysis on a dataset of leukaemia (example of individual study) and on a dataset of skeletal muscle diseases (meta-analysis approach). In leukaemia, we identified the Hox gene set, a gene set closely related to the pathology that other algorithms of gene set analysis do not identify, while the meta-analysis approach on muscular disease discriminates between related pathologies and correlates similar ones from different studies. CONCLUSIONS: STEPath is a new method that integrates gene expression profiles, genomic co-expressed regions and the information about the biological function of genes. The usage of the STEPath-computed gene set scores overcomes batch effects in the meta-analysis approaches allowing the direct comparison of different pathologies and different studies on a gene set activation level. Paolo G. V. Martini, Davide Risso, Gabriele Sales, Chiara Romualdi, Gerolamo Lanfranchi, Stefano Cagnin |
BMC Bioinform. | 4 |
| 2009 | A modified LOESS normalization applied to microRNA arrays: a comparative evaluationabstractMOTIVATION: Microarray normalization is a fundamental step in removing systematic bias and noise variability caused by technical and experimental artefacts. Several approaches, suitable for large-scale genome arrays, have been proposed and shown to be effective in the reduction of systematic errors. Most of these methodologies are based on specific assumptions that are reasonable for whole-genome arrays, but possibly unsuitable for small microRNA (miRNA) platforms. In this work, we propose a novel normalization (loessM), and we investigate, through simulated and real datasets, the influence that normalizations for two-colour miRNA arrays have on the identification of differentially expressed genes. RESULTS: We show that normalizations usually applied to large-scale arrays, in several cases, modify the actual structure of miRNA data, leading to large portions of false positives and false negatives. Nevertheless, loessM is able to outperform other techniques in most experimental scenarios. Moreover, when usual assumptions on differential expression distribution are missed, channel effect has a strikingly negative influence on small arrays, bias that cannot be removed by normalizations but rather by an appropriate experimental design. We find that the combination of loessM with eCADS, an experimental design based on biological replicates dye-swap recently proposed for channel-effect reduction, gives better results in most of the experimental conditions in terms of specificity/sensitivity both on simulated and real data. AVAILABILITY: LoessM R function is freely available at http://gefu.cribi.unipd.it/papers/miRNA-simulation/ Davide Risso, Maria Sofia Massa, Monica Chiogna, Chiara Romualdi |
Bioinform. | 4 |
| 2009 | A-MADMAN: Annotation-based microarray data meta-analysis toolabstractBACKGROUND: Publicly available datasets of microarray gene expression signals represent an unprecedented opportunity for extracting genomic relevant information and validating biological hypotheses. However, the exploitation of this exceptionally rich mine of information is still hampered by the lack of appropriate computational tools, able to overcome the critical issues raised by meta-analysis. RESULTS: This work presents A-MADMAN, an open source web application which allows the retrieval, annotation, organization and meta-analysis of gene expression datasets obtained from Gene Expression Omnibus. A-MADMAN addresses and resolves several open issues in the meta-analysis of gene expression data. CONCLUSION: A-MADMAN allows i) the batch retrieval from Gene Expression Omnibus and the local organization of raw data files and of any related meta-information, ii) the re-annotation of samples to fix incomplete, or otherwise inadequate, metadata and to create user-defined batches of data, iii) the integrative analysis of data obtained from different Affymetrix platforms through custom chip definition files and meta-normalization. Software and documentation are available on-line at http://compgen.bio.unipd.it/bioinfo/amadman/. Andrea Bisognin, Alessandro Coppe, Francesco Ferrari, Davide Risso, Chiara Romualdi, Silvio Bicciato, Stefania Bortoluzzi |
BMC Bioinform. | 5 |
| 2009 | A comparison on effects of normalisations in the detection of differentially expressed genesabstractBACKGROUND: Various normalisation techniques have been developed in the context of microarray analysis to try to correct expression measurements for experimental bias and random fluctuations. Major techniques include: total intensity normalisation; intensity dependent normalisation; and variance stabilising normalisation. The aim of this paper is to discuss the impact of normalisation techniques for two-channel array technology on the process of identification of differentially expressed genes. RESULTS: Through three precise simulation plans, we quantify the impact of normalisations: (a) on the sensitivity and specificity of a specified test statistic for the identification of deregulated genes, (b) on the gene ranking induced by the statistic. CONCLUSION: Although we found a limited difference of sensitivities and specificities for the test after each normalisation, the study highlights a strong impact in terms of gene ranking agreement, resulting in different levels of agreement between competing normalisations. However, we show that the combination of two normalisations, such as glog and lowess, that handle different aspects of microarray data, is able to outperform other individual techniques. Monica Chiogna, Maria Sofia Massa, Davide Risso, Chiara Romualdi |
BMC Bioinform. | 4 |
| 2007 | A global gene evolution analysis on Vibrionaceae family using phylogenetic profileabstractBACKGROUND: Vibrionaceae represent a significant portion of the cultivable heterotrophic sea bacteria; they strongly affect nutrient cycling and some species are devastating pathogens. In this work we propose an improved phylogenetic profile analysis on 14 Vibrionaceae genomes, to study the evolution of this family on the basis of gene content. The phylogenetic profile is based on the observation that genes involved in the same process (e.g. metabolic pathway or structural complex) tend to be concurrently present or absent within different genomes. This allows the prediction of hypothetical functions on the basis of a shared phylogenetic profiles. Moreover this approach is useful to identify putative laterally transferred elements on the basis of their presence on distantly phylogenetically related bacteria. RESULTS: Vibrionaceae ORFs were aligned against all the available bacterial proteomes. Phylogenetic profile is defined as an array of distances, based on aminoacid substitution matrixes, from single genes to all their orthologues. Final phylogenetic profiles, derived from non-redundant list of all ORFs, was defined as the median of all the profiles belonging to the cluster. The resulting phylogenetic profiles matrix contains gene clusters on the rows and organisms on the columns. Cluster analysis identified groups of "core genes" with a widespread high similarity across all the organisms and several clusters that contain genes homologous only to a limited set of organisms. On each of these clusters, COG class enrichment has been calculated. The analysis reveals that clusters of core genes have the highest number of enriched classes, while the others are enriched just for few of them like DNA replication, recombination and repair. CONCLUSION: We found that mobile elements have heterogeneous profiles not only across the entire set of organisms, but also within Vibrionaceae; this confirms their great influence on bacteria evolution even inside the same family. Furthermore, several hypothetical proteins highly correlate with mobile elements profiles suggesting a possible horizontal transfer mechanism for the evolution of these genes. Finally, we suggested the putative role of some ORFs having an unknown function on the basis of their phylogenetic profile similarity to well characterized genes. Nicola Vitulo, Alessandro Vezzi, Chiara Romualdi, Stefano Campanaro, Giorgio Valle |
BMC Bioinform. | 3 |
| 2005 | RAP: a new computer program for de novo identification of repeated sequences in whole genomesabstractMOTIVATION: DNA repeats are a common feature of most genomic sequences. Their de novo identification is still difficult despite being a crucial step in genomic analysis and oligonucleotides design. Several efficient algorithms based on word counting are available, but too short words decrease specificity while long words decrease sensitivity, particularly in degenerated repeats. RESULTS: The Repeat Analysis Program (RAP) is based on a new word-counting algorithm optimized for high resolution repeat identification using gapped words. Many different overlapping gapped words can be counted at the same genomic position, thus producing a better signal than the single ungapped word. This results in better specificity both in terms of low-frequency detection, being able to identify sequences repeated only once, and highly divergent detection, producing a generally high score in most intron sequences. AVAILABILITY: The program is freely available for non-profit organizations, upon request to the authors. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: The program has been tested on the Caenorhabditis elegans genome using word lengths of 12, 14 and 16 bases. The full analysis has been implemented in the UCSC Genome Browser and is accessible at http://genome.cribi.unipd.it. Davide Campagna, Chiara Romualdi, Nicola Vitulo, Micky Del Favero, Matej Lexa, Nicola Cannata, Giorgio Valle |
Bioinform. | 2 |
| 2003 | TRAIT (TRAnscript Integrated Table): a knowledgebase of human skeletal muscle transcriptsabstractAbstract Summary: TRAIT is a knowledgebase integrating information on transcripts with related data from genome, proteins, ortholog genes and diseases. It was initially built as a system to manage an EST-based gene discovery project on human skeletal muscle, which yielded over 4500 independent sequence clusters. Transcripts are annotated using automatic as well as manual procedures, linking known transcripts to public databases and unknown transcripts to tables of predicted features. Data are stored in a MySQL database. Complex queries are automatically built by means of a user-friendly web interface that allows the concurrent selection of many fields such as ontology, expression level, map position and protein domains. The results are parsed by the system and returned in a ranked order, in respect to the number of satisfied criteria. Availability: http://muscle.cribi.unipd.it and http://muscle.cribi.unipd.it/features/querystrait.html Contact: [email protected]; [email protected] * To whom correspondence should be addressed. Stefano Toppo, Nicola Cannata, Paolo Fontana, Chiara Romualdi, Paolo Laveder, Emanuela Bertocco, Gerolamo Lanfranchi, Giorgio Valle |
Bioinform. | 4 |
| 2002 | Simplifying amino acid alphabets by means of a branch and bound algorithm and substitution matricesabstractMOTIVATION: Protein and DNA are generally represented by sequences of letters. In a number of circumstances simplified alphabets (where one or more letters would be represented by the same symbol) have proved their potential utility in several fields of bioinformatics including searching for patterns occurring at an unexpected rate, studying protein folding and finding consensus sequences in multiple alignments. The main issue addressed in this paper is the possibility of finding a general approach that would allow an exhaustive analysis of all the possible simplified alphabets, using substitution matrices like PAM and BLOSUM as a measure for scoring. RESULTS: The computational approach presented in this paper has led to a computer program called AlphaSimp (Alphabet Simplifier) that can perform an exhaustive analysis of the possible simplified amino acid alphabets, using a branch and bound algorithm together with standard or user-defined substitution matrices. The program returns a ranked list of the highest-scoring simplified alphabets. When the extent of the simplification is limited and the simplified alphabets are maintained above ten symbols the program is able to complete the analysis in minutes or even seconds on a personal computer. However, the performance becomes worse, taking up to several hours, for highly simplified alphabets. AVAILABILITY: AlphaSimp and other accessory programs are available at http://bioinformatics.cribi.unipd.it/alphasimp Nicola Cannata, Stefano Toppo, Chiara Romualdi, Giorgio Valle |
Bioinform. | 3 |
| 2001 | Differential expression of genes coding for ribosomal proteins in different human tissuesabstractMOTIVATION: To perform a computational and statistical study on a large set of gene expression data pertaining six adult human tissues (brain, liver, skeletal muscle, ovary, retina and uterus) for analyzing the expression of ribosomal protein genes. RESULTS: Unexpectedly, in each of the considered tissues large variations in the expression of ribosomal protein genes were observed. Moreover, when comparing the expression levels of 89 ribosomal protein genes in six different tissues, 13 genes appeared differentially expressed among tissues. AVAILABILITY: The expression data of the ribosomal protein genes together with supplementary material (complete transcriptional profiles of the considered human tissues) are freely available at the site GETProfiles (http://telethon.bio.unipd.it/GETProfiles/). CONTACT: [email protected] Stefania Bortoluzzi, Fabio d'Alessi, Chiara Romualdi, Gian Antonio Danieli |
Bioinform. | 3 |