VLDB 2026 Research / reviewers in the wild / expert
Francesca Cordero
dblp:55/1163
· DBLP profile ↗
28ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-3143-3330ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UnifiedGreatMod: a new holistic modelling paradigm for studying biological systems on a complete and harmonious scaleabstractMOTIVATION: Computational models are crucial for addressing critical questions about systems evolution and deciphering system connections. The pivotal feature of making this concept recognizable from the biological and clinical community is the possibility of quickly inspecting the whole system, bearing in mind the different granularity levels of its components. This holistic view of system behaviour expands the evolution study by identifying the heterogeneous behaviours applicable, e.g. to the cancer evolution study. RESULTS: To address this aspect, we propose a new modelling paradigm, UnifiedGreatMod, which allows modellers to integrate fine-grained and coarse-grained biological information into a unique model. It enables functional studies by combining the analysis of the system's multi-level stable states with its fluctuating conditions. This approach helps to investigate the functional relationships and dependencies among biological entities. This is achieved, thanks to the hybridization of two analysis approaches that capture a system's different granularity levels. The proposed paradigm was then implemented into the open-source, general modelling framework GreatMod, in which a graphical meta-formalism is exploited to simplify the model creation phase and R languages to define user-defined analysis workflows. The proposal's effectiveness was demonstrated by mechanistically simulating the metabolic output of Escherichia coli under environmental nutrient perturbations and integrating a gene expression dataset. Additionally, the UnifiedGreatMod was used to examine the responses of luminal epithelial cells to Clostridium difficile infection. AVAILABILITY AND IMPLEMENTATION: GreatMod https://qbioturin.github.io/epimod/, epimod_FBAfunctions https://github.com/qBioTurin/epimod_FBAfunctions, first case study E. coli https://github.com/qBioTurin/Ec_coli_modelling, second case study C. difficile https://github.com/qBioTurin/EpiCell_CDifficile. Riccardo Aucello, Simone Pernice, Dora Tortarolo, Raffaele A. Calogero, Celia Herrera-Rincon, Giulia Ronchi, Stefano Geuna, Francesca Cordero, Pietro Liò, Marco Beccuti |
Bioinform. | 8 |
| 2023 | OmniReprodubileCellAnalysis: a comprehensive toolbox for the analysis of cellular biology dataabstractOpen science and reproducibility are two key pillars of modern scientific research. Open science is making scientific research and data accessible and transparent to the broader scientific community and the public. Reproducibility, on the other hand, is the ability to replicate and confirm research results by following the same methods and procedures. Reproducibility is thus crucial because it ensures the reliability and validity of scientific findings. The relationship between open science and reproducibility is intertwined; indeed open science practices, such as sharing raw data, detailed methodologies, and code, greatly facilitate the reproducibility of research. In recent years, concerns about the reproducibility of scientific research have gained prominence, and indeed scientists still lament the lack of details in the methods sections of published papers and the unavailability of raw data from the authors.To assist cellular biologists and immunologists and to promote a more transparent, open and reproducible research practice, we developed OmniReproducibleCellAnalysis (ORCA), a new Shiny Application based in R, for the semi-automated analysis of Western Blot (WB), Reverse Transcription-quantitative PCR (RT-qPCR), Enzyme-Linked ImmunoSorbent Assay (ELISA), Endocytosis and Cytotoxicity experiments. ORCA is open-source and approachable by scientists without advanced R language knowledge. Our application automatically compiles a report containing the finalized data analysis and all its preliminary and intermediate steps, ensuring data analysis standardization and reproducibility. Furthermore, ORCA allows to upload raw data and results directly on the data repository Harvard Dataverse, a valuable tool for promoting transparency and data accessibility in scientific research.By employing ORCA, scientists will cut down analysis time and human-dependent errors, while taking a step towards a research practice compliant with Open Science and FAIR principle. Dora Tortarolo, Simone Pernice, Fabiana Clapero, Donatella Valdembri, Guido Serini, Federica Riccardo, Lidia Tarone, Chiara Enrico Bena, Carla Bosia, Sandro Gepiro Contaldo, Marco Beccuti, Marzio Pennisi, Francesca Cordero |
BIBM | 13 |
| 2023 | CONNECTOR, fitting and clustering of longitudinal data to reveal a new risk stratification systemabstractMOTIVATION: The transition from evaluating a single time point to examining the entire dynamic evolution of a system is possible only in the presence of the proper framework. The strong variability of dynamic evolution makes the definition of an explanatory procedure for data fitting and clustering challenging. RESULTS: We developed CONNECTOR, a data-driven framework able to analyze and inspect longitudinal data in a straightforward and revealing way. When used to analyze tumor growth kinetics over time in 1599 patient-derived xenograft growth curves from ovarian and colorectal cancers, CONNECTOR allowed the aggregation of time-series data through an unsupervised approach in informative clusters. We give a new perspective of mechanism interpretation, specifically, we define novel model aggregations and we identify unanticipated molecular associations with response to clinically approved therapies. AVAILABILITY AND IMPLEMENTATION: CONNECTOR is freely available under GNU GPL license at https://qbioturin.github.io/connector and https://doi.org/10.17504/protocols.io.8epv56e74g1b/v1. Simone Pernice, Roberta Sirovich, Elena Grassi, Marco Viviani 0002, Martina Ferri, Francesco Sassi, Luca Alessandrì, Dora Tortarolo, Raffaele A. Calogero, Livio Trusolino, Andrea Bertotti, Marco Beccuti, Martina Olivero, Francesca Cordero |
Bioinform. | 14 |
| 2023 | A new computational workflow to guide personalized drug therapyabstractOBJECTIVE: Computational models are at the forefront of the pursuit of personalized medicine thanks to their descriptive and predictive abilities. In the presence of complex and heterogeneous data, patient stratification is a prerequisite for effective precision medicine, since disease development is often driven by individual variability and unpredictable environmental events. Herein, we present GreatNectorworkflow as a valuable tool for (i) the analysis and clustering of patient-derived longitudinal data, and (ii) the simulation of the resulting model of patient-specific disease dynamics. METHODS: GreatNectoris designed by combining an analytic strategy composed of CONNECTOR, a data-driven framework for the inspection of longitudinal data, and an unsupervised methodology to stratify the subjects with GreatMod, a quantitative modeling framework based on the Petri Net formalism and its generalizations. RESULTS: To illustrate GreatNectorcapabilities, we exploited longitudinal data of four immune cell populations collected from Multiple Sclerosis patients. Our main results report that the T-cell dynamics after alemtuzumab treatment separate non-responders versus responders patients, and the patients in the non-responders group are characterized by an increase of the Th17 concentration around 36 months. CONCLUSION: GreatNectoranalysis was able to stratify individual patients into three model meta-patients whose dynamics suggested insight into patient-tailored interventions. Simone Pernice, Alessandro Maglione, Dora Tortarolo, Roberta Sirovich, Marinella Clerico, Simona Rolla, Marco Beccuti, Francesca Cordero |
J. Biomed. Informatics | 8 |
| 2020 | Computational modeling of the immune response in multiple sclerosis using epimod frameworkabstractBACKGROUND: Multiple Sclerosis (MS) represents nowadays in Europe the leading cause of non-traumatic disabilities in young adults, with more than 700,000 EU cases. Although huge strides have been made over the years, MS etiology remains partially unknown. Furthermore, the presence of various endogenous and exogenous factors can greatly influence the immune response of different individuals, making it difficult to study and understand the disease. This becomes more evident in a personalized-fashion when medical doctors have to choose the best therapy for patient well-being. In this optics, the use of stochastic models, capable of taking into consideration all the fluctuations due to unknown factors and individual variability, is highly advisable. RESULTS: We propose a new model to study the immune response in relapsing remitting MS (RRMS), the most common form of MS that is characterized by alternate episodes of symptom exacerbation (relapses) with periods of disease stability (remission). In this new model, both the peripheral lymph node/blood vessel and the central nervous system are explicitly represented. The model was created and analysed using Epimod, our recently developed general framework for modeling complex biological systems. Then the effectiveness of our model was shown by modeling the complex immunological mechanisms characterizing RRMS during its course and under the DAC administration. CONCLUSIONS: Simulation results have proven the ability of the model to reproduce in silico the immune T cell balance characterizing RRMS course and the DAC effects. Furthermore, they confirmed the importance of a timely intervention on the disease course. Simone Pernice, Laura Follia, Alessandro Maglione, Marzio Pennisi, Francesco Pappalardo 0001, Francesco Novelli, Marinella Clerico, Marco Beccuti, Francesca Cordero, Simona Rolla |
BMC Bioinform. | 9 |
| 2020 | Integrating Petri Nets and Flux Balance Methods in Computational Biology Models: a Methodological and Computational PracticeabstractComputational Biology is a fast-growing field that is enriched by different data-driven methodological approaches and by findings and applications in a broad range of biological areas. Fundamental to these approaches are the mathematical and computational models used to describe the different state s at microscopic (for example a biochemical reaction), mesoscopic (the signalling effects at tissue level), and macroscopic levels (physiological and pathological effects) of biological processes. In this paper we address the problem of combining two powerful classes of methodologies: Flux Balance Analysis (FBA) methods which are now producing a revolution in biotechnology and medicine, and Petri Nets (PNs) which allow system generalisation and are central to various mathematical treatments, for example Ordinary Differential Equation (ODE) specification of the biosystem under study. While the former is limited to modelling metabolic networks, i.e. does not account for intermittent dynamical signalling events, the latter is hampered by the need for a large amount of metabolic data. A first result presented in this paper is the identification of three types of cross-talks between PNs and FBA methods and their dependencies on available data. We exemplify our insights with the analysis of a pancreatic cancer model. We discuss how our reasoning framework provides a biologically and mathematically grounded decision making setting for the integration of regulatory, signalling, and metabolic networks and greatly increases model interpretability and reusability. We discuss how the parameters of PN and FBA models can be tuned and combined together so to highlight the computational effort needed to perform this task. We conclude with speculations and suggestions on this new promising research direction. Simone Pernice, Laura Follia, Gianfranco Balbo, Luciano Milanesi, Giulia Sartini, Niccoló Totis, Pietro Liò, Ivan Merelli, Francesca Cordero, Marco Beccuti |
Fundam. Informaticae | 9 |
| 2019 | BITS2018: the fifteenth annual meeting of the Italian Society of BioinformaticsabstractThis preface introduces the content of the BioMed Central Bioinformatics journal Supplement related to the 15th annual meeting of the Bioinformatics Italian Society, BITS2018. The Conference was held in Torino, Italy, from June 27th to 29th, 2018. Francesca Cordero, Raffaele A. Calogero, Michele Caselle |
BMC Bioinform. | 1 |
| 2019 | A computational approach based on the colored Petri net formalism for studying multiple sclerosisabstractBACKGROUND: Multiple Sclerosis (MS) is an immune-mediated inflammatory disease of the Central Nervous System (CNS) which damages the myelin sheath enveloping nerve cells thus causing severe physical disability in patients. Relapsing Remitting Multiple Sclerosis (RRMS) is one of the most common form of MS in adults and is characterized by a series of neurologic symptoms, followed by periods of remission. Recently, many treatments were proposed and studied to contrast the RRMS progression. Among these drugs, daclizumab (commercial name Zinbryta), an antibody tailored against the Interleukin-2 receptor of T cells, exhibited promising results, but its efficacy was accompanied by an increased frequency of serious adverse events. Manifested side effects consisted of infections, encephalitis, and liver damages. Therefore daclizumab has been withdrawn from the market worldwide. Another interesting case of RRMS regards its progression in pregnant women where a smaller incidence of relapses until the delivery has been observed. RESULTS: In this paper we propose a new methodology for studying RRMS, which we implemented in GreatSPN, a state-of-the-art open-source suite for modelling and analyzing complex systems through the Petri Net (PN) formalism. This methodology exploits: (a) an extended Colored PN formalism to provide a compact graphical description of the system and to automatically derive a set of ODEs encoding the system dynamics and (b) the Latin Hypercube Sampling with PRCC index to calibrate ODE parameters for reproducing the real behaviours in healthy and MS subjects.To show the effectiveness of such methodology a model of RRMS has been constructed and studied. Two different scenarios of RRMS were thus considered. In the former scenario the effect of the daclizumab administration is investigated, while in the latter one RRMS was studied in pregnant women. CONCLUSIONS: We propose a new computational methodology to study RRMS disease. Moreover, we show that model generated and calibrated according to this methodology is able to reproduce the expected behaviours. Simone Pernice, Marzio Pennisi, Greta Romano, Alessandro Maglione, Santina Cutrupi, Francesco Pappalardo 0001, Gianfranco Balbo, Marco Beccuti, Francesca Cordero, Raffaele A. Calogero |
BMC Bioinform. | 9 |
| 2018 | SeqBox: RNAseq/ChIPseq reproducible analysis on a consumer game computerabstractSummary: Short reads sequencing technology has been used for more than a decade now. However, the analysis of RNAseq and ChIPseq data is still computational demanding and the simple access to raw data does not guarantee results reproducibility between laboratories. To address these two aspects, we developed SeqBox, a cheap, efficient and reproducible RNAseq/ChIPseq hardware/software solution based on NUC6I7KYK mini-PC (an Intel consumer game computer with a fast processor and a high performance SSD disk), and Docker container platform. In SeqBox the analysis of RNAseq and ChIPseq data is supported by a friendly GUI. This allows access to fast and reproducible analysis also to scientists with/without scripting experience. Availability and implementation: Docker container images, docker4seq package and the GUI are available at http://www.bioinformatica.unito.it/reproducibile.bioinformatics.html. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Marco Beccuti, Francesca Cordero, Maddalena Arigoni, Riccardo Panero, Elvio Gilberto Amparore, Susanna Donatelli, Raffaele A. Calogero |
Bioinform. | 2 |
| 2018 | Reproducible bioinformatics project: a community for reproducible bioinformatics analysis pipelinesabstractBACKGROUND: Reproducibility of a research is a key element in the modern science and it is mandatory for any industrial application. It represents the ability of replicating an experiment independently by the location and the operator. Therefore, a study can be considered reproducible only if all used data are available and the exploited computational analysis workflow is clearly described. However, today for reproducing a complex bioinformatics analysis, the raw data and the list of tools used in the workflow could be not enough to guarantee the reproducibility of the results obtained. Indeed, different releases of the same tools and/or of the system libraries (exploited by such tools) might lead to sneaky reproducibility issues. RESULTS: To address this challenge, we established the Reproducible Bioinformatics Project (RBP), which is a non-profit and open-source project, whose aim is to provide a schema and an infrastructure, based on docker images and R package, to provide reproducible results in Bioinformatics. One or more Docker images are then defined for a workflow (typically one for each task), while the workflow implementation is handled via R-functions embedded in a package available at github repository. Thus, a bioinformatician participating to the project has firstly to integrate her/his workflow modules into Docker image(s) exploiting an Ubuntu docker image developed ad hoc by RPB to make easier this task. Secondly, the workflow implementation must be realized in R according to an R-skeleton function made available by RPB to guarantee homogeneity and reusability among different RPB functions. Moreover she/he has to provide the R vignette explaining the package functionality together with an example dataset which can be used to improve the user confidence in the workflow utilization. CONCLUSIONS: Reproducible Bioinformatics Project provides a general schema and an infrastructure to distribute robust and reproducible workflows. Thus, it guarantees to final users the ability to repeat consistently any analysis independently by the used UNIX-like architecture. Neha Kulkarni, Luca Alessandrì, Riccardo Panero, Maddalena Arigoni, Martina Olivero, Giulio Ferrero, Francesca Cordero, Marco Beccuti, Raffaele A. Calogero |
BMC Bioinform. | 7 |
| 2017 | A mathematical model to study breast cancer growthabstractThe aim of this paper is (i)to study breast cancer growth by mean of a mathematical model describing cell population dynamics during cancer growth, and (ii)to use this model to reproduce and explain experimental data. We started from a linear model describing cancer subpopulations evolution based on the Cancer Stem Cell (CSC) theory, and we added feedback mechanisms from the cell populations to mimic micro-environment effects in cancer growth. In details, we hypothesized two feedback mechanisms and we studied their effects both separately and combined together. In this way we obtained three new models that we tuned using data derived by TUBO Cancer cell line and describing the evolution of the total cell population and the subpopulations over time. Finally, we exploited these three models to understand which combination of feedback mechanisms better describe the experimental data. Giorgia Chivassa, Chiara Fornari, Roberta Sirovich, Marzio Pennisi, Marco Beccuti, Francesca Cordero |
BIBM | 6 |
| 2017 | HashClone: a new tool to quantify the minimal residual disease in B-cell lymphoma from deep sequencing dataabstractBACKGROUND: Mantle Cell Lymphoma (MCL) is a B cell aggressive neoplasia accounting for about the 6% of all lymphomas. The most common molecular marker of clonality in MCL, as in other B lymphoproliferative disorders, is the ImmunoGlobulin Heavy chain (IGH) rearrangement, occurring in B-lymphocytes. The patient-specific IGH rearrangement is extensively used to monitor the Minimal Residual Disease (MRD) after treatment through the standardized Allele-Specific Oligonucleotides Quantitative Polymerase Chain Reaction based technique. Recently, several studies have suggested that the IGH monitoring through deep sequencing techniques can produce not only comparable results to Polymerase Chain Reaction-based methods, but also might overcome the classical technique in terms of feasibility and sensitivity. However, no standard bioinformatics tool is available at the moment for data analysis in this context. RESULTS: In this paper we present HashClone, an easy-to-use and reliable bioinformatics tool that provides B-cells clonality assessment and MRD monitoring over time analyzing data from Next-Generation Sequencing (NGS) technique. The HashClone strategy-based is composed of three steps: the first and second steps implement an alignment-free prediction method that identifies a set of putative clones belonging to the repertoire of the patient under study. In the third step the IGH variable region, diversity region, and joining region identification is obtained by the alignment of rearrangements with respect to the international ImMunoGenetics information system database. Moreover, a provided graphical user interface for HashClone execution and clonality visualization over time facilitate the tool use and the results interpretation. The HashClone performance was tested on the NGS data derived from MCL patients to assess the major B-cell clone in the diagnostic samples and to monitor the MRD in the real and artificial follow up samples. CONCLUSIONS: Our experiments show that in all the experimental settings, HashClone was able to correctly detect the major B-cell clones and to precisely follow them in several samples showing better accuracy than the state-of-art tool. Marco Beccuti, Elisa Genuardi, Greta Romano, Luigia Monitillo, Daniela Barbero, Mario Boccadoro, Marco Ladetto, Raffaele A. Calogero, Simone Ferrero, Francesca Cordero |
BMC Bioinform. | 10 |
| 2016 | Erratum to: 'NETTAB 2014: From high-throughput structural bioinformatics to integrative systems biology'abstractUnfortunately, the original version of this article [1] contained a few errors. The editorial department of BMC Bioinformatics would like to apologize and inform its readers of the following revisions.
The first sentence of the last paragraph “The manuscript “Weighted integration of multi-omic layers of conditions in genome-scale models” [12]” should be “The manuscript “Multiplex methods provide effective integration of multi-omic data in genome-scale models” [12]”.
In the reference list, the title of reference [12] “Weighted integration of multi-omic layers of conditions in genome-scale models” should be “Multiplex methods provide effective integration of multi-omic data in genome-scale models”.
Reference 7 has 2016;17 Suppl 3:S2 as its year, volume and supplement number and this should be 2016;17(Suppl 4):54 instead.
Reference 8 has 2016;17 Suppl 3:S3 as its year, volume and supplement number and this should be 2016;17(Suppl 4):57 instead.
Reference 9 has 2016;17 Suppl 3:S4 as its year, volume and supplement number and this should be 2016;17(Suppl 4):69 instead.
Reference 11 has 2016;17 Suppl 3:S5 as its year, volume and supplement number and this should be 2016;17(Suppl 4):61 instead.
Reference 12 has 2016;17 Suppl 3:S6 as its year, volume and supplement number and this should be 2016;17(Suppl 4):83 instead. Paolo Romano 0001, Francesca Cordero |
BMC Bioinform. | 2 |
| 2016 | NETTAB 2014: From high-throughput structural bioinformatics to integrative systems biologyabstractThe fourteenth NETTAB workshop, NETTAB 2014, was devoted to a range of disciplines going from structural bioinformatics, to proteomics and to integrative systems biology. The topics of the workshop were centred around bioinformatics methods, tools, applications, and perspectives for models, standards and management of high-throughput biological data, structural bioinformatics, functional proteomics, mass spectrometry, drug discovery, and systems biology.43 scientific contributions were presented at NETTAB 2014, including keynote, special guest and tutorial talks, oral communications, and posters. Full papers from some of the best contributions presented at the workshop were later submitted to a special Call for this Supplement.Here, we provide an overview of the workshop and introduce manuscripts that have been accepted for publication in this Supplement. Paolo Romano 0001, Francesca Cordero |
BMC Bioinform. | 2 |
| 2015 | Scrible: Ultra-Accurate Error-Correction of Pooled Sequenced Reads
Denise Duma, Francesca Cordero, Marco Beccuti, Gianfranco Ciardo, Timothy J. Close, Stefano Lonardi |
WABI | 2 |
| 2015 | Alternative splicing detection workflow needs a careful combination of sample prep and bioinformatics analysisabstractBACKGROUND: RNA-Seq provides remarkable power in the area of biomarkers discovery and disease characterization. Two crucial steps that affect RNA-Seq experiment results are Library Sample Preparation (LSP) and Bioinformatics Analysis (BA). This work describes an evaluation of the combined effect of LSP methods and BA tools in the detection of splice variants. RESULTS: Different LSPs (TruSeq unstranded/stranded, ScriptSeq, NuGEN) allowed the detection of a large common set of splice variants. However, each LSP also detected a small set of unique transcripts that are characterized by a low coverage and/or FPKM. This effect was particularly evident using the low input RNA NuGEN v2 protocol. A benchmark dataset, in which synthetic reads as well as reads generated from standard (Illumina TruSeq 100) and low input (NuGEN) LSPs were spiked-in was used to evaluate the effect of LSP on the statistical detection of alternative splicing events (AltDE). Statistical detection of AltDE was done using as prototypes for splice variant-quantification Cuffdiff2 and RSEM-EBSeq. As prototype for exon-level analysis DEXSeq was used. Exon-level analysis performed slightly better than splice variant-quantification approaches, although at most only 50% of the spiked-in transcripts was detected. The performances of both splice variant-quantification and exon-level analysis improved when raising the number of input reads. CONCLUSION: Data, derived from NuGEN v2, were not the ideal input for AltDE, especially when the exon-level approach was used. We observed that both splice variant-quantification and exon-level analysis performances were strongly dependent on the number of input reads. Moreover, the ribosomal RNA depletion protocol was less sensitive in detecting splicing variants, possibly due to the significant percentage of the reads mapping to non-coding transcripts. Matteo Carrara, Josephine Lum, Francesca Cordero, Marco Beccuti, Michael Poidinger, Susanna Donatelli, Raffaele A. Calogero, Francesca Zolezzi |
BMC Bioinform. | 3 |
| 2014 | Chimera: a Bioconductor package for secondary analysis of fusion productsabstractAbstract Summary: Chimera is a Bioconductor package that organizes, annotates, analyses and validates fusions reported by different fusion detection tools; current implementation can deal with output from bellerophontes, chimeraScan, deFuse, fusionCatcher, FusionFinder, FusionHunter, FusionMap, mapSplice, Rsubread, tophat-fusion and STAR. The core of Chimera is a fusion data structure that can store fusion events detected with any of the aforementioned tools. Fusions are then easily manipulated with standard R functions or through the set of functionalities specifically developed in Chimera with the aim of supporting the user in managing fusions and discriminating false-positive results. Availability and implementation: Chimera is implemented as a Bioconductor package in R. The package and the vignette can be downloaded at bioconductor.org. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Marco Beccuti, Matteo Carrara, Francesca Cordero, Fulvio Lazzarato, Susanna Donatelli, Francesca Nadalin, Alberto Policriti, Raffaele A. Calogero |
Bioinform. | 3 |
| 2014 | Leveraging additional knowledge to support coherent bicluster discovery in gene expression dataabstractThe increasing availability of gene expression data has encouraged the development of purposely-built intelligent data analysis techniques. Grouping genes characterized by similar expression patterns is a widely accepted – and often mandatory – analysis step. Despite the fact that a number of biclu stering methods have been developed to discover clusters of genes exhibiting a similar expression profile under a subgroup of experimental conditions, approaches driven by similarity measures based on expression profiles alone may lead to groups that are biologically meaningless. The integration of additional information, such as functional annotations, into biclustering algorithms can instead provide an effective support for identifying meaningful gene associations. In this paper we propose a new biclustering approach called Additional Information Driven Iterative Signature Algorithm, AID-ISA. It supports the extraction of biologically relevant biclusters by leveraging additional knowledge. We show that AID-ISA allows the discovery of coherent biclusters in baker's yeast and human gene expression data sets. Alessia Visconti, Francesca Cordero, Ruggero G. Pensa |
Intell. Data Anal. | 2 |
| 2013 | State of art fusion-finder algorithms are suitable to detect transcription-induced chimeras in normal tissues?abstractBACKGROUND: RNA-seq has the potential to discover genes created by chromosomal rearrangements. Fusion genes, also known as "chimeras", are formed by the breakage and re-joining of two different chromosomes. It is known that chimeras have been implicated in the development of cancer. Few publications in the past showed the presence of fusion events also in normal tissue, but with very limited overlaps between their results. More recently, two fusion genes in normal tissues were detected using both RNA-seq and protein data.Due to heterogeneous results in identifying chimeras in normal tissue, we decided to evaluate the efficacy of state of the art fusion finders in detecting chimeras in RNA-seq data from normal tissues. RESULTS: We compared the performance of six fusion-finder tools: FusionHunter, FusionMap, FusionFinder, MapSplice, deFuse and TopHat-fusion. To evaluate the sensitivity we used a synthetic dataset of fusion-products, called positive dataset; in these experiments FusionMap, FusionFinder, MapSplice, and TopHat-fusion are able to detect more than 78% of fusion genes. All tools were error prone with high variability among the tools, identifying some fusion genes not present in the synthetic dataset. To better investigate the false discovery chimera detection rate, synthetic datasets free of fusion-products, called negative datasets, were used. The negative datasets have different read lengths and quality scores, which allow detecting dependency of the tools on both these features. FusionMap, FusionFinder, mapSplice, deFuse and TopHat-fusion were error-prone. Only FusionHunter results were free of false positive. FusionMap gave the best compromise in terms of specificity in the negative dataset and of sensitivity in the positive dataset. CONCLUSIONS: We have observed a dependency of the tools on read length, quality score and on the number of reads supporting each chimera. Thus, it is important to carefully select the software on the basis of the structure of the RNA-seq data under analysis. Furthermore, the sensitivity of chimera detection tools does not seem to be sufficient to provide results consistent with those obtained in normal tissues on the basis of fusion events extracted from published data. Matteo Carrara, Marco Beccuti, Federica Cavallo, Susanna Donatelli, Fulvio Lazzarato, Francesca Cordero, Raffaele A. Calogero |
BMC Bioinform. | 6 |
| 2013 | Multi-level model for the investigation of oncoantigen-driven vaccination effectabstractBACKGROUND: Cancer stem cell theory suggests that cancers are derived by a population of cells named Cancer Stem Cells (CSCs) that are involved in the growth and in the progression of tumors, and lead to a hierarchical structure characterized by differentiated cell population. This cell heterogeneity affects the choice of cancer therapies, since many current cancer treatments have limited or no impact at all on CSC population, while they reveal a positive effect on the differentiated cell populations. RESULTS: In this paper we investigated the effect of vaccination on a cancer hierarchical structure through a multi-level model representing both population and molecular aspects. The population level is modeled by a system of Ordinary Differential Equations (ODEs) describing the cancer population's dynamics. The molecular level is modeled using the Petri Net (PN) formalism to detail part of the proliferation pathway. Moreover, we propose a new methodology which exploits the temporal behavior derived from the molecular level to parameterize the ODE system modeling populations. Using this multi-level model we studied the ErbB2-driven vaccination effect in breast cancer. CONCLUSIONS: We propose a multi-level model that describes the inter-dependencies between population and genetic levels, and that can be efficiently used to estimate the efficacy of drug and vaccine therapies in cancer models, given the availability of molecular data on the cancer driving force. Francesca Cordero, Marco Beccuti, Chiara Fornari, Stefania Lanzardo, Laura Conti, Federica Cavallo, Gianfranco Balbo, Raffaele A. Calogero |
BMC Bioinform. | 1 |
| 2013 | Guest EditorialabstractModeling and analyzing networks is a major emerging topic in different research areas, such as computational biology, social science, document retrieval and social web applications.By connecting objects, it is possible to obtain an intuitive and global view of the relationships among components of a complex system.Nowadays, scientific communities have access to huge volume of network-structured data, such as social networks, gene/proteins/metabolic networks, sensor networks, and peer-to-peer networks.Often, data is collected at different time points allowing capturing a dynamic trend of the observed network.Consequently, the time component plays a key role in the comprehension of the evolutionary behavior of the studied network (evolution of the network structure and/or of flows within the system).Time can help to determine the real causal relationships within, for instance, gene activations, link creation, and information flow.Handling such data is a major challenge for current research in machine learning and data mining, and it has led to the development of recent innovative techniques that consider complex/multi-level networks, time-evolving graphs, heterogeneous information (nodes and links), and requires scalable algorithms that are able to manage large-scale complex networks.This special issue is the follow-up of the Dynamic Networks and Knowledge Discovery workshop (DyNaK) 1 that has been held in conjunction to ECML-PKDD 2011 at Barcelona on September 24th 2011.The workshop was motivated by the interest of providing a meeting point for scientists with different backgrounds who are interested in the study of large-scale dynamic complex networks.The workshop has attracted 18 submissions out of which 9 papers has been accepted.The workshop has gathered more than 30 participants and was also the host of three highly appreciated invited keynotes and one industrial talk.Building on the success of the DyNaK workshop, an open call for papers has been issued for this special issue, focusing on the major topic discussed in the workshop: analyzing, modeling and mining large-scale real network.15 high quality papers have been received; each of which has been reviewed by three reviewers.Only 7 contributions were finally selected.These contributions show the vitality of the field: a broad panel of techniques are applied to modeling the dynamics of complex systems, using a wide set of formalisms ranging from descriptive rules to Probabilistic Real-Time Automata.Application fields are also wide: vision, opinion diffusion in social network, business process modeling and text mining.In Internal link prediction: a new approach for predicting links in bipartite graphs, Allali et al. present an algorithm for predicting internal link in bipartite graph.They address the problem of predicting Ruggero G. Pensa, Francesca Cordero, Céline Rouveirol, Rushed Kanawati |
Intell. Data Anal. | 2 |
| 2013 | Combinatorial Pooling Enables Selective Sequencing of the Barley Gene SpaceabstractFor the vast majority of species - including many economically or ecologically important organisms, progress in biological research is hampered due to the lack of a reference genome sequence. Despite recent advances in sequencing technologies, several factors still limit the availability of such a critical resource. At the same time, many research groups and international consortia have already produced BAC libraries and physical maps and now are in a position to proceed with the development of whole-genome sequences organized around a physical map anchored to a genetic map. We propose a BAC-by-BAC sequencing protocol that combines combinatorial pooling design and second-generation sequencing technology to efficiently approach denovo selective genome sequencing. We show that combinatorial pooling is a cost-effective and practical alternative to exhaustive DNA barcoding when preparing sequencing libraries for hundreds or thousands of DNA samples, such as in this case gene-bearing minimum-tiling-path BAC clones. The novelty of the protocol hinges on the computational ability to efficiently compare hundred millions of short reads and assign them to the correct BAC clones (deconvolution) so that the assembly can be carried out clone-by-clone. Experimental results on simulated data for the rice genome show that the deconvolution is very accurate, and the resulting BAC assemblies have high quality. Results on real data for a gene-rich subset of the barley genome confirm that the deconvolution is accurate and the BAC assemblies have good quality. While our method cannot provide the level of completeness that one would achieve with a comprehensive whole-genome sequencing project, we show that it is quite successful in reconstructing the gene sequences within BACs. In the case of plants such as barley, this level of sequence knowledge is sufficient to support critical end-point objectives such as map-based cloning and marker-assisted breeding. Stefano Lonardi, Denisa Duma, Matthew Alpert, Francesca Cordero, Marco Beccuti, Prasanna Bhat, Gianfranco Ciardo, Burair Alsaihati, Yaqin Ma, Steve Wanamaker, Josh Resnik, Serdar Bozdag, Ming-Cheng Luo, Timothy J. Close |
PLoS Comput. Biol. | 4 |
| 2011 | Simplification of a complex signal transduction model using invariants and flow equivalent servers
Francesca Cordero, András Horváth, Daniele Manini, Lucia Napione, Massimiliano De Pierro, Simona Pavan, Andrea Picco, Andrea Veglio, Matteo Sereno, Federico Bussolino, Gianfranco Balbo |
Theor. Comput. Sci. | 1 |
| 2010 | Gene Ontology Rewritten for Computing Gene Functional SimilarityabstractDiscovery biological organisation of the cell in modules network is a challenging task. Currently, approaches based on a controlled vocabulary, as Gene Ontology, to identify the function similarity among a pair of genes have developed. We present RGO: a rewriting of the GO aiming at obtaining a more compact and informative ontology, leading to closer biological regulated GO terms. The RGO will help researcher to easily identify the genes belonging to the same network module without the need of additional data. Alessia Visconti, Francesca Cordero, Marco Botta, Raffaele A. Calogero |
CISIS | 2 |
| 2009 | Genome-Wide Search for Splicing Defects Associated with Amyotrophic Lateral Sclerosis (ALS)abstractAmyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disease caused by the degeneration of motor neurons. Although the cause of ALS is unknown, mutations in the gene that produces the SOD1 enzyme are associated with some cases of familial ALS. SOD1 is a powerful antioxidant that protects the body from damage caused by superoxide, a toxic free radical. It has been proposed that defects in splicing of some mRNAs, induced by oxidative stress, can play a role in ALS pathogenesis. Alterations of splicing patterns have also been observed in ALS patients and in ALS murine models, suggesting that alterations in the splicing events can contribute to ALS progression. Using Exon 1.0 ST GeneChips, which allow the definition of alternative splicing events (ASEs) , the SH-SY5Y neuroblastoma cell line has been profiled after treatment with paraquat, which by inducing oxidative stress alters the patterns of alternative splicing. Furthermore, the same cell line stably transfected with wt and ALS mutant SOD has also been profiled. The integration of the two ALS models efficiently moderates ASE false discovery rate, one of the most critical issues in high-throughput ASEs detection. This approach allowed the identification of a total of 14 splicing events affecting respectively both internal coding exons and 5' UTR of known gene isoforms. Silvia C. Lenzken, Silvia Vivarelli, Francesca Zolezzi, Francesca Cordero, Cristina Della Beffa, Raffaele A. Calogero, Silvia Barabino |
CISIS | 4 |
| 2007 | oneChannelGUI: a graphical interface to Bioconductor tools, designed for life scientists who are not familiar with R languageabstractUNLABELLED: OneChannelGUI is an add-on Bioconductor package providing a new set of functions extending the capability of the affylmGUI package. This library provides a graphical interface (GUI) for Bioconductor libraries to be used for quality control, normalization, filtering, statistical validation and data mining for single channel microarrays. Affymetrix 3' expression (IVT) arrays as well as the new whole transcript expression arrays, i.e. gene/exon 1.0 ST, are actually implemented. oneChannelGUI is available for most platforms on which R runs, i.e. Windows and Unix-like machines. AVAILABILITY: http://www.bioconductor.org/packages/2.0/bioc/html/oneChannelGUI.html Remo Sanges, Francesca Cordero, Raffaele A. Calogero |
Bioinform. | 2 |
| 2005 | An integrated approach of immunogenomics and bioinformatics to identify new Tumor Associated Antigens (TAA) for mammary cancer immunological preventionabstractBACKGROUND: Neoplastic transformation is a multistep process in which distinct gene products of specific cell regulatory pathways are involved at each stage. Identification of overexpressed genes provides an unprecedented opportunity to address the immune system against antigens typical of defined stages of neoplastic transformation. HER-2/neu/ERBB2 (Her2) oncogene is a prototype of deregulated oncogenic protein kinase membrane receptors. Mice transgenic for rat Her2 (BALB-neuT mice) were studied to evaluate the stage in which vaccines can prevent the onset of Her2 driven mammary carcinomas. As Her2 is not overexpressed in all mammary carcinomas, definition of an additional set of tumor associated antigens (TAAs) expressed at defined stages by most breast carcinomas would allow a broader coverage of vaccination. To address this question, a meta-analysis was performed on two transcription profile studies to identify a set of new TAA targets to be used instead of or in conjunction with Her2. RESULTS: The five TAAs identified (Tes, Rcn2, Rnf4, Cradd, Galnt3) are those whose expression is linearly related to the tumor mass increase in BALB-neuT mammary glands. Moreover, they have a low expression in normal tissues and are generally expressed in human breast tumors, though at a lower level than Her2. CONCLUSION: Although the number of putative TAAs identified is limited, this pilot study suggests that meta-analysis of expression profiles produces results that could assist in the designing of pre-clinical immunopreventive vaccines. Federica Cavallo, Annalisa Astolfi, Manuela Iezzi, Francesca Cordero, Pierluigi Lollini, Guido Forni, Raffaele A. Calogero |
BMC Bioinform. | 4 |
| 2004 | RRE: a tool for the extraction of non-coding regions surrounding annotated genes from genomic datasetsabstractUNLABELLED: RRE allows the extraction of non-coding regions surrounding a coding sequence [i.e. gene upstream region, 5'-untranslated region (5'-UTR), introns, 3'-UTR, downstream region] from annotated genomic datasets available at NCBI. AVAILABILITY: RRE parser and web-based interface are accessible at http://www.bioinformatica.unito.it/bioinformatics/rre/rre.html Fulvio Lazzarato, Giuliana Franceschinis, Marco Botta, Francesca Cordero, Raffaele A. Calogero |
Bioinform. | 4 |