EDBT 2026 Demo / reviewers in the wild / expert
Dario Greco
dblp:32/8253
· DBLP profile ↗
26ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0001-9195-9003ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 11 since 2021Artificial intelligence and machine learning · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MUUMI: an R package for statistical and network-based meta-analysis for multi-omics data integrationabstractBACKGROUND: Disentangling physiopathological mechanisms of biological systems through high-level integration of omics data has become a standard procedure in life sciences. However, platform heterogeneity, batch effects, and the lack of unified methods for single- and multi-omics analyses represent relevant drawbacks that hinder the extrapolation of a meaningful biological interpretation. While statistical meta-analysis is widely used to integrate several omics datasets of the same type, it does not allow the integration of multi-modal data deriving from multi-omics experiments. Network science is at the forefront of systems biology, where the inference of molecular interactomes allowed the investigation of perturbed biological systems, by shedding light on the disrupted relationships that keep the homeostasis of complex systems. RESULTS: Here, we present MUUMI, an R package that unifies statistical meta-analysis and network-based omics data integration within a single analytical framework. MUUMI allows the identification of robust molecular signatures through multiple meta-analytical methods, inference and analysis of molecular interactomes and the integration of multiple omics layers through similarity network fusion. We demonstrate the functionalities of MUUMI by presenting two case studies in which we analysed (1) 17 transcriptomic datasets on idiopathic pulmonary fibrosis (IPF) from both microarray and RNA-Seq platforms and (2) multi-omics data of THP-1 macrophages exposed to different polarising stimuli. In both examples, MUUMI revealed biologically coherent signatures, underscoring its value in elucidating complex biological processes. CONCLUSIONS: MUUMI leverages omics data meta-analysis, integration and interpretation that implements both traditional and network-based approaches to unleash the power of multi-study datasets. Statistical and network-based approaches are integrated in a unique framework, allowing the user to derive robust and biologically meaningful results from different studies and datasets. MUUMI is an open-source package and is freely available at https://github.com/fhaive/muumi . Simo Iisakki Inkala, Michele Fratello, Giusy del Giudice, Giorgia Migliaccio, Angela Serra, Dario Greco, Antonio Federico |
BMC Bioinform. | 6 |
| 2025 | OpTiles: an R package for adaptive tiling and methylation variability profilingabstractSUMMARY: OpTiles is an R package that dynamically defines tiling windows based on the distribution of sequenced CpGs, addressing the limitations of traditional fixed-tiling approaches in targeted methylation datasets. By integrating CpG density with intra-region methylation variability, it provides a reliability metric and extended functionality for annotating, prioritizing, and interpreting complex methylation data. AVAILABILITY AND IMPLEMENTATION: OpTiles is implemented in R and source code is freely available at https://github.com/fhaive/OpTiles. Data are available on Zenodo at https://doi.org/10.5281/zenodo.16961292. Giorgia Migliaccio, Lena Möbus, Giusy del Giudice, Jack Morikka, Antonio Federico, Angela Serra, Dario Greco |
Bioinform. | 7 |
| 2023 | DREAM: an R package for druggability evaluation of human complex diseasesabstractMOTIVATION: De novo drug development is a long and expensive process that poses significant challenges from the design to the preclinical testing, making the introduction into the market slow and difficult. This limitation paved the way to the development of drug repurposing, which consists in the re-usage of already approved drugs, developed for other therapeutic indications. Although several efforts have been carried out in the last decade in order to achieve clinically relevant drug repurposing predictions, the amount of repurposed drugs that have been employed in actual pharmacological therapies is still limited. On one hand, mechanistic approaches, including profile-based and network-based methods, exploit the wealth of data about drug sensitivity and perturbational profiles as well as disease transcriptomics profiles. On the other hand, chemocentric approaches, including structure-based methods, take into consideration the intrinsic structural properties of the drugs and their molecular targets. The poor integration between mechanistic and chemocentric approaches is one of the main limiting factors behind the poor translatability of drug repurposing predictions into the clinics. RESULTS: In this work, we introduce DREAM, an R package aimed to integrate mechanistic and chemocentric approaches in a unified computational workflow. DREAM is devoted to the druggability evaluation of pathological conditions of interest, leveraging robust drug repurposing predictions. In addition, the user can derive optimized sets of drugs putatively suitable for combination therapy. In order to show the functionalities of the DREAM package, we report a case study on atopic dermatitis. AVAILABILITY AND IMPLEMENTATION: DREAM is freely available at https://github.com/fhaive/dream. The docker image of DREAM is available at: https://hub.docker.com/r/fhaive/dream. Antonio Federico, Michele Fratello, Alisa Pavel, Lena Möbus, Giusy del Giudice, Angela Serra, Dario Greco |
Bioinform. | 7 |
| 2023 | ESPERANTO: a GLP-field sEmi-SuPERvised toxicogenomics metadAta curatioN TOolabstractSUMMARY: Biological data repositories are an invaluable source of publicly available research evidence. Unfortunately, the lack of convergence of the scientific community on a common metadata annotation strategy has resulted in large amounts of data with low FAIRness (Findable, Accessible, Interoperable and Reusable). The possibility of generating high-quality insights from their integration relies on data curation, which is typically an error-prone process while also being expensive in terms of time and human labour. Here, we present ESPERANTO, an innovative framework that enables a standardized semi-supervised harmonization and integration of toxicogenomics metadata and increases their FAIRness in a Good Laboratory Practice-compliant fashion. The harmonization across metadata is guaranteed with the definition of an ad hoc vocabulary. The tool interface is designed to support the user in metadata harmonization in a user-friendly manner, regardless of the background and the type of expertise. AVAILABILITY AND IMPLEMENTATION: ESPERANTO and its user manual are freely available for academic purposes at https://github.com/fhaive/esperanto. The input and the results showcased in Supplementary File S1 are available at the same link. Emanuele Di Lieto, Angela Serra, Simo Iisakki Inkala, Laura Aliisa Saarimäki, Giusy del Giudice, Michele Fratello, Veera Hautanen, Maria Annala, Antonio Federico, Dario Greco |
Bioinform. | 10 |
| 2023 | KNeMAP: a network mapping approach for knowledge-driven comparison of transcriptomic profilesabstractMOTIVATION: Transcriptomic data can be used to describe the mechanism of action (MOA) of a chemical compound. However, omics data tend to be complex and prone to noise, making the comparison of different datasets challenging. Often, transcriptomic profiles are compared at the level of individual gene expression values, or sets of differentially expressed genes. Such approaches can suffer from underlying technical and biological variance, such as the biological system exposed on or the machine/method used to measure gene expression data, technical errors and further neglect the relationships between the genes. We propose a network mapping approach for knowledge-driven comparison of transcriptomic profiles (KNeMAP), which combines genes into similarity groups based on multiple levels of prior information, hence adding a higher-level view onto the individual gene view. When comparing KNeMAP with fold change (expression) based and deregulated gene set-based methods, KNeMAP was able to group compounds with higher accuracy with respect to prior information as well as is less prone to noise corrupted data. RESULT: We applied KNeMAP to analyze the Connectivity Map dataset, where the gene expression changes of three cell lines were analyzed after treatment with 676 drugs as well as the Fortino et al. dataset where two cell lines with 31 nanomaterials were analyzed. Although the expression profiles across the biological systems are highly different, KNeMAP was able to identify sets of compounds that induce similar molecular responses when exposed on the same biological system. AVAILABILITY AND IMPLEMENTATION: Relevant data and the KNeMAP function is available at: https://github.com/fhaive/KNeMAP and 10.5281/zenodo.7334711. Alisa Pavel, Giusy del Giudice, Michele Fratello, Leo Ghemtio, Antonio Di Lieto, Jari Yli-Kauhaluoma, Henri Xhaard, Antonio Federico, Angela Serra, Dario Greco |
Bioinform. | 10 |
| 2022 | Computationally prioritized drugs inhibit SARS-CoV-2 infection and syncytia formationabstractThe pharmacological arsenal against the COVID-19 pandemic is largely based on generic anti-inflammatory strategies or poorly scalable solutions. Moreover, as the ongoing vaccination campaign is rolling slower than wished, affordable and effective therapeutics are needed. To this end, there is increasing attention toward computational methods for drug repositioning and de novo drug design. Here, multiple data-driven computational approaches are systematically integrated to perform a virtual screening and prioritize candidate drugs for the treatment of COVID-19. From the list of prioritized drugs, a subset of representative candidates to test in human cells is selected. Two compounds, 7-hydroxystaurosporine and bafetinib, show synergistic antiviral effects in vitro and strongly inhibit viral-induced syncytia formation. Moreover, since existing drug repositioning methods provide limited usable information for de novo drug design, the relevant chemical substructures of the identified drugs are extracted to provide a chemical vocabulary that may help to design new effective drugs. Angela Serra, Michele Fratello, Antonio Federico, Ravi Ojha, Riccardo Provenzani, Ervin Tasnádi, Luca Cattelani, Giusy del Giudice, Pia Anneli Sofia Kinaret, Laura Aliisa Saarimäki, Alisa Pavel, Suvi Kuivanen, Vincenzo Cerullo, Olli Vapalahti, Peter Horváth, Antonio Di Lieto, Jari Yli-Kauhaluoma, Giuseppe Balistreri, Dario Greco |
Briefings Bioinform. | 19 |
| 2021 | Integrated network analysis reveals new genes suggesting COVID-19 chronic effects and treatmentabstractThe COVID-19 disease led to an unprecedented health emergency, still ongoing worldwide. Given the lack of a vaccine or a clear therapeutic strategy to counteract the infection as well as its secondary effects, there is currently a pressing need to generate new insights into the SARS-CoV-2 induced host response. Biomedical data can help to investigate new aspects of the COVID-19 pathogenesis, but source heterogeneity represents a major drawback and limitation. In this work, we applied data integration methods to develop a Unified Knowledge Space (UKS) and used it to identify a new set of genes associated with SARS-CoV-2 host response, both in vitro and in vivo. Functional analysis of these genes reveals possible long-term systemic effects of the infection, such as vascular remodelling and fibrosis. Finally, we identified a set of potentially relevant drugs targeting proteins involved in multiple steps of the host response to the virus. Alisa Pavel, Giusy del Giudice, Antonio Federico, Antonio Di Lieto, Pia Anneli Sofia Kinaret, Angela Serra, Dario Greco |
Briefings Bioinform. | 7 |
| 2021 | A systematic comparison of data- and knowledge-driven approaches to disease subtype discoveryabstractTypical clustering analysis for large-scale genomics data combines two unsupervised learning techniques: dimensionality reduction and clustering (DR-CL) methods. It has been demonstrated that transforming gene expression to pathway-level information can improve the robustness and interpretability of disease grouping results. This approach, referred to as biological knowledge-driven clustering (BK-CL) approach, is often neglected, due to a lack of tools enabling systematic comparisons with more established DR-based methods. Moreover, classic clustering metrics based on group separability tend to favor the DR-CL paradigm, which may increase the risk of identifying less actionable disease subtypes that have ambiguous biological and clinical explanations. Hence, there is a need for developing metrics that assess biological and clinical relevance. To facilitate the systematic analysis of BK-CL methods, we propose a computational protocol for quantitative analysis of clustering results derived from both DR-CL and BK-CL methods. Moreover, we propose a new BK-CL method that combines prior knowledge of disease relevant genes, network diffusion algorithms and gene set enrichment analysis to generate robust pathway-level information. Benchmarking studies were conducted to compare the grouping results from different DR-CL and BK-CL approaches with respect to standard clustering evaluation metrics, concordance with known subtypes, association with clinical outcomes and disease modules in co-expression networks of genes. No single approach dominated every metric, showing the importance multi-objective evaluation in clustering analysis. However, we demonstrated that, on gene expression data sets derived from TCGA samples, the BK-CL approach can find groupings that provide significant prognostic value in both breast and prostate cancers. Teemu J. Rintala, Antonio Federico, Leena Latonen, Dario Greco, Vittorio Fortino |
Briefings Bioinform. | 4 |
| 2021 | VOLTA: adVanced mOLecular neTwork AnalysisabstractMOTIVATION: Network analysis is a powerful approach to investigate biological systems. It is often applied to study gene co-expression patterns derived from transcriptomics experiments. Even though co-expression analysis is widely used, there is still a lack of tools that are open and customizable on the basis of different network types and analysis scenarios (e.g. through function accessibility), but are also suitable for novice users by providing complete analysis pipelines. RESULTS: We developed VOLTA, a Python package suited for complex co-expression network analysis. VOLTA is designed to allow users direct access to the individual functions, while they are also provided with complete analysis pipelines. Moreover, VOLTA offers when possible multiple algorithms applicable to each analytical step (e.g. multiple community detection or clustering algorithms are provided), hence providing the user with the possibility to perform analysis tailored to their needs. This makes VOLTA highly suitable for experienced users who wish to build their own analysis pipelines for a wide range of networks as well as for novice users for which a 'plug and play' system is provided. AVAILABILITY AND IMPLEMENTATION: The package and used data are available at GitHub: https://github.com/fhaive/VOLTA and 10.5281/zenodo.5171719. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alisa Pavel, Antonio Federico, Giusy del Giudice, Angela Serra, Dario Greco |
Bioinform. | 5 |
| 2021 | CpGmotifs: a tool to discover DNA motifs associated to CpG methylation eventsabstractBACKGROUND: The investigation of molecular alterations associated with the conservation and variation of DNA methylation in eukaryotes is gaining interest in the biomedical research community. Among the different determinants of methylation stability, the DNA composition of the CpG surrounding regions has been shown to have a crucial role in the maintenance and establishment of methylation statuses. This aspect has been previously characterized in a quantitative manner by inspecting the nucleotidic composition in the region. Research in this field still lacks a qualitative perspective, linked to the identification of certain sequences (or DNA motifs) related to particular DNA methylation phenomena. RESULTS: Here we present a novel computational strategy based on short DNA motif discovery in order to characterize sequence patterns related to aberrant CpG methylation events. We provide our framework as a user-friendly, shiny-based application, CpGmotifs, to easily retrieve and characterize DNA patterns related to CpG methylation in the human genome. Our tool supports the functional interpretation of deregulated methylation events by predicting transcription factors binding sites (TFBS) encompassing the identified motifs. CONCLUSIONS: CpGmotifs is an open source software. Its source code is available on GitHub https://github.com/Greco-Lab/CpGmotifs and a ready-to-use docker image is provided on DockerHub at https://hub.docker.com/r/grecolab/cpgmotifs . Giovanni Scala, Antonio Federico, Dario Greco |
BMC Bioinform. | 3 |
| 2021 | Clustering based approach for population level identification of condition-associated T-cell receptor β-chain CDR3 sequencesabstractBACKGROUND: Deep immune receptor sequencing, RepSeq, provides unprecedented opportunities for identifying and studying condition-associated T-cell clonotypes, represented by T-cell receptor (TCR) CDR3 sequences. However, due to the immense diversity of the immune repertoire, identification of condition relevant TCR CDR3s from total repertoires has mostly been limited to either "public" CDR3 sequences or to comparisons of CDR3 frequencies observed in a single individual. A methodology for the identification of condition-associated TCR CDR3s by direct population level comparison of RepSeq samples is currently lacking. RESULTS: We present a method for direct population level comparison of RepSeq samples using immune repertoire sub-units (or sub-repertoires) that are shared across individuals. The method first performs unsupervised clustering of CDR3s within each sample. It then finds matching clusters across samples, called immune sub-repertoires, and performs statistical differential abundance testing at the level of the identified sub-repertoires. It finally ranks CDR3s in differentially abundant sub-repertoires for relevance to the condition. We applied the method on total TCR CDR3β RepSeq datasets of celiac disease patients, as well as on public datasets of yellow fever vaccination. The method successfully identified celiac disease associated CDR3β sequences, as evidenced by considerable agreement of TRBV-gene and positional amino acid usage patterns in the detected CDR3β sequences with previously known CDR3βs specific to gluten in celiac disease. It also successfully recovered significantly high numbers of previously known CDR3β sequences relevant to each condition than would be expected by chance. CONCLUSION: We conclude that immune sub-repertoires of similar immuno-genomic features shared across unrelated individuals can serve as viable units of immune repertoire comparison, serving as proxy for identification of condition-associated CDR3s. Dawit A. Yohannes, Katri Kaukinen, Kalle Kurppa, Päivi Saavalainen, Dario Greco |
BMC Bioinform. | 5 |
| 2020 | Feature set optimization in biomarker discovery from genome-scale dataabstractMOTIVATION: Omics technologies have the potential to facilitate the discovery of new biomarkers. However, only few omics-derived biomarkers have been successfully translated into clinical applications to date. Feature selection is a crucial step in this process that identifies small sets of features with high predictive power. Models consisting of a limited number of features are not only more robust in analytical terms, but also ensure cost effectiveness and clinical translatability of new biomarker panels. Here we introduce GARBO, a novel multi-island adaptive genetic algorithm to simultaneously optimize accuracy and set size in omics-driven biomarker discovery problems. RESULTS: Compared to existing methods, GARBO enables the identification of biomarker sets that best optimize the trade-off between classification accuracy and number of biomarkers. We tested GARBO and six alternative selection methods with two high relevant topics in precision medicine: cancer patient stratification and drug sensitivity prediction. We found multivariate biomarker models from different omics data types such as mRNA, miRNA, copy number variation, mutation and DNA methylation. The top performing models were evaluated by using two different strategies: the Pareto-based selection, and the weighted sum between accuracy and set size (w = 0.5). Pareto-based preferences show the ability of the proposed algorithm to search minimal subsets of relevant features that can be used to model accurate random forest-based classification systems. Moreover, GARBO systematically identified, on larger omics data types, such as gene expression and DNA methylation, biomarker panels exhibiting higher classification accuracy or employing a number of features much lower than those discovered with other methods. These results were confirmed on independent datasets. AVAILABILITY AND IMPLEMENTATION: github.com/Greco-Lab/GARBO. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vittorio Fortino, Giovanni Scala, Dario Greco |
Bioinform. | 3 |
| 2020 | MaNGA: a novel multi-niche multi-objective genetic algorithm for QSAR modellingabstractSUMMARY: Quantitative structure-activity relationship (QSAR) modelling is currently used in multiple fields to relate structural properties of compounds to their biological activities. This technique is also used for drug design purposes with the aim of predicting parameters that determine drug behaviour. To this end, a sophisticated process, involving various analytical steps concatenated in series, is employed to identify and fine-tune the optimal set of predictors from a large dataset of molecular descriptors (MDs). The search of the optimal model requires to optimize multiple objectives at the same time, as the aim is to obtain the minimal set of features that maximizes the goodness of fit and the applicability domain (AD). Hence, a multi-objective optimization strategy, improving multiple parameters in parallel, can be applied. Here we propose a new multi-niche multi-objective genetic algorithm that simultaneously enables stable feature selection as well as obtaining robust and validated regression models with maximized AD. We benchmarked our method on two simulated datasets. Moreover, we analyzed an aquatic acute toxicity dataset and compared the performances of single- and multi-objective fitness functions on different regression models. Our results show that our multi-objective algorithm is a valid alternative to classical QSAR modelling strategy, for continuous response values, since it automatically finds the model with the best compromise between statistical robustness, predictive performance, widest AD, and the smallest number of MDs. AVAILABILITY AND IMPLEMENTATION: The python implementation of MaNGA is available at https://github.com/Greco-Lab/MaNGA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Angela Serra, Serli Önlü, Paola Festa, Vittorio Fortino, Dario Greco |
Bioinform. | 5 |
| 2020 | BMDx: a graphical Shiny application to perform Benchmark Dose analysis for transcriptomics dataabstractMOTIVATION: The analysis of dose-dependent effects on the gene expression is gaining attention in the field of toxicogenomics. Currently available computational methods are usually limited to specific omics platforms or biological annotations and are able to analyse only one experiment at a time. RESULTS: We developed the software BMDx with a graphical user interface for the Benchmark Dose (BMD) analysis of transcriptomics data. We implemented an approach based on the fitting of multiple models and the selection of the optimal model based on the Akaike Information Criterion. The BMDx tool takes as an input a gene expression matrix and a phenotype table, computes the BMD, its related values, and IC50/EC50 estimations. It reports interactive tables and plots that the user can investigate for further details of the fitting, dose effects and functional enrichment. BMDx allows a fast and convenient comparison of the BMD values of a transcriptomics experiment at different time points and an effortless way to interpret the results. Furthermore, BMDx allows to analyse and to compare multiple experiments at once. AVAILABILITY AND IMPLEMENTATION: BMDx is implemented as an R/Shiny software and is available at https://github.com/Greco-Lab/BMDx/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Angela Serra, Laura Aliisa Saarimäki, Michele Fratello, Veer Singh Marwah, Dario Greco |
Bioinform. | 5 |
| 2019 | FunMappOne: a tool to hierarchically organize and visually navigate functional gene annotations in multiple experimentsabstractBACKGROUND: Functional annotation of genes is an essential step in omics data analysis. Multiple databases and methods are currently available to summarize the functions of sets of genes into higher level representations, such as ontologies and molecular pathways. Annotating results from omics experiments into functional categories is essential not only to understand the underlying regulatory dynamics but also to compare multiple experimental conditions at a higher level of abstraction. Several tools are already available to the community to represent and compare functional profiles of omics experiments. However, when the number of experiments and/or enriched functional terms is high, it becomes difficult to interpret the results even when graphically represented. Therefore, there is currently a need for interactive and user-friendly tools to graphically navigate and further summarize annotations in order to facilitate results interpretation also when the dimensionality is high. RESULTS: We developed an approach that exploits the intrinsic hierarchical structure of several functional annotations to summarize the results obtained through enrichment analyses to higher levels of interpretation and to map gene related information at each summarized level. We built a user-friendly graphical interface that allows to visualize the functional annotations of one or multiple experiments at once. The tool is implemented as a R-Shiny application called FunMappOne and is available at https://github.com/grecolab/FunMappOne . CONCLUSION: FunMappOne is a R-shiny graphical tool that takes in input multiple lists of human or mouse genes, optionally along with their related modification magnitudes, computes the enriched annotations from Gene Ontology, Kyoto Encyclopedia of Genes and Genomes, or Reactome databases, and reports interactive maps of functional terms and pathways organized in rational groups. FunMappOne allows a fast and convenient comparison of multiple experiments and an easy way to interpret results. Giovanni Scala, Angela Serra, Veer Singh Marwah, Laura Aliisa Saarimäki, Dario Greco |
BMC Bioinform. | 5 |
| 2018 | INfORM: Inference of NetwOrk Response ModulesabstractSummary: Detecting and interpreting responsive modules from gene expression data by using network-based approaches is a common but laborious task. It often requires the application of several computational methods implemented in different software packages, forcing biologists to compile complex analytical pipelines. Here we introduce INfORM (Inference of NetwOrk Response Modules), an R shiny application that enables non-expert users to detect, evaluate and select gene modules with high statistical and biological significance. INfORM is a comprehensive tool for the identification of biologically meaningful response modules from consensus gene networks inferred by using multiple algorithms. It is accessible through an intuitive graphical user interface allowing for a level of abstraction from the computational steps. Availability and implementation: INfORM is freely available for academic use at https://github.com/Greco-Lab/INfORM. Supplementary information: Supplementary data are available at Bioinformatics online. Veer Singh Marwah, Pia Anneli Sofia Kinaret, Angela Serra, Giovanni Scala, Antti Lauerma, Vittorio Fortino, Dario Greco |
Bioinform. | 7 |
| 2018 | IntEREst: intron-exon retention estimatorabstractBACKGROUND: In-depth study of the intron retention levels of transcripts provide insights on the mechanisms regulating pre-mRNA splicing efficiency. Additionally, detailed analysis of retained introns can link these introns to post-transcriptional regulation or identify aberrant splicing events in human diseases. RESULTS: We present IntEREst, Intron-Exon Retention Estimator, an R package that supports rigorous analysis of non-annotated intron retention events (in addition to the ones annotated by RefSeq or similar databases), and support intra-sample in addition to inter-sample comparisons. It accepts binary sequence alignment/map (.bam) files as input and determines genome-wide estimates of intron retention or exon-exon junction levels. Moreover, it includes functions for comparing subsets of user-defined introns (e.g. U12-type vs U2-type) and its plotting functions allow visualization of the distribution of the retention levels of the introns. Statistical methods are adapted from the DESeq2, edgeR and DEXSeq R packages to extract the significantly more or less retained introns. Analyses can be performed either sequentially (on single core) or in parallel (on multiple cores). We used IntEREst to investigate the U12- and U2-type intron retention in human and plant RNAseq dataset with defects in the U12-dependent spliceosome due to mutations in the ZRSR2 component of this spliceosome. Additionally, we compared the retained introns discovered by IntEREst with that of other methods and studies. CONCLUSION: IntEREst is an R package for Intron retention and exon-exon junction levels analysis of RNA-seq data. Both the human and plant analyses show that the U12-type introns are retained at higher level compared to the U2-type introns already in the control samples, but the retention is exacerbated in patient or plant samples carrying a mutated ZRSR2 gene. Intron retention events caused by ZRSR2 mutation that we discovered using IntEREst (DESeq2 based function) show considerable overlap with the retained introns discovered by other methods (e.g. IRFinder and edgeR based function of IntEREst). Our results indicate that increase in both the number of biological replicates and the depth of sequencing library promote the discovery of retained introns, but the effect of library size gradually decreases with more than 35 million reads mapped to the introns. Ali Oghabian, Dario Greco, Mikko J. Frilander |
BMC Bioinform. | 2 |
| 2016 | Time aware knowledge extraction to analyze nanosafety cluster scientific activitiesabstractWith the rapid development ot biomedical sciences, a growing amount of papers reporting new scientific findings are published and indexed in different unstructured biomedical data sources. In order to really appreciate and effectively benefit from the availability of this amount of data there is an urgent need to support the deployment of intelligent information services, such as: temporal trends and group detection, expert finding, review experts, link prediction, and so on. This need is even more stressed if we analyze dissemination activity of emerging scientific communities that are working on specific research topics in the field of biomedical science. Motivated by the fact that nanotechnologies are one of the key enabling technologies nowadays, in this paper we instantiate and contextualize the Time Aware Knowledge Extraction (TAKE) methodology, introduced in previous work, as a tool to analyze the activities of the nano-safety scientific community coordinated by the EU NanoSafety Cluster (NSC). This methodology enables us to extract timed association rules. To validate and give evidence of the goodness of these rules, a summary of the so obtained results of the analysis is provided identifying distinguishing features, and detecting emerging collaboration among the NSC's Working Groups and their members over the timeline. Carmen De Maio, Mimmo Parente, Giuseppe Fenza, Dario Greco |
CEC | 4 |
| 2016 | Data integration in genomics and systems biologyabstractMulti-view learning is the branch of machine learning that deals with multi modal data, i.e. with patterns represented by different sets of features. The fast spread of this learning technique is motivated by the continuing increase of real applications based on multi-view data. For example, in bioinformatics multiple experiments can be available (mRNA, miRNA and protein expression, genome wide association studies (GWAS) and others) for a set of samples. In bioinformatics multi-view approaches are useful since heterogeneous genome-wide data sources capture information on different aspects of complex biological systems. Each view provides a distinct facet of the same domain, encoding different biologically-relevant patterns. The integration of such views can provide a richer model of the underlying system than those produced by a single view alone. This paper provides a review of the literature with respect to bioinformatics, with the purpose to understand the principles and operation modes of the existing methods and their possible applications. In order to organize the proposed methods in literature and to find similarities between them, these approaches are organized according to three categories: the type of data used in the papers, the statistical problem and the stage of integration. Angela Serra, Michele Fratello, Dario Greco, Roberto Tagliaferri |
CEC | 3 |
| 2016 | CONDOP: an R package for CONdition-Dependent Operon PredictionsabstractThe use of high-throughput RNA sequencing to predict dynamic operon structures in prokaryotic genomes has recently gained popularity in bioinformatics. We provide the R implementation of a novel method that uses transcriptomic features extracted from RNA-seq transcriptome profiles to develop ensemble classifiers for condition-dependent operon predictions. The CONDOP package provides a deeper insight into RNA-seq data analysis and allows scientists to highlight the operon organization in the context of transcriptional regulation with a few lines of code. AVAILABILITY AND IMPLEMENTATION: CONDOP is implemented in R and is freely available at CRAN. CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Vittorio Fortino, Roberto Tagliaferri, Dario Greco |
Bioinform. | 3 |
| 2015 | Impact of different metrics on multi-view clusteringabstractClustering of patients allows to find groups of subjects with similar characteristics. This categorization can facilitate diagnosis, treatment decision and prognosis prediction. Heterogeneous genome-wide data sources capture different biological aspects that can be integrated in order to better categorize the patients. Clustering methods work by comparing how patients are similar or dissimilar in a suitable similarity space. While several clustering methods have been proposed, there is no systematic comparative study concerning the impact of similarity metrics on the cluster quality. We compared seven popular similarity measures (Pearson, Spearman and Kendall Correlations; Euclidean, Canberra, Minkowski and Manhattan Distances) in conjunction with two classical single-view clustering algorithms and a late integration approach (partitioning around medoids, hierarchical clustering and matrix factorization approaches), on high dimensional multi-view cancer data coming from the TCGA repository. Performance was measured against tumour subcategories classification. Only Euclidean and Minkowski distances showed similar results in terms of clustering similarity indexes. On the other hand, an absolute best similarity measure did not emerge in terms of misclassification, but it strongly depends on the data. Angela Serra, Dario Greco, Roberto Tagliaferri |
IJCNN | 2 |
| 2015 | BACA: bubble chArt to compare annotationsabstractBACKGROUND: DAVID is the most popular tool for interpreting large lists of gene/proteins classically produced in high-throughput experiments. However, the use of DAVID website becomes difficult when analyzing multiple gene lists, for it does not provide an adequate visualization tool to show/compare multiple enrichment results in a concise and informative manner. RESULT: We implemented a new R-based graphical tool, BACA (Bubble chArt to Compare Annotations), which uses the DAVID web service for cross-comparing enrichment analysis results derived from multiple large gene lists. BACA is implemented in R and is freely available at the CRAN repository ( http://cran.r-project.org/web/packages/BACA/ ). CONCLUSION: The package BACA allows R users to combine multiple annotation charts into one output graph by passing DAVID website. Vittorio Fortino, Harri Alenius, Dario Greco |
BMC Bioinform. | 3 |
| 2015 | A multi-view genomic data simulatorabstractBACKGROUND: OMICs technologies allow to assay the state of a large number of different features (e.g., mRNA expression, miRNA expression, copy number variation, DNA methylation, etc.) from the same samples. The objective of these experiments is usually to find a reduced set of significant features, which can be used to differentiate the conditions assayed. In terms of development of novel feature selection computational methods, this task is challenging for the lack of fully annotated biological datasets to be used for benchmarking. A possible way to tackle this problem is generating appropriate synthetic datasets, whose composition and behaviour are fully controlled and known a priori. RESULTS: Here we propose a novel method centred on the generation of networks of interactions among different biological molecules, especially involved in regulating gene expression. Synthetic datasets are obtained from ordinary differential equations based models with known parameters. Our results show that the generated datasets are well mimicking the behaviour of real data, for popular data analysis methods are able to selectively identify existing interactions. CONCLUSIONS: The proposed method can be used in conjunction to real biological datasets in the assessment of data mining techniques. The main strength of this method consists in the full control on the simulated data while retaining coherence with the real biological processes. The R package MVBioDataSim is freely available to the scientific community at http://neuronelab.unisa.it/?p=1722. Michele Fratello, Angela Serra, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco |
BMC Bioinform. | 6 |
| 2015 | MVDA: a multi-view genomic data integration methodologyabstractBACKGROUND: Multiple high-throughput molecular profiling by omics technologies can be collected for the same individuals. Combining these data, rather than exploiting them separately, can significantly increase the power of clinically relevant patients subclassifications. RESULTS: We propose a multi-view approach in which the information from different data layers (views) is integrated at the levels of the results of each single view clustering iterations. It works by factorizing the membership matrices in a late integration manner. We evaluated the effectiveness and the performance of our method on six multi-view cancer datasets. In all the cases, we found patient sub-classes with statistical significance, identifying novel sub-groups previously not emphasized in literature. Our method performed better as compared to other multi-view clustering algorithms and, unlike other existing methods, it is able to quantify the contribution of single views on the final results. CONCLUSION: Our observations suggest that integration of prior information with genomic features in the subtyping analysis is an effective strategy in identifying disease subgroups. The methodology is implemented in R and the source code is available online at http://neuronelab.unisa.it/a-multi-view-genomic-data-integration-methodology/ . Angela Serra, Michele Fratello, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco |
BMC Bioinform. | 6 |
| 2014 | Transcriptome dynamics-based operon prediction in prokaryotesabstractBACKGROUND: Inferring operon maps is crucial to understanding the regulatory networks of prokaryotic genomes. Recently, RNA-seq based transcriptome studies revealed that in many bacterial species the operon structure vary with the change of environmental conditions. Therefore, new computational solutions that use both static and dynamic data are necessary to create condition specific operon predictions. RESULTS: In this work, we propose a novel classification method that integrates RNA-seq based transcriptome profiles with genomic sequence features to accurately identify the operons that are expressed under a measured condition. The classifiers are trained on a small set of confirmed operons and then used to classify the remaining gene pairs of the organism studied. Finally, by linking consecutive gene pairs classified as operons, our computational approach produces condition-dependent operon maps. We evaluated our approach on various RNA-seq expression profiles of the bacteria Haemophilus somni, Porphyromonas gingivalis, Escherichia coli and Salmonella enterica. Our results demonstrate that, using features depending on both transcriptome dynamics and genome sequence characteristics, we can identify operon pairs with high accuracy. Moreover, the combination of DNA sequence and expression data results in more accurate predictions than each one alone. CONCLUSION: We present a computational strategy for the comprehensive analysis of condition-dependent operon maps in prokaryotes. Our method can be used to generate condition specific operon maps of many bacterial organisms for which high-resolution transcriptome data is available. Vittorio Fortino, Olli-Pekka Smolander, Petri Auvinen, Roberto Tagliaferri, Dario Greco |
BMC Bioinform. | 5 |
| 2010 | Bayesian integrated modeling of expression data: a case study on RhoGabstractBACKGROUND: DNA microarrays provide an efficient method for measuring activity of genes in parallel and even covering all the known transcripts of an organism on a single array. This has to be balanced against that analyzing data emerging from microarrays involves several consecutive steps, and each of them is a potential source of errors. Errors tend to accumulate when moving from the lower level towards the higher level analyses because of the sequential nature. Eliminating such errors does not seem feasible without completely changing the technologies, but one should nevertheless try to meet the goal of being able to realistically assess degree of the uncertainties that are involved when drawing the final conclusions from such analyses. RESULTS: We present a Bayesian hierarchical model for finding differentially expressed genes between two experimental conditions, proposing an integrated statistical approach where correcting signal saturation, systematic array effects, dye effects, and finding differentially expressed genes, are all modeled jointly. The integration allows all these components, and also the associated errors, to be considered simultaneously. The inference is based on full posterior distribution of gene expression indices and on quantities derived from them rather than on point estimates. The model was applied and tested on two different datasets. CONCLUSIONS: The method presents a way of integrating various steps of microarray analysis into a single joint analysis, and thereby enables extracting information on differential expression in a manner, which properly accounts for various sources of potential error in the process. Rashi Gupta, Dario Greco, Petri Auvinen, Elja Arjas |
BMC Bioinform. | 2 |