Michele Fratello

dblp:162/7475 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-3997-2339ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3
YearPublicationVenuePosition
2026 MUUMI: an R package for statistical and network-based meta-analysis for multi-omics data integration
abstract
BACKGROUND: Disentangling physiopathological mechanisms of biological systems through high-level integration of omics data has become a standard procedure in life sciences. However, platform heterogeneity, batch effects, and the lack of unified methods for single- and multi-omics analyses represent relevant drawbacks that hinder the extrapolation of a meaningful biological interpretation. While statistical meta-analysis is widely used to integrate several omics datasets of the same type, it does not allow the integration of multi-modal data deriving from multi-omics experiments. Network science is at the forefront of systems biology, where the inference of molecular interactomes allowed the investigation of perturbed biological systems, by shedding light on the disrupted relationships that keep the homeostasis of complex systems. RESULTS: Here, we present MUUMI, an R package that unifies statistical meta-analysis and network-based omics data integration within a single analytical framework. MUUMI allows the identification of robust molecular signatures through multiple meta-analytical methods, inference and analysis of molecular interactomes and the integration of multiple omics layers through similarity network fusion. We demonstrate the functionalities of MUUMI by presenting two case studies in which we analysed (1) 17 transcriptomic datasets on idiopathic pulmonary fibrosis (IPF) from both microarray and RNA-Seq platforms and (2) multi-omics data of THP-1 macrophages exposed to different polarising stimuli. In both examples, MUUMI revealed biologically coherent signatures, underscoring its value in elucidating complex biological processes. CONCLUSIONS: MUUMI leverages omics data meta-analysis, integration and interpretation that implements both traditional and network-based approaches to unleash the power of multi-study datasets. Statistical and network-based approaches are integrated in a unique framework, allowing the user to derive robust and biologically meaningful results from different studies and datasets. MUUMI is an open-source package and is freely available at https://github.com/fhaive/muumi .
Simo Iisakki Inkala, Michele Fratello, Giusy del Giudice, Giorgia Migliaccio, Angela Serra, Dario Greco, Antonio Federico
BMC Bioinform.2
2023 DREAM: an R package for druggability evaluation of human complex diseases
abstract
MOTIVATION: De novo drug development is a long and expensive process that poses significant challenges from the design to the preclinical testing, making the introduction into the market slow and difficult. This limitation paved the way to the development of drug repurposing, which consists in the re-usage of already approved drugs, developed for other therapeutic indications. Although several efforts have been carried out in the last decade in order to achieve clinically relevant drug repurposing predictions, the amount of repurposed drugs that have been employed in actual pharmacological therapies is still limited. On one hand, mechanistic approaches, including profile-based and network-based methods, exploit the wealth of data about drug sensitivity and perturbational profiles as well as disease transcriptomics profiles. On the other hand, chemocentric approaches, including structure-based methods, take into consideration the intrinsic structural properties of the drugs and their molecular targets. The poor integration between mechanistic and chemocentric approaches is one of the main limiting factors behind the poor translatability of drug repurposing predictions into the clinics. RESULTS: In this work, we introduce DREAM, an R package aimed to integrate mechanistic and chemocentric approaches in a unified computational workflow. DREAM is devoted to the druggability evaluation of pathological conditions of interest, leveraging robust drug repurposing predictions. In addition, the user can derive optimized sets of drugs putatively suitable for combination therapy. In order to show the functionalities of the DREAM package, we report a case study on atopic dermatitis. AVAILABILITY AND IMPLEMENTATION: DREAM is freely available at https://github.com/fhaive/dream. The docker image of DREAM is available at: https://hub.docker.com/r/fhaive/dream.
Antonio Federico, Michele Fratello, Alisa Pavel, Lena Möbus, Giusy del Giudice, Angela Serra, Dario Greco
Bioinform.2
2023 ESPERANTO: a GLP-field sEmi-SuPERvised toxicogenomics metadAta curatioN TOol
abstract
SUMMARY: Biological data repositories are an invaluable source of publicly available research evidence. Unfortunately, the lack of convergence of the scientific community on a common metadata annotation strategy has resulted in large amounts of data with low FAIRness (Findable, Accessible, Interoperable and Reusable). The possibility of generating high-quality insights from their integration relies on data curation, which is typically an error-prone process while also being expensive in terms of time and human labour. Here, we present ESPERANTO, an innovative framework that enables a standardized semi-supervised harmonization and integration of toxicogenomics metadata and increases their FAIRness in a Good Laboratory Practice-compliant fashion. The harmonization across metadata is guaranteed with the definition of an ad hoc vocabulary. The tool interface is designed to support the user in metadata harmonization in a user-friendly manner, regardless of the background and the type of expertise. AVAILABILITY AND IMPLEMENTATION: ESPERANTO and its user manual are freely available for academic purposes at https://github.com/fhaive/esperanto. The input and the results showcased in Supplementary File S1 are available at the same link.
Emanuele Di Lieto, Angela Serra, Simo Iisakki Inkala, Laura Aliisa Saarimäki, Giusy del Giudice, Michele Fratello, Veera Hautanen, Maria Annala, Antonio Federico, Dario Greco
Bioinform.6
2023 KNeMAP: a network mapping approach for knowledge-driven comparison of transcriptomic profiles
abstract
MOTIVATION: Transcriptomic data can be used to describe the mechanism of action (MOA) of a chemical compound. However, omics data tend to be complex and prone to noise, making the comparison of different datasets challenging. Often, transcriptomic profiles are compared at the level of individual gene expression values, or sets of differentially expressed genes. Such approaches can suffer from underlying technical and biological variance, such as the biological system exposed on or the machine/method used to measure gene expression data, technical errors and further neglect the relationships between the genes. We propose a network mapping approach for knowledge-driven comparison of transcriptomic profiles (KNeMAP), which combines genes into similarity groups based on multiple levels of prior information, hence adding a higher-level view onto the individual gene view. When comparing KNeMAP with fold change (expression) based and deregulated gene set-based methods, KNeMAP was able to group compounds with higher accuracy with respect to prior information as well as is less prone to noise corrupted data. RESULT: We applied KNeMAP to analyze the Connectivity Map dataset, where the gene expression changes of three cell lines were analyzed after treatment with 676 drugs as well as the Fortino et al. dataset where two cell lines with 31 nanomaterials were analyzed. Although the expression profiles across the biological systems are highly different, KNeMAP was able to identify sets of compounds that induce similar molecular responses when exposed on the same biological system. AVAILABILITY AND IMPLEMENTATION: Relevant data and the KNeMAP function is available at: https://github.com/fhaive/KNeMAP and 10.5281/zenodo.7334711.
Alisa Pavel, Giusy del Giudice, Michele Fratello, Leo Ghemtio, Antonio Di Lieto, Jari Yli-Kauhaluoma, Henri Xhaard, Antonio Federico, Angela Serra, Dario Greco
Bioinform.3
2022 Computationally prioritized drugs inhibit SARS-CoV-2 infection and syncytia formation
abstract
The pharmacological arsenal against the COVID-19 pandemic is largely based on generic anti-inflammatory strategies or poorly scalable solutions. Moreover, as the ongoing vaccination campaign is rolling slower than wished, affordable and effective therapeutics are needed. To this end, there is increasing attention toward computational methods for drug repositioning and de novo drug design. Here, multiple data-driven computational approaches are systematically integrated to perform a virtual screening and prioritize candidate drugs for the treatment of COVID-19. From the list of prioritized drugs, a subset of representative candidates to test in human cells is selected. Two compounds, 7-hydroxystaurosporine and bafetinib, show synergistic antiviral effects in vitro and strongly inhibit viral-induced syncytia formation. Moreover, since existing drug repositioning methods provide limited usable information for de novo drug design, the relevant chemical substructures of the identified drugs are extracted to provide a chemical vocabulary that may help to design new effective drugs.
Angela Serra, Michele Fratello, Antonio Federico, Ravi Ojha, Riccardo Provenzani, Ervin Tasnádi, Luca Cattelani, Giusy del Giudice, Pia Anneli Sofia Kinaret, Laura Aliisa Saarimäki, Alisa Pavel, Suvi Kuivanen, Vincenzo Cerullo, Olli Vapalahti, Peter Horváth, Antonio Di Lieto, Jari Yli-Kauhaluoma, Giuseppe Balistreri, Dario Greco
Briefings Bioinform.2
2020 BMDx: a graphical Shiny application to perform Benchmark Dose analysis for transcriptomics data
abstract
MOTIVATION: The analysis of dose-dependent effects on the gene expression is gaining attention in the field of toxicogenomics. Currently available computational methods are usually limited to specific omics platforms or biological annotations and are able to analyse only one experiment at a time. RESULTS: We developed the software BMDx with a graphical user interface for the Benchmark Dose (BMD) analysis of transcriptomics data. We implemented an approach based on the fitting of multiple models and the selection of the optimal model based on the Akaike Information Criterion. The BMDx tool takes as an input a gene expression matrix and a phenotype table, computes the BMD, its related values, and IC50/EC50 estimations. It reports interactive tables and plots that the user can investigate for further details of the fitting, dose effects and functional enrichment. BMDx allows a fast and convenient comparison of the BMD values of a transcriptomics experiment at different time points and an effortless way to interpret the results. Furthermore, BMDx allows to analyse and to compare multiple experiments at once. AVAILABILITY AND IMPLEMENTATION: BMDx is implemented as an R/Shiny software and is available at https://github.com/Greco-Lab/BMDx/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Angela Serra, Laura Aliisa Saarimäki, Michele Fratello, Veer Singh Marwah, Dario Greco
Bioinform.3
2019 Strong-Weak Pruning for Brain Network Identification in Connectome-Wide Neuroimaging: Application to Amyotrophic Lateral Sclerosis Disease Stage Characterization
abstract
Magnetic resonance imaging allows acquiring functional and structural connectivity data from which high-density whole-brain networks can be derived to carry out connectome-wide analyses in normal and clinical populations. Graph theory has been widely applied to investigate the modular structure of brain connections by using centrality measures to identify the "hub" of human connectomes, and community detection methods to delineate subnetworks associated with diverse cognitive and sensorimotor functions. These analyses typically rely on a preprocessing step (pruning) to reduce computational complexity and remove the weakest edges that are most likely affected by experimental noise. However, weak links may contain relevant information about brain connectivity, therefore, the identification of the optimal trade-off between retained and discarded edges is a subject of active research. We introduce a pruning algorithm to identify edges that carry the highest information content. The algorithm selects both strong edges (i.e. edges belonging to shortest paths) and weak edges that are topologically relevant in weakly connected subnetworks. The newly developed "strong-weak" pruning (SWP) algorithm was validated on simulated networks that mimic the structure of human brain networks. It was then applied for the analysis of a real dataset of subjects affected by amyotrophic lateral sclerosis (ALS), both at the early (ALS2) and late (ALS3) stage of the disease, and of healthy control subjects. SWP preprocessing allowed identifying statistically significant differences in the path length of networks between patients and healthy subjects. ALS patients showed a decrease of connectivity between frontal cortex to temporal cortex and parietal cortex and between temporal and occipital cortex. Moreover, degree of centrality measures revealed significantly different hub and centrality scores between patient subgroups. These findings suggest a widespread alteration of network topology in ALS associated with disease progression.
Angela Serra, Paola Galdi, Emanuele Pesce, Michele Fratello, Francesca Trojsi, Gioacchino Tedeschi, Roberto Tagliaferri, Fabrizio Esposito
Int. J. Neural Syst.4
2018 Robust and sparse correlation matrix estimation for the analysis of high-dimensional genomics data
abstract
Motivation: Microarray technology can be used to study the expression of thousands of genes across a number of different experimental conditions, usually hundreds. The underlying principle is that genes sharing similar expression patterns, across different samples, can be part of the same co-expression system, or they may share the same biological functions. Groups of genes are usually identified based on cluster analysis. Clustering methods rely on the similarity matrix between genes. A common choice to measure similarity is to compute the sample correlation matrix. Dimensionality reduction is another popular data analysis task which is also based on covariance/correlation matrix estimates. Unfortunately, covariance/correlation matrix estimation suffers from the intrinsic noise present in high-dimensional data. Sources of noise are: sampling variations, presents of outlying sample units, and the fact that in most cases the number of units is much larger than the number of genes. Results: In this paper, we propose a robust correlation matrix estimator that is regularized based on adaptive thresholding. The resulting method jointly tames the effects of the high-dimensionality, and data contamination. Computations are easy to implement and do not require hand tunings. Both simulated and real data are analyzed. A Monte Carlo experiment shows that the proposed method is capable of remarkable performances. Our correlation metric is more robust to outliers compared with the existing alternatives in two gene expression datasets. It is also shown how the regularization allows to automatically detect and filter spurious correlations. The same regularization is also extended to other less robust correlation measures. Finally, we apply the ARACNE algorithm on the SyNTreN gene expression data. Sensitivity and specificity of the reconstructed network is compared with the gold standard. We show that ARACNE performs better when it takes the proposed correlation matrix estimator as input. Availability and implementation: The R software is available at https://github.com/angy89/RobustSparseCorrelation. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Angela Serra, Pietro Coretto, Michele Fratello, Roberto Tagliaferri
Bioinform.3
2018 Consensus-based feature extraction in rs-fMRI data analysis
Paola Galdi, Michele Fratello, Francesca Trojsi, Gioacchino Tedeschi, Roberto Tagliaferri, Fabrizio Esposito
Soft Comput.2
2016 Data integration in genomics and systems biology
abstract
Multi-view learning is the branch of machine learning that deals with multi modal data, i.e. with patterns represented by different sets of features. The fast spread of this learning technique is motivated by the continuing increase of real applications based on multi-view data. For example, in bioinformatics multiple experiments can be available (mRNA, miRNA and protein expression, genome wide association studies (GWAS) and others) for a set of samples. In bioinformatics multi-view approaches are useful since heterogeneous genome-wide data sources capture information on different aspects of complex biological systems. Each view provides a distinct facet of the same domain, encoding different biologically-relevant patterns. The integration of such views can provide a richer model of the underlying system than those produced by a single view alone. This paper provides a review of the literature with respect to bioinformatics, with the purpose to understand the principles and operation modes of the existing methods and their possible applications. In order to organize the proposed methods in literature and to find similarities between them, these approaches are organized according to three categories: the type of data used in the papers, the statistical problem and the stage of integration.
Angela Serra, Michele Fratello, Dario Greco, Roberto Tagliaferri
CEC2
2015 A multi-view genomic data simulator
abstract
BACKGROUND: OMICs technologies allow to assay the state of a large number of different features (e.g., mRNA expression, miRNA expression, copy number variation, DNA methylation, etc.) from the same samples. The objective of these experiments is usually to find a reduced set of significant features, which can be used to differentiate the conditions assayed. In terms of development of novel feature selection computational methods, this task is challenging for the lack of fully annotated biological datasets to be used for benchmarking. A possible way to tackle this problem is generating appropriate synthetic datasets, whose composition and behaviour are fully controlled and known a priori. RESULTS: Here we propose a novel method centred on the generation of networks of interactions among different biological molecules, especially involved in regulating gene expression. Synthetic datasets are obtained from ordinary differential equations based models with known parameters. Our results show that the generated datasets are well mimicking the behaviour of real data, for popular data analysis methods are able to selectively identify existing interactions. CONCLUSIONS: The proposed method can be used in conjunction to real biological datasets in the assessment of data mining techniques. The main strength of this method consists in the full control on the simulated data while retaining coherence with the real biological processes. The R package MVBioDataSim is freely available to the scientific community at http://neuronelab.unisa.it/?p=1722.
Michele Fratello, Angela Serra, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco
BMC Bioinform.1
2015 MVDA: a multi-view genomic data integration methodology
abstract
BACKGROUND: Multiple high-throughput molecular profiling by omics technologies can be collected for the same individuals. Combining these data, rather than exploiting them separately, can significantly increase the power of clinically relevant patients subclassifications. RESULTS: We propose a multi-view approach in which the information from different data layers (views) is integrated at the levels of the results of each single view clustering iterations. It works by factorizing the membership matrices in a late integration manner. We evaluated the effectiveness and the performance of our method on six multi-view cancer datasets. In all the cases, we found patient sub-classes with statistical significance, identifying novel sub-groups previously not emphasized in literature. Our method performed better as compared to other multi-view clustering algorithms and, unlike other existing methods, it is able to quantify the contribution of single views on the final results. CONCLUSION: Our observations suggest that integration of prior information with genomic features in the subtyping analysis is an effective strategy in identifying disease subgroups. The methodology is implemented in R and the source code is available online at http://neuronelab.unisa.it/a-multi-view-genomic-data-integration-methodology/ .
Angela Serra, Michele Fratello, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco
BMC Bioinform.2