Angela Serra

dblp:162/7408 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-3374-1492ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MUUMI: an R package for statistical and network-based meta-analysis for multi-omics data integration
abstract
BACKGROUND: Disentangling physiopathological mechanisms of biological systems through high-level integration of omics data has become a standard procedure in life sciences. However, platform heterogeneity, batch effects, and the lack of unified methods for single- and multi-omics analyses represent relevant drawbacks that hinder the extrapolation of a meaningful biological interpretation. While statistical meta-analysis is widely used to integrate several omics datasets of the same type, it does not allow the integration of multi-modal data deriving from multi-omics experiments. Network science is at the forefront of systems biology, where the inference of molecular interactomes allowed the investigation of perturbed biological systems, by shedding light on the disrupted relationships that keep the homeostasis of complex systems. RESULTS: Here, we present MUUMI, an R package that unifies statistical meta-analysis and network-based omics data integration within a single analytical framework. MUUMI allows the identification of robust molecular signatures through multiple meta-analytical methods, inference and analysis of molecular interactomes and the integration of multiple omics layers through similarity network fusion. We demonstrate the functionalities of MUUMI by presenting two case studies in which we analysed (1) 17 transcriptomic datasets on idiopathic pulmonary fibrosis (IPF) from both microarray and RNA-Seq platforms and (2) multi-omics data of THP-1 macrophages exposed to different polarising stimuli. In both examples, MUUMI revealed biologically coherent signatures, underscoring its value in elucidating complex biological processes. CONCLUSIONS: MUUMI leverages omics data meta-analysis, integration and interpretation that implements both traditional and network-based approaches to unleash the power of multi-study datasets. Statistical and network-based approaches are integrated in a unique framework, allowing the user to derive robust and biologically meaningful results from different studies and datasets. MUUMI is an open-source package and is freely available at https://github.com/fhaive/muumi .
Simo Iisakki Inkala, Michele Fratello, Giusy del Giudice, Giorgia Migliaccio, Angela Serra, Dario Greco, Antonio Federico
BMC Bioinform.5
2025 OpTiles: an R package for adaptive tiling and methylation variability profiling
abstract
SUMMARY: OpTiles is an R package that dynamically defines tiling windows based on the distribution of sequenced CpGs, addressing the limitations of traditional fixed-tiling approaches in targeted methylation datasets. By integrating CpG density with intra-region methylation variability, it provides a reliability metric and extended functionality for annotating, prioritizing, and interpreting complex methylation data. AVAILABILITY AND IMPLEMENTATION: OpTiles is implemented in R and source code is freely available at https://github.com/fhaive/OpTiles. Data are available on Zenodo at https://doi.org/10.5281/zenodo.16961292.
Giorgia Migliaccio, Lena Möbus, Giusy del Giudice, Jack Morikka, Antonio Federico, Angela Serra, Dario Greco
Bioinform.6
2023 DREAM: an R package for druggability evaluation of human complex diseases
abstract
MOTIVATION: De novo drug development is a long and expensive process that poses significant challenges from the design to the preclinical testing, making the introduction into the market slow and difficult. This limitation paved the way to the development of drug repurposing, which consists in the re-usage of already approved drugs, developed for other therapeutic indications. Although several efforts have been carried out in the last decade in order to achieve clinically relevant drug repurposing predictions, the amount of repurposed drugs that have been employed in actual pharmacological therapies is still limited. On one hand, mechanistic approaches, including profile-based and network-based methods, exploit the wealth of data about drug sensitivity and perturbational profiles as well as disease transcriptomics profiles. On the other hand, chemocentric approaches, including structure-based methods, take into consideration the intrinsic structural properties of the drugs and their molecular targets. The poor integration between mechanistic and chemocentric approaches is one of the main limiting factors behind the poor translatability of drug repurposing predictions into the clinics. RESULTS: In this work, we introduce DREAM, an R package aimed to integrate mechanistic and chemocentric approaches in a unified computational workflow. DREAM is devoted to the druggability evaluation of pathological conditions of interest, leveraging robust drug repurposing predictions. In addition, the user can derive optimized sets of drugs putatively suitable for combination therapy. In order to show the functionalities of the DREAM package, we report a case study on atopic dermatitis. AVAILABILITY AND IMPLEMENTATION: DREAM is freely available at https://github.com/fhaive/dream. The docker image of DREAM is available at: https://hub.docker.com/r/fhaive/dream.
Antonio Federico, Michele Fratello, Alisa Pavel, Lena Möbus, Giusy del Giudice, Angela Serra, Dario Greco
Bioinform.6
2023 ESPERANTO: a GLP-field sEmi-SuPERvised toxicogenomics metadAta curatioN TOol
abstract
SUMMARY: Biological data repositories are an invaluable source of publicly available research evidence. Unfortunately, the lack of convergence of the scientific community on a common metadata annotation strategy has resulted in large amounts of data with low FAIRness (Findable, Accessible, Interoperable and Reusable). The possibility of generating high-quality insights from their integration relies on data curation, which is typically an error-prone process while also being expensive in terms of time and human labour. Here, we present ESPERANTO, an innovative framework that enables a standardized semi-supervised harmonization and integration of toxicogenomics metadata and increases their FAIRness in a Good Laboratory Practice-compliant fashion. The harmonization across metadata is guaranteed with the definition of an ad hoc vocabulary. The tool interface is designed to support the user in metadata harmonization in a user-friendly manner, regardless of the background and the type of expertise. AVAILABILITY AND IMPLEMENTATION: ESPERANTO and its user manual are freely available for academic purposes at https://github.com/fhaive/esperanto. The input and the results showcased in Supplementary File S1 are available at the same link.
Emanuele Di Lieto, Angela Serra, Simo Iisakki Inkala, Laura Aliisa Saarimäki, Giusy del Giudice, Michele Fratello, Veera Hautanen, Maria Annala, Antonio Federico, Dario Greco
Bioinform.2
2023 KNeMAP: a network mapping approach for knowledge-driven comparison of transcriptomic profiles
abstract
MOTIVATION: Transcriptomic data can be used to describe the mechanism of action (MOA) of a chemical compound. However, omics data tend to be complex and prone to noise, making the comparison of different datasets challenging. Often, transcriptomic profiles are compared at the level of individual gene expression values, or sets of differentially expressed genes. Such approaches can suffer from underlying technical and biological variance, such as the biological system exposed on or the machine/method used to measure gene expression data, technical errors and further neglect the relationships between the genes. We propose a network mapping approach for knowledge-driven comparison of transcriptomic profiles (KNeMAP), which combines genes into similarity groups based on multiple levels of prior information, hence adding a higher-level view onto the individual gene view. When comparing KNeMAP with fold change (expression) based and deregulated gene set-based methods, KNeMAP was able to group compounds with higher accuracy with respect to prior information as well as is less prone to noise corrupted data. RESULT: We applied KNeMAP to analyze the Connectivity Map dataset, where the gene expression changes of three cell lines were analyzed after treatment with 676 drugs as well as the Fortino et al. dataset where two cell lines with 31 nanomaterials were analyzed. Although the expression profiles across the biological systems are highly different, KNeMAP was able to identify sets of compounds that induce similar molecular responses when exposed on the same biological system. AVAILABILITY AND IMPLEMENTATION: Relevant data and the KNeMAP function is available at: https://github.com/fhaive/KNeMAP and 10.5281/zenodo.7334711.
Alisa Pavel, Giusy del Giudice, Michele Fratello, Leo Ghemtio, Antonio Di Lieto, Jari Yli-Kauhaluoma, Henri Xhaard, Antonio Federico, Angela Serra, Dario Greco
Bioinform.9
2022 Computationally prioritized drugs inhibit SARS-CoV-2 infection and syncytia formation
abstract
The pharmacological arsenal against the COVID-19 pandemic is largely based on generic anti-inflammatory strategies or poorly scalable solutions. Moreover, as the ongoing vaccination campaign is rolling slower than wished, affordable and effective therapeutics are needed. To this end, there is increasing attention toward computational methods for drug repositioning and de novo drug design. Here, multiple data-driven computational approaches are systematically integrated to perform a virtual screening and prioritize candidate drugs for the treatment of COVID-19. From the list of prioritized drugs, a subset of representative candidates to test in human cells is selected. Two compounds, 7-hydroxystaurosporine and bafetinib, show synergistic antiviral effects in vitro and strongly inhibit viral-induced syncytia formation. Moreover, since existing drug repositioning methods provide limited usable information for de novo drug design, the relevant chemical substructures of the identified drugs are extracted to provide a chemical vocabulary that may help to design new effective drugs.
Angela Serra, Michele Fratello, Antonio Federico, Ravi Ojha, Riccardo Provenzani, Ervin Tasnádi, Luca Cattelani, Giusy del Giudice, Pia Anneli Sofia Kinaret, Laura Aliisa Saarimäki, Alisa Pavel, Suvi Kuivanen, Vincenzo Cerullo, Olli Vapalahti, Peter Horváth, Antonio Di Lieto, Jari Yli-Kauhaluoma, Giuseppe Balistreri, Dario Greco
Briefings Bioinform.1
2022 Deep learning for volatility forecasting in asset management
abstract
Abstract Predicting volatility is a critical activity for taking risk- adjusted decisions in asset trading and allocation. In order to provide effective decision-making support, in this paper we investigate the profitability of a deep Long Short-Term Memory (LSTM) Neural Network for forecasting daily stock market volatility using a panel of 28 assets representative of the Dow Jones Industrial Average index combined with the market factor proxied by the SPY and, separately, a panel of 92 assets belonging to the NASDAQ 100 index. The Dow Jones plus SPY data are from January 2002 to August 2008, while the NASDAQ 100 is from December 2012 to November 2017. If, on the one hand, we expect that this evolutionary behavior can be effectively captured adaptively through the use of Artificial Intelligence (AI) flexible methods, on the other, in this setting, standard parametric approaches could fail to provide optimal predictions. We compared the volatility forecasts generated by the LSTM approach to those obtained through use of widely recognized benchmarks models in this field, in particular, univariate parametric models such as the Realized Generalized Autoregressive Conditionally Heteroskedastic (R-GARCH) and the Glosten–Jagannathan–Runkle Multiplicative Error Models (GJR-MEM). The results demonstrate the superiority of the LSTM over the widely popular R-GARCH and GJR-MEM univariate parametric methods, when forecasting in condition of high volatility, while still producing comparable predictions for more tranquil periods.
Alessio Petrozziello, Luigi Troiano, Angela Serra, Ivan Jordanov, Giuseppe Storti, Roberto Tagliaferri, Michele La Rocca 0001
Soft Comput.3
2021 Integrated network analysis reveals new genes suggesting COVID-19 chronic effects and treatment
abstract
The COVID-19 disease led to an unprecedented health emergency, still ongoing worldwide. Given the lack of a vaccine or a clear therapeutic strategy to counteract the infection as well as its secondary effects, there is currently a pressing need to generate new insights into the SARS-CoV-2 induced host response. Biomedical data can help to investigate new aspects of the COVID-19 pathogenesis, but source heterogeneity represents a major drawback and limitation. In this work, we applied data integration methods to develop a Unified Knowledge Space (UKS) and used it to identify a new set of genes associated with SARS-CoV-2 host response, both in vitro and in vivo. Functional analysis of these genes reveals possible long-term systemic effects of the infection, such as vascular remodelling and fibrosis. Finally, we identified a set of potentially relevant drugs targeting proteins involved in multiple steps of the host response to the virus.
Alisa Pavel, Giusy del Giudice, Antonio Federico, Antonio Di Lieto, Pia Anneli Sofia Kinaret, Angela Serra, Dario Greco
Briefings Bioinform.6
2021 VOLTA: adVanced mOLecular neTwork Analysis
abstract
MOTIVATION: Network analysis is a powerful approach to investigate biological systems. It is often applied to study gene co-expression patterns derived from transcriptomics experiments. Even though co-expression analysis is widely used, there is still a lack of tools that are open and customizable on the basis of different network types and analysis scenarios (e.g. through function accessibility), but are also suitable for novice users by providing complete analysis pipelines. RESULTS: We developed VOLTA, a Python package suited for complex co-expression network analysis. VOLTA is designed to allow users direct access to the individual functions, while they are also provided with complete analysis pipelines. Moreover, VOLTA offers when possible multiple algorithms applicable to each analytical step (e.g. multiple community detection or clustering algorithms are provided), hence providing the user with the possibility to perform analysis tailored to their needs. This makes VOLTA highly suitable for experienced users who wish to build their own analysis pipelines for a wide range of networks as well as for novice users for which a 'plug and play' system is provided. AVAILABILITY AND IMPLEMENTATION: The package and used data are available at GitHub: https://github.com/fhaive/VOLTA and 10.5281/zenodo.5171719. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alisa Pavel, Antonio Federico, Giusy del Giudice, Angela Serra, Dario Greco
Bioinform.4
2020 MaNGA: a novel multi-niche multi-objective genetic algorithm for QSAR modelling
abstract
SUMMARY: Quantitative structure-activity relationship (QSAR) modelling is currently used in multiple fields to relate structural properties of compounds to their biological activities. This technique is also used for drug design purposes with the aim of predicting parameters that determine drug behaviour. To this end, a sophisticated process, involving various analytical steps concatenated in series, is employed to identify and fine-tune the optimal set of predictors from a large dataset of molecular descriptors (MDs). The search of the optimal model requires to optimize multiple objectives at the same time, as the aim is to obtain the minimal set of features that maximizes the goodness of fit and the applicability domain (AD). Hence, a multi-objective optimization strategy, improving multiple parameters in parallel, can be applied. Here we propose a new multi-niche multi-objective genetic algorithm that simultaneously enables stable feature selection as well as obtaining robust and validated regression models with maximized AD. We benchmarked our method on two simulated datasets. Moreover, we analyzed an aquatic acute toxicity dataset and compared the performances of single- and multi-objective fitness functions on different regression models. Our results show that our multi-objective algorithm is a valid alternative to classical QSAR modelling strategy, for continuous response values, since it automatically finds the model with the best compromise between statistical robustness, predictive performance, widest AD, and the smallest number of MDs. AVAILABILITY AND IMPLEMENTATION: The python implementation of MaNGA is available at https://github.com/Greco-Lab/MaNGA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Angela Serra, Serli Önlü, Paola Festa, Vittorio Fortino, Dario Greco
Bioinform.1
2020 BMDx: a graphical Shiny application to perform Benchmark Dose analysis for transcriptomics data
abstract
MOTIVATION: The analysis of dose-dependent effects on the gene expression is gaining attention in the field of toxicogenomics. Currently available computational methods are usually limited to specific omics platforms or biological annotations and are able to analyse only one experiment at a time. RESULTS: We developed the software BMDx with a graphical user interface for the Benchmark Dose (BMD) analysis of transcriptomics data. We implemented an approach based on the fitting of multiple models and the selection of the optimal model based on the Akaike Information Criterion. The BMDx tool takes as an input a gene expression matrix and a phenotype table, computes the BMD, its related values, and IC50/EC50 estimations. It reports interactive tables and plots that the user can investigate for further details of the fitting, dose effects and functional enrichment. BMDx allows a fast and convenient comparison of the BMD values of a transcriptomics experiment at different time points and an effortless way to interpret the results. Furthermore, BMDx allows to analyse and to compare multiple experiments at once. AVAILABILITY AND IMPLEMENTATION: BMDx is implemented as an R/Shiny software and is available at https://github.com/Greco-Lab/BMDx/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Angela Serra, Laura Aliisa Saarimäki, Michele Fratello, Veer Singh Marwah, Dario Greco
Bioinform.1
2019 FunMappOne: a tool to hierarchically organize and visually navigate functional gene annotations in multiple experiments
abstract
BACKGROUND: Functional annotation of genes is an essential step in omics data analysis. Multiple databases and methods are currently available to summarize the functions of sets of genes into higher level representations, such as ontologies and molecular pathways. Annotating results from omics experiments into functional categories is essential not only to understand the underlying regulatory dynamics but also to compare multiple experimental conditions at a higher level of abstraction. Several tools are already available to the community to represent and compare functional profiles of omics experiments. However, when the number of experiments and/or enriched functional terms is high, it becomes difficult to interpret the results even when graphically represented. Therefore, there is currently a need for interactive and user-friendly tools to graphically navigate and further summarize annotations in order to facilitate results interpretation also when the dimensionality is high. RESULTS: We developed an approach that exploits the intrinsic hierarchical structure of several functional annotations to summarize the results obtained through enrichment analyses to higher levels of interpretation and to map gene related information at each summarized level. We built a user-friendly graphical interface that allows to visualize the functional annotations of one or multiple experiments at once. The tool is implemented as a R-Shiny application called FunMappOne and is available at https://github.com/grecolab/FunMappOne . CONCLUSION: FunMappOne is a R-shiny graphical tool that takes in input multiple lists of human or mouse genes, optionally along with their related modification magnitudes, computes the enriched annotations from Gene Ontology, Kyoto Encyclopedia of Genes and Genomes, or Reactome databases, and reports interactive maps of functional terms and pathways organized in rational groups. FunMappOne allows a fast and convenient comparison of multiple experiments and an easy way to interpret results.
Giovanni Scala, Angela Serra, Veer Singh Marwah, Laura Aliisa Saarimäki, Dario Greco
BMC Bioinform.2
2019 Strong-Weak Pruning for Brain Network Identification in Connectome-Wide Neuroimaging: Application to Amyotrophic Lateral Sclerosis Disease Stage Characterization
abstract
Magnetic resonance imaging allows acquiring functional and structural connectivity data from which high-density whole-brain networks can be derived to carry out connectome-wide analyses in normal and clinical populations. Graph theory has been widely applied to investigate the modular structure of brain connections by using centrality measures to identify the "hub" of human connectomes, and community detection methods to delineate subnetworks associated with diverse cognitive and sensorimotor functions. These analyses typically rely on a preprocessing step (pruning) to reduce computational complexity and remove the weakest edges that are most likely affected by experimental noise. However, weak links may contain relevant information about brain connectivity, therefore, the identification of the optimal trade-off between retained and discarded edges is a subject of active research. We introduce a pruning algorithm to identify edges that carry the highest information content. The algorithm selects both strong edges (i.e. edges belonging to shortest paths) and weak edges that are topologically relevant in weakly connected subnetworks. The newly developed "strong-weak" pruning (SWP) algorithm was validated on simulated networks that mimic the structure of human brain networks. It was then applied for the analysis of a real dataset of subjects affected by amyotrophic lateral sclerosis (ALS), both at the early (ALS2) and late (ALS3) stage of the disease, and of healthy control subjects. SWP preprocessing allowed identifying statistically significant differences in the path length of networks between patients and healthy subjects. ALS patients showed a decrease of connectivity between frontal cortex to temporal cortex and parietal cortex and between temporal and occipital cortex. Moreover, degree of centrality measures revealed significantly different hub and centrality scores between patient subgroups. These findings suggest a widespread alteration of network topology in ALS associated with disease progression.
Angela Serra, Paola Galdi, Emanuele Pesce, Michele Fratello, Francesca Trojsi, Gioacchino Tedeschi, Roberto Tagliaferri, Fabrizio Esposito
Int. J. Neural Syst.1
2018 Robust clustering of noisy high-dimensional gene expression data for patients subtyping
abstract
Motivation: One of the most important research areas in personalized medicine is the discovery of disease sub-types with relevance in clinical applications. This is usually accomplished by exploring gene expression data with unsupervised clustering methodologies. Then, with the advent of multiple omics technologies, data integration methodologies have been further developed to obtain better performances in patient separability. However, these methods do not guarantee the survival separability of the patients in different clusters. Results: We propose a new methodology that first computes a robust and sparse correlation matrix of the genes, then decomposes it and projects the patient data onto the first m spectral components of the correlation matrix. After that, a robust and adaptive to noise clustering algorithm is applied. The clustering is set up to optimize the separation between survival curves estimated cluster-wise. The method is able to identify clusters that have different omics signatures and also statistically significant differences in survival time. The proposed methodology is tested on five cancer datasets downloaded from The Cancer Genome Atlas repository. The proposed method is compared with the Similarity Network Fusion (SNF) approach, and model based clustering based on Student's t-distribution (TMIX). Our method obtains a better performance in terms of survival separability, even if it uses a single gene expression view compared to the multi-view approach of the SNF method. Finally, a pathway based analysis is accomplished to highlight the biological processes that differentiate the obtained patient groups. Availability and implementation: Our R source code is available online at https://github.com/angy89/RobustClusteringPatientSubtyping. Supplementary information: Supplementary data are available at Bioinformatics online.
Pietro Coretto, Angela Serra, Roberto Tagliaferri
Bioinform.2
2018 INfORM: Inference of NetwOrk Response Modules
abstract
Summary: Detecting and interpreting responsive modules from gene expression data by using network-based approaches is a common but laborious task. It often requires the application of several computational methods implemented in different software packages, forcing biologists to compile complex analytical pipelines. Here we introduce INfORM (Inference of NetwOrk Response Modules), an R shiny application that enables non-expert users to detect, evaluate and select gene modules with high statistical and biological significance. INfORM is a comprehensive tool for the identification of biologically meaningful response modules from consensus gene networks inferred by using multiple algorithms. It is accessible through an intuitive graphical user interface allowing for a level of abstraction from the computational steps. Availability and implementation: INfORM is freely available for academic use at https://github.com/Greco-Lab/INfORM. Supplementary information: Supplementary data are available at Bioinformatics online.
Veer Singh Marwah, Pia Anneli Sofia Kinaret, Angela Serra, Giovanni Scala, Antti Lauerma, Vittorio Fortino, Dario Greco
Bioinform.3
2018 Robust and sparse correlation matrix estimation for the analysis of high-dimensional genomics data
abstract
Motivation: Microarray technology can be used to study the expression of thousands of genes across a number of different experimental conditions, usually hundreds. The underlying principle is that genes sharing similar expression patterns, across different samples, can be part of the same co-expression system, or they may share the same biological functions. Groups of genes are usually identified based on cluster analysis. Clustering methods rely on the similarity matrix between genes. A common choice to measure similarity is to compute the sample correlation matrix. Dimensionality reduction is another popular data analysis task which is also based on covariance/correlation matrix estimates. Unfortunately, covariance/correlation matrix estimation suffers from the intrinsic noise present in high-dimensional data. Sources of noise are: sampling variations, presents of outlying sample units, and the fact that in most cases the number of units is much larger than the number of genes. Results: In this paper, we propose a robust correlation matrix estimator that is regularized based on adaptive thresholding. The resulting method jointly tames the effects of the high-dimensionality, and data contamination. Computations are easy to implement and do not require hand tunings. Both simulated and real data are analyzed. A Monte Carlo experiment shows that the proposed method is capable of remarkable performances. Our correlation metric is more robust to outliers compared with the existing alternatives in two gene expression datasets. It is also shown how the regularization allows to automatically detect and filter spurious correlations. The same regularization is also extended to other less robust correlation measures. Finally, we apply the ARACNE algorithm on the SyNTreN gene expression data. Sensitivity and specificity of the reconstructed network is compared with the gold standard. We show that ARACNE performs better when it takes the proposed correlation matrix estimator as input. Availability and implementation: The R software is available at https://github.com/angy89/RobustSparseCorrelation. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Angela Serra, Pietro Coretto, Michele Fratello, Roberto Tagliaferri
Bioinform.1
2016 Data integration in genomics and systems biology
abstract
Multi-view learning is the branch of machine learning that deals with multi modal data, i.e. with patterns represented by different sets of features. The fast spread of this learning technique is motivated by the continuing increase of real applications based on multi-view data. For example, in bioinformatics multiple experiments can be available (mRNA, miRNA and protein expression, genome wide association studies (GWAS) and others) for a set of samples. In bioinformatics multi-view approaches are useful since heterogeneous genome-wide data sources capture information on different aspects of complex biological systems. Each view provides a distinct facet of the same domain, encoding different biologically-relevant patterns. The integration of such views can provide a richer model of the underlying system than those produced by a single view alone. This paper provides a review of the literature with respect to bioinformatics, with the purpose to understand the principles and operation modes of the existing methods and their possible applications. In order to organize the proposed methods in literature and to find similarities between them, these approaches are organized according to three categories: the type of data used in the papers, the statistical problem and the stage of integration.
Angela Serra, Michele Fratello, Dario Greco, Roberto Tagliaferri
CEC1
2015 Impact of different metrics on multi-view clustering
abstract
Clustering of patients allows to find groups of subjects with similar characteristics. This categorization can facilitate diagnosis, treatment decision and prognosis prediction. Heterogeneous genome-wide data sources capture different biological aspects that can be integrated in order to better categorize the patients. Clustering methods work by comparing how patients are similar or dissimilar in a suitable similarity space. While several clustering methods have been proposed, there is no systematic comparative study concerning the impact of similarity metrics on the cluster quality. We compared seven popular similarity measures (Pearson, Spearman and Kendall Correlations; Euclidean, Canberra, Minkowski and Manhattan Distances) in conjunction with two classical single-view clustering algorithms and a late integration approach (partitioning around medoids, hierarchical clustering and matrix factorization approaches), on high dimensional multi-view cancer data coming from the TCGA repository. Performance was measured against tumour subcategories classification. Only Euclidean and Minkowski distances showed similar results in terms of clustering similarity indexes. On the other hand, an absolute best similarity measure did not emerge in terms of misclassification, but it strongly depends on the data.
Angela Serra, Dario Greco, Roberto Tagliaferri
IJCNN1
2015 A multi-view genomic data simulator
abstract
BACKGROUND: OMICs technologies allow to assay the state of a large number of different features (e.g., mRNA expression, miRNA expression, copy number variation, DNA methylation, etc.) from the same samples. The objective of these experiments is usually to find a reduced set of significant features, which can be used to differentiate the conditions assayed. In terms of development of novel feature selection computational methods, this task is challenging for the lack of fully annotated biological datasets to be used for benchmarking. A possible way to tackle this problem is generating appropriate synthetic datasets, whose composition and behaviour are fully controlled and known a priori. RESULTS: Here we propose a novel method centred on the generation of networks of interactions among different biological molecules, especially involved in regulating gene expression. Synthetic datasets are obtained from ordinary differential equations based models with known parameters. Our results show that the generated datasets are well mimicking the behaviour of real data, for popular data analysis methods are able to selectively identify existing interactions. CONCLUSIONS: The proposed method can be used in conjunction to real biological datasets in the assessment of data mining techniques. The main strength of this method consists in the full control on the simulated data while retaining coherence with the real biological processes. The R package MVBioDataSim is freely available to the scientific community at http://neuronelab.unisa.it/?p=1722.
Michele Fratello, Angela Serra, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco
BMC Bioinform.2
2015 MVDA: a multi-view genomic data integration methodology
abstract
BACKGROUND: Multiple high-throughput molecular profiling by omics technologies can be collected for the same individuals. Combining these data, rather than exploiting them separately, can significantly increase the power of clinically relevant patients subclassifications. RESULTS: We propose a multi-view approach in which the information from different data layers (views) is integrated at the levels of the results of each single view clustering iterations. It works by factorizing the membership matrices in a late integration manner. We evaluated the effectiveness and the performance of our method on six multi-view cancer datasets. In all the cases, we found patient sub-classes with statistical significance, identifying novel sub-groups previously not emphasized in literature. Our method performed better as compared to other multi-view clustering algorithms and, unlike other existing methods, it is able to quantify the contribution of single views on the final results. CONCLUSION: Our observations suggest that integration of prior information with genomic features in the subtyping analysis is an effective strategy in identifying disease subgroups. The methodology is implemented in R and the source code is available online at http://neuronelab.unisa.it/a-multi-view-genomic-data-integration-methodology/ .
Angela Serra, Michele Fratello, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco
BMC Bioinform.1