Federico Manuel Giorgi

dblp:59/8708 · also Federico M. Giorgi · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-7325-9908ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Differential causal networks highlight sex-based differences in human tissues
abstract
Sex differences appear in healthy and pathological conditions and may influence sex-specific therapeutic responses. Understanding such differences is a key activity for developing precision medicine strategies. This study investigates sex differences in gene expression across 40 human tissues by applying a Differential Causal Network (DCN) analysis using data from the Genotype-Tissue Expression project. We identified sex-based DCNs that highlight distinct molecular mechanisms influencing both health and disease in men and women. For example, in pancreas tissue, genes associated with immune system show significant differences in their regulatory patterns between sexes, demonstrating a possible different response to diseases such as diabetes mellitus and cancer. Our findings provide valuable information on the biological underpinnings of sex differences, offering potential pathways for the development of precision medicine strategies.
Annamaria Defilippo, Kimberly Glass, Federico Manuel Giorgi, Tamer Kahveci, Pierangelo Veltri, Pietro H. Guzzi
Briefings Bioinform.3
2025 Utargetome: A targetome prediction tool for modified U1-snRNAs to identify distal-target positions with improved selectivity
abstract
The endogenous U1 small nuclear RNA (U1-snRNA) plays a crucial role in splicing initiation through base-pairing to donor splice sites (5'-SSs). Likewise, modified U1s that carry a mutation-adapted 5'-terminal sequence have been demonstrated to rescue exon splicing when this is disrupted by genetic mutations within the 5'-SS. Given the base-pairing flexibility of the endogenous U1, the selectivity of modified U1s requires investigation. We developed a computational pipeline (Utargetome) that considers combinations of mismatches and alternative annealing registers to predict the transcriptome-wide binding sites (or targetome) of a U1. The pipeline accuracy was tested by recapitulating well-established alternative annealing registers and specificity for 5'-SSs in the predicted targetome of the human endogenous U1. It was then applied to analyse the targetome of 54 modified U1s that have been demonstrated to restore exon inclusion when affected by 5'-SS pathogenic mutations. While the targetome size was found to be wide-ranging, the off-target load appeared to be reduced for U1s targeting distal sites from the canonical U1-binding position. This feature was predicted also for a large set of 30,204 newly designed U1s targeting 839 5'-SS pathogenic mutations that were expected to affect exon inclusion. Targetome analysis indeed revealed an optimal distal-targeting position at 3 nucleotides downstream from the canonical 5'-SS, for which a modified U1 is likely to have minimal off-targets at 5'-SSs and acceptor splice sites (3'-SSs). Based on these insights, we propose to implement targetome prediction in the design and optimization of therapeutic U1s with improved selectivity.
Paolo Pigini, Federico Manuel Giorgi, Keng Boon Wee
PLoS Comput. Biol.2
2022 Detection of pan-cancer surface protein biomarkers via a network-based approach on transcriptomics data
abstract
Cell surface proteins have been used as diagnostic and prognostic markers in cancer research and as targets for the development of anticancer agents. Many of these proteins lie at the top of signaling cascades regulating cell responses and gene expression, therefore acting as 'signaling hubs'. It has been previously demonstrated that the integrated network analysis on transcriptomic data is able to infer cell surface protein activity in breast cancer. Such an approach has been implemented in a publicly available method called 'SURFACER'. SURFACER implements a network-based analysis of transcriptomic data focusing on the overall activity of curated surface proteins, with the final aim to identify those proteins driving major phenotypic changes at a network level, named surface signaling hubs. Here, we show the ability of SURFACER to discover relevant knowledge within and across cancer datasets. We also show how different cancers can be stratified in surface-activity-specific groups. Our strategy may identify cancer-wide markers to design targeted therapies and biomarker-based diagnostic approaches.
Daniele Mercatelli, Chiara Cabrelle, Pierangelo Veltri, Federico Manuel Giorgi, Pietro H. Guzzi
Briefings Bioinform.4
2021 Web tools to fight pandemics: the COVID-19 experience
abstract
The current outbreak of COVID-19 has generated an unprecedented scientific response worldwide, with the generation of vast amounts of publicly available epidemiological, biological and clinical data. Bioinformatics scientists have quickly produced online methods to provide non-computational users with the opportunity of analyzing such data. In this review, we report the results of this effort, by cataloguing the currently most popular web tools for COVID-19 research and analysis. Our focus was driven on tools drawing data from the fields of epidemiology, genomics, interactomics and pharmacology, in order to provide a meaningful depiction of the current state of the art of COVID-19 online resources.
Daniele Mercatelli, Andrew N. Holding, Federico Manuel Giorgi
Briefings Bioinform.3
2021 svpluscnv: analysis and visualization of complex structural variation data
abstract
MOTIVATION: Despite widespread prevalence of somatic structural variations (SVs) across most tumor types, understanding of their molecular implications often remains poor. SVs are extremely heterogeneous in size and complexity, hindering the interpretation of their pathogenic role. Tools integrating large SV datasets across platforms are required to fully characterize the cancer's somatic landscape. RESULTS: svpluscnv R package is a swiss army knife for the integration and interpretation of orthogonal datasets including copy number variant segmentation profiles and sequencing-based structural variant calls. The package implements analysis and visualization tools to evaluate chromosomal instability and ploidy, identify genes harboring recurrent SVs and detects complex rearrangements such as chromothripsis and chromoplexia. Further, it allows systematic identification of hot-spot shattered genomic regions, showing reproducibility across alternative detection methods and datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/ccbiolab/svpluscnv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gonzalo López, Laura E. Egolf, Federico Manuel Giorgi, Sharon J. Diskin, Adam A. Margolin
Bioinform.3
2020 Spathial: an R package for the evolutionary analysis of biological data
abstract
SUMMARY: A primary problem in high-throughput genomics experiments is finding the most important genes involved in biological processes (e.g. tumor progression). In this applications note, we introduce spathial, an R package for navigating high-dimensional data spaces. spathial implements the Principal Path algorithm, which is a topological method for locally navigating on the data manifold. The package, together with the core algorithm, provides several high-level functions for interpreting the results. One of the analyses we propose is the extraction of the genes that are mainly involved in the progress from one state to another. We show a possible application in the context of tumor progression using RNA-Seq and single-cell datasets, and we compare our results with two commonly used algorithms, edgeR and monocle3, respectively. AVAILABILITY AND IMPLEMENTATION: The R package spathial is available on the Comprehensive R Archive Network (https://cran.r-project.org/web/packages/spathial/index.html) and on GitHub (https://github.com/erikagardini/spathial). It is distributed under the GNU General Public License (version 3). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erika Gardini, Federico Manuel Giorgi, Sergio Decherchi, Andrea Cavalli
Bioinform.2
2020 corto: a lightweight R package for gene network inference and master regulator analysis
abstract
MOTIVATION: Gene network inference and master regulator analysis (MRA) have been widely adopted to define specific transcriptional perturbations from gene expression signatures. Several tools exist to perform such analyses but most require a computer cluster or large amounts of RAM to be executed. RESULTS: We developed corto, a fast and lightweight R package to infer gene networks and perform MRA from gene expression data, with optional corrections for copy-number variations and able to run on signatures generated from RNA-Seq or ATAC-Seq data. We extensively benchmarked it to infer context-specific gene networks in 39 human tumor and 27 normal tissue datasets. AVAILABILITY AND IMPLEMENTATION: Cross-platform and multi-threaded R package on CRAN (stable version) https://cran.r-project.org/package=corto and Github (development release) https://github.com/federicogiorgi/corto. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniele Mercatelli, Gonzalo Lopez-Garcia, Federico Manuel Giorgi
Bioinform.3
2016 Detection and removal of spatial bias in multiwell assays
abstract
MOTIVATION: Multiplex readout assays are now increasingly being performed using microfluidic automation in multiwell format. For instance, the Library of Integrated Network-based Cellular Signatures (LINCS) has produced gene expression measurements for tens of thousands of distinct cell perturbations using a 384-well plate format. This dataset is by far the largest 384-well gene expression measurement assay ever performed. We investigated the gene expression profiles of a million samples from the LINCS dataset and found that the vast majority (96%) of the tested plates were affected by a significant 2D spatial bias. RESULTS: Using a novel algorithm combining spatial autocorrelation detection and principal component analysis, we could remove most of the spatial bias from the LINCS dataset and show in parallel a dramatic improvement of similarity between biological replicates assayed in different plates. The proposed methodology is fully general and can be applied to any highly multiplexed assay performed in multiwell format. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander Lachmann, Federico Manuel Giorgi, Mariano J. Alvarez, Andrea Califano
Bioinform.2
2016 ARACNe-AP: gene network reverse engineering through adaptive partitioning inference of mutual information
abstract
UNLABELLED: The accurate reconstruction of gene regulatory networks from large scale molecular profile datasets represents one of the grand challenges of Systems Biology. The Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) represents one of the most effective tools to accomplish this goal. However, the initial Fixed Bandwidth (FB) implementation is both inefficient and unable to deal with sample sets providing largely uneven coverage of the probability density space. Here, we present a completely new implementation of the algorithm, based on an Adaptive Partitioning strategy (AP) for estimating the Mutual Information. The new AP implementation (ARACNe-AP) achieves a dramatic improvement in computational performance (200× on average) over the previous methodology, while preserving the Mutual Information estimator and the Network inference accuracy of the original algorithm. Given that the previous version of ARACNe is extremely demanding, the new version of the algorithm will allow even researchers with modest computational resources to build complex regulatory networks from hundreds of gene expression profiles. AVAILABILITY AND IMPLEMENTATION: A JAVA cross-platform command line executable of ARACNe, together with all source code and a detailed usage guide are freely available on Sourceforge (http://sourceforge.net/projects/aracne-ap). JAVA version 8 or higher is required. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander Lachmann, Federico Manuel Giorgi, Gonzalo López, Andrea Califano
Bioinform.2
2013 The NGS WikiBook: a dynamic collaborative online training effort with long-term sustainability
abstract
Next-generation sequencing (NGS) is increasingly being adopted as the backbone of biomedical research. With the commercialization of various affordable desktop sequencers, NGS will be reached by increasing numbers of cellular and molecular biologists, necessitating community consensus on bioinformatics protocols to tackle the exponential increase in quantity of sequence data. The current resources for NGS informatics are extremely fragmented. Finding a centralized synthesis is difficult. A multitude of tools exist for NGS data analysis; however, none of these satisfies all possible uses and needs. This gap in functionality could be filled by integrating different methods in customized pipelines, an approach helped by the open-source nature of many NGS programmes. Drawing from community spirit and with the use of the Wikipedia framework, we have initiated a collaborative NGS resource: The NGS WikiBook. We have collected a sufficient amount of text to incentivize a broader community to contribute to it. Users can search, browse, edit and create new content, so as to facilitate self-learning and feedback to the community. The overall structure and style for this dynamic material is designed for the bench biologists and non-bioinformaticians. The flexibility of online material allows the readers to ignore details in a first read, yet have immediate access to the information they need. Each chapter comes with practical exercises so readers may familiarize themselves with each step. The NGS WikiBook aims to create a collective laboratory book and protocol that explains the key concepts and describes best practices in this fast-evolving field.
Jing-Woei Li, Dan M. Bolser, Magnus Manske, Federico Manuel Giorgi, Nikolay Vyahhi, Björn Usadel, Bernardo J. Clavijo, Ting-Fung Chan, Nathalie Wong, Daniel R. Zerbino, Maria Victoria Schneider
Briefings Bioinform.4
2013 Comparative study of RNA-seq- and Microarray-derived coexpression networks in Arabidopsis thaliana
abstract
MOTIVATION: Coexpression networks are data-derived representations of genes behaving in a similar way across tissues and experimental conditions. They have been used for hypothesis generation and guilt-by-association approaches for inferring functions of previously unknown genes. So far, the main platform for expression data has been DNA microarrays; however, the recent development of RNA-seq allows for higher accuracy and coverage of transcript populations. It is therefore important to assess the potential for biological investigation of coexpression networks derived from this novel technique in a condition-independent dataset. RESULTS: We collected 65 publicly available Illumina RNA-seq high quality Arabidopsis thaliana samples and generated Pearson correlation coexpression networks. These networks were then compared with those derived from analogous microarray data. We show how Variance-Stabilizing Transformed (VST) RNA-seq data samples are the most similar to microarray ones, with respect to inter-sample variation, correlation coefficient distribution and network topological architecture. Microarray networks show a slightly higher score in biology-derived quality assessments such as overlap with the known protein-protein interaction network and edge ontological agreement. Different coexpression network centralities are investigated; in particular, we show how betweenness centrality is generally a positive marker for essential genes in A.thaliana, regardless of the platform originating the data. In the end, we focus on a specific gene network case, showing that although microarray data seem more suited for gene network reverse engineering, RNA-seq offers the great advantage of extending coexpression analyses to the entire transcriptome.
Federico Manuel Giorgi, Cristian Del Fabbro, Francesco Licausi
Bioinform.1
2010 Algorithm-driven Artifacts in median polish summarization of Microarray data
abstract
BACKGROUND: High-throughput measurement of transcript intensities using Affymetrix type oligonucleotide microarrays has produced a massive quantity of data during the last decade. Different preprocessing techniques exist to convert the raw signal intensities measured by these chips into gene expression estimates. Although these techniques have been widely benchmarked in the context of differential gene expression analysis, there are only few examples where their performance has been assessed in respect to coexpression-based studies such as sample classification. RESULTS: In the present paper we benchmark the three most used normalization procedures (MAS5, RMA and GCRMA) in the context of inter-array correlation analysis, confirming and extending the finding that RMA and GCRMA consistently overestimate sample similarity upon normalization. We determine that median polish summarization is responsible for generating a large proportion of these over-similarity artifacts. Furthermore, we show that most affected probesets show also internal signal disagreement, and tend to be composed by individual probes hitting different gene transcripts. We finally provide a correction to the RMA/GCRMA summarization procedure that massively reduces inter-array correlation artifacts, without affecting the detection of differentially expressed genes. CONCLUSIONS: We propose tRMA as a modification of RMA to normalize microarray experiments for correlation-based analysis.
Federico Manuel Giorgi, Anthony M. Bolger, Marc Lohse, Björn Usadel
BMC Bioinform.1