Giacomo Baruzzo

dblp:227/1612 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-6129-5007ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Biologically Informed procedure for Feature Summarization in Spatial Transcriptomics
abstract
Modern machine learning approaches have shown remarkable success in extracting patterns from high-dimensional biological data. However, when applied to spatial transcriptomics, these methods face significant challenges due to the sparsity of spatially resolved measurements and the complex, nonlinear relationships between molecular features.To address these challenges, we propose a procedure that integrates single-cell and spatial transcriptomics by considering biologically meaningful regulatory factors as an interpretable feature space. These factors act as latent variables that encode transcriptional programmes, reducing dimensionality and preserving mechanistic relevance.This approach improves interpretability by shifting from raw gene expression to a structured representation of regulatory activity, providing a scalable and biologically interpretable framework for spatial transcriptomic analysis.
Matteo Baldan, Giulia Cesaro, Giacomo Baruzzo, Barbara Di Camillo
IJCNN3
2025 quickSparseM: a library for memory- and time-efficient computation on large, sparse matrices with application to omics data
abstract
Omics data have revolutionized molecular biology by introducing large-scale data analysis, pushing the field into the realm of big data and presenting substantial challenges in data storage and analysis. Despite describing distinct aspects of molecular biology, most omics data share common characteristics, such as being representable as large, sparse matrices, and requiring similar computational approaches, mainly involving embarrassing parallel tasks across rows or columns. While R is a popular choice for omics analysis, it encounters performance bottlenecks when handling large datasets due to its reliance on dense data formats and constraints like 32-bit indexing in some structures. Even when sparse representations are utilized, the inherent limitations of R lead to inefficiencies. Additionally, its lack of native support for shared-memory parallelism prevents it from fully utilizing modern parallel computing architectures. Similarly, many other data-intensive fields that rely on R face similar challenges with large, sparse data requiring fast and memory-efficient row-wise and column-wise operations. To address these challenges, we introduce quickSparseM, a time- and memory-efficient library for storing and processing large, sparse matrices, available as an R package. Developed in C++ with OpenMP for parallelism, quickSparseM achieves efficient performance while remaining compatible with existing R-based workflows. The library utilizes the R dgCMatrix format to represent sparse matrices in a compressed, column-oriented format and provide functions to compute basic statistics and operations commonly used in omics analyses. Experiments varying dataset sizes and core counts, as well as two case studies using omics data, demonstrate the library’s efficiency and scalability. The results indicate that quickSparseM outperforms state-of-the-art R packages for sparse matrix computation in terms of time, memory usage, and scalability.
Giacomo Baruzzo, Giulia Cesaro, Barbara Di Camillo
PDP1
2025 Advances and challenges in cell-cell communication inference: a comprehensive review of tools, resources, and future directions
abstract
Recent advancements in high-resolution and high-throughput sequencing technologies have significantly enhanced the study of cell-cell communication inference using single-cell and spatial transcriptomics data. Over the past 6 years, this growing interest has led to the development of more than 100 bioinformatics tools and nearly 50 resources, primarily in the form of ligand-receptor databases. These tools vary widely in their requirements, scoring approaches, ability to infer inter- and/or intra-cellular communication, assumptions, and limitations. Similarly, cell-cell communication resources differ in many aspects, mainly in the number of annotated interactions, species coverage, and their focus on inter-cellular signaling or both inter- and intra-cellular communication. This abundance and diversity create challenges in identifying compatible and suitable tools and resources to meet specific user needs. In this collaborative effort, we aim to provide a comprehensive report on the current state of cell-cell communication analysis derived from single-cell or spatial transcriptomics data. The report reviews existing methods and resources, addressing all relevant aspects from the user's perspective. It also explores current limitations, pitfalls, and unresolved issues in cell-cell communication inference, offering an aggregated analysis of the existing literature on the topic. Furthermore, we highlight potential future directions in the field and consolidate the collected knowledge into CCC-Catalog (https://sysbiobig.gitlab.io/ccc-catalog), a centralized web platform designed to serve as a hub for bioinformaticians and researchers interested in cell-cell communication inference.
Giulia Cesaro, James Shiniti Nagai, Nicolò Gnoato, Alice Chiodi, Gaia Tussardi, Vanessa Klöker, Carmelo Vittorio Musumarra, Ettore Mosca, Ivan G. Costa, Barbara Di Camillo, Enrica Calura, Giacomo Baruzzo
Briefings Bioinform.12
2025 MOV&RSim: computational modelling of cancer-specific variants and sequencing reads characteristics for realistic tumoral sample simulation
abstract
BACKGROUND: Bioinformatics pipelines for variant calling have undergone significant advancements due to the decreasing costs of next-generation sequencing. Accurate mutation detection is crucial for personalised medicine in cancer, particularly in assignment of therapy. Somatic variant calling, however, remains challenging due to diverse cancer types, heterogeneity, complex mutational profiles, and unpredictable sequencing errors. A dataset of fully characterised tumoral genomes and sequencing reads, large enough to represent the variability inherent in different cancer types, is still lacking, even considering synthetic data. The lack of such datasets hampers rigorous evaluation, benchmarking and optimization of variant callers for specific cancer types. RESULTS: The contribution of this work is twofold. First, we conducted a comprehensive analysis of nine somatic sample simulators (Synggen, BAMSurgeon, SVEngine, VarSim, Xome-Blender, tHapMix, Pysim-sv, SCNVSim, HeteroGenesis) assessing their ability to control biological parameters, including variants characteristics (type, number, position, length, content, zygosity), and sample characteristics (clonality, contamination); and technical parameters, including reads characteristics (sequencing errors, coverage, base qualities). No single simulator provided complete control over both biological and technical parameters, nor guidance on tuning biological parameters for cancer-specific simulations. Consequently, we developed MOV&RSim, a novel simulator that leverages data-driven information to set variants and reads characteristics, producing realistic tumoral samples, and providing full control on biological and technical parameters. Additionally, we leveraged well-annotated variant databases to create cancer-specific presets that inform the simulator’s parameters for 21 cancer types. CONCLUSION: This new simulator, containerised with Docker and freely available for academic use, empowers users to define each biological parameter of a tumoral genome and faithfully replicates the variability of technical noise observed in real sequencing reads. The proposed simulator and presets represent the most adaptable and comprehensive framework currently available for generating tumor samples, enabling comprehensive benchmarking and, ultimately, the optimization of somatic variant callers across diverse cancer types.
Francesca Longhin, Giacomo Baruzzo, Enidia Hazizaj, Diego Boscarino, Dino Paladin, Barbara Di Camillo
BMC Bioinform.2
2022 Identify, quantify and characterize cellular communication from single-cell RNA sequencing data with scSeqComm
abstract
MOTIVATION: Recently, single-cell RNA-seq (scRNA-seq) data have been used to study cellular communication. Most bioinformatics methods infer only the intercellular signaling between groups of cells, mainly exploiting ligand-receptor expression levels. Only few methods consider the entire intercellular + intracellular signaling, mainly inferring lists/networks of signaling involved genes. RESULTS: Here, we present scSeqComm, a computational method to identify and quantify the evidence of ongoing intercellular and intracellular signaling from scRNA-seq data, and at the same time providing a functional characterization of the inferred cellular communication. The possibility to quantify the evidence of ongoing communication assists the prioritization of the results, while the combined evidence of both intercellular and intracellular signaling increase the reliability of inferred communication. The application to a scRNA-seq dataset of tumor microenvironment, the agreement with independent bioinformatics analysis, the validation using spatial transcriptomics data and the comparison with state-of-the-art intercellular scoring schemes confirmed the robustness and reliability of the proposed method. AVAILABILITY AND IMPLEMENTATION: scSeqComm R package is freely available at https://gitlab.com/sysbiobig/scseqcomm and https://sysbiobig.dei.unipd.it/software/#scSeqComm. Submitted software version and test data are available in Zenodo, at https://dx.doi.org/10.5281/zenodo.5833298. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giacomo Baruzzo, Giulia Cesaro, Barbara Di Camillo
Bioinform.1
2022 Investigating differential abundance methods in microbiome data: A benchmark study
abstract
The development of increasingly efficient and cost-effective high throughput DNA sequencing techniques has enhanced the possibility of studying complex microbial systems. Recently, researchers have shown great interest in studying the microorganisms that characterise different ecological niches. Differential abundance analysis aims to find the differences in the abundance of each taxa between two classes of subjects or samples, assigning a significance value to each comparison. Several bioinformatic methods have been specifically developed, taking into account the challenges of microbiome data, such as sparsity, the different sequencing depth constraint between samples and compositionality. Differential abundance analysis has led to important conclusions in different fields, from health to the environment. However, the lack of a known biological truth makes it difficult to validate the results obtained. In this work we exploit metaSPARSim, a microbial sequencing count data simulator, to simulate data with differential abundance features between experimental groups. We perform a complete comparison of recently developed and established methods on a common benchmark with great effort to the reliability of both the simulated scenarios and the evaluation metrics. The performance overview includes the investigation of numerous scenarios, studying the effect on methods' results on the main covariates such as sample size, percentage of differentially abundant features, sequencing depth, feature variability, normalisation approach and ecological niches. Mainly, we find that methods show a good control of the type I error and, generally, also of the false discovery rate at high sample size, while recall seem to depend on the dataset and sample size.
Marco Cappellato, Giacomo Baruzzo, Barbara Di Camillo
PLoS Comput. Biol.2
2021 ITSoneWB: profiling global taxonomic diversity of eukaryotic communities on Galaxy
abstract
SUMMARY: ITSoneWB (ITSone WorkBench) is a Galaxy-based bioinformatic environment where comprehensive and high-quality reference data are connected with established pipelines and new tools in an automated and easy-to-use service targeted at global taxonomic analysis of eukaryotic communities based on Internal Transcribed Spacer 1 variants high-throughput sequencing. AVAILABILITY AND IMPLEMENTATION: ITSoneWB has been deployed on the INFN-Bari ReCaS cloud facility and is freely available on the web at http://itsonewb.cloud.ba.infn.it/galaxy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco Antonio Tangaro, Giuseppe Defazio, Bruno Fosso, Flavio Licciulli, Giorgio Grillo, Giacinto Donvito, Enrico Lavezzo, Giacomo Baruzzo, Graziano Pesole, Monica Santamaria
Bioinform.8
2021 Beware to ignore the rare: how imputing zero-values can improve the quality of 16S rRNA gene studies results
abstract
BACKGROUND: 16S rRNA-gene sequencing is a valuable approach to characterize the taxonomic content of the whole bacterial population inhabiting a metabolic and spatial niche, providing an important opportunity to study bacteria and their role in many health and environmental mechanisms. The analysis of data produced by amplicon sequencing, however, brings very specific methodological issues that need to be properly addressed to obtain reliable biological conclusions. Among these, 16S count data tend to be very sparse, with many null values reflecting species that are present but got unobserved due to the multiplexing constraints. However, current data workflows do not consider a step in which the information about unobserved species is recovered. RESULTS: In this work, we evaluate for the first time the effects of introducing in the 16S data workflow a new preprocessing step, zero-imputation, to recover this lost information. Due to the lack of published zero-imputation methods specifically designed for 16S count data, we considered a set of zero-imputation strategies available for other frameworks, and benchmarked them using in silico 16S count data reflecting different experimental designs. Additionally, we assessed the effect of combining zero-imputation and normalization, i.e. the only preprocessing step in current 16S workflow. Overall, we benchmarked 35 16S preprocessing pipelines assessing their ability to handle data sparsity, identify species presence/absence, recovery sample proportional abundance distributions, and improve typical downstream analyses such as computation of alpha and beta diversity indices and differential abundance analysis. CONCLUSIONS: The results clearly show that 16S data analysis greatly benefits from a properly-performed zero-imputation step, despite the choice of the right zero-imputation method having a pivotal role. In addition, we identify a set of best-performing pipelines that could be a valuable indication for data analysts.
Giacomo Baruzzo, Ilaria Patuzzi, Barbara Di Camillo
BMC Bioinform.1
2020 SPARSim single cell: a count data simulator for scRNA-seq data
abstract
MOTIVATION: Single cell RNA-seq (scRNA-seq) count data show many differences compared with bulk RNA-seq count data, making the application of many RNA-seq pre-processing/analysis methods not straightforward or even inappropriate. For this reason, the development of new methods for handling scRNA-seq count data is currently one of the most active research fields in bioinformatics. To help the development of such new methods, the availability of simulated data could play a pivotal role. However, only few scRNA-seq count data simulators are available, often showing poor or not demonstrated similarity with real data. RESULTS: In this article we present SPARSim, a scRNA-seq count data simulator based on a Gamma-Multivariate Hypergeometric model. We demonstrate that SPARSim allows to generate count data that resemble real data in terms of count intensity, variability and sparsity, performing comparably or better than one of the most used scRNA-seq simulator, Splat. In particular, SPARSim simulated count matrices well resemble the distribution of zeros across different expression intensities observed in real count data. AVAILABILITY AND IMPLEMENTATION: SPARSim R package is freely available at http://sysbiobig.dei.unipd.it/? q=SPARSim and at https://gitlab.com/sysbiobig/sparsim. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giacomo Baruzzo, Ilaria Patuzzi, Barbara Di Camillo
Bioinform.1
2019 metaSPARSim: a 16S rRNA gene sequencing count data simulator
abstract
BACKGROUND: In the last few years, 16S rRNA gene sequencing (16S rDNA-seq) has seen a surprisingly rapid increase in election rate as a methodology to perform microbial community studies. Despite the considerable popularity of this technique, an exiguous number of specific tools are currently available for proper 16S rDNA-seq count data preprocessing and simulation. Indeed, the great majority of tools have been developed adapting methodologies previously used for bulk RNA-seq data, with poor assessment of their applicability in the metagenomics field. For such tools and the few ones specifically developed for 16S rDNA-seq data, performance assessment is challenging, mainly due to the complex nature of the data and the lack of realistic simulation models. In fact, to the best of our knowledge, no software thought for data simulation are available to directly obtain synthetic 16S rDNA-seq count tables that properly model heavy sparsity and compositionality typical of these data. RESULTS: In this paper we present metaSPARSim, a sparse count matrix simulator intended for usage in development of 16S rDNA-seq metagenomic data processing pipelines. metaSPARSim implements a new generative process that models the sequencing process with a Multivariate Hypergeometric distribution in order to realistically simulate 16S rDNA-seq count table, resembling real experimental data compositionality and sparsity. It provides ready-to-use count matrices and comes with the possibility to reproduce different pre-coded scenarios and to estimate simulation parameters from real experimental data. The tool is made available at http://sysbiobig.dei.unipd.it/?q=Software#metaSPARSimand https://gitlab.com/sysbiobig/metasparsim. CONCLUSION: metaSPARSim is able to generate count matrices resembling real 16S rDNA-seq data. The availability of count data simulators is extremely valuable both for methods developers, for which a ground truth for tools validation is needed, and for users who want to assess state of the art analysis tools for choosing the most accurate one. Thus, we believe that metaSPARSim is a valuable tool for researchers involved in developing, testing and using robust and reliable data analysis methods in the context of 16S rRNA gene sequencing.
Ilaria Patuzzi, Giacomo Baruzzo, Carmen Losasso, Antonia Ricci, Barbara Di Camillo
BMC Bioinform.2
2018 Optimizing PCR primers targeting the bacterial 16S ribosomal RNA gene
abstract
BACKGROUND: Targeted amplicon sequencing of the 16S ribosomal RNA gene is one of the key tools for studying microbial diversity. The accuracy of this approach strongly depends on the choice of primer pairs and, in particular, on the balance between efficiency, specificity and sensitivity in the amplification of the different bacterial 16S sequences contained in a sample. There is thus the need for computational methods to design optimal bacterial 16S primers able to take into account the knowledge provided by the new sequencing technologies. RESULTS: We propose here a computational method for optimizing the choice of primer sets, based on multi-objective optimization, which simultaneously: 1) maximizes efficiency and specificity of target amplification; 2) maximizes the number of different bacterial 16S sequences matched by at least one primer; 3) minimizes the differences in the number of primers matching each bacterial 16S sequence. Our algorithm can be applied to any desired amplicon length without affecting computational performance. The source code of the developed algorithm is released as the mopo16S software tool (Multi-Objective Primer Optimization for 16S experiments) under the GNU General Public License and is available at http://sysbiobig.dei.unipd.it/?q=Software#mopo16S . CONCLUSIONS: Results show that our strategy is able to find better primer pairs than the ones available in the literature according to all three optimization criteria. We also experimentally validated three of the primer pairs identified by our method on multiple bacterial species, belonging to different genera and phyla. Results confirm the predicted efficiency and the ability to maximize the number of different bacterial 16S sequences matched by primers.
Francesco Sambo, Francesca Finotello, Enrico Lavezzo, Giacomo Baruzzo, Giulia Masi, Elektra Peta, Marco Falda, Stefano Toppo, Luisa Barzon, Barbara Di Camillo
BMC Bioinform.4