Francesca Buffa

dblp:98/10784 · also Francesca M. Buffa · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0003-0409-406XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Knowledge-Enriched Cell-Type Annotation in Single-Cell Transcriptomics via LLM Embeddings
abstract
Single-cell RNA sequencing (scRNA-seq) has profoundly reshaped our understanding of cellular diversity and functionality; however, accurate cell-type annotation is required for biological interpretation. Current annotation methods, which are predominantly reliant on gene expression alone or manual curation, suffer from subjectivity and a limited biological context. Here, we introduce a novel approach that integrates textual biological knowledge via gene embeddings, derived from fine-tuning Large Language Models (LLMs), with gene counts to enrich the input space for supervised models in automatic cell-type classification. In particular, we trained an XGBoost model and a multi-layer perceptron (MLP) to automatically classify the cell populations. We demonstrate that combining Modern-BERT embeddings and raw counts enhances the performance of MLPs, particularly in complex classification scenarios that involve subtle cell-subtype distinctions. Our results also show that ModernBERT generated better embeddings than smaller LLM architectures, underlining the value of enriched, biologically informed embeddings. By embedding prior knowledge from curated biological databases and literature, our approach enhances the MLP’s ability to distinguish sub-cell populations and biological signals. This work provides a scalable framework for integrating broader biological context into scRNA-seq analyses, offering new opportunities for downstream tasks such as gene regulatory network inference and cross-species annotation.
Andrea Fabbricatore, Francesca Buffa, Andrea Tangherloni
CIBCB2
2025 GENESIS: Generating scRNA-Seq data from Multiome Gene Expression
abstract
Single-cell technologies have significantly advanced our understanding of cellular heterogeneity by allowing the examination of individual cells at high resolution. Traditional single-cell RNA sequencing (scRNA-Seq) methods, which utilise whole cells, capture comprehensive RNA content. In contrast, emerging Multiome technologies, which simultaneously profile multiple omics such as gene expression (GEX) and chromatin accessibility, rely on nuclear RNA, potentially missing key cytoplasmic information. This discrepancy results in substantial technical and biological differences between GEX and scRNA-Seq datasets, making it challenging to integrate the data and perform downstream tasks, such as cell-type classification. To address this challenge, we introduce GENESIS (Gene Expression Normalisation and Enhancement for Single-cell Integrated Sequencing), a novel computational framework designed to transform GEX data from Multiome experiments into enhanced, scRNA-Seq like profiles. Utilising advanced generative models—including Variational Autoencoders, Generative Adversarial Networks, and a tailored VAE_UNet architecture—GENESIS can generate high-quality data by modelling and compensating for the inherent differences between nuclear and cytoplasmic RNA. Our comprehensive evaluations show that GENESIS, particularly through the VAE_UNet model, generates synthetic scRNA-Seq data that closely resembles the resolution and biological accuracy of whole-cell sequencing, thereby improving downstream tasks, especially cell-type classification.
Simone G. Riva, Brynelle Myers, Francesca Buffa, Andrea Tangherloni
CIBCB3
2025 GeneFEAST: the pivotal, gene-centric step in functional enrichment analysis interpretation
abstract
SUMMARY: GeneFEAST, implemented in Python, is a gene-centric functional enrichment analysis summarization and visualization tool that can be applied to large functional enrichment analysis (FEA) results arising from upstream FEA pipelines. It produces a systematic, navigable HTML report, making it easy to identify sets of genes putatively driving multiple enrichments and to explore gene-level quantitative data first used to identify input genes. Further, GeneFEAST can juxtapose FEA results from multiple studies, making it possible to highlight patterns of gene expression amongst genes that are differentially expressed in at least one of multiple conditions, and which give rise to shared enrichments under those conditions. Thus, GeneFEAST offers a novel, effective way to address the complexities of linking up many overlapping FEA results to their underlying genes and data, advancing gene-centric hypotheses, and providing pivotal information for downstream validation experiments. AVAILABILITY AND IMPLEMENTATION: GeneFEAST GitHub repository: https://github.com/avigailtaylor/GeneFEAST; Zenodo record: 10.5281/zenodo.14753734; Python Package Index: https://pypi.org/project/genefeast; Docker container: ghcr.io/avigailtaylor/genefeast.
Avigail Taylor, Valentine M. Macaulay, Matthieu J. Miossec, Anand K. Maurya, Francesca Buffa
Bioinform.5
2024 A Modified EACOP Implementation for Real-Parameter Single Objective Optimization Problems
abstract
Evolutionary algorithms are effective techniques for optimizing non-linear and complex high-dimensional problems. However, most of them require a precise fine-tuning of their functioning settings to achieve satisfactory results. In this work, we propose a modified version of an evolutionary approach called the Evolutionary Algorithm for COmplex-process oPtimization (EACOP), designed to have a limited number of hyper-parameters. The base version of EACOP (bEACOP) combines different strategies, including the scatter search methodology, local searches, and a novel combination method based on path relinking to balance the exploration and exploitation phases. Our improved version (iEACOP) intensifies the exploration phase to escape from suboptimal search space areas where, on the contrary, bEACOP gets stuck. Our results show that iEACOP outperforms bEACOP on 27 out of 29 CEC 2017 test suite benchmark functions, exhibiting comparable performance against the three best algorithms of the CEC 2017 competition on single-objective bound-constrained real-parameter numerical optimization. The source code of bEACOP and iEACOP will be made publicly available on GitHub upon acceptance.
Andrea Tangherloni, Vasco Coelho, Francesca Buffa, Paolo Cazzaniga
CEC3
2024 Forest-based Evolutionary Algorithm for Reconstructing Boolean Gene Regulatory Networks
abstract
Gene Regulatory Networks (GRNs) play a fundamental role in orchestrating the expression of our genes through complex interactions between DNA, RNA, proteins, and other molecules. Accurately reconstructing such networks from gene expression data is a critical yet challenging task in Systems Biology due to their intricate nature and limited data availability. In this work, we introduce a novel Forest-based Evolutionary Algorithm (FP) designed for reconstructing Boolean GRNs from time series data of gene expressions. Unlike traditional methods that struggle with scalability and accurate representation of regulatory interactions, FP utilizes a forest structure where each tree represents the logical relationships between genes, enhancing the model’s capacity to depict complex networks efficiently. Our comprehensive testing indicates that FP rapidly converges towards potential solutions within a limited number of generations, although a higher fitness score does not always equate to a more accurate GRN representation. Implementing mini-batching techniques, inspired by their effectiveness in gradient descent optimization, shows promise in improving computational efficiency without sacrificing performance. A comparative analysis against the main state-of-the-art approaches reveals FP’s tendency towards conservative predictions, emphasizing precision over recall, making it particularly suitable for contexts where the cost of false positives is high. These initial results suggest that FP is a robust and efficient tool for GRN inference.
Nicolò Stranieri, Francesca Buffa, Andrea Tangherloni
CIBCB2
2024 A Fast Feature Selection for Interpretable Modeling Based on Fuzzy Inference Systems
abstract
Large datasets are often beneficial for the generation of predictive models using machine learning approaches. However, it is often the case that not all variables in the dataset contain useful information. In fact, some variables might be useless, redundant, misleading, or even harmful to performance, both in terms of accuracy and computational effort. Because of that, Feature Selection (FS) is one of the most delicate and important steps in machine learning. This is even more relevant in the case of interpretable models based on Fuzzy Inference Systems (FIS). The reasons are two-fold: on the one hand, FIS are generally built on top of a data partitioning based on clustering, which can suffer from high dimensionality; on the other hand, the knowledge base of the FIS, to be concretely understandable, should not contain rules involving too many variables. FS can be performed using multiple approaches, most notably filter and wrapper methods. The latter are often based on evolutionary algorithms, where a population of candidate solutions (each representing a possible set of selected variables) evolves towards the optimal selection. Although wrapper methods can be effective, they are, in general, computationally expensive. In this work, we propose a completely different – and more computationally effective – algorithm based on Random Forest (RF) models. Specifically, we exploit RFs to rank variables according to their importance. Then, we use that information to perform a statistical analysis and determine the minimal set of features necessary to build an accurate FIS. We show the effectiveness of our approach by using two (semi)synthetic datasets built on real-world datasets, and we validate our approach by applying the FS method to a medical dataset.
Andrea Tangherloni, Paolo Cazzaniga, Nicolò Stranieri, Francesca Buffa, Marco S. Nobile
CIBCB4
2024 Metabolic symbiosis between oxygenated and hypoxic tumour cells: An agent-based modelling study
abstract
Deregulated metabolism is one of the hallmarks of cancer. It is well-known that tumour cells tend to metabolize glucose via glycolysis even when oxygen is available and mitochondrial respiration is functional. However, the lower energy efficiency of aerobic glycolysis with respect to mitochondrial respiration makes this behaviour, namely the Warburg effect, counter-intuitive, although it has now been recognized as source of anabolic precursors. On the other hand, there is evidence that oxygenated tumour cells could be fuelled by exogenous lactate produced from glycolysis. We employed a multi-scale approach that integrates multi-agent modelling, diffusion-reaction, stoichiometric equations, and Boolean networks to study metabolic cooperation between hypoxic and oxygenated cells exposed to varying oxygen, nutrient, and inhibitor concentrations. The results show that the cooperation reduces the depletion of environmental glucose, resulting in an overall advantage of using aerobic glycolysis. In addition, the oxygen level was found to be decreased by symbiosis, promoting a further shift towards anaerobic glycolysis. However, the oxygenated and hypoxic populations may gradually reach quasi-equilibrium. A sensitivity analysis using Latin hypercube sampling and partial rank correlation shows that the symbiotic dynamics depends on properties of the specific cell such as the minimum glucose level needed for glycolysis. Our results suggest that strategies that block glucose transporters may be more effective to reduce tumour growth than those blocking lactate intake transporters.
Pahala Gedara Jayathilake, Pedro Victori, Clara E. Pavillet, Chang Heon Lee, Dimitrios Voukantsis, Ana Miar, Anjali Arora, Adrian L. Harris, Karl J. Morten, Francesca Buffa
PLoS Comput. Biol.10
2023 Consensus Clustering Strategy for Cell Type Assignments of scRNA-seq Data
abstract
Cell type annotation is a crucial step for analyzing single-cell RNA sequencing data. Among others, single-cell Automatic Labeling of cell POpulations (scALPO) is a computational pipeline developed to automatically assign the cell types to the identified clusters in scRNA-seq data. Different from most of the approaches, scALPO relies only on the information on marker genes from published literature. Specifically, after the definition of the dataset obtained from gene information retrieved from online databases, the Leiden clustering algorithm is executed to partition cells that are finally annotated. Since the Leiden algorithm might struggle to obtain a reliable outcome under certain circumstances, in this work, we include several clustering algorithms in scALPO, and we propose a pseudo-voting consensus approach that combines the outcome of a set of clustering algorithms. The results obtained on three different datasets show that the consensus approach can improve the cell type annotation without selecting a specific clustering algorithm that best suits the data under investigation.
Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Francesca Buffa, Andrea Tangherloni
CIBCB4
2023 MAGNETO: Cell type marker panel generator from single-cell transcriptomic data
abstract
Single-cell RNA sequencing experiments produce data useful to identify different cell types, including uncharacterized and rare ones. This enables us to study the specific functional roles of these cells in different microenvironments and contexts. After identifying a (novel) cell type of interest, it is essential to build succinct marker panels, composed of a few genes referring to cell surface proteins and clusters of differentiation molecules, able to discriminate the desired cells from the other cell populations. In this work, we propose a fully-automatic framework called MAGNETO, which can help construct optimal marker panels starting from a single-cell gene expression matrix and a cell type identity for each cell. MAGNETO builds effective marker panels solving a tailored bi-objective optimization problem, where the first objective regards the identification of the genes able to isolate a specific cell type, while the second conflicting objective concerns the minimization of the total number of genes included in the panel. Our results on three public datasets show that MAGNETO can identify marker panels that identify the cell populations of interest better than state-of-the-art approaches. Finally, by fine-tuning MAGNETO, our results demonstrate that it is possible to obtain marker panels with different specificity levels.
Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Francesca Buffa, Paolo Cazzaniga
J. Biomed. Informatics4
2020 Guest Editorial Data Science in Smart Healthcare: Challenges and Opportunities
abstract
The fifteen articles in this special section focus on data science used in smart healthcare applications. A shift toward a data-driven socio-economic health model is occurring. This is the result of the increased volume, velocity and variety of data collected from the public and private sector in healthcare, and biology in general. In the past five-years, there has been an impressive development of computational intelligence and informatics methods for application to health and biomedical science. However, the effective use of data to address the scale and scope of human health problems has yet to realize its full potential. The barriers limiting the impact of practical application of standard data mining and machine learning methods have been inherent to the characteristics of health data. Besides the volume of the data (‘big data’), these are challenging due to their heterogeneity, complexity, variability and dynamic nature. Finally, data management and interpretability of the results have been limited by practical challenges in implementing new and also existing standards across the different health providers and research institutions. The scope of this Special issue is to discuss some of these challenges and opportunities in health and biological data science, with particular focus on the infrastructure, software, methods and algorithms needed to analyze large datasets in biological and clinical research.
Barbara Di Camillo, Giuseppe Nicosia, Francesca Buffa, Benny P. L. Lo
IEEE J. Biomed. Health Informatics3