VLDB 2026 Research / reviewers in the wild / expert
Simone G. Riva
dblp:299/2904
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-4400-9328ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GENESIS: Generating scRNA-Seq data from Multiome Gene ExpressionabstractSingle-cell technologies have significantly advanced our understanding of cellular heterogeneity by allowing the examination of individual cells at high resolution. Traditional single-cell RNA sequencing (scRNA-Seq) methods, which utilise whole cells, capture comprehensive RNA content. In contrast, emerging Multiome technologies, which simultaneously profile multiple omics such as gene expression (GEX) and chromatin accessibility, rely on nuclear RNA, potentially missing key cytoplasmic information. This discrepancy results in substantial technical and biological differences between GEX and scRNA-Seq datasets, making it challenging to integrate the data and perform downstream tasks, such as cell-type classification. To address this challenge, we introduce GENESIS (Gene Expression Normalisation and Enhancement for Single-cell Integrated Sequencing), a novel computational framework designed to transform GEX data from Multiome experiments into enhanced, scRNA-Seq like profiles. Utilising advanced generative models—including Variational Autoencoders, Generative Adversarial Networks, and a tailored VAE_UNet architecture—GENESIS can generate high-quality data by modelling and compensating for the inherent differences between nuclear and cytoplasmic RNA. Our comprehensive evaluations show that GENESIS, particularly through the VAE_UNet model, generates synthetic scRNA-Seq data that closely resembles the resolution and biological accuracy of whole-cell sequencing, thereby improving downstream tasks, especially cell-type classification. Simone G. Riva, Brynelle Myers, Francesca Buffa, Andrea Tangherloni |
CIBCB | 1 |
| 2025 | REnformer, a single-cell ATAC-seq predicting model to investigate open chromatin sitesabstractGenome regulatory elements are fundamental to cellular identity and cell type specific gene expression. Understanding how the underlying genetic code is differentially utilised by different cell types is central to understanding human health and disease. To better understand how DNA encodes genome regulatory elements such as promoters, enhancers, and boundary elements, we leverage the Enformer gene expression and epigenetic prediction model. We used transfer learning with high quality single cell ATAC datasets to develop REnformer, a model to predict chromatin accessibility. By introducing a benchmark for comparing performances against Enformer model, REnformer significantly outperformed Enformer in terms of higher prediction outcomes and lower error rates in all extensive analyses shown; introducing these benchmarks allowed us, and possibly future works, to fairly compare such models. We further tested REnformer by predicting the effects of a well characterised α-thalassemia variant and found that the prediction aligned with the observed change in genome regulatory element, previously validated. We conclude that REnformer is and can be a state-of-the-art tool to predict cell type specific regulatory elements and interrogate the effect of genome variation in health and disease. Simone G. Riva, Edward Sanders, Nicolò Stranieri, E. Ravza Gür, Matthew Baxter, Jim R. Hughes |
CIBCB | 1 |
| 2025 | Modelling constant and cell-type specific CTCF sites by using Convolutional Neural NetworksabstractUnderstanding gene regulation is crucial to understanding human health and disease. One of the most important regulators of gene expression is the DNA binding protein CTCF. Many CTCF binding sites are consistent across cell types, however some are highly cell-type specific. The mechanisms by which CTCF can bind different genomic sites in different cells are poorly understood. One of the challenges is the vast number of potential CTCF binding sites across the 3 billion base pairs of the human genome. Finding patterns in datasets of this size is difficult for the human brain, but may be amenable to modelling using convolutional neural networks (CNN). We therefore designed and trained a CNN to model cell-type constant and cell-type specific CTCF binding sites across 33 distinct human cell types. The model achieved a micro and macro averaged AUC of 0.91 and 0.90 respectively, demonstrating a high level of accuracy in predicting CTCF binding across the different cell types. To test the effectiveness of the model we compared CTCF predictions for two cell types with highly different biological phenotypes (endothelial cells and neutrophils). Overall the model attained 81% accuracy for endothelial cells and 77% for neutrophils. These results demonstrate that it is possible to accurately predict cell-type specific CTCF binding sites based on genetic code alone. We believe this model will have future applications in understanding CTCF-mediated gene regulation in healthy and diseased cell states. Edward Sanders, Simone G. Riva, Emily Georgiades, Matthew Baxter, Jim R. Hughes |
IJCNN | 2 |
| 2023 | Consensus Clustering Strategy for Cell Type Assignments of scRNA-seq DataabstractCell type annotation is a crucial step for analyzing single-cell RNA sequencing data. Among others, single-cell Automatic Labeling of cell POpulations (scALPO) is a computational pipeline developed to automatically assign the cell types to the identified clusters in scRNA-seq data. Different from most of the approaches, scALPO relies only on the information on marker genes from published literature. Specifically, after the definition of the dataset obtained from gene information retrieved from online databases, the Leiden clustering algorithm is executed to partition cells that are finally annotated. Since the Leiden algorithm might struggle to obtain a reliable outcome under certain circumstances, in this work, we include several clustering algorithms in scALPO, and we propose a pseudo-voting consensus approach that combines the outcome of a set of clustering algorithms. The results obtained on three different datasets show that the consensus approach can improve the cell type annotation without selecting a specific clustering algorithm that best suits the data under investigation. Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Francesca Buffa, Andrea Tangherloni |
CIBCB | 1 |
| 2023 | MAGNETO: Cell type marker panel generator from single-cell transcriptomic dataabstractSingle-cell RNA sequencing experiments produce data useful to identify different cell types, including uncharacterized and rare ones. This enables us to study the specific functional roles of these cells in different microenvironments and contexts. After identifying a (novel) cell type of interest, it is essential to build succinct marker panels, composed of a few genes referring to cell surface proteins and clusters of differentiation molecules, able to discriminate the desired cells from the other cell populations. In this work, we propose a fully-automatic framework called MAGNETO, which can help construct optimal marker panels starting from a single-cell gene expression matrix and a cell type identity for each cell. MAGNETO builds effective marker panels solving a tailored bi-objective optimization problem, where the first objective regards the identification of the genes able to isolate a specific cell type, while the second conflicting objective concerns the minimization of the total number of genes included in the panel. Our results on three public datasets show that MAGNETO can identify marker panels that identify the cell populations of interest better than state-of-the-art approaches. Finally, by fine-tuning MAGNETO, our results demonstrate that it is possible to obtain marker panels with different specificity levels. Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Francesca Buffa, Paolo Cazzaniga |
J. Biomed. Informatics | 2 |
| 2022 | Predicting and Characterizing Legal Claims of Hospitals with Computational Intelligence: the Legal and Ethical ImplicationsabstractIn this paper we propose a fuzzy logic-based approach to analyze UK National Health Service (NHS) public administrative data related to pre- and post-pandemic claims filed by patients, analyzing the legal and ethical issues connected to the use of Artificial Intelligence systems, including our own, to take critical decisions having a significant impact on patients, such as employing computational intelligence to justify the management choices related to Intensive Care Unit (ICU) bed allocation. Differently from previous papers, in this work we follow an unsupervised approach and, specifically, we perform an analysis of UK hospitals by means of a computational intelligence algorithm integrating Fuzzy C- Means and swarm intelligence. The dataset that we analyse allows us to compare pre- and post-pandemic data, to analyze the ethical and legal challenges of the use of computational intelligence for critical decision-making in the health care field. Chiara Gallese, Caro Fuchs, Simone G. Riva, Emanuela Foglia, Fabrizio Schettini, Lucrezia Ferrario, Elena Falletti, Marco S. Nobile |
CIBCB | 3 |
| 2022 | A Deep Learning Pipeline for the Automatic cell type Assignment of scRNA-seq DataabstractThe increasing number of single-cell transcriptomics and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism, as well as the onset of pathologies. In this context, cell type annotation represents a crucial step for the analysis of single-cell RNA sequencing data, which is usually performed by means of time-consuming and possibly biased manual processes, carried out by expert biologists. Recently, alternative computational tools have been proposed to realize an automatic cell identification either based on supervised or unsupervised Machine Learning approaches. These methods typically exploit gene expression data of curated marker gene databases to associate gene expression profiles of single cells with a cell type. In this paper, we propose a novel fully-automatic computational pipeline, named single-cell Automatic Labeling of cell POpulations (scALPO), which leverages a Long Short-Term Memory Neural Network to assign the cell types. Specifically, scALPO can label the provided clusters by simply relying on marker genes rather than gene expressions. Our results, obtained by considering two different datasets, show that scALPO outperforms the most promising state-of-the-art approaches (i.e., SCSA and scType), achieving a cell type annotation more similar to the manually-created ground truth. Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Andrea Tangherloni |
CIBCB | 1 |
| 2022 | Multi-objective Optimization for Marker Panel Identification in Single-cell DataabstractThe computational analyses of single-cell data, aimed at elucidating and characterizing the functional roles of known and putative novel cell types, are enabling a thorough understanding of the processes driving cell development and pathology progression. The isolation of specific cell types is a crucial step to perform detailed analyses but requires the identification of succinct marker panels, which include genes that refer to cell surface proteins and clusters of differentiation molecules. This still represents a challenging NP-hard computational problem, which can be tackled through global optimization techniques. In this work, we formulate the marker panel identification problem as a bi-objective optimization problem, where the first objective regards the capability of the marker panels to accurately discriminate different cell types, while the second objective is related to the number of genes to include in the panel. In particular, we compared the performance of two multi-objective optimization algorithms, as well as of Genetic Algorithms (GAs) when considering only the first objective, employing two different representations for the candidate solutions. Our results show that the multi-objective optimization algorithms are better than GAs, considering both the quality and the consistency of the obtained marker panels; moreover, the collected results point out that different representations of the candidate solutions have a relevant impact on the performance of the optimization algorithms. Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Paolo Cazzaniga |
CIBCB | 2 |
| 2022 | SMaSH: a scalable, general marker gene identification framework for single-cell RNA-sequencingabstractBACKGROUND: Single-cell RNA-sequencing is revolutionising the study of cellular and tissue-wide heterogeneity in a large number of biological scenarios, from highly tissue-specific studies of disease to human-wide cell atlases. A central task in single-cell RNA-sequencing analysis design is the calculation of cell type-specific genes in order to study the differential impact of different replicates (e.g. tumour vs. non-tumour environment) on the regulation of those genes and their associated networks. The crucial task is the efficient and reliable calculation of such cell type-specific 'marker' genes. These optimise the ability of the experiment to isolate highly-specific cell phenotypes of interest to the analyser. However, while methods exist that can calculate marker genes from single-cell RNA-sequencing, no such method places emphasise on specific cell phenotypes for downstream study in e.g. differential gene expression or other experimental protocols (spatial transcriptomics protocols for example). Here we present SMaSH, a general computational framework for extracting key marker genes from single-cell RNA-sequencing data which reliably characterise highly-specific and niche populations of cells in numerous different biological data-sets. RESULTS: SMaSH extracts robust and biologically well-motivated marker genes, which characterise a given single-cell RNA-sequencing data-set better than existing computational approaches for general marker gene calculation. We demonstrate the utility of SMaSH through its substantial performance improvement over several existing methods in the field. Furthermore, we evaluate the SMaSH markers on spatial transcriptomics data, demonstrating they identify highly localised compartments of the mouse cortex. CONCLUSION: SMaSH is a new methodology for calculating robust markers genes from large single-cell RNA-sequencing data-sets, and has implications for e.g. effective gene identification for probe design in downstream analyses spatial transcriptomics experiments. SMaSH has been fully-integrated with the ScanPy framework and provides a valuable bioinformatics tool for cell type characterisation and validation in every-growing data-sets spanning over 50 different cell types across hundreds of thousands of cells. M. E. Nelson, Simone G. Riva, Ana Cvejic |
BMC Bioinform. | 2 |
| 2021 | Integration of Multiple scRNA-Seq Datasets on the Autoencoder Latent SpaceabstractThe application of single-cell transcriptomic sequencing technologies, such as single-cell RNA sequencing (scRNA-Seq), have witnessed in recent years a dramatic increase, allowing for the elucidation of the molecular processes driving both normal cell development and the onset of pathologies. In particular, scRNA-Seq can be exploited to investigate cell heterogeneity at single-cell resolution, and to identify the variety of known and putatively novel cell populations, which can potentially have different functional roles in different contexts. However, the heterogeneity among cells of the same cell-type can make the integration of multiple scRNA-Seq datasets a challenging task. In this context, technical non-negligible batch effects in the datasets—which may arise from the sequencing technology employed and from the size of the experiment—must be considered to realize a correct data integration. In this work, we present a novel strategy based on Autoencoders (AEs) for the integration of multiple scRNA-Seq datasets, whose performance is compared with different integration strategies that do not exploit a batch effect removal step, which might introduce artifacts in the datasets. Our results, obtained by considering 3 different datasets, suggest that AEs represent a suitable strategy for the integration of scRNA-Seq datasets, achieving better performance than other approaches, i.e., Scanorama, Ingest, and Seurat, in most of the cases. Simone G. Riva, Paolo Cazzaniga, Andrea Tangherloni |
BIBM | 1 |
| 2021 | The Impact of Representation on the Optimization of Marker Panels for Single-cell RNA DataabstractThe increasing number of single-cell transcriptomic and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism as well as the onset of pathologies. These studies continuously refine the functional roles of known cell populations, and provide their characterization as soon as putatively novel cell populations are detected. In order to isolate the cell populations for further tailored analysis, succinct marker panels—composed of a few cell surface proteins and clusters of differentiation molecules—must be identified. The identification of these marker panels is a challenging computational problem due to its intrinsic combinatorial nature, which makes it an NP-hard problem. Genetic Algorithms (GAs) have been successfully used in Bioinformatics and other biomedical applications to tackle combinatorial problems. We present here a GA-based approach to solve the problem of the identification of succinct marker panels. Since the performance of a GA is strictly related to the representation of the candidate solutions, we propose and compare three alternative representations, able to implicitly introduce different constraints on the search space. For each representation, we perform a fine-tuning of the parameter settings to calibrate the GA, and we show that different representations yield different performance, where the most relaxed representations— in which the GA can also evolve the number of genes in the panel—turn out to be the more effective, especially in the case of 0-knowledge problems. Our results also show that the marker panels identified by GAs can outperform manually curated solutions. Andrea Tangherloni, Simone G. Riva, Simone Spolaor, Daniela Besozzi, Marco S. Nobile, Paolo Cazzaniga |
CEC | 2 |