Andrea Tangherloni

dblp:190/2536 · DBLP profile ↗
← Back
30ranked-venue papers
13as first author
17since 2021 · last 2025
0000-0002-5856-4453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Knowledge-Enriched Cell-Type Annotation in Single-Cell Transcriptomics via LLM Embeddings
abstract
Single-cell RNA sequencing (scRNA-seq) has profoundly reshaped our understanding of cellular diversity and functionality; however, accurate cell-type annotation is required for biological interpretation. Current annotation methods, which are predominantly reliant on gene expression alone or manual curation, suffer from subjectivity and a limited biological context. Here, we introduce a novel approach that integrates textual biological knowledge via gene embeddings, derived from fine-tuning Large Language Models (LLMs), with gene counts to enrich the input space for supervised models in automatic cell-type classification. In particular, we trained an XGBoost model and a multi-layer perceptron (MLP) to automatically classify the cell populations. We demonstrate that combining Modern-BERT embeddings and raw counts enhances the performance of MLPs, particularly in complex classification scenarios that involve subtle cell-subtype distinctions. Our results also show that ModernBERT generated better embeddings than smaller LLM architectures, underlining the value of enriched, biologically informed embeddings. By embedding prior knowledge from curated biological databases and literature, our approach enhances the MLP’s ability to distinguish sub-cell populations and biological signals. This work provides a scalable framework for integrating broader biological context into scRNA-seq analyses, offering new opportunities for downstream tasks such as gene regulatory network inference and cross-species annotation.
Andrea Fabbricatore, Francesca Buffa, Andrea Tangherloni
CIBCB3
2025 GENESIS: Generating scRNA-Seq data from Multiome Gene Expression
abstract
Single-cell technologies have significantly advanced our understanding of cellular heterogeneity by allowing the examination of individual cells at high resolution. Traditional single-cell RNA sequencing (scRNA-Seq) methods, which utilise whole cells, capture comprehensive RNA content. In contrast, emerging Multiome technologies, which simultaneously profile multiple omics such as gene expression (GEX) and chromatin accessibility, rely on nuclear RNA, potentially missing key cytoplasmic information. This discrepancy results in substantial technical and biological differences between GEX and scRNA-Seq datasets, making it challenging to integrate the data and perform downstream tasks, such as cell-type classification. To address this challenge, we introduce GENESIS (Gene Expression Normalisation and Enhancement for Single-cell Integrated Sequencing), a novel computational framework designed to transform GEX data from Multiome experiments into enhanced, scRNA-Seq like profiles. Utilising advanced generative models—including Variational Autoencoders, Generative Adversarial Networks, and a tailored VAE_UNet architecture—GENESIS can generate high-quality data by modelling and compensating for the inherent differences between nuclear and cytoplasmic RNA. Our comprehensive evaluations show that GENESIS, particularly through the VAE_UNet model, generates synthetic scRNA-Seq data that closely resembles the resolution and biological accuracy of whole-cell sequencing, thereby improving downstream tasks, especially cell-type classification.
Simone G. Riva, Brynelle Myers, Francesca Buffa, Andrea Tangherloni
CIBCB4
2025 We Are Sending You Back... to the Optimum! Fuzzy Time Travel Particle Swarm Optimization
Daniele M. Papetti, Andrea Tangherloni, Vasco Coelho, Daniela Besozzi, Paolo Cazzaniga, Marco S. Nobile
EvoApplications (2)2
2024 A Modified EACOP Implementation for Real-Parameter Single Objective Optimization Problems
abstract
Evolutionary algorithms are effective techniques for optimizing non-linear and complex high-dimensional problems. However, most of them require a precise fine-tuning of their functioning settings to achieve satisfactory results. In this work, we propose a modified version of an evolutionary approach called the Evolutionary Algorithm for COmplex-process oPtimization (EACOP), designed to have a limited number of hyper-parameters. The base version of EACOP (bEACOP) combines different strategies, including the scatter search methodology, local searches, and a novel combination method based on path relinking to balance the exploration and exploitation phases. Our improved version (iEACOP) intensifies the exploration phase to escape from suboptimal search space areas where, on the contrary, bEACOP gets stuck. Our results show that iEACOP outperforms bEACOP on 27 out of 29 CEC 2017 test suite benchmark functions, exhibiting comparable performance against the three best algorithms of the CEC 2017 competition on single-objective bound-constrained real-parameter numerical optimization. The source code of bEACOP and iEACOP will be made publicly available on GitHub upon acceptance.
Andrea Tangherloni, Vasco Coelho, Francesca Buffa, Paolo Cazzaniga
CEC1
2024 Forest-based Evolutionary Algorithm for Reconstructing Boolean Gene Regulatory Networks
abstract
Gene Regulatory Networks (GRNs) play a fundamental role in orchestrating the expression of our genes through complex interactions between DNA, RNA, proteins, and other molecules. Accurately reconstructing such networks from gene expression data is a critical yet challenging task in Systems Biology due to their intricate nature and limited data availability. In this work, we introduce a novel Forest-based Evolutionary Algorithm (FP) designed for reconstructing Boolean GRNs from time series data of gene expressions. Unlike traditional methods that struggle with scalability and accurate representation of regulatory interactions, FP utilizes a forest structure where each tree represents the logical relationships between genes, enhancing the model’s capacity to depict complex networks efficiently. Our comprehensive testing indicates that FP rapidly converges towards potential solutions within a limited number of generations, although a higher fitness score does not always equate to a more accurate GRN representation. Implementing mini-batching techniques, inspired by their effectiveness in gradient descent optimization, shows promise in improving computational efficiency without sacrificing performance. A comparative analysis against the main state-of-the-art approaches reveals FP’s tendency towards conservative predictions, emphasizing precision over recall, making it particularly suitable for contexts where the cost of false positives is high. These initial results suggest that FP is a robust and efficient tool for GRN inference.
Nicolò Stranieri, Francesca Buffa, Andrea Tangherloni
CIBCB3
2024 A Fast Feature Selection for Interpretable Modeling Based on Fuzzy Inference Systems
abstract
Large datasets are often beneficial for the generation of predictive models using machine learning approaches. However, it is often the case that not all variables in the dataset contain useful information. In fact, some variables might be useless, redundant, misleading, or even harmful to performance, both in terms of accuracy and computational effort. Because of that, Feature Selection (FS) is one of the most delicate and important steps in machine learning. This is even more relevant in the case of interpretable models based on Fuzzy Inference Systems (FIS). The reasons are two-fold: on the one hand, FIS are generally built on top of a data partitioning based on clustering, which can suffer from high dimensionality; on the other hand, the knowledge base of the FIS, to be concretely understandable, should not contain rules involving too many variables. FS can be performed using multiple approaches, most notably filter and wrapper methods. The latter are often based on evolutionary algorithms, where a population of candidate solutions (each representing a possible set of selected variables) evolves towards the optimal selection. Although wrapper methods can be effective, they are, in general, computationally expensive. In this work, we propose a completely different – and more computationally effective – algorithm based on Random Forest (RF) models. Specifically, we exploit RFs to rank variables according to their importance. Then, we use that information to perform a statistical analysis and determine the minimal set of features necessary to build an accurate FIS. We show the effectiveness of our approach by using two (semi)synthetic datasets built on real-world datasets, and we validate our approach by applying the FS method to a medical dataset.
Andrea Tangherloni, Paolo Cazzaniga, Nicolò Stranieri, Francesca Buffa, Marco S. Nobile
CIBCB1
2023 The Domination Game: Dilating Bubbles to Fill Up Pareto Fronts
abstract
Multi-objective optimization algorithms might struggle in finding optimal dominating solutions, especially in real-case scenarios where problems are generally characterized by non-separability, non-differentiability, and multi-modality issues. An effective strategy that already showed to improve the outcome of optimization algorithms consists in manipulating the search space, in order to explore its most promising areas. In this work, starting from a Pareto front identified by an optimization strategy, we exploit Local Bubble Dilation Functions (LBDFs) to manipulate a locally bounded region of the search space containing non-dominated solutions. We tested our approach on the benchmark functions included in the DTLZ and WFG suites, showing that the Pareto front obtained after the application of LBDFs is most of the time characterized by an increased hyper-volume value. Our results confirm that LBDFs are an effective means to identify additional non-dominated solutions that can improve the quality of the Pareto front.
Vasco Coelho, Daniele M. Papetti, Andrea Tangherloni, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CEC3
2023 Consensus Clustering Strategy for Cell Type Assignments of scRNA-seq Data
abstract
Cell type annotation is a crucial step for analyzing single-cell RNA sequencing data. Among others, single-cell Automatic Labeling of cell POpulations (scALPO) is a computational pipeline developed to automatically assign the cell types to the identified clusters in scRNA-seq data. Different from most of the approaches, scALPO relies only on the information on marker genes from published literature. Specifically, after the definition of the dataset obtained from gene information retrieved from online databases, the Leiden clustering algorithm is executed to partition cells that are finally annotated. Since the Leiden algorithm might struggle to obtain a reliable outcome under certain circumstances, in this work, we include several clustering algorithms in scALPO, and we propose a pseudo-voting consensus approach that combines the outcome of a set of clustering algorithms. The results obtained on three different datasets show that the consensus approach can improve the cell type annotation without selecting a specific clustering algorithm that best suits the data under investigation.
Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Francesca Buffa, Andrea Tangherloni
CIBCB5
2023 MAGNETO: Cell type marker panel generator from single-cell transcriptomic data
abstract
Single-cell RNA sequencing experiments produce data useful to identify different cell types, including uncharacterized and rare ones. This enables us to study the specific functional roles of these cells in different microenvironments and contexts. After identifying a (novel) cell type of interest, it is essential to build succinct marker panels, composed of a few genes referring to cell surface proteins and clusters of differentiation molecules, able to discriminate the desired cells from the other cell populations. In this work, we propose a fully-automatic framework called MAGNETO, which can help construct optimal marker panels starting from a single-cell gene expression matrix and a cell type identity for each cell. MAGNETO builds effective marker panels solving a tailored bi-objective optimization problem, where the first objective regards the identification of the genes able to isolate a specific cell type, while the second conflicting objective concerns the minimization of the total number of genes included in the panel. Our results on three public datasets show that MAGNETO can identify marker panels that identify the cell populations of interest better than state-of-the-art approaches. Finally, by fine-tuning MAGNETO, our results demonstrate that it is possible to obtain marker panels with different specificity levels.
Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Francesca Buffa, Paolo Cazzaniga
J. Biomed. Informatics1
2022 A Deep Learning Pipeline for the Automatic cell type Assignment of scRNA-seq Data
abstract
The increasing number of single-cell transcriptomics and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism, as well as the onset of pathologies. In this context, cell type annotation represents a crucial step for the analysis of single-cell RNA sequencing data, which is usually performed by means of time-consuming and possibly biased manual processes, carried out by expert biologists. Recently, alternative computational tools have been proposed to realize an automatic cell identification either based on supervised or unsupervised Machine Learning approaches. These methods typically exploit gene expression data of curated marker gene databases to associate gene expression profiles of single cells with a cell type. In this paper, we propose a novel fully-automatic computational pipeline, named single-cell Automatic Labeling of cell POpulations (scALPO), which leverages a Long Short-Term Memory Neural Network to assign the cell types. Specifically, scALPO can label the provided clusters by simply relying on marker genes rather than gene expressions. Our results, obtained by considering two different datasets, show that scALPO outperforms the most promising state-of-the-art approaches (i.e., SCSA and scType), achieving a cell type annotation more similar to the manually-created ground truth.
Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Andrea Tangherloni
CIBCB4
2022 Multi-objective Optimization for Marker Panel Identification in Single-cell Data
abstract
The computational analyses of single-cell data, aimed at elucidating and characterizing the functional roles of known and putative novel cell types, are enabling a thorough understanding of the processes driving cell development and pathology progression. The isolation of specific cell types is a crucial step to perform detailed analyses but requires the identification of succinct marker panels, which include genes that refer to cell surface proteins and clusters of differentiation molecules. This still represents a challenging NP-hard computational problem, which can be tackled through global optimization techniques. In this work, we formulate the marker panel identification problem as a bi-objective optimization problem, where the first objective regards the capability of the marker panels to accurately discriminate different cell types, while the second objective is related to the number of genes to include in the panel. In particular, we compared the performance of two multi-objective optimization algorithms, as well as of Genetic Algorithms (GAs) when considering only the first objective, employing two different representations for the candidate solutions. Our results show that the multi-objective optimization algorithms are better than GAs, considering both the quality and the consistency of the obtained marker panels; moreover, the collected results point out that different representations of the candidate solutions have a relevant impact on the performance of the optimization algorithms.
Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Paolo Cazzaniga
CIBCB1
2022 Salp Swarm Optimization: A critical review
Mauro Castelli, Luca Manzoni, Luca Mariot, Marco S. Nobile, Andrea Tangherloni
Expert Syst. Appl.5
2021 Integration of Multiple scRNA-Seq Datasets on the Autoencoder Latent Space
abstract
The application of single-cell transcriptomic sequencing technologies, such as single-cell RNA sequencing (scRNA-Seq), have witnessed in recent years a dramatic increase, allowing for the elucidation of the molecular processes driving both normal cell development and the onset of pathologies. In particular, scRNA-Seq can be exploited to investigate cell heterogeneity at single-cell resolution, and to identify the variety of known and putatively novel cell populations, which can potentially have different functional roles in different contexts. However, the heterogeneity among cells of the same cell-type can make the integration of multiple scRNA-Seq datasets a challenging task. In this context, technical non-negligible batch effects in the datasets—which may arise from the sequencing technology employed and from the size of the experiment—must be considered to realize a correct data integration. In this work, we present a novel strategy based on Autoencoders (AEs) for the integration of multiple scRNA-Seq datasets, whose performance is compared with different integration strategies that do not exploit a batch effect removal step, which might introduce artifacts in the datasets. Our results, obtained by considering 3 different datasets, suggest that AEs represent a suitable strategy for the integration of scRNA-Seq datasets, achieving better performance than other approaches, i.e., Scanorama, Ingest, and Seurat, in most of the cases.
Simone G. Riva, Paolo Cazzaniga, Andrea Tangherloni
BIBM3
2021 The Impact of Representation on the Optimization of Marker Panels for Single-cell RNA Data
abstract
The increasing number of single-cell transcriptomic and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism as well as the onset of pathologies. These studies continuously refine the functional roles of known cell populations, and provide their characterization as soon as putatively novel cell populations are detected. In order to isolate the cell populations for further tailored analysis, succinct marker panels—composed of a few cell surface proteins and clusters of differentiation molecules—must be identified. The identification of these marker panels is a challenging computational problem due to its intrinsic combinatorial nature, which makes it an NP-hard problem. Genetic Algorithms (GAs) have been successfully used in Bioinformatics and other biomedical applications to tackle combinatorial problems. We present here a GA-based approach to solve the problem of the identification of succinct marker panels. Since the performance of a GA is strictly related to the representation of the candidate solutions, we propose and compare three alternative representations, able to implicitly introduce different constraints on the search space. For each representation, we perform a fine-tuning of the parameter settings to calibrate the GA, and we show that different representations yield different performance, where the most relaxed representations— in which the GA can also evolve the number of genes in the panel—turn out to be the more effective, especially in the case of 0-knowledge problems. Our results also show that the marker panels identified by GAs can outperform manually curated solutions.
Andrea Tangherloni, Simone G. Riva, Simone Spolaor, Daniela Besozzi, Marco S. Nobile, Paolo Cazzaniga
CEC1
2021 Analysis of single-cell RNA sequencing data based on autoencoders
abstract
BACKGROUND: Single-cell RNA sequencing (scRNA-Seq) experiments are gaining ground to study the molecular processes that drive normal development as well as the onset of different pathologies. Finding an effective and efficient low-dimensional representation of the data is one of the most important steps in the downstream analysis of scRNA-Seq data, as it could provide a better identification of known or putatively novel cell-types. Another step that still poses a challenge is the integration of different scRNA-Seq datasets. Though standard computational pipelines to gain knowledge from scRNA-Seq data exist, a further improvement could be achieved by means of machine learning approaches. RESULTS: Autoencoders (AEs) have been effectively used to capture the non-linearities among gene interactions of scRNA-Seq data, so that the deployment of AE-based tools might represent the way forward in this context. We introduce here scAEspy, a unifying tool that embodies: (1) four of the most advanced AEs, (2) two novel AEs that we developed on purpose, (3) different loss functions. We show that scAEspy can be coupled with various batch-effect removal tools to integrate data by different scRNA-Seq platforms, in order to better identify the cell-types. We benchmarked scAEspy against the most used batch-effect removal tools, showing that our AE-based strategies outperform the existing solutions. CONCLUSIONS: scAEspy is a user-friendly tool that enables using the most recent and promising AEs to analyse scRNA-Seq data by only setting up two user-defined parameters. Thanks to its modularity, scAEspy can be easily extended to accommodate new AEs to further improve the downstream analysis of scRNA-Seq data. Considering the relevant results we achieved, scAEspy can be considered as a starting point to build a more comprehensive toolkit designed to integrate multi single-cell omics.
Andrea Tangherloni, Federico Ricciuti, Daniela Besozzi, Pietro Liò, Ana Cvejic
BMC Bioinform.1
2021 FiCoS: A fine-grained and coarse-grained GPU-powered deterministic simulator for biochemical networks
abstract
Mathematical models of biochemical networks can largely facilitate the comprehension of the mechanisms at the basis of cellular processes, as well as the formulation of hypotheses that can be tested by means of targeted laboratory experiments. However, two issues might hamper the achievement of fruitful outcomes. On the one hand, detailed mechanistic models can involve hundreds or thousands of molecular species and their intermediate complexes, as well as hundreds or thousands of chemical reactions, a situation generally occurring in rule-based modeling. On the other hand, the computational analysis of a model typically requires the execution of a large number of simulations for its calibration, or to test the effect of perturbations. As a consequence, the computational capabilities of modern Central Processing Units can be easily overtaken, possibly making the modeling of biochemical networks a worthless or ineffective effort. To the aim of overcoming the limitations of the current state-of-the-art simulation approaches, we present in this paper FiCoS, a novel "black-box" deterministic simulator that effectively realizes both a fine-grained and a coarse-grained parallelization on Graphics Processing Units. In particular, FiCoS exploits two different integration methods, namely, the Dormand-Prince and the Radau IIA, to efficiently solve both non-stiff and stiff systems of coupled Ordinary Differential Equations. We tested the performance of FiCoS against different deterministic simulators, by considering models of increasing size and by running analyses with increasing computational demands. FiCoS was able to dramatically speedup the computations up to 855×, showing to be a promising solution for the simulation and analysis of large-scale models of complex biological processes.
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Giulia Capitoli, Simone Spolaor, Leonardo Rundo, Giancarlo Mauri, Daniela Besozzi
PLoS Comput. Biol.1
2021 A CUDA-powered method for the feature extraction and unsupervised analysis of medical images
abstract
Abstract Image texture extraction and analysis are fundamental steps in computer vision. In particular, considering the biomedical field, quantitative imaging methods are increasingly gaining importance because they convey scientifically and clinically relevant information for prediction, prognosis, and treatment response assessment. In this context, radiomic approaches are fostering large-scale studies that can have a significant impact in the clinical practice. In this work, we present a novel method, called CHASM (Cuda, HAralick & SoM), which is accelerated on the graphics processing unit (GPU) for quantitative imaging analyses based on Haralick features and on the self-organizing map (SOM). The Haralick features extraction step relies upon the gray-level co-occurrence matrix, which is computationally burdensome on medical images characterized by a high bit depth. The downstream analyses exploit the SOM with the goal of identifying the underlying clusters of pixels in an unsupervised manner. CHASM is conceived to leverage the parallel computation capabilities of modern GPUs. Analyzing ovarian cancer computed tomography images, CHASM achieved up to $$\sim 19.5\times $$ ∼ 19.5 × and $$\sim 37\times $$ ∼ 37 × speed-up factors for the Haralick feature extraction and for the SOM execution, respectively, compared to the corresponding C++ coded sequential versions. Such computational results point out the potential of GPUs in the clinical research.
Leonardo Rundo, Andrea Tangherloni, Paolo Cazzaniga, Matteo Mistri, Simone Galimberti, Ramona Woitek, Evis Sala, Giancarlo Mauri, Marco S. Nobile
J. Supercomput.2
2020 Computational Intelligence for Life Sciences
abstract
Computational Intelligence (CI) is a computer science discipline encompassing the theory, design, development and application of biologically and linguistically derived computational paradigms. Traditionally, the main elements of CI are Evolutionary Computation, Swarm Intelligence, Fuzzy Logic, and Neural Networks. CI aims at proposing new algorithms able to solve complex computational problems by taking inspiration from natural phenomena. In an intriguing turn of events, these nature-inspired methods have been widely adopted to investigate a plethora of problems related to nature itself. In this paper we present a variety of CI methods applied to three problems in life sciences, highlighting their effectiveness: we describe how protein folding can be faced by exploiting Genetic Programming, the inference of haplotypes can be tackled using Genetic Algorithms, and the estimation of biochemical kinetic parameters can be performed by means of Swarm Intelligence. We show that CI methods can generate very high quality solutions, providing a sound methodology to solve complex optimization problems in life sciences.
Daniela Besozzi, Luca Manzoni, Marco S. Nobile, Simone Spolaor, Mauro Castelli, Leonardo Vanneschi, Paolo Cazzaniga, Stefano Ruberto, Leonardo Rundo, Andrea Tangherloni
Fundam. Informaticae10
2019 GenHap: a novel computational method based on genetic algorithms for haplotype assembly
abstract
BACKGROUND: In order to fully characterize the genome of an individual, the reconstruction of the two distinct copies of each chromosome, called haplotypes, is essential. The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the two chromosomes. Indeed, the knowledge of complete haplotypes is generally more informative than analyzing single SNPs and plays a fundamental role in many medical applications. RESULTS: To reconstruct the two haplotypes, we addressed the weighted Minimum Error Correction (wMEC) problem, which is a successful approach for haplotype assembly. This NP-hard problem consists in computing the two haplotypes that partition the sequencing reads into two disjoint sub-sets, with the least number of corrections to the SNP values. To this aim, we propose here GenHap, a novel computational method for haplotype assembly based on Genetic Algorithms, yielding optimal solutions by means of a global search process. In order to evaluate the effectiveness of our approach, we run GenHap on two synthetic (yet realistic) datasets, based on the Roche/454 and PacBio RS II sequencing technologies. We compared the performance of GenHap against HapCol, an efficient state-of-the-art algorithm for haplotype phasing. Our results show that GenHap always obtains high accuracy solutions (in terms of haplotype error rate), and is up to 4× faster than HapCol in the case of Roche/454 instances and up to 20× faster when compared on the PacBio RS II dataset. Finally, we assessed the performance of GenHap on two different real datasets. CONCLUSIONS: Future-generation sequencing technologies, producing longer reads with higher coverage, can highly benefit from GenHap, thanks to its capability of efficiently solving large instances of the haplotype assembly problem. Moreover, the optimization approach proposed in GenHap can be extended to the study of allele-specific genomic features, such as expression, methylation and chromatin conformation, by exploiting multi-objective optimization techniques. The source code and the full documentation are available at the following GitHub repository: https://github.com/andrea-tango/GenHap .
Andrea Tangherloni, Simone Spolaor, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Pietro Liò, Ivan Merelli, Daniela Besozzi
BMC Bioinform.1
2019 MedGA: A novel evolutionary method for image enhancement in medical imaging systems
Leonardo Rundo, Andrea Tangherloni, Marco S. Nobile, Carmelo Militello, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
Expert Syst. Appl.2
2019 USE-Net: Incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets
Leonardo Rundo, Yudai Nagano, Ryuichiro Hataya, Carmelo Militello, Andrea Tangherloni, Marco S. Nobile, Claudio Ferretti, Daniela Besozzi, Maria Carla Gilardi, Salvatore Vitabile, Giancarlo Mauri, Hideki Nakayama, Paolo Cazzaniga
Neurocomputing7
2018 Computational Intelligence for Parameter Estimation of Biochemical Systems
abstract
In the field of Systems Biology, simulating the dynamics of biochemical models represents one of the most effective methodologies to understand the functioning of cellular processes in normal or altered conditions. However, the lack of kinetic rates, necessary to perform accurate simulations, strongly limits the scope of these analyses. Parameter Estimation (PE), which consists in identifying a proper model parameterization, is a non-linear, non-convex and multi-modal optimization problem, typically tackled by means of Computational Intelligence techniques, such as Evolutionary Computation and Swarm Intelligence. In this work, we perform a thorough investigation of the most widespread methods for PE-namely, Artificial Bee Colony (ABC), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Differential Evolution (DE), Estimation of Distribution Algorithm (EDA), Genetic Algorithms (GAs), Particle Swarm Optimization (PSO), and Fuzzy Self-Tuning PSO (FST-PSO)-comparing their performances on a set of synthetic (yet realistic) biochemical models of increasing size and complexity. Our results show that a variant of the settings-free FST-PSO algorithm can consistently outperform all other methods; ABC and GAs represent the most performing alternatives, while methods based on multivariate normal distributions (e.g., CMA-ES, EDA) struggle to keep pace with the other approaches.
Marco S. Nobile, Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
CEC2
2018 GPU-Powered Multi-Swarm Parameter Estimation of Biological Systems: A Master-Slave Approach
abstract
In silico investigation of biological systems requires the knowledge of numerical parameters that cannot be easily measured in laboratory experiments, leading to the Parameter Estimation (PE) problem, in which the unknown parameters are automatically inferred by means of optimization algorithms exploiting the available experimental data. Here we present MS 2 PSO, an efficient parallel and distributed implementation of a PE method based on Particle Swarm Optimization (PSO) for the estimation of reaction constants in mathematical models of biological systems, considering as target for the estimation a set of discrete-time measurements of molecular species amounts. In particular, such PE method accounts for the availability of experimental data typically measured under different experimental conditions, by considering a multi-swarm PSO in which the best particles of the swarms can migrate. This strategy allows to infer a common set of reaction constants that simultaneously fits all target data used in the PE. To the aim of efficiently tackling the PE problem, MS 2 PSO embeds the execution of cupSODA, a deterministic simulator that relies on Graphics Processing Units to achieve a massive parallelization of the simulations required in the fitness evaluation of particles. In addition, a further level of parallelism is realized by exploiting the Master-Slave distributed programming paradigm. We apply MS 2 PSO for the PE of synthetic biochemical models with 10, 20 and 30 parameters to be estimated, and compare the performances obtained with different GPUs and different configurations (i.e., numbers of processes) of the Master-Slave.
Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Paolo Cazzaniga, Marco S. Nobile
PDP1
2017 Proactive Particles in Swarm Optimization: A settings-free algorithm for real-parameter single objective optimization problems
abstract
Particle Swarm Optimization (PSO) is an effective Swarm Intelligence technique for the optimization of non-linear and complex high-dimensional problems. Since PSO's performance is strongly dependent on the choice of its functioning settings, in this work we consider a self-tuning version of PSO, called Proactive Particles in Swarm Optimization (PPSO). PPSO leverages Fuzzy Logic to dynamically determine the best settings for the inertia weight, cognitive factor and social factor. The PPSO algorithm significantly differs from other versions of PSO relying on Fuzzy Logic, because specific settings are assigned to each particle according to its history, instead of being globally assigned to the whole swarm. In such a way, PPSO's particles gain a limited autonomous and proactive intelligence with respect to the reactive agents proposed by PSO. Our results show that PPSO achieves overall good optimization performances on the benchmark functions proposed in the CEC 2017 test suite, with the exception of those based on the Schwefel function, whose fitness landscape seems to mislead the fuzzy reasoning. Moreover, with many benchmark functions, PPSO is characterized by a higher speed of convergence than PSO in the case of high-dimensional problems.
Andrea Tangherloni, Leonardo Rundo, Marco S. Nobile
CEC1
2017 Reboot strategies in particle swarm optimization and their impact on parameter estimation of biochemical systems
abstract
Computational methods adopted in the field of Systems Biology require the complete knowledge of reaction kinetic constants to perform simulations of the dynamics and understand the emergent behavior of biochemical systems. However, kinetic parameters of biochemical reactions are often difficult or impossible to measure, thus they are generally inferred from experimental data, in a process known as Parameter Estimation (PE). We consider here a PE methodology that exploits Particle Swarm Optimization (PSO) to estimate an appropriate kinetic parameterization, by comparing experimental time-series target data with in silica dynamics, simulated by using the parameterization encoded by each particle. In this work we present three different reboot strategies for PSO, whose aim is to reinitialize particle positions to avoid particles to get trapped in local optima, and we compare the performance of PSO coupled with the reboot strategies with respect to standard PSO in the case of the PE of two biochemical systems. Since the PE requires a huge number of simulations at each iteration, in this work we exploit a GPU-powered deterministic simulator, cupSODA, which performs in a parallel fashion all simulations and fitness evaluations. Finally, we show that the performances of our implementation scale sublinearly with respect to the swarm size, even on outdated GPUs.
Simone Spolaor, Andrea Tangherloni, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga
CIBCB2
2017 Graphics processing units in bioinformatics, computational biology and systems biology
abstract
Several studies in Bioinformatics, Computational Biology and Systems Biology rely on the definition of physico-chemical or mathematical models of biological systems at different scales and levels of complexity, ranging from the interaction of atoms in single molecules up to genome-wide interaction networks. Traditional computational methods and software tools developed in these research fields share a common trait: they can be computationally demanding on Central Processing Units (CPUs), therefore limiting their applicability in many circumstances. To overcome this issue, general-purpose Graphics Processing Units (GPUs) are gaining an increasing attention by the scientific community, as they can considerably reduce the running time required by standard CPU-based software, and allow more intensive investigations of biological systems. In this review, we present a collection of GPU tools recently developed to perform computational analyses in life science disciplines, emphasizing the advantages and the drawbacks in the use of these parallel architectures. The complete list of GPU-powered tools here reviewed is available at http://bit.ly/gputools.
Marco S. Nobile, Paolo Cazzaniga, Andrea Tangherloni, Daniela Besozzi
Briefings Bioinform.3
2017 LASSIE: simulating large-scale models of biochemical systems on GPUs
abstract
BACKGROUND: Mathematical modeling and in silico analysis are widely acknowledged as complementary tools to biological laboratory methods, to achieve a thorough understanding of emergent behaviors of cellular processes in both physiological and perturbed conditions. Though, the simulation of large-scale models-consisting in hundreds or thousands of reactions and molecular species-can rapidly overtake the capabilities of Central Processing Units (CPUs). The purpose of this work is to exploit alternative high-performance computing solutions, such as Graphics Processing Units (GPUs), to allow the investigation of these models at reduced computational costs. RESULTS: LASSIE is a "black-box" GPU-accelerated deterministic simulator, specifically designed for large-scale models and not requiring any expertise in mathematical modeling, simulation algorithms or GPU programming. Given a reaction-based model of a cellular process, LASSIE automatically generates the corresponding system of Ordinary Differential Equations (ODEs), assuming mass-action kinetics. The numerical solution of the ODEs is obtained by automatically switching between the Runge-Kutta-Fehlberg method in the absence of stiffness, and the Backward Differentiation Formulae of first order in presence of stiffness. The computational performance of LASSIE are assessed using a set of randomly generated synthetic reaction-based models of increasing size, ranging from 64 to 8192 reactions and species, and compared to a CPU-implementation of the LSODA numerical integration algorithm. CONCLUSIONS: LASSIE adopts a novel fine-grained parallelization strategy to distribute on the GPU cores all the calculations required to solve the system of ODEs. By virtue of this implementation, LASSIE achieves up to 92× speed-up with respect to LSODA, therefore reducing the running time from approximately 1 month down to 8 h to simulate models consisting in, for instance, four thousands of reactions and species. Notably, thanks to its smaller memory footprint, LASSIE is able to perform fast simulations of even larger models, whereby the tested CPU-implementation of LSODA failed to reach termination. LASSIE is therefore expected to make an important breakthrough in Systems Biology applications, for the execution of faster and in-depth computational analyses of large-scale models of complex biological systems.
Andrea Tangherloni, Marco S. Nobile, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
BMC Bioinform.1
2017 Gillespie's Stochastic Simulation Algorithm on MIC coprocessors
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.1
2016 GPU-powered and settings-free parameter estimation of biochemical systems
abstract
To understand the emergent behavior of biochemical systems, computational analyses generally require the inference of unknown reaction kinetic constants, a problem known as parameter estimation (PE). In this work we propose a PE methodology that exploits Particle Swarm Optimization (PSO) to examine a set of candidate kinetic parameterizations, whose fitness is evaluated by comparing given target time-series of experimental data with in silico dynamics, simulated by using the parameterization encoded by each particle. In particular, we consider a Fuzzy Logic-based version of PSO - called Proactive Particles in Swarm Optimization (PPSO) - that automatically tunes the setting (inertia, cognitive and social factors) of each particle, independently from all other particles in the swarm. Since the optimization phase requires a large number of simulations for each particle at each iteration, we exploit a GPU-accelerated deterministic simulator, called cupSODA, that automatically generates the system of Ordinary Differential Equations associated with the biochemical system and performs its simulation for each candidate parameterization. We compare the performance of PPSO with respect to PSO for the PE problem by considering two biochemical systems as test cases. In addition, we evaluate the impact on PE of different strategies adopted, both in PPSO and PSO, for the selection of the initial positions of particles within the search space. We prove the effectiveness of our settings-free PE methodology by showing that PPSO outperforms PSO with respect to the computational time required to execute the optimization, achieving comparable results concerning the fitness of the best parameterization found.
Marco S. Nobile, Andrea Tangherloni, Daniela Besozzi, Paolo Cazzaniga
CEC2
2016 GPU-powered Bat Algorithm for the parameter estimation of biochemical kinetic values
abstract
The emergent behavior of biochemical systems can be investigated by means of mathematical modeling and computational analyses, which usually require the automatic inference of the unknown values of the model's parameters. This problem, known as Parameter Estimation (PE), is usually tackled with bio-inspired meta-heuristics for global optimization, most notably Particle Swarm Optimization (PSO). In this work we assess the performances of PSO and Bat Algorithm with differential operator and Lévy flights trajectories (DLBA). In particular, we compared these meta-heuristics for the PE using two biochemical models: the expression of genes in prokaryotes and the heat shock response in eukaryotes. In our tests, we also evaluated the impact on PE of different strategies for the initial positioning of individuals within the search space. Our results show that DLBA achieves comparable results with respect to PSO, but it converges to better results when a uniform initialization is employed. Since every iteration of DLBA requires three fitness evaluations for each bat, the whole methodology is built around a GPU-powered biochemical simulator (cupSODA) which is able to parallelize the process. We show that the acceleration achieved with cupSODA strongly reduces the running time, with an empirical 61× speedup that has been obtained comparing a Nvidia GeForce Titan GTX with respect to a CPU Intel Core i7-4790K. Moreover, we show that DLBA always outperforms PSO with respect to the computational time required to execute the optimization process.
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga
CIBCB1