VLDB 2026 Research / reviewers in the wild / expert
Paolo Cazzaniga
dblp:25/3374
· DBLP profile ↗
54ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0001-7780-0434ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 20 · 8 since 2021Systems, architecture and hardware · 4 · 1 since 2021Theory of computation · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyCAPS: A Settings-Free Optimization Heuristics Integrating Evolutionary Computation and Swarm Intelligence
Daniele M. Papetti, Marco S. Nobile, Matteo Grazioso, Paolo Cazzaniga, Leonardo Vanneschi, Daniela Besozzi |
EvoApplications | 4 |
| 2025 | Automated Phenotype-Based Clustering of Clinical Reports Using Large Language Models
Martina Saletta, Andrea Bombarda, Matteo Bellini, Lucrezia Goisis, Paolo Cazzaniga, Maria Iascone, Domenico Fabio Savo |
AIME (2) | 5 |
| 2025 | We Are Sending You Back... to the Optimum! Fuzzy Time Travel Particle Swarm Optimization
Daniele M. Papetti, Andrea Tangherloni, Vasco Coelho, Daniela Besozzi, Paolo Cazzaniga, Marco S. Nobile |
EvoApplications (2) | 5 |
| 2024 | A Modified EACOP Implementation for Real-Parameter Single Objective Optimization ProblemsabstractEvolutionary algorithms are effective techniques for optimizing non-linear and complex high-dimensional problems. However, most of them require a precise fine-tuning of their functioning settings to achieve satisfactory results. In this work, we propose a modified version of an evolutionary approach called the Evolutionary Algorithm for COmplex-process oPtimization (EACOP), designed to have a limited number of hyper-parameters. The base version of EACOP (bEACOP) combines different strategies, including the scatter search methodology, local searches, and a novel combination method based on path relinking to balance the exploration and exploitation phases. Our improved version (iEACOP) intensifies the exploration phase to escape from suboptimal search space areas where, on the contrary, bEACOP gets stuck. Our results show that iEACOP outperforms bEACOP on 27 out of 29 CEC 2017 test suite benchmark functions, exhibiting comparable performance against the three best algorithms of the CEC 2017 competition on single-objective bound-constrained real-parameter numerical optimization. The source code of bEACOP and iEACOP will be made publicly available on GitHub upon acceptance. Andrea Tangherloni, Vasco Coelho, Francesca Buffa, Paolo Cazzaniga |
CEC | 4 |
| 2024 | A Fast Feature Selection for Interpretable Modeling Based on Fuzzy Inference SystemsabstractLarge datasets are often beneficial for the generation of predictive models using machine learning approaches. However, it is often the case that not all variables in the dataset contain useful information. In fact, some variables might be useless, redundant, misleading, or even harmful to performance, both in terms of accuracy and computational effort. Because of that, Feature Selection (FS) is one of the most delicate and important steps in machine learning. This is even more relevant in the case of interpretable models based on Fuzzy Inference Systems (FIS). The reasons are two-fold: on the one hand, FIS are generally built on top of a data partitioning based on clustering, which can suffer from high dimensionality; on the other hand, the knowledge base of the FIS, to be concretely understandable, should not contain rules involving too many variables. FS can be performed using multiple approaches, most notably filter and wrapper methods. The latter are often based on evolutionary algorithms, where a population of candidate solutions (each representing a possible set of selected variables) evolves towards the optimal selection. Although wrapper methods can be effective, they are, in general, computationally expensive. In this work, we propose a completely different – and more computationally effective – algorithm based on Random Forest (RF) models. Specifically, we exploit RFs to rank variables according to their importance. Then, we use that information to perform a statistical analysis and determine the minimal set of features necessary to build an accurate FIS. We show the effectiveness of our approach by using two (semi)synthetic datasets built on real-world datasets, and we validate our approach by applying the FS method to a medical dataset. Andrea Tangherloni, Paolo Cazzaniga, Nicolò Stranieri, Francesca Buffa, Marco S. Nobile |
CIBCB | 2 |
| 2023 | The Domination Game: Dilating Bubbles to Fill Up Pareto FrontsabstractMulti-objective optimization algorithms might struggle in finding optimal dominating solutions, especially in real-case scenarios where problems are generally characterized by non-separability, non-differentiability, and multi-modality issues. An effective strategy that already showed to improve the outcome of optimization algorithms consists in manipulating the search space, in order to explore its most promising areas. In this work, starting from a Pareto front identified by an optimization strategy, we exploit Local Bubble Dilation Functions (LBDFs) to manipulate a locally bounded region of the search space containing non-dominated solutions. We tested our approach on the benchmark functions included in the DTLZ and WFG suites, showing that the Pareto front obtained after the application of LBDFs is most of the time characterized by an increased hyper-volume value. Our results confirm that LBDFs are an effective means to identify additional non-dominated solutions that can improve the quality of the Pareto front. Vasco Coelho, Daniele M. Papetti, Andrea Tangherloni, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile |
CEC | 4 |
| 2023 | Consensus Clustering Strategy for Cell Type Assignments of scRNA-seq DataabstractCell type annotation is a crucial step for analyzing single-cell RNA sequencing data. Among others, single-cell Automatic Labeling of cell POpulations (scALPO) is a computational pipeline developed to automatically assign the cell types to the identified clusters in scRNA-seq data. Different from most of the approaches, scALPO relies only on the information on marker genes from published literature. Specifically, after the definition of the dataset obtained from gene information retrieved from online databases, the Leiden clustering algorithm is executed to partition cells that are finally annotated. Since the Leiden algorithm might struggle to obtain a reliable outcome under certain circumstances, in this work, we include several clustering algorithms in scALPO, and we propose a pseudo-voting consensus approach that combines the outcome of a set of clustering algorithms. The results obtained on three different datasets show that the consensus approach can improve the cell type annotation without selecting a specific clustering algorithm that best suits the data under investigation. Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Francesca Buffa, Andrea Tangherloni |
CIBCB | 3 |
| 2023 | Unsupervised neural networks as a support tool for pathology diagnosis in MALDI-MSI experiments: A case study on thyroid biopsiesabstractArtificial intelligence is getting a foothold in medicine for disease screening and diagnosis. While typical machine learning methods require large labeled datasets for training and validation, their application is limited in clinical fields since ground truth information can hardly be obtained on a sizeable cohort of patients. Unsupervised neural networks – such as Self-Organizing Maps (SOMs) – represent an alternative approach to identifying hidden patterns in biomedical data. Here we investigate the feasibility of SOMs for the identification of malignant and non-malignant regions in liquid biopsies of thyroid nodules, on a patient-specific basis. MALDI-ToF (Matrix Assisted Laser Desorption Ionization - Time of Flight) mass spectrometry-imaging (MSI) was used to measure the spectral profile of bioptic samples. SOMs were then applied for the analysis of MALDI-MSI data of individual patients’ samples, also testing various pre-processing and agglomerative clustering methods to investigate their impact on SOMs’ discrimination efficacy. The final clustering was compared against the sample’s probability to be malignant, hyperplastic or related to Hashimoto thyroiditis as quantified by multinomial regression with LASSO. Our results show that SOMs are effective in separating the areas of a sample containing benign cells from those containing malignant cells. Moreover, they allow to overlap the different areas of cytological glass slides with the corresponding proteomic profile image, and inspect the specific weight of every cellular component in bioptic samples. We envision that this approach could represent an effective means to assist pathologists in diagnostic tasks, avoiding the need to manually annotate cytological images and the effort in creating labeled datasets. Marco S. Nobile, Giulia Capitoli, Virgil Sowirono, Francesca Clerici, Isabella Piga, Kirsten van Abeelen, Fulvio Magni, Fabio Pagni, Stefania Galimberti, Paolo Cazzaniga, Daniela Besozzi |
Expert Syst. Appl. | 10 |
| 2023 | MAGNETO: Cell type marker panel generator from single-cell transcriptomic dataabstractSingle-cell RNA sequencing experiments produce data useful to identify different cell types, including uncharacterized and rare ones. This enables us to study the specific functional roles of these cells in different microenvironments and contexts. After identifying a (novel) cell type of interest, it is essential to build succinct marker panels, composed of a few genes referring to cell surface proteins and clusters of differentiation molecules, able to discriminate the desired cells from the other cell populations. In this work, we propose a fully-automatic framework called MAGNETO, which can help construct optimal marker panels starting from a single-cell gene expression matrix and a cell type identity for each cell. MAGNETO builds effective marker panels solving a tailored bi-objective optimization problem, where the first objective regards the identification of the genes able to isolate a specific cell type, while the second conflicting objective concerns the minimization of the total number of genes included in the panel. Our results on three public datasets show that MAGNETO can identify marker panels that identify the cell populations of interest better than state-of-the-art approaches. Finally, by fine-tuning MAGNETO, our results demonstrate that it is possible to obtain marker panels with different specificity levels. Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Francesca Buffa, Paolo Cazzaniga |
J. Biomed. Informatics | 5 |
| 2022 | Local Bubble Dilation Functions: Hypersphere-bounded Landscape Deformations Simplify Global OptimizationabstractSolving optimization problems is one of the most complex and widespread task in Computer Science. In many scenarios, finding the global optimum of a function is hampered by several features that characterize the fitness landscapes, such as noisiness, multi-modality, non-convexity, non-separability, and non-differentiability. In order to facilitate the optimization process, a variety of methods have been proposed to manipulate either the search space or the fitness landscape. Among these, Dilation Functions (DFs) were introduced to expand regions of the search space that are characterized by promising fitness values. In this work, we extend the family of DFs by introducing Local Bubble Dilation Functions (LBDFs), a novel approach that generates local distortions bounded by hyper-spheres. By performing an appropriate mapping of the search space, LBDFs can improve the optimization performance, since they expand and reveal the promising regions around the global optimum, while leaving the rest of the fitness landscape untouched. The additional advantage of LBDFs, with respect to DFs, is that different dilations can be applied to each dimension of the search space, which is useful in the case of asymmetric landscapes. In order to show the benefits of local dilations, we executed several tests on the Michalewicz benchmark function, with different settings for the LBDFs. Our results show that a properly designed LBDF can lead to statistically significant better results than using vanilla optimization. Finally, we investigated the use of LBDFs to facilitate the solution of the parameter estimation problem in Systems Biology by analyzing the landscape related to a stochastic model of enzyme kinetics. Daniele M. Papetti, Vasco Coelho, Dan Ashlock, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Marco S. Nobile |
CIBCB | 4 |
| 2022 | A Deep Learning Pipeline for the Automatic cell type Assignment of scRNA-seq DataabstractThe increasing number of single-cell transcriptomics and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism, as well as the onset of pathologies. In this context, cell type annotation represents a crucial step for the analysis of single-cell RNA sequencing data, which is usually performed by means of time-consuming and possibly biased manual processes, carried out by expert biologists. Recently, alternative computational tools have been proposed to realize an automatic cell identification either based on supervised or unsupervised Machine Learning approaches. These methods typically exploit gene expression data of curated marker gene databases to associate gene expression profiles of single cells with a cell type. In this paper, we propose a novel fully-automatic computational pipeline, named single-cell Automatic Labeling of cell POpulations (scALPO), which leverages a Long Short-Term Memory Neural Network to assign the cell types. Specifically, scALPO can label the provided clusters by simply relying on marker genes rather than gene expressions. Our results, obtained by considering two different datasets, show that scALPO outperforms the most promising state-of-the-art approaches (i.e., SCSA and scType), achieving a cell type annotation more similar to the manually-created ground truth. Simone G. Riva, Brynelle Myers, Paolo Cazzaniga, Andrea Tangherloni |
CIBCB | 3 |
| 2022 | Multi-objective Optimization for Marker Panel Identification in Single-cell DataabstractThe computational analyses of single-cell data, aimed at elucidating and characterizing the functional roles of known and putative novel cell types, are enabling a thorough understanding of the processes driving cell development and pathology progression. The isolation of specific cell types is a crucial step to perform detailed analyses but requires the identification of succinct marker panels, which include genes that refer to cell surface proteins and clusters of differentiation molecules. This still represents a challenging NP-hard computational problem, which can be tackled through global optimization techniques. In this work, we formulate the marker panel identification problem as a bi-objective optimization problem, where the first objective regards the capability of the marker panels to accurately discriminate different cell types, while the second objective is related to the number of genes to include in the panel. In particular, we compared the performance of two multi-objective optimization algorithms, as well as of Genetic Algorithms (GAs) when considering only the first objective, employing two different representations for the candidate solutions. Our results show that the multi-objective optimization algorithms are better than GAs, considering both the quality and the consistency of the obtained marker panels; moreover, the collected results point out that different representations of the candidate solutions have a relevant impact on the performance of the optimization algorithms. Andrea Tangherloni, Simone G. Riva, Brynelle Myers, Paolo Cazzaniga |
CIBCB | 4 |
| 2021 | Integration of Multiple scRNA-Seq Datasets on the Autoencoder Latent SpaceabstractThe application of single-cell transcriptomic sequencing technologies, such as single-cell RNA sequencing (scRNA-Seq), have witnessed in recent years a dramatic increase, allowing for the elucidation of the molecular processes driving both normal cell development and the onset of pathologies. In particular, scRNA-Seq can be exploited to investigate cell heterogeneity at single-cell resolution, and to identify the variety of known and putatively novel cell populations, which can potentially have different functional roles in different contexts. However, the heterogeneity among cells of the same cell-type can make the integration of multiple scRNA-Seq datasets a challenging task. In this context, technical non-negligible batch effects in the datasets—which may arise from the sequencing technology employed and from the size of the experiment—must be considered to realize a correct data integration. In this work, we present a novel strategy based on Autoencoders (AEs) for the integration of multiple scRNA-Seq datasets, whose performance is compared with different integration strategies that do not exploit a batch effect removal step, which might introduce artifacts in the datasets. Our results, obtained by considering 3 different datasets, suggest that AEs represent a suitable strategy for the integration of scRNA-Seq datasets, achieving better performance than other approaches, i.e., Scanorama, Ingest, and Seurat, in most of the cases. Simone G. Riva, Paolo Cazzaniga, Andrea Tangherloni |
BIBM | 2 |
| 2021 | If You Can't Beat It, Squash It: Simplify Global Optimization by Evolving Dilation FunctionsabstractOptimization problems represent a class of pervasive and complex tasks in Computer Science, aimed at identifying the global optimum of a given objective function. Optimization problems are typically noisy, multi-modal, non-convex, non-separable, and often non-differentiable. Because of these features, they mandate the use of sophisticated population-based meta-heuristics to effectively explore the search space. Additionally, computational techniques based on the manipulation of the optimization landscape, such as Dilation Functions (DFs), can be effectively exploited to either "compress" or "dilate" some target regions of the search space, in order to improve the exploration and exploitation capabilities of any meta-heuristic. The main limitation of DFs is that they must be tailored on the specific optimization problem under investigation. In this work, we propose a solution to this issue, based on the idea of evolving the DFs. Specifically, we introduce a two-layered evolutionary framework, which combines Evolutionary Computation and Swarm Intelligence to solve the meta-problem of optimizing both the structure and the parameters of DFs. We evolved optimal DFs on a variety of benchmark problems, showing that this approach yields extremely simpler versions of the original optimization problems. Daniele M. Papetti, Dan Ashlock, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile |
CEC | 3 |
| 2021 | The Impact of Representation on the Optimization of Marker Panels for Single-cell RNA DataabstractThe increasing number of single-cell transcriptomic and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism as well as the onset of pathologies. These studies continuously refine the functional roles of known cell populations, and provide their characterization as soon as putatively novel cell populations are detected. In order to isolate the cell populations for further tailored analysis, succinct marker panels—composed of a few cell surface proteins and clusters of differentiation molecules—must be identified. The identification of these marker panels is a challenging computational problem due to its intrinsic combinatorial nature, which makes it an NP-hard problem. Genetic Algorithms (GAs) have been successfully used in Bioinformatics and other biomedical applications to tackle combinatorial problems. We present here a GA-based approach to solve the problem of the identification of succinct marker panels. Since the performance of a GA is strictly related to the representation of the candidate solutions, we propose and compare three alternative representations, able to implicitly introduce different constraints on the search space. For each representation, we perform a fine-tuning of the parameter settings to calibrate the GA, and we show that different representations yield different performance, where the most relaxed representations— in which the GA can also evolve the number of genes in the panel—turn out to be the more effective, especially in the case of 0-knowledge problems. Our results also show that the marker panels identified by GAs can outperform manually curated solutions. Andrea Tangherloni, Simone G. Riva, Simone Spolaor, Daniela Besozzi, Marco S. Nobile, Paolo Cazzaniga |
CEC | 6 |
| 2021 | A comparison of multi-objective optimization algorithms to identify drug target combinationsabstractCombination therapies represent one of the most effective strategy in inducing cancer cell death and reducing the risk to develop drug resistance. The identification of putative novel drug combinations, which typically requires the execution of expensive and time consuming lab experiments, can be supported by the synergistic use of mathematical models and multi-objective optimization algorithms. The computational approach allows to automatically search for potential therapeutic combinations and to test their effectiveness in silico, thus reducing the costs of time and money, and driving the experiments toward the most promising therapies. In this work, we couple dynamic fuzzy modeling of cancer cells with different multi-objective optimization algorithm, and we compare their performance in identifying drug target combinations. Specifically, we perform batches of optimizations with 3 and 4 objective functions defined to achieve a desired behavior of the system (e.g., maximize apop-tosis while minimizing necrosis and survival), and we compare the quality of the solutions included in the Pareto fronts. Our results show that both the choice of the multi-objective algorithm and the formulation of the optimization problem have an impact on the identified solutions, highlighting the strengths as well as the limitations of this approach. Simone Spolaor, Daniele M. Papetti, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile |
CIBCB | 3 |
| 2021 | Selected papers from the 15th and 16th international conference on Computational Intelligence Methods for Bioinformatics and BiostatisticsabstractThis supplement contains seven revised and extended papers selected from CIBB 2018 and CIBB 2019, the 15th and 16th editions of the international conference on Computational Intelligence Methods for Bioinformatics and Biostatistics. CIBB is a venue that embraces researchers with different backgrounds, ranging from mathematics to computer science, from materials science to medicine, and from engineering to biology, all interested in the investigation and application of computational intelligence methods to open problems in bioinformatics, biostatistics, systems biology, synthetic biology, and medical informatics. Paolo Cazzaniga, Maria Raposo, Daniela Besozzi, Ivan Merelli, Antonino Staiano, Angelo Ciaramella, Riccardo Rizzo, Luca Manzoni |
BMC Bioinform. | 1 |
| 2021 | Investigating the performance of multi-objective optimization when learning Bayesian Networks
Marco S. Nobile, Paolo Cazzaniga, Daniele Ramazzotti |
Neurocomputing | 2 |
| 2021 | FiCoS: A fine-grained and coarse-grained GPU-powered deterministic simulator for biochemical networksabstractMathematical models of biochemical networks can largely facilitate the comprehension of the mechanisms at the basis of cellular processes, as well as the formulation of hypotheses that can be tested by means of targeted laboratory experiments. However, two issues might hamper the achievement of fruitful outcomes. On the one hand, detailed mechanistic models can involve hundreds or thousands of molecular species and their intermediate complexes, as well as hundreds or thousands of chemical reactions, a situation generally occurring in rule-based modeling. On the other hand, the computational analysis of a model typically requires the execution of a large number of simulations for its calibration, or to test the effect of perturbations. As a consequence, the computational capabilities of modern Central Processing Units can be easily overtaken, possibly making the modeling of biochemical networks a worthless or ineffective effort. To the aim of overcoming the limitations of the current state-of-the-art simulation approaches, we present in this paper FiCoS, a novel "black-box" deterministic simulator that effectively realizes both a fine-grained and a coarse-grained parallelization on Graphics Processing Units. In particular, FiCoS exploits two different integration methods, namely, the Dormand-Prince and the Radau IIA, to efficiently solve both non-stiff and stiff systems of coupled Ordinary Differential Equations. We tested the performance of FiCoS against different deterministic simulators, by considering models of increasing size and by running analyses with increasing computational demands. FiCoS was able to dramatically speedup the computations up to 855×, showing to be a promising solution for the simulation and analysis of large-scale models of complex biological processes. Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Giulia Capitoli, Simone Spolaor, Leonardo Rundo, Giancarlo Mauri, Daniela Besozzi |
PLoS Comput. Biol. | 3 |
| 2021 | A CUDA-powered method for the feature extraction and unsupervised analysis of medical imagesabstractAbstract Image texture extraction and analysis are fundamental steps in computer vision. In particular, considering the biomedical field, quantitative imaging methods are increasingly gaining importance because they convey scientifically and clinically relevant information for prediction, prognosis, and treatment response assessment. In this context, radiomic approaches are fostering large-scale studies that can have a significant impact in the clinical practice. In this work, we present a novel method, called CHASM (Cuda, HAralick & SoM), which is accelerated on the graphics processing unit (GPU) for quantitative imaging analyses based on Haralick features and on the self-organizing map (SOM). The Haralick features extraction step relies upon the gray-level co-occurrence matrix, which is computationally burdensome on medical images characterized by a high bit depth. The downstream analyses exploit the SOM with the goal of identifying the underlying clusters of pixels in an unsupervised manner. CHASM is conceived to leverage the parallel computation capabilities of modern GPUs. Analyzing ovarian cancer computed tomography images, CHASM achieved up to $$\sim 19.5\times $$ ∼ 19.5 × and $$\sim 37\times $$ ∼ 37 × speed-up factors for the Haralick feature extraction and for the SOM execution, respectively, compared to the corresponding C++ coded sequential versions. Such computational results point out the potential of GPUs in the clinical research. Leonardo Rundo, Andrea Tangherloni, Paolo Cazzaniga, Matteo Mistri, Simone Galimberti, Ramona Woitek, Evis Sala, Giancarlo Mauri, Marco S. Nobile |
J. Supercomput. | 3 |
| 2020 | Which random is the best random? A study on sampling methods in Fourier surrogate modelingabstractGlobal optimization problems can be effectively solved by means of Computational Intelligence methods. However, there are several areas in which the effectiveness of these algorithms can be hampered by the computational costs of the fitness evaluations, or by specific features of the fitness landscape that can be characterized by noise and by the presence of several (even infinite) local optima. These issues bring about the necessity of defining specific techniques to replace the original problem with a surrogate representation. Fourier surrogate modeling represents a novel and effective approach to generate smoother, and possibly easier to explore, fitness landscapes, and to reduce the computational effort. Fourier surrogates require an initial sampling of the search space that must be performed to calculate the Fourier transforms. In this paper we investigate the impact on the quality of the surrogate models of the hyper-parameters of the methodology, and of several methods that can be employed for the initial sampling of the fitness landscape (i.e., pseudorandom numbers, low discrepancy sequences, a logistic map in chaotic regime, true random positions generated by a quantum computer, and point packing). Our results show that semistructured approaches like quasi-random sequences and point packing can outperform the other sampling methods. Marco S. Nobile, Simone Spolaor, Paolo Cazzaniga, Daniele M. Papetti, Daniela Besozzi, Dan Ashlock, Luca Manzoni |
CEC | 3 |
| 2020 | Fourier Surrogate Models of Dilated Fitness Landscapes in Systems Biology : or how we learned to torture optimization problems until they confessabstractOne of the most complex problems in Systems Biology is Parameter estimation (PE), which consists in inferring the kinetic parameters of biochemical systems. The identification of an accurate parameterization, able to reproduce any observed experimental behavior, is fundamental for the definition of predictive models. PE is a non-convex, multi-modal, and non-separable problem that is usually tackled by using Computational Intelligence methods. When the biochemical species appear in the system in a very low amount, the intrinsic noise due to the randomness of molecular collisions cannot be neglected. In this case, stochastic simulation algorithms should be employed to properly reproduce the system dynamics. Stochastic fluctuations make the PE problem even more complicated, as they can lead to radically different values of the fitness function for the same candidate parameterization. In addition, the kinetic parameters generally follow a log-uniform distribution, so that global optima tend to be localized in the lowest orders of magnitude of the search space. To simultaneously tackle all the aforementioned issues, in this work we investigate a novel approach based on the combination of dilation functions with Fourier surrogate modeling and filtering on the fitness landscape. The results show that our approach is able to strongly simplify the PE problem for low-dimensional optimization instances. Marco S. Nobile, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Luca Manzoni |
CIBCB | 2 |
| 2020 | On the automatic calibration of fully analogical spiking neuromorphic chipsabstractNowadays, understanding the topology of biological neural networks and sampling their activity is possible thanks to various laboratory protocols that provide a large amount of experimental data, thus paving the way to accurate modeling and simulation. Neuromorphic systems were developed to simulate the dynamics of biological neural networks by means of electronic circuits, offering an efficient alternative to classic simulations based on systems of differential equations, from both the points of view of the energy consumed and the overall computational effort. Spikey is a configurable neuromorphic chip based on the Leaky Integrate-And-Fire model, which gives the user the possibility to model an arbitrary neural topology and simulate the temporal evolution of membrane potentials. To accurately reproduce the behavior of a specific biological network, a detailed parameterization of all neurons in the neuromorphic chip is necessary. Determining such parameters is a hard, error-prone, and generally time consuming task. In this work, we propose a novel methodology for the automatic calibration of neuromorphic chips that exploits a given neural activity as target. Our results show that, in the case of small networks with a low complexity, the method can estimate a vector of parameters capable of reproducing the target activity. Conversely, in the case of more complex networks, the simulations with Spikey can be highly affected by noise, which causes small variations in the simulations outcome even when identical networks are simulated, hindering the convergence to optimal parameterizations. Daniele M. Papetti, Simone Spolaor, Daniela Besozzi, Paolo Cazzaniga, Marco Antoniotti, Marco S. Nobile |
IJCNN | 4 |
| 2020 | Fuzzy modeling and global optimization to predict novel therapeutic targets in cancer cellsabstractMOTIVATION: The elucidation of dysfunctional cellular processes that can induce the onset of a disease is a challenging issue from both the experimental and computational perspectives. Here we introduce a novel computational method based on the coupling between fuzzy logic modeling and a global optimization algorithm, whose aims are to (1) predict the emergent dynamical behaviors of highly heterogeneous systems in unperturbed and perturbed conditions, regardless of the availability of quantitative parameters, and (2) determine a minimal set of system components whose perturbation can lead to a desired system response, therefore facilitating the design of a more appropriate experimental strategy. RESULTS: We applied this method to investigate what drives K-ras-induced cancer cells, displaying the typical Warburg effect, to death or survival upon progressive glucose depletion. The optimization analysis allowed to identify new combinations of stimuli that maximize pro-apoptotic processes. Namely, our results provide different evidences of an important protective role for protein kinase A in cancer cells under several cellular stress conditions mimicking tumor behavior. The predictive power of this method could facilitate the assessment of the response of other complex heterogeneous systems to drugs or mutations in fields as medicine and pharmacology, therefore paving the way for the development of novel therapeutic treatments. AVAILABILITY AND IMPLEMENTATION: The source code of FUMOSO is available under the GPL 2.0 license on GitHub at the following URL: https://github.com/aresio/FUMOSO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marco S. Nobile, Giuseppina Votta, Roberta Palorini, Simone Spolaor, Humberto De Vitto, Paolo Cazzaniga, Francesca Ricciardiello, Giancarlo Mauri, Lilia Alberghina, Ferdinando Chiaradonna, Daniela Besozzi |
Bioinform. | 6 |
| 2020 | Computational Intelligence for Life SciencesabstractComputational Intelligence (CI) is a computer science discipline encompassing the theory, design, development and application of biologically and linguistically derived computational paradigms. Traditionally, the main elements of CI are Evolutionary Computation, Swarm Intelligence, Fuzzy Logic, and Neural Networks. CI aims at proposing new algorithms able to solve complex computational problems by taking inspiration from natural phenomena. In an intriguing turn of events, these nature-inspired methods have been widely adopted to investigate a plethora of problems related to nature itself. In this paper we present a variety of CI methods applied to three problems in life sciences, highlighting their effectiveness: we describe how protein folding can be faced by exploiting Genetic Programming, the inference of haplotypes can be tackled using Genetic Algorithms, and the estimation of biochemical kinetic parameters can be performed by means of Swarm Intelligence. We show that CI methods can generate very high quality solutions, providing a sound methodology to solve complex optimization problems in life sciences. Daniela Besozzi, Luca Manzoni, Marco S. Nobile, Simone Spolaor, Mauro Castelli, Leonardo Vanneschi, Paolo Cazzaniga, Stefano Ruberto, Leonardo Rundo, Andrea Tangherloni |
Fundam. Informaticae | 7 |
| 2020 | Coupling Mechanistic Approaches and Fuzzy Logic to Model and Simulate Complex SystemsabstractSeveral mathematical formalisms can be exploited to model complex systems, in order to capture different features of their dynamic behavior and leverage any available quantitative or qualitative data. Correspondingly, either quantitative models or qualitative models can be defined; bridging the gap between these two worlds would allow us to simultaneously exploit the peculiar advantages provided by each modeling approach. However, to date, the attempts in this direction have been limited to specific fields of research. In this paper, we propose a novel, general-purpose computational framework, named Fuzzy-mechanistic modeling of compleX systems (FuzzX), for the analysis of hybrid models consisting of a quantitative (or mechanistic) module and a qualitative module that can reciprocally control each other's dynamic behavior through a common interface. FuzzX takes advantage of precise quantitative information about the system through the definition and simulation of the mechanistic module. At the same time, it describes the behavior of components and their interactions that are not known in full details, by exploiting fuzzy logic for the definition of the qualitative module. We applied FuzzX for the analysis of a hybrid model of a complex biochemical system, characterized by the presence of positive and negative feedback regulations. We show that FuzzX is able to correctly reproduce known emergent behaviors of this system in normal and perturbed conditions. We envision that FuzzX could be employed to analyze any kind of complex system when quantitative information is limited, as well as to extend existing mechanistic models with fuzzy modules to describe those components and interactions of the system that are not fully characterized. Simone Spolaor, Marco S. Nobile, Giancarlo Mauri, Paolo Cazzaniga, Daniela Besozzi |
IEEE Trans. Fuzzy Syst. | 4 |
| 2020 | cuProCell: GPU-Accelerated Analysis of Cell Proliferation With Flow Cytometry DataabstractThe investigation of cell proliferation can provide useful insights for the comprehension of cancer progression, resistance to chemotherapy and relapse. To this aim, computational methods and experimental measurements based on in vivo label-retaining assays can be coupled to explore the dynamic behavior of tumoral cells. ProCell is a software that exploits flow cytometry data to model and simulate the kinetics of fluorescence loss that is due to stochastic events of cell division. Since the rate of cell division is not known, ProCell embeds a calibration process that might require thousands of stochastic simulations to properly infer the parameterization of cell proliferation models. To mitigate the high computational costs, in this paper we introduce a parallel implementation of ProCell's simulation algorithm, named cuProCell, which leverages Graphics Processing Units (GPUs). Dynamic Parallelism was used to efficiently manage the cell duplication events, in a radically different way with respect to common computing architectures. We present the advantages of cuProCell for the analysis of different models of cell proliferation in Acute Myeloid Leukemia (AML), using data collected from the spleen of human xenografts in mice. We show that, by exploiting GPUs, our method is able to not only automatically infer the models' parameterization, but it is also 237× faster than the sequential implementation. This study highlights the presence of a relevant percentage of quiescent and potentially chemoresistant cells in AML in vivo, and suggests that maintaining a dynamic equilibrium among the different proliferating cell populations might play an important role in disease progression. Marco S. Nobile, Eric Nisoli, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Dilation Functions in Global OptimizationabstractComplex tasks in Computer Science can be reformulated as optimization problems, in which the global optimum of a given function must be identified. Such problems are typically noisy, multi-modal, non-convex and non-separable, and they require the application of population-based global search metaheuristics to effectively explore the search space. In this work, we address the issue of manipulating the search space of these complex optimization problems to the aim of improving the exploration and exploitation capabilities of metaheuristics. In particular, we show that the implicit assumption in global optimization problems, i.e., that candidate solutions are represented by vectors of values whose meaning has a straightforward interpretation, is not always adequate and that the semantics of parameters can be modified by re-mapping their values in the search space by means of user-defined Dilation Functions. Dilation Functions are general purpose transformations that can be applied to any metaheuristics and optimization problem to "compress" or "dilate" some regions of the search space, allowing to improve the quality of the initial population and the exploitation of promising areas, especially in the case of Swarm Intelligence algorithms. The advantages given by the application of Dilation Functions have been observed by running experiments with Fuzzy Self-Tuning Particle Swarm Optimization and Covariance Matrix Adaptation Evolution Strategies, for the optimization of the Ackley benchmark function and for the parameter estimation of a "synthetic" model of a biochemical system. Marco S. Nobile, Paolo Cazzaniga, Dan Ashlock |
CEC | 2 |
| 2019 | ProCell: Investigating cell proliferation with Swarm IntelligenceabstractComputational methods represent an effective mean for the analysis of complex biological processes, such as cell proliferation, especially when combined to well established experimental protocols. In particular, mathematical modeling coupled with computational intelligence algorithms can be successfully exploited to investigate different aspects of cell population dynamics in the context of tumor growth. To this aim, we defined ProCell, a modeling and simulation framework specifically designed for the investigation of cell proliferation, which makes use of Fuzzy Self-Tuning Particle Swarm Optimization to estimate the unknown parameters of cell population models. ProCell is here applied to the analysis of cell proliferation in acute myeloid leukemia, a hematological malignancy characterized by an inherent intra-tumoral heterogeneity that plays an important role in disease recurrence and resistance to chemotherapy. ProCell allowed to provide new insights on the intricate organization of cells with highly heterogeneous proliferative potential, and to highlight the important role of different cell types in the progression and evolution of the disease. ProCell is available under the GPL 2.0 license on GitHub at https://github.com/aresio/ProCell. Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi |
CIBCB | 4 |
| 2019 | Modeling cell proliferation in human acute myeloid leukemia xenograftsabstractMOTIVATION: Acute myeloid leukemia (AML) is one of the most common hematological malignancies, characterized by high relapse and mortality rates. The inherent intra-tumor heterogeneity in AML is thought to play an important role in disease recurrence and resistance to chemotherapy. Although experimental protocols for cell proliferation studies are well established and widespread, they are not easily applicable to in vivo contexts, and the analysis of related time-series data is often complex to achieve. To overcome these limitations, model-driven approaches can be exploited to investigate different aspects of cell population dynamics. RESULTS: In this work, we present ProCell, a novel modeling and simulation framework to investigate cell proliferation dynamics that, differently from other approaches, takes into account the inherent stochasticity of cell division events. We apply ProCell to compare different models of cell proliferation in AML, notably leveraging experimental data derived from human xenografts in mice. ProCell is coupled with Fuzzy Self-Tuning Particle Swarm Optimization, a swarm-intelligence settings-free algorithm used to automatically infer the models parameterizations. Our results provide new insights on the intricate organization of AML cells with highly heterogeneous proliferative potential, highlighting the important role played by quiescent cells and proliferating cells characterized by different rates of division in the progression and evolution of the disease, thus hinting at the necessity to further characterize tumor cell subpopulations. AVAILABILITY AND IMPLEMENTATION: The source code of ProCell and the experimental data used in this work are available under the GPL 2.0 license on GITHUB at the following URL: https://github.com/aresio/ProCell. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Daniela Bossi, Paolo Cazzaniga, Luisa Lanfrancone, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi |
Bioinform. | 5 |
| 2019 | GenHap: a novel computational method based on genetic algorithms for haplotype assemblyabstractBACKGROUND: In order to fully characterize the genome of an individual, the reconstruction of the two distinct copies of each chromosome, called haplotypes, is essential. The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the two chromosomes. Indeed, the knowledge of complete haplotypes is generally more informative than analyzing single SNPs and plays a fundamental role in many medical applications. RESULTS: To reconstruct the two haplotypes, we addressed the weighted Minimum Error Correction (wMEC) problem, which is a successful approach for haplotype assembly. This NP-hard problem consists in computing the two haplotypes that partition the sequencing reads into two disjoint sub-sets, with the least number of corrections to the SNP values. To this aim, we propose here GenHap, a novel computational method for haplotype assembly based on Genetic Algorithms, yielding optimal solutions by means of a global search process. In order to evaluate the effectiveness of our approach, we run GenHap on two synthetic (yet realistic) datasets, based on the Roche/454 and PacBio RS II sequencing technologies. We compared the performance of GenHap against HapCol, an efficient state-of-the-art algorithm for haplotype phasing. Our results show that GenHap always obtains high accuracy solutions (in terms of haplotype error rate), and is up to 4× faster than HapCol in the case of Roche/454 instances and up to 20× faster when compared on the PacBio RS II dataset. Finally, we assessed the performance of GenHap on two different real datasets. CONCLUSIONS: Future-generation sequencing technologies, producing longer reads with higher coverage, can highly benefit from GenHap, thanks to its capability of efficiently solving large instances of the haplotype assembly problem. Moreover, the optimization approach proposed in GenHap can be extended to the study of allele-specific genomic features, such as expression, methylation and chromatin conformation, by exploiting multi-objective optimization techniques. The source code and the full documentation are available at the following GitHub repository: https://github.com/andrea-tango/GenHap . Andrea Tangherloni, Simone Spolaor, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Pietro Liò, Ivan Merelli, Daniela Besozzi |
BMC Bioinform. | 5 |
| 2019 | MedGA: A novel evolutionary method for image enhancement in medical imaging systems
Leonardo Rundo, Andrea Tangherloni, Marco S. Nobile, Carmelo Militello, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga |
Expert Syst. Appl. | 7 |
| 2019 | USE-Net: Incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets
Leonardo Rundo, Yudai Nagano, Ryuichiro Hataya, Carmelo Militello, Andrea Tangherloni, Marco S. Nobile, Claudio Ferretti, Daniela Besozzi, Maria Carla Gilardi, Salvatore Vitabile, Giancarlo Mauri, Hideki Nakayama, Paolo Cazzaniga |
Neurocomputing | 15 |
| 2019 | ginSODA: massive parallel integration of stiff ODE systems on GPUs
Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri |
J. Supercomput. | 2 |
| 2018 | Computational Intelligence for Parameter Estimation of Biochemical SystemsabstractIn the field of Systems Biology, simulating the dynamics of biochemical models represents one of the most effective methodologies to understand the functioning of cellular processes in normal or altered conditions. However, the lack of kinetic rates, necessary to perform accurate simulations, strongly limits the scope of these analyses. Parameter Estimation (PE), which consists in identifying a proper model parameterization, is a non-linear, non-convex and multi-modal optimization problem, typically tackled by means of Computational Intelligence techniques, such as Evolutionary Computation and Swarm Intelligence. In this work, we perform a thorough investigation of the most widespread methods for PE-namely, Artificial Bee Colony (ABC), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Differential Evolution (DE), Estimation of Distribution Algorithm (EDA), Genetic Algorithms (GAs), Particle Swarm Optimization (PSO), and Fuzzy Self-Tuning PSO (FST-PSO)-comparing their performances on a set of synthetic (yet realistic) biochemical models of increasing size and complexity. Our results show that a variant of the settings-free FST-PSO algorithm can consistently outperform all other methods; ABC and GAs represent the most performing alternatives, while methods based on multivariate normal distributions (e.g., CMA-ES, EDA) struggle to keep pace with the other approaches. Marco S. Nobile, Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga |
CEC | 7 |
| 2018 | GPU-Powered Multi-Swarm Parameter Estimation of Biological Systems: A Master-Slave ApproachabstractIn silico investigation of biological systems requires the knowledge of numerical parameters that cannot be easily measured in laboratory experiments, leading to the Parameter Estimation (PE) problem, in which the unknown parameters are automatically inferred by means of optimization algorithms exploiting the available experimental data. Here we present MS 2 PSO, an efficient parallel and distributed implementation of a PE method based on Particle Swarm Optimization (PSO) for the estimation of reaction constants in mathematical models of biological systems, considering as target for the estimation a set of discrete-time measurements of molecular species amounts. In particular, such PE method accounts for the availability of experimental data typically measured under different experimental conditions, by considering a multi-swarm PSO in which the best particles of the swarms can migrate. This strategy allows to infer a common set of reaction constants that simultaneously fits all target data used in the PE. To the aim of efficiently tackling the PE problem, MS 2 PSO embeds the execution of cupSODA, a deterministic simulator that relies on Graphics Processing Units to achieve a massive parallelization of the simulations required in the fitness evaluation of particles. In addition, a further level of parallelism is realized by exploiting the Master-Slave distributed programming paradigm. We apply MS 2 PSO for the PE of synthetic biochemical models with 10, 20 and 30 parameters to be estimated, and compare the performances obtained with different GPUs and different configurations (i.e., numbers of processes) of the Master-Slave. Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Paolo Cazzaniga, Marco S. Nobile |
PDP | 4 |
| 2017 | Reboot strategies in particle swarm optimization and their impact on parameter estimation of biochemical systemsabstractComputational methods adopted in the field of Systems Biology require the complete knowledge of reaction kinetic constants to perform simulations of the dynamics and understand the emergent behavior of biochemical systems. However, kinetic parameters of biochemical reactions are often difficult or impossible to measure, thus they are generally inferred from experimental data, in a process known as Parameter Estimation (PE). We consider here a PE methodology that exploits Particle Swarm Optimization (PSO) to estimate an appropriate kinetic parameterization, by comparing experimental time-series target data with in silica dynamics, simulated by using the parameterization encoded by each particle. In this work we present three different reboot strategies for PSO, whose aim is to reinitialize particle positions to avoid particles to get trapped in local optima, and we compare the performance of PSO coupled with the reboot strategies with respect to standard PSO in the case of the PE of two biochemical systems. Since the PE requires a huge number of simulations at each iteration, in this work we exploit a GPU-powered deterministic simulator, cupSODA, which performs in a parallel fashion all simulations and fitness evaluations. Finally, we show that the performances of our implementation scale sublinearly with respect to the swarm size, even on outdated GPUs. Simone Spolaor, Andrea Tangherloni, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga |
CIBCB | 5 |
| 2017 | Graphics processing units in bioinformatics, computational biology and systems biologyabstractSeveral studies in Bioinformatics, Computational Biology and Systems Biology rely on the definition of physico-chemical or mathematical models of biological systems at different scales and levels of complexity, ranging from the interaction of atoms in single molecules up to genome-wide interaction networks. Traditional computational methods and software tools developed in these research fields share a common trait: they can be computationally demanding on Central Processing Units (CPUs), therefore limiting their applicability in many circumstances. To overcome this issue, general-purpose Graphics Processing Units (GPUs) are gaining an increasing attention by the scientific community, as they can considerably reduce the running time required by standard CPU-based software, and allow more intensive investigations of biological systems. In this review, we present a collection of GPU tools recently developed to perform computational analyses in life science disciplines, emphasizing the advantages and the drawbacks in the use of these parallel architectures. The complete list of GPU-powered tools here reviewed is available at http://bit.ly/gputools. Marco S. Nobile, Paolo Cazzaniga, Andrea Tangherloni, Daniela Besozzi |
Briefings Bioinform. | 2 |
| 2017 | GPU-powered model analysis with PySB/cupSODAabstractSUMMARY: A major barrier to the practical utilization of large, complex models of biochemical systems is the lack of open-source computational tools to evaluate model behaviors over high-dimensional parameter spaces. This is due to the high computational expense of performing thousands to millions of model simulations required for statistical analysis. To address this need, we have implemented a user-friendly interface between cupSODA, a GPU-powered kinetic simulator, and PySB, a Python-based modeling and simulation framework. For three example models of varying size, we show that for large numbers of simulations PySB/cupSODA achieves order-of-magnitude speedups relative to a CPU-based ordinary differential equation integrator. AVAILABILITY AND IMPLEMENTATION: The PySB/cupSODA interface has been integrated into the PySB modeling framework (version 1.4.0), which can be installed from the Python Package Index (PyPI) using a Python package manager such as pip. cupSODA source code and precompiled binaries (Linux, Mac OS/X, Windows) are available at github.com/aresio/cupSODA (requires an Nvidia GPU; developer.nvidia.com/cuda-gpus). Additional information about PySB is available at pysb.org. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leonard A. Harris, Marco S. Nobile, James C. Pino, Alexander L. R. Lubbock, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga, Carlos F. Lopez |
Bioinform. | 7 |
| 2017 | LASSIE: simulating large-scale models of biochemical systems on GPUsabstractBACKGROUND: Mathematical modeling and in silico analysis are widely acknowledged as complementary tools to biological laboratory methods, to achieve a thorough understanding of emergent behaviors of cellular processes in both physiological and perturbed conditions. Though, the simulation of large-scale models-consisting in hundreds or thousands of reactions and molecular species-can rapidly overtake the capabilities of Central Processing Units (CPUs). The purpose of this work is to exploit alternative high-performance computing solutions, such as Graphics Processing Units (GPUs), to allow the investigation of these models at reduced computational costs. RESULTS: LASSIE is a "black-box" GPU-accelerated deterministic simulator, specifically designed for large-scale models and not requiring any expertise in mathematical modeling, simulation algorithms or GPU programming. Given a reaction-based model of a cellular process, LASSIE automatically generates the corresponding system of Ordinary Differential Equations (ODEs), assuming mass-action kinetics. The numerical solution of the ODEs is obtained by automatically switching between the Runge-Kutta-Fehlberg method in the absence of stiffness, and the Backward Differentiation Formulae of first order in presence of stiffness. The computational performance of LASSIE are assessed using a set of randomly generated synthetic reaction-based models of increasing size, ranging from 64 to 8192 reactions and species, and compared to a CPU-implementation of the LSODA numerical integration algorithm. CONCLUSIONS: LASSIE adopts a novel fine-grained parallelization strategy to distribute on the GPU cores all the calculations required to solve the system of ODEs. By virtue of this implementation, LASSIE achieves up to 92× speed-up with respect to LSODA, therefore reducing the running time from approximately 1 month down to 8 h to simulate models consisting in, for instance, four thousands of reactions and species. Notably, thanks to its smaller memory footprint, LASSIE is able to perform fast simulations of even larger models, whereby the tested CPU-implementation of LSODA failed to reach termination. LASSIE is therefore expected to make an important breakthrough in Systems Biology applications, for the execution of faster and in-depth computational analyses of large-scale models of complex biological systems. Andrea Tangherloni, Marco S. Nobile, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga |
BMC Bioinform. | 5 |
| 2017 | Efficient Simulation of Reaction Systems on Graphics Processing UnitsabstractReaction systems represent a theoretical framework based on the regulation mechanisms of facilitation and inhibition of biochemical reactions. The dynamic process defined by a reaction system is typically derived by hand, starting from the set of reactions and a given context sequence. However, thi s procedure may be error-prone and time-consuming, especially when the size of the reaction system increases. Here we present HERESY, a simulator of reaction systems accelerated on Graphics Processing Units (GPUs). HERESY is based on a fine-grained parallelization strategy, whereby all reactions are simultaneously executed on the GPU, therefore reducing the overall running time of the simulation. HERESY is particularly advantageous for the simulation of large-scale reaction systems, consisting of hundreds or thousands of reactions. By considering as test case some reaction systems with an increasing number of reactions and entities, as well as an increasing number of entities per reaction, we show that HERESY allows up to 29× speed-up with respect to a CPU-based simulator of reaction systems. Finally, we provide some directions for the optimization of HERESY, considering minimal reaction systems in normal form. Marco S. Nobile, Antonio E. Porreca, Simone Spolaor, Luca Manzoni, Paolo Cazzaniga, Giancarlo Mauri, Daniela Besozzi |
Fundam. Informaticae | 5 |
| 2017 | Gillespie's Stochastic Simulation Algorithm on MIC coprocessors
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri |
J. Supercomput. | 3 |
| 2016 | GPU-powered and settings-free parameter estimation of biochemical systemsabstractTo understand the emergent behavior of biochemical systems, computational analyses generally require the inference of unknown reaction kinetic constants, a problem known as parameter estimation (PE). In this work we propose a PE methodology that exploits Particle Swarm Optimization (PSO) to examine a set of candidate kinetic parameterizations, whose fitness is evaluated by comparing given target time-series of experimental data with in silico dynamics, simulated by using the parameterization encoded by each particle. In particular, we consider a Fuzzy Logic-based version of PSO - called Proactive Particles in Swarm Optimization (PPSO) - that automatically tunes the setting (inertia, cognitive and social factors) of each particle, independently from all other particles in the swarm. Since the optimization phase requires a large number of simulations for each particle at each iteration, we exploit a GPU-accelerated deterministic simulator, called cupSODA, that automatically generates the system of Ordinary Differential Equations associated with the biochemical system and performs its simulation for each candidate parameterization. We compare the performance of PPSO with respect to PSO for the PE problem by considering two biochemical systems as test cases. In addition, we evaluate the impact on PE of different strategies adopted, both in PPSO and PSO, for the selection of the initial positions of particles within the search space. We prove the effectiveness of our settings-free PE methodology by showing that PPSO outperforms PSO with respect to the computational time required to execute the optimization, achieving comparable results concerning the fitness of the best parameterization found. Marco S. Nobile, Andrea Tangherloni, Daniela Besozzi, Paolo Cazzaniga |
CEC | 4 |
| 2016 | Parallel implementation of efficient search schemes for the inference of cancer progression modelsabstractThe emergence and development of cancer is a consequence of the accumulation over time of genomic mutations involving a specific set of genes, which provides the cancer clones with a functional selective advantage. In this work, we model the order of accumulation of such mutations during the progression, which eventually leads to the disease, by means of probabilistic graphic models, i.e., Bayesian Networks (BNs). We investigate how to perform the task of learning the structure of such BNs, according to experimental evidence, adopting a global optimization meta-heuristics. In particular, in this work we rely on Genetic Algorithms, and to strongly reduce the execution time of the inference-which can also involve multiple repetitions to collect statistically significant assessments of the data-we distribute the calculations using both multi-threading and a multi-node architecture. The results show that our approach is characterized by good accuracy and specificity; we also demonstrate its feasibility, thanks to a 84× reduction of the overall execution time with respect to a traditional sequential implementation. Daniele Ramazzotti, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Marco Antoniotti |
CIBCB | 3 |
| 2016 | GPU-powered Bat Algorithm for the parameter estimation of biochemical kinetic valuesabstractThe emergent behavior of biochemical systems can be investigated by means of mathematical modeling and computational analyses, which usually require the automatic inference of the unknown values of the model's parameters. This problem, known as Parameter Estimation (PE), is usually tackled with bio-inspired meta-heuristics for global optimization, most notably Particle Swarm Optimization (PSO). In this work we assess the performances of PSO and Bat Algorithm with differential operator and Lévy flights trajectories (DLBA). In particular, we compared these meta-heuristics for the PE using two biochemical models: the expression of genes in prokaryotes and the heat shock response in eukaryotes. In our tests, we also evaluated the impact on PE of different strategies for the initial positioning of individuals within the search space. Our results show that DLBA achieves comparable results with respect to PSO, but it converges to better results when a uniform initialization is employed. Since every iteration of DLBA requires three fitness evaluations for each bat, the whole methodology is built around a GPU-powered biochemical simulator (cupSODA) which is able to parallelize the process. We show that the acceleration achieved with cupSODA strongly reduces the running time, with an empirical 61× speedup that has been obtained comparing a Nvidia GeForce Titan GTX with respect to a CPU Intel Core i7-4790K. Moreover, we show that DLBA always outperforms PSO with respect to the computational time required to execute the optimization process. Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga |
CIBCB | 3 |
| 2015 | The impact of particles initialization in PSO: Parameter estimation as a case in pointabstractDespite the intense research focused on the investigation of the functioning settings of Particle Swarm Optimization, the particles initialization functions - determining the initial positions in the search space - are generally ignored, especially in the case of real-world applications. As a matter of fact, almost all works exploit uniform distributions to randomly generate the particles coordinates. In this article, we analyze the impact on the optimization performances of alternative initialization functions based on logarithmic, normal, and lognormal distributions. Our results show how different initialization strategies can affect - and in some cases largely improve - the convergence speed, both in the case of benchmark functions and in the optimization of the kinetic constants of biochemical systems. Paolo Cazzaniga, Marco S. Nobile, Daniela Besozzi |
CIBCB | 1 |
| 2015 | Proactive Particles in Swarm Optimization: A self-tuning algorithm based on Fuzzy LogicabstractAmong the existing global optimization algorithms, Particle Swarm Optimization (PSO) is one of the most effective when dealing with non-linear and complex high-dimensional problems. However, the performance of PSO is strongly dependent on the choice of its settings. In this work we propose a novel and self-tuning PSO algorithm - called Proactive Particles in Swarm Optimization (PPSO) - which exploits Fuzzy Logic to calculate the best setting for the inertia, cognitive factor and social factor. Thanks to additional heuristics, PPSO automatically determines also the best setting for the swarm size and for the particles maximum velocity. PPSO significantly differs from other versions of PSO that exploit Fuzzy Logic, since specific settings are assigned to each particle according to its history, instead of being globally defined for the whole swarm. Thus, the novelty of PPSO is that particles gain a limited autonomous and proactive intelligence, instead of being simple reactive agents. Our results show that PPSO outperforms the standard PSO, both in terms of convergence speed and average quality of solutions, remarkably without the need for any user setting. Marco S. Nobile, Gabriella Pasi, Paolo Cazzaniga, Daniela Besozzi, Riccardo Colombo, Giancarlo Mauri |
FUZZ-IEEE | 3 |
| 2014 | A memetic hybrid method for the Molecular Distance Geometry Problem with incomplete informationabstractThe definition of computational methodologies for the inference of molecular structural information plays a relevant role in disciplines as drug discovery and metabolic engineering, since the functionality of a biochemical molecule is determined by its three-dimensional structure. In this work, we present an automatic methodology to solve the Molecular Distance Geometry Problem, that is, to determine the best three-dimensional shape that satisfies a given set of target inter-atomic distances. In particular, our method is designed to cope with incomplete distance information derived from Nuclear Magnetic Resonance measurements. To tackle this problem, that is known to be NP-hard, we present a memetic method that combines two soft-computing algorithms - Particle Swarm Optimization and Genetic Algorithms - with a local search approach, to improve the effectiveness of the crossover mechanism. We show the validity of our method on a set of reference molecules with a length ranging from 402 to 1003 atoms. Marco S. Nobile, Andrea G. Citrolo, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri |
IEEE Congress on Evolutionary Computation | 3 |
| 2014 | Simulation and Analysis of the Blood Coagulation Cascade Accelerated on GPUabstractThe use of Graphics Processing Units (GPUs) has recently witnessed ever growing applications for different computational analyses in the field of Life Sciences. In this work we present a CUDA-powered computational tool, named coagSODA, that was purposely developed and applied for the analysis of a large model of the blood coagulation cascade defined as a system of ordinary differential equations, based on both mass-action kinetics and Hill functions. We discuss the biological results of the parameter sweep analyses of this model, and show that GPUs can boost the computational performances up to 177x speedup. Matteo Bellini, Daniela Besozzi, Paolo Cazzaniga, Giancarlo Mauri, Marco S. Nobile |
PDP | 3 |
| 2014 | GPU-accelerated simulations of mass-action kinetics models with cupSODA
Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri |
J. Supercomput. | 2 |
| 2013 | Reverse engineering of kinetic reaction networks by means of Cartesian Genetic Programming and Particle Swarm OptimizationabstractThe modeling of biochemical reaction networks is a fundamental but complex task in Systems Biology, which is traditionally performed exploiting human expertise and the available experimental data. Because of the general lack of knowledge on the molecular mechanisms occurring in living cells, an intense research activity focused on the development of reverse engineering methodologies is currently underway. This problem is further complicated by the fact that a proper parameterization needs to be associated to the reaction network, in order to investigate its dynamical behavior. In this work we propose a novel computational methodology for the reverse engineering of fully parameterized kinetic networks, based on the combined use of two evolutionary programming techniques: Cartesian Genetic Programming (CGP) and Particle Swarm Optimization (PSO). In particular, CGP is used to infer the network topology, while PSO performs the parameter estimation task. To the purpose of applying our methodology in routine laboratory environments, we designed it to exploit a small set of experimental time series as target. We show that our methodology is able to reconstruct kinetic networks that perfectly fit with the target data. Marco S. Nobile, Daniela Besozzi, Paolo Cazzaniga, Dario Pescini, Giancarlo Mauri |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Computing with energy and chemical reactions
Alberto Leporati, Daniela Besozzi, Paolo Cazzaniga, Dario Pescini, Claudio Ferretti |
Nat. Comput. | 3 |
| 2009 | (Tissue) P systems with cell polarityabstractWe consider the structure of the intestinal epithelial tissue and of cell–cell junctions as the biological model inspiring a new class of P systems. First we define the concept of cell polarity, a formal property derived from epithelial cells, which present morphologically and functionally distinct regions of the plasma membrane. Then we show two preliminary results for this new model of computation: on the theoretical side, we show that P systems with cell polarity are computationally (Turing) complete; on the modelling side, we show that the transepithelial movement of glucose from the intestinal lumen into the blood can be described by such a formal system. Finally, we define tissue P systems with cell polarity, where each cell has fixed connections to the neighbouring cells and to the environment, according to both the cell polarity and specific cell–cell junctions. Daniela Besozzi, Nadia Busi, Paolo Cazzaniga, Claudio Ferretti, Alberto Leporati, Giancarlo Mauri, Dario Pescini, Claudio Zandron |
Math. Struct. Comput. Sci. | 3 |
| 2007 | Cycles and communicating classes in membrane systems and molecular dynamics
Michael Muskulus, Daniela Besozzi, Robert Brijder, Paolo Cazzaniga, Sanne Houweling, Dario Pescini, Grzegorz Rozenberg |
Theor. Comput. Sci. | 4 |