Daniela Besozzi

dblp:75/2051 · DBLP profile ↗
← Back
44ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-5532-3059ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 7 since 2021Theory of computation · 7 · 4 first-authorSystems, architecture and hardware · 3
YearPublicationVenuePosition
2026 HyCAPS: A Settings-Free Optimization Heuristics Integrating Evolutionary Computation and Swarm Intelligence
Daniele M. Papetti, Marco S. Nobile, Matteo Grazioso, Paolo Cazzaniga, Leonardo Vanneschi, Daniela Besozzi
EvoApplications6
2025 We Are Sending You Back... to the Optimum! Fuzzy Time Travel Particle Swarm Optimization
Daniele M. Papetti, Andrea Tangherloni, Vasco Coelho, Daniela Besozzi, Paolo Cazzaniga, Marco S. Nobile
EvoApplications (2)4
2023 The Domination Game: Dilating Bubbles to Fill Up Pareto Fronts
abstract
Multi-objective optimization algorithms might struggle in finding optimal dominating solutions, especially in real-case scenarios where problems are generally characterized by non-separability, non-differentiability, and multi-modality issues. An effective strategy that already showed to improve the outcome of optimization algorithms consists in manipulating the search space, in order to explore its most promising areas. In this work, starting from a Pareto front identified by an optimization strategy, we exploit Local Bubble Dilation Functions (LBDFs) to manipulate a locally bounded region of the search space containing non-dominated solutions. We tested our approach on the benchmark functions included in the DTLZ and WFG suites, showing that the Pareto front obtained after the application of LBDFs is most of the time characterized by an increased hyper-volume value. Our results confirm that LBDFs are an effective means to identify additional non-dominated solutions that can improve the quality of the Pareto front.
Vasco Coelho, Daniele M. Papetti, Andrea Tangherloni, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CEC5
2023 Unsupervised neural networks as a support tool for pathology diagnosis in MALDI-MSI experiments: A case study on thyroid biopsies
abstract
Artificial intelligence is getting a foothold in medicine for disease screening and diagnosis. While typical machine learning methods require large labeled datasets for training and validation, their application is limited in clinical fields since ground truth information can hardly be obtained on a sizeable cohort of patients. Unsupervised neural networks – such as Self-Organizing Maps (SOMs) – represent an alternative approach to identifying hidden patterns in biomedical data. Here we investigate the feasibility of SOMs for the identification of malignant and non-malignant regions in liquid biopsies of thyroid nodules, on a patient-specific basis. MALDI-ToF (Matrix Assisted Laser Desorption Ionization - Time of Flight) mass spectrometry-imaging (MSI) was used to measure the spectral profile of bioptic samples. SOMs were then applied for the analysis of MALDI-MSI data of individual patients’ samples, also testing various pre-processing and agglomerative clustering methods to investigate their impact on SOMs’ discrimination efficacy. The final clustering was compared against the sample’s probability to be malignant, hyperplastic or related to Hashimoto thyroiditis as quantified by multinomial regression with LASSO. Our results show that SOMs are effective in separating the areas of a sample containing benign cells from those containing malignant cells. Moreover, they allow to overlap the different areas of cytological glass slides with the corresponding proteomic profile image, and inspect the specific weight of every cellular component in bioptic samples. We envision that this approach could represent an effective means to assist pathologists in diagnostic tasks, avoiding the need to manually annotate cytological images and the effort in creating labeled datasets.
Marco S. Nobile, Giulia Capitoli, Virgil Sowirono, Francesca Clerici, Isabella Piga, Kirsten van Abeelen, Fulvio Magni, Fabio Pagni, Stefania Galimberti, Paolo Cazzaniga, Daniela Besozzi
Expert Syst. Appl.11
2022 Local Bubble Dilation Functions: Hypersphere-bounded Landscape Deformations Simplify Global Optimization
abstract
Solving optimization problems is one of the most complex and widespread task in Computer Science. In many scenarios, finding the global optimum of a function is hampered by several features that characterize the fitness landscapes, such as noisiness, multi-modality, non-convexity, non-separability, and non-differentiability. In order to facilitate the optimization process, a variety of methods have been proposed to manipulate either the search space or the fitness landscape. Among these, Dilation Functions (DFs) were introduced to expand regions of the search space that are characterized by promising fitness values. In this work, we extend the family of DFs by introducing Local Bubble Dilation Functions (LBDFs), a novel approach that generates local distortions bounded by hyper-spheres. By performing an appropriate mapping of the search space, LBDFs can improve the optimization performance, since they expand and reveal the promising regions around the global optimum, while leaving the rest of the fitness landscape untouched. The additional advantage of LBDFs, with respect to DFs, is that different dilations can be applied to each dimension of the search space, which is useful in the case of asymmetric landscapes. In order to show the benefits of local dilations, we executed several tests on the Michalewicz benchmark function, with different settings for the LBDFs. Our results show that a properly designed LBDF can lead to statistically significant better results than using vanilla optimization. Finally, we investigated the use of LBDFs to facilitate the solution of the parameter estimation problem in Systems Biology by analyzing the landscape related to a stochastic model of enzyme kinetics.
Daniele M. Papetti, Vasco Coelho, Dan Ashlock, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Marco S. Nobile
CIBCB6
2021 If You Can't Beat It, Squash It: Simplify Global Optimization by Evolving Dilation Functions
abstract
Optimization problems represent a class of pervasive and complex tasks in Computer Science, aimed at identifying the global optimum of a given objective function. Optimization problems are typically noisy, multi-modal, non-convex, non-separable, and often non-differentiable. Because of these features, they mandate the use of sophisticated population-based meta-heuristics to effectively explore the search space. Additionally, computational techniques based on the manipulation of the optimization landscape, such as Dilation Functions (DFs), can be effectively exploited to either "compress" or "dilate" some target regions of the search space, in order to improve the exploration and exploitation capabilities of any meta-heuristic. The main limitation of DFs is that they must be tailored on the specific optimization problem under investigation. In this work, we propose a solution to this issue, based on the idea of evolving the DFs. Specifically, we introduce a two-layered evolutionary framework, which combines Evolutionary Computation and Swarm Intelligence to solve the meta-problem of optimizing both the structure and the parameters of DFs. We evolved optimal DFs on a variety of benchmark problems, showing that this approach yields extremely simpler versions of the original optimization problems.
Daniele M. Papetti, Dan Ashlock, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CEC4
2021 The Impact of Representation on the Optimization of Marker Panels for Single-cell RNA Data
abstract
The increasing number of single-cell transcriptomic and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism as well as the onset of pathologies. These studies continuously refine the functional roles of known cell populations, and provide their characterization as soon as putatively novel cell populations are detected. In order to isolate the cell populations for further tailored analysis, succinct marker panels—composed of a few cell surface proteins and clusters of differentiation molecules—must be identified. The identification of these marker panels is a challenging computational problem due to its intrinsic combinatorial nature, which makes it an NP-hard problem. Genetic Algorithms (GAs) have been successfully used in Bioinformatics and other biomedical applications to tackle combinatorial problems. We present here a GA-based approach to solve the problem of the identification of succinct marker panels. Since the performance of a GA is strictly related to the representation of the candidate solutions, we propose and compare three alternative representations, able to implicitly introduce different constraints on the search space. For each representation, we perform a fine-tuning of the parameter settings to calibrate the GA, and we show that different representations yield different performance, where the most relaxed representations— in which the GA can also evolve the number of genes in the panel—turn out to be the more effective, especially in the case of 0-knowledge problems. Our results also show that the marker panels identified by GAs can outperform manually curated solutions.
Andrea Tangherloni, Simone G. Riva, Simone Spolaor, Daniela Besozzi, Marco S. Nobile, Paolo Cazzaniga
CEC4
2021 A comparison of multi-objective optimization algorithms to identify drug target combinations
abstract
Combination therapies represent one of the most effective strategy in inducing cancer cell death and reducing the risk to develop drug resistance. The identification of putative novel drug combinations, which typically requires the execution of expensive and time consuming lab experiments, can be supported by the synergistic use of mathematical models and multi-objective optimization algorithms. The computational approach allows to automatically search for potential therapeutic combinations and to test their effectiveness in silico, thus reducing the costs of time and money, and driving the experiments toward the most promising therapies. In this work, we couple dynamic fuzzy modeling of cancer cells with different multi-objective optimization algorithm, and we compare their performance in identifying drug target combinations. Specifically, we perform batches of optimizations with 3 and 4 objective functions defined to achieve a desired behavior of the system (e.g., maximize apop-tosis while minimizing necrosis and survival), and we compare the quality of the solutions included in the Pareto fronts. Our results show that both the choice of the multi-objective algorithm and the formulation of the optimization problem have an impact on the identified solutions, highlighting the strengths as well as the limitations of this approach.
Simone Spolaor, Daniele M. Papetti, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CIBCB4
2021 Selected papers from the 15th and 16th international conference on Computational Intelligence Methods for Bioinformatics and Biostatistics
abstract
This supplement contains seven revised and extended papers selected from CIBB 2018 and CIBB 2019, the 15th and 16th editions of the international conference on Computational Intelligence Methods for Bioinformatics and Biostatistics. CIBB is a venue that embraces researchers with different backgrounds, ranging from mathematics to computer science, from materials science to medicine, and from engineering to biology, all interested in the investigation and application of computational intelligence methods to open problems in bioinformatics, biostatistics, systems biology, synthetic biology, and medical informatics.
Paolo Cazzaniga, Maria Raposo, Daniela Besozzi, Ivan Merelli, Antonino Staiano, Angelo Ciaramella, Riccardo Rizzo, Luca Manzoni
BMC Bioinform.3
2021 Analysis of single-cell RNA sequencing data based on autoencoders
abstract
BACKGROUND: Single-cell RNA sequencing (scRNA-Seq) experiments are gaining ground to study the molecular processes that drive normal development as well as the onset of different pathologies. Finding an effective and efficient low-dimensional representation of the data is one of the most important steps in the downstream analysis of scRNA-Seq data, as it could provide a better identification of known or putatively novel cell-types. Another step that still poses a challenge is the integration of different scRNA-Seq datasets. Though standard computational pipelines to gain knowledge from scRNA-Seq data exist, a further improvement could be achieved by means of machine learning approaches. RESULTS: Autoencoders (AEs) have been effectively used to capture the non-linearities among gene interactions of scRNA-Seq data, so that the deployment of AE-based tools might represent the way forward in this context. We introduce here scAEspy, a unifying tool that embodies: (1) four of the most advanced AEs, (2) two novel AEs that we developed on purpose, (3) different loss functions. We show that scAEspy can be coupled with various batch-effect removal tools to integrate data by different scRNA-Seq platforms, in order to better identify the cell-types. We benchmarked scAEspy against the most used batch-effect removal tools, showing that our AE-based strategies outperform the existing solutions. CONCLUSIONS: scAEspy is a user-friendly tool that enables using the most recent and promising AEs to analyse scRNA-Seq data by only setting up two user-defined parameters. Thanks to its modularity, scAEspy can be easily extended to accommodate new AEs to further improve the downstream analysis of scRNA-Seq data. Considering the relevant results we achieved, scAEspy can be considered as a starting point to build a more comprehensive toolkit designed to integrate multi single-cell omics.
Andrea Tangherloni, Federico Ricciuti, Daniela Besozzi, Pietro Liò, Ana Cvejic
BMC Bioinform.3
2021 FiCoS: A fine-grained and coarse-grained GPU-powered deterministic simulator for biochemical networks
abstract
Mathematical models of biochemical networks can largely facilitate the comprehension of the mechanisms at the basis of cellular processes, as well as the formulation of hypotheses that can be tested by means of targeted laboratory experiments. However, two issues might hamper the achievement of fruitful outcomes. On the one hand, detailed mechanistic models can involve hundreds or thousands of molecular species and their intermediate complexes, as well as hundreds or thousands of chemical reactions, a situation generally occurring in rule-based modeling. On the other hand, the computational analysis of a model typically requires the execution of a large number of simulations for its calibration, or to test the effect of perturbations. As a consequence, the computational capabilities of modern Central Processing Units can be easily overtaken, possibly making the modeling of biochemical networks a worthless or ineffective effort. To the aim of overcoming the limitations of the current state-of-the-art simulation approaches, we present in this paper FiCoS, a novel "black-box" deterministic simulator that effectively realizes both a fine-grained and a coarse-grained parallelization on Graphics Processing Units. In particular, FiCoS exploits two different integration methods, namely, the Dormand-Prince and the Radau IIA, to efficiently solve both non-stiff and stiff systems of coupled Ordinary Differential Equations. We tested the performance of FiCoS against different deterministic simulators, by considering models of increasing size and by running analyses with increasing computational demands. FiCoS was able to dramatically speedup the computations up to 855×, showing to be a promising solution for the simulation and analysis of large-scale models of complex biological processes.
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Giulia Capitoli, Simone Spolaor, Leonardo Rundo, Giancarlo Mauri, Daniela Besozzi
PLoS Comput. Biol.8
2020 Which random is the best random? A study on sampling methods in Fourier surrogate modeling
abstract
Global optimization problems can be effectively solved by means of Computational Intelligence methods. However, there are several areas in which the effectiveness of these algorithms can be hampered by the computational costs of the fitness evaluations, or by specific features of the fitness landscape that can be characterized by noise and by the presence of several (even infinite) local optima. These issues bring about the necessity of defining specific techniques to replace the original problem with a surrogate representation. Fourier surrogate modeling represents a novel and effective approach to generate smoother, and possibly easier to explore, fitness landscapes, and to reduce the computational effort. Fourier surrogates require an initial sampling of the search space that must be performed to calculate the Fourier transforms. In this paper we investigate the impact on the quality of the surrogate models of the hyper-parameters of the methodology, and of several methods that can be employed for the initial sampling of the fitness landscape (i.e., pseudorandom numbers, low discrepancy sequences, a logistic map in chaotic regime, true random positions generated by a quantum computer, and point packing). Our results show that semistructured approaches like quasi-random sequences and point packing can outperform the other sampling methods.
Marco S. Nobile, Simone Spolaor, Paolo Cazzaniga, Daniele M. Papetti, Daniela Besozzi, Dan Ashlock, Luca Manzoni
CEC5
2020 Fourier Surrogate Models of Dilated Fitness Landscapes in Systems Biology : or how we learned to torture optimization problems until they confess
abstract
One of the most complex problems in Systems Biology is Parameter estimation (PE), which consists in inferring the kinetic parameters of biochemical systems. The identification of an accurate parameterization, able to reproduce any observed experimental behavior, is fundamental for the definition of predictive models. PE is a non-convex, multi-modal, and non-separable problem that is usually tackled by using Computational Intelligence methods. When the biochemical species appear in the system in a very low amount, the intrinsic noise due to the randomness of molecular collisions cannot be neglected. In this case, stochastic simulation algorithms should be employed to properly reproduce the system dynamics. Stochastic fluctuations make the PE problem even more complicated, as they can lead to radically different values of the fitness function for the same candidate parameterization. In addition, the kinetic parameters generally follow a log-uniform distribution, so that global optima tend to be localized in the lowest orders of magnitude of the search space. To simultaneously tackle all the aforementioned issues, in this work we investigate a novel approach based on the combination of dilation functions with Fourier surrogate modeling and filtering on the fitness landscape. The results show that our approach is able to strongly simplify the PE problem for low-dimensional optimization instances.
Marco S. Nobile, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Luca Manzoni
CIBCB4
2020 On the automatic calibration of fully analogical spiking neuromorphic chips
abstract
Nowadays, understanding the topology of biological neural networks and sampling their activity is possible thanks to various laboratory protocols that provide a large amount of experimental data, thus paving the way to accurate modeling and simulation. Neuromorphic systems were developed to simulate the dynamics of biological neural networks by means of electronic circuits, offering an efficient alternative to classic simulations based on systems of differential equations, from both the points of view of the energy consumed and the overall computational effort. Spikey is a configurable neuromorphic chip based on the Leaky Integrate-And-Fire model, which gives the user the possibility to model an arbitrary neural topology and simulate the temporal evolution of membrane potentials. To accurately reproduce the behavior of a specific biological network, a detailed parameterization of all neurons in the neuromorphic chip is necessary. Determining such parameters is a hard, error-prone, and generally time consuming task. In this work, we propose a novel methodology for the automatic calibration of neuromorphic chips that exploits a given neural activity as target. Our results show that, in the case of small networks with a low complexity, the method can estimate a vector of parameters capable of reproducing the target activity. Conversely, in the case of more complex networks, the simulations with Spikey can be highly affected by noise, which causes small variations in the simulations outcome even when identical networks are simulated, hindering the convergence to optimal parameterizations.
Daniele M. Papetti, Simone Spolaor, Daniela Besozzi, Paolo Cazzaniga, Marco Antoniotti, Marco S. Nobile
IJCNN3
2020 Fuzzy modeling and global optimization to predict novel therapeutic targets in cancer cells
abstract
MOTIVATION: The elucidation of dysfunctional cellular processes that can induce the onset of a disease is a challenging issue from both the experimental and computational perspectives. Here we introduce a novel computational method based on the coupling between fuzzy logic modeling and a global optimization algorithm, whose aims are to (1) predict the emergent dynamical behaviors of highly heterogeneous systems in unperturbed and perturbed conditions, regardless of the availability of quantitative parameters, and (2) determine a minimal set of system components whose perturbation can lead to a desired system response, therefore facilitating the design of a more appropriate experimental strategy. RESULTS: We applied this method to investigate what drives K-ras-induced cancer cells, displaying the typical Warburg effect, to death or survival upon progressive glucose depletion. The optimization analysis allowed to identify new combinations of stimuli that maximize pro-apoptotic processes. Namely, our results provide different evidences of an important protective role for protein kinase A in cancer cells under several cellular stress conditions mimicking tumor behavior. The predictive power of this method could facilitate the assessment of the response of other complex heterogeneous systems to drugs or mutations in fields as medicine and pharmacology, therefore paving the way for the development of novel therapeutic treatments. AVAILABILITY AND IMPLEMENTATION: The source code of FUMOSO is available under the GPL 2.0 license on GitHub at the following URL: https://github.com/aresio/FUMOSO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco S. Nobile, Giuseppina Votta, Roberta Palorini, Simone Spolaor, Humberto De Vitto, Paolo Cazzaniga, Francesca Ricciardiello, Giancarlo Mauri, Lilia Alberghina, Ferdinando Chiaradonna, Daniela Besozzi
Bioinform.11
2020 Computational Intelligence for Life Sciences
abstract
Computational Intelligence (CI) is a computer science discipline encompassing the theory, design, development and application of biologically and linguistically derived computational paradigms. Traditionally, the main elements of CI are Evolutionary Computation, Swarm Intelligence, Fuzzy Logic, and Neural Networks. CI aims at proposing new algorithms able to solve complex computational problems by taking inspiration from natural phenomena. In an intriguing turn of events, these nature-inspired methods have been widely adopted to investigate a plethora of problems related to nature itself. In this paper we present a variety of CI methods applied to three problems in life sciences, highlighting their effectiveness: we describe how protein folding can be faced by exploiting Genetic Programming, the inference of haplotypes can be tackled using Genetic Algorithms, and the estimation of biochemical kinetic parameters can be performed by means of Swarm Intelligence. We show that CI methods can generate very high quality solutions, providing a sound methodology to solve complex optimization problems in life sciences.
Daniela Besozzi, Luca Manzoni, Marco S. Nobile, Simone Spolaor, Mauro Castelli, Leonardo Vanneschi, Paolo Cazzaniga, Stefano Ruberto, Leonardo Rundo, Andrea Tangherloni
Fundam. Informaticae1
2020 Coupling Mechanistic Approaches and Fuzzy Logic to Model and Simulate Complex Systems
abstract
Several mathematical formalisms can be exploited to model complex systems, in order to capture different features of their dynamic behavior and leverage any available quantitative or qualitative data. Correspondingly, either quantitative models or qualitative models can be defined; bridging the gap between these two worlds would allow us to simultaneously exploit the peculiar advantages provided by each modeling approach. However, to date, the attempts in this direction have been limited to specific fields of research. In this paper, we propose a novel, general-purpose computational framework, named Fuzzy-mechanistic modeling of compleX systems (FuzzX), for the analysis of hybrid models consisting of a quantitative (or mechanistic) module and a qualitative module that can reciprocally control each other's dynamic behavior through a common interface. FuzzX takes advantage of precise quantitative information about the system through the definition and simulation of the mechanistic module. At the same time, it describes the behavior of components and their interactions that are not known in full details, by exploiting fuzzy logic for the definition of the qualitative module. We applied FuzzX for the analysis of a hybrid model of a complex biochemical system, characterized by the presence of positive and negative feedback regulations. We show that FuzzX is able to correctly reproduce known emergent behaviors of this system in normal and perturbed conditions. We envision that FuzzX could be employed to analyze any kind of complex system when quantitative information is limited, as well as to extend existing mechanistic models with fuzzy modules to describe those components and interactions of the system that are not fully characterized.
Simone Spolaor, Marco S. Nobile, Giancarlo Mauri, Paolo Cazzaniga, Daniela Besozzi
IEEE Trans. Fuzzy Syst.5
2020 cuProCell: GPU-Accelerated Analysis of Cell Proliferation With Flow Cytometry Data
abstract
The investigation of cell proliferation can provide useful insights for the comprehension of cancer progression, resistance to chemotherapy and relapse. To this aim, computational methods and experimental measurements based on in vivo label-retaining assays can be coupled to explore the dynamic behavior of tumoral cells. ProCell is a software that exploits flow cytometry data to model and simulate the kinetics of fluorescence loss that is due to stochastic events of cell division. Since the rate of cell division is not known, ProCell embeds a calibration process that might require thousands of stochastic simulations to properly infer the parameterization of cell proliferation models. To mitigate the high computational costs, in this paper we introduce a parallel implementation of ProCell's simulation algorithm, named cuProCell, which leverages Graphics Processing Units (GPUs). Dynamic Parallelism was used to efficiently manage the cell duplication events, in a radically different way with respect to common computing architectures. We present the advantages of cuProCell for the analysis of different models of cell proliferation in Acute Myeloid Leukemia (AML), using data collected from the spleen of human xenografts in mice. We show that, by exploiting GPUs, our method is able to not only automatically infer the models' parameterization, but it is also 237× faster than the sequential implementation. This study highlights the presence of a relevant percentage of quiescent and potentially chemoresistant cells in AML in vivo, and suggests that maintaining a dynamic equilibrium among the different proliferating cell populations might play an important role in disease progression.
Marco S. Nobile, Eric Nisoli, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
IEEE J. Biomed. Health Informatics8
2019 ProCell: Investigating cell proliferation with Swarm Intelligence
abstract
Computational methods represent an effective mean for the analysis of complex biological processes, such as cell proliferation, especially when combined to well established experimental protocols. In particular, mathematical modeling coupled with computational intelligence algorithms can be successfully exploited to investigate different aspects of cell population dynamics in the context of tumor growth. To this aim, we defined ProCell, a modeling and simulation framework specifically designed for the investigation of cell proliferation, which makes use of Fuzzy Self-Tuning Particle Swarm Optimization to estimate the unknown parameters of cell population models. ProCell is here applied to the analysis of cell proliferation in acute myeloid leukemia, a hematological malignancy characterized by an inherent intra-tumoral heterogeneity that plays an important role in disease recurrence and resistance to chemotherapy. ProCell allowed to provide new insights on the intricate organization of cells with highly heterogeneous proliferative potential, and to highlight the important role of different cell types in the progression and evolution of the disease. ProCell is available under the GPL 2.0 license on GitHub at https://github.com/aresio/ProCell.
Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
CIBCB7
2019 Modeling cell proliferation in human acute myeloid leukemia xenografts
abstract
MOTIVATION: Acute myeloid leukemia (AML) is one of the most common hematological malignancies, characterized by high relapse and mortality rates. The inherent intra-tumor heterogeneity in AML is thought to play an important role in disease recurrence and resistance to chemotherapy. Although experimental protocols for cell proliferation studies are well established and widespread, they are not easily applicable to in vivo contexts, and the analysis of related time-series data is often complex to achieve. To overcome these limitations, model-driven approaches can be exploited to investigate different aspects of cell population dynamics. RESULTS: In this work, we present ProCell, a novel modeling and simulation framework to investigate cell proliferation dynamics that, differently from other approaches, takes into account the inherent stochasticity of cell division events. We apply ProCell to compare different models of cell proliferation in AML, notably leveraging experimental data derived from human xenografts in mice. ProCell is coupled with Fuzzy Self-Tuning Particle Swarm Optimization, a swarm-intelligence settings-free algorithm used to automatically infer the models parameterizations. Our results provide new insights on the intricate organization of AML cells with highly heterogeneous proliferative potential, highlighting the important role played by quiescent cells and proliferating cells characterized by different rates of division in the progression and evolution of the disease, thus hinting at the necessity to further characterize tumor cell subpopulations. AVAILABILITY AND IMPLEMENTATION: The source code of ProCell and the experimental data used in this work are available under the GPL 2.0 license on GITHUB at the following URL: https://github.com/aresio/ProCell. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Daniela Bossi, Paolo Cazzaniga, Luisa Lanfrancone, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
Bioinform.9
2019 GenHap: a novel computational method based on genetic algorithms for haplotype assembly
abstract
BACKGROUND: In order to fully characterize the genome of an individual, the reconstruction of the two distinct copies of each chromosome, called haplotypes, is essential. The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the two chromosomes. Indeed, the knowledge of complete haplotypes is generally more informative than analyzing single SNPs and plays a fundamental role in many medical applications. RESULTS: To reconstruct the two haplotypes, we addressed the weighted Minimum Error Correction (wMEC) problem, which is a successful approach for haplotype assembly. This NP-hard problem consists in computing the two haplotypes that partition the sequencing reads into two disjoint sub-sets, with the least number of corrections to the SNP values. To this aim, we propose here GenHap, a novel computational method for haplotype assembly based on Genetic Algorithms, yielding optimal solutions by means of a global search process. In order to evaluate the effectiveness of our approach, we run GenHap on two synthetic (yet realistic) datasets, based on the Roche/454 and PacBio RS II sequencing technologies. We compared the performance of GenHap against HapCol, an efficient state-of-the-art algorithm for haplotype phasing. Our results show that GenHap always obtains high accuracy solutions (in terms of haplotype error rate), and is up to 4× faster than HapCol in the case of Roche/454 instances and up to 20× faster when compared on the PacBio RS II dataset. Finally, we assessed the performance of GenHap on two different real datasets. CONCLUSIONS: Future-generation sequencing technologies, producing longer reads with higher coverage, can highly benefit from GenHap, thanks to its capability of efficiently solving large instances of the haplotype assembly problem. Moreover, the optimization approach proposed in GenHap can be extended to the study of allele-specific genomic features, such as expression, methylation and chromatin conformation, by exploiting multi-objective optimization techniques. The source code and the full documentation are available at the following GitHub repository: https://github.com/andrea-tango/GenHap .
Andrea Tangherloni, Simone Spolaor, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Pietro Liò, Ivan Merelli, Daniela Besozzi
BMC Bioinform.9
2019 MedGA: A novel evolutionary method for image enhancement in medical imaging systems
Leonardo Rundo, Andrea Tangherloni, Marco S. Nobile, Carmelo Militello, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
Expert Syst. Appl.5
2019 USE-Net: Incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets
Leonardo Rundo, Yudai Nagano, Ryuichiro Hataya, Carmelo Militello, Andrea Tangherloni, Marco S. Nobile, Claudio Ferretti, Daniela Besozzi, Maria Carla Gilardi, Salvatore Vitabile, Giancarlo Mauri, Hideki Nakayama, Paolo Cazzaniga
Neurocomputing10
2019 ginSODA: massive parallel integration of stiff ODE systems on GPUs
Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.3
2018 Computational Intelligence for Parameter Estimation of Biochemical Systems
abstract
In the field of Systems Biology, simulating the dynamics of biochemical models represents one of the most effective methodologies to understand the functioning of cellular processes in normal or altered conditions. However, the lack of kinetic rates, necessary to perform accurate simulations, strongly limits the scope of these analyses. Parameter Estimation (PE), which consists in identifying a proper model parameterization, is a non-linear, non-convex and multi-modal optimization problem, typically tackled by means of Computational Intelligence techniques, such as Evolutionary Computation and Swarm Intelligence. In this work, we perform a thorough investigation of the most widespread methods for PE-namely, Artificial Bee Colony (ABC), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Differential Evolution (DE), Estimation of Distribution Algorithm (EDA), Genetic Algorithms (GAs), Particle Swarm Optimization (PSO), and Fuzzy Self-Tuning PSO (FST-PSO)-comparing their performances on a set of synthetic (yet realistic) biochemical models of increasing size and complexity. Our results show that a variant of the settings-free FST-PSO algorithm can consistently outperform all other methods; ABC and GAs represent the most performing alternatives, while methods based on multivariate normal distributions (e.g., CMA-ES, EDA) struggle to keep pace with the other approaches.
Marco S. Nobile, Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
CEC5
2017 Graphics processing units in bioinformatics, computational biology and systems biology
abstract
Several studies in Bioinformatics, Computational Biology and Systems Biology rely on the definition of physico-chemical or mathematical models of biological systems at different scales and levels of complexity, ranging from the interaction of atoms in single molecules up to genome-wide interaction networks. Traditional computational methods and software tools developed in these research fields share a common trait: they can be computationally demanding on Central Processing Units (CPUs), therefore limiting their applicability in many circumstances. To overcome this issue, general-purpose Graphics Processing Units (GPUs) are gaining an increasing attention by the scientific community, as they can considerably reduce the running time required by standard CPU-based software, and allow more intensive investigations of biological systems. In this review, we present a collection of GPU tools recently developed to perform computational analyses in life science disciplines, emphasizing the advantages and the drawbacks in the use of these parallel architectures. The complete list of GPU-powered tools here reviewed is available at http://bit.ly/gputools.
Marco S. Nobile, Paolo Cazzaniga, Andrea Tangherloni, Daniela Besozzi
Briefings Bioinform.4
2017 GPU-powered model analysis with PySB/cupSODA
abstract
SUMMARY: A major barrier to the practical utilization of large, complex models of biochemical systems is the lack of open-source computational tools to evaluate model behaviors over high-dimensional parameter spaces. This is due to the high computational expense of performing thousands to millions of model simulations required for statistical analysis. To address this need, we have implemented a user-friendly interface between cupSODA, a GPU-powered kinetic simulator, and PySB, a Python-based modeling and simulation framework. For three example models of varying size, we show that for large numbers of simulations PySB/cupSODA achieves order-of-magnitude speedups relative to a CPU-based ordinary differential equation integrator. AVAILABILITY AND IMPLEMENTATION: The PySB/cupSODA interface has been integrated into the PySB modeling framework (version 1.4.0), which can be installed from the Python Package Index (PyPI) using a Python package manager such as pip. cupSODA source code and precompiled binaries (Linux, Mac OS/X, Windows) are available at github.com/aresio/cupSODA (requires an Nvidia GPU; developer.nvidia.com/cuda-gpus). Additional information about PySB is available at pysb.org. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leonard A. Harris, Marco S. Nobile, James C. Pino, Alexander L. R. Lubbock, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga, Carlos F. Lopez
Bioinform.5
2017 LASSIE: simulating large-scale models of biochemical systems on GPUs
abstract
BACKGROUND: Mathematical modeling and in silico analysis are widely acknowledged as complementary tools to biological laboratory methods, to achieve a thorough understanding of emergent behaviors of cellular processes in both physiological and perturbed conditions. Though, the simulation of large-scale models-consisting in hundreds or thousands of reactions and molecular species-can rapidly overtake the capabilities of Central Processing Units (CPUs). The purpose of this work is to exploit alternative high-performance computing solutions, such as Graphics Processing Units (GPUs), to allow the investigation of these models at reduced computational costs. RESULTS: LASSIE is a "black-box" GPU-accelerated deterministic simulator, specifically designed for large-scale models and not requiring any expertise in mathematical modeling, simulation algorithms or GPU programming. Given a reaction-based model of a cellular process, LASSIE automatically generates the corresponding system of Ordinary Differential Equations (ODEs), assuming mass-action kinetics. The numerical solution of the ODEs is obtained by automatically switching between the Runge-Kutta-Fehlberg method in the absence of stiffness, and the Backward Differentiation Formulae of first order in presence of stiffness. The computational performance of LASSIE are assessed using a set of randomly generated synthetic reaction-based models of increasing size, ranging from 64 to 8192 reactions and species, and compared to a CPU-implementation of the LSODA numerical integration algorithm. CONCLUSIONS: LASSIE adopts a novel fine-grained parallelization strategy to distribute on the GPU cores all the calculations required to solve the system of ODEs. By virtue of this implementation, LASSIE achieves up to 92× speed-up with respect to LSODA, therefore reducing the running time from approximately 1 month down to 8 h to simulate models consisting in, for instance, four thousands of reactions and species. Notably, thanks to its smaller memory footprint, LASSIE is able to perform fast simulations of even larger models, whereby the tested CPU-implementation of LSODA failed to reach termination. LASSIE is therefore expected to make an important breakthrough in Systems Biology applications, for the execution of faster and in-depth computational analyses of large-scale models of complex biological systems.
Andrea Tangherloni, Marco S. Nobile, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
BMC Bioinform.3
2017 Efficient Simulation of Reaction Systems on Graphics Processing Units
abstract
Reaction systems represent a theoretical framework based on the regulation mechanisms of facilitation and inhibition of biochemical reactions. The dynamic process defined by a reaction system is typically derived by hand, starting from the set of reactions and a given context sequence. However, thi s procedure may be error-prone and time-consuming, especially when the size of the reaction system increases. Here we present HERESY, a simulator of reaction systems accelerated on Graphics Processing Units (GPUs). HERESY is based on a fine-grained parallelization strategy, whereby all reactions are simultaneously executed on the GPU, therefore reducing the overall running time of the simulation. HERESY is particularly advantageous for the simulation of large-scale reaction systems, consisting of hundreds or thousands of reactions. By considering as test case some reaction systems with an increasing number of reactions and entities, as well as an increasing number of entities per reaction, we show that HERESY allows up to 29× speed-up with respect to a CPU-based simulator of reaction systems. Finally, we provide some directions for the optimization of HERESY, considering minimal reaction systems in normal form.
Marco S. Nobile, Antonio E. Porreca, Simone Spolaor, Luca Manzoni, Paolo Cazzaniga, Giancarlo Mauri, Daniela Besozzi
Fundam. Informaticae7
2017 Gillespie's Stochastic Simulation Algorithm on MIC coprocessors
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.4
2016 GPU-powered and settings-free parameter estimation of biochemical systems
abstract
To understand the emergent behavior of biochemical systems, computational analyses generally require the inference of unknown reaction kinetic constants, a problem known as parameter estimation (PE). In this work we propose a PE methodology that exploits Particle Swarm Optimization (PSO) to examine a set of candidate kinetic parameterizations, whose fitness is evaluated by comparing given target time-series of experimental data with in silico dynamics, simulated by using the parameterization encoded by each particle. In particular, we consider a Fuzzy Logic-based version of PSO - called Proactive Particles in Swarm Optimization (PPSO) - that automatically tunes the setting (inertia, cognitive and social factors) of each particle, independently from all other particles in the swarm. Since the optimization phase requires a large number of simulations for each particle at each iteration, we exploit a GPU-accelerated deterministic simulator, called cupSODA, that automatically generates the system of Ordinary Differential Equations associated with the biochemical system and performs its simulation for each candidate parameterization. We compare the performance of PPSO with respect to PSO for the PE problem by considering two biochemical systems as test cases. In addition, we evaluate the impact on PE of different strategies adopted, both in PPSO and PSO, for the selection of the initial positions of particles within the search space. We prove the effectiveness of our settings-free PE methodology by showing that PPSO outperforms PSO with respect to the computational time required to execute the optimization, achieving comparable results concerning the fitness of the best parameterization found.
Marco S. Nobile, Andrea Tangherloni, Daniela Besozzi, Paolo Cazzaniga
CEC3
2016 Reaction-Based Models of Biochemical Networks
Daniela Besozzi
CiE1
2015 The impact of particles initialization in PSO: Parameter estimation as a case in point
abstract
Despite the intense research focused on the investigation of the functioning settings of Particle Swarm Optimization, the particles initialization functions - determining the initial positions in the search space - are generally ignored, especially in the case of real-world applications. As a matter of fact, almost all works exploit uniform distributions to randomly generate the particles coordinates. In this article, we analyze the impact on the optimization performances of alternative initialization functions based on logarithmic, normal, and lognormal distributions. Our results show how different initialization strategies can affect - and in some cases largely improve - the convergence speed, both in the case of benchmark functions and in the optimization of the kinetic constants of biochemical systems.
Paolo Cazzaniga, Marco S. Nobile, Daniela Besozzi
CIBCB3
2015 Proactive Particles in Swarm Optimization: A self-tuning algorithm based on Fuzzy Logic
abstract
Among the existing global optimization algorithms, Particle Swarm Optimization (PSO) is one of the most effective when dealing with non-linear and complex high-dimensional problems. However, the performance of PSO is strongly dependent on the choice of its settings. In this work we propose a novel and self-tuning PSO algorithm - called Proactive Particles in Swarm Optimization (PPSO) - which exploits Fuzzy Logic to calculate the best setting for the inertia, cognitive factor and social factor. Thanks to additional heuristics, PPSO automatically determines also the best setting for the swarm size and for the particles maximum velocity. PPSO significantly differs from other versions of PSO that exploit Fuzzy Logic, since specific settings are assigned to each particle according to its history, instead of being globally defined for the whole swarm. Thus, the novelty of PPSO is that particles gain a limited autonomous and proactive intelligence, instead of being simple reactive agents. Our results show that PPSO outperforms the standard PSO, both in terms of convergence speed and average quality of solutions, remarkably without the need for any user setting.
Marco S. Nobile, Gabriella Pasi, Paolo Cazzaniga, Daniela Besozzi, Riccardo Colombo, Giancarlo Mauri
FUZZ-IEEE4
2014 A memetic hybrid method for the Molecular Distance Geometry Problem with incomplete information
abstract
The definition of computational methodologies for the inference of molecular structural information plays a relevant role in disciplines as drug discovery and metabolic engineering, since the functionality of a biochemical molecule is determined by its three-dimensional structure. In this work, we present an automatic methodology to solve the Molecular Distance Geometry Problem, that is, to determine the best three-dimensional shape that satisfies a given set of target inter-atomic distances. In particular, our method is designed to cope with incomplete distance information derived from Nuclear Magnetic Resonance measurements. To tackle this problem, that is known to be NP-hard, we present a memetic method that combines two soft-computing algorithms - Particle Swarm Optimization and Genetic Algorithms - with a local search approach, to improve the effectiveness of the crossover mechanism. We show the validity of our method on a set of reference molecules with a length ranging from 402 to 1003 atoms.
Marco S. Nobile, Andrea G. Citrolo, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
IEEE Congress on Evolutionary Computation4
2014 Simulation and Analysis of the Blood Coagulation Cascade Accelerated on GPU
abstract
The use of Graphics Processing Units (GPUs) has recently witnessed ever growing applications for different computational analyses in the field of Life Sciences. In this work we present a CUDA-powered computational tool, named coagSODA, that was purposely developed and applied for the analysis of a large model of the blood coagulation cascade defined as a system of ordinary differential equations, based on both mass-action kinetics and Hill functions. We discuss the biological results of the parameter sweep analyses of this model, and show that GPUs can boost the computational performances up to 177x speedup.
Matteo Bellini, Daniela Besozzi, Paolo Cazzaniga, Giancarlo Mauri, Marco S. Nobile
PDP2
2014 GPU-accelerated simulations of mass-action kinetics models with cupSODA
Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.3
2013 Reverse engineering of kinetic reaction networks by means of Cartesian Genetic Programming and Particle Swarm Optimization
abstract
The modeling of biochemical reaction networks is a fundamental but complex task in Systems Biology, which is traditionally performed exploiting human expertise and the available experimental data. Because of the general lack of knowledge on the molecular mechanisms occurring in living cells, an intense research activity focused on the development of reverse engineering methodologies is currently underway. This problem is further complicated by the fact that a proper parameterization needs to be associated to the reaction network, in order to investigate its dynamical behavior. In this work we propose a novel computational methodology for the reverse engineering of fully parameterized kinetic networks, based on the combined use of two evolutionary programming techniques: Cartesian Genetic Programming (CGP) and Particle Swarm Optimization (PSO). In particular, CGP is used to infer the network topology, while PSO performs the parameter estimation task. To the purpose of applying our methodology in routine laboratory environments, we designed it to exploit a small set of experimental time series as target. We show that our methodology is able to reconstruct kinetic networks that perfectly fit with the target data.
Marco S. Nobile, Daniela Besozzi, Paolo Cazzaniga, Dario Pescini, Giancarlo Mauri
IEEE Congress on Evolutionary Computation2
2012 An excursion in reaction systems: From computer science to biology
Luca Corolli, Carlo Maj, Fabrizio Marini, Daniela Besozzi, Giancarlo Mauri
Theor. Comput. Sci.4
2010 Computing with energy and chemical reactions
Alberto Leporati, Daniela Besozzi, Paolo Cazzaniga, Dario Pescini, Claudio Ferretti
Nat. Comput.2
2009 (Tissue) P systems with cell polarity
abstract
We consider the structure of the intestinal epithelial tissue and of cell–cell junctions as the biological model inspiring a new class of P systems. First we define the concept of cell polarity, a formal property derived from epithelial cells, which present morphologically and functionally distinct regions of the plasma membrane. Then we show two preliminary results for this new model of computation: on the theoretical side, we show that P systems with cell polarity are computationally (Turing) complete; on the modelling side, we show that the transepithelial movement of glucose from the intestinal lumen into the blood can be described by such a formal system. Finally, we define tissue P systems with cell polarity, where each cell has fixed connections to the neighbouring cells and to the environment, according to both the cell polarity and specific cell–cell junctions.
Daniela Besozzi, Nadia Busi, Paolo Cazzaniga, Claudio Ferretti, Alberto Leporati, Giancarlo Mauri, Dario Pescini, Claudio Zandron
Math. Struct. Comput. Sci.1
2007 Cycles and communicating classes in membrane systems and molecular dynamics
Michael Muskulus, Daniela Besozzi, Robert Brijder, Paolo Cazzaniga, Sanne Houweling, Dario Pescini, Grzegorz Rozenberg
Theor. Comput. Sci.2
2005 On the power and size of extended gemmating P systems
Daniela Besozzi, Erzsébet Csuhaj-Varjú, Giancarlo Mauri, Claudio Zandron
Soft Comput.1
2003 Gemmating P systems: collapsing hierarchies
Daniela Besozzi, Giancarlo Mauri, Gheorghe Paun, Claudio Zandron
Theor. Comput. Sci.1