Simone Spolaor

dblp:204/5069 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-3383-367XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 1 since 2021Theory of computation · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 Ten quick tips for fuzzy logic modeling of biomedical systems
abstract
Fuzzy logic is useful tool to describe and represent biological or medical scenarios, where often states and outcomes are not only completely true or completely false, but rather partially true or partially false. Despite its usefulness and spread, fuzzy logic modeling might easily be done in the wrong way, especially by beginners and unexperienced researchers, who might overlook some important aspects or might make common mistakes. Malpractices and pitfalls, in turn, can lead to wrong or overoptimistic, inflated results, with negative consequences to the biomedical research community trying to comprehend a particular phenomenon, or even to patients suffering from the investigated disease. To avoid common mistakes, we present here a list of quick tips for fuzzy logic modeling any biomedical scenario: some guidelines which should be taken into account by any fuzzy logic practitioner, including experts. We believe our best practices can have a strong impact in the scientific community, allowing researchers who follow them to obtain better, more reliable results and outcomes in biomedical contexts.
Davide Chicco, Simone Spolaor, Marco S. Nobile
PLoS Comput. Biol.2
2022 The Impact of Variable Selection and Transformation on the Interpretability and Accuracy of Fuzzy Models
abstract
Data transformation is an important step in Machine Learning pipelines which can strongly improve their performance. For instance, min-max normalization is often used to make all variables lie in the same range, while log-transformation is used to map data that is scattered across several orders of magnitude to a logarithmic space. Such transformations can be beneficial when the machine learning approach measures distance in a metric space, such as cluster-based approaches. These two transformation approaches can be combined to reveal hidden patterns in the data in the case of log-normally distributed data points, which commonly occur in biological and medical data. In this work we introduce a novel evolutionary approach designed to automatically determine the optimal log-transformation and selection of variables. Our approach is built around an interpretable AI system (created by pyFUME), so that all transformations are followed by inverse transformations to map back the values into the original universe of discourse, and preserve the interpretability of the results. We test our approach on two synthetic datasets, designed to reproduce a condition in which some variables are normally distributed, some variables are log-normally distributed, and some variables are just noise in the dataset. Our results show that our approach yields better performing models compared to conventional methods, and that the resulting model is also characterised by a better interpretability, making such approach particularly useful to study biomedical datasets.
Caro Fuchs, Simone Spolaor, Uzay Kaymak, Marco S. Nobile
CIBCB2
2022 Local Bubble Dilation Functions: Hypersphere-bounded Landscape Deformations Simplify Global Optimization
abstract
Solving optimization problems is one of the most complex and widespread task in Computer Science. In many scenarios, finding the global optimum of a function is hampered by several features that characterize the fitness landscapes, such as noisiness, multi-modality, non-convexity, non-separability, and non-differentiability. In order to facilitate the optimization process, a variety of methods have been proposed to manipulate either the search space or the fitness landscape. Among these, Dilation Functions (DFs) were introduced to expand regions of the search space that are characterized by promising fitness values. In this work, we extend the family of DFs by introducing Local Bubble Dilation Functions (LBDFs), a novel approach that generates local distortions bounded by hyper-spheres. By performing an appropriate mapping of the search space, LBDFs can improve the optimization performance, since they expand and reveal the promising regions around the global optimum, while leaving the rest of the fitness landscape untouched. The additional advantage of LBDFs, with respect to DFs, is that different dilations can be applied to each dimension of the search space, which is useful in the case of asymmetric landscapes. In order to show the benefits of local dilations, we executed several tests on the Michalewicz benchmark function, with different settings for the LBDFs. Our results show that a properly designed LBDF can lead to statistically significant better results than using vanilla optimization. Finally, we investigated the use of LBDFs to facilitate the solution of the parameter estimation problem in Systems Biology by analyzing the landscape related to a stochastic model of enzyme kinetics.
Daniele M. Papetti, Vasco Coelho, Dan Ashlock, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Marco S. Nobile
CIBCB5
2021 The Impact of Representation on the Optimization of Marker Panels for Single-cell RNA Data
abstract
The increasing number of single-cell transcriptomic and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism as well as the onset of pathologies. These studies continuously refine the functional roles of known cell populations, and provide their characterization as soon as putatively novel cell populations are detected. In order to isolate the cell populations for further tailored analysis, succinct marker panels—composed of a few cell surface proteins and clusters of differentiation molecules—must be identified. The identification of these marker panels is a challenging computational problem due to its intrinsic combinatorial nature, which makes it an NP-hard problem. Genetic Algorithms (GAs) have been successfully used in Bioinformatics and other biomedical applications to tackle combinatorial problems. We present here a GA-based approach to solve the problem of the identification of succinct marker panels. Since the performance of a GA is strictly related to the representation of the candidate solutions, we propose and compare three alternative representations, able to implicitly introduce different constraints on the search space. For each representation, we perform a fine-tuning of the parameter settings to calibrate the GA, and we show that different representations yield different performance, where the most relaxed representations— in which the GA can also evolve the number of genes in the panel—turn out to be the more effective, especially in the case of 0-knowledge problems. Our results also show that the marker panels identified by GAs can outperform manually curated solutions.
Andrea Tangherloni, Simone G. Riva, Simone Spolaor, Daniela Besozzi, Marco S. Nobile, Paolo Cazzaniga
CEC3
2021 A comparison of multi-objective optimization algorithms to identify drug target combinations
abstract
Combination therapies represent one of the most effective strategy in inducing cancer cell death and reducing the risk to develop drug resistance. The identification of putative novel drug combinations, which typically requires the execution of expensive and time consuming lab experiments, can be supported by the synergistic use of mathematical models and multi-objective optimization algorithms. The computational approach allows to automatically search for potential therapeutic combinations and to test their effectiveness in silico, thus reducing the costs of time and money, and driving the experiments toward the most promising therapies. In this work, we couple dynamic fuzzy modeling of cancer cells with different multi-objective optimization algorithm, and we compare their performance in identifying drug target combinations. Specifically, we perform batches of optimizations with 3 and 4 objective functions defined to achieve a desired behavior of the system (e.g., maximize apop-tosis while minimizing necrosis and survival), and we compare the quality of the solutions included in the Pareto fronts. Our results show that both the choice of the multi-objective algorithm and the formulation of the optimization problem have an impact on the identified solutions, highlighting the strengths as well as the limitations of this approach.
Simone Spolaor, Daniele M. Papetti, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CIBCB1
2021 FiCoS: A fine-grained and coarse-grained GPU-powered deterministic simulator for biochemical networks
abstract
Mathematical models of biochemical networks can largely facilitate the comprehension of the mechanisms at the basis of cellular processes, as well as the formulation of hypotheses that can be tested by means of targeted laboratory experiments. However, two issues might hamper the achievement of fruitful outcomes. On the one hand, detailed mechanistic models can involve hundreds or thousands of molecular species and their intermediate complexes, as well as hundreds or thousands of chemical reactions, a situation generally occurring in rule-based modeling. On the other hand, the computational analysis of a model typically requires the execution of a large number of simulations for its calibration, or to test the effect of perturbations. As a consequence, the computational capabilities of modern Central Processing Units can be easily overtaken, possibly making the modeling of biochemical networks a worthless or ineffective effort. To the aim of overcoming the limitations of the current state-of-the-art simulation approaches, we present in this paper FiCoS, a novel "black-box" deterministic simulator that effectively realizes both a fine-grained and a coarse-grained parallelization on Graphics Processing Units. In particular, FiCoS exploits two different integration methods, namely, the Dormand-Prince and the Radau IIA, to efficiently solve both non-stiff and stiff systems of coupled Ordinary Differential Equations. We tested the performance of FiCoS against different deterministic simulators, by considering models of increasing size and by running analyses with increasing computational demands. FiCoS was able to dramatically speedup the computations up to 855×, showing to be a promising solution for the simulation and analysis of large-scale models of complex biological processes.
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Giulia Capitoli, Simone Spolaor, Leonardo Rundo, Giancarlo Mauri, Daniela Besozzi
PLoS Comput. Biol.5
2020 Which random is the best random? A study on sampling methods in Fourier surrogate modeling
abstract
Global optimization problems can be effectively solved by means of Computational Intelligence methods. However, there are several areas in which the effectiveness of these algorithms can be hampered by the computational costs of the fitness evaluations, or by specific features of the fitness landscape that can be characterized by noise and by the presence of several (even infinite) local optima. These issues bring about the necessity of defining specific techniques to replace the original problem with a surrogate representation. Fourier surrogate modeling represents a novel and effective approach to generate smoother, and possibly easier to explore, fitness landscapes, and to reduce the computational effort. Fourier surrogates require an initial sampling of the search space that must be performed to calculate the Fourier transforms. In this paper we investigate the impact on the quality of the surrogate models of the hyper-parameters of the methodology, and of several methods that can be employed for the initial sampling of the fitness landscape (i.e., pseudorandom numbers, low discrepancy sequences, a logistic map in chaotic regime, true random positions generated by a quantum computer, and point packing). Our results show that semistructured approaches like quasi-random sequences and point packing can outperform the other sampling methods.
Marco S. Nobile, Simone Spolaor, Paolo Cazzaniga, Daniele M. Papetti, Daniela Besozzi, Dan Ashlock, Luca Manzoni
CEC2
2020 Fourier Surrogate Models of Dilated Fitness Landscapes in Systems Biology : or how we learned to torture optimization problems until they confess
abstract
One of the most complex problems in Systems Biology is Parameter estimation (PE), which consists in inferring the kinetic parameters of biochemical systems. The identification of an accurate parameterization, able to reproduce any observed experimental behavior, is fundamental for the definition of predictive models. PE is a non-convex, multi-modal, and non-separable problem that is usually tackled by using Computational Intelligence methods. When the biochemical species appear in the system in a very low amount, the intrinsic noise due to the randomness of molecular collisions cannot be neglected. In this case, stochastic simulation algorithms should be employed to properly reproduce the system dynamics. Stochastic fluctuations make the PE problem even more complicated, as they can lead to radically different values of the fitness function for the same candidate parameterization. In addition, the kinetic parameters generally follow a log-uniform distribution, so that global optima tend to be localized in the lowest orders of magnitude of the search space. To simultaneously tackle all the aforementioned issues, in this work we investigate a novel approach based on the combination of dilation functions with Fourier surrogate modeling and filtering on the fitness landscape. The results show that our approach is able to strongly simplify the PE problem for low-dimensional optimization instances.
Marco S. Nobile, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Luca Manzoni
CIBCB3
2020 pyFUME: a Python Package for Fuzzy Model Estimation
abstract
Living in the era of "data deluge" demands for an increase in the application and development of machine learning methods, both in basic and applied research. Among these methods, in the last decades fuzzy inference systems carved out their own niche as (light) grey box models, which are considered more interpretable and transparent than other commonly employed methods, such as artificial neural networks. Although commercially distributed alternatives are available, software able to assist practitioners and researchers in each step of the estimation of a fuzzy model from data are still limited in scope and applicability. This is especially true when looking at software developed in Python, a programming language that quickly gained popularity among data scientists and it is often considered their language of choice. To fill this gap, we introduce pyFUME, a Python library for automatically estimating fuzzy models from data. pyFUME contains a set of classes and methods to estimate the antecedent sets and the consequent parameters of a Takagi-Sugeno fuzzy model from data, and then create an executable fuzzy model exploiting the Simpful library. pyFUME can be beneficial to practitioners, thanks to its pre-implemented and user-friendly pipelines, but also to researchers that want to fine-tune each step of the estimation process.
Caro Fuchs, Simone Spolaor, Marco S. Nobile, Uzay Kaymak
FUZZ-IEEE2
2020 On the automatic calibration of fully analogical spiking neuromorphic chips
abstract
Nowadays, understanding the topology of biological neural networks and sampling their activity is possible thanks to various laboratory protocols that provide a large amount of experimental data, thus paving the way to accurate modeling and simulation. Neuromorphic systems were developed to simulate the dynamics of biological neural networks by means of electronic circuits, offering an efficient alternative to classic simulations based on systems of differential equations, from both the points of view of the energy consumed and the overall computational effort. Spikey is a configurable neuromorphic chip based on the Leaky Integrate-And-Fire model, which gives the user the possibility to model an arbitrary neural topology and simulate the temporal evolution of membrane potentials. To accurately reproduce the behavior of a specific biological network, a detailed parameterization of all neurons in the neuromorphic chip is necessary. Determining such parameters is a hard, error-prone, and generally time consuming task. In this work, we propose a novel methodology for the automatic calibration of neuromorphic chips that exploits a given neural activity as target. Our results show that, in the case of small networks with a low complexity, the method can estimate a vector of parameters capable of reproducing the target activity. Conversely, in the case of more complex networks, the simulations with Spikey can be highly affected by noise, which causes small variations in the simulations outcome even when identical networks are simulated, hindering the convergence to optimal parameterizations.
Daniele M. Papetti, Simone Spolaor, Daniela Besozzi, Paolo Cazzaniga, Marco Antoniotti, Marco S. Nobile
IJCNN2
2020 A Graph Theory Approach to Fuzzy Rule Base Simplification
Caro Fuchs, Simone Spolaor, Marco S. Nobile, Uzay Kaymak
IPMU (1)2
2020 Fuzzy modeling and global optimization to predict novel therapeutic targets in cancer cells
abstract
MOTIVATION: The elucidation of dysfunctional cellular processes that can induce the onset of a disease is a challenging issue from both the experimental and computational perspectives. Here we introduce a novel computational method based on the coupling between fuzzy logic modeling and a global optimization algorithm, whose aims are to (1) predict the emergent dynamical behaviors of highly heterogeneous systems in unperturbed and perturbed conditions, regardless of the availability of quantitative parameters, and (2) determine a minimal set of system components whose perturbation can lead to a desired system response, therefore facilitating the design of a more appropriate experimental strategy. RESULTS: We applied this method to investigate what drives K-ras-induced cancer cells, displaying the typical Warburg effect, to death or survival upon progressive glucose depletion. The optimization analysis allowed to identify new combinations of stimuli that maximize pro-apoptotic processes. Namely, our results provide different evidences of an important protective role for protein kinase A in cancer cells under several cellular stress conditions mimicking tumor behavior. The predictive power of this method could facilitate the assessment of the response of other complex heterogeneous systems to drugs or mutations in fields as medicine and pharmacology, therefore paving the way for the development of novel therapeutic treatments. AVAILABILITY AND IMPLEMENTATION: The source code of FUMOSO is available under the GPL 2.0 license on GitHub at the following URL: https://github.com/aresio/FUMOSO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco S. Nobile, Giuseppina Votta, Roberta Palorini, Simone Spolaor, Humberto De Vitto, Paolo Cazzaniga, Francesca Ricciardiello, Giancarlo Mauri, Lilia Alberghina, Ferdinando Chiaradonna, Daniela Besozzi
Bioinform.4
2020 Computational Intelligence for Life Sciences
abstract
Computational Intelligence (CI) is a computer science discipline encompassing the theory, design, development and application of biologically and linguistically derived computational paradigms. Traditionally, the main elements of CI are Evolutionary Computation, Swarm Intelligence, Fuzzy Logic, and Neural Networks. CI aims at proposing new algorithms able to solve complex computational problems by taking inspiration from natural phenomena. In an intriguing turn of events, these nature-inspired methods have been widely adopted to investigate a plethora of problems related to nature itself. In this paper we present a variety of CI methods applied to three problems in life sciences, highlighting their effectiveness: we describe how protein folding can be faced by exploiting Genetic Programming, the inference of haplotypes can be tackled using Genetic Algorithms, and the estimation of biochemical kinetic parameters can be performed by means of Swarm Intelligence. We show that CI methods can generate very high quality solutions, providing a sound methodology to solve complex optimization problems in life sciences.
Daniela Besozzi, Luca Manzoni, Marco S. Nobile, Simone Spolaor, Mauro Castelli, Leonardo Vanneschi, Paolo Cazzaniga, Stefano Ruberto, Leonardo Rundo, Andrea Tangherloni
Fundam. Informaticae4
2020 Coupling Mechanistic Approaches and Fuzzy Logic to Model and Simulate Complex Systems
abstract
Several mathematical formalisms can be exploited to model complex systems, in order to capture different features of their dynamic behavior and leverage any available quantitative or qualitative data. Correspondingly, either quantitative models or qualitative models can be defined; bridging the gap between these two worlds would allow us to simultaneously exploit the peculiar advantages provided by each modeling approach. However, to date, the attempts in this direction have been limited to specific fields of research. In this paper, we propose a novel, general-purpose computational framework, named Fuzzy-mechanistic modeling of compleX systems (FuzzX), for the analysis of hybrid models consisting of a quantitative (or mechanistic) module and a qualitative module that can reciprocally control each other's dynamic behavior through a common interface. FuzzX takes advantage of precise quantitative information about the system through the definition and simulation of the mechanistic module. At the same time, it describes the behavior of components and their interactions that are not known in full details, by exploiting fuzzy logic for the definition of the qualitative module. We applied FuzzX for the analysis of a hybrid model of a complex biochemical system, characterized by the presence of positive and negative feedback regulations. We show that FuzzX is able to correctly reproduce known emergent behaviors of this system in normal and perturbed conditions. We envision that FuzzX could be employed to analyze any kind of complex system when quantitative information is limited, as well as to extend existing mechanistic models with fuzzy modules to describe those components and interactions of the system that are not fully characterized.
Simone Spolaor, Marco S. Nobile, Giancarlo Mauri, Paolo Cazzaniga, Daniela Besozzi
IEEE Trans. Fuzzy Syst.1
2020 cuProCell: GPU-Accelerated Analysis of Cell Proliferation With Flow Cytometry Data
abstract
The investigation of cell proliferation can provide useful insights for the comprehension of cancer progression, resistance to chemotherapy and relapse. To this aim, computational methods and experimental measurements based on in vivo label-retaining assays can be coupled to explore the dynamic behavior of tumoral cells. ProCell is a software that exploits flow cytometry data to model and simulate the kinetics of fluorescence loss that is due to stochastic events of cell division. Since the rate of cell division is not known, ProCell embeds a calibration process that might require thousands of stochastic simulations to properly infer the parameterization of cell proliferation models. To mitigate the high computational costs, in this paper we introduce a parallel implementation of ProCell's simulation algorithm, named cuProCell, which leverages Graphics Processing Units (GPUs). Dynamic Parallelism was used to efficiently manage the cell duplication events, in a radically different way with respect to common computing architectures. We present the advantages of cuProCell for the analysis of different models of cell proliferation in Acute Myeloid Leukemia (AML), using data collected from the spleen of human xenografts in mice. We show that, by exploiting GPUs, our method is able to not only automatically infer the models' parameterization, but it is also 237× faster than the sequential implementation. This study highlights the presence of a relevant percentage of quiescent and potentially chemoresistant cells in AML in vivo, and suggests that maintaining a dynamic equilibrium among the different proliferating cell populations might play an important role in disease progression.
Marco S. Nobile, Eric Nisoli, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
IEEE J. Biomed. Health Informatics4
2019 ProCell: Investigating cell proliferation with Swarm Intelligence
abstract
Computational methods represent an effective mean for the analysis of complex biological processes, such as cell proliferation, especially when combined to well established experimental protocols. In particular, mathematical modeling coupled with computational intelligence algorithms can be successfully exploited to investigate different aspects of cell population dynamics in the context of tumor growth. To this aim, we defined ProCell, a modeling and simulation framework specifically designed for the investigation of cell proliferation, which makes use of Fuzzy Self-Tuning Particle Swarm Optimization to estimate the unknown parameters of cell population models. ProCell is here applied to the analysis of cell proliferation in acute myeloid leukemia, a hematological malignancy characterized by an inherent intra-tumoral heterogeneity that plays an important role in disease recurrence and resistance to chemotherapy. ProCell allowed to provide new insights on the intricate organization of cells with highly heterogeneous proliferative potential, and to highlight the important role of different cell types in the progression and evolution of the disease. ProCell is available under the GPL 2.0 license on GitHub at https://github.com/aresio/ProCell.
Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
CIBCB3
2019 A Swarm Intelligence Approach to Avoid Local Optima in Fuzzy C-Means Clustering
abstract
Clustering analysis is an important computational task that has applications in many domains. One of the most popular algorithms to solve the clustering problem is fuzzy c-means, which exploits notions from fuzzy logic to provide a smooth partitioning of the data into classes, allowing the possibility of multiple membership for each data sample. The fuzzy c-means algorithm is based on the optimization of a partitioning function, which minimizes inter-cluster similarity. This optimization problem is known to be NP-hard and it is generally tackled using a hill climbing method, a local optimizer that provides acceptable but sub-optimal solutions, since it is sensitive to initialization and tends to get stuck in local optima. In this work we propose an alternative approach based on the swarm intelligence global optimization method Fuzzy Self-Tuning Particle Swarm Optimization (FST-PSO). We solve the fuzzy clustering task by optimizing fuzzy c-means' partitioning function using FST-PSO. We show that this population-based metaheuristics is more effective than hill climbing, providing high quality solutions with the cost of an additional computational complexity. It is noteworthy that, since this particle swarm optimization algorithm is self-tuning, the user does not have to specify additional hyperparameters for the optimization process.
Caro Fuchs, Simone Spolaor, Marco S. Nobile, Uzay Kaymak
FUZZ-IEEE2
2019 Modeling cell proliferation in human acute myeloid leukemia xenografts
abstract
MOTIVATION: Acute myeloid leukemia (AML) is one of the most common hematological malignancies, characterized by high relapse and mortality rates. The inherent intra-tumor heterogeneity in AML is thought to play an important role in disease recurrence and resistance to chemotherapy. Although experimental protocols for cell proliferation studies are well established and widespread, they are not easily applicable to in vivo contexts, and the analysis of related time-series data is often complex to achieve. To overcome these limitations, model-driven approaches can be exploited to investigate different aspects of cell population dynamics. RESULTS: In this work, we present ProCell, a novel modeling and simulation framework to investigate cell proliferation dynamics that, differently from other approaches, takes into account the inherent stochasticity of cell division events. We apply ProCell to compare different models of cell proliferation in AML, notably leveraging experimental data derived from human xenografts in mice. ProCell is coupled with Fuzzy Self-Tuning Particle Swarm Optimization, a swarm-intelligence settings-free algorithm used to automatically infer the models parameterizations. Our results provide new insights on the intricate organization of AML cells with highly heterogeneous proliferative potential, highlighting the important role played by quiescent cells and proliferating cells characterized by different rates of division in the progression and evolution of the disease, thus hinting at the necessity to further characterize tumor cell subpopulations. AVAILABILITY AND IMPLEMENTATION: The source code of ProCell and the experimental data used in this work are available under the GPL 2.0 license on GITHUB at the following URL: https://github.com/aresio/ProCell. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Daniela Bossi, Paolo Cazzaniga, Luisa Lanfrancone, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
Bioinform.3
2019 GenHap: a novel computational method based on genetic algorithms for haplotype assembly
abstract
BACKGROUND: In order to fully characterize the genome of an individual, the reconstruction of the two distinct copies of each chromosome, called haplotypes, is essential. The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the two chromosomes. Indeed, the knowledge of complete haplotypes is generally more informative than analyzing single SNPs and plays a fundamental role in many medical applications. RESULTS: To reconstruct the two haplotypes, we addressed the weighted Minimum Error Correction (wMEC) problem, which is a successful approach for haplotype assembly. This NP-hard problem consists in computing the two haplotypes that partition the sequencing reads into two disjoint sub-sets, with the least number of corrections to the SNP values. To this aim, we propose here GenHap, a novel computational method for haplotype assembly based on Genetic Algorithms, yielding optimal solutions by means of a global search process. In order to evaluate the effectiveness of our approach, we run GenHap on two synthetic (yet realistic) datasets, based on the Roche/454 and PacBio RS II sequencing technologies. We compared the performance of GenHap against HapCol, an efficient state-of-the-art algorithm for haplotype phasing. Our results show that GenHap always obtains high accuracy solutions (in terms of haplotype error rate), and is up to 4× faster than HapCol in the case of Roche/454 instances and up to 20× faster when compared on the PacBio RS II dataset. Finally, we assessed the performance of GenHap on two different real datasets. CONCLUSIONS: Future-generation sequencing technologies, producing longer reads with higher coverage, can highly benefit from GenHap, thanks to its capability of efficiently solving large instances of the haplotype assembly problem. Moreover, the optimization approach proposed in GenHap can be extended to the study of allele-specific genomic features, such as expression, methylation and chromatin conformation, by exploiting multi-objective optimization techniques. The source code and the full documentation are available at the following GitHub repository: https://github.com/andrea-tango/GenHap .
Andrea Tangherloni, Simone Spolaor, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Pietro Liò, Ivan Merelli, Daniela Besozzi
BMC Bioinform.2
2018 Computational Intelligence for Parameter Estimation of Biochemical Systems
abstract
In the field of Systems Biology, simulating the dynamics of biochemical models represents one of the most effective methodologies to understand the functioning of cellular processes in normal or altered conditions. However, the lack of kinetic rates, necessary to perform accurate simulations, strongly limits the scope of these analyses. Parameter Estimation (PE), which consists in identifying a proper model parameterization, is a non-linear, non-convex and multi-modal optimization problem, typically tackled by means of Computational Intelligence techniques, such as Evolutionary Computation and Swarm Intelligence. In this work, we perform a thorough investigation of the most widespread methods for PE-namely, Artificial Bee Colony (ABC), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Differential Evolution (DE), Estimation of Distribution Algorithm (EDA), Genetic Algorithms (GAs), Particle Swarm Optimization (PSO), and Fuzzy Self-Tuning PSO (FST-PSO)-comparing their performances on a set of synthetic (yet realistic) biochemical models of increasing size and complexity. Our results show that a variant of the settings-free FST-PSO algorithm can consistently outperform all other methods; ABC and GAs represent the most performing alternatives, while methods based on multivariate normal distributions (e.g., CMA-ES, EDA) struggle to keep pace with the other approaches.
Marco S. Nobile, Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
CEC4
2018 GPU-Powered Multi-Swarm Parameter Estimation of Biological Systems: A Master-Slave Approach
abstract
In silico investigation of biological systems requires the knowledge of numerical parameters that cannot be easily measured in laboratory experiments, leading to the Parameter Estimation (PE) problem, in which the unknown parameters are automatically inferred by means of optimization algorithms exploiting the available experimental data. Here we present MS 2 PSO, an efficient parallel and distributed implementation of a PE method based on Particle Swarm Optimization (PSO) for the estimation of reaction constants in mathematical models of biological systems, considering as target for the estimation a set of discrete-time measurements of molecular species amounts. In particular, such PE method accounts for the availability of experimental data typically measured under different experimental conditions, by considering a multi-swarm PSO in which the best particles of the swarms can migrate. This strategy allows to infer a common set of reaction constants that simultaneously fits all target data used in the PE. To the aim of efficiently tackling the PE problem, MS 2 PSO embeds the execution of cupSODA, a deterministic simulator that relies on Graphics Processing Units to achieve a massive parallelization of the simulations required in the fitness evaluation of particles. In addition, a further level of parallelism is realized by exploiting the Master-Slave distributed programming paradigm. We apply MS 2 PSO for the PE of synthetic biochemical models with 10, 20 and 30 parameters to be estimated, and compare the performances obtained with different GPUs and different configurations (i.e., numbers of processes) of the Master-Slave.
Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Paolo Cazzaniga, Marco S. Nobile
PDP3
2017 Reboot strategies in particle swarm optimization and their impact on parameter estimation of biochemical systems
abstract
Computational methods adopted in the field of Systems Biology require the complete knowledge of reaction kinetic constants to perform simulations of the dynamics and understand the emergent behavior of biochemical systems. However, kinetic parameters of biochemical reactions are often difficult or impossible to measure, thus they are generally inferred from experimental data, in a process known as Parameter Estimation (PE). We consider here a PE methodology that exploits Particle Swarm Optimization (PSO) to estimate an appropriate kinetic parameterization, by comparing experimental time-series target data with in silica dynamics, simulated by using the parameterization encoded by each particle. In this work we present three different reboot strategies for PSO, whose aim is to reinitialize particle positions to avoid particles to get trapped in local optima, and we compare the performance of PSO coupled with the reboot strategies with respect to standard PSO in the case of the PE of two biochemical systems. Since the PE requires a huge number of simulations at each iteration, in this work we exploit a GPU-powered deterministic simulator, cupSODA, which performs in a parallel fashion all simulations and fitness evaluations. Finally, we show that the performances of our implementation scale sublinearly with respect to the swarm size, even on outdated GPUs.
Simone Spolaor, Andrea Tangherloni, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga
CIBCB1
2017 Efficient Simulation of Reaction Systems on Graphics Processing Units
abstract
Reaction systems represent a theoretical framework based on the regulation mechanisms of facilitation and inhibition of biochemical reactions. The dynamic process defined by a reaction system is typically derived by hand, starting from the set of reactions and a given context sequence. However, thi s procedure may be error-prone and time-consuming, especially when the size of the reaction system increases. Here we present HERESY, a simulator of reaction systems accelerated on Graphics Processing Units (GPUs). HERESY is based on a fine-grained parallelization strategy, whereby all reactions are simultaneously executed on the GPU, therefore reducing the overall running time of the simulation. HERESY is particularly advantageous for the simulation of large-scale reaction systems, consisting of hundreds or thousands of reactions. By considering as test case some reaction systems with an increasing number of reactions and entities, as well as an increasing number of entities per reaction, we show that HERESY allows up to 29× speed-up with respect to a CPU-based simulator of reaction systems. Finally, we provide some directions for the optimization of HERESY, considering minimal reaction systems in normal form.
Marco S. Nobile, Antonio E. Porreca, Simone Spolaor, Luca Manzoni, Paolo Cazzaniga, Giancarlo Mauri, Daniela Besozzi
Fundam. Informaticae3