Marco S. Nobile

dblp:61/11116 · also Marco Salvatore Nobile · DBLP profile ↗
← Back
67ranked-venue papers
21as first author
30since 2021 · last 2026
0000-0002-7692-7203ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 35 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 29 · 10 first-author · 12 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HyCAPS: A Settings-Free Optimization Heuristics Integrating Evolutionary Computation and Swarm Intelligence
Daniele M. Papetti, Marco S. Nobile, Matteo Grazioso, Paolo Cazzaniga, Leonardo Vanneschi, Daniela Besozzi
EvoApplications2
2025 Improving the Efficiency and the Validity of Molecular Transformers
abstract
Since their advent, Transformer models have been applied across a wide range of fields, including cheminformatics. In this context, drug discovery has benefited from using Molecular Transformers by leveraging diverse string representations of molecules, such as the Simplified Molecular Input Line Entry Systems (SMILES), for a variety of tasks. In this study, we present a model focused on the optimization of a formerly developed Molecular Transformer specifically dedicated to metabolism prediction. Metabolism refers to all the biotransformations a drug undergoes once inside the human body, directly influencing its therapeutic effect and potential toxicity, and therefore represents a key topic in medicinal chemistry. Framing molecular transformation prediction as a sequence-to-sequence translation task has shown promise, but suffers from limitations such as low validity of generated molecules and high computational cost. To address this limitation, we here propose an optimized model that integrates pre-training, transfer learning, and fine-tuning techniques, already improving validity and reducing computation time. Finally, by separating the metabolism prediction task from the SMILES syntax learning, we ensure broader applicability of the proposed model across diverse datasets and a variety of SMILES-based tasks beyond metabolic transformations, expanding its potential utility.
Leone Bacciu, Matteo Grazioso, Silvia Multari, Francesca Grisoni, Angelica Mazzolari, Marco S. Nobile
CIBCB6
2025 Assessing Cardiac Functionality by Means of Interpretable AI and Myocardial Strain
abstract
Cardiac Imaging is a powerful methodology for the accurate assessment of heart functionality. Among the possible approaches, Myocardial Strain assesses the functionality of the heart by tracking the movement and deformation of myocardium during the cardiac cycle. This information, that can be acquired also by means of Cardiac Magnetic Resonance, can pave the way to the development of predictive models using machine learning. In this work, we developed a predictive model of left ventricular ejection fraction, which is a measure of the heart’s function to pump oxygen-rich blood to the body, trained using strain data. Specifically, we developed a fully interpretable model based on a rule-based Fuzzy Inference System, coupled with a novel methodology for the disambiguation of the rules. Our results show that the developed model is able to accurately estimate the ejection fraction, and can provide physicians with additional insights about the role of strain features.
Marco S. Nobile, Amalia Lupi, Leone Bacciu, Matteo Grazioso, Chiara Gallese, Emilio Quaia, Alesia Pepe
CIBCB1
2025 A Survey of Modern Hybrid Particle Swarm Optimization Algorithms
Matteo Grazioso, Chiara Gallese, Leonardo Vanneschi, Marco S. Nobile
EvoApplications (2)4
2025 We Are Sending You Back... to the Optimum! Fuzzy Time Travel Particle Swarm Optimization
Daniele M. Papetti, Andrea Tangherloni, Vasco Coelho, Daniela Besozzi, Paolo Cazzaniga, Marco S. Nobile
EvoApplications (2)6
2024 Predicting Metabolic Reactions with a Molecular Transformer for Drug Design Optimization
abstract
Metabolism prediction is a crucial step of drug development, as the biotransformations a drug candidate undergoes inside the human body can affect the clinical outcome. Computer-aided drug design has been extensively employed to speed up the process and enhance its efficiency and effectiveness, but among the investigated areas, metabolism has received less attention. This project aimed at leveraging machine learning to analyze large metabolic datasets, make predictions, and recognize patterns, in order to fill this knowledge gap and enhance our understanding of metabolism and its impact on drug development. To achieve this goal, we developed a Deep Learning model for metabolism prediction using natural language processing techniques trained on molecular string representations, i.e., Simplified Molecular Input Line Entry Systems (SMILES) strings. To this end, we employ a Molecular Transformer, because of its ability to capture sequential and contextual information within strings (in this case, SMILES) enabling the learning of complex relationships. The transformer was trained using a high-quality dataset, MetaQSAR, from which we derived approximately 100000 instances of metabolic reactions. In this work, we investigate whether the Transformer architecture bears the potential to learn a mapping between the input molecular structures and their corresponding metabolites, in order to expedite drug discovery and improve patient safety.
Silvia Multari, Riza Özçelik, Angelica Mazzolari, Marco S. Nobile, Francesca Grisoni
CIBCB4
2024 A Fast Feature Selection for Interpretable Modeling Based on Fuzzy Inference Systems
abstract
Large datasets are often beneficial for the generation of predictive models using machine learning approaches. However, it is often the case that not all variables in the dataset contain useful information. In fact, some variables might be useless, redundant, misleading, or even harmful to performance, both in terms of accuracy and computational effort. Because of that, Feature Selection (FS) is one of the most delicate and important steps in machine learning. This is even more relevant in the case of interpretable models based on Fuzzy Inference Systems (FIS). The reasons are two-fold: on the one hand, FIS are generally built on top of a data partitioning based on clustering, which can suffer from high dimensionality; on the other hand, the knowledge base of the FIS, to be concretely understandable, should not contain rules involving too many variables. FS can be performed using multiple approaches, most notably filter and wrapper methods. The latter are often based on evolutionary algorithms, where a population of candidate solutions (each representing a possible set of selected variables) evolves towards the optimal selection. Although wrapper methods can be effective, they are, in general, computationally expensive. In this work, we propose a completely different – and more computationally effective – algorithm based on Random Forest (RF) models. Specifically, we exploit RFs to rank variables according to their importance. Then, we use that information to perform a statistical analysis and determine the minimal set of features necessary to build an accurate FIS. We show the effectiveness of our approach by using two (semi)synthetic datasets built on real-world datasets, and we validate our approach by applying the FS method to a medical dataset.
Andrea Tangherloni, Paolo Cazzaniga, Nicolò Stranieri, Francesca Buffa, Marco S. Nobile
CIBCB5
2024 An explainable data-driven decision support framework for strategic customer development
abstract
Financial institutions benefit from the advanced predictive performance of machine learning algorithms in automatic decision-making for credit scoring. However, two main challenges hamper machine learning algorithms’ applicability in practice: the complex and black-box nature of algorithms that hinder their understandability and the inability to guide rejected customers to have a successful application. Regarding customer relationship management is one of the main responsibilities of financial institutions; they must clarify the decision-making process to guide them. However, financial institutions are not willing to disclose their decision-making procedure to prevent potential risks from customers or competitors side. Hence, in this study, a decision support framework is proposed to clarify the decision-making process and model strategic decision-making to guide rejected customers simultaneously. To do so, after classifying customers in their corresponding groups, the capability of Shapley additive exPlanations method is exploited to extract the most impactful features to the prediction’s outcome globally and locally. Then, based on the benchmarking approach , the equivalent approved peer is found for the rejected customer for target setting to modify the application. To find the optimal modified values for a counterfactual prediction, a multi-objective gamed-based counterfactual explanation model is developed using the prisoner’s dilemma game as the constraint to simulate strategic decision-making. After optimization, the decision is reported to the customers concerning the credential background. A public data set is used to elaborate on the proposed framework. This framework can generate counterfactual predictions successfully by modifying perspective features.
Mohsen Abbaspour Onari, Mustafa Jahangoshai Rezaee, Morteza Saberi, Marco S. Nobile
Knowl. Based Syst.4
2023 Trustworthy Machine Learning Predictions to Support Clinical Research and Decisions
abstract
Nowadays, physicians have at their hands a huge amount of data produced by a large set of diagnostic and instrumental tests integrated with data obtained by high-throughput technologies. If such data were opportunely linked and analysed, they might be used to strengthen predictions, so that to improve the prevention and the time-to-diagnosis, reduce the costs of the health system, and bring out hidden knowledge. Machine learning is the principal technique used nowadays to leverage data and gain useful information. However, it has led to various challenges, such as improving the interpretability and explainability of the employed predictive models and integrating expert knowledge into the final system. Solving those challenges is of paramount importance to enhance the trust of both clinicians and patients in the system predictions. To solve the aforementioned issues, in this paper we propose a software workflow able to cope with the trustworthiness aspects of machine learning models and considering a multitude of heterogeneous data and models.
Andrea Bianchi, Antinisca Di Marco, Francesca Marzi, Giovanni Stilo, Cristina Pellegrini, Stefano Masi, Alessandro Mengozzi, Agostino Virdis, Marco S. Nobile, Marta Simeoni
CBMS9
2023 The Domination Game: Dilating Bubbles to Fill Up Pareto Fronts
abstract
Multi-objective optimization algorithms might struggle in finding optimal dominating solutions, especially in real-case scenarios where problems are generally characterized by non-separability, non-differentiability, and multi-modality issues. An effective strategy that already showed to improve the outcome of optimization algorithms consists in manipulating the search space, in order to explore its most promising areas. In this work, starting from a Pareto front identified by an optimization strategy, we exploit Local Bubble Dilation Functions (LBDFs) to manipulate a locally bounded region of the search space containing non-dominated solutions. We tested our approach on the benchmark functions included in the DTLZ and WFG suites, showing that the Pareto front obtained after the application of LBDFs is most of the time characterized by an increased hyper-volume value. Our results confirm that LBDFs are an effective means to identify additional non-dominated solutions that can improve the quality of the Pareto front.
Vasco Coelho, Daniele M. Papetti, Andrea Tangherloni, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CEC6
2023 Investigating Semi-Automatic Assessment of Data Sets Fairness by Means of Fuzzy Logic
abstract
Research has shown how data sets convey social bias in AI systems, especially those based on machine learning. A biased data set is not representative of reality and might contribute to perpetuate societal biases within the model. To tackle this problem, it is important to understand how to avoid biases, errors, and unethical practices while creating the data sets. In this work we offer a preliminary framework for the semi-automated evaluation of fairness in data sets, by combining statistical information about data with qualitative consideration. We address the issue of how much (un)fairness can be included in a data set used for machine learning research, focusing on classification issues. In order to provide guidance for the use of data sets in contexts of critical decision-making, such as health decisions, we identify six fundamental features (balance, numerosity, unevenness, compliance, quality, incompleteness) that could affect model fairness. We developed a rule-based approach based on fuzzy logic that combines these characteristics into a single score and enables a semi-automatic evaluation of a data set in algorithmic fairness research.
Chiara Gallese, Teresa Scantamburlo, Luca Manzoni, Marco S. Nobile
CIBCB4
2023 Trustworthy Artificial Intelligence in Medical Applications: A Mini Survey
abstract
Nowadays, a large amount of structured and unstructured data is being produced in various fields, creating tremendous opportunities to implement Machine Learning (ML) algorithms for decision-making. Although ML algorithms can outperform human performance in some fields, the black-box inherent characteristics of advanced models can hinder experts from exploiting them in sensitive domains such as medicine. The black-box nature of advanced ML models shadows the transparency of these algorithms, which could hamper their fair and robust performance due to the complexity of the algorithms. Consequently, individuals, organizations, and societies will not be able to achieve the full potential of ML without establishing trust in its development, deployment, and use. The field of eXplainable Artificial Intelligence (XAI) endeavors to solve this problem by providing human-understandable explanations for black-box models as a potential solution to acquire trustworthy AI. However, explainability is one of many requirements to fulfill trustworthy AI, and other prerequisites must also be met. Hence, this survey analyzes the fulfillment of five algorithmic requirements of accuracy, transparency, trust, robustness, and fairness through the lens of the literature in the medical domain. Regarding that medical experts are reluctant to put their judgment aside in favor of a machine, trustworthy AI algorithmic fulfillment could be a way to convince them to use ML. The results show there is still a long way to implement the algorithmic requirements in practice, and scholars need to consider them in future studies.
Mohsen Abbaspour Onari, Isel Grau, Marco S. Nobile, Yingqian Zhang 0001
CIBCB3
2023 Estimation of Fuzzy Models from Mixed Data Sets with pyFUME
abstract
pyFUME is a python package for the automatic estimation of fuzzy inference systems. Fuzzy models are considered among the most interpretable, understandable, and transparent methods that are currently available, making them ideal for the development of Interpretable AI systems. Such models are suitable for the creation of decision support systems in extremely sensitive domains where the right to an explanation is particularly important, like medicine and healthcare. pyFUME can automatically estimate the antecedent sets and the consequent parameters of a Takagi-Sugeno fuzzy model directly from data, and deliver an executable fuzzy model implemented with the Simpful python library. The main limitation of pyFUME was that it was not well-equipped to deal with purely categorical, non-ordinal variables since it used distance metrics suitable for continuous variables to cluster the data for determining the fuzzy model’s structure. In this paper, we introduce a new version of pyFUME that supports mixed (i.e., continuous and categorical) data sets, relying on a novel version of fuzzy Cprototypes clustering. Our results show that our new approach is effective, leading to better fitting with respect to models based only on continuous features. We also present alternative plotting methods tailored for categorical variables, which improves the overall interpretability of the estimated discrete fuzzy sets.
Daniele M. Papetti, Caro Fuchs, Vasco Coelho, Uzay Kaymak, Marco S. Nobile
CIBCB5
2023 A Theoretical Framework for AI Models Explainability with Application in Biomedicine
abstract
EXplainable Artificial Intelligence (XAI) is a vibrant research topic in the artificial intelligence community. It is raising growing interest across methods and domains, especially those involving high stake decision-making, such as the biomedical sector. Much has been written about the subject, yet XAI still lacks shared terminology and a framework capable of providing structural soundness to explanations. In our work, we address these issues by proposing a novel definition of explanation that synthesizes what can be found in the literature. We recognize that explanations are not atomic but the combination of evidence stemming from the model and its input-output mapping, and the human interpretation of this evidence. Furthermore, we fit explanations into the properties of faithfulness (i.e., the explanation is an accurate description of the model’s inner workings and decision-making process) and plausibility (i.e., how much the explanation seems convincing to the user). Our theoretical framework simplifies how these properties are operationalized, and it provides new insights into common explanation methods that we analyze as case studies. We also discuss the impact that our framework could have in biomedicine, a very sensitive application domain where XAI can have a central role in generating trust.
Matteo Rizzo, Alberto Veneri, Andrea Albarelli, Claudio Lucchese, Marco S. Nobile, Cristina Conati
CIBCB5
2023 Unsupervised neural networks as a support tool for pathology diagnosis in MALDI-MSI experiments: A case study on thyroid biopsies
abstract
Artificial intelligence is getting a foothold in medicine for disease screening and diagnosis. While typical machine learning methods require large labeled datasets for training and validation, their application is limited in clinical fields since ground truth information can hardly be obtained on a sizeable cohort of patients. Unsupervised neural networks – such as Self-Organizing Maps (SOMs) – represent an alternative approach to identifying hidden patterns in biomedical data. Here we investigate the feasibility of SOMs for the identification of malignant and non-malignant regions in liquid biopsies of thyroid nodules, on a patient-specific basis. MALDI-ToF (Matrix Assisted Laser Desorption Ionization - Time of Flight) mass spectrometry-imaging (MSI) was used to measure the spectral profile of bioptic samples. SOMs were then applied for the analysis of MALDI-MSI data of individual patients’ samples, also testing various pre-processing and agglomerative clustering methods to investigate their impact on SOMs’ discrimination efficacy. The final clustering was compared against the sample’s probability to be malignant, hyperplastic or related to Hashimoto thyroiditis as quantified by multinomial regression with LASSO. Our results show that SOMs are effective in separating the areas of a sample containing benign cells from those containing malignant cells. Moreover, they allow to overlap the different areas of cytological glass slides with the corresponding proteomic profile image, and inspect the specific weight of every cellular component in bioptic samples. We envision that this approach could represent an effective means to assist pathologists in diagnostic tasks, avoiding the need to manually annotate cytological images and the effort in creating labeled datasets.
Marco S. Nobile, Giulia Capitoli, Virgil Sowirono, Francesca Clerici, Isabella Piga, Kirsten van Abeelen, Fulvio Magni, Fabio Pagni, Stefania Galimberti, Paolo Cazzaniga, Daniela Besozzi
Expert Syst. Appl.1
2023 Ten quick tips for fuzzy logic modeling of biomedical systems
abstract
Fuzzy logic is useful tool to describe and represent biological or medical scenarios, where often states and outcomes are not only completely true or completely false, but rather partially true or partially false. Despite its usefulness and spread, fuzzy logic modeling might easily be done in the wrong way, especially by beginners and unexperienced researchers, who might overlook some important aspects or might make common mistakes. Malpractices and pitfalls, in turn, can lead to wrong or overoptimistic, inflated results, with negative consequences to the biomedical research community trying to comprehend a particular phenomenon, or even to patients suffering from the investigated disease. To avoid common mistakes, we present here a list of quick tips for fuzzy logic modeling any biomedical scenario: some guidelines which should be taken into account by any fuzzy logic practitioner, including experts. We believe our best practices can have a strong impact in the scientific community, allowing researchers who follow them to obtain better, more reliable results and outcomes in biomedical contexts.
Davide Chicco, Simone Spolaor, Marco S. Nobile
PLoS Comput. Biol.3
2022 The Impact of Variable Selection and Transformation on the Interpretability and Accuracy of Fuzzy Models
abstract
Data transformation is an important step in Machine Learning pipelines which can strongly improve their performance. For instance, min-max normalization is often used to make all variables lie in the same range, while log-transformation is used to map data that is scattered across several orders of magnitude to a logarithmic space. Such transformations can be beneficial when the machine learning approach measures distance in a metric space, such as cluster-based approaches. These two transformation approaches can be combined to reveal hidden patterns in the data in the case of log-normally distributed data points, which commonly occur in biological and medical data. In this work we introduce a novel evolutionary approach designed to automatically determine the optimal log-transformation and selection of variables. Our approach is built around an interpretable AI system (created by pyFUME), so that all transformations are followed by inverse transformations to map back the values into the original universe of discourse, and preserve the interpretability of the results. We test our approach on two synthetic datasets, designed to reproduce a condition in which some variables are normally distributed, some variables are log-normally distributed, and some variables are just noise in the dataset. Our results show that our approach yields better performing models compared to conventional methods, and that the resulting model is also characterised by a better interpretability, making such approach particularly useful to study biomedical datasets.
Caro Fuchs, Simone Spolaor, Uzay Kaymak, Marco S. Nobile
CIBCB4
2022 Predicting and Characterizing Legal Claims of Hospitals with Computational Intelligence: the Legal and Ethical Implications
abstract
In this paper we propose a fuzzy logic-based approach to analyze UK National Health Service (NHS) public administrative data related to pre- and post-pandemic claims filed by patients, analyzing the legal and ethical issues connected to the use of Artificial Intelligence systems, including our own, to take critical decisions having a significant impact on patients, such as employing computational intelligence to justify the management choices related to Intensive Care Unit (ICU) bed allocation. Differently from previous papers, in this work we follow an unsupervised approach and, specifically, we perform an analysis of UK hospitals by means of a computational intelligence algorithm integrating Fuzzy C- Means and swarm intelligence. The dataset that we analyse allows us to compare pre- and post-pandemic data, to analyze the ethical and legal challenges of the use of computational intelligence for critical decision-making in the health care field.
Chiara Gallese, Caro Fuchs, Simone G. Riva, Emanuela Foglia, Fabrizio Schettini, Lucrezia Ferrario, Elena Falletti, Marco S. Nobile
CIBCB8
2022 Comparing Interpretable AI Approaches for the Clinical Environment: an Application to COVID-19
abstract
Machine Learning (ML) models play an important role in healthcare thanks to their remarkable performance in predicting complex phenomena. During the COVID-19 pandemic, different ML models were implemented to support decisions in the medical settings. However, clinical experts need to ensure that these models are valid, provide clinically useful information, and are implemented and used correctly. In this vein, they need to understand the logic behind the models to be able to trust them. Hence, developing transparent and interpretable models has increasing relevance. In this work, we applied four interpretable ML models including logistic regression, decision tree, pyFUME, and RIPPER to classify suspected COVID-19 patients based on clinical data collected from blood samples. After preprocessing the data set and training the models, we evaluate the models based on their predictive performance. Then, we illustrate that interpretability can be achieved in different ways. First, SHAP explanations are built from logistic regression and decision trees to obtain the features' importance. Then, the potential of pyFUME and RIPPER in providing inherent interpretability are reflected. Finally, potential ways to achieve trust in future studies are briefly discussed.
Mohsen Abbaspour Onari, Marco S. Nobile, Isel Grau, Caro Fuchs, Yingqian Zhang 0001, Arjen-Kars Boer, Volkher Scharnhorst
CIBCB2
2022 Local Bubble Dilation Functions: Hypersphere-bounded Landscape Deformations Simplify Global Optimization
abstract
Solving optimization problems is one of the most complex and widespread task in Computer Science. In many scenarios, finding the global optimum of a function is hampered by several features that characterize the fitness landscapes, such as noisiness, multi-modality, non-convexity, non-separability, and non-differentiability. In order to facilitate the optimization process, a variety of methods have been proposed to manipulate either the search space or the fitness landscape. Among these, Dilation Functions (DFs) were introduced to expand regions of the search space that are characterized by promising fitness values. In this work, we extend the family of DFs by introducing Local Bubble Dilation Functions (LBDFs), a novel approach that generates local distortions bounded by hyper-spheres. By performing an appropriate mapping of the search space, LBDFs can improve the optimization performance, since they expand and reveal the promising regions around the global optimum, while leaving the rest of the fitness landscape untouched. The additional advantage of LBDFs, with respect to DFs, is that different dilations can be applied to each dimension of the search space, which is useful in the case of asymmetric landscapes. In order to show the benefits of local dilations, we executed several tests on the Michalewicz benchmark function, with different settings for the LBDFs. Our results show that a properly designed LBDF can lead to statistically significant better results than using vanilla optimization. Finally, we investigated the use of LBDFs to facilitate the solution of the parameter estimation problem in Systems Biology by analyzing the landscape related to a stochastic model of enzyme kinetics.
Daniele M. Papetti, Vasco Coelho, Dan Ashlock, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Marco S. Nobile
CIBCB7
2022 Building Interpretable and Parsimonious Fuzzy Models using a Multi-Objective Approach
abstract
Nowadays, the growing amounts of collected data enable the training of machine learning models that can be used to extract insights from the data and make better-informed decisions. Among the possible models that can be learned from data are fuzzy rule-based models, which are transparent and enable – when properly designed – interpretable artificial intelligence. One of the requirements of interpretability is a simple model structure, which can be achieved by performing feature selection and by limiting the number of rules in the model. However, the chosen feature set and the number of rules may interact and strongly affect the model’s accuracy. In this study, we employ techniques from the field of evolutionary computation to perform feature and rule number selection simultaneously. To ensure the developed models do not only perform well but are also interpretable and have good generalization capabilities, we adopt a multi-objective approach in which we train the models focusing on three objectives: performance, complexity, and model stability. In this way, we strive to develop simple, well-performing parsimonious fuzzy models. We show the effectiveness of our approach on three benchmark data sets.
Caro Fuchs, Uzay Kaymak, Marco S. Nobile
FUZZ-IEEE3
2022 Salp Swarm Optimization: A critical review
Mauro Castelli, Luca Manzoni, Luca Mariot, Marco S. Nobile, Andrea Tangherloni
Expert Syst. Appl.4
2021 If You Can't Beat It, Squash It: Simplify Global Optimization by Evolving Dilation Functions
abstract
Optimization problems represent a class of pervasive and complex tasks in Computer Science, aimed at identifying the global optimum of a given objective function. Optimization problems are typically noisy, multi-modal, non-convex, non-separable, and often non-differentiable. Because of these features, they mandate the use of sophisticated population-based meta-heuristics to effectively explore the search space. Additionally, computational techniques based on the manipulation of the optimization landscape, such as Dilation Functions (DFs), can be effectively exploited to either "compress" or "dilate" some target regions of the search space, in order to improve the exploration and exploitation capabilities of any meta-heuristic. The main limitation of DFs is that they must be tailored on the specific optimization problem under investigation. In this work, we propose a solution to this issue, based on the idea of evolving the DFs. Specifically, we introduce a two-layered evolutionary framework, which combines Evolutionary Computation and Swarm Intelligence to solve the meta-problem of optimizing both the structure and the parameters of DFs. We evolved optimal DFs on a variety of benchmark problems, showing that this approach yields extremely simpler versions of the original optimization problems.
Daniele M. Papetti, Dan Ashlock, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CEC5
2021 The Impact of Representation on the Optimization of Marker Panels for Single-cell RNA Data
abstract
The increasing number of single-cell transcriptomic and single-cell RNA sequencing studies are allowing for a deeper understanding of the molecular processes underlying the normal development of an organism as well as the onset of pathologies. These studies continuously refine the functional roles of known cell populations, and provide their characterization as soon as putatively novel cell populations are detected. In order to isolate the cell populations for further tailored analysis, succinct marker panels—composed of a few cell surface proteins and clusters of differentiation molecules—must be identified. The identification of these marker panels is a challenging computational problem due to its intrinsic combinatorial nature, which makes it an NP-hard problem. Genetic Algorithms (GAs) have been successfully used in Bioinformatics and other biomedical applications to tackle combinatorial problems. We present here a GA-based approach to solve the problem of the identification of succinct marker panels. Since the performance of a GA is strictly related to the representation of the candidate solutions, we propose and compare three alternative representations, able to implicitly introduce different constraints on the search space. For each representation, we perform a fine-tuning of the parameter settings to calibrate the GA, and we show that different representations yield different performance, where the most relaxed representations— in which the GA can also evolve the number of genes in the panel—turn out to be the more effective, especially in the case of 0-knowledge problems. Our results also show that the marker panels identified by GAs can outperform manually curated solutions.
Andrea Tangherloni, Simone G. Riva, Simone Spolaor, Daniela Besozzi, Marco S. Nobile, Paolo Cazzaniga
CEC5
2021 A comparison of multi-objective optimization algorithms to identify drug target combinations
abstract
Combination therapies represent one of the most effective strategy in inducing cancer cell death and reducing the risk to develop drug resistance. The identification of putative novel drug combinations, which typically requires the execution of expensive and time consuming lab experiments, can be supported by the synergistic use of mathematical models and multi-objective optimization algorithms. The computational approach allows to automatically search for potential therapeutic combinations and to test their effectiveness in silico, thus reducing the costs of time and money, and driving the experiments toward the most promising therapies. In this work, we couple dynamic fuzzy modeling of cancer cells with different multi-objective optimization algorithm, and we compare their performance in identifying drug target combinations. Specifically, we perform batches of optimizations with 3 and 4 objective functions defined to achieve a desired behavior of the system (e.g., maximize apop-tosis while minimizing necrosis and survival), and we compare the quality of the solutions included in the Pareto fronts. Our results show that both the choice of the multi-objective algorithm and the formulation of the optimization problem have an impact on the identified solutions, highlighting the strengths as well as the limitations of this approach.
Simone Spolaor, Daniele M. Papetti, Paolo Cazzaniga, Daniela Besozzi, Marco S. Nobile
CIBCB5
2021 Tremor assessment using smartphone sensor data and fuzzy reasoning
abstract
BACKGROUND: Tremor severity assessment is an important step for the diagnosis and treatment decision-making of essential tremor (ET) patients. Traditionally, tremor severity is assessed by using questionnaires (e.g., ETRS and QUEST surveys). In this work we assume the possibility of assessing tremor severity using sensor data and computerized analyses. The goal of this work is to assess severity of tremor objectively, to be better able to asses improvement in ET patients due to deep brain stimulation or other treatments. METHODS: We collect tremor data by strapping smartphones to the wrists of ET patients. The resulting raw sensor data is then pre-processed to remove any artifact due to patient's intentional movement. Finally, this data is exploited to automatically build a transparent, interpretable, and succinct fuzzy model for the severity assessment of ET. For this purpose, we exploit pyFUME, a tool for the data-driven estimation of fuzzy models. It leverages the FST-PSO swarm intelligence meta-heuristic to identify optimal clusters in data, reducing the possibility of a premature convergence in local minima which would result in a sub-optimal model. pyFUME was also combined with GRABS, a novel methodology for the automatic simplification of fuzzy rules. RESULTS: Our model is able to assess tremor severity of patients suffering from Essential Tremor, notably without the need for subjective questionnaires nor interviews. The fuzzy model improves the mean absolute error (MAE) metric by 78-81% compared to linear models and by 71-74% compared to a model based on decision trees. CONCLUSION: This study confirms that tremor data gathered using the smartphones is useful for the constructing of machine learning models that can be used to support the diagnosis and monitoring of patients who suffer from Essential Tremor. The model produced by our methodology is easy to inspect and, notably, characterized by a lower error with respect to approaches based on linear models or decision trees.
Caro Fuchs, Marco S. Nobile, Guillaume Zamora, Aurélie Degeneffe, Pieter Leonard Kubben, Uzay Kaymak
BMC Bioinform.2
2021 Accelerated global sensitivity analysis of genome-wide constraint-based metabolic models
abstract
BACKGROUND: Genome-wide reconstructions of metabolism opened the way to thorough investigations of cell metabolism for health care and industrial purposes. However, the predictions offered by Flux Balance Analysis (FBA) can be strongly affected by the choice of flux boundaries, with particular regard to the flux of reactions that sink nutrients into the system. To mitigate possible errors introduced by a poor selection of such boundaries, a rational approach suggests to focus the modeling efforts on the pivotal ones. METHODS: In this work, we present a methodology for the automatic identification of the key fluxes in genome-wide constraint-based models, by means of variance-based sensitivity analysis. The goal is to identify the parameters for which a small perturbation entails a large variation of the model outcomes, also referred to as sensitive parameters. Due to the high number of FBA simulations that are necessary to assess sensitivity coefficients on genome-wide models, our method exploits a master-slave methodology that distributes the computation on massively multi-core architectures. We performed the following steps: (1) we determined the putative parameterizations of the genome-wide metabolic constraint-based model, using Saltelli's method; (2) we applied FBA to each parameterized model, distributing the massive amount of calculations over multiple nodes by means of MPI; (3) we then recollected and exploited the results of all FBA runs to assess a global sensitivity analysis. RESULTS: We show a proof-of-concept of our approach on latest genome-wide reconstructions of human metabolism Recon2.2 and Recon3D. We report that most sensitive parameters are mainly associated with the intake of essential amino acids in Recon2.2, whereas in Recon 3D they are associated largely with phospholipids. We also illustrate that in most cases there is a significant contribution of higher order effects. CONCLUSION: Our results indicate that interaction effects between different model parameters exist, which should be taken into account especially at the stage of calibration of genome-wide models, supporting the importance of a global strategy of sensitivity analysis.
Marco S. Nobile, Vasco Coelho, Dario Pescini, Chiara Damiani
BMC Bioinform.1
2021 Investigating the performance of multi-objective optimization when learning Bayesian Networks
Marco S. Nobile, Paolo Cazzaniga, Daniele Ramazzotti
Neurocomputing1
2021 FiCoS: A fine-grained and coarse-grained GPU-powered deterministic simulator for biochemical networks
abstract
Mathematical models of biochemical networks can largely facilitate the comprehension of the mechanisms at the basis of cellular processes, as well as the formulation of hypotheses that can be tested by means of targeted laboratory experiments. However, two issues might hamper the achievement of fruitful outcomes. On the one hand, detailed mechanistic models can involve hundreds or thousands of molecular species and their intermediate complexes, as well as hundreds or thousands of chemical reactions, a situation generally occurring in rule-based modeling. On the other hand, the computational analysis of a model typically requires the execution of a large number of simulations for its calibration, or to test the effect of perturbations. As a consequence, the computational capabilities of modern Central Processing Units can be easily overtaken, possibly making the modeling of biochemical networks a worthless or ineffective effort. To the aim of overcoming the limitations of the current state-of-the-art simulation approaches, we present in this paper FiCoS, a novel "black-box" deterministic simulator that effectively realizes both a fine-grained and a coarse-grained parallelization on Graphics Processing Units. In particular, FiCoS exploits two different integration methods, namely, the Dormand-Prince and the Radau IIA, to efficiently solve both non-stiff and stiff systems of coupled Ordinary Differential Equations. We tested the performance of FiCoS against different deterministic simulators, by considering models of increasing size and by running analyses with increasing computational demands. FiCoS was able to dramatically speedup the computations up to 855×, showing to be a promising solution for the simulation and analysis of large-scale models of complex biological processes.
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Giulia Capitoli, Simone Spolaor, Leonardo Rundo, Giancarlo Mauri, Daniela Besozzi
PLoS Comput. Biol.2
2021 A CUDA-powered method for the feature extraction and unsupervised analysis of medical images
abstract
Abstract Image texture extraction and analysis are fundamental steps in computer vision. In particular, considering the biomedical field, quantitative imaging methods are increasingly gaining importance because they convey scientifically and clinically relevant information for prediction, prognosis, and treatment response assessment. In this context, radiomic approaches are fostering large-scale studies that can have a significant impact in the clinical practice. In this work, we present a novel method, called CHASM (Cuda, HAralick & SoM), which is accelerated on the graphics processing unit (GPU) for quantitative imaging analyses based on Haralick features and on the self-organizing map (SOM). The Haralick features extraction step relies upon the gray-level co-occurrence matrix, which is computationally burdensome on medical images characterized by a high bit depth. The downstream analyses exploit the SOM with the goal of identifying the underlying clusters of pixels in an unsupervised manner. CHASM is conceived to leverage the parallel computation capabilities of modern GPUs. Analyzing ovarian cancer computed tomography images, CHASM achieved up to $$\sim 19.5\times $$ ∼ 19.5 × and $$\sim 37\times $$ ∼ 37 × speed-up factors for the Haralick feature extraction and for the SOM execution, respectively, compared to the corresponding C++ coded sequential versions. Such computational results point out the potential of GPUs in the clinical research.
Leonardo Rundo, Andrea Tangherloni, Paolo Cazzaniga, Matteo Mistri, Simone Galimberti, Ramona Woitek, Evis Sala, Giancarlo Mauri, Marco S. Nobile
J. Supercomput.9
2020 Preventing litigation with a predictive model of COVID-19 ICUs occupancy
abstract
The COVID-19 pandemic has generated an overall slowdown in hospital activities that might lead to delays in healthcare interventions, and the scarcity of resources can raise concerns about ventilators allocation criteria. These circumstances could lead to lawsuits against hospitals and healthcare professionals: together with Regions and States, they may be vulnerable to legal actions, due to the breach of right to health, to physical integrity and right to life, to the manifestation of the informed consent in the medical field or on the basis of contractual or Aquilian obligations. In this context, predicting the litigation rate could be useful to assess the economic impact of a dispute at a local and national level, so that hospital managers and public institutions can perform multi-dimensional and cost/benefit evaluations to decide whether to invest resources to increase critical care surge capacity. In this work we present CLIP (COVID-19 LItigation Prediction), a modeling approach supported by swarm intelligence designed to forecast the occupancy of intensive care units using COVID-19 time-series. CLIP fits a logistic model of COVID-19 patients admission in order to estimate the future number of patients, and then exploits a probabilistic model to predict the number of occupied intensive care beds, whose parameters are calibrated by means of Fuzzy Self-Tuning Particle Swarm Optimization. We assume that each individual rejected from an intensive care unit due to the lack of resources should be considered a potential plaintiff. The development and the availability of such a predictive model, that could further be used within other clinical conditions and important diseases, could help policy-makers in taking decisions under conditions of uncertainty.
Chiara Gallese, Elena Falletti, Marco S. Nobile, Lucrezia Ferrario, Fabrizio Schettini, Emanuela Foglia
IEEE BigData3
2020 Which random is the best random? A study on sampling methods in Fourier surrogate modeling
abstract
Global optimization problems can be effectively solved by means of Computational Intelligence methods. However, there are several areas in which the effectiveness of these algorithms can be hampered by the computational costs of the fitness evaluations, or by specific features of the fitness landscape that can be characterized by noise and by the presence of several (even infinite) local optima. These issues bring about the necessity of defining specific techniques to replace the original problem with a surrogate representation. Fourier surrogate modeling represents a novel and effective approach to generate smoother, and possibly easier to explore, fitness landscapes, and to reduce the computational effort. Fourier surrogates require an initial sampling of the search space that must be performed to calculate the Fourier transforms. In this paper we investigate the impact on the quality of the surrogate models of the hyper-parameters of the methodology, and of several methods that can be employed for the initial sampling of the fitness landscape (i.e., pseudorandom numbers, low discrepancy sequences, a logistic map in chaotic regime, true random positions generated by a quantum computer, and point packing). Our results show that semistructured approaches like quasi-random sequences and point packing can outperform the other sampling methods.
Marco S. Nobile, Simone Spolaor, Paolo Cazzaniga, Daniele M. Papetti, Daniela Besozzi, Dan Ashlock, Luca Manzoni
CEC1
2020 Fourier Surrogate Models of Dilated Fitness Landscapes in Systems Biology : or how we learned to torture optimization problems until they confess
abstract
One of the most complex problems in Systems Biology is Parameter estimation (PE), which consists in inferring the kinetic parameters of biochemical systems. The identification of an accurate parameterization, able to reproduce any observed experimental behavior, is fundamental for the definition of predictive models. PE is a non-convex, multi-modal, and non-separable problem that is usually tackled by using Computational Intelligence methods. When the biochemical species appear in the system in a very low amount, the intrinsic noise due to the randomness of molecular collisions cannot be neglected. In this case, stochastic simulation algorithms should be employed to properly reproduce the system dynamics. Stochastic fluctuations make the PE problem even more complicated, as they can lead to radically different values of the fitness function for the same candidate parameterization. In addition, the kinetic parameters generally follow a log-uniform distribution, so that global optima tend to be localized in the lowest orders of magnitude of the search space. To simultaneously tackle all the aforementioned issues, in this work we investigate a novel approach based on the combination of dilation functions with Fourier surrogate modeling and filtering on the fitness landscape. The results show that our approach is able to strongly simplify the PE problem for low-dimensional optimization instances.
Marco S. Nobile, Paolo Cazzaniga, Simone Spolaor, Daniela Besozzi, Luca Manzoni
CIBCB1
2020 pyFUME: a Python Package for Fuzzy Model Estimation
abstract
Living in the era of "data deluge" demands for an increase in the application and development of machine learning methods, both in basic and applied research. Among these methods, in the last decades fuzzy inference systems carved out their own niche as (light) grey box models, which are considered more interpretable and transparent than other commonly employed methods, such as artificial neural networks. Although commercially distributed alternatives are available, software able to assist practitioners and researchers in each step of the estimation of a fuzzy model from data are still limited in scope and applicability. This is especially true when looking at software developed in Python, a programming language that quickly gained popularity among data scientists and it is often considered their language of choice. To fill this gap, we introduce pyFUME, a Python library for automatically estimating fuzzy models from data. pyFUME contains a set of classes and methods to estimate the antecedent sets and the consequent parameters of a Takagi-Sugeno fuzzy model from data, and then create an executable fuzzy model exploiting the Simpful library. pyFUME can be beneficial to practitioners, thanks to its pre-implemented and user-friendly pipelines, but also to researchers that want to fine-tune each step of the estimation process.
Caro Fuchs, Simone Spolaor, Marco S. Nobile, Uzay Kaymak
FUZZ-IEEE3
2020 On the automatic calibration of fully analogical spiking neuromorphic chips
abstract
Nowadays, understanding the topology of biological neural networks and sampling their activity is possible thanks to various laboratory protocols that provide a large amount of experimental data, thus paving the way to accurate modeling and simulation. Neuromorphic systems were developed to simulate the dynamics of biological neural networks by means of electronic circuits, offering an efficient alternative to classic simulations based on systems of differential equations, from both the points of view of the energy consumed and the overall computational effort. Spikey is a configurable neuromorphic chip based on the Leaky Integrate-And-Fire model, which gives the user the possibility to model an arbitrary neural topology and simulate the temporal evolution of membrane potentials. To accurately reproduce the behavior of a specific biological network, a detailed parameterization of all neurons in the neuromorphic chip is necessary. Determining such parameters is a hard, error-prone, and generally time consuming task. In this work, we propose a novel methodology for the automatic calibration of neuromorphic chips that exploits a given neural activity as target. Our results show that, in the case of small networks with a low complexity, the method can estimate a vector of parameters capable of reproducing the target activity. Conversely, in the case of more complex networks, the simulations with Spikey can be highly affected by noise, which causes small variations in the simulations outcome even when identical networks are simulated, hindering the convergence to optimal parameterizations.
Daniele M. Papetti, Simone Spolaor, Daniela Besozzi, Paolo Cazzaniga, Marco Antoniotti, Marco S. Nobile
IJCNN6
2020 A Graph Theory Approach to Fuzzy Rule Base Simplification
Caro Fuchs, Simone Spolaor, Marco S. Nobile, Uzay Kaymak
IPMU (1)3
2020 Fuzzy modeling and global optimization to predict novel therapeutic targets in cancer cells
abstract
MOTIVATION: The elucidation of dysfunctional cellular processes that can induce the onset of a disease is a challenging issue from both the experimental and computational perspectives. Here we introduce a novel computational method based on the coupling between fuzzy logic modeling and a global optimization algorithm, whose aims are to (1) predict the emergent dynamical behaviors of highly heterogeneous systems in unperturbed and perturbed conditions, regardless of the availability of quantitative parameters, and (2) determine a minimal set of system components whose perturbation can lead to a desired system response, therefore facilitating the design of a more appropriate experimental strategy. RESULTS: We applied this method to investigate what drives K-ras-induced cancer cells, displaying the typical Warburg effect, to death or survival upon progressive glucose depletion. The optimization analysis allowed to identify new combinations of stimuli that maximize pro-apoptotic processes. Namely, our results provide different evidences of an important protective role for protein kinase A in cancer cells under several cellular stress conditions mimicking tumor behavior. The predictive power of this method could facilitate the assessment of the response of other complex heterogeneous systems to drugs or mutations in fields as medicine and pharmacology, therefore paving the way for the development of novel therapeutic treatments. AVAILABILITY AND IMPLEMENTATION: The source code of FUMOSO is available under the GPL 2.0 license on GitHub at the following URL: https://github.com/aresio/FUMOSO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco S. Nobile, Giuseppina Votta, Roberta Palorini, Simone Spolaor, Humberto De Vitto, Paolo Cazzaniga, Francesca Ricciardiello, Giancarlo Mauri, Lilia Alberghina, Ferdinando Chiaradonna, Daniela Besozzi
Bioinform.1
2020 Computational Intelligence for Life Sciences
abstract
Computational Intelligence (CI) is a computer science discipline encompassing the theory, design, development and application of biologically and linguistically derived computational paradigms. Traditionally, the main elements of CI are Evolutionary Computation, Swarm Intelligence, Fuzzy Logic, and Neural Networks. CI aims at proposing new algorithms able to solve complex computational problems by taking inspiration from natural phenomena. In an intriguing turn of events, these nature-inspired methods have been widely adopted to investigate a plethora of problems related to nature itself. In this paper we present a variety of CI methods applied to three problems in life sciences, highlighting their effectiveness: we describe how protein folding can be faced by exploiting Genetic Programming, the inference of haplotypes can be tackled using Genetic Algorithms, and the estimation of biochemical kinetic parameters can be performed by means of Swarm Intelligence. We show that CI methods can generate very high quality solutions, providing a sound methodology to solve complex optimization problems in life sciences.
Daniela Besozzi, Luca Manzoni, Marco S. Nobile, Simone Spolaor, Mauro Castelli, Leonardo Vanneschi, Paolo Cazzaniga, Stefano Ruberto, Leonardo Rundo, Andrea Tangherloni
Fundam. Informaticae3
2020 Coupling Mechanistic Approaches and Fuzzy Logic to Model and Simulate Complex Systems
abstract
Several mathematical formalisms can be exploited to model complex systems, in order to capture different features of their dynamic behavior and leverage any available quantitative or qualitative data. Correspondingly, either quantitative models or qualitative models can be defined; bridging the gap between these two worlds would allow us to simultaneously exploit the peculiar advantages provided by each modeling approach. However, to date, the attempts in this direction have been limited to specific fields of research. In this paper, we propose a novel, general-purpose computational framework, named Fuzzy-mechanistic modeling of compleX systems (FuzzX), for the analysis of hybrid models consisting of a quantitative (or mechanistic) module and a qualitative module that can reciprocally control each other's dynamic behavior through a common interface. FuzzX takes advantage of precise quantitative information about the system through the definition and simulation of the mechanistic module. At the same time, it describes the behavior of components and their interactions that are not known in full details, by exploiting fuzzy logic for the definition of the qualitative module. We applied FuzzX for the analysis of a hybrid model of a complex biochemical system, characterized by the presence of positive and negative feedback regulations. We show that FuzzX is able to correctly reproduce known emergent behaviors of this system in normal and perturbed conditions. We envision that FuzzX could be employed to analyze any kind of complex system when quantitative information is limited, as well as to extend existing mechanistic models with fuzzy modules to describe those components and interactions of the system that are not fully characterized.
Simone Spolaor, Marco S. Nobile, Giancarlo Mauri, Paolo Cazzaniga, Daniela Besozzi
IEEE Trans. Fuzzy Syst.2
2020 cuProCell: GPU-Accelerated Analysis of Cell Proliferation With Flow Cytometry Data
abstract
The investigation of cell proliferation can provide useful insights for the comprehension of cancer progression, resistance to chemotherapy and relapse. To this aim, computational methods and experimental measurements based on in vivo label-retaining assays can be coupled to explore the dynamic behavior of tumoral cells. ProCell is a software that exploits flow cytometry data to model and simulate the kinetics of fluorescence loss that is due to stochastic events of cell division. Since the rate of cell division is not known, ProCell embeds a calibration process that might require thousands of stochastic simulations to properly infer the parameterization of cell proliferation models. To mitigate the high computational costs, in this paper we introduce a parallel implementation of ProCell's simulation algorithm, named cuProCell, which leverages Graphics Processing Units (GPUs). Dynamic Parallelism was used to efficiently manage the cell duplication events, in a radically different way with respect to common computing architectures. We present the advantages of cuProCell for the analysis of different models of cell proliferation in Acute Myeloid Leukemia (AML), using data collected from the spleen of human xenografts in mice. We show that, by exploiting GPUs, our method is able to not only automatically infer the models' parameterization, but it is also 237× faster than the sequential implementation. This study highlights the presence of a relevant percentage of quiescent and potentially chemoresistant cells in AML in vivo, and suggests that maintaining a dynamic equilibrium among the different proliferating cell populations might play an important role in disease progression.
Marco S. Nobile, Eric Nisoli, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
IEEE J. Biomed. Health Informatics1
2019 Dilation Functions in Global Optimization
abstract
Complex tasks in Computer Science can be reformulated as optimization problems, in which the global optimum of a given function must be identified. Such problems are typically noisy, multi-modal, non-convex and non-separable, and they require the application of population-based global search metaheuristics to effectively explore the search space. In this work, we address the issue of manipulating the search space of these complex optimization problems to the aim of improving the exploration and exploitation capabilities of metaheuristics. In particular, we show that the implicit assumption in global optimization problems, i.e., that candidate solutions are represented by vectors of values whose meaning has a straightforward interpretation, is not always adequate and that the semantics of parameters can be modified by re-mapping their values in the search space by means of user-defined Dilation Functions. Dilation Functions are general purpose transformations that can be applied to any metaheuristics and optimization problem to "compress" or "dilate" some regions of the search space, allowing to improve the quality of the initial population and the exploitation of promising areas, especially in the case of Swarm Intelligence algorithms. The advantages given by the application of Dilation Functions have been observed by running experiments with Fuzzy Self-Tuning Particle Swarm Optimization and Covariance Matrix Adaptation Evolution Strategies, for the optimization of the Ackley benchmark function and for the parameter estimation of a "synthetic" model of a biochemical system.
Marco S. Nobile, Paolo Cazzaniga, Dan Ashlock
CEC1
2019 ProCell: Investigating cell proliferation with Swarm Intelligence
abstract
Computational methods represent an effective mean for the analysis of complex biological processes, such as cell proliferation, especially when combined to well established experimental protocols. In particular, mathematical modeling coupled with computational intelligence algorithms can be successfully exploited to investigate different aspects of cell population dynamics in the context of tumor growth. To this aim, we defined ProCell, a modeling and simulation framework specifically designed for the investigation of cell proliferation, which makes use of Fuzzy Self-Tuning Particle Swarm Optimization to estimate the unknown parameters of cell population models. ProCell is here applied to the analysis of cell proliferation in acute myeloid leukemia, a hematological malignancy characterized by an inherent intra-tumoral heterogeneity that plays an important role in disease recurrence and resistance to chemotherapy. ProCell allowed to provide new insights on the intricate organization of cells with highly heterogeneous proliferative potential, and to highlight the important role of different cell types in the progression and evolution of the disease. ProCell is available under the GPL 2.0 license on GitHub at https://github.com/aresio/ProCell.
Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Paolo Cazzaniga, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
CIBCB1
2019 A Swarm Intelligence Approach to Avoid Local Optima in Fuzzy C-Means Clustering
abstract
Clustering analysis is an important computational task that has applications in many domains. One of the most popular algorithms to solve the clustering problem is fuzzy c-means, which exploits notions from fuzzy logic to provide a smooth partitioning of the data into classes, allowing the possibility of multiple membership for each data sample. The fuzzy c-means algorithm is based on the optimization of a partitioning function, which minimizes inter-cluster similarity. This optimization problem is known to be NP-hard and it is generally tackled using a hill climbing method, a local optimizer that provides acceptable but sub-optimal solutions, since it is sensitive to initialization and tends to get stuck in local optima. In this work we propose an alternative approach based on the swarm intelligence global optimization method Fuzzy Self-Tuning Particle Swarm Optimization (FST-PSO). We solve the fuzzy clustering task by optimizing fuzzy c-means' partitioning function using FST-PSO. We show that this population-based metaheuristics is more effective than hill climbing, providing high quality solutions with the cost of an additional computational complexity. It is noteworthy that, since this particle swarm optimization algorithm is self-tuning, the user does not have to specify additional hyperparameters for the optimization process.
Caro Fuchs, Simone Spolaor, Marco S. Nobile, Uzay Kaymak
FUZZ-IEEE3
2019 Modeling cell proliferation in human acute myeloid leukemia xenografts
abstract
MOTIVATION: Acute myeloid leukemia (AML) is one of the most common hematological malignancies, characterized by high relapse and mortality rates. The inherent intra-tumor heterogeneity in AML is thought to play an important role in disease recurrence and resistance to chemotherapy. Although experimental protocols for cell proliferation studies are well established and widespread, they are not easily applicable to in vivo contexts, and the analysis of related time-series data is often complex to achieve. To overcome these limitations, model-driven approaches can be exploited to investigate different aspects of cell population dynamics. RESULTS: In this work, we present ProCell, a novel modeling and simulation framework to investigate cell proliferation dynamics that, differently from other approaches, takes into account the inherent stochasticity of cell division events. We apply ProCell to compare different models of cell proliferation in AML, notably leveraging experimental data derived from human xenografts in mice. ProCell is coupled with Fuzzy Self-Tuning Particle Swarm Optimization, a swarm-intelligence settings-free algorithm used to automatically infer the models parameterizations. Our results provide new insights on the intricate organization of AML cells with highly heterogeneous proliferative potential, highlighting the important role played by quiescent cells and proliferating cells characterized by different rates of division in the progression and evolution of the disease, thus hinting at the necessity to further characterize tumor cell subpopulations. AVAILABILITY AND IMPLEMENTATION: The source code of ProCell and the experimental data used in this work are available under the GPL 2.0 license on GITHUB at the following URL: https://github.com/aresio/ProCell. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco S. Nobile, Thalia Vlachou, Simone Spolaor, Daniela Bossi, Paolo Cazzaniga, Luisa Lanfrancone, Giancarlo Mauri, Pier Giuseppe Pelicci, Daniela Besozzi
Bioinform.1
2019 GenHap: a novel computational method based on genetic algorithms for haplotype assembly
abstract
BACKGROUND: In order to fully characterize the genome of an individual, the reconstruction of the two distinct copies of each chromosome, called haplotypes, is essential. The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the two chromosomes. Indeed, the knowledge of complete haplotypes is generally more informative than analyzing single SNPs and plays a fundamental role in many medical applications. RESULTS: To reconstruct the two haplotypes, we addressed the weighted Minimum Error Correction (wMEC) problem, which is a successful approach for haplotype assembly. This NP-hard problem consists in computing the two haplotypes that partition the sequencing reads into two disjoint sub-sets, with the least number of corrections to the SNP values. To this aim, we propose here GenHap, a novel computational method for haplotype assembly based on Genetic Algorithms, yielding optimal solutions by means of a global search process. In order to evaluate the effectiveness of our approach, we run GenHap on two synthetic (yet realistic) datasets, based on the Roche/454 and PacBio RS II sequencing technologies. We compared the performance of GenHap against HapCol, an efficient state-of-the-art algorithm for haplotype phasing. Our results show that GenHap always obtains high accuracy solutions (in terms of haplotype error rate), and is up to 4× faster than HapCol in the case of Roche/454 instances and up to 20× faster when compared on the PacBio RS II dataset. Finally, we assessed the performance of GenHap on two different real datasets. CONCLUSIONS: Future-generation sequencing technologies, producing longer reads with higher coverage, can highly benefit from GenHap, thanks to its capability of efficiently solving large instances of the haplotype assembly problem. Moreover, the optimization approach proposed in GenHap can be extended to the study of allele-specific genomic features, such as expression, methylation and chromatin conformation, by exploiting multi-objective optimization techniques. The source code and the full documentation are available at the following GitHub repository: https://github.com/andrea-tango/GenHap .
Andrea Tangherloni, Simone Spolaor, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Pietro Liò, Ivan Merelli, Daniela Besozzi
BMC Bioinform.4
2019 MedGA: A novel evolutionary method for image enhancement in medical imaging systems
Leonardo Rundo, Andrea Tangherloni, Marco S. Nobile, Carmelo Militello, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
Expert Syst. Appl.3
2019 USE-Net: Incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets
Leonardo Rundo, Yudai Nagano, Ryuichiro Hataya, Carmelo Militello, Andrea Tangherloni, Marco S. Nobile, Claudio Ferretti, Daniela Besozzi, Maria Carla Gilardi, Salvatore Vitabile, Giancarlo Mauri, Hideki Nakayama, Paolo Cazzaniga
Neurocomputing8
2019 ginSODA: massive parallel integration of stiff ODE systems on GPUs
Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.1
2018 Computational Intelligence for Parameter Estimation of Biochemical Systems
abstract
In the field of Systems Biology, simulating the dynamics of biochemical models represents one of the most effective methodologies to understand the functioning of cellular processes in normal or altered conditions. However, the lack of kinetic rates, necessary to perform accurate simulations, strongly limits the scope of these analyses. Parameter Estimation (PE), which consists in identifying a proper model parameterization, is a non-linear, non-convex and multi-modal optimization problem, typically tackled by means of Computational Intelligence techniques, such as Evolutionary Computation and Swarm Intelligence. In this work, we perform a thorough investigation of the most widespread methods for PE-namely, Artificial Bee Colony (ABC), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Differential Evolution (DE), Estimation of Distribution Algorithm (EDA), Genetic Algorithms (GAs), Particle Swarm Optimization (PSO), and Fuzzy Self-Tuning PSO (FST-PSO)-comparing their performances on a set of synthetic (yet realistic) biochemical models of increasing size and complexity. Our results show that a variant of the settings-free FST-PSO algorithm can consistently outperform all other methods; ABC and GAs represent the most performing alternatives, while methods based on multivariate normal distributions (e.g., CMA-ES, EDA) struggle to keep pace with the other approaches.
Marco S. Nobile, Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
CEC1
2018 GPU-Powered Multi-Swarm Parameter Estimation of Biological Systems: A Master-Slave Approach
abstract
In silico investigation of biological systems requires the knowledge of numerical parameters that cannot be easily measured in laboratory experiments, leading to the Parameter Estimation (PE) problem, in which the unknown parameters are automatically inferred by means of optimization algorithms exploiting the available experimental data. Here we present MS 2 PSO, an efficient parallel and distributed implementation of a PE method based on Particle Swarm Optimization (PSO) for the estimation of reaction constants in mathematical models of biological systems, considering as target for the estimation a set of discrete-time measurements of molecular species amounts. In particular, such PE method accounts for the availability of experimental data typically measured under different experimental conditions, by considering a multi-swarm PSO in which the best particles of the swarms can migrate. This strategy allows to infer a common set of reaction constants that simultaneously fits all target data used in the PE. To the aim of efficiently tackling the PE problem, MS 2 PSO embeds the execution of cupSODA, a deterministic simulator that relies on Graphics Processing Units to achieve a massive parallelization of the simulations required in the fitness evaluation of particles. In addition, a further level of parallelism is realized by exploiting the Master-Slave distributed programming paradigm. We apply MS 2 PSO for the PE of synthetic biochemical models with 10, 20 and 30 parameters to be estimated, and compare the performances obtained with different GPUs and different configurations (i.e., numbers of processes) of the Master-Slave.
Andrea Tangherloni, Leonardo Rundo, Simone Spolaor, Paolo Cazzaniga, Marco S. Nobile
PDP5
2017 Proactive Particles in Swarm Optimization: A settings-free algorithm for real-parameter single objective optimization problems
abstract
Particle Swarm Optimization (PSO) is an effective Swarm Intelligence technique for the optimization of non-linear and complex high-dimensional problems. Since PSO's performance is strongly dependent on the choice of its functioning settings, in this work we consider a self-tuning version of PSO, called Proactive Particles in Swarm Optimization (PPSO). PPSO leverages Fuzzy Logic to dynamically determine the best settings for the inertia weight, cognitive factor and social factor. The PPSO algorithm significantly differs from other versions of PSO relying on Fuzzy Logic, because specific settings are assigned to each particle according to its history, instead of being globally assigned to the whole swarm. In such a way, PPSO's particles gain a limited autonomous and proactive intelligence with respect to the reactive agents proposed by PSO. Our results show that PPSO achieves overall good optimization performances on the benchmark functions proposed in the CEC 2017 test suite, with the exception of those based on the Schwefel function, whose fitness landscape seems to mislead the fuzzy reasoning. Moreover, with many benchmark functions, PPSO is characterized by a higher speed of convergence than PSO in the case of high-dimensional problems.
Andrea Tangherloni, Leonardo Rundo, Marco S. Nobile
CEC3
2017 Reboot strategies in particle swarm optimization and their impact on parameter estimation of biochemical systems
abstract
Computational methods adopted in the field of Systems Biology require the complete knowledge of reaction kinetic constants to perform simulations of the dynamics and understand the emergent behavior of biochemical systems. However, kinetic parameters of biochemical reactions are often difficult or impossible to measure, thus they are generally inferred from experimental data, in a process known as Parameter Estimation (PE). We consider here a PE methodology that exploits Particle Swarm Optimization (PSO) to estimate an appropriate kinetic parameterization, by comparing experimental time-series target data with in silica dynamics, simulated by using the parameterization encoded by each particle. In this work we present three different reboot strategies for PSO, whose aim is to reinitialize particle positions to avoid particles to get trapped in local optima, and we compare the performance of PSO coupled with the reboot strategies with respect to standard PSO in the case of the PE of two biochemical systems. Since the PE requires a huge number of simulations at each iteration, in this work we exploit a GPU-powered deterministic simulator, cupSODA, which performs in a parallel fashion all simulations and fitness evaluations. Finally, we show that the performances of our implementation scale sublinearly with respect to the swarm size, even on outdated GPUs.
Simone Spolaor, Andrea Tangherloni, Leonardo Rundo, Marco S. Nobile, Paolo Cazzaniga
CIBCB4
2017 Graphics processing units in bioinformatics, computational biology and systems biology
abstract
Several studies in Bioinformatics, Computational Biology and Systems Biology rely on the definition of physico-chemical or mathematical models of biological systems at different scales and levels of complexity, ranging from the interaction of atoms in single molecules up to genome-wide interaction networks. Traditional computational methods and software tools developed in these research fields share a common trait: they can be computationally demanding on Central Processing Units (CPUs), therefore limiting their applicability in many circumstances. To overcome this issue, general-purpose Graphics Processing Units (GPUs) are gaining an increasing attention by the scientific community, as they can considerably reduce the running time required by standard CPU-based software, and allow more intensive investigations of biological systems. In this review, we present a collection of GPU tools recently developed to perform computational analyses in life science disciplines, emphasizing the advantages and the drawbacks in the use of these parallel architectures. The complete list of GPU-powered tools here reviewed is available at http://bit.ly/gputools.
Marco S. Nobile, Paolo Cazzaniga, Andrea Tangherloni, Daniela Besozzi
Briefings Bioinform.1
2017 GPU-powered model analysis with PySB/cupSODA
abstract
SUMMARY: A major barrier to the practical utilization of large, complex models of biochemical systems is the lack of open-source computational tools to evaluate model behaviors over high-dimensional parameter spaces. This is due to the high computational expense of performing thousands to millions of model simulations required for statistical analysis. To address this need, we have implemented a user-friendly interface between cupSODA, a GPU-powered kinetic simulator, and PySB, a Python-based modeling and simulation framework. For three example models of varying size, we show that for large numbers of simulations PySB/cupSODA achieves order-of-magnitude speedups relative to a CPU-based ordinary differential equation integrator. AVAILABILITY AND IMPLEMENTATION: The PySB/cupSODA interface has been integrated into the PySB modeling framework (version 1.4.0), which can be installed from the Python Package Index (PyPI) using a Python package manager such as pip. cupSODA source code and precompiled binaries (Linux, Mac OS/X, Windows) are available at github.com/aresio/cupSODA (requires an Nvidia GPU; developer.nvidia.com/cuda-gpus). Additional information about PySB is available at pysb.org. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leonard A. Harris, Marco S. Nobile, James C. Pino, Alexander L. R. Lubbock, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga, Carlos F. Lopez
Bioinform.2
2017 LASSIE: simulating large-scale models of biochemical systems on GPUs
abstract
BACKGROUND: Mathematical modeling and in silico analysis are widely acknowledged as complementary tools to biological laboratory methods, to achieve a thorough understanding of emergent behaviors of cellular processes in both physiological and perturbed conditions. Though, the simulation of large-scale models-consisting in hundreds or thousands of reactions and molecular species-can rapidly overtake the capabilities of Central Processing Units (CPUs). The purpose of this work is to exploit alternative high-performance computing solutions, such as Graphics Processing Units (GPUs), to allow the investigation of these models at reduced computational costs. RESULTS: LASSIE is a "black-box" GPU-accelerated deterministic simulator, specifically designed for large-scale models and not requiring any expertise in mathematical modeling, simulation algorithms or GPU programming. Given a reaction-based model of a cellular process, LASSIE automatically generates the corresponding system of Ordinary Differential Equations (ODEs), assuming mass-action kinetics. The numerical solution of the ODEs is obtained by automatically switching between the Runge-Kutta-Fehlberg method in the absence of stiffness, and the Backward Differentiation Formulae of first order in presence of stiffness. The computational performance of LASSIE are assessed using a set of randomly generated synthetic reaction-based models of increasing size, ranging from 64 to 8192 reactions and species, and compared to a CPU-implementation of the LSODA numerical integration algorithm. CONCLUSIONS: LASSIE adopts a novel fine-grained parallelization strategy to distribute on the GPU cores all the calculations required to solve the system of ODEs. By virtue of this implementation, LASSIE achieves up to 92× speed-up with respect to LSODA, therefore reducing the running time from approximately 1 month down to 8 h to simulate models consisting in, for instance, four thousands of reactions and species. Notably, thanks to its smaller memory footprint, LASSIE is able to perform fast simulations of even larger models, whereby the tested CPU-implementation of LSODA failed to reach termination. LASSIE is therefore expected to make an important breakthrough in Systems Biology applications, for the execution of faster and in-depth computational analyses of large-scale models of complex biological systems.
Andrea Tangherloni, Marco S. Nobile, Daniela Besozzi, Giancarlo Mauri, Paolo Cazzaniga
BMC Bioinform.2
2017 Efficient Simulation of Reaction Systems on Graphics Processing Units
abstract
Reaction systems represent a theoretical framework based on the regulation mechanisms of facilitation and inhibition of biochemical reactions. The dynamic process defined by a reaction system is typically derived by hand, starting from the set of reactions and a given context sequence. However, thi s procedure may be error-prone and time-consuming, especially when the size of the reaction system increases. Here we present HERESY, a simulator of reaction systems accelerated on Graphics Processing Units (GPUs). HERESY is based on a fine-grained parallelization strategy, whereby all reactions are simultaneously executed on the GPU, therefore reducing the overall running time of the simulation. HERESY is particularly advantageous for the simulation of large-scale reaction systems, consisting of hundreds or thousands of reactions. By considering as test case some reaction systems with an increasing number of reactions and entities, as well as an increasing number of entities per reaction, we show that HERESY allows up to 29× speed-up with respect to a CPU-based simulator of reaction systems. Finally, we provide some directions for the optimization of HERESY, considering minimal reaction systems in normal form.
Marco S. Nobile, Antonio E. Porreca, Simone Spolaor, Luca Manzoni, Paolo Cazzaniga, Giancarlo Mauri, Daniela Besozzi
Fundam. Informaticae1
2017 Gillespie's Stochastic Simulation Algorithm on MIC coprocessors
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.2
2016 GPU-powered and settings-free parameter estimation of biochemical systems
abstract
To understand the emergent behavior of biochemical systems, computational analyses generally require the inference of unknown reaction kinetic constants, a problem known as parameter estimation (PE). In this work we propose a PE methodology that exploits Particle Swarm Optimization (PSO) to examine a set of candidate kinetic parameterizations, whose fitness is evaluated by comparing given target time-series of experimental data with in silico dynamics, simulated by using the parameterization encoded by each particle. In particular, we consider a Fuzzy Logic-based version of PSO - called Proactive Particles in Swarm Optimization (PPSO) - that automatically tunes the setting (inertia, cognitive and social factors) of each particle, independently from all other particles in the swarm. Since the optimization phase requires a large number of simulations for each particle at each iteration, we exploit a GPU-accelerated deterministic simulator, called cupSODA, that automatically generates the system of Ordinary Differential Equations associated with the biochemical system and performs its simulation for each candidate parameterization. We compare the performance of PPSO with respect to PSO for the PE problem by considering two biochemical systems as test cases. In addition, we evaluate the impact on PE of different strategies adopted, both in PPSO and PSO, for the selection of the initial positions of particles within the search space. We prove the effectiveness of our settings-free PE methodology by showing that PPSO outperforms PSO with respect to the computational time required to execute the optimization, achieving comparable results concerning the fitness of the best parameterization found.
Marco S. Nobile, Andrea Tangherloni, Daniela Besozzi, Paolo Cazzaniga
CEC1
2016 Parallel implementation of efficient search schemes for the inference of cancer progression models
abstract
The emergence and development of cancer is a consequence of the accumulation over time of genomic mutations involving a specific set of genes, which provides the cancer clones with a functional selective advantage. In this work, we model the order of accumulation of such mutations during the progression, which eventually leads to the disease, by means of probabilistic graphic models, i.e., Bayesian Networks (BNs). We investigate how to perform the task of learning the structure of such BNs, according to experimental evidence, adopting a global optimization meta-heuristics. In particular, in this work we rely on Genetic Algorithms, and to strongly reduce the execution time of the inference-which can also involve multiple repetitions to collect statistically significant assessments of the data-we distribute the calculations using both multi-threading and a multi-node architecture. The results show that our approach is characterized by good accuracy and specificity; we also demonstrate its feasibility, thanks to a 84× reduction of the overall execution time with respect to a traditional sequential implementation.
Daniele Ramazzotti, Marco S. Nobile, Paolo Cazzaniga, Giancarlo Mauri, Marco Antoniotti
CIBCB2
2016 GPU-powered Bat Algorithm for the parameter estimation of biochemical kinetic values
abstract
The emergent behavior of biochemical systems can be investigated by means of mathematical modeling and computational analyses, which usually require the automatic inference of the unknown values of the model's parameters. This problem, known as Parameter Estimation (PE), is usually tackled with bio-inspired meta-heuristics for global optimization, most notably Particle Swarm Optimization (PSO). In this work we assess the performances of PSO and Bat Algorithm with differential operator and Lévy flights trajectories (DLBA). In particular, we compared these meta-heuristics for the PE using two biochemical models: the expression of genes in prokaryotes and the heat shock response in eukaryotes. In our tests, we also evaluated the impact on PE of different strategies for the initial positioning of individuals within the search space. Our results show that DLBA achieves comparable results with respect to PSO, but it converges to better results when a uniform initialization is employed. Since every iteration of DLBA requires three fitness evaluations for each bat, the whole methodology is built around a GPU-powered biochemical simulator (cupSODA) which is able to parallelize the process. We show that the acceleration achieved with cupSODA strongly reduces the running time, with an empirical 61× speedup that has been obtained comparing a Nvidia GeForce Titan GTX with respect to a CPU Intel Core i7-4790K. Moreover, we show that DLBA always outperforms PSO with respect to the computational time required to execute the optimization process.
Andrea Tangherloni, Marco S. Nobile, Paolo Cazzaniga
CIBCB2
2015 A double swarm methodology for parameter estimation in oscillating Gene Regulatory Networks
abstract
S-systems are mathematical models based on the power-law formalism, which are widely employed for the investigation of Gene Regulatory Networks (GRNs). Because of their complex dynamics - characterized by multi-modality and nonlinearity-the parameterization of S-systems is far from straightforward, demanding global optimization techniques. The problem of parameter estimation of S-systems is further complicated when the desired dynamics is characterized by oscillations. In this work, we describe a novel methodology based on Particle Swarm Optimization for the automatic parameterization of oscillating Ssystems. In this methodology, two swarms perform independent optimizations, and cooperate by periodically exchanging the best particles. The two swarms exploit two different fitness functions: a traditional point-to-point distance, and a spectra-based fitness function. We show that this cooperative approach allows the double swarm to outperform the common methodology, based on a single swarm exploiting a single fitness function. We demonstrate the effectiveness of our method using a GRN of five genes, performing tests of increasing complexity, up to the simultaneous inference of 17 parameters.
Marco S. Nobile, Hitoshi Iba
CEC1
2015 The impact of particles initialization in PSO: Parameter estimation as a case in point
abstract
Despite the intense research focused on the investigation of the functioning settings of Particle Swarm Optimization, the particles initialization functions - determining the initial positions in the search space - are generally ignored, especially in the case of real-world applications. As a matter of fact, almost all works exploit uniform distributions to randomly generate the particles coordinates. In this article, we analyze the impact on the optimization performances of alternative initialization functions based on logarithmic, normal, and lognormal distributions. Our results show how different initialization strategies can affect - and in some cases largely improve - the convergence speed, both in the case of benchmark functions and in the optimization of the kinetic constants of biochemical systems.
Paolo Cazzaniga, Marco S. Nobile, Daniela Besozzi
CIBCB2
2015 Proactive Particles in Swarm Optimization: A self-tuning algorithm based on Fuzzy Logic
abstract
Among the existing global optimization algorithms, Particle Swarm Optimization (PSO) is one of the most effective when dealing with non-linear and complex high-dimensional problems. However, the performance of PSO is strongly dependent on the choice of its settings. In this work we propose a novel and self-tuning PSO algorithm - called Proactive Particles in Swarm Optimization (PPSO) - which exploits Fuzzy Logic to calculate the best setting for the inertia, cognitive factor and social factor. Thanks to additional heuristics, PPSO automatically determines also the best setting for the swarm size and for the particles maximum velocity. PPSO significantly differs from other versions of PSO that exploit Fuzzy Logic, since specific settings are assigned to each particle according to its history, instead of being globally defined for the whole swarm. Thus, the novelty of PPSO is that particles gain a limited autonomous and proactive intelligence, instead of being simple reactive agents. Our results show that PPSO outperforms the standard PSO, both in terms of convergence speed and average quality of solutions, remarkably without the need for any user setting.
Marco S. Nobile, Gabriella Pasi, Paolo Cazzaniga, Daniela Besozzi, Riccardo Colombo, Giancarlo Mauri
FUZZ-IEEE1
2014 A memetic hybrid method for the Molecular Distance Geometry Problem with incomplete information
abstract
The definition of computational methodologies for the inference of molecular structural information plays a relevant role in disciplines as drug discovery and metabolic engineering, since the functionality of a biochemical molecule is determined by its three-dimensional structure. In this work, we present an automatic methodology to solve the Molecular Distance Geometry Problem, that is, to determine the best three-dimensional shape that satisfies a given set of target inter-atomic distances. In particular, our method is designed to cope with incomplete distance information derived from Nuclear Magnetic Resonance measurements. To tackle this problem, that is known to be NP-hard, we present a memetic method that combines two soft-computing algorithms - Particle Swarm Optimization and Genetic Algorithms - with a local search approach, to improve the effectiveness of the crossover mechanism. We show the validity of our method on a set of reference molecules with a length ranging from 402 to 1003 atoms.
Marco S. Nobile, Andrea G. Citrolo, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
IEEE Congress on Evolutionary Computation1
2014 Simulation and Analysis of the Blood Coagulation Cascade Accelerated on GPU
abstract
The use of Graphics Processing Units (GPUs) has recently witnessed ever growing applications for different computational analyses in the field of Life Sciences. In this work we present a CUDA-powered computational tool, named coagSODA, that was purposely developed and applied for the analysis of a large model of the blood coagulation cascade defined as a system of ordinary differential equations, based on both mass-action kinetics and Hill functions. We discuss the biological results of the parameter sweep analyses of this model, and show that GPUs can boost the computational performances up to 177x speedup.
Matteo Bellini, Daniela Besozzi, Paolo Cazzaniga, Giancarlo Mauri, Marco S. Nobile
PDP5
2014 GPU-accelerated simulations of mass-action kinetics models with cupSODA
Marco S. Nobile, Paolo Cazzaniga, Daniela Besozzi, Giancarlo Mauri
J. Supercomput.1
2013 Reverse engineering of kinetic reaction networks by means of Cartesian Genetic Programming and Particle Swarm Optimization
abstract
The modeling of biochemical reaction networks is a fundamental but complex task in Systems Biology, which is traditionally performed exploiting human expertise and the available experimental data. Because of the general lack of knowledge on the molecular mechanisms occurring in living cells, an intense research activity focused on the development of reverse engineering methodologies is currently underway. This problem is further complicated by the fact that a proper parameterization needs to be associated to the reaction network, in order to investigate its dynamical behavior. In this work we propose a novel computational methodology for the reverse engineering of fully parameterized kinetic networks, based on the combined use of two evolutionary programming techniques: Cartesian Genetic Programming (CGP) and Particle Swarm Optimization (PSO). In particular, CGP is used to infer the network topology, while PSO performs the parameter estimation task. To the purpose of applying our methodology in routine laboratory environments, we designed it to exploit a small set of experimental time series as target. We show that our methodology is able to reconstruct kinetic networks that perfectly fit with the target data.
Marco S. Nobile, Daniela Besozzi, Paolo Cazzaniga, Dario Pescini, Giancarlo Mauri
IEEE Congress on Evolutionary Computation1