Paola Stolfi

dblp:258/5049 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
6since 2021 · last 2023
0000-0003-3688-5464ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 7 first-author · 6 since 2021
YearPublicationVenuePosition
2023 NIAPU: network-informed adaptive positive-unlabeled learning for disease gene identification
abstract
MOTIVATION: Gene-disease associations are fundamental for understanding disease etiology and developing effective interventions and treatments. Identifying genes not yet associated with a disease due to a lack of studies is a challenging task in which prioritization based on prior knowledge is an important element. The computational search for new candidate disease genes may be eased by positive-unlabeled learning, the machine learning (ML) setting in which only a subset of instances are labeled as positive while the rest of the dataset is unlabeled. In this work, we propose a set of effective network-based features to be used in a novel Markov diffusion-based multi-class labeling strategy for putative disease gene discovery. RESULTS: The performances of the new labeling algorithm and the effectiveness of the proposed features have been tested on 10 different disease datasets using three ML algorithms. The new features have been compared against classical topological and functional/ontological features and a set of network- and biological-derived features already used in gene discovery tasks. The predictive power of the integrated methodology in searching for new disease genes has been found to be competitive against state-of-the-art algorithms. AVAILABILITY AND IMPLEMENTATION: The source code of NIAPU can be accessed at https://github.com/AndMastro/NIAPU. The source data used in this study are available online on the respective websites. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Paola Stolfi, Andrea Mastropietro, Giuseppe Pasculli, Paolo Tieri, Davide Vergni
Bioinform.1
2022 DruSiLa: an integrated, in-silico disease similarity-based approach for drug repurposing
abstract
The importance of faster drug development has never been more evident than in present time when the whole world is struggling to cope up with the COVID-19 pandemic. At times when timely development of effective drugs and treatment plans could potentially save millions of lives, drug repurposing is one area of medicine that has garnered much of research interest. Apart from experimental drug repurposing studies that happen within wet labs, lot many new quantitative methods have been proposed in the literature. In this paper, one such quantitative methods for drug repurposing is implemented and evaluated. DruSiLa (DRUg in-SIlico LAboratory) is an in-silico drug repurposing method that leverages disease similarity measures to quantitatively rank existing drugs for their potential therapeutic efficacy against novel diseases. The proposed method makes use of available, manually curated, and open datasets on diseases, their genetic origins, and disease-related patho-phenotypes. DruSiLa evaluates pairwise disease similarity scores of any given target disease to each known disease in our dataset. Such similarity scores are then propagated through disease-drug associations, and aggregated at drug nodes to rank them for their predicted effectiveness against the target disease.
Pratuat Amatya, Paola Stolfi, Flavio Lombardi, Paolo Tieri
BIBM2
2022 An agent-based multi-level model to study the spread of antimicrobial-resistant gonorrhoea
abstract
Antimicrobial resistance (AMR) is a major public health problem of the 21st century. The ability of some bacteria to develop resistance to specific antibiotics is the cause of an increased morbidity, mortality and health expenditure. Several surveillance programms have been introduced in the last decades to monitor the spread of antimicrobial resistance. The present work has been conducted within the JPIAMR-project MAGIcIAN whose aim is to support the sustainable introduction of novel class and last-resort antimicrobial drugs minimising the emergence of AMR. Within this project we developed a multi-level model to describe the spread of the sexually transmitted disease of gonorrhoea, caused by the Neisseria gonorrhoeae bacterium, a multidrug resistant bacteria who has progressively developed resistance to many treatment options. The multi-level model includes a dynamic sexual contact network, that describes the dynamic of sexual partnerships, a transmission model that describes the probability of infection during intercourse, and a within-host model, that describes the dynamic of gonorrhoea infection within an individual. The novelty of the proposed model is in including communities having different sexual orientations and behaviour and the possibility of these communities to interact in a dynamic framework. In this work, we calibrate the model using data coming from several clinics located in Amsterdam.
Paola Stolfi, Davide Vergni, Rik Oldenkamp, Constance Schultsz, Emiliano Mancini, Filippo Castiglione
BIBM1
2021 A data-driven model for the generation of Virtual Cohorts
abstract
In silico trials are emerging as a valuable tool for improving both study design and outcomes. A key component of this process is the definition of a virtual cohort, i.e., a set of virtual patients with plausible physiological characteristics (covariates). Building on the NHANES study (2017-2020), we developed a statistical model to infer immunological parameters and a technique to generate a population of plausible immunological virtual patients. A thorough statistical analysis showed that the most appropriate model to represent our data is a conditional multivariate model. Compared to others, it is able to reproduce asymmetric distributions more accurately and is therefore more suitable in cases where there is no prior knowledge of the relationships between covariates. Our analysis also demonstrates the inter-variability and inter-dependence of the different covariates of interest. For example, age has a negative impact on the number of lymphocytes and, surprisingly, ethnicity has a minor influence on the other immunological covariates.
Enrico Mastrostefano, Paola Stolfi, Filippo Castiglione
BIBM2
2021 A functional data analysis approach to assess the prognostic value of SARS-CoV-2 infections surrogate data
abstract
COVID-19 is characterised by quite diverse prognosis. While the majority of infected individuals present no or very mild symptoms, some individuals develop severe disease requiring intensive care. This work leverages the parameters of a virtual cohort of infected individuals generated by a computational immunology model. In so doing we identify the most relevant immunological parameters for the classification of severe COVID-19 cases. The functional data analysis approach used turns out to be appropriate to analyse the output of the computational model. In this work, we classify the disease prognosis using both statistical models and machine learning algorithms adapted from functional data analysis and we compare their performances.
Paola Stolfi, Filippo Castiglione
BIBM1
2021 Emulating complex simulations by machine learning methods
abstract
BACKGROUND: The aim of the present paper is to construct an emulator of a complex biological system simulator using a machine learning approach. More specifically, the simulator is a patient-specific model that integrates metabolic, nutritional, and lifestyle data to predict the metabolic and inflammatory processes underlying the development of type-2 diabetes in absence of familiarity. Given the very high incidence of type-2 diabetes, the implementation of this predictive model on mobile devices could provide a useful instrument to assess the risk of the disease for aware individuals. The high computational cost of the developed model, being a mixture of agent-based and ordinary differential equations and providing a dynamic multivariate output, makes the simulator executable only on powerful workstations but not on mobile devices. Hence the need to implement an emulator with a reduced computational cost that can be executed on mobile devices to provide real-time self-monitoring. RESULTS: Similarly to our previous work, we propose an emulator based on a machine learning algorithm but here we consider a different approach which turn out to have better performances, indeed in terms of root mean square error we have an improvement of two order magnitude. We tested the proposed emulator on samples containing different number of simulated trajectories, and it turned out that the fitted trajectories are able to predict with high accuracy the entire dynamics of the simulator output variables. We apply the emulator to control the level of inflammation while leveraging on the nutritional input. CONCLUSION: The proposed emulator can be implemented and executed on mobile health devices to perform quick-and-easy self-monitoring assessments.
Paola Stolfi, Filippo Castiglione
BMC Bioinform.1
2020 Emulation of dynamic multi-output simulator of risk of type-2 diabetes
abstract
We have recently developed and validated multilevel patient-specific model able to integrate metabolic, nutritional and lifestyle data for the prediction of the metabolic and inflammatory processes underlying the development of type-2 diabetes in the absence of familiarity. Given the incidence of type-2 diabete, which accounts for 85-90% of all cases of diabetes in the world, the implementation of this predictive model on mobile devices could provide a useful instrument to assess the risk of type-2 diabete by informed and aware individuals. However, give the high computational cost of this model, being a mixture of agent-based and ordinary differential equations and providing a dynamic multivariate output, it can run only on powerful workstations but not on mobile devices. The aim of the present paper it to construct an emulator model, using a machine learning approach, with reduced computational cost so to run on mobile devices to provide real time self-monitoring.
Paola Stolfi, Filippo Castiglione
BIBM1
2020 Potential predictors of type-2 diabetes risk: machine learning, synthetic data and wearable health devices
abstract
BACKGROUND: The aim of a recent research project was the investigation of the mechanisms involved in the onset of type 2 diabetes in the absence of familiarity. This has led to the development of a computational model that recapitulates the aetiology of the disease and simulates the immunological and metabolic alterations linked to type-2 diabetes subjected to clinical, physiological, and behavioural features of prototypical human individuals. RESULTS: We analysed the time course of 46,170 virtual subjects, experiencing different lifestyle conditions. We then set up a statistical model able to recapitulate the simulated outcomes. CONCLUSIONS: The resulting machine learning model adequately predicts the synthetic dataset and can, therefore, be used as a computationally-cheaper version of the detailed mathematical model, ready to be implemented on mobile devices to allow self-assessment by informed and aware individuals. The computational model used to generate the dataset of this work is available as a web-service at the following address: http://kraken.iac.rm.cnr.it/T2DM .
Paola Stolfi, Ilaria Valentini, Maria Concetta Palumbo, Paolo Tieri, Andrea Grignolio, Filippo Castiglione
BMC Bioinform.1
2019 Potential predictors of type-2 diabetes risk: machine learning, synthetic data and wearable health devices
abstract
Investigation about the mechanisms involved in the onset of type 2 diabetes in absence of familiarity is the focus of a research project which has led to the development of a computational model that recapitulates the aetiology of the disease. The model simulates the metabolic and immunological alterations related to type-2 diabetes associated to several clinical, physiological and behavioural characteristics of representative virtual patients. In this study, the results of 46170 simulations corresponding to the same number of virtual subjects, experiencing different lifestyle conditions, are analysed for the construction of a statistical model able to recapitulate the simulated dynamics. The resulting machine learning model adequately predicts the synthetic data and can therefore be used as a computationally-cheaper version of the detailed mathematical model, ready to be implemented on mobile devices to allow self assessment by informed and aware individuals.
Paola Stolfi, Ilaria Valentini, Maria Concetta Palumbo, Paolo Tieri, Andrea Grignolio, Filippo Castiglione
BIBM1