Erica Tavazzi

dblp:253/8563 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0001-6188-6413ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Robust end-to-end stratification of amyotrophic lateral sclerosis patients via recurrent variational autoencoder and consensus clustering
abstract
OBJECTIVE: This study aims to develop a data-driven methodology for stratifying Amyotrophic Lateral Sclerosis (ALS) patients based on longitudinal disease progression patterns, using a novel deep learning framework that combines a Recurrent Variational Autoencoder (RVA) with consensus clustering to identify clinically meaningful subgroups. METHODS: The RVA integrates Peephole Long Short-Term Memory networks within the Variational Deep Embedding (VaDE) architecture to simultaneously learn latent representations and cluster assignments from multivariate time-series data. The approach incorporates hyperparameter optimization via prediction strength with two-fold cross-validation, consensus clustering, and internal validation metrics (Silhouette Coefficient, Davies-Bouldin index, Calinski-Harabasz index) for optimal cluster selection. The methodology was validated on simulated data and applied to 3076 ALS patients from the PRO-ACT dataset, using ALSFRS-R total scores, domain subscores, and MiToS staging from the first six months of observation. RESULTS: Simulation experiments demonstrated that consensus clustering consistently outperformed single-model predictions across all noise levels. Applied to the PRO-ACT real data, the framework identified five distinct patient subgroups. These clusters exhibited distinct progression patterns and statistically significant differences in baseline clinical features, disease onset characteristics, and survival outcomes, with median survival ranging from 12.8 months to 27.5 months. CONCLUSION: The proposed deep learning framework effectively captures the heterogeneous nature of ALS progression and identifies clinically relevant patient subgroups using routine clinical assessments. The stratification provides a foundation for personalized prognosis, optimized clinical trial design, and tailored therapeutic strategies, representing a practical tool for improving ALS patient management.
Federico De Mori Bajolin, Erica Tavazzi, Anna M. Bianchi, Martin O. Méndez
J. Biomed. Informatics2
2025 Exploring the Use of Projecting Conflicting Gradients in Multi-task Neural Networks with an Application to Amyotrophic Lateral Sclerosis
Davide Dei Cas, Enrico Longato, Erica Tavazzi, Umberto Manera, Adriano Chiò, Marta Gromicho, Inês Alves, Mamede de Carvalho, Barbara Di Camillo
AIME (1)3
2025 Towards Distributed Process Discovery in Healthcare: Testing and Proving the Feasibility of the Federated Alpha+ Algorithm
Leonardo Nucciarelli, Roberto Gatta, Andrada Mihaela Tudor, Erica Tavazzi, Giovanni Arcuri, Mauro Vallati, Gema Ibáñez-Sánchez, Zoe Valero-Ramon, Carlos Fernández-Llatas, Andrea Damiani
AIME (2)4
2025 Deep Learning Model Predicts Relapse Occurrence in Multiple Sclerosis Via Sequences of Environmental Data
abstract
Air pollution is a known risk factor for the exacerbation of many diseases. Among these, is multiple sclerosis (MS), a chronic, autoimmune, neurological disease, characterised by transient episodes of neurological impairment known as relapses. Although the link between environmental factors and relapses has been a subject of investigation in the medical and biostatistical literature, its implications for predictive modelling are still unclear. Thus, in this work, we develop a deep learning model that is able to combine four weeks of environmental data, collected by pollutant-monitoring and weather stations, with patient information to predict an imminent relapse in the following week. Specifically, we cast the task as distinguishing between 4-week sequences followed by a relapse vs. 4-week sequences followed by another relapse-free week, the latter of which were extracted from MS patients who were never observed to have had a relapse. The 1556 sequences were collected in the context of the H2020 BRAINTEASER (”Bringing Artificial Intelligence Home for a Better Care of Amyotrophic Lateral Sclerosis and Multiple Sclerosis”) project. The best-performing model was a recurrent neural network, which yielded an encouraging test-set area under the receiveroperating characteristic curve (AUROC) of 0.70. It also performed adequately (AUROC$=0.60$) on a modified version of the test set where the 4-week relapse-free sequences followed by another relapse-free week were extracted from the same subjects from whom the test sequences followed by a relapse came. Thus, our results, albeit preliminary, suggest that the inclusion of environmental data as the basis of predictive models of MS relapses is a promising direction to obtain short-term predictions, which may be helpful for therapy and life planning. It is especially encouraging that better-than-random performance was preserved on the modified test set, where environmental factors were, by construction, the most informative predictors.
Enrico Longato, Erica Tavazzi, Anna Milani, Elena Marinello, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo
BIBM2
2025 A Dynamic Bayesian Network Approach for Generating Synthetic Longitudinal Clinical Data: A Case Study on Long-Term Diabetes Outcomes
abstract
Synthetic clinical data offer several advantages, including the possibility to simulate patient trajectories and investigate long-term outcomes that would otherwise require extensive time and resources to investigate through traditional clinical trials. In this work, we propose a modelling approach based on dynamic Bayesian networks (DBNs) to generate reliable synthetic data that faithfully reproduce the characteristics of original longitudinal clinical datasets. The proposed pipeline includes two main steps: i) learning a DBN model using a layered variable structure informed by domain knowledge, and ii) employing the trained model to simulate longitudinal clinical data. We applied this approach to generate a synthetic version of the LEADER trial dataset, which includes longitudinal data of patients with type 2 diabetes and high cardiovascular risk. After preprocessing, the dataset consisted of 45,556 observations across 82 variables from 8,301 patients. The general utility of the synthetic data was assessed by comparing the distribution of each variable between the synthetic and original datasets. In addition, we evaluated the similarity of Kaplan-Meier survival curves for the major adverse cardiovascular events (MACE), that was the primary endpoints of the trial, between synthetic and real data. Overall, the proposed method demonstrated satisfactory performance, with the synthetic data closely replicating both the variable distributions and time-to-event outcomes observed in the original dataset.
Sara Poletto, Noemi Gonzato, Erica Tavazzi, Enrico Longato, Amanda Adler, Mari-Anne Gall, Matthias Müllenborn, Barbara Di Camillo, Martina Vettoretti
BIBM3
2024 Machine Learning Models Highlight the Impact of Pollution and Weather Patterns on Relapse Occurrence in Multiple Sclerosis Patients
abstract
Multiple Sclerosis (MS) is a chronic autoimmune and inflammatory neurological disorder characterised by episodes of symptom exacerbation, known as relapses. Relapses have been linked to environmental factors such as the weather and pollutant concentrations in the air, but the exact relationship between these phenomena is still unclear. In this study, we investigated the role of environmental factors in predicting imminent relapse occurrence in MS patients, leveraging clinical and environmental data collected over a period of one week preceeding the possible event, using data collected in the context of the H2020 BRAINTEASER project. To do this, we developed and tested a range of combinations of predictive models (logistic regression, LR; and random forest, RF) and feature selection schemes, both manual and data-driven. The RF model trained after a data-driven feature selection process based on the Variable Importance in Projection (VIP) metric yielded the best results, i.e., an AUC-ROC of 0.713 and an AUC-PR of 0.639. We identified several key predictors, including clinical variables such as time since MS onset, age at onset, diagnostic delay, and the Expanded Disability Status Scale (EDSS) score, and environmental variables such as wind speed, precipitation, NO2, PM10, average and maximum temperatures, and humidity. These findings suggest that environmental factors may be viable predictors of imminent relapse occurrence in MS.
Elena Marinello, Erica Tavazzi, Enrico Longato, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo
BIBM2
2024 DYNAMITE: Integrating Archetypal Analysis and Process Mining for Interpretable Disease Progression Modelling
abstract
DYNAMITE, an acronym for DYNamic Archetypal analysis for MIning disease TrajEctories, is a new methodology developed specifically to model disease progression by exploiting information available in longitudinal clinical datasets. First, archetypal analysis is applied to data organised in matrix form, with the aim of finding extreme and representative disease states (archetypes) linked to the original data through convex coefficients. Then, each original observation is associated with a single archetype based on their similarity; finally, an event log is created encoding the progression of disease states for each patient in terms of archetype states. In the last stage of the procedure, archetypal analysis is coupled with process mining, which allows the event log archetypes to be visualised graphically as sequences of disease states, allowing the clinical trajectories of patients to be extracted and examined. As a proof of concept, we applied the proposed method to data from a cohort of amyotrophic lateral sclerosis patients whose progression was monitored using the 12-item ALSFRS-R questionnaire. Without any a priori knowledge, DYNAMITE identified six archetypes clearly describing different types and severity of impairment and provided reliable clinical trajectories consistent with the prognosis of amyotrophic lateral sclerosis patients. DYNAMITE offers high interpretability at every stage of the analysis, which makes it particularly suitable for use in healthcare where explainability is paramount, and enables analysis of clinical trajectories at both individual and population levels.
Isotta Trescato, Erica Tavazzi, Martina Vettoretti, Roberto Gatta, Rosario Vasta, Adriano Chiò, Barbara Di Camillo
IEEE J. Biomed. Health Informatics2
2023 Dealing with Data Scarcity in Rare Diseases: Dynamic Bayesian Networks and Transfer Learning to Develop Prognostic Models of Amyotrophic Lateral Sclerosis
Enrico Longato, Erica Tavazzi, Adriano Chiò, Gabriele Mora, Giovanni Sparacino, Barbara Di Camillo
AIME2
2023 Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic review
abstract
BACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors.
Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo
Artif. Intell. Medicine1
2022 Eleven quick tips for data cleaning and feature engineering
abstract
Applying computational statistics or machine learning methods to data is a key component of many scientific studies, in any field, but alone might not be sufficient to generate robust and reliable outcomes and results. Before applying any discovery method, preprocessing steps are necessary to prepare the data to the computational analysis. In this framework, data cleaning and feature engineering are key pillars of any scientific study involving data analysis and that should be adequately designed and performed since the first phases of the project. We call "feature" a variable describing a particular trait of a person or an observation, recorded usually as a column in a dataset. Even if pivotal, these data cleaning and feature engineering steps sometimes are done poorly or inefficiently, especially by beginners and unexperienced researchers. For this reason, we propose here our quick tips for data cleaning and feature engineering on how to carry out these important preprocessing steps correctly avoiding common mistakes and pitfalls. Although we designed these guidelines with bioinformatics and health informatics scenarios in mind, we believe they can more in general be applied to any scientific area. We therefore target these guidelines to any researcher or practitioners wanting to perform data cleaning or feature engineering. We believe our simple recommendations can help researchers and scholars perform better computational analyses that can lead, in turn, to more solid outcomes and more reliable discoveries.
Davide Chicco, Luca Oneto, Erica Tavazzi
PLoS Comput. Biol.3