Riccardo Bellazzi

dblp:27/3072 · DBLP profile ↗
← Back
165ranked-venue papers
30as first author
36since 2021 · last 2026
0000-0002-6974-9808ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 126 · 13 first-author · 30 since 2021Artificial intelligence and machine learning · 46 · 17 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 Is Heavier Better? Benchmarking Optical Flow Time-Series vs. Video Transformers for In Vitro Fertilization Counseling
Lorenzo Corso, Andrea Fantinato, Emirhan Kayar, Margherita Isernia, Giulia Fiorentino, Marilena Taggi, Federica Innocenti, Marcos Meseguer, Alberto Vaiarelli, Laura Rienzi, Maurizio Zuccotti, Giovanni Coticchio, Riccardo Bellazzi, Danilo Cimadomo, Giovanna Nicora
AIME (1)13
2026 Automated Pulmonary Hypertension Subtype Discrimination from CTPA Scans
Matteo Dallera, Carlotta Pairazzi, Riccardo Bellazzi, Stefano Ghio, Adele Valentini, Lucia Sacchi
AIME (1)3
2026 Biological Plausibility Assessment of Viral Sequences Generated by a Genomic Language Model
Pablo Arozarena Donelli, Simone Rancati, Giovanna Nicora, Riccardo Bellazzi, Enea Parimbelli, Luigi Portinale
AIME (2)4
2026 Informative Missingness to Generate Irregular Clinical Time Series
Hadi Mehdizavareh, Gabriele Santangelo, Giovanna Nicora, Simon Lebech Cichosz, Arianna Dagliati, Arijit Khan 0001, Riccardo Bellazzi
AIME (2)7
2026 A Scoring Strategy to Assess AI Prediction Reliability: Validation and Impact on Medical Decision Making
Lorenzo Peracchio, Laura Bergomi, Ana Isabel Hernáiz Ferrer, Chandra Bortolotto, Valentina Zuccaro, Francesco Salinaro, Lorenzo Preda, Riccardo Bellazzi, Giovanna Nicora
AIME (1)8
2026 Epistemologically Guided LLM Reasoning for Differential Diagnosis
Simone Rancati, Laura Bergomi, Enea Parimbelli, Giovanna Nicora, Riccardo Bellazzi
AIME (1)5
2026 Forecasting Gait Dynamics with Foundation Models
Federico Colelli Riano, Alberto Malovini, Armando Coccia, Federica Amitrano, Gianni D'Addio, Riccardo Bellazzi
AIME (2)6
2026 Artificial intelligence use and performance in detecting and predicting healthcare-associated infections: A systematic review
abstract
The increasing digitisation of healthcare data and the rapid development of Artificial Intelligence (AI) pave the way for innovative strategies for infectious disease management. This study aimed to systematically retrieve and summarize current evidence on the use and performance of AI-based models for healthcare-associated infection (HAI) detection (i.e., identifying infections already present in available data) and prediction (i.e., estimating future risk based on earlier patient information). PubMed, Embase, Scopus and Web of Science were searched for experimental and observational studies published between 1 July 2018 and 12 February 2024. Primary outcomes included technical performance metrics for HAI detection and prediction (e.g. recall, precision, AUROC). Any reported clinical, organisational or economic impacts were evaluated as secondary outcomes. Of 4489 records initially identified, 121 studies were included. Twenty-five studies (20.6 %) focused on HAI detection, with more than half achieving an AUROC above 0.90. In contrast, studies on HAI prediction ( n = 93, 76.9 %) reported more heterogeneous performance. Among studies comparing AI with traditional methods ( n = 32), AI models outperformed conventional approaches in 81.3 % of cases ( n = 26). A growing body of evidence suggests that AI models are equal to or superior to traditional methods for HAI detection and prediction, but challenges remain in evaluating performance, with many studies lacking comparators, few prospective evaluations, and limited assessment of organisational impact. • We observed a significant increase in the number of published studies since 2018 • AI models appear to be equal or superior to traditional methods in HAI control • Overall, AI models show high sensibility and specificity, but low precision • Detection models generally outperform prediction models in terms of AUROC • Many studies lack comparators and prospective assessment of organisational impact
Chiara Barbati, Luca Viviani, Riccardo Vecchio, Guglielmo Arzilli, Luigi De Angelis, Francesco Baglivo, Lucia Sacchi, Riccardo Bellazzi, Caterina Rizzo, Anna Odone
Artif. Intell. Medicine8
2026 Rehabilitation movement simulation via joint angle-based generative AI
abstract
In recent years, generative models have shown remarkable capabilities in synthesizing realistic human motion, with applications ranging from animation to virtual reality. However, their potential in clinical and rehabilitation settings remains underexplored. In this work, we introduce a conditional diffusion-based generative framework for rehabilitation-oriented motion synthesis, which directly operates on joint-angle representations of full-body movement. Unlike most existing approaches that rely on joint positions, our method generates motion in a clinically meaningful space that explicitly encodes joint range of motion, aligning the generation process with how motor performance is assessed in rehabilitation practice. This design enables subject-independent modeling while improving the interpretability of the generated movements from a clinical perspective. We propose a comprehensive evaluation protocol by combining qualitative and quantitative metrics, including simulation visualizations, similarity analysis, and automated assessment of simulations adherence to users input. Experiments based on cross-subject and leave-one-combination-out settings demonstrate the model's ability to generate plausible, contextually accurate motion sequences, with improved generalization when using joint angle representations, achieving superior performance compared to a position-based approach. Despite limitations due to dataset size and gesture diversity, results support the feasibility of generating rehabilitation-oriented motion simulations, motivating future investigation in personalized rehabilitation scenarios.
Gabriele Santangelo, Chiara Alessi, Giovanna Nicora, Nikolas Sacchi, Samuele Pe, Antonella Ferrara, Riccardo Bellazzi, Arianna Dagliati
Artif. Intell. Medicine7
2025 Combining Clinical and Gene Expression Variables via Knowledge Graph Embedding for Prediction of Coronary Artery Stenosis
Giuseppe Albi, Arianna Dagliati, Chiara Vavassori, Laura Pisani, Mattia Chiesa, Luca Piacentini, Saima Mushtaq, Gianluca Pontone, Riccardo Bellazzi, Gualtiero Colombo 0002
AIME (1)9
2025 Enhancing RAGs for Rheumatology Triage: Strategies for Optimized Knowledge Retrieval
Tommaso Mario Buonocore, Emanuele Cardinale, Garifallia Sakellariou, Riccardo Bellazzi, Lucia Sacchi
AIME (2)4
2025 Upper Limb Movements Simulations with Generative Diffusion Models
Gabriele Santangelo, Chiara Alessi, Nikolas Sacchi, Giovanna Nicora, Riccardo Bellazzi, Antonella Ferrara, Arianna Dagliati
AIME (2)5
2025 Deep Learning Model Predicts Relapse Occurrence in Multiple Sclerosis Via Sequences of Environmental Data
abstract
Air pollution is a known risk factor for the exacerbation of many diseases. Among these, is multiple sclerosis (MS), a chronic, autoimmune, neurological disease, characterised by transient episodes of neurological impairment known as relapses. Although the link between environmental factors and relapses has been a subject of investigation in the medical and biostatistical literature, its implications for predictive modelling are still unclear. Thus, in this work, we develop a deep learning model that is able to combine four weeks of environmental data, collected by pollutant-monitoring and weather stations, with patient information to predict an imminent relapse in the following week. Specifically, we cast the task as distinguishing between 4-week sequences followed by a relapse vs. 4-week sequences followed by another relapse-free week, the latter of which were extracted from MS patients who were never observed to have had a relapse. The 1556 sequences were collected in the context of the H2020 BRAINTEASER (”Bringing Artificial Intelligence Home for a Better Care of Amyotrophic Lateral Sclerosis and Multiple Sclerosis”) project. The best-performing model was a recurrent neural network, which yielded an encouraging test-set area under the receiveroperating characteristic curve (AUROC) of 0.70. It also performed adequately (AUROC$=0.60$) on a modified version of the test set where the 4-week relapse-free sequences followed by another relapse-free week were extracted from the same subjects from whom the test sequences followed by a relapse came. Thus, our results, albeit preliminary, suggest that the inclusion of environmental data as the basis of predictive models of MS relapses is a promising direction to obtain short-term predictions, which may be helpful for therapy and life planning. It is especially encouraging that better-than-random performance was preserved on the modified test set, where environmental factors were, by construction, the most informative predictors.
Enrico Longato, Erica Tavazzi, Anna Milani, Elena Marinello, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo
BIBM8
2025 SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike protein
abstract
BACKGROUND: The COVID-19 pandemic has caused over 776 million infections and 7 million deaths globally between December 2019 and November 2024. Since the emergence of the original Wuhan strain, SARS-CoV-2 has evolved into multiple variants-including Alpha, Delta, and Omicron-primarily through mutations in the Spike glycoprotein. The S1 subunit, which binds the human angiotensin-converting enzyme 2 (ACE2) receptor, mutates frequently and plays a key role in infectivity and immune escape, while the more conserved S2 subunit mediates membrane fusion. Anticipating future mutations is essential for guiding vaccine design and therapeutic strategies. Generative Large Language Models (LLMs) have shown promise in protein sequence modeling due to their capacity to produce realistic and functional synthetic sequences. Here, we introduce SARITA, a GPT-3-based LLM with up to 1.2 billion parameters, fine-tuned via continual learning on the protein model RITA trained on 107 017 high-quality SARS-CoV-2 Spike sequences (up to March 1st 2021) to generate high-quality synthetic SARS-CoV-2 Spike S1 subunits. RESULTS: SARITA is able to generate realistic, full-length synthetic S1 subunits starting from a 14-amino-acid prompt. When evaluated on unseen sequences collected between March 2021 and November 2023-including major Variants of Concern (VOCs) such as Delta and Omicron, and Variants of Interest such as Iota-SARITA outperforms baseline and state-of-the-art LLMs in terms of sequence quality, biological plausibility, and similarity to real-world viral evolution. SARITA generates high-quality sequences in over 97% of cases, with markedly lower False Mutation Rate and higher similarity scores (PAM30, Levenshtein distance) compared to alternative approaches. It also accurately reproduces key mutations characteristic of future variants-such as L212I, R158L, T95P, and E406K-which were not present in the training data but emerged later in VOCs like Omicron and Delta. Structure-based analysis confirms the functional plausibility of these substitutions, with ΔΔG values within experimentally supported thresholds for ACE2 and antibody binding. Furthermore, SARITA anticipates immune-evasive mutations and accurately captures the positional and statistical distribution of mutations found in post- March 1st 2021 variants, highlighting its potential as a predictive tool for viral evolution. CONCLUSION: These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments.
Simone Rancati, Giovanna Nicora, Laura Bergomi, Tommaso Mario Buonocore, Daniel M. Czyz, Enea Parimbelli, Riccardo Bellazzi, Marco Salemi, Mattia Prosperi, Simone Marini
Briefings Bioinform.7
2024 Do You Trust Your Model Explanations? An Analysis of XAI Performance Under Dataset Shift
Lorenzo Peracchio, Giovanna Nicora, Tommaso Mario Buonocore, Riccardo Bellazzi, Enea Parimbelli
AIME (2)4
2024 Assessing a Personalized, Hybrid, and Generic Approach for Glucose Prediction in Type 1 Diabetes
abstract
This study investigates the potential of using generic, hybrid, and personalized neural network models for glucose prediction in individuals with Type 1 Diabetes (T1D). Data from 194 participants in the Wireless Innovations for Seniors with Diabetes Mellitus (WISDM) study, totaling over 46 million minutes of Continuous Glucose Monitoring (CGM), were used to develop and evaluate the models. A baseline reference model, Last Observation Carried Forward (LOCF), was also included for comparison. Models were trained using data from 70% of the participants and tested on the remaining 30%, with prediction horizons (PH) set at 30 and 60 minutes. At the 30-minute PH, the generic model achieved a Root Mean Square Error (RMSE) of 19.6 mg/dL and a Mean Absolute Relative Difference (MARD) of 9.2%. These results were slightly worse than those of the hybrid and personalized models, which yielded RMSEs of 19.4 mg/dL and 19.5 mg/dL, respectively, and MARDs of 9.6% for both. However, the differences were not statistically significant. At the 60-minute PH, the generic model showed the best performance, with an RMSE of 34.9 mg/dL and MARD of 16.2%. The hybrid and personalized models exhibited slightly higher RMSEs (35.3 mg/dL and 35.5 mg/dL, respectively) and MARDs (17.8% and 18.1%, respectively). These findings suggest that both generic and individualized models can provide satisfactory glucose forecasting results based solely on CGM data. Nonetheless, the potential benefits of individualized approaches deserve further investigation, particularly when substantial training data are available.
Pietro Bosoni, Morten Hasselstrøm Jensen, Riccardo Bellazzi, Simon Lebech Cichosz
BIBM3
2024 Machine Learning Models Highlight the Impact of Pollution and Weather Patterns on Relapse Occurrence in Multiple Sclerosis Patients
abstract
Multiple Sclerosis (MS) is a chronic autoimmune and inflammatory neurological disorder characterised by episodes of symptom exacerbation, known as relapses. Relapses have been linked to environmental factors such as the weather and pollutant concentrations in the air, but the exact relationship between these phenomena is still unclear. In this study, we investigated the role of environmental factors in predicting imminent relapse occurrence in MS patients, leveraging clinical and environmental data collected over a period of one week preceeding the possible event, using data collected in the context of the H2020 BRAINTEASER project. To do this, we developed and tested a range of combinations of predictive models (logistic regression, LR; and random forest, RF) and feature selection schemes, both manual and data-driven. The RF model trained after a data-driven feature selection process based on the Variable Importance in Projection (VIP) metric yielded the best results, i.e., an AUC-ROC of 0.713 and an AUC-PR of 0.639. We identified several key predictors, including clinical variables such as time since MS onset, age at onset, diagnostic delay, and the Expanded Disability Status Scale (EDSS) score, and environmental variables such as wind speed, precipitation, NO2, PM10, average and maximum temperatures, and humidity. These findings suggest that environmental factors may be viable predictors of imminent relapse occurrence in MS.
Elena Marinello, Erica Tavazzi, Enrico Longato, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo
BIBM7
2024 Land Use Regression on Interpolated Urban Graphs to Assess Personal Exposure to Air Pollution
abstract
Past research has demonstrated that continuous exposure to pollutants, such as PM2.5 and PM10, is associated with an increased risk of developing and worsening respiratory and neurodegenerative diseases. Calculating and reducing exposure to these pollutants is crucial to assess these risks and perform proper prevention. In this study, we estimate personal exposure to PM2.5 based on the integration of sensors measurements, meteorological data and land use parameters, which could impact on actual pollution levels, especially in areas located far from the sensors. Pollution data have been collected from a dense network of sensors located in Pavia, Italy, meteorological and geographical data have been collected from public sources. We used geographical data to create graphs that model the city road structure, and applied Land Use Regression methods to estimate air pollution on its nodes, adjusting the measurements interpolated from the sensors with the effects of weather data, land use parameters such as the distance from the closest high-traffic road, and additional temporal information such as weekends/holidays and working days. We tested several regression methods: linear regression, both simple and with regularization (Ridge, LASSO and ElasticNet), Random Forest regression, Gradient Boosting and Support Vector Regression (SVR). Results show that meteorological variables, namely temperature and humidity, and temporal factors do contribute significantly in obtaining pollution values in the graph nodes that differ from values obtained exclusively through sensors interpolation.
Daniele Pala, Giacomo Zagami, Pietro Bosoni, Mahin Vazifehdan, Riccardo Bellazzi, Arianna Dagliati
BIBM5
2024 Sequencing Efforts and Epidemiological Trends: Analyzing SARS-CoV-2 Dynamics Across European Nations
abstract
The COVID-19 pandemic has profoundly impacted global health, leading to millions of deaths and overwhelming healthcare systems worldwide. This study investigates the relationship between SARS-CoV-2 sequencing rates and critical epidemiological parameters, such as cases, deaths, and ICU admissions, across 25 European countries from January 2020 to November 2023. By analyzing these relationships, we aim to determine whether sequencing efforts were reactive—in response to epidemiological pressures—or proactive, guided by public health strategies. The analysis used publicly available data from GISAID, OxCGRT, and ECDC, and included weekly aggregation, correlation analysis, and the application of TimeGPT for predictive modeling. Results show that sequencing rates were significantly correlated with ICU admissions, hospitalizations, case numbers, and deaths, though with variability between countries and over different pandemic phases. TimeGPT analysis revealed that sequencing rates were often the most informative feature for predicting future COVID-19 cases in many countries. These findings highlight the potential of sequencing rates to serve as early indicators for severe pandemic outcomes and underscore the importance of context-specific approaches for managing future health crises.
Simone Rancati, Daniele Pala, Simone Marini, Marco Salemi, Riccardo Bellazzi, Giovanna Nicora
BIBM5
2024 Reshaping free-text radiology notes into structured reports with generative question answering transformers
abstract
BACKGROUND: Radiology reports are typically written in a free-text format, making clinical information difficult to extract and use. Recently, the adoption of structured reporting (SR) has been recommended by various medical societies thanks to the advantages it offers, e.g. standardization, completeness, and information retrieval. We propose a pipeline to extract information from Italian free-text radiology reports that fits with the items of the reference SR registry proposed by a national society of interventional and medical radiology, focusing on CT staging of patients with lymphoma. METHODS: Our work aims to leverage the potential of Natural Language Processing and Transformer-based models to deal with automatic SR registry filling. With the availability of 174 Italian radiology reports, we investigate a rule-free generative Question Answering approach based on the Italian-specific version of T5: IT5. To address information content discrepancies, we focus on the six most frequently filled items in the annotations made on the reports: three categorical (multichoice), one free-text (free-text), and two continuous numerical (factual). In the preprocessing phase, we encode also information that is not supposed to be entered. Two strategies (batch-truncation and ex-post combination) are implemented to comply with the IT5 context length limitations. Performance is evaluated in terms of strict accuracy, f1, and format accuracy, and compared with the widely used GPT-3.5 Large Language Model. Unlike multichoice and factual, free-text answers do not have 1-to-1 correspondence with their reference annotations. For this reason, we collect human-expert feedback on the similarity between medical annotations and generated free-text answers, using a 5-point Likert scale questionnaire (evaluating the criteria of correctness and completeness). RESULTS: The combination of fine-tuning and batch splitting allows IT5 ex-post combination to achieve notable results in terms of information extraction of different types of structured data, performing on par with GPT-3.5. Human-based assessment scores of free-text answers show a high correlation with the AI performance metrics f1 (Spearman's correlation coefficients>0.5, p-values<0.001) for both IT5 ex-post combination and GPT-3.5. The latter is better at generating plausible human-like statements, even if it systematically provides answers even when they are not supposed to be given. CONCLUSIONS: In our experimental setting, a fine-tuned Transformer-based model with a modest number of parameters (i.e., IT5, 220 M) performs well as a clinical information extraction system for automatic SR registry filling task. It can extract information from more than one place in the report, elaborating it in a manner that complies with the response specifications provided by the SR registry (for multichoice and factual items), or that closely approximates the work of a human-expert (free-text items); with the ability to discern when an answer is supposed to be given or not to a user query.
Laura Bergomi, Tommaso Mario Buonocore, Paolo Antonazzo, Lorenzo Alberghi, Riccardo Bellazzi, Lorenzo Preda, Chandra Bortolotto, Enea Parimbelli
Artif. Intell. Medicine5
2024 Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncoders
abstract
The COVID-19 pandemic is marked by the successive emergence of new SARS-CoV-2 variants, lineages, and sublineages that outcompete earlier strains, largely due to factors like increased transmissibility and immune escape. We propose DeepAutoCoV, an unsupervised deep learning anomaly detection system, to predict future dominant lineages (FDLs). We define FDLs as viral (sub)lineages that will constitute >10% of all the viral sequences added to the GISAID, a public database supporting viral genetic sequence sharing, in a given week. DeepAutoCoV is trained and validated by assembling global and country-specific data sets from over 16 million Spike protein sequences sampled over a period of ~4 years. DeepAutoCoV successfully flags FDLs at very low frequencies (0.01%-3%), with median lead times of 4-17 weeks, and predicts FDLs between ~5 and ~25 times better than a baseline approach. For example, the B.1.617.2 vaccine reference strain was flagged as FDL when its frequency was only 0.01%, more than a year before it was considered for an updated COVID-19 vaccine. Furthermore, DeepAutoCoV outputs interpretable results by pinpointing specific mutations potentially linked to increased fitness and may provide significant insights for the optimization of public health 'pre-emptive' intervention strategies.
Simone Rancati, Giovanna Nicora, Mattia Prosperi, Riccardo Bellazzi, Marco Salemi, Simone Marini
Briefings Bioinform.4
2023 A Topological Data Analysis Framework for Computational Phenotyping
Giuseppe Albi, Alessia Gerbasi, Mattia Chiesa, Gualtiero Colombo 0002, Riccardo Bellazzi, Arianna Dagliati
AIME5
2023 A Rule-Free Approach for Cardiological Registry Filling from Italian Clinical Notes with Question Answering Transformers
Tommaso Mario Buonocore, Enea Parimbelli, Valentina Tibollo, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi
AIME6
2023 Why did AI get this one wrong? - Tree-based explanations of machine learning model predictions
abstract
Increasingly complex learning methods such as boosting, bagging and deep learning have made ML models more accurate, but harder to interpret and explain, culminating in black-box machine learning models. Model developers and users alike are often presented with a trade-off between performance and intelligibility, especially in high-stakes applications like medicine. In the present article we propose a novel methodological approach for generating explanations for the predictions of a generic machine learning model, given a specific instance for which the prediction has been made. The method, named AraucanaXAI, is based on surrogate, locally-fitted classification and regression trees that are used to provide post-hoc explanations of the prediction of a generic machine learning model. Advantages of the proposed XAI approach include superior fidelity to the original model, ability to deal with non-linear decision boundaries, and native support to both classification and regression problems. We provide a packaged, open-source implementation of the AraucanaXAI method and evaluate its behaviour in a number of different settings that are commonly encountered in medical applications of AI. These include potential disagreement between the model prediction and physician's expert opinion and low reliability of the prediction due to data scarcity.
Enea Parimbelli, Tommaso Mario Buonocore, Giovanna Nicora, Wojtek Michalowski, Szymon Wilk, Riccardo Bellazzi
Artif. Intell. Medicine6
2023 Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic review
abstract
BACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors.
Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo
Artif. Intell. Medicine21
2023 Localizing in-domain adaptation of transformer-based biomedical language models
abstract
In the era of digital healthcare, the huge volumes of textual information generated every day in hospitals constitute an essential but underused asset that could be exploited with task-specific, fine-tuned biomedical language representation models, improving patient care and management. For such specialized domains, previous research has shown that fine-tuning models stemming from broad-coverage checkpoints can largely benefit additional training rounds over large-scale in-domain resources. However, these resources are often unreachable for less-resourced languages like Italian, preventing local medical institutions to employ in-domain adaptation. In order to reduce this gap, our work investigates two accessible approaches to derive biomedical language models in languages other than English, taking Italian as a concrete use-case: one based on neural machine translation of English resources, favoring quantity over quality; the other based on a high-grade, narrow-scoped corpus natively written in Italian, thus preferring quality over quantity. Our study shows that data quantity is a harder constraint than data quality for biomedical adaptation, but the concatenation of high-quality data can improve model performance even when dealing with relatively size-limited corpora. The models published from our investigations have the potential to unlock important research opportunities for Italian hospitals and academia. Finally, the set of lessons learned from the study constitutes valuable insights towards a solution to build biomedical language models that are generalizable to other less-resourced languages and different domain settings.
Tommaso Mario Buonocore, Claudio Crema, Alberto Redolfi, Riccardo Bellazzi, Enea Parimbelli
J. Biomed. Informatics4
2023 Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application
abstract
The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation by using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Transformers-based model. Moreover, we collected and leveraged three external independent datasets to implement an effective multicenter model, with overall F1-score 84.77 %, Precision 83.16 %, Recall 86.44 %. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a "low-resource" approach. This allowed us to establish methodological guidelines that pave the way for Natural Language Processing studies in less-resourced languages.
Claudio Crema, Tommaso Mario Buonocore, Silvia Fostinelli, Enea Parimbelli, Federico Verde, Cira Fundarò, Marina Manera, Matteo Cotta Ramusino, Marco Capelli, Alfredo Costa, Giuliano Binetti, Riccardo Bellazzi, Alberto Redolfi
J. Biomed. Informatics12
2022 The PERISCOPE Data Atlas: A Demonstration of Release v1.2
Enea Parimbelli, Cristiana Larizza, Vladimir Urosevic, Andrea Pogliaghi, Manuel Ottaviano, Cindy Cheng, Vincent Benoit, Daniele Pala, Vittorio Casella, Riccardo Bellazzi, Paolo Giudici
AIME10
2022 A Deductive Data-Driven Pipeline Powered by MLHO for Post-Acute Sequelae of COVID-19 (PASC) Phenotyping
Arianna Dagliati, Zachary H. Strasser, Rebecca Mesa, Zahra Shakeri, Alaleh Azhir, Riccardo Bellazzi, Shawn N. Murphy, Hossein Estiri
AMIA6
2022 A manifesto on explainability for artificial intelligence in medicine
abstract
The rapid increase of interest in, and use of, artificial intelligence (AI) in computer applications has raised a parallel concern about its ability (or lack thereof) to provide understandable, or explainable, output to users. This concern is especially legitimate in biomedical contexts, where patient safety is of paramount importance. This position paper brings together seven researchers working in the field with different roles and perspectives, to explore in depth the concept of explainable AI, or XAI, offering a functional definition and conceptual framework or model that can be used when considering XAI. This is followed by a series of desiderata for attaining explainability in AI, each of which touches upon a key domain in biomedicine.
Carlo Combi, Beatrice Amico, Riccardo Bellazzi, Andreas Holzinger, Jason H. Moore, Marinka Zitnik, John H. Holmes
Artif. Intell. Medicine3
2022 Evaluating pointwise reliability of machine learning prediction
abstract
Interest in Machine Learning applications to tackle clinical and biological problems is increasing. This is driven by promising results reported in many research papers, the increasing number of AI-based software products, and by the general interest in Artificial Intelligence to solve complex problems. It is therefore of importance to improve the quality of machine learning output and add safeguards to support their adoption. In addition to regulatory and logistical strategies, a crucial aspect is to detect when a Machine Learning model is not able to generalize to new unseen instances, which may originate from a population distant to that of the training population or from an under-represented subpopulation. As a result, the prediction of the machine learning model for these instances may be often wrong, given that the model is applied outside its "reliable" space of work, leading to a decreasing trust of the final users, such as clinicians. For this reason, when a model is deployed in practice, it would be important to advise users when the model's predictions may be unreliable, especially in high-stakes applications, including those in healthcare. Yet, reliability assessment of each machine learning prediction is still poorly addressed. Here, we review approaches that can support the identification of unreliable predictions, we harmonize the notation and terminology of relevant concepts, and we highlight and extend possible interrelationships and overlap among concepts. We then demonstrate, on simulated and real data for ICU in-hospital death prediction, a possible integrative framework for the identification of reliable and unreliable predictions. To do so, our proposed approach implements two complementary principles, namely the density principle and the local fit principle. The density principle verifies that the instance we want to evaluate is similar to the training set. The local fit principle verifies that the trained model performs well on training subsets that are more similar to the instance under evaluation. Our work can contribute to consolidating work in machine learning especially in medicine.
Giovanna Nicora, Miguel Ángel Ríos-Gaona, Ameen Abu-Hanna, Riccardo Bellazzi
J. Biomed. Informatics4
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics13
2021 A Topological Data Analysis Mapper of the Ovarian Folliculogenesis Based on MALDI Mass Spectrometry Imaging Proteomics
Giulia Campi, Giovanna Nicora, Giulia Fiorentino, Fulvio Magni, Silvia Garagna, Maurizio Zuccotti, Riccardo Bellazzi
AIME8
2021 Temporal Phenotypic Pathways of Post-Acute Sequelae of SARS-CoV-2 by an International Consortium for Clinical Characterization of COVID-19 (4CE)
Shawn N. Murphy, Hossein Estiri, Arianna Dagliati, Riccardo Bellazzi, John H. Holmes
AMIA4
2021 Health informatics and EHR to support clinical research in the COVID-19 pandemic: an overview
abstract
The coronavirus disease 2019 (COVID-19) pandemic has clearly shown that major challenges and threats for humankind need to be addressed with global answers and shared decisions. Data and their analytics are crucial components of such decision-making activities. Rather interestingly, one of the most difficult aspects is reusing and sharing of accurate and detailed clinical data collected by Electronic Health Records (EHR), even if these data have a paramount importance. EHR data, in fact, are not only essential for supporting day-by-day activities, but also they can leverage research and support critical decisions about effectiveness of drugs and therapeutic strategies. In this paper, we will concentrate our attention on collaborative data infrastructures to support COVID-19 research and on the open issues of data sharing and data governance that COVID-19 had made emerge. Data interoperability, healthcare processes modelling and representation, shared procedures to deal with different data privacy regulations, and data stewardship and governance are seen as the most important aspects to boost collaborative research. Lessons learned from COVID-19 pandemic can be a strong element to improve international research and our future capability of dealing with fast developing emergencies and needs, which are likely to be more frequent in the future in our connected and intertwined world.
Arianna Dagliati, Alberto Malovini, Valentina Tibollo, Riccardo Bellazzi
Briefings Bioinform.4
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.25
2020 Deep Learning Applied to Blood Glucose Prediction from Flash Glucose Monitoring and Fitbit Data
Pietro Bosoni, Marco Meccariello, Valeria Calcaterra, Cristiana Larizza, Lucia Sacchi, Riccardo Bellazzi
AIME6
2020 Explainable Artificial Intelligence (XAI): Current Approaches and Paths to the Future
John H. Holmes, Riccardo Bellazzi, Carlo Combi, Jason H. Moore, Niels Peek
AMIA2
2020 A Reliable Machine Learning Approach applied to Single-Cell Classification in Acute Myeloid Leukemia
Giovanna Nicora, Riccardo Bellazzi
AMIA2
2020 The PULSE Project: A Case of Use of Big Data Uses Toward a Cohomprensive Health Vision of City Well Being
abstract
Despite the silent effects sometimes hidden to the major audience, air pollution is becoming one of the most impactful threat to global health. Cities are the places where deaths due to air pollution are concentrated most. In order to correctly address intervention and prevention thus is essential to assest the risk and the impacts of air pollution spatially and temporally inside the urban spaces. PULSE aims to design and build a large-scale data management system enabling real time analytics of health, behaviour and environmental data on air quality. The objective is to reduce the environmental and behavioral risk of chronic disease incidence to allow timely and evidence-driven management of epidemiological episodes linked in particular to two pathologies; asthma and type 2 diabetes in adult populations. developing a policy-making across the domains of health, environment, transport, planning in the PULSE test bed cities.
Domenico Vito, Manuel Ottaviano, Riccardo Bellazzi, Cristiana Larizza, Vittorio Casella, Daniele Pala, Marica Franzini
ICOST3
2020 Mining post-surgical care processes in breast cancer patients
Lorenzo Chiudinelli, Arianna Dagliati, Valentina Tibollo, Sara Albasini, Nophar Geifman, Niels Peek, John H. Holmes, Fabio Corsi, Riccardo Bellazzi, Lucia Sacchi
Artif. Intell. Medicine9
2020 Using topological data analysis and pseudo time series to infer temporal phenotypes from electronic health records
abstract
Temporal phenotyping enables clinicians to better understand observable characteristics of a disease as it progresses. Modelling disease progression that captures interactions between phenotypes is inherently challenging. Temporal models that capture change in disease over time can identify the key features that characterize disease subtypes that underpin these trajectories. These models will enable clinicians to identify early warning signs of progression in specific sub-types and therefore to make informed decisions tailored to individual patients. In this paper, we explore two approaches to building temporal phenotypes based on the topology of data: topological data analysis and pseudo time-series. Using type 2 diabetes data, we show that the topological data analysis approach is able to identify disease trajectories and that pseudo time-series can infer a state space model characterized by transitions between hidden states that represent distinct temporal phenotypes. Both approaches highlight lipid profiles as key factors in distinguishing the phenotypes.
Arianna Dagliati, Nophar Geifman, Niels Peek, John H. Holmes, Lucia Sacchi, Riccardo Bellazzi, Seyed Erfan Sajjadi, Allan Tucker
Artif. Intell. Medicine6
2020 A Bayesian data fusion based approach for learning genome-wide transcriptional regulatory networks
abstract
BACKGROUND: Reverse engineering of transcriptional regulatory networks (TRN) from genomics data has always represented a computational challenge in System Biology. The major issue is modeling the complex crosstalk among transcription factors (TFs) and their target genes, with a method able to handle both the high number of interacting variables and the noise in the available heterogeneous experimental sources of information. RESULTS: In this work, we propose a data fusion approach that exploits the integration of complementary omics-data as prior knowledge within a Bayesian framework, in order to learn and model large-scale transcriptional networks. We develop a hybrid structure-learning algorithm able to jointly combine TFs ChIP-Sequencing data and gene expression compendia to reconstruct TRNs in a genome-wide perspective. Applying our method to high-throughput data, we verified its ability to deal with the complexity of a genomic TRN, providing a snapshot of the synergistic TFs regulatory activity. Given the noisy nature of data-driven prior knowledge, which potentially contains incorrect information, we also tested the method's robustness to false priors on a benchmark dataset, comparing the proposed approach to other regulatory network reconstruction algorithms. We demonstrated the effectiveness of our framework by evaluating structural commonalities of our learned genomic network with other existing networks inferred by different DNA binding information-based methods. CONCLUSIONS: This Bayesian omics-data fusion based methodology allows to gain a genome-wide picture of the transcriptional interplay, helping to unravel key hierarchical transcriptional interactions, which could be subsequently investigated, and it represents a promising learning approach suitable for multi-layered genomic data integration, given its robustness to noisy sources and its tailored framework for handling high dimensional data.
Elisabetta Sauta, Andrea Demartini, Francesca Vitali, Alberto Riva, Riccardo Bellazzi
BMC Bioinform.5
2020 SCOR: A secure international informatics infrastructure to investigate COVID-19
abstract
Global pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale.
Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux
J. Am. Medical Informatics Assoc.5
2020 A survey on single and multi omics data mining methods in cancer data classification
Zahra Momeni, Esmail Hassanzadeh, Mohammad Saniee Abadeh, Riccardo Bellazzi
J. Biomed. Informatics4
2020 A continuous-time Markov model approach for modeling myelodysplastic syndromes progression from cross-sectional data
Giovanna Nicora, F. Moretti, Elisabetta Sauta, Matteo Giovanni Della Porta, Luca Malcovati, Mario Cazzola, Silvana Quaglini, Riccardo Bellazzi
J. Biomed. Informatics8
2019 A Rule-Based Expert System for Automatic Implementation of Somatic Variant Clinical Interpretation Guidelines
Giovanna Nicora, Ivan Limongelli, Riccardo Cova, Matteo Giovanni Della Porta, Luca Malcovati, Mario Cazzola, Riccardo Bellazzi
AIME7
2019 A Semi-supervised Learning Approach for Pan-Cancer Somatic Genomic Variant Classification
Giovanna Nicora, Simone Marini, Ivan Limongelli, Ettore Rizzo, Stefano Montoli, Francesca Floriana Tricomi, Riccardo Bellazzi
AIME7
2019 Agent-Based Models and Spatial Enablement: A Simulation Tool to Improve Health and Wellbeing in Big Cities
Daniele Pala, John H. Holmes, José Pagán, Enea Parimbelli, Marica Teresa Rocca, Vittorio Casella, Riccardo Bellazzi
AIME7
2019 Latent Class Multi-Label Classification to Identify Subclasses of Disease for Improved Prediction
abstract
Disease subtyping can assist the development of precision medicine but remains a challenge in data analysis by reason of the many different methods to group individuals depending on their data. However, identification of subclasses of disease will help to produce better models which are more specific to patients and will improve prediction and interpretation of underlying characteristics of disease. This paper presents a novel algorithm that integrates latent class models with supervised learning. The new algorithm uses latent class models to cluster patients within groups that results in improved classification as well as aiding the understanding of the dissimilarities of the discovered groups. The methods are tested on data from patients with Systemic Sclerosis (SSc), a rare potentially fatal condition. Results show that the "Latent Class Multi-Label Classification Model" improves accuracy when compared with competitive similar methods.
Awad Alsaid Alyousef, Svetlana I. Nihtyanova, Christopher P. Denton, Pietro Bosoni, Riccardo Bellazzi, Allan Tucker
CBMS5
2019 Transfer Learning for Urban Landscape Clustering and Correlation with Health Indexes
abstract
Within the EU-funded Pulse project, we are implementing a data analytic platform designed to provide public health decision makers with advanced approaches to jointly analyze maps and geospatial information with health care data and air pollution measurements. In this paper we describe a component of such platform, designed to couple deep learning analysis of geospatial images of cities and some healthcare and behavioral indexes collected by the 500 cities US project, showing that, in New York City, urban landscape significantly correlates with the access to healthcare services.
Riccardo Bellazzi, Alessandro Aldo Caldarone, Daniele Pala, Marica Franzini, Alberto Malovini, Cristiana Larizza, Vittorio Casella
ICOST1
2019 Supervised methods to extract clinical events from cardiology reports in Italian
Natalia Viani, Timothy A. Miller, Carlo Napolitano, Silvia G. Priori, Guergana K. Savova, Riccardo Bellazzi, Lucia Sacchi
J. Biomed. Informatics6
2018 Predicting Disease Complications Using a Stepwise Hidden Variable Approach for Learning Dynamic Bayesian Networks
abstract
Predicting Diabetes Type 2 Mellitus (T2DM) complications such as retinopathy and liver disease is still a challenge despite being a growing public health concern worldwide. This is due to the complex interactions between complications and other features, as well as between the different complications, themselves. What is more, there are likely to be many unmeasured effects that impact the disease progression of different patients. Probabilistic graphical models such as Dynamic Bayesian Networks (DBNs) have demonstrated much promise in the modeling of disease progression and they can naturally incorporate hidden (latent) variables using the EM algorithm. Unlike deep learning approaches that attempt to model complex interactions in data by using a large number of hidden variables, we adopt a different approach. We are interested in models that not only capture unmeasured effects but are also transparent in how they model data so that knowledge about disease processes can be extracted and trust in the model can be maintained by clinicians. As a result, we have developed a step-wise hidden variable structure learning process that incrementally adds hidden variables based on the IC* algorithm. To the best of our knowledge, this is the first study for classifying disease complication using a step-wise learning methodology for identifying hidden and T2DM features with a DBN structure from clinical data. Our extensive set of experiments show that the proposed method improves classification accuracy, identifying the correct number of hidden variables, and targeting their precise location within the network structure.
Leila Yousefi, Allan Tucker, Mashael Al-Luhaybi, Lucia Sacchi, Riccardo Bellazzi, Luca Chiovato
CBMS5
2018 A dashboard-based system for supporting diabetes care
abstract
Objective: To describe the development, as part of the European Union MOSAIC (Models and Simulation Techniques for Discovering Diabetes Influence Factors) project, of a dashboard-based system for the management of type 2 diabetes and assess its impact on clinical practice. Methods: The MOSAIC dashboard system is based on predictive modeling, longitudinal data analytics, and the reuse and integration of data from hospitals and public health repositories. Data are merged into an i2b2 data warehouse, which feeds a set of advanced temporal analytic models, including temporal abstractions, care-flow mining, drug exposure pattern detection, and risk-prediction models for type 2 diabetes complications. The dashboard has 2 components, designed for (1) clinical decision support during follow-up consultations and (2) outcome assessment on populations of interest. To assess the impact of the clinical decision support component, a pre-post study was conducted considering visit duration, number of screening examinations, and lifestyle interventions. A pilot sample of 700 Italian patients was investigated. Judgments on the outcome assessment component were obtained via focus groups with clinicians and health care managers. Results: The use of the decision support component in clinical activities produced a reduction in visit duration (P ≪ .01) and an increase in the number of screening exams for complications (P < .01). We also observed a relevant, although nonstatistically significant, increase in the proportion of patients receiving lifestyle interventions (from 69% to 77%). Regarding the outcome assessment component, focus groups highlighted the system's capability of identifying and understanding the characteristics of patient subgroups treated at the center. Conclusion: Our study demonstrates that decision support tools based on the integration of multiple-source data and visual and predictive analytics do improve the management of a chronic disease such as type 2 diabetes by enacting a successful implementation of the learning health care system cycle.
Arianna Dagliati, Lucia Sacchi, Valentina Tibollo, Giulia Cogni, Marsida Teliti, Antonio Martinez-Millana, Vicente Traver 0001, Daniele Segagni, Manuel Ottaviano, Giuseppe Fico, María Teresa Arredondo, Pasquale De Cata, Luca Chiovato, Riccardo Bellazzi
J. Am. Medical Informatics Assoc.15
2018 Incorporating repeating temporal association rules in Naïve Bayes classifiers for coronary heart disease diagnosis
Kalia Orphanou, Arianna Dagliati, Lucia Sacchi, Athena Stassopoulou, Elpida T. Keravnou, Riccardo Bellazzi
J. Biomed. Informatics6
2018 Patient similarity for precision medicine: A systematic review
Enea Parimbelli, Simone Marini, Lucia Sacchi, Riccardo Bellazzi
J. Biomed. Informatics4
2017 Data Fusion Approach for Learning Transcriptional Bayesian Networks
Elisabetta Sauta, Andrea Demartini, Francesca Vitali, Alberto Riva, Riccardo Bellazzi
AIME5
2017 Recurrent Neural Network Architectures for Event Extraction from Italian Medical Reports
Natalia Viani, Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi, Lucia Sacchi, Guergana K. Savova
AIME7
2017 Predicting Comorbidities Using Resampling and Dynamic Bayesian Networks with Latent Variables
abstract
Comorbidities such as hypertension and lipid metabolism are often associated in diseases such as diabetes, and the early prediction of these is of great value when trying to manage progression. This is the start of a project to model multiple comorbidities in diabetes using dynamic Bayesian networks with latent variables in order to stratify patient cohorts. In this paper, we demonstrate some initial results on a dataset where the class imbalance problem poses an issue due to the rare occurrence of different individual comorbidities on a visit-by-visit basis. This is dealt with using a bootstrap technique that has been specifically designed for longitudinal data where the occurrence of the positive class occurs far less than the negative.
Leila Yousefi, Lucia Sacchi, Riccardo Bellazzi, Luca Chiovato, Allan Tucker
CBMS3
2017 Artificial Intelligence in Medicine AIME 2015
John H. Holmes, Lucia Sacchi, Riccardo Bellazzi, Niels Peek
Artif. Intell. Medicine3
2017 Temporal electronic phenotyping by mining careflows of breast cancer patients
Arianna Dagliati, Lucia Sacchi, Alberto Zambelli, Valentina Tibollo, L. Pavesi, John H. Holmes, Riccardo Bellazzi
J. Biomed. Informatics7
2016 Hierarchical Bayesian Logistic Regression to forecast metabolic control in type 2 DM patients
Arianna Dagliati, Alberto Malovini, Pasquale De Cata, Giulia Cogni, Marsida Teliti, Lucia Sacchi, Carlo Cerra, Luca Chiovato, Riccardo Bellazzi
AMIA9
2016 Information Extraction from Italian medical reports: first steps towards clinical timelines development
Natalia Viani, Valentina Tibollo, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi, Cristiana Larizza, Lucia Sacchi
AMIA5
2016 Combining Unsupervised and Supervised Learning for Discovering Disease Subclasses
abstract
Diseases are often umbrella terms for many subcategories of disease. The identification of these subcategories is vital if we are to develop personalised treatments that are better focussed on individual patients. In this short paper, we explore the use of a combination of unsupervised learning to identify potential subclasses, and supervised learning to build models for better predicting a number of different health outcomes for patients that suffer from systemic sclerosis, a rare chronic connective tissue disorder - but one that shares many characteristics with other diseases. We explore a number of different algorithms for constructing models that simultaneously predict health outcomes and identify subcategories.
Pietro Bosoni, Allan Tucker, Riccardo Bellazzi, Svetlana I. Nihtyanova, Christopher P. Denton
CBMS3
2016 Out-of-Home Activity Recognition from GPS Data in Schizophrenic Patients
abstract
Risk of psychotic relapse in schizophrenic patients is commonly measured by social functioning (SF), which focuses on patients' daily activities. Monitoring of SF usually relies on infrequent clinic visits, limiting the capacity to detect sudden changes. GPS data that is passively collected with smartphones introduce new opportunities to monitor SF. We conducted a five-day pilot study with five schizophrenic patients to assess the feasibility of this approach. Participants used a smartphone to continuously record their GPS location, and completed a paper-based SF diary to register out-of-home activities. We implemented a time-based method and a density-based method to identify the geolocations visited and then we clustered geolocations visited in places visited. Finally, we used semantic enrichment to classify places types and associated activities. We evaluated the performance of the two approaches by comparing the activities detected from the GPS data with those recorded in the SF diary. Recall was better for the density-based method, ranging from 0.686 (Standard Deviation [SD] 0.168) to 0.771 (SD 0.264) while precision was better for the time-based method (0.722 (SD 0.197) to 0.954 (SD 0.093)). To conclude, using routinely collected GPS data and relatively simple analytical methods we detected patients' out-of-home activities with moderate recall, more sophisticated analytical methods may obtain better performance.
Sonia Difrancesco, Paolo Fraccaro, Sabine van der Veer, Bader Alshoumr, John D. Ainsworth, Riccardo Bellazzi, Niels Peek
CBMS6
2016 A computational method for designing diverse linear epitopes including citrullinated peptides with desired binding affinities to intravenous immunoglobulin
abstract
BACKGROUND: Understanding the interactions between antibodies and the linear epitopes that they recognize is an important task in the study of immunological diseases. We present a novel computational method for the design of linear epitopes of specified binding affinity to Intravenous Immunoglobulin (IVIg). RESULTS: We show that the method, called Pythia-design can accurately design peptides with both high-binding affinity and low binding affinity to IVIg. To show this, we experimentally constructed and tested the computationally constructed designs. We further show experimentally that these designed peptides are more accurate that those produced by a recent method for the same task. Pythia-design is based on combining random walks with an ensemble of probabilistic support vector machines (SVM) classifiers, and we show that it produces a diverse set of designed peptides, an important property to develop robust sets of candidates for construction. We show that by combining Pythia-design and the method of (PloS ONE 6(8):23616, 2011), we are able to produce an even more accurate collection of designed peptides. Analysis of the experimental validation of Pythia-design peptides indicates that binding of IVIg is favored by epitopes that contain trypthophan and cysteine. CONCLUSIONS: Our method, Pythia-design, is able to generate a diverse set of binding and non-binding peptides, and its designs have been experimentally shown to be accurate.
Rob Patro, Raquel Norel, Robert J. Prill, Julio Saez-Rodriguez, Peter Lorenz, Felix Steinbeck, Bjoern Ziems, Mitja Lustrek, Nicola Barbarini, Alessandra Tiengo, Riccardo Bellazzi, Hans-Jürgen Thiesen, Gustavo Stolovitzky, Carl Kingsford
BMC Bioinform.11
2016 Guest Editorial IEEE EMBC 2015
abstract
The papers in this special issue were presented at the 37th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2015), which was held in Milano, Italy, from August 25-29.
Elsa D. Angelini, Riccardo Bellazzi, Walter G. Besio, Marius George Linguraru
IEEE J. Biomed. Health Informatics2
2015 Comparison of Probabilistic versus Non-probabilistic Electronic Nose Classification Methods in an Animal Model
Camilla Colombo, Jan Hendrik Leopold, Lieuwe D. J. Bos, Riccardo Bellazzi, Ameen Abu-Hanna
AIME4
2015 Running Genome Wide Data Analysis Using a Parallel Approach on a Cloud Platform
Andrea Demartini, Davide Capozzi, Alberto Malovini, Riccardo Bellazzi
AIME4
2015 A Genomic Data Fusion Framework to Exploit Rare and Common Variants for Association Discovery
Simone Marini, Ivan Limongelli, Ettore Rizzo, Tan Da, Riccardo Bellazzi
AIME5
2015 Collaborative Filtering for Estimating Health Related Utilities in Decision Support Systems
Enea Parimbelli, Silvana Quaglini, Riccardo Bellazzi, John H. Holmes
AIME3
2015 Inferring air quality maps from remotely sensed data to exploit georeferenced clinical onsets: The Pavia 2013 case
abstract
Recent developments in data acquisition, storage, mining and maintenance have allowed the flourishing of several multi-disciplinary research fields, which can be stated, defined and carried out according to the so-called Big Data paradigm. In this environment, the investigation and analysis of interactions between human phenomena and natural events play a key-role, as they can be fundamental for several applications, from sustainable development to community policy design and short-, medium- and long-range resource allocation planning. In this paper, we provide a study of the interplay between air pollution (as estimated by remotely sensed data processing) and clinical records, so that inferences and correlations among black particulate concentration, micro- and macro-vascular disease onsets and hospitalization tracks can be efficiently drawn. We focused on the second order administrative area of the city of Pavia, Italy, on 2013. Experimental results show how effective connections between the estimated air quality and the hospitalizations behavior can be accurately drawn and derived.
Andrea Marinoni, Arianna Dagliati, Riccardo Bellazzi, Paolo Gamba
IGARSS3
2015 Thirty years of artificial intelligence in medicine (AIME) conferences: A review of research themes
Niels Peek, Carlo Combi, Roque Marín, Riccardo Bellazzi
Artif. Intell. Medicine4
2015 BigQ: a NoSQL based framework to handle genomic variants in i2b2
abstract
BACKGROUND: Precision medicine requires the tight integration of clinical and molecular data. To this end, it is mandatory to define proper technological solutions able to manage the overwhelming amount of high throughput genomic data needed to test associations between genomic signatures and human phenotypes. The i2b2 Center (Informatics for Integrating Biology and the Bedside) has developed a widely internationally adopted framework to use existing clinical data for discovery research that can help the definition of precision medicine interventions when coupled with genetic data. i2b2 can be significantly advanced by designing efficient management solutions of Next Generation Sequencing data. RESULTS: We developed BigQ, an extension of the i2b2 framework, which integrates patient clinical phenotypes with genomic variant profiles generated by Next Generation Sequencing. A visual programming i2b2 plugin allows retrieving variants belonging to the patients in a cohort by applying filters on genomic variant annotations. We report an evaluation of the query performance of our system on more than 11 million variants, showing that the implemented solution scales linearly in terms of query time and disk space with the number of variants. CONCLUSIONS: In this paper we describe a new i2b2 web service composed of an efficient and scalable document-based database that manages annotations of genomic variants and of a visual programming plug-in designed to dynamically perform queries on clinical and genetic data. The system therefore allows managing the fast growing volume of genomic variants and can be used to integrate heterogeneous genomic annotations.
Matteo Gabetta, Ivan Limongelli, Ettore Rizzo, Alberto Riva, Daniele Segagni, Riccardo Bellazzi
BMC Bioinform.6
2015 PaPI: pseudo amino acid composition to score human protein-coding variants
abstract
BACKGROUND: High throughput sequencing technologies are able to identify the whole genomic variation of an individual. Gene-targeted and whole-exome experiments are mainly focused on coding sequence variants related to a single or multiple nucleotides. The analysis of the biological significance of this multitude of genomic variant is challenging and computational demanding. RESULTS: We present PaPI, a new machine-learning approach to classify and score human coding variants by estimating the probability to damage their protein-related function. The novelty of this approach consists in using pseudo amino acid composition through which wild and mutated protein sequences are represented in a discrete model. A machine learning classifier has been trained on a set of known deleterious and benign coding variants with the aim to score unobserved variants by taking into account hidden sequence patterns in human genome potentially leading to diseases. We show how the combination of amphiphilic pseudo amino acid composition, evolutionary conservation and homologous proteins based methods outperforms several prediction algorithms and it is also able to score complex variants such as deletions, insertions and indels. CONCLUSIONS: This paper describes a machine-learning approach to predict the deleteriousness of human coding variants. A freely available web application (http://papi.unipv.it) has been developed with the presented method, able to score up to thousands variants in a single run.
Ivan Limongelli, Simone Marini, Riccardo Bellazzi
BMC Bioinform.3
2015 Optimal marker placement in hadrontherapy: Intelligent optimization strategies with augmented Lagrangian pattern search
Cristina Altomare, Raffaella Guglielmann, Marco Riboldi, Riccardo Bellazzi, Guido Baroni
J. Biomed. Informatics4
2015 A Dynamic Bayesian Network model for long-term simulation of clinical complications in type 1 diabetes
Simone Marini, Emanuele Trifoglio, Nicola Barbarini, Francesco Sambo, Barbara Di Camillo, Alberto Malovini, Marco Manfrini, Claudio Cobelli, Riccardo Bellazzi
J. Biomed. Informatics9
2015 A kinetic model-based algorithm to classify NGS short reads by their allele origin
Andrea Marinoni, Ettore Rizzo, Ivan Limongelli, Paolo Gamba, Riccardo Bellazzi
J. Biomed. Informatics5
2014 Handling Clinical and Next Generation Sequencing data: new strategies in i2b2 and tranSMART
Shawn N. Murphy, Riccardo Bellazzi, Matteo Gabetta, Paul Avillach, Lori C. Phillips
AMIA2
2014 Exposome informatics: considerations for the design of future biomedical research information systems
abstract
The environment's contribution to health has been conceptualized as the exposome. Biomedical research interest in environmental exposures as a determinant of physiopathological processes is rising as such data increasingly become available. The panoply of miniaturized sensing devices now accessible and affordable for individuals to use to monitor a widening range of parameters opens up a new world of research data. Biomedical informatics (BMI) must provide a coherent framework for dealing with multi-scale population data including the phenome, the genome, the exposome, and their interconnections. The combination of these more continuous, comprehensive, and personalized data sources requires new research and development approaches to data management, analysis, and visualization. This article analyzes the implications of a new paradigm for the discipline of BMI, one that recognizes genome, phenome, and exposome data and their intricate interactions as the basis for biomedical research now and for clinical care in the near future.
Fernando Martín-Sánchez, Kathleen Gray, Riccardo Bellazzi, Guillermo López-Campos
J. Am. Medical Informatics Assoc.3
2013 Knowledge-Based Identification of Multicomponent Therapies
Francesca Vitali, Francesca Mulas, Pietro Marini, Riccardo Bellazzi
AIME4
2013 Biomedical and Healthcare Analytics on Big Data
Niels Peek, Jimeng Sun 0001, John H. Holmes, Fernando Martín-Sánchez, Riccardo Bellazzi
AMIA5
2013 Mining Careflow Patterns in data warehouses of breast cancer patients
Lucia Sacchi, Daniele Segagni, Arianna Dagliati, Alberto Zambelli, Riccardo Bellazzi
AMIA5
2013 Network-based target ranking for polypharmacological therapies
Francesca Vitali, Francesca Mulas, Pietro Marini, Riccardo Bellazzi
J. Biomed. Informatics4
2012 Report From European Summit On Trustworthy Reuse Of Health Data
Charles Safran, Antoine Geissbühler, Riccardo Bellazzi, Iain E. Buchan, Steven E. Labkoff
AMIA3
2012 Clinical Bioinformatics: challenges and opportunities
abstract
BACKGROUND: Network Tools and Applications in Biology (NETTAB) Workshops are a series of meetings focused on the most promising and innovative ICT tools and to their usefulness in Bioinformatics. The NETTAB 2011 workshop, held in Pavia, Italy, in October 2011 was aimed at presenting some of the most relevant methods, tools and infrastructures that are nowadays available for Clinical Bioinformatics (CBI), the research field that deals with clinical applications of bioinformatics. METHODS: In this editorial, the viewpoints and opinions of three world CBI leaders, who have been invited to participate in a panel discussion of the NETTAB workshop on the next challenges and future opportunities of this field, are reported. These include the development of data warehouses and ICT infrastructures for data sharing, the definition of standards for sharing phenotypic data and the implementation of novel tools to implement efficient search computing solutions. RESULTS: Some of the most important design features of a CBI-ICT infrastructure are presented, including data warehousing, modularity and flexibility, open-source development, semantic interoperability, integrated search and retrieval of -omics information. CONCLUSIONS: Clinical Bioinformatics goals are ambitious. Many factors, including the availability of high-throughput "-omics" technologies and equipment, the widespread availability of clinical data warehouses and the noteworthy increase in data storage and computational power of the most recent ICT systems, justify research and efforts in this domain, which promises to be a crucial leveraging factor for biomedical research.
Riccardo Bellazzi, Marco Masseroli, Shawn N. Murphy, Amnon Shabo, Paolo Romano 0001
BMC Bioinform.1
2012 Hierarchical Naive Bayes for genetic association studies
abstract
BACKGROUND: Genome Wide Association Studies represent powerful approaches that aim at disentangling the genetic and molecular mechanisms underlying complex traits. The usual "one-SNP-at-the-time" testing strategy cannot capture the multi-factorial nature of this kind of disorders. We propose a Hierarchical Naïve Bayes classification model for taking into account associations in SNPs data characterized by Linkage Disequilibrium. Validation shows that our model reaches classification performances superior to those obtained by the standard Naïve Bayes classifier for simulated and real datasets. METHODS: In the Hierarchical Naïve Bayes implemented, the SNPs mapping to the same region of Linkage Disequilibrium are considered as "details" or "replicates" of the locus, each contributing to the overall effect of the region on the phenotype. A latent variable for each block, which models the "population" of correlated SNPs, can be then used to summarize the available information. The classification is thus performed relying on the latent variables conditional probability distributions and on the SNPs data available. RESULTS: The developed methodology has been tested on simulated datasets, each composed by 300 cases, 300 controls and a variable number of SNPs. Our approach has been also applied to two real datasets on the genetic bases of Type 1 Diabetes and Type 2 Diabetes generated by the Wellcome Trust Case Control Consortium. CONCLUSIONS: The approach proposed in this paper, called Hierarchical Naïve Bayes, allows dealing with classification of examples for which genetic information of structurally correlated SNPs are available. It improves the Naïve Bayes performances by properly handling the within-loci variability.
Alberto Malovini, Nicola Barbarini, Riccardo Bellazzi, Francesca Demichelis
BMC Bioinform.3
2012 An ICT infrastructure to integrate clinical and molecular data in oncology research
abstract
BACKGROUND: The ONCO-i2b2 platform is a bioinformatics tool designed to integrate clinical and research data and support translational research in oncology. It is implemented by the University of Pavia and the IRCCS Fondazione Maugeri hospital (FSM), and grounded on the software developed by the Informatics for Integrating Biology and the Bedside (i2b2) research center. I2b2 has delivered an open source suite based on a data warehouse, which is efficiently interrogated to find sets of interesting patients through a query tool interface. METHODS: Onco-i2b2 integrates data coming from multiple sources and allows the users to jointly query them. I2b2 data are then stored in a data warehouse, where facts are hierarchically structured as ontologies. Onco-i2b2 gathers data from the FSM pathology unit (PU) database and from the hospital biobank and merges them with the clinical information from the hospital information system. Our main effort was to provide a robust integrated research environment, giving a particular emphasis to the integration process and facing different challenges, consecutively listed: biospecimen samples privacy and anonymization; synchronization of the biobank database with the i2b2 data warehouse through a series of Extract, Transform, Load (ETL) operations; development and integration of a Natural Language Processing (NLP) module, to retrieve coded information, such as SNOMED terms and malignant tumors (TNM) classifications, and clinical tests results from unstructured medical records. Furthermore, we have developed an internal SNOMED ontology rested on the NCBO BioPortal web services. RESULTS: Onco-i2b2 manages data of more than 6,500 patients with breast cancer diagnosis collected between 2001 and 2011 (over 390 of them have at least one biological sample in the cancer biobank), more than 47,000 visits and 96,000 observations over 960 medical concepts. CONCLUSIONS: Onco-i2b2 is a concrete example of how integrated Information and Communication Technology architecture can be implemented to support translational research. The next steps of our project will involve the extension of its capabilities by implementing new plug-in devoted to bioinformatics data analysis as well as a temporal query module.
Daniele Segagni, Valentina Tibollo, Arianna Dagliati, Alberto Zambelli, Silvia G. Priori, Riccardo Bellazzi
BMC Bioinform.6
2012 Stochastic model search with binary outcomes for genome-wide association studies
abstract
OBJECTIVE: The spread of case-control genome-wide association studies (GWASs) has stimulated the development of new variable selection methods and predictive models. We introduce a novel Bayesian model search algorithm, Binary Outcome Stochastic Search (BOSS), which addresses the model selection problem when the number of predictors far exceeds the number of binary responses. MATERIALS AND METHODS: Our method is based on a latent variable model that links the observed outcomes to the underlying genetic variables. A Markov Chain Monte Carlo approach is used for model search and to evaluate the posterior probability of each predictor. RESULTS: BOSS is compared with three established methods (stepwise regression, logistic lasso, and elastic net) in a simulated benchmark. Two real case studies are also investigated: a GWAS on the genetic bases of longevity, and the type 2 diabetes study from the Wellcome Trust Case Control Consortium. Simulations show that BOSS achieves higher precisions than the reference methods while preserving good recall rates. In both experimental studies, BOSS successfully detects genetic polymorphisms previously reported to be associated with the analyzed phenotypes. DISCUSSION: BOSS outperforms the other methods in terms of F-measure on simulated data. In the two real studies, BOSS successfully detects biologically relevant features, some of which are missed by univariate analysis and the three reference techniques. CONCLUSION: The proposed algorithm is an advance in the methodology for model selection with a large number of features. Our simulated and experimental results showed that BOSS proves effective in detecting relevant markers while providing a parsimonious model.
Alberto Russu, Alberto Malovini, Annibale A. Puca, Riccardo Bellazzi
J. Am. Medical Informatics Assoc.4
2011 A Data Mining Library for miRNA Annotation and Analysis
Angelo Nuzzo, Riccardo Beretta, Francesca Mulas, Valerie Roobrouck, Catherine M. Verfaillie, Blaz Zupan, Riccardo Bellazzi
AIME7
2011 Ranking and 1-Dimensional Projection of Cell Development Transcription Profiles
Lan Zagar, Francesca Mulas, Riccardo Bellazzi, Blaz Zupan
AIME3
2011 Mario Stefanelli, 1945-2010
Riccardo Bellazzi, Carlo Combi, Silvana Quaglini
Artif. Intell. Medicine1
2011 Stage prediction of embryonic stem cell differentiation from genome-wide expression data
abstract
MOTIVATION: The developmental stage of a cell can be determined by cellular morphology or various other observable indicators. Such classical markers could be complemented with modern surrogates, like whole-genome transcription profiles, that can encode the state of the entire organism and provide increased quantitative resolution. Recent findings suggest that such profiles provide sufficient information to reliably predict the cell's developmental stage. RESULTS: We use whole-genome transcription data and several data projection methods to infer differentiation stage prediction models for embryonic cells. Given a transcription profile of an uncharacterized cell, these models can then predict its developmental stage. In a series of experiments comprising 14 datasets from the Gene Expression Omnibus, we demonstrate that the approach is robust and has excellent prediction ability both within a specific cell line and across different cell lines. AVAILABILITY: Model inference and computational evaluation procedures in the form of Python scripts and accompanying datasets are available at http://www.biolab.si/supp/stagerank. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lan Zagar, Francesca Mulas, Silvia Garagna, Maurizio Zuccotti, Riccardo Bellazzi, Blaz Zupan
Bioinform.5
2011 R Engine Cell: integrating R into the i2b2 software infrastructure
abstract
Informatics for Integrating Biology and the Bedside (i2b2) is an initiative funded by the NIH that aims at building an informatics infrastructure to support biomedical research. The University of Pavia has recently integrated i2b2 infrastructure with a registry of inherited arrhythmogenic diseases. Within this project, the authors created a novel i2b2 cell, named R Engine Cell, which allows the communication between i2b2 and the R statistical software. As survival analyses are routinely performed by cardiology researchers, the authors have first concentrated on making Kaplan-Meier analyses available within the i2b2 web interface. To this aim, the authors developed a web-client plug-in to select the patient set on which to perform the analysis and to display the results in a graphical, intuitive way. R Engine Cell has been designed to easily support the integration of other R-based statistical analyses into i2b2.
Daniele Segagni, Fulvia Ferrazzi, Cristiana Larizza, Valentina Tibollo, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi
J. Am. Medical Informatics Assoc.7
2011 Inferring cell cycle feedback regulation from gene expression data
Fulvia Ferrazzi, Felix B. Engel, Erxi Wu, Annie P. Moseman, Isaac S. Kohane, Riccardo Bellazzi, Marco Ramoni
J. Biomed. Informatics6
2011 Computer-based genealogy reconstruction in founder populations
Giuseppe Milani, Corrado Masciullo, Cinzia Sala, Riccardo Bellazzi, Iwan Buetti, Giorgio Pistis, Michela Traglia, Daniela Toniolo, Cristiana Larizza
J. Biomed. Informatics4
2011 Corrigendum to "In Memoriam: 'Professor Mario Stefanelli (1945-2010)'" [J Biomed Inform (2010) 859-60]
Vimla L. Patel, Riccardo Bellazzi, Silvana Quaglini
J. Biomed. Informatics2
2010 Translational Bioinformatics: Challenges and Opportunities for Case-Based Reasoning and Decision Support
Riccardo Bellazzi, Cristiana Larizza, Matteo Gabetta, Giuseppe Milani, Angelo Nuzzo, Valentina Favalli, Eloisa Arbustini
ICCBR1
2010 Professor Mario Stefanelli (1945-2010)
Vimla L. Patel, Riccardo Bellazzi, Silvana Quaglini
J. Biomed. Informatics2
2010 An automated reasoning framework for translational research
Alberto Riva, Angelo Nuzzo, Mario Stefanelli, Riccardo Bellazzi
J. Biomed. Informatics4
2009 Temporal Data Mining of HIV Registries: Results from a 25 Years Follow-Up
Paloma Chausa, César Cáceres, Lucia Sacchi, Agathe León, Felipe García, Riccardo Bellazzi, Enrique J. Gómez
AIME6
2009 Mining Healthcare Data with Temporal Association Rules: Improvements and Assessment for a Practical Use
Stefano Concaro, Lucia Sacchi, Carlo Cerra, Pietro Fratino, Riccardo Bellazzi
AIME5
2009 On Quality of Different Annotation Sources for Gene Expression Analysis
Francesca Mulas, Tomaz Curk, Riccardo Bellazzi, Blaz Zupan
AIME3
2009 An Architecture for Automated Reasoning Systems for Genome-Wide Studies
Angelo Nuzzo, Alberto Riva, Mario Stefanelli, Riccardo Bellazzi
AIME4
2009 A Temporal Abstraction Framework for Classifying Clinical Temporal Data
Iyad Batal, Lucia Sacchi, Riccardo Bellazzi, Milos Hauskrecht
AMIA3
2009 Temporal Data Mining for the Assessment of the Costs Related to Diabetes Mellitus Pharmacological Treatment
Stefano Concaro, Lucia Sacchi, Carlo Cerra, Mario Stefanelli, Pietro Fratino, Riccardo Bellazzi
AMIA6
2009 Artificial intelligence in medicine AIME'07
Riccardo Bellazzi, Ameen Abu-Hanna
Artif. Intell. Medicine1
2009 The coming of age of artificial intelligence in medicine
Vimla L. Patel, Edward H. Shortliffe, Mario Stefanelli, Peter Szolovits, Michael R. Berthold, Riccardo Bellazzi, Ameen Abu-Hanna
Artif. Intell. Medicine6
2009 Phenotype forecasting with SNPs data through gene-based Bayesian networks
abstract
BACKGROUND: Bayesian networks are powerful instruments to learn genetic models from association studies data. They are able to derive the existing correlation between genetic markers and phenotypic traits and, at the same time, to find the relationships between the markers themselves. However, learning Bayesian networks is often non-trivial due to the high number of variables to be taken into account in the model with respect to the instances of the dataset. Therefore, it becomes very interesting to use an abstraction of the variable space that suitably reduces its dimensionality without losing information. In this paper we present a new strategy to achieve this goal by mapping the SNPs related to the same gene to one meta-variable. In order to assign states to the meta-variables we employ an approach based on classification trees. RESULTS: We applied our approach to data coming from a genome-wide scan on 288 individuals affected by arterial hypertension and 271 nonagenarians without history of hypertension. After pre-processing, we focused on a subset of 24 SNPs. We compared the performance of the proposed approach with the Bayesian network learned with SNPs as variables and with the network learned with haplotypes as meta-variables. The results were obtained by running a hold-out experiment five times. The mean accuracy of the new method was 64.28%, while the mean accuracy of the SNPs network was 58.99% and the mean accuracy of the haplotype network was 54.57%. CONCLUSION: The new approach presented in this paper is able to derive a gene-based predictive model based on SNPs data. Such model is more parsimonious than the one based on single SNPs, while preserving the capability of highlighting predictive SNPs configurations. The prediction performance of this approach was consistently superior to the SNP-based and the haplotype-based one in all the test sets of the evaluation procedure. The method can be then considered as an alternative way to analyze the data coming from association studies.
Alberto Malovini, Angelo Nuzzo, Fulvia Ferrazzi, Annibale A. Puca, Riccardo Bellazzi
BMC Bioinform.5
2009 Phenotypic and genotypic data integration and exploration through a web-service architecture
abstract
BACKGROUND: Linking genotypic and phenotypic information is one of the greatest challenges of current genetics research. The definition of an Information Technology infrastructure to support this kind of studies, and in particular studies aimed at the analysis of complex traits, which require the definition of multifaceted phenotypes and the integration genotypic information to discover the most prevalent diseases, is a paradigmatic goal of Biomedical Informatics. This paper describes the use of Information Technology methods and tools to develop a system for the management, inspection and integration of phenotypic and genotypic data. RESULTS: We present the design and architecture of the Phenotype Miner, a software system able to flexibly manage phenotypic information, and its extended functionalities to retrieve genotype information from external repositories and to relate it to phenotypic data. For this purpose we developed a module to allow customized data upload by the user and a SOAP-based communications layer to retrieve data from existing biomedical knowledge management tools. In this paper we also demonstrate the system functionality by an example application of the system in which we analyze two related genomic datasets. CONCLUSION: In this paper we show how a comprehensive, integrated and automated workbench for genotype and phenotype integration can facilitate and improve the hypothesis generation process underlying modern genetic studies.
Angelo Nuzzo, Alberto Riva, Riccardo Bellazzi
BMC Bioinform.3
2008 Analysis and Visualization of Spatial Proteomic Data for Tissue Characterization
abstract
Spatial proteomic profiling of tissue sections provides in situ molecular analysis of proteins and peptides. Analysis and visualization of these high-dimensional data cubes is challenging. We present a methodology for this task based on a novel developed algorithm for the feature identification and reduction step. To show the validity of our approach, we analyzed prostate cancer tissue sections with an adapted kernel-density based clustering algorithm.
Christian Fuchsberger, Heidi Hübl, Georg Schäfer, Alexandre Pelzer, Georg Bartsch, Helmut Klocker, Nicola Barbarini, Riccardo Bellazzi, Wolfgang Wieder, Günther Bonn
CBMS8
2008 TimeClust: a clustering tool for gene expression time series
abstract
Abstract Summary: TimeClust is a user-friendly software package to cluster genes according to their temporal expression profiles. It can be conveniently used to analyze data obtained from DNA microarray time-course experiments. It implements two original algorithms specifically designed for clustering short time series together with hierarchical clustering and self-organizing maps. Availability: TimeClust executable files for Windows and LINUX platforms can be downloaded free of charge for non-profit institutions from the following web site: http://aimed11.unipv.it/TimeClust. Contact: [email protected] or for software support [email protected] Supplementary information: A simple user's guide (example.pdf) is available in the download area together with two trial data sets.
Paolo Magni, Fulvia Ferrazzi, Lucia Sacchi, Riccardo Bellazzi
Bioinform.4
2008 Building a Normative Decision Support System for Clinical and Operational Risk Management in Hemodialysis
abstract
This paper describes the design and implementation of a decision support system for risk management in hemodialysis (HD) departments. The proposed system exploits a domain ontology to formalize the problem as a Bayesian network. It also relies on a software tool, able to automatically collect HD data, to learn the network conditional probabilities. By merging prior knowledge and the available data, the system allows to estimate risk profiles both for patients and HD departments. The risk management process is completed by an influence diagram that enables scenario analysis to choose the optimal decisions that mitigate a patient's risk. The methods and design of the decision support tool are described in detail, and the derived decision model is presented. Examples and case studies are also shown. The tool is one of the few examples of normative system explicitly conceived to manage operational and clinical risks in health care environments.
Chiara Cornalba, Roberto G. Bellazzi, Riccardo Bellazzi
IEEE Trans. Inf. Technol. Biomed.3
2007 Temporal abstraction for feature extraction: A comparative case study in prediction from intensive care monitoring data
Marion Verduijn, Lucia Sacchi, Niels Peek, Riccardo Bellazzi, Evert de Jonge, Bas A. de Mol
Artif. Intell. Medicine4
2007 A procedure to decompose high resolution mass spectra
Nicola Barbarini, Paolo Magni, Riccardo Bellazzi
BMC Bioinform.3
2007 Bayesian approaches to reverse engineer cellular systems: a simulation study on nonlinear Gaussian networks
abstract
BACKGROUND: Reverse engineering cellular networks is currently one of the most challenging problems in systems biology. Dynamic Bayesian networks (DBNs) seem to be particularly suitable for inferring relationships between cellular variables from the analysis of time series measurements of mRNA or protein concentrations. As evaluating inference results on a real dataset is controversial, the use of simulated data has been proposed. However, DBN approaches that use continuous variables, thus avoiding the information loss associated with discretization, have not yet been extensively assessed, and most of the proposed approaches have dealt with linear Gaussian models. RESULTS: We propose a generalization of dynamic Gaussian networks to accommodate nonlinear dependencies between variables. As a benchmark dataset to test the new approach, we used data from a mathematical model of cell cycle control in budding yeast that realistically reproduces the complexity of a cellular system. We evaluated the ability of the networks to describe the dynamics of cellular systems and their precision in reconstructing the true underlying causal relationships between variables. We also tested the robustness of the results by analyzing the effect of noise on the data, and the impact of a different sampling time. CONCLUSION: The results confirmed that DBNs with Gaussian models can be effectively exploited for a first level analysis of data from complex cellular systems. The inferred models are parsimonious and have a satisfying goodness of fit. Furthermore, the networks not only offer a phenomenological description of the dynamics of cellular systems, but are also able to suggest hypotheses concerning the causal interactions between variables. The proposed nonlinear generalization of Gaussian models yielded models characterized by a slightly lower goodness of fit than the linear model, but a better ability to recover the true underlying connections between variables.
Fulvia Ferrazzi, Paola Sebastiani, Marco Ramoni, Riccardo Bellazzi
BMC Bioinform.4
2007 Data mining with Temporal Abstractions: learning rules from time series
Lucia Sacchi, Cristiana Larizza, Carlo Combi, Riccardo Bellazzi
Data Min. Knowl. Discov.4
2007 Towards knowledge-based gene expression data mining
Riccardo Bellazzi, Blaz Zupan
J. Biomed. Informatics1
2007 Precedence Temporal Networks to represent temporal relationships in gene expression data
Lucia Sacchi, Cristiana Larizza, Paolo Magni, Riccardo Bellazzi
J. Biomed. Informatics4
2006 A New Approach for the Analysis of Mass Spectrometry Data for Biomarker Discovery
Nicola Barbarini, Paolo Magni, Riccardo Bellazzi
AMIA3
2006 Dynamic Bayesian Networks in Modelling Cellular Systems: a Critical Appraisal on Simulated Data
abstract
Dynamic Bayesian networks offer a powerful modelling tool to unravel cellular mechanisms. In particular, Gaussian networks have recently been used to model gene expression data, thanks to their capability to avoid information loss associated with discretization and their good computational efficiency. Gaussian networks typically describe the conditional mean of a node as a linear regression of the parent variables. Such model can be generalized by using a linear regression of nonlinear transformations of the parent values. In this paper we investigate the use of both models and evaluate the performance of Gaussian networks in learning the complex dynamic interactions among genes and proteins. To this aim, we analyzed simulated data produced by a mathematical model of cell cycle control in budding yeast. The results obtained allowed us to appraise the performance of the different models and confirmed the suitability of dynamic Bayesian networks for a first level, genome-wide analysis of high throughput dynamic data
Fulvia Ferrazzi, Paola Sebastiani, Isaac S. Kohane, Marco Ramoni, Riccardo Bellazzi
CBMS5
2006 Case-based retrieval to support the treatment of end stage renal failure patients
Stefania Montani, Luigi Portinale, Giorgio Leonardi, Riccardo Bellazzi, Roberto G. Bellazzi
Artif. Intell. Medicine4
2006 Knowledge-based data analysis and interpretation
Blaz Zupan, John H. Holmes, Riccardo Bellazzi
Artif. Intell. Medicine3
2006 A hierarchical Naïve Bayes Model for handling sample heterogeneity in classification problems: an application to tissue microarrays
abstract
BACKGROUND: Uncertainty often affects molecular biology experiments and data for different reasons. Heterogeneity of gene or protein expression within the same tumor tissue is an example of biological uncertainty which should be taken into account when molecular markers are used in decision making. Tissue Microarray (TMA) experiments allow for large scale profiling of tissue biopsies, investigating protein patterns characterizing specific disease states. TMA studies deal with multiple sampling of the same patient, and therefore with multiple measurements of same protein target, to account for possible biological heterogeneity. The aim of this paper is to provide and validate a classification model taking into consideration the uncertainty associated with measuring replicate samples. RESULTS: We propose an extension of the well-known Naïve Bayes classifier, which accounts for biological heterogeneity in a probabilistic framework, relying on Bayesian hierarchical models. The model, which can be efficiently learned from the training dataset, exploits a closed-form of classification equation, thus providing no additional computational cost with respect to the standard Naïve Bayes classifier. We validated the approach on several simulated datasets comparing its performances with the Naïve Bayes classifier. Moreover, we demonstrated that explicitly dealing with heterogeneity can improve classification accuracy on a TMA prostate cancer dataset. CONCLUSION: The proposed Hierarchical Naïve Bayes classifier can be conveniently applied in problems where within sample heterogeneity must be taken into account, such as TMA experiments and biological contexts where several measurements (replicates) are available for the same biological sample. The performance of the new approach is better than the standard Naïve Bayes model, in particular when the within sample heterogeneity is different in the different classes.
Francesca Demichelis, Paolo Magni, Paolo Piergiorgi, Mark A. Rubin, Riccardo Bellazzi
BMC Bioinform.5
2005 Learning Rules with Complex Temporal Patterns in Biomedical Domains
Lucia Sacchi, Riccardo Bellazzi, Cristiana Larizza, Riccardo Porreca, Paolo Magni
AIME2
2005 Comparison of two temporal abstraction procedures: a case study in prediction from monitoring data
Marion Verduijn, Arianna Dagliati, Lucia Sacchi, Niels Peek, Riccardo Bellazzi, Evert de Jonge, Bas A. de Mol
AMIA5
2005 Precedence Temporal Networks from Gene Expression Data
abstract
In this paper we introduce a novel method to extract from data and graphically represent the temporal relationships between events, called precedence temporal network. The new approach first derives events from time series by exploiting the temporal abstraction technique, then derives temporal precedence between abstractions in terms of association rules and finally expresses the relationships as a labeled graph. The method is applied to the problem of representing the temporal behavior of gene expressions, as they are collected by DNA microarrays. In particular, in this paper we present the results obtained from the analysis of the expression of a subset of the genes involved in cell-cycle regulation.
Lucia Sacchi, Riccardo Bellazzi, Riccardo Porreca, Cristiana Larizza, Paolo Magni
CBMS2
2005 Temporal data mining for the quality assessment of hemodialysis services
Riccardo Bellazzi, Cristiana Larizza, Paolo Magni, Roberto G. Bellazzi
Artif. Intell. Medicine1
2003 Quality Assessment of Hemodialysis Services through Temporal Data Mining
Riccardo Bellazzi, Cristiana Larizza, Paolo Magni, Roberto G. Bellazzi
AIME1
2003 Integrating model-based decision support in a multi-modal reasoning system for managing type 1 diabetic patients
Stefania Montani, Paolo Magni, Riccardo Bellazzi, Cristiana Larizza, Abdul V. Roudsari, Ewart R. Carson
Artif. Intell. Medicine3
2002 Multi-access Services for the Management of Diabetes Mellitus: The M2DM Project
Riccardo Bellazzi, Giuliana Bensa, Eulalia Brugués, Ewart R. Carson, Claudio Cobelli, Derek G. Cramp, Giuseppe d'Annunzio, Pasquale De Cata, Alberto de Leiva, Tibor Deutsch, Pietro Fratino, Carmine Gazzaruso, Angel Garcia, Tamás Gergely, Enrique J. Gómez, Fiona E. Harvey, Pietro Ferrari, Christiane Harras Friederich, María Elena Hernando, Maged N. Kamel Boulos, Cristiana Larizza, Hans Ludekke, Monika Luebker, Alberto Maran, Gianluca Nucci, Fernando Ortiz Garcia, Cristina Pennati, Abdul V. Roudsari, Mercedes Rigla, Karsten Schutte, Mario Stefanelli
AMIA1
2001 Mining Data from a Knowledge Management Perspective: An Application to Outcome Prediction in Patients with Resectable Hepatocellular Carcinoma
Riccardo Bellazzi, Ivano Azzini, Gianna Toffolo, Stefano Bacchetti, Mario Lise
AIME1
2001 Integrating Different Methodologies for Insulin Therapy Support in Type 1 Diabetic Patients
Stefania Montani, Paolo Magni, Abdul V. Roudsari, Ewart R. Carson, Riccardo Bellazzi
AIME5
2001 Supervised Implementation of Guidelines for Diabetes Management on the World Wide Web
Riccardo Bellazzi, Stefania Montani, M. Arcelloni, Pasquale De Cata, Carmine Gazzaruso, R. Giacchero, Pietro Fratino
AMIA1
2001 Learning from biomedical time series through the integration of qualitative models and fuzzy systems
Riccardo Bellazzi, Raffaella Guglielmann, Liliana Ironi
Artif. Intell. Medicine1
2001 A Hybrid Input-Output Approach to Model Metabolic Systems: An Application to Intracellular Thiamine Kinetics
Riccardo Bellazzi, Raffaella Guglielmann, Liliana Ironi, Cesare Patrini
J. Biomed. Informatics1
2000 Exploiting multi-modal reasoning for knowledge management and decision support: an evaluation study
Stefania Montani, Riccardo Bellazzi
AMIA2
2000 Artificial Intelligence Techniques for Diabetes Management: the T-IDDM Project
Stefania Montani, Riccardo Bellazzi, Alberto Riva, Cristiana Larizza, Luigi Portinale, Mario Stefanelli
ECAI2
2000 Intelligent analysis of clinical time series: an application in the diabetes mellitus domain
Riccardo Bellazzi, Cristiana Larizza, Paolo Magni, Stefania Montani, Mario Stefanelli
Artif. Intell. Medicine1
2000 How to improve fuzzy-neural system modeling by means of qualitative simulation
abstract
The main problem in efficiently building robust fuzzy-neural models of nonlinear systems lies in the difficulty to define a "meaningful" fuzzy rule-base. Our approach to the solution of such a problem is based on a hybrid method which integrates fuzzy systems with qualitative models. We introduce qualitative models to exploit the available, although incomplete, a priori physical knowledge on the system with the goal to infer, through qualitative simulation, all of its possible behaviors.We show here that a rule-base, which captures all of the distinctions in the system states, is automatically generated by encoding the knowledge of the system dynamics described by the outcomes of its qualitative simulation. Such a rule-base properly initializes a fuzzy identifier, which is then tuned to a set of experimental data. Our method has shown good performance when applied both as a predictor and as a simulator.
Riccardo Bellazzi, Raffaella Guglielmann, Liliana Ironi
IEEE Trans. Neural Networks Learn. Syst.1
1999 Integrating case based and rule based reasoning in a decision support system: evaluation with simulated patients
Stefania Montani, Riccardo Bellazzi
AMIA2
1999 Integrating Rule-Based and Case-Based Decision Making in Diabetic Patient Management
Riccardo Bellazzi, Stefania Montani, Luigi Portinale, Alberto Riva
ICCBR1
1999 A Qualitative-Fuzzy Framework for Nonlinear Black-Box System Identification
Riccardo Bellazzi, Raffaella Guglielmann, Liliana Ironi
IJCAI1
1998 Mining biomedical time series by combining structural analysis and temporal abstractions
Riccardo Bellazzi, Paolo Magni, Cristiana Larizza, Giuseppe De Nicolao, Alberto Riva, Mario Stefanelli
AMIA1
1998 A Web-Based System for Diabetes Management: The Technical and Clinical Infrastructure
Riccardo Bellazzi, Alberto Riva, Stefania Montani, Cristiana Larizza, Stefano Fiocchi, Giuseppe d'Annunzio, Renata Lorini, A. Monteforte, Mario Stefanelli
AMIA1
1998 Qualitative models and fuzzy systems: an integrated approach for learning from data
Riccardo Bellazzi, Liliana Ironi, Raffaella Guglielmann, Mario Stefanelli
Artif. Intell. Medicine1
1998 A development environment for knowledge-based medical applications on the world-wide web
Alberto Riva, Riccardo Bellazzi, Giordano Lanzola, Mario Stefanelli
Artif. Intell. Medicine2
1998 Temporal Abstractions for Interpreting Diabetic Patients Monitoring Data
abstract
In this article we present a new approach for the intelligent analysis of longitudinal data coming from chronic patients home monitoring. This approach exploits temporal abstractions to pre-process the raw data and to obtain a new time series of abstract episodes, whose features are then interpreted through statistical and probabilistic techniques. We describe in detail an application of the presented technique to the analysis of diabetic patients' data, showing some results obtained on a real case monitored for six months.
Riccardo Bellazzi, Cristiana Larizza, Alberto Riva
Intell. Data Anal.1
1998 Bayesian Function Learning Using MCMC Methods
abstract
The paper deals with the problem of reconstructing a continuous 1D function from discrete noisy samples. The measurements may also be indirect in the sense that the samples may be the output of a linear operator applied to the function. Bayesian estimation provides a unified treatment of this class of problems. We show that a rigorous Bayesian solution can be efficiently implemented by resorting to a Markov chain Monte Carlo (MCMC) simulation scheme. In particular, we discuss how the structure of the problem can be exploited in order to improve the computational and convergence performances. The effectiveness of the proposed scheme is demonstrated on two classical benchmark problems as well as on the analysis of IVGTT (IntraVenous glucose tolerance test) data, a complex identification-deconvolution problem concerning the estimation of the insulin secretion rate following the administration of an intravenous glucose injection.
Paolo Magni, Riccardo Bellazzi, Giuseppe De Nicolao
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Learning Bayesian networks probabilities from longitudinal data
abstract
Many real applications of Bayesian networks (BN) concern problems in which several observations are collected over time on a certain number of similar plants. This situation is typical of the context of medical monitoring, in which several measurements of the relevant physiological quantities are available over time on a population of patients under treatment, and the conditional probabilities that describe the model are usually obtained from the available data through a suitable learning algorithm. In situations with small data sets for each plant, it is useful to reinforce the parameter estimation process of the BN by taking into account the observations obtained from other similar plants. On the other hand, a desirable feature to be preserved is the ability to learn individualized conditional probability tables, rather than pooling together all the available data. In this work we apply a Bayesian hierarchical model able to preserve individual parameterization, and, at the same time, to allow the conditionals of each plant to borrow strength from all the experience contained in the data-base. A testing example and an application in the context of diabetes monitoring will be shown.
Riccardo Bellazzi, Alberto Riva
IEEE Trans. Syst. Man Cybern. Part A1
1997 Learning from Data Through the Integration of Qualitative Models and Fuzzy Systems
Riccardo Bellazzi, Liliana Ironi, Raffaella Guglielmann, Mario Stefanelli
AIME1
1997 Temporal Abstractions for Diabetic Patients Management
Cristiana Larizza, Riccardo Bellazzi, Alberto Riva
AIME2
1997 Interpreting Longitudinal Data through Temporal Abstractions: An Application to Diabetic Patients Monitoring
Riccardo Bellazzi, Cristiana Larizza, Alberto Riva
IDA1
1997 Dynamic Probabilistic Networks for Modelling and Identifying Dynamic Systems: a MCMC Approach
abstract
In this article we deal with the problem of interpreting data coming from a dynamic system by using causal probabilistic (CPN), a probabilistic graphical model particularly appealing in Intelligent Data Analysis. We discuss the different approaches presented in the literature, outlining their pros and cons through a simple training example. Then, we present a new method for reconstructing the state of the dynamic system, based on Markov Chain Monte Carlo algorithms, called dynamic probabilistic network smoothing (DPN-smoothing). Finally, we present an example of the application of DPN-smoothing in the field of signal deconvolution.
Riccardo Bellazzi, Paolo Magni, Giuseppe De Nicolao
Intell. Data Anal.1
1996 Causal Probabilistic Networks for Dynamic Modeling
Riccardo Bellazzi
ECAI1
1996 Learning temporal probabilistic causal models from longitudinal data
Alberto Riva, Riccardo Bellazzi
Artif. Intell. Medicine2
1995 High Level Control Strategies for Diabetes Therapy
Alberto Riva, Riccardo Bellazzi
AIME2
1995 Adaptive controllers for intelligent monitoring
Riccardo Bellazzi, Carlo Siviero, Mario Stefanelli, Giuseppe De Nicolao
Artif. Intell. Medicine1
1994 Reusable influence diagrams
Riccardo Bellazzi, Silvana Quaglini
Artif. Intell. Medicine1
1992 GAMEES II: an environment for building probabilistic expert systems based on arrays of Bayesian belief networks
abstract
The authors describe GAMEES II (Graphical Modeling Environment for Expert Systems II), a computer system built to manage arrays of Bayesian belief networks (BBNs). With regard to biomedical applications, the system makes it possible to represent probabilistic knowledge both for the individual patient and for population, enlarging the class of problems actually represented through BBNs. In addition, mathematical models can be represented through the BBN formalism. These characteristics were exploited to perform model-based patient monitoring and population pharmacokinetic/pharmacodynamic studies. BBNs are built within GAMEES II through a graphical interface which is provided with some statistical knowledge, in order to facilitate the construction of sound statistical models. Probabilistic inference is performed by stochastic simulation algorithms that do not give rise to computational problems when dealing with complex networks.>
Riccardo Bellazzi, Silvana Quaglini, Carlo Berzuini
CBMS1
1992 Bayesian networks for patient monitoring
Carlo Berzuini, Riccardo Bellazzi, Silvana Quaglini, David J. Spiegelhalter
Artif. Intell. Medicine2
1991 Cytotoxic Chemotherapy Monitoring Using Stochastic Simulation on Graphical Models
Riccardo Bellazzi, Carlo Berzuini, Silvana Quaglini, David J. Spiegelhalter, Mark Leaning
AIME1
1991 A Blackboard Control Architecture for Therapy Planning
Silvana Quaglini, Riccardo Bellazzi, Carlo Berzuini, Mario Stefanelli, Giovanni Barosi
AIME2
1991 Bayesian Networks Applied to Therapy Monitoring
Carlo Berzuini, David J. Spiegelhalter, Riccardo Bellazzi
UAI3
1989 Therapy Planning by Combining Ai and Decision Theoretic Techniques
Silvana Quaglini, Carlo Berzuini, Riccardo Bellazzi, Mario Stefanelli, Giovanni Barosi
AIME3