Arianna Dagliati

dblp:123/7575 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-5041-0409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
YearPublicationVenuePosition
2026 A Benchmark Study for Reporting Feasibility in AI-Based Infant Hearing Screening: Exploring the Limits of Passive Sensing
Lorenzo Corso, Samuele Pe, Anisa Visram, Iain Jackson, Michael Stone, Kevin J. Munro, Enea Parimbelli, Giovanna Nicora, Arianna Dagliati
AIME (1)9
2026 Informative Missingness to Generate Irregular Clinical Time Series
Hadi Mehdizavareh, Gabriele Santangelo, Giovanna Nicora, Simon Lebech Cichosz, Arianna Dagliati, Arijit Khan 0001, Riccardo Bellazzi
AIME (2)5
2026 Rehabilitation movement simulation via joint angle-based generative AI
abstract
In recent years, generative models have shown remarkable capabilities in synthesizing realistic human motion, with applications ranging from animation to virtual reality. However, their potential in clinical and rehabilitation settings remains underexplored. In this work, we introduce a conditional diffusion-based generative framework for rehabilitation-oriented motion synthesis, which directly operates on joint-angle representations of full-body movement. Unlike most existing approaches that rely on joint positions, our method generates motion in a clinically meaningful space that explicitly encodes joint range of motion, aligning the generation process with how motor performance is assessed in rehabilitation practice. This design enables subject-independent modeling while improving the interpretability of the generated movements from a clinical perspective. We propose a comprehensive evaluation protocol by combining qualitative and quantitative metrics, including simulation visualizations, similarity analysis, and automated assessment of simulations adherence to users input. Experiments based on cross-subject and leave-one-combination-out settings demonstrate the model's ability to generate plausible, contextually accurate motion sequences, with improved generalization when using joint angle representations, achieving superior performance compared to a position-based approach. Despite limitations due to dataset size and gesture diversity, results support the feasibility of generating rehabilitation-oriented motion simulations, motivating future investigation in personalized rehabilitation scenarios.
Gabriele Santangelo, Chiara Alessi, Giovanna Nicora, Nikolas Sacchi, Samuele Pe, Antonella Ferrara, Riccardo Bellazzi, Arianna Dagliati
Artif. Intell. Medicine8
2025 Combining Clinical and Gene Expression Variables via Knowledge Graph Embedding for Prediction of Coronary Artery Stenosis
Giuseppe Albi, Arianna Dagliati, Chiara Vavassori, Laura Pisani, Mattia Chiesa, Luca Piacentini, Saima Mushtaq, Gianluca Pontone, Riccardo Bellazzi, Gualtiero Colombo 0002
AIME (1)2
2025 Domain-Adversarial Neural Networks to Explore Biases in the Diagnosis of Multiple Eye Conditions from Fundus Image Data
Chiara Pullega, Arianna Dagliati, Allan Tucker
AIME (2)2
2025 Upper Limb Movements Simulations with Generative Diffusion Models
Gabriele Santangelo, Chiara Alessi, Nikolas Sacchi, Giovanna Nicora, Riccardo Bellazzi, Antonella Ferrara, Arianna Dagliati
AIME (2)7
2025 Deep Learning Model Predicts Relapse Occurrence in Multiple Sclerosis Via Sequences of Environmental Data
abstract
Air pollution is a known risk factor for the exacerbation of many diseases. Among these, is multiple sclerosis (MS), a chronic, autoimmune, neurological disease, characterised by transient episodes of neurological impairment known as relapses. Although the link between environmental factors and relapses has been a subject of investigation in the medical and biostatistical literature, its implications for predictive modelling are still unclear. Thus, in this work, we develop a deep learning model that is able to combine four weeks of environmental data, collected by pollutant-monitoring and weather stations, with patient information to predict an imminent relapse in the following week. Specifically, we cast the task as distinguishing between 4-week sequences followed by a relapse vs. 4-week sequences followed by another relapse-free week, the latter of which were extracted from MS patients who were never observed to have had a relapse. The 1556 sequences were collected in the context of the H2020 BRAINTEASER (”Bringing Artificial Intelligence Home for a Better Care of Amyotrophic Lateral Sclerosis and Multiple Sclerosis”) project. The best-performing model was a recurrent neural network, which yielded an encouraging test-set area under the receiveroperating characteristic curve (AUROC) of 0.70. It also performed adequately (AUROC$=0.60$) on a modified version of the test set where the 4-week relapse-free sequences followed by another relapse-free week were extracted from the same subjects from whom the test sequences followed by a relapse came. Thus, our results, albeit preliminary, suggest that the inclusion of environmental data as the basis of predictive models of MS relapses is a promising direction to obtain short-term predictions, which may be helpful for therapy and life planning. It is especially encouraging that better-than-random performance was preserved on the modified test set, where environmental factors were, by construction, the most informative predictors.
Enrico Longato, Erica Tavazzi, Anna Milani, Elena Marinello, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo
BIBM6
2024 Machine Learning Models Highlight the Impact of Pollution and Weather Patterns on Relapse Occurrence in Multiple Sclerosis Patients
abstract
Multiple Sclerosis (MS) is a chronic autoimmune and inflammatory neurological disorder characterised by episodes of symptom exacerbation, known as relapses. Relapses have been linked to environmental factors such as the weather and pollutant concentrations in the air, but the exact relationship between these phenomena is still unclear. In this study, we investigated the role of environmental factors in predicting imminent relapse occurrence in MS patients, leveraging clinical and environmental data collected over a period of one week preceeding the possible event, using data collected in the context of the H2020 BRAINTEASER project. To do this, we developed and tested a range of combinations of predictive models (logistic regression, LR; and random forest, RF) and feature selection schemes, both manual and data-driven. The RF model trained after a data-driven feature selection process based on the Variable Importance in Projection (VIP) metric yielded the best results, i.e., an AUC-ROC of 0.713 and an AUC-PR of 0.639. We identified several key predictors, including clinical variables such as time since MS onset, age at onset, diagnostic delay, and the Expanded Disability Status Scale (EDSS) score, and environmental variables such as wind speed, precipitation, NO2, PM10, average and maximum temperatures, and humidity. These findings suggest that environmental factors may be viable predictors of imminent relapse occurrence in MS.
Elena Marinello, Erica Tavazzi, Enrico Longato, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo
BIBM5
2024 Land Use Regression on Interpolated Urban Graphs to Assess Personal Exposure to Air Pollution
abstract
Past research has demonstrated that continuous exposure to pollutants, such as PM2.5 and PM10, is associated with an increased risk of developing and worsening respiratory and neurodegenerative diseases. Calculating and reducing exposure to these pollutants is crucial to assess these risks and perform proper prevention. In this study, we estimate personal exposure to PM2.5 based on the integration of sensors measurements, meteorological data and land use parameters, which could impact on actual pollution levels, especially in areas located far from the sensors. Pollution data have been collected from a dense network of sensors located in Pavia, Italy, meteorological and geographical data have been collected from public sources. We used geographical data to create graphs that model the city road structure, and applied Land Use Regression methods to estimate air pollution on its nodes, adjusting the measurements interpolated from the sensors with the effects of weather data, land use parameters such as the distance from the closest high-traffic road, and additional temporal information such as weekends/holidays and working days. We tested several regression methods: linear regression, both simple and with regularization (Ridge, LASSO and ElasticNet), Random Forest regression, Gradient Boosting and Support Vector Regression (SVR). Results show that meteorological variables, namely temperature and humidity, and temporal factors do contribute significantly in obtaining pollution values in the graph nodes that differ from values obtained exclusively through sensors interpolation.
Daniele Pala, Giacomo Zagami, Pietro Bosoni, Mahin Vazifehdan, Riccardo Bellazzi, Arianna Dagliati
BIBM6
2024 iDPP@CLEF 2024: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi
ECIR (6)5
2023 A Topological Data Analysis Framework for Computational Phenotyping
Giuseppe Albi, Alessia Gerbasi, Mattia Chiesa, Gualtiero Colombo 0002, Riccardo Bellazzi, Arianna Dagliati
AIME6
2023 iDPP@CLEF 2023: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi
ECIR (3)5
2023 Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic review
abstract
BACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors.
Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo
Artif. Intell. Medicine22
2023 Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes
J. Biomed. Informatics7
2022 A Deductive Data-Driven Pipeline Powered by MLHO for Post-Acute Sequelae of COVID-19 (PASC) Phenotyping
Arianna Dagliati, Zachary H. Strasser, Rebecca Mesa, Zahra Shakeri, Alaleh Azhir, Riccardo Bellazzi, Shawn N. Murphy, Hossein Estiri
AMIA1
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics16
2021 Temporal Phenotypic Pathways of Post-Acute Sequelae of SARS-CoV-2 by an International Consortium for Clinical Characterization of COVID-19 (4CE)
Shawn N. Murphy, Hossein Estiri, Arianna Dagliati, Riccardo Bellazzi, John H. Holmes
AMIA3
2021 Health informatics and EHR to support clinical research in the COVID-19 pandemic: an overview
abstract
The coronavirus disease 2019 (COVID-19) pandemic has clearly shown that major challenges and threats for humankind need to be addressed with global answers and shared decisions. Data and their analytics are crucial components of such decision-making activities. Rather interestingly, one of the most difficult aspects is reusing and sharing of accurate and detailed clinical data collected by Electronic Health Records (EHR), even if these data have a paramount importance. EHR data, in fact, are not only essential for supporting day-by-day activities, but also they can leverage research and support critical decisions about effectiveness of drugs and therapeutic strategies. In this paper, we will concentrate our attention on collaborative data infrastructures to support COVID-19 research and on the open issues of data sharing and data governance that COVID-19 had made emerge. Data interoperability, healthcare processes modelling and representation, shared procedures to deal with different data privacy regulations, and data stewardship and governance are seen as the most important aspects to boost collaborative research. Lessons learned from COVID-19 pandemic can be a strong element to improve international research and our future capability of dealing with fast developing emergencies and needs, which are likely to be more frequent in the future in our connected and intertwined world.
Arianna Dagliati, Alberto Malovini, Valentina Tibollo, Riccardo Bellazzi
Briefings Bioinform.1
2020 Mining post-surgical care processes in breast cancer patients
Lorenzo Chiudinelli, Arianna Dagliati, Valentina Tibollo, Sara Albasini, Nophar Geifman, Niels Peek, John H. Holmes, Fabio Corsi, Riccardo Bellazzi, Lucia Sacchi
Artif. Intell. Medicine2
2020 Using topological data analysis and pseudo time series to infer temporal phenotypes from electronic health records
abstract
Temporal phenotyping enables clinicians to better understand observable characteristics of a disease as it progresses. Modelling disease progression that captures interactions between phenotypes is inherently challenging. Temporal models that capture change in disease over time can identify the key features that characterize disease subtypes that underpin these trajectories. These models will enable clinicians to identify early warning signs of progression in specific sub-types and therefore to make informed decisions tailored to individual patients. In this paper, we explore two approaches to building temporal phenotypes based on the topology of data: topological data analysis and pseudo time-series. Using type 2 diabetes data, we show that the topological data analysis approach is able to identify disease trajectories and that pseudo time-series can infer a state space model characterized by transitions between hidden states that represent distinct temporal phenotypes. Both approaches highlight lipid profiles as key factors in distinguishing the phenotypes.
Arianna Dagliati, Nophar Geifman, Niels Peek, John H. Holmes, Lucia Sacchi, Riccardo Bellazzi, Seyed Erfan Sajjadi, Allan Tucker
Artif. Intell. Medicine1
2020 The use of missing values in proteomic data-independent acquisition mass spectrometry to enable disease activity discrimination
abstract
MOTIVATION: Data-independent acquisition mass spectrometry allows for comprehensive peptide detection and relative quantification than standard data-dependent approaches. While less prone to missing values, these still exist. Current approaches for handling the so-called missingness have challenges. We hypothesized that non-random missingness is a useful biological measure and demonstrate the importance of analysing missingness for proteomic discovery within a longitudinal study of disease activity. RESULTS: The magnitude of missingness did not correlate with mean peptide concentration. The magnitude of missingness for each protein strongly correlated between collection time points (baseline, 3 months, 6 months; R = 0.95-0.97, confidence interval = 0.94-0.97) indicating little time-dependent effect. This allowed for the identification of proteins with outlier levels of missingness that differentiate between the patient groups characterized by different patterns of disease activity. The association of these proteins with disease activity was confirmed by machine learning techniques. Our novel approach complements analyses on complete observations and other missing value strategies in biomarker prediction of disease activity. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kathryn A. McGurk, Arianna Dagliati, Davide Chiasserini, Dave Lee, Darren Plant, Ivona Baricevic-Jones, Janet Kelsall, Rachael Eineman, Rachel Reed, Bethany Geary, Richard D. Unwin, Anna Nicolaou, Bernard D. Keavney, Anne Barton, Anthony D. Whetton, Nophar Geifman
Bioinform.2
2019 Inferring Temporal Phenotypes with Topological Data Analysis and Pseudo Time-Series
Arianna Dagliati, Nophar Geifman, Niels Peek, John H. Holmes, Lucia Sacchi, Seyed Erfan Sajjadi, Allan Tucker
AIME1
2018 A dashboard-based system for supporting diabetes care
abstract
Objective: To describe the development, as part of the European Union MOSAIC (Models and Simulation Techniques for Discovering Diabetes Influence Factors) project, of a dashboard-based system for the management of type 2 diabetes and assess its impact on clinical practice. Methods: The MOSAIC dashboard system is based on predictive modeling, longitudinal data analytics, and the reuse and integration of data from hospitals and public health repositories. Data are merged into an i2b2 data warehouse, which feeds a set of advanced temporal analytic models, including temporal abstractions, care-flow mining, drug exposure pattern detection, and risk-prediction models for type 2 diabetes complications. The dashboard has 2 components, designed for (1) clinical decision support during follow-up consultations and (2) outcome assessment on populations of interest. To assess the impact of the clinical decision support component, a pre-post study was conducted considering visit duration, number of screening examinations, and lifestyle interventions. A pilot sample of 700 Italian patients was investigated. Judgments on the outcome assessment component were obtained via focus groups with clinicians and health care managers. Results: The use of the decision support component in clinical activities produced a reduction in visit duration (P ≪ .01) and an increase in the number of screening exams for complications (P < .01). We also observed a relevant, although nonstatistically significant, increase in the proportion of patients receiving lifestyle interventions (from 69% to 77%). Regarding the outcome assessment component, focus groups highlighted the system's capability of identifying and understanding the characteristics of patient subgroups treated at the center. Conclusion: Our study demonstrates that decision support tools based on the integration of multiple-source data and visual and predictive analytics do improve the management of a chronic disease such as type 2 diabetes by enacting a successful implementation of the learning health care system cycle.
Arianna Dagliati, Lucia Sacchi, Valentina Tibollo, Giulia Cogni, Marsida Teliti, Antonio Martinez-Millana, Vicente Traver 0001, Daniele Segagni, Manuel Ottaviano, Giuseppe Fico, María Teresa Arredondo, Pasquale De Cata, Luca Chiovato, Riccardo Bellazzi
J. Am. Medical Informatics Assoc.1
2018 Incorporating repeating temporal association rules in Naïve Bayes classifiers for coronary heart disease diagnosis
Kalia Orphanou, Arianna Dagliati, Lucia Sacchi, Athena Stassopoulou, Elpida T. Keravnou, Riccardo Bellazzi
J. Biomed. Informatics2
2017 pMineR: An Innovative R Library for Performing Process Mining in Medicine
Roberto Gatta, Jacopo Lenkowicz, Mauro Vallati, Eric Rojas Cordoba, Andrea Damiani, Lucia Sacchi, Berardino De Bari, Arianna Dagliati, Carlos Fernández-Llatas, Matteo Montesi, Antonio Marchetti, Maurizio Castellano, Vincenzo Valentini
AIME8
2017 Generating and Comparing Knowledge Graphs of Medical Processes Using pMineR
abstract
Process mining focuses on extracting knowledge, under the form of models, from data generated and stored in information systems. The analysis of generated models can provide useful insights to domain experts. In addition, models of processes can be used to test if a considered process complies with some given specifications. For these reasons, process mining is gaining significant importance in the healthcare domain, where the complexity and flexibility of processes makes extremely hard to evaluate and assess how patients have been treated.
Roberto Gatta, Mauro Vallati, Jacopo Lenkowicz, Eric Rojas Cordoba, Andrea Damiani, Lucia Sacchi, Berardino De Bari, Arianna Dagliati, Carlos Fernández-Llatas, Matteo Montesi, Antonio Marchetti, Maurizio Castellano, Vincenzo Valentini
K-CAP8
2017 Temporal electronic phenotyping by mining careflows of breast cancer patients
Arianna Dagliati, Lucia Sacchi, Alberto Zambelli, Valentina Tibollo, L. Pavesi, John H. Holmes, Riccardo Bellazzi
J. Biomed. Informatics1
2016 Hierarchical Bayesian Logistic Regression to forecast metabolic control in type 2 DM patients
Arianna Dagliati, Alberto Malovini, Pasquale De Cata, Giulia Cogni, Marsida Teliti, Lucia Sacchi, Carlo Cerra, Luca Chiovato, Riccardo Bellazzi
AMIA1
2015 Inferring air quality maps from remotely sensed data to exploit georeferenced clinical onsets: The Pavia 2013 case
abstract
Recent developments in data acquisition, storage, mining and maintenance have allowed the flourishing of several multi-disciplinary research fields, which can be stated, defined and carried out according to the so-called Big Data paradigm. In this environment, the investigation and analysis of interactions between human phenomena and natural events play a key-role, as they can be fundamental for several applications, from sustainable development to community policy design and short-, medium- and long-range resource allocation planning. In this paper, we provide a study of the interplay between air pollution (as estimated by remotely sensed data processing) and clinical records, so that inferences and correlations among black particulate concentration, micro- and macro-vascular disease onsets and hospitalization tracks can be efficiently drawn. We focused on the second order administrative area of the city of Pavia, Italy, on 2013. Experimental results show how effective connections between the estimated air quality and the hospitalizations behavior can be accurately drawn and derived.
Andrea Marinoni, Arianna Dagliati, Riccardo Bellazzi, Paolo Gamba
IGARSS2
2013 Mining Careflow Patterns in data warehouses of breast cancer patients
Lucia Sacchi, Daniele Segagni, Arianna Dagliati, Alberto Zambelli, Riccardo Bellazzi
AMIA3
2012 An ICT infrastructure to integrate clinical and molecular data in oncology research
abstract
BACKGROUND: The ONCO-i2b2 platform is a bioinformatics tool designed to integrate clinical and research data and support translational research in oncology. It is implemented by the University of Pavia and the IRCCS Fondazione Maugeri hospital (FSM), and grounded on the software developed by the Informatics for Integrating Biology and the Bedside (i2b2) research center. I2b2 has delivered an open source suite based on a data warehouse, which is efficiently interrogated to find sets of interesting patients through a query tool interface. METHODS: Onco-i2b2 integrates data coming from multiple sources and allows the users to jointly query them. I2b2 data are then stored in a data warehouse, where facts are hierarchically structured as ontologies. Onco-i2b2 gathers data from the FSM pathology unit (PU) database and from the hospital biobank and merges them with the clinical information from the hospital information system. Our main effort was to provide a robust integrated research environment, giving a particular emphasis to the integration process and facing different challenges, consecutively listed: biospecimen samples privacy and anonymization; synchronization of the biobank database with the i2b2 data warehouse through a series of Extract, Transform, Load (ETL) operations; development and integration of a Natural Language Processing (NLP) module, to retrieve coded information, such as SNOMED terms and malignant tumors (TNM) classifications, and clinical tests results from unstructured medical records. Furthermore, we have developed an internal SNOMED ontology rested on the NCBO BioPortal web services. RESULTS: Onco-i2b2 manages data of more than 6,500 patients with breast cancer diagnosis collected between 2001 and 2011 (over 390 of them have at least one biological sample in the cancer biobank), more than 47,000 visits and 96,000 observations over 960 medical concepts. CONCLUSIONS: Onco-i2b2 is a concrete example of how integrated Information and Communication Technology architecture can be implemented to support translational research. The next steps of our project will involve the extension of its capabilities by implementing new plug-in devoted to bioinformatics data analysis as well as a temporal query module.
Daniele Segagni, Valentina Tibollo, Arianna Dagliati, Alberto Zambelli, Silvia G. Priori, Riccardo Bellazzi
BMC Bioinform.3
2005 Comparison of two temporal abstraction procedures: a case study in prediction from monitoring data
Marion Verduijn, Arianna Dagliati, Lucia Sacchi, Niels Peek, Riccardo Bellazzi, Evert de Jonge, Bas A. de Mol
AMIA2