Lorenzo Famiglini

dblp:295/4111 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-1934-5899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Conformal Prediction for ECG Interpretation: A Study on Human-AI Collaboration in Clinical Decision Support
Duarte Folgado, Lorenzo Famiglini, Andrea Campagner, Hélder Dores, Marília Barandas, Hugo Gamboa, Federico Cabitza
AIME (1)2
2025 From Oracular to Judicial: Enhancing Clinical Decision Making through Contrasting Explanations and a Novel Interaction Protocol
abstract
Clinical Decision Support Systems (CDSS) utilizing machine learning (ML) classifiers have demonstrated substantial potential for improving diagnostic accuracy across various medical domains. However, concerns regarding automation bias, diminished sense of agency, and over-reliance on these systems remain, particularly in clinical settings where decision-making autonomy is critical.To address these challenges, we propose "Judicial AI,"an innovative interaction protocol aimed at reducing automation bias and preserving a sense of agency. This system presents contrasting explanations to medical professionals rather than definitive recommendations, encouraging user engagement and critical evaluation.Before adopting interaction protocols that avoid definitive recommendations, it is important to assess whether such an approach impacts diagnostic accuracy, and if so, how. This paper reports an exploratory study investigating the efficacy of a Judicial CDSS in the diagnosis of vertebral fractures from X-ray images. Sixteen medical professionals, comprising spine surgeons and radiologists, participated in the diagnosis of 18 X-ray images, which were carefully selected to represent particularly difficult and complex cases. Diagnosticians first recorded their decisions independently and then with support from the Judicial AI, which provided activation maps for opposing diagnoses.Our findings show a significant improvement in diagnostic accuracy for complex cases among experienced users (p =.045), with an overall accuracy increase of 0.24. Confidence levels also rose, particularly in the case of complex diagnoses (p =.034). However, the protocol was less beneficial for less experienced users, suggesting that cognitive load might be a limiting factor.These results suggest that Judicial AI, which frames decision-makers as the ultimate authority in the decision-making process, may be an effective tool for mitigating automation bias and preserving a sense of agency in clinical environments.
Federico Cabitza, Lorenzo Famiglini, Caterina Fregosi, Samuele Pe, Enea Parimbelli, Giovanni Andrea La Maida, Enrico Gallazzi
IUI2
2024 Dissimilar Similarities: Comparing Human and Statistical Similarity Evaluation in Medical AI
Federico Cabitza, Lorenzo Famiglini, Andrea Campagner, Luca Maria Sconfienza, Stefano Fusco, Valerio Caccavella, Enrico Gallazzi
MDAI2
2024 Never tell me the odds: Investigating pro-hoc explanations in medical decision making
abstract
This paper examines a kind of explainable AI, centered around what we term pro-hoc explanations, that is a form of support that consists of offering alternative explanations (one for each possible outcome) instead of a specific post-hoc explanation following specific advice. Specifically, our support mechanism utilizes explanations by examples, featuring analogous cases for each category in a binary setting. Pro-hoc explanations are an instance of what we called frictional AI, a general class of decision support aimed at achieving a useful compromise between the increase of decision effectiveness and the mitigation of cognitive risks, such as over-reliance, automation bias and deskilling. To illustrate an instance of frictional AI, we conducted an empirical user study to investigate its impact on the task of radiological detection of vertebral fractures in x-rays. Our study engaged 16 orthopedists in a 'human-first, second-opinion' interaction protocol. In this protocol, clinicians first made initial assessments of the x-rays without AI assistance and then provided their final diagnosis after considering the pro-hoc explanations. Our findings indicate that physicians, particularly those with less experience, perceived pro-hoc XAI support as significantly beneficial, even though it did not notably enhance their diagnostic accuracy. However, their increased confidence in final diagnoses suggests a positive overall impact. Given the promisingly high effect size observed, our results advocate for further research into pro-hoc explanations specifically, and into the broader concept of frictional AI.
Federico Cabitza, Chiara Natali, Lorenzo Famiglini, Andrea Campagner, Valerio Caccavella, Enrico Gallazzi
Artif. Intell. Medicine3
2023 Let Me Think! Investigating the Effect of Explanations Feeding Doubts About the AI Advice
Federico Cabitza, Andrea Campagner, Lorenzo Famiglini, Chiara Natali, Valerio Caccavella, Enrico Gallazzi
CD-MAKE3
2023 Towards a Rigorous Calibration Assessment Framework: Advancements in Metrics, Methods, and Use
abstract
Calibration is paramount in developing and validating Machine Learning models, particularly in sensitive domains such as medicine. Despite its significance, existing metrics to assess calibration have been found to have shortcomings in regard to their interpretation and theoretical properties. This article introduces a novel and comprehensive framework to assess the calibration of Machine and Deep Learning models that addresses the above limitations. The proposed framework is based on a modification of the Expected Calibration Error (ECE), called the Estimated Calibration Index (ECI), which grounds on and extends prior research. ECI was initially formulated for binary settings, and we adapted it to fit multiclass settings. ECI offers a more nuanced, both locally and globally, and informative measure of a model’s tendency towards over/underconfidence. The paper first outlines the issues related to the prevalent definitions of ECE, including potential biases that may arise in the evaluation of their measures. Then, we present the results of a series of experiments conducted to demonstrate the effectiveness of the proposed framework in supporting a more accurate understanding of a model’s calibration level. Additionally, we discuss how to address and potentially mitigate some biases in calibration assessment.
Lorenzo Famiglini, Andrea Campagner, Federico Cabitza
ECAI1
2022 Global Interpretable Calibration Index, a New Metric to Estimate Machine Learning Models' Calibration
Federico Cabitza, Andrea Campagner, Lorenzo Famiglini
CD-MAKE3
2022 Color Shadows (Part I): Exploratory Usability Evaluation of Activation Maps in Radiological Machine Learning
Federico Cabitza, Andrea Campagner, Lorenzo Famiglini, Enrico Gallazzi, Giovanni Andrea La Maida
CD-MAKE3
2022 Re-calibrating Machine Learning Models Using Confidence Interval Bounds
Andrea Campagner, Lorenzo Famiglini, Federico Cabitza
MDAI2
2021 Prediction of ICU admission for COVID-19 patients: a Machine Learning approach based on Complete Blood Count data
abstract
In this article we discuss the development of prognostic Machine Learning (ML) models for COVID-19 progression: specifically, we address the task of predicting intensive care unit (ICU) admission in the next 5 days. We developed three ML models on the basis of 4995 Complete Blood Count (CBC) tests. We propose three ML models that differ in terms of interpretability: two fully interpretable models and a black-box one. We report an AUC of. 81 and. 83 for the interpretable models (the decision tree and logistic regression, respectively), and an AUC of. 88 for the black-box model (an ensemble). This shows that CBC data and ML methods can be used for cost-effective prediction of ICU admission of COVID-19 patients: in particular, as the CBC can be acquired rapidly through routine blood exams, our models could also be applied in resource-limited settings and to get fast indications at triage and daily rounds.
Lorenzo Famiglini, Giorgio Bini, Anna Carobene, Andrea Campagner, Federico Cabitza
CBMS1
2021 On the Generalization of Figurative Language Detection: The Case of Irony and Sarcasm
Lorenzo Famiglini, Elisabetta Fersini, Paolo Rosso
NLDB1