Federico Cabitza

dblp:20/4501 · DBLP profile ↗
← Back
71ranked-venue papers
44as first author
38since 2021 · last 2026
0000-0002-4065-3415ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 43 · 30 first-author · 17 since 2021Artificial intelligence and machine learning · 29 · 15 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Calibrating Reliance: Addressing Misuse and Disuse in AI-Based Second-Opinion Systems for Medical Diagnosis
abstract
AI systems are widely proposed as second-opinion advisors in clinical diagnosis, offering the promise of enhancing decision accuracy and clinician confidence while preserving human oversight. However, successful deployment in real-world practice faces a critical barrier: clinicians' reliance on AI is often miscalibrated, manifesting as misuse (over-reliance driven by automation bias) and disuse (under-utilization driven by self-anchoring bias). This paper addresses these deployment challenges by systematically analyzing how such reliance patterns affect diagnostic accuracy, confidence, and decision-making across diverse medical specialties. We report results from controlled simulations involving over 300 medical professionals across six diagnostic settings—including knee MRI analysis, spinal X-rays, cardiac ECG evaluation, and gastrointestinal endoscopy—using a human-first, AI-second workflow. Although AI advice improved average diagnostic accuracy (+2 percentage points) and clinician confidence (+3 points on a normalized scale), overall levels of appropriate reliance remained well below 50%, with disuse emerging as the more prevalent and consequential barrier. We introduce and validate Appropriate Reliance as an actionable metric for assessing and improving human-AI collaboration, providing practical guidance for developers, healthcare institutions, and policymakers seeking to deploy second-opinion AI systems safely and effectively. By identifying the sociotechnical barriers and offering evidence-based design insights, this work supports the emerging application of AI as a collaborative advisor in clinical workflows, charting a clear path toward deployment that enhances diagnostic safety, accountability, and patient care. Specifically, we propose integrating the Appropriate Reliance metric into system development workflows, clinician training, and regulatory evaluations to enable safe and effective deployment of second-opinion AI systems.
Federico Cabitza, Andrea Campagner, Gian Eugenio Tontini
AAAI1
2026 Too Sure for Our Own Good: A User Study on AI Confidence and Human Reliance
abstract
Achieving appropriate human reliance on Artificial Intelligence (AI) systems remains a central challenge in Human-Computer Interaction. Confidence scores—indicators of an AI system’s certainty in its recommendations—have been proposed as a means to help users calibrate their trust and reliance on AI Decision Support Systems (DSS). However, limited research has explored how well-calibrated versus miscalibrated confidence scores affect human decision-making. We report a study examining the effects of confidence calibration on user reliance, decision accuracy, and perceived utility of an AI DSS. In a within-subjects experiment involving 184 participants solving logic puzzles, we found that well-calibrated confidence scores significantly improved decision accuracy (+20%, 95% CI: [0.18, 0.23]), whereas miscalibrated scores yielded minimal accuracy gains (+2%, 95% CI: [-0.00, 0.04]) and increased vulnerability to automation bias and conservatism bias. Participants were more likely to accept AI recommendations when high confidence was expressed, even when those recommendations were incorrect, resulting in errors. Conversely, miscalibrated and low-confidence recommendations increased conservatism bias, leading users to reject even accurate AI suggestions. Perceived utility of the AI system was higher when confidence levels were high (p < 0.001) and when confidence was well-calibrated (p = 0.002). These findings underscore the importance of designing AI systems with properly calibrated confidence cues to improve human-AI collaboration and mitigate reliance-related biases.
Caterina Fregosi, Lucia Vicente, Andrea Campagner, Federico Cabitza
AAAI4
2026 Clinical Deployability: A Socio-technical Construct for Evaluating Real-World Readiness of Medical AI
Federico Cabitza, Mauro Dragoni
AIME (2)1
2026 How Explanation Framing Shapes Reliance on AI in Clinical Decision Support
Federico Cabitza, Alessia Papale, Lucia Vicente, Rossella Tomaiuolo
AIME (1)1
2026 Unity Is Strength: Hybridizing Neuro-Symbolic AI and Agentic AI through Symbolic Coordination Mechanisms
Alessia Papale, Federico Cabitza
ICAART (1)2
2026 Calibration-informed metrics for instance-level predictive reliability in medical AI
abstract
Conventional performance metrics in clinical decision support systems, such as accuracy or sensitivity, fail to reflect the reliability of individual predictions—an essential concern for clinicians operating in high-stakes environments. We introduce a calibration-informed framework featuring two novel metrics: the Local Predictive Value (LPV) and the Credible Predictive Value (CPV). LPV estimates the empirical reliability of a prediction by assessing the observed correctness frequency in the neighborhood of its confidence score. CPV refines this estimate using a Bayesian approach, integrating global predictive values as priors to produce a posterior distribution over correctness probabilities. LPV offers a descriptive, data-driven view of local reliability, while CPV provides a belief-adjusted estimate that mitigates overfitting to sparse local data. Applied to benchmark medical imaging datasets, these metrics yielded locally adaptive, interpretable reliability estimates. Divergences between LPV and CPV identified cases where local evidence was insufficient or misleading, highlighting how Bayesian smoothing improves stability against sparse or misleading local evidence. By combining local calibration with Bayesian inference, LPV and CPV advance the development of medical AI systems that are not only accurate but also interpretable and trustworthy at the individual case level. • Introduced two metrics to assess prediction reliability in clinical AI decision support. • Local Predictive Value estimates trust based on observed outcomes near the score. • Credible Predictive Value uses Bayesian inference to blend local and global evidence. • Metrics reveal where local evidence is misleading and priors improve robustness. • Applied the method to clinical datasets for interpretable, case-specific evaluation. • Proposed a calibration-aware framework for trustworthy prediction in medical AI.
Federico Cabitza
Artif. Intell. Medicine1
2026 Machine learning systems as mentors in human learning: The role of AI output in diagnostic knowledge acquisition
abstract
Artificial intelligence (AI) systems are increasingly integrated into decision-making across high-stakes domains, influencing not only task performance but potentially human learning. While prior research has focused on AI’s impact on accuracy, its role in supporting long-term knowledge acquisition remains underexplored. This study investigates whether AI-based decision support systems can serve as implicit mentors, enabling users to internalize novel decision strategies through repeated interaction, a phenomenon we term machine mentoring. In a simulated diagnostic task, 289 medical students received different forms of AI support. Participants in the Feedback condition, who received trial-by-trial feedback on the correctness of both their own and the AI’s decisions, significantly improved their accuracy on cases involving a hidden diagnostic criterion and retained these gains in subsequent unaided trials. By contrast, AI advice alone, or with confidence indicators, did not lead to significant learning outcomes. These findings indicate that evaluative, trial-by-trial feedback is the key driver of AI-supported knowledge transfer: advice alone, even when paired with calibrated confidence cues, did not yield durable learning. Rather than attributing the effect to AI per se, our results indicate that what drives durable learning is access to trial-level, ground-truth corrective feedback—something competent AI systems can help deliver at scale. This is especially relevant in medical education and simulation contexts, where feedback can be reliably validated.
Federico Cabitza, Lucia Vicente
Int. J. Hum. Comput. Stud.1
2026 Under what influence: Measuring AI influence to fit user profiles in decision-making
abstract
Artificial Intelligence (AI) has become a pivotal tool in augmenting human decision-making across various domains, yet its influence on user decisions often lacks comprehensive evaluation. While technical performance metrics such as accuracy and efficiency dominate AI design, integrating human-centered approaches that consider trust and reliance remains underexplored. This study addresses the knowledge gap in understanding how AI systems influence decision-making quality, calibrated to user profiles, including their expertise, skills, professional role, confidence, and reliance tendencies. We present a novel and comprehensive metric framework for evaluating AI influence, emphasizing behavioral patterns and measurable improvements in decision outcomes beyond simple alignment with AI recommendations. The framework is applied to four medical domain case studies—MRI, ECG, X-ray, and ENDO – with user groups spanning specialists, sub-specialists, and trainees. Results reveal that while human and AI systems achieve high agreement rates (up to 81%), AI influence on decision quality varies significantly. Notably, X-ray decision-making showed the highest influence index (0.27), while MRI decisions exhibited substantial self-anchoring bias (6.94), undermining the potential positive impact of AI. Influence metrics unveiled nuances missed by agreement scores, highlighting domain-specific biases and opportunities to optimize AI-human interaction. This research underscores the necessity for adapting the type of AI system and affordance to user characteristics and attitudes of reliance to foster calibrated trust and improve decision outcomes. Our findings inform the design of AI systems that better support diverse user needs and align with human decisions, driving progress toward human-centered AI integration in high-stakes domains. • Developed novel metrics to evaluate AI’s influence beyond user agreement with AI. • Identified biases impacting AI influence, such as self-anchoring and automation bias. • Applied framework to four medical studies with 330 clinicians and 15,000 decisions. • Revealed up to 81% alignment but variances in appropriate reliance and influence. • Highlighted need for adaptive AI systems to match user expertise. • Demonstrated that influence metrics uncover dynamics missed by traditional reliance.
Andrea Campagner, Caterina Fregosi, Chiara Natali, Federico Cabitza
Int. J. Hum. Comput. Stud.4
2025 Who Knocks on Heaven's Door: Measuring Augmentation and Outperformance in Human-AI Diagnostic Teams
Federico Cabitza
AIME (2)1
2025 Conformal Prediction for ECG Interpretation: A Study on Human-AI Collaboration in Clinical Decision Support
Duarte Folgado, Lorenzo Famiglini, Andrea Campagner, Hélder Dores, Marília Barandas, Hugo Gamboa, Federico Cabitza
AIME (1)7
2025 Explainable Machine Learning for Neonatal Screening: A Fast&Frugal Decision Tree for Rare Metabolic Disease Detection
Gloria Lopiano, Andrea Campagner, Cristina Cereda, Stephana Carelli, Federico Cabitza
AIME (1)5
2025 An Evidence-Theoretic Framework for Online Learning from Expert Advice
abstract
The use of belief function theory (BFT) in machine learning has gained attention as researchers seek more principled foundations for decision-making in uncertain environments. However, research has mostly focused on the setting of batch learning. In this article, in contrast and to our knowledge for the first time in the literature, we study the application of BFT to the setting of online (machine) learning. Within this context, online learning from expert advice (LEA) offers a framework where learners iteratively update their predictions based on experts’ input and (adversarially labeled) observed outcomes. Despite extensive study and strong theoretical results, the epistemological underpinnings of LEA remain largely heuristic. This work addresses this gap by proposing belief function theory (BFT) as a formal foundation for LEA. Here we report a theoretical and algorithmic integration of BFT into LEA, showing that classical LEA algorithms such as Halving and Weighted Majority can be derived as special cases of evidential reasoning. We further introduce two novel LEA algorithms—Evidential Halving and Evidential Weighted Majority—which fully exploit BFT and support cautious prediction through abstention. These new algorithms demonstrate improved regret bounds over traditional methods, under mild assumptions. These findings open a new direction in online learning by leveraging the full expressive power of BFT to design theoretically grounded algorithms.
Andrea Campagner, Francesca Arredondo, Davide Ciucci, Federico Cabitza
ECAI4
2025 From Oracular to Judicial: Enhancing Clinical Decision Making through Contrasting Explanations and a Novel Interaction Protocol
abstract
Clinical Decision Support Systems (CDSS) utilizing machine learning (ML) classifiers have demonstrated substantial potential for improving diagnostic accuracy across various medical domains. However, concerns regarding automation bias, diminished sense of agency, and over-reliance on these systems remain, particularly in clinical settings where decision-making autonomy is critical.To address these challenges, we propose "Judicial AI,"an innovative interaction protocol aimed at reducing automation bias and preserving a sense of agency. This system presents contrasting explanations to medical professionals rather than definitive recommendations, encouraging user engagement and critical evaluation.Before adopting interaction protocols that avoid definitive recommendations, it is important to assess whether such an approach impacts diagnostic accuracy, and if so, how. This paper reports an exploratory study investigating the efficacy of a Judicial CDSS in the diagnosis of vertebral fractures from X-ray images. Sixteen medical professionals, comprising spine surgeons and radiologists, participated in the diagnosis of 18 X-ray images, which were carefully selected to represent particularly difficult and complex cases. Diagnosticians first recorded their decisions independently and then with support from the Judicial AI, which provided activation maps for opposing diagnoses.Our findings show a significant improvement in diagnostic accuracy for complex cases among experienced users (p =.045), with an overall accuracy increase of 0.24. Confidence levels also rose, particularly in the case of complex diagnoses (p =.034). However, the protocol was less beneficial for less experienced users, suggesting that cognitive load might be a limiting factor.These results suggest that Judicial AI, which frames decision-makers as the ultimate authority in the decision-making process, may be an effective tool for mitigating automation bias and preserving a sense of agency in clinical environments.
Federico Cabitza, Lorenzo Famiglini, Caterina Fregosi, Samuele Pe, Enea Parimbelli, Giovanni Andrea La Maida, Enrico Gallazzi
IUI1
2025 Dimensions of Human-Machine Combination: Prompting the Development of Deployable Intelligent Decision Systems for Situated Clinical Contexts
abstract
Abstract Whilst it is commonly reported that healthcare is set to benefit from advances in Artificial Intelligence (AI), there is a consensus that, for clinical AI, a gulf exists between conception and implementation. Here we advocate the increased use of situated design and evaluation to close this gap, showing that in the literature there are comparatively few prospective situated studies. Focusing on the combined human-machine decision-making process - modelling, exchanging and resolving - we highlight the need for advances in exchanging and resolving. We present a novel relational space - contextual dimensions of combination - a means by which researchers, developers and clinicians can begin to frame the issues that must be addressed in order to close the chasm. We introduce a space of eight initial dimensions, namely participating agents, control relations, task overlap, temporal patterning, informational proximity, informational overlap, input influence and output representation coverage. We propose that our awareness of where we are in this space of combination will drive the development of interactions and the designs of AI models themselves. Designs that take account of how user-centered they will need to be for their performance to be translated into societal and individual benefit.
Benjamin Wilson 0002, Chiara Natali, Matthew Roach 0001, Darren Scott, Alma As-Aad Mohammad Rahat, David Rawlinson 0003, Federico Cabitza
Comput. Support. Cooperative Work.7
2025 Machine learning systems as mentors in human learning: A user study on machine bias transmission in medical training
abstract
While accurate AI systems can enhance human performance, exerting both an augmentation and good mentoring effect, imperfect systems may act as poor mentors, transmitting biases and systematic errors to users. However, there is still limited research on the potential for AI to transmit biases to humans, an effect that could be even more pronounced for less experienced users, such as novices or trainees, making decisions supported by AI-based systems. To investigate the bias transmission effect and the potential of AI to serve as a mentor, we involved eighty-six medical students, dividing them into an AI-assisted group and a control group. We tasked them with classifying simulated tissue samples for a fictitious disease. In the first phase of the task, the AI group received diagnostic advice from a simulated AI system that made systematic errors for a specific type of case, while being accurate for all other types. The control group did not receive any assistance. In the second phase, participants in both groups classified new tissue samples, including ambiguous cases, without any support to test the residual impact of AI bias. The results showed that the AI-assisted group exhibited a higher error rate when classifying cases where the AI provided systematically erroneous advice, both in the AI-assisted and the subsequent unassisted phase, suggesting the persistence of AI-induced bias. Our study emphasizes the need for careful implementation and continuous evaluation of AI systems in education and training to mitigate potential negative impacts on trainee learning outcomes. • Machine-human knowledge transmission received little attention from prior research. • Our study explored AI-driven upskilling and bias transmission in medical trainees. • Students relied on the AI and mimicked its bias even time after the AI was removed. • Trainees learnt from the AI not only biased but also correct patterns of response. • The research highlights AI’s potential to be both a good and a bad mentor.
Lucia Vicente, Helena Matute, Caterina Fregosi, Federico Cabitza
Int. J. Hum. Comput. Stud.4
2025 Uncovering hidden subtypes in dementia: An unsupervised machine learning approach to dementia diagnosis and personalization of care
Andrea Campagner, Luca Marconi, Edoardo Bianchi, Beatrice Arosio, Paolo Rossi, Giorgio Annoni, Tiziano Angelo Lucchi, Nicola Montano, Federico Cabitza
J. Biomed. Informatics9
2025 Five Degrees of Separation: Investigating the Unexpected Potential of Displaced Human-AI Collaboration Protocols for Apter AI Support
abstract
The integration of AI into decision-making processes offers substantial benefits, particularly in enhancing accuracy and efficiency. However, long-term consequences, such as over-reliance, skill erosion, and loss of human agency, present significant challenges. This study investigates various human-AI collaboration protocols~-~traditional, inhibition, displacement, and replacement~-~across multiple medical settings, including radiological imaging, ECG, and endoscopy. We introduce a novel framework that includes a choice nomogram and qualitative assessment tool, designed to optimize both decision accuracy and socio-technical impacts. Our findings reveal that the displacement protocol consistently outperformed others in several contexts, achieving 87% accuracy in MRI analysis, 89% in x-ray reading and 85% in endoscopy; conversely, the traditional protocol was most effective only in ECG analysis, with 82% accuracy. These results demonstrate that no single protocol is universally optimal, highlighting the need for context-specific selection to ensure effective and sustainable AI-supported decision-making, with a focus on balancing short-term performance with long-term human factors.
Federico Cabitza, Andrea Campagner, Caterina Fregosi, Matteo Cameli, Enrico Gallazzi, Luca Maria Sconfienza, Gian Eugenio Tontini
Proc. ACM Hum. Comput. Interact.1
2024 Dissimilar Similarities: Comparing Human and Statistical Similarity Evaluation in Medical AI
Federico Cabitza, Lorenzo Famiglini, Andrea Campagner, Luca Maria Sconfienza, Stefano Fusco, Valerio Caccavella, Enrico Gallazzi
MDAI1
2024 Never tell me the odds: Investigating pro-hoc explanations in medical decision making
abstract
This paper examines a kind of explainable AI, centered around what we term pro-hoc explanations, that is a form of support that consists of offering alternative explanations (one for each possible outcome) instead of a specific post-hoc explanation following specific advice. Specifically, our support mechanism utilizes explanations by examples, featuring analogous cases for each category in a binary setting. Pro-hoc explanations are an instance of what we called frictional AI, a general class of decision support aimed at achieving a useful compromise between the increase of decision effectiveness and the mitigation of cognitive risks, such as over-reliance, automation bias and deskilling. To illustrate an instance of frictional AI, we conducted an empirical user study to investigate its impact on the task of radiological detection of vertebral fractures in x-rays. Our study engaged 16 orthopedists in a 'human-first, second-opinion' interaction protocol. In this protocol, clinicians first made initial assessments of the x-rays without AI assistance and then provided their final diagnosis after considering the pro-hoc explanations. Our findings indicate that physicians, particularly those with less experience, perceived pro-hoc XAI support as significantly beneficial, even though it did not notably enhance their diagnostic accuracy. However, their increased confidence in final diagnoses suggests a positive overall impact. Given the promisingly high effect size observed, our results advocate for further research into pro-hoc explanations specifically, and into the broader concept of frictional AI.
Federico Cabitza, Chiara Natali, Lorenzo Famiglini, Andrea Campagner, Valerio Caccavella, Enrico Gallazzi
Artif. Intell. Medicine1
2024 Invisible to Machines: Designing AI that Supports Vision Work in Radiology
abstract
Abstract In this article we provide an analysis focusing on clinical use of two deep learning-based automatic detection tools in the field of radiology. The value of these technologies conceived to assist the physicians in the reading of imaging data (like X-rays) is generally assessed by the human-machine performance comparison, which does not take into account the complexity of the interpretation process of radiologists in its social, tacit and emotional dimensions. In this radiological vision work, data which informs the physician about the context surrounding a visible anomaly are essential to the definition of its pathological nature. Likewise, experiential data resulting from the contextual tacit knowledge that regulates professional conduct allows for the assessment of an anomaly according to the radiologist’s, and patient’s, experience. These data, which remain excluded from artificial intelligence processing, question the gap between the norms incorporated by the machine and those leveraged in the daily work of radiologists. The possibility that automated detection may modify the incorporation or the exercise of tacit knowledge raises questions about the impact of AI technologies on medical work. This article aims to highlight how the standards that emerge from the observation practices of radiologists challenge the automation of their vision work, but also under what conditions AI technologies are considered “objective” and trustworthy by professionals.
Giulia Anichini, Chiara Natali, Federico Cabitza
Comput. Support. Cooperative Work.3
2024 Ensemble Predictors: Possibilistic Combination of Conformal Predictors for Multivariate Time Series Classification
abstract
In this article we propose a conceptual framework to study ensembles of conformal predictors (CP), that we call Ensemble Predictors (EP). Our approach is inspired by the application of imprecise probabilities in information fusion. Based on the proposed framework, we study, for the first time in the literature, the theoretical properties of CP ensembles in a general setting, by focusing on simple and commonly used possibilistic combination rules. We also illustrate the applicability of the proposed methods in the setting of multivariate time-series classification, showing that these methods provide better performance (in terms of both robustness, conservativeness, accuracy and running time) than both standard classification algorithms and other combination rules proposed in the literature, on a large set of benchmarks from the UCR time series archive.
Andrea Campagner, Marília Barandas, Duarte Folgado, Hugo Gamboa, Federico Cabitza
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Toward a Perspectivist Turn in Ground Truthing for Predictive Computing
abstract
Most current Artificial Intelligence applications are based on supervised Machine Learning (ML), which ultimately grounds on data annotated by small teams of experts or large ensemble of volunteers. The annotation process is often performed in terms of a majority vote, however this has been proved to be often problematic by recent evaluation studies. In this article, we describe and advocate for a different paradigm, which we call perspectivism: this counters the removal of disagreement and, consequently, the assumption of correctness of traditionally aggregated gold-standard datasets, and proposes the adoption of methods that preserve divergence of opinions and integrate multiple perspectives in the ground truthing process of ML development. Drawing on previous works which inspired it, mainly from the crowdsourcing and multi-rater labeling settings, we survey the state-of-the-art and describe the potential of our proposal for not only the more subjective tasks (e.g. those related to human language) but also those tasks commonly understood as objective (e.g. medical decision making). We present the main benefits of adopting a perspectivist stance in ML, as well as possible disadvantages, and various ways in which such a stance can be implemented in practice. Finally, we share a set of recommendations and outline a research agenda to advance the perspectivist stance in ML.
Federico Cabitza, Andrea Campagner, Valerio Basile
AAAI1
2023 Let Me Think! Investigating the Effect of Explanations Feeding Doubts About the AI Advice
Federico Cabitza, Andrea Campagner, Lorenzo Famiglini, Chiara Natali, Valerio Caccavella, Enrico Gallazzi
CD-MAKE1
2023 Controllable AI - An Alternative to Trustworthiness in Complex AI Systems?
abstract
Abstract The release of ChatGPT to the general public has sparked discussions about the dangers of artificial intelligence (AI) among the public. The European Commission’s draft of the AI Act has further fueled these discussions, particularly in relation to the definition of AI and the assignment of risk levels to different technologies. Security concerns in AI systems arise from the need to protect against potential adversaries and to safeguard individuals from AI decisions that may harm their well-being. However, ensuring secure and trustworthy AI systems is challenging, especially with deep learning models that lack explainability. This paper proposes the concept of Controllable AI as an alternative to Trustworthy AI and explores the major differences between the two. The aim is to initiate discussions on securing complex AI systems without sacrificing practical capabilities or transparency. The paper provides an overview of techniques that can be employed to achieve Controllable AI. It discusses the background definitions of explainability, Trustworthy AI, and the AI Act. The principles and techniques of Controllable AI are detailed, including detecting and managing control loss, implementing transparent AI decisions, and addressing intentional bias or backdoors. The paper concludes by discussing the potential applications of Controllable AI and its implications for real-world scenarios.
Peter Kieseberg, Edgar R. Weippl, A Min Tjoa, Federico Cabitza, Andrea Campagner, Andreas Holzinger
CD-MAKE4
2023 The Tower of Babel in Explainable Artificial Intelligence (XAI)
abstract
Abstract As machine learning (ML) has emerged as the predominant technological paradigm for artificial intelligence (AI), complex black box models such as GPT-4 have gained widespread adoption. Concurrently, explainable AI (XAI) has risen in significance as a counterbalancing force. But the rapid expansion of this research domain has led to a proliferation of terminology and an array of diverse definitions, making it increasingly challenging to maintain coherence. This confusion of languages also stems from the plethora of different perspectives on XAI, e.g. ethics, law, standardization and computer science. This situation threatens to create a “tower of Babel” effect, whereby a multitude of languages impedes the establishment of a common (scientific) ground. In response, this paper first maps different vocabularies, used in ethics, law and standardization. It shows that despite a quest for standardized, uniform XAI definitions, there is still a confusion of languages. Drawing lessons from these viewpoints, it subsequently proposes a methodology for identifying a unified lexicon from a scientific standpoint. This could aid the scientific community in presenting a more unified front to better influence ongoing definition efforts in law and standardization, often without enough scientific representation, which will shape the nature of AI and XAI in the future.
David Schneeberger, Richard Röttger, Federico Cabitza, Andrea Campagner, Markus Plass, Heimo Müller, Andreas Holzinger
CD-MAKE3
2023 AI Shall Have No Dominion: on How to Measure Technology Dominance in AI-supported Human decision-making
abstract
In this article, we propose a conceptual and methodological framework for measuring the impact of the introduction of AI systems in decision settings, based on the concept of technological dominance, i.e. the influence that an AI system can exert on human judgment and decisions. We distinguish between a negative component of dominance (automation bias) and a positive one (algorithm appreciation) by focusing on and systematizing the patterns of interaction between human judgment and AI support, or reliance patterns, and their associated cognitive effects. We then define statistical approaches for measuring these dimensions of dominance, as well as corresponding qualitative visualizations. By reporting about four medical case studies, we illustrate how the proposed methods can be used to inform assessments of dominance and of related cognitive biases in real-world settings. Our study lays the groundwork for future investigations into the effects of introducing AI support into naturalistic and collaborative decision-making.
Federico Cabitza, Andrea Campagner, Riccardo Angius, Chiara Natali, Carlo Reverberi
CHI1
2023 Towards a Rigorous Calibration Assessment Framework: Advancements in Metrics, Methods, and Use
abstract
Calibration is paramount in developing and validating Machine Learning models, particularly in sensitive domains such as medicine. Despite its significance, existing metrics to assess calibration have been found to have shortcomings in regard to their interpretation and theoretical properties. This article introduces a novel and comprehensive framework to assess the calibration of Machine and Deep Learning models that addresses the above limitations. The proposed framework is based on a modification of the Expected Calibration Error (ECE), called the Estimated Calibration Index (ECI), which grounds on and extends prior research. ECI was initially formulated for binary settings, and we adapted it to fit multiclass settings. ECI offers a more nuanced, both locally and globally, and informative measure of a model’s tendency towards over/underconfidence. The paper first outlines the issues related to the prevalent definitions of ECE, including potential biases that may arise in the evaluation of their measures. Then, we present the results of a series of experiments conducted to demonstrate the effectiveness of the proposed framework in supporting a more accurate understanding of a model’s calibration level. Additionally, we discuss how to address and potentially mitigate some biases in calibration assessment.
Lorenzo Famiglini, Andrea Campagner, Federico Cabitza
ECAI3
2023 The Impact of Gender and Personality in Human-AI Teaming: The Case of Collaborative Question Answering
Frida Milella, Chiara Natali, Teresa Scantamburlo, Andrea Campagner, Federico Cabitza
INTERACT (2)5
2023 Rams, hounds and white boxes: Investigating human-AI collaboration protocols in medical diagnosis
abstract
In this paper, we study human-AI collaboration protocols, a design-oriented construct aimed at establishing and evaluating how humans and AI can collaborate in cognitive tasks. We applied this construct in two user studies involving 12 specialist radiologists (the knee MRI study) and 44 ECG readers of varying expertise (the ECG study), who evaluated 240 and 20 cases, respectively, in different collaboration configurations. We confirm the utility of AI support but find that XAI can be associated with a "white-box paradox", producing a null or detrimental effect. We also find that the order of presentation matters: AI-first protocols are associated with higher diagnostic accuracy than human-first protocols, and with higher accuracy than both humans and AI alone. Our findings identify the best conditions for AI to augment human diagnostic skills, rather than trigger dysfunctional responses and cognitive biases that can undermine decision effectiveness.
Federico Cabitza, Andrea Campagner, Luca Ronzio, Matteo Cameli, Giulia Elena Mandoli, Maria Concetta Pastore, Luca Maria Sconfienza, Duarte Folgado, Marília Barandas, Hugo Gamboa
Artif. Intell. Medicine1
2023 Quod erat demonstrandum? - Towards a typology of the concept of explanation for the design of explainable AI
abstract
In this paper, we present a fundamental framework for defining different types of explanations of AI systems and the criteria for evaluating their quality. Starting from a structural view of how explanations can be constructed, i.e., in terms of an explanandum (what needs to be explained), multiple explanantia (explanations, clues, or parts of information that explain), and a relationship linking explanandum and explanantia, we propose an explanandum-based typology and point to other possible typologies based on how explanantia are presented and how they relate to explanandia. We also highlight two broad and complementary perspectives for defining possible quality criteria for assessing explainability: epistemological and psychological (cognitive). These definition attempts aim to support the three main functions that we believe should attract the interest and further research of XAI scholars: clear inventories, clear verification criteria, and clear validation methods.
Federico Cabitza, Andrea Campagner, Gianclaudio Malgieri, Chiara Natali, David Schneeberger, Karl Stöger, Andreas Holzinger
Expert Syst. Appl.1
2022 Global Interpretable Calibration Index, a New Metric to Estimate Machine Learning Models' Calibration
Federico Cabitza, Andrea Campagner, Lorenzo Famiglini
CD-MAKE1
2022 Color Shadows (Part I): Exploratory Usability Evaluation of Activation Maps in Radiological Machine Learning
Federico Cabitza, Andrea Campagner, Lorenzo Famiglini, Enrico Gallazzi, Giovanni Andrea La Maida
CD-MAKE1
2022 Re-calibrating Machine Learning Models Using Confidence Interval Bounds
Andrea Campagner, Lorenzo Famiglini, Federico Cabitza
MDAI3
2021 Prediction of ICU admission for COVID-19 patients: a Machine Learning approach based on Complete Blood Count data
abstract
In this article we discuss the development of prognostic Machine Learning (ML) models for COVID-19 progression: specifically, we address the task of predicting intensive care unit (ICU) admission in the next 5 days. We developed three ML models on the basis of 4995 Complete Blood Count (CBC) tests. We propose three ML models that differ in terms of interpretability: two fully interpretable models and a black-box one. We report an AUC of. 81 and. 83 for the interpretable models (the decision tree and logistic regression, respectively), and an AUC of. 88 for the black-box model (an ensemble). This shows that CBC data and ML methods can be used for cost-effective prediction of ICU admission of COVID-19 patients: in particular, as the CBC can be acquired rapidly through routine blood exams, our models could also be applied in resource-limited settings and to get fast indications at triage and daily rounds.
Lorenzo Famiglini, Giorgio Bini, Anna Carobene, Andrea Campagner, Federico Cabitza
CBMS5
2021 Weighted Utility: A Utility Metric Based on the Case-Wise Raters' Perceptions
Andrea Campagner, Enrico Conte, Federico Cabitza
CD-MAKE3
2021 The need to move away from agential-AI: Empirical investigations, useful concepts and open issues
Federico Cabitza, Andrea Campagner, Carla Simone
Int. J. Hum. Comput. Stud.1
2021 Three-way decision and conformal prediction: Isomorphisms, differences and theoretical properties of cautious learning approaches
abstract
The aim of this article is to study the relationship between two popular Cautious Learning approaches, namely: Three-way decision (TWD) and conformal prediction (CP). Based on the novel proposal of a technique to transform three-way decision classifiers into conformal predictors, and vice versa, we provide conditions for the equivalence between TWD and CP. These theoretical results provide error-bound guarantees for TWD, together with a formal construction to define cost-sensitive cautious classifiers based on CP. The proposed techniques are then applied and evaluated on a collection of benchmark and real-world datasets. The results of the experiments show that the proposed techniques can be used to obtain cautious learning classifiers that are competitive with, and often out-perform, state-of-the-art approaches. Further, through a qualitative medical case study we discuss the usefulness of cautious learning in the development of robust Machine Learning.
Andrea Campagner, Federico Cabitza, Pedro Berjano, Davide Ciucci
Inf. Sci.2
2021 Ground truthing from multi-rater labeling with three-way decision and possibility theory
Andrea Campagner, Davide Ciucci, Carl-Magnus Svensson, Marc Thilo Figge, Federico Cabitza
Inf. Sci.5
2020 Back to the Feature: A Neural-Symbolic Perspective on Explainable AI
Andrea Campagner, Federico Cabitza
CD-MAKE2
2020 Ensemble Learning, Social Choice and Collective Intelligence - An Experimental Comparison of Aggregation Techniques
Andrea Campagner, Davide Ciucci, Federico Cabitza
MDAI3
2020 Trading off between control and autonomy: a narrative review around de-design
abstract
In this work, we provide an overview of contemporary perspectives of design that may challenge the traditional design of IT and socio-technical systems. Our starting metaphor is that of ‘wicked problems’, where the singularity, incompleteness and intrinsic uncertainty of real world settings foregrounds how the worldview that designers offer to practitioners may be optimal in theory but useless in practice. To go beyond traditional notions of design and designer, we intercepted insights coming from minoritarian voices in both theoretic and practice-based design fields. ‘De-design’ is a term we coined to encompass this wide spectrum of approaches that make more resilient and sustainable information artifact, de-emphasize design as a theoretical construct, and reconsider practice as the leading principle of digital innovation. This paper is a narrative review of voices in an extensive array of fields: from Information Systems to Human-Computer Interaction, from End-User Development to Critical Design, from Software Design to Design Studies. Our contribution retraces the motivational roots of de-design and tries to characterise de-design by filling relational gaps between disparate approaches and by bringing them back to IT and socio-technical design, to make digital artifacts sustainable in all of the new environmental, organisational and cultural spaces near to come.
Federico Cabitza, Angela Locoro, Aurelio Ravarini
Behav. Inf. Technol.1
2020 The three-way-in and three-way-out framework to treat and exploit ambiguity in data
Andrea Campagner, Federico Cabitza, Davide Ciucci
Int. J. Approx. Reason.2
2019 Digitizing the Informed Consent: the Challenges to Design for Practices
abstract
This paper reports a user study performed to assess the usability of a Web-based electronic informed consent application called DICE, which is aimed at supporting patients in the process of reading, understanding and using the informed consent as a trigger for further interaction with the team of care givers. In particular, we performed a questionnaire-based study and a series of individual semi-structured interviews to understand whether the application is usable and can be used in real-world settings, respectively. We found that patients could appreciate the availability of interactive tools like DICE, but health professionals believe that its actual adoption in current workflows and practices could be hampered by the chronic lack of time and health operators who could timely address the licit requests that such a tool could bring to light.
Michela Assale, Erica Barbero, Federico Cabitza
CBMS3
2019 New Frontiers in Explainable AI: Understanding the GI to Interpret the GO
Federico Cabitza, Andrea Campagner, Davide Ciucci
CD-MAKE1
2019 Biases Affecting Human Decision Making in AI-Supported Second Opinion Settings
Federico Cabitza
MDAI1
2019 Programmed Inefficiencies in DSS-Supported Human Decision Making
Federico Cabitza, Andrea Campagner, Davide Ciucci, Andrea Seveso
MDAI1
2019 Repetita Iuvant: Exploring and Supporting Redundancy in Hospital Practices
Federico Cabitza, Gunnar Ellingsen, Angela Locoro, Carla Simone
Comput. Support. Cooperative Work.1
2019 Personal Health Records and Patient-Oriented Infrastructures: Building Technology, Shaping (New) Patients, and Healthcare Practitioners
Enrico Maria Piras, Federico Cabitza, Myriam Lewkowicz, Liam J. Bannon
Comput. Support. Cooperative Work.2
2017 Exploiting collective knowledge with three-way decision theory: Cases from the questionnaire-based research
Federico Cabitza, Davide Ciucci, Angela Locoro
Int. J. Approx. Reason.1
2017 Rule-based tools for the configuration of ambient intelligence systems: a comparative user study
Federico Cabitza, Daniela Fogli, Rosa Lanzilotti, Antonio Piccinno
Multim. Tools Appl.1
2016 Valuable Visualization of Healthcare Information: From the Quantified Self Data to Conversations
abstract
Big data analytics in healthcare would be almost useless, without suitable tools allowing users "see" them, and gain insight for their situated decisions. The VVH (Valuable Visualization in Healthcare) workshop focuses on the role of interactive data visualization tools by which people can make sense of healthcare data; these data include sensor data, the messages exchanged in social media, the emails between patients and their doctors, the content of patient records as well as the discussions among different specialists that led to such record content. All these data are used by different types of users, like doctors, nurses, policy makers and common citizens. The VVH workshop aims at contributing on: the assessment of the usability of advanced interactive tools of health-related data visualization; the assessment of the quality of the information and value for insight that these tools make available to their users; the collection of reports of either success stories or failures in the appropriation and use of complex and multidimensional healthcare datasets; the collection of methodological and design-oriented contributions that could share methods, techniques, and heuristics for the design of interactive tools and applications supporting data work, data telling and data interpretation in healthcare.
Federico Cabitza, Angela Locoro, Daniela Fogli, Massimiliano Giacomin
AVI1
2016 Moving Western Neighborliness to East? A study on Local Exchange in Bangladesh
abstract
This paper focuses on the main question whether social media specifically conceived to enable local exchange trading schema can be adopted in different contexts than the western digitized society, where those systems have been considered a feasible alternative to money-based capitalism. We report a qualitative study employing focus groups to study the factors which may affect the adoption of these social media in Bangladesh, a developing country that exhibits characteristics such as strong young unemployment, gender-oriented underemployment, aging population, but also a reduced access to the service economy due to the lack of spare time. The benefits of local exchange seem to be particularly fitting the urban and social structure of Bangladesh.
Federico Cabitza, Angela Locoro, Carla Simone, Tunazzina Sultana
CSCW1
2016 Touch&Screen: widget collection for large screens controlled through smartphones
abstract
We present Touch&Screen, a wide set of interaction techniques for the remote control of widgets (menu, lists, videos, maps etc.) for large screens through smartphones. After presenting the design of these widgets and the related control interfaces of the smartphone, we evaluated the interaction through two user studies. The first study (48 users) aimed to evaluate the user experience (naturalness, usability, etc.) and spontaneity of use. The second user study (12 users) aimed to evaluate our interaction techniques comparatively with direct touch and a commercial smartphone application designed to control cursors on distant screens. The evaluation results show that Touch&Screen is faster and more natural than existing solutions to control distant screens.
Alessio Bellino, Federico Cabitza, Giorgio De Michelis, Flavio De Paoli
MUM2
2015 On a QUESt for a web-based tool promoting knowledge-sharing in medical communities
abstract
The paper reports on the design and development of QUESt, a platform that is aimed at enabling lay users to deploy web-based multi-page dynamic questionnaires. The platform requires little effort and no programming skills, as it uses a simple configuration file that can be expressed in an almost unstructured and text-based manner. We have validated the approach and platform in the healthcare domain, where the questionnaires were intended to solicit and collect both structured and unstructured feedback from large communities of practitioners in response to the sharing and dissemination of relevant case studies; more specifically in this paper we use a qualitative research approach, encompassing evaluation questionnaires and a particular kind of focus group, and the incremental prototype-based development that led to the current release of QUESt. The paper also reports on the experimentation of the platform in a real context, involving almost 100 orthopaedics, including the post-use evaluation of the participants. In light of this evaluation, we discuss the specific requirements of openness and flexibility that end-users ask for in order to be autonomous in developing their own tools for knowledge-sharing; in particular, we discuss the role of lightweight tools, like QUESt to support the dissemination and discussion of clinical case reports.
Federico Cabitza
Behav. Inf. Technol.1
2014 Knowledge artifacts within knowing communities to foster collective knowledge
abstract
In this paper we present a novel model of knowledge creation and diffusion (viz. the Knowledge-Stream model) in communities of knowledgeable citizens (viz. knowing communities). This model takes into account the individual, social and cultural dimensions of knowledge (what we denote as co-knowledge) to account for the various ways knowledge is "circulated" among people (i.e., members of any social structure); we also propose the concept of IT Knowledge Artifact, as the technological driver enabling such a circulation, and exemplify its main roles in a citizen science project that we are going to undertake in the domains of the urban cultural heritage and the food- and diet-related traditions.
Federico Cabitza, Andrea Cerroni, Carla Simone
AVI1
2013 "Drops Hollowing the Stone": Workarounds as Resources for Better Task-Artifact Fit
Federico Cabitza, Carla Simone
ECSCW1
2013 Determining factors in ICT adoption by MSME's in agriculture clusters: An exploratory case study
abstract
In this paper we consider the case of the ICT adoption and use in an agriculture cluster in Lombardy, a northern region of Italy. At the state of the art, relationships among key factors of adoption and use of ICT in agriculture area received little attention by the academic literature. Thus, in this paper we aim to identify a research model in order to provide evidence of four different research questions concerning the determining factors for ICT adoption. The proposed case study reports and discusses the results obtained by analysing data from a survey of about 600 agricultural farms. Finally, Belief Bayesian Networks (BBNs) are used to analyse the complex influence relationships detected between research variables.
Gianluigi Viscusi, Federico Cabitza, Andrea Maurino, Fabio Stella
RCIS2
2013 Computational Coordination Mechanisms: A tale of a struggle for flexibility
Federico Cabitza, Carla Simone
Comput. Support. Cooperative Work.1
2013 Leveraging underspecification in knowledge artifacts to foster collaborative activities in professional communities
Federico Cabitza, Gianluca Colombo, Carla Simone
Int. J. Hum. Comput. Stud.1
2012 Supporting artifact-mediated discourses through a recursive annotation tool
abstract
This paper focuses on tight communities and specifically on the distributed, mediated discourses that their members articulate around documents and inscribed material artifacts. The paper presents a prototype-based design experience toward the definition of a collaborative annotation tool that is endowed with discourse oriented functionalities whose main characteristics have emerged from case studies we undertook in the healthcare and agricultural domains. In latter domain an initial prototype was proposed and progressively tuned to help users propose modifications to an institutional document through the expression of comments gathered around and about a common artifact, and then build a representative summary of the opinions emerging within the community as a result of this distributed discussion. In light of the reported case study, we discuss a new perspective on this class of annotating applications and the related functionalities that could realize a new simplified model of discourse and foster its adoption in distributed settings and communities of practice.
Federico Cabitza, Carla Simone, Marco P. Locatelli
GROUP1
2012 Providing end-users with a visual editor to make their electronic documents active
abstract
In recent years, visual languages have become increasingly popular in the educational domain. But also researchers involved in other application domains are progressively looking at the potential of visual languages to make difficult tasks, as programming is, easier. In this paper we present the research we have conducted to provide end-users of a document management system with a visual language and the related editor by which to define content-based and context-aware rules. Rules are intended to be the proactive components of electronic documents and the visual rule editor we developed is aimed at empowering end-users and domain experts in making their documents more active with respect to both content, context and the users' interactions.
Federico Cabitza, Iade Gesso, Carla Simone
VL/HCC1
2012 Affording Mechanisms: An Integrated View of Coordination and Knowledge Management
Federico Cabitza, Carla Simone
Comput. Support. Cooperative Work.1
2011 "Remain Faithful to the Earth!"*: Reporting Experiences of Artifact-Centered Design in Healthcare
Federico Cabitza
Comput. Support. Cooperative Work.1
2009 PRODOC: an Electronic Patient Record to Foster Process-Oriented Practices
Federico Cabitza, Carla Simone, Giovanni Zorzato
ECSCW1
2009 Leveraging Coordinative Conventions to Promote Collaboration Awareness
Federico Cabitza, Carla Simone, Marcello Sarini
Comput. Support. Cooperative Work.1
2008 Supporting Practices of Positive Redundancy for Seamless Care
abstract
The paper shows how redundancy can be put at work to play a positive role in facilitating the cognitive and coordinative tasks of clinicians in a ward setting. The main requirement to accomplish this positive function is to allow and support clinicians in annotating the clinical record and making correlations between redundant data explicit. We report an observational study we undertook in the design of coordination mechanisms based on a minimal set of meaningful correlations.
Federico Cabitza, Carla Simone
CBMS1
2007 "...and do it the usual way": fostering awareness of work conventions in document-mediated collaboration
Federico Cabitza, Carla Simone
ECSCW1
2007 Providing awareness through situated process maps: the hospital care case
abstract
Clinical Pathways (CPs) are artifacts that clinicians are increasingly introducing in their practices in order to deal with health problems in the most effective, efficient and agreed way. As a result of an observational study at a Neonatology Intensive Care Unit, we found that most CPs are still paper-based. Although perceived useful even on paper, the physicians advocated a system integrating CPs with the clinical record. Based on their requirements, we present a proposal on how to conceive a computational system that can promote awareness in order to achieve better coordination and committed inclusion of pathways in daily clinical practice.
Federico Cabitza, Marcello Sarini, Carla Simone
GROUP1
2006 Designing Computational Places for Communities within Organizations
abstract
The paper focuses on collaboration that is achieved inside communities thriving within organizations. The nature of the communities and their connections with the institutional technologies embedding the organizational prescriptions require the construction of places where the communities can construct and maintain their local memory, policies and conventions. The paper proposes a framework (i.e., a model and the related language) supporting the construction of such places (i.e., fulcra) and illustrates its use by means of a scenario derived from previous empirical investigations on communities of practice
Federico Cabitza, Marco P. Locatelli, Carla Simone
CollaborateCom1
2006 CASMAS: Supporting Collaboration in Pervasive Environments
abstract
The paper proposes a model to design collaboration supports in pervasive environments through the notion of community. As in human communities, the degree of participation of a member can dynamically change in relation both to her physical location and to her position in the logical space of the applications the community uses to cooperate. The paper shows how this approach is able to make pervasive environments reactive to user behavior on the basis of a semantically rich context representation
Federico Cabitza, Marco P. Locatelli, Marcello Sarini, Carla Simone
PerCom1
2005 When once is not enough: the role of redundancy in a hospital ward setting
abstract
The paper discusses the role of redundancy in hospital ward work on the basis of a field study that focuses on the use of paper artifacts supporting healthcare and its coordination. On the basis of literature and direct observations, we identified different kinds of redundancy, i.e. redundancy of effort, functions and data. Hence, we analyzed how these different forms of redundancy may affect each other and the coordination inside hospital wards. Redundancy plays a positive or negative role depending on various circumstances. This twofold nature defines different requirements for a technology to support healthcare and ward work by preserving practices linked to paper-based artifacts and by unobtrusively augmenting them with computational capabilities.
Federico Cabitza, Marcello Sarini, Carla Simone, Michele Telaro
GROUP1