Caterina Fregosi

dblp:388/3156 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0004-7626-8131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Too Sure for Our Own Good: A User Study on AI Confidence and Human Reliance
abstract
Achieving appropriate human reliance on Artificial Intelligence (AI) systems remains a central challenge in Human-Computer Interaction. Confidence scores—indicators of an AI system’s certainty in its recommendations—have been proposed as a means to help users calibrate their trust and reliance on AI Decision Support Systems (DSS). However, limited research has explored how well-calibrated versus miscalibrated confidence scores affect human decision-making. We report a study examining the effects of confidence calibration on user reliance, decision accuracy, and perceived utility of an AI DSS. In a within-subjects experiment involving 184 participants solving logic puzzles, we found that well-calibrated confidence scores significantly improved decision accuracy (+20%, 95% CI: [0.18, 0.23]), whereas miscalibrated scores yielded minimal accuracy gains (+2%, 95% CI: [-0.00, 0.04]) and increased vulnerability to automation bias and conservatism bias. Participants were more likely to accept AI recommendations when high confidence was expressed, even when those recommendations were incorrect, resulting in errors. Conversely, miscalibrated and low-confidence recommendations increased conservatism bias, leading users to reject even accurate AI suggestions. Perceived utility of the AI system was higher when confidence levels were high (p < 0.001) and when confidence was well-calibrated (p = 0.002). These findings underscore the importance of designing AI systems with properly calibrated confidence cues to improve human-AI collaboration and mitigate reliance-related biases.
Caterina Fregosi, Lucia Vicente, Andrea Campagner, Federico Cabitza
AAAI1
2026 Under what influence: Measuring AI influence to fit user profiles in decision-making
abstract
Artificial Intelligence (AI) has become a pivotal tool in augmenting human decision-making across various domains, yet its influence on user decisions often lacks comprehensive evaluation. While technical performance metrics such as accuracy and efficiency dominate AI design, integrating human-centered approaches that consider trust and reliance remains underexplored. This study addresses the knowledge gap in understanding how AI systems influence decision-making quality, calibrated to user profiles, including their expertise, skills, professional role, confidence, and reliance tendencies. We present a novel and comprehensive metric framework for evaluating AI influence, emphasizing behavioral patterns and measurable improvements in decision outcomes beyond simple alignment with AI recommendations. The framework is applied to four medical domain case studies—MRI, ECG, X-ray, and ENDO – with user groups spanning specialists, sub-specialists, and trainees. Results reveal that while human and AI systems achieve high agreement rates (up to 81%), AI influence on decision quality varies significantly. Notably, X-ray decision-making showed the highest influence index (0.27), while MRI decisions exhibited substantial self-anchoring bias (6.94), undermining the potential positive impact of AI. Influence metrics unveiled nuances missed by agreement scores, highlighting domain-specific biases and opportunities to optimize AI-human interaction. This research underscores the necessity for adapting the type of AI system and affordance to user characteristics and attitudes of reliance to foster calibrated trust and improve decision outcomes. Our findings inform the design of AI systems that better support diverse user needs and align with human decisions, driving progress toward human-centered AI integration in high-stakes domains. • Developed novel metrics to evaluate AI’s influence beyond user agreement with AI. • Identified biases impacting AI influence, such as self-anchoring and automation bias. • Applied framework to four medical studies with 330 clinicians and 15,000 decisions. • Revealed up to 81% alignment but variances in appropriate reliance and influence. • Highlighted need for adaptive AI systems to match user expertise. • Demonstrated that influence metrics uncover dynamics missed by traditional reliance.
Andrea Campagner, Caterina Fregosi, Chiara Natali, Federico Cabitza
Int. J. Hum. Comput. Stud.2
2025 From Oracular to Judicial: Enhancing Clinical Decision Making through Contrasting Explanations and a Novel Interaction Protocol
abstract
Clinical Decision Support Systems (CDSS) utilizing machine learning (ML) classifiers have demonstrated substantial potential for improving diagnostic accuracy across various medical domains. However, concerns regarding automation bias, diminished sense of agency, and over-reliance on these systems remain, particularly in clinical settings where decision-making autonomy is critical.To address these challenges, we propose "Judicial AI,"an innovative interaction protocol aimed at reducing automation bias and preserving a sense of agency. This system presents contrasting explanations to medical professionals rather than definitive recommendations, encouraging user engagement and critical evaluation.Before adopting interaction protocols that avoid definitive recommendations, it is important to assess whether such an approach impacts diagnostic accuracy, and if so, how. This paper reports an exploratory study investigating the efficacy of a Judicial CDSS in the diagnosis of vertebral fractures from X-ray images. Sixteen medical professionals, comprising spine surgeons and radiologists, participated in the diagnosis of 18 X-ray images, which were carefully selected to represent particularly difficult and complex cases. Diagnosticians first recorded their decisions independently and then with support from the Judicial AI, which provided activation maps for opposing diagnoses.Our findings show a significant improvement in diagnostic accuracy for complex cases among experienced users (p =.045), with an overall accuracy increase of 0.24. Confidence levels also rose, particularly in the case of complex diagnoses (p =.034). However, the protocol was less beneficial for less experienced users, suggesting that cognitive load might be a limiting factor.These results suggest that Judicial AI, which frames decision-makers as the ultimate authority in the decision-making process, may be an effective tool for mitigating automation bias and preserving a sense of agency in clinical environments.
Federico Cabitza, Lorenzo Famiglini, Caterina Fregosi, Samuele Pe, Enea Parimbelli, Giovanni Andrea La Maida, Enrico Gallazzi
IUI3
2025 Machine learning systems as mentors in human learning: A user study on machine bias transmission in medical training
abstract
While accurate AI systems can enhance human performance, exerting both an augmentation and good mentoring effect, imperfect systems may act as poor mentors, transmitting biases and systematic errors to users. However, there is still limited research on the potential for AI to transmit biases to humans, an effect that could be even more pronounced for less experienced users, such as novices or trainees, making decisions supported by AI-based systems. To investigate the bias transmission effect and the potential of AI to serve as a mentor, we involved eighty-six medical students, dividing them into an AI-assisted group and a control group. We tasked them with classifying simulated tissue samples for a fictitious disease. In the first phase of the task, the AI group received diagnostic advice from a simulated AI system that made systematic errors for a specific type of case, while being accurate for all other types. The control group did not receive any assistance. In the second phase, participants in both groups classified new tissue samples, including ambiguous cases, without any support to test the residual impact of AI bias. The results showed that the AI-assisted group exhibited a higher error rate when classifying cases where the AI provided systematically erroneous advice, both in the AI-assisted and the subsequent unassisted phase, suggesting the persistence of AI-induced bias. Our study emphasizes the need for careful implementation and continuous evaluation of AI systems in education and training to mitigate potential negative impacts on trainee learning outcomes. • Machine-human knowledge transmission received little attention from prior research. • Our study explored AI-driven upskilling and bias transmission in medical trainees. • Students relied on the AI and mimicked its bias even time after the AI was removed. • Trainees learnt from the AI not only biased but also correct patterns of response. • The research highlights AI’s potential to be both a good and a bad mentor.
Lucia Vicente, Helena Matute, Caterina Fregosi, Federico Cabitza
Int. J. Hum. Comput. Stud.3
2025 Five Degrees of Separation: Investigating the Unexpected Potential of Displaced Human-AI Collaboration Protocols for Apter AI Support
abstract
The integration of AI into decision-making processes offers substantial benefits, particularly in enhancing accuracy and efficiency. However, long-term consequences, such as over-reliance, skill erosion, and loss of human agency, present significant challenges. This study investigates various human-AI collaboration protocols~-~traditional, inhibition, displacement, and replacement~-~across multiple medical settings, including radiological imaging, ECG, and endoscopy. We introduce a novel framework that includes a choice nomogram and qualitative assessment tool, designed to optimize both decision accuracy and socio-technical impacts. Our findings reveal that the displacement protocol consistently outperformed others in several contexts, achieving 87% accuracy in MRI analysis, 89% in x-ray reading and 85% in endoscopy; conversely, the traditional protocol was most effective only in ECG analysis, with 82% accuracy. These results demonstrate that no single protocol is universally optimal, highlighting the need for context-specific selection to ensure effective and sustainable AI-supported decision-making, with a focus on balancing short-term performance with long-term human factors.
Federico Cabitza, Andrea Campagner, Caterina Fregosi, Matteo Cameli, Enrico Gallazzi, Luca Maria Sconfienza, Gian Eugenio Tontini
Proc. ACM Hum. Comput. Interact.3
2024 Algorithmic Authority & AI Influence in Decision Settings: Theories and Implications for Design
abstract
This workshop explores the influence of AI systems on human decision-making - algorithmic authority - and the broader concept of technology dominance, which includes both positive and negative impacts of AI reliance. Drawing from diverse fields such as Human-AI Interaction, Sociology, Epistemology, and Cognitive Science, the workshop will discuss theoretical foundations, empirical studies, and design implications of AI’s role in shaping human judgment and behavior. The objectives are to examine in-depth the concepts of algorithmic authority and technology dominance, and identify metrics for their assessment. The workshop aims to foster interdisciplinary collaboration and produce practical design principles that help to counter risks associated to AI technology dominance and thus foster a responsible use of AI systems.
Alessandro Facchini, Caterina Fregosi, Chiara Natali, Alberto Termine, Benjamin Wilson 0002
HAI2