Simone Stumpf

dblp:47/3113 · DBLP profile ↗
← Back
52ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-6482-1973ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 37 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Predicting Fibromyalgia Pain From Self-reported Health Data Using Time-Series and Deep Learning Models
Delnia Alipour, Olga Perepelkina, Alessandro Vinciarelli, Tahir Janmohamed, Binh P. Nguyen, Simone Stumpf
AIME (1)6
2026 "I think this is fair": Uncovering the Complexities of Stakeholder Decision-Making in AI Fairness Assessment
abstract
Assessing fairness in artificial intelligence (AI) typically involves AI experts who select protected features, fairness metrics, and set fairness thresholds to assess outcome fairness. However, little is known about how stakeholders, particularly those affected by AI outcomes but lacking AI expertise, assess fairness. To address this gap, we conducted a qualitative study with 26 stakeholders without AI expertise, representing potential decision subjects in a credit rating scenario, to examine how they assess fairness when placed in the role of deciding on features with priority, metrics, and thresholds. We reveal that stakeholders' fairness decisions are more complex than typical AI expert practices: they considered features far beyond legally protected features, tailored metrics for specific contexts, set diverse yet stricter fairness thresholds, and even preferred designing customized fairness. Our results extend the understanding of how stakeholders can meaningfully contribute to AI fairness governance and mitigation, underscoring the importance of incorporating stakeholders' nuanced fairness judgments.
Yuri Nakao, Mathieu Chollet, Hiroya Inakoshi, Simone Stumpf
CHI5
2026 Empowering Stakeholders with Participatory Auditing of Predictive AI: Perspectives from End-Users and Decision Subjects without AI Expertise
abstract
Artificial intelligence (AI) applications have become ubiquitous in their impact on individuals and society, highlighting a crucial need for their responsible development. Recent research has called for participatory AI auditing, empowering individuals without AI expertise to audit AI applications throughout the entire AI development pipeline. Our work focuses on investigating how to support these kinds of auditors through participatory AI auditing tools and processes. We conducted a series of co-design workshops, using two health-related predictive AI applications as examples. Our results show that participants wanted to be part of AI audits, and were insightful in identifying the potential impacts of applications, but needed to be assisted in conducting audits, especially how to measure impacts. Importantly, participants provided examples of impacts not considered in current risk/harm taxonomies. Our findings provide implications for the design of tools and processes to empower everyone to contribute to responsible AI development in the future.
Patrizia Di Campli San Vito, Evangelia Fringi, Penny S. Johnston, Leonardo C. T. Bezerra, Marios Aristodemou, Siamak F. Shahandashti, Emily O'Hara, Laura Fiona Whyte, Mark Wong, Ayah Soufan, Yashar Moshfeghi, Simone Stumpf
CHI13
2026 "It's a Bit of Me": Predicting Persuasion from Writing Style and Persuader-Receiver Similarity
abstract
Personalising persuasive technologies is essential for driving effective behaviour change. Personalisation often involves building user profiles from personal data, such as demographics or personality. Drawing from previous work exploring the influence of writing style on agreement in negotiations and debates, this paper investigates whether writing style and persuader-receiver similarity predicts persuasion. We asked 239 participants to rate 217 persuasive statements about sustainable behaviour and captured their own writing style through a writing task. Our results showed that the receiver’s writing style, for example negations and social references, significantly predicted persuasion, outperforming widely used profiling criteria such as demographics and personality. Persuader-receiver similarity across various writing features, such as tone and analytic language, was a stronger predictor of persuasion than the statement’s features alone, suggesting that persuasion is more about the match between a message and its receiver than the message alone. For almost half of the features, higher similarity predicted lower persuasion, suggesting that people may prefer complementary qualities over exact linguistic matches. Our work paves the way for persuasive technologies which adapt their writing style to individual users.
Elena Minucci, Simone Stumpf, Martin Lages
UMAP2
2025 Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems
abstract
Representation bias is one of the most common types of biases in artificial intelligence (AI) systems, causing AI models to perform poorly on underrepresented data segments. Although AI practitioners use various methods to reduce representation bias, their effectiveness is often constrained by insufficient domain knowledge in the debiasing process. To address this gap, this paper introduces a set of generic design guidelines for effectively involving domain experts in representation debiasing. We instantiated our proposed guidelines in a healthcare-focused application and evaluated them through a comprehensive mixed-methods user study with 35 healthcare experts. Our findings show that involving domain experts can reduce representation bias without compromising model accuracy. Based on our findings, we also offer recommendations for developers to build robust debiasing systems guided by our generic design guidelines, ensuring more effective inclusion of domain experts in the debiasing process.
Aditya Bhattacharya, Simone Stumpf, Robin De Croon, Katrien Verbert
CHI2
2025 (Un)sustainable Personalities: The Role of Personality When Persuading to Adopt Sustainable Behaviours
Elena Minucci, Martin Lages, Simone Stumpf
PERSUASIVE3
2025 EARN Fairness: Explaining, Asking, Reviewing, and Negotiating Artificial Intelligence Fairness Metrics Among Stakeholders
abstract
Numerous fairness metrics have been proposed and employed by artificial intelligence (AI) experts to quantitatively measure bias and define fairness in AI models. Recognizing the need to accommodate stakeholders' diverse fairness understandings, efforts are underway to solicit their input. However, conveying AI fairness metrics to stakeholders without AI expertise, capturing their personal preferences, and seeking a collective consensus remain challenging and underexplored. To bridge this gap, we propose a new framework, EARN ( Explain, Ask, Review, and Negotiate ) Fairness, which facilitates collective metric decisions among stakeholders without requiring AI expertise. The framework features an adaptable interactive system and a stakeholder-centered EARN Fairness process to Explain fairness metrics, Ask stakeholders' personal metric preferences, Review metrics collectively, and Negotiate a consensus on metric selection. To gather empirical results, we applied the framework to a credit rating scenario and conducted a user study involving 18 decision subjects without AI knowledge. We elicited their personal metric preferences and subsequently we studied how they reached metric consensus in team sessions. Our work shows that the EARN Fairness framework supports stakeholders to express and negotiate fairness preferences, and we provide practical guidance for implementing human-centered AI fairness in high-risk contexts. Through this approach, we aim to reach consensus of fairness perspectives, fostering more equitable and inclusive AI fairness.
Yuri Nakao, Mathieu Chollet, Hiroya Inakoshi, Simone Stumpf
Proc. ACM Hum. Comput. Interact.5
2025 Assessing and Mitigating the Privacy Implications of Eye Tracking on Handheld Mobile Devices
abstract
While gaze data brings benefits like allowing hands-free interaction, it can also reveal sensitive information about people, such as their gender, age, and geographical origin. Privacy leakage and safeguards have been explored for gaze data collected through headsets and stationary eye trackers, but never for gaze data collected through handheld mobile devices, like smartphones. Eye tracking on handheld mobile devices has the potential to be ubiquitous, but gaze data is typically of lower quality, compounded by additional noise and instability due to less controlled environments and screen size constraints. To address this gap, we provide the first evidence of privacy leakage through gaze data collected on handheld mobile devices. In a user study (N=35), we collected our novel SmartEyePhone dataset of gaze data using a smartphone’s front-facing camera. Second, we present the first evaluation and comparison of three Differential Privacy (DP) techniques against our dataset and the TüEyeQ dataset, which was collected in prior work using a stationary remote eye tracker. We found that SmartEyePhone dataset leaks on average 65.5% of private data. DP mechanisms reduce privacy leakage in our data by 22.33% using Laplace mechanism, 10.43% using Exponential mechanism, 22.60% using Gaussian mechanism, and 28.34% using AI model perturbation. However, this also reduces the accuracy of the main task prediction, and is impacted by the choice of privacy parameter values. We present insights on the tradeoff between privacy preservation and the practical usefulness of gaze data. Our insights advance the understanding of privacy in mobile settings and pave the way for privacy preserving gaze-enabled handheld mobile devices.
Noora Alsakar, Norah Mohsen T. Alotaibi, Mohamed Khamis, Simone Stumpf
ACM Trans. Priv. Secur.4
2024 EXMOS: Explanatory Model Steering through Multifaceted Explanations and Data Configurations
abstract
Explanations in interactive machine-learning systems facilitate debugging and improving prediction models. However, the effectiveness of various global model-centric and data-centric explanations in aiding domain experts to detect and resolve potential data issues for model improvement remains unexplored. This research investigates the influence of data-centric and model-centric global explanations in systems that support healthcare experts in optimising models through automated and manual data configurations. We conducted quantitative (n=70) and qualitative (n=30) studies with healthcare experts to explore the impact of different explanations on trust, understandability and model improvement. Our results reveal the insufficiency of global model-centric explanations for guiding users during data configuration. Although data-centric explanations enhanced understanding of post-configuration system changes, a hybrid fusion of both explanation types demonstrated the highest effectiveness. Based on our study results, we also present design implications for effective explanation-driven interactive machine-learning systems.
Aditya Bhattacharya, Simone Stumpf, Lucija Gosak, Gregor Stiglic, Katrien Verbert
CHI2
2024 User Characteristics in Explainable AI: The Rabbit Hole of Personalization?
abstract
As Artificial Intelligence (AI) becomes ubiquitous, the need for Explainable AI (XAI) has become critical for transparency and trust among users. A significant challenge in XAI is catering to diverse users, such as data scientists, domain experts, and end-users. Recent research has started to investigate how users’ characteristics impact interactions with and user experience of explanations, with a view to personalizing XAI. However, are we heading down a rabbit hole by focusing on unimportant details? Our research aimed to investigate how user characteristics are related to using, understanding, and trusting an AI system that provides explanations. Our empirical study with 149 participants who interacted with an XAI system that flagged inappropriate comments showed that very few user characteristics mattered; only age and the personality trait openness influenced actual understanding. Our work provides evidence to reorient user-focused XAI research and question the pursuit of personalized XAI based on fine-grained user characteristics.
Robert Nimmo, Marios Constantinides, Ke Zhou 0003, Daniele Quercia, Simone Stumpf
CHI5
2023 A Multi-perspective Panel on User-Centred Transparency, Explainability, and Controllability in Automations
Philippe A. Palanque, Fabio Paternò, Virpi Roto, Albrecht Schmidt 0001, Simone Stumpf, Jürgen Ziegler 0001
INTERACT (4)5
2023 Towards Responsible AI: A Design Space Exploration of Human-Centered Artificial Intelligence User Interfaces to Investigate Fairness
abstract
With Artificial intelligence (AI) to aid or automate decision-making advancing rapidly, a particular concern is its fairness. In order to create reliable, safe and trustworthy systems through human-centred artificial intelligence (HCAI) design, recent efforts have produced user interfaces (UIs) for AI experts to investigate the fairness of AI models. In this work, we provide a design space exploration that supports not only data scientists but also domain experts to investigate AI fairness. Using loan applications as an example, we held a series of workshops with loan officers and data scientists to elicit their requirements. We instantiated these requirements into FairHIL, a UI to support human-in-the-loop fairness investigations, and describe how this UI could be generalized to other use cases. We evaluated FairHIL through a think-aloud user study. Our work contributes better designs to investigate an AI model’s fairness—and move closer towards responsible AI.
Yuri Nakao, Lorenzo Strappelli, Simone Stumpf, Aisha Naseer, Daniele Regoli, Giulia Del Gamba
Int. J. Hum. Comput. Interact.3
2023 Investigating Privacy Perceptions and Subjective Acceptance of Eye Tracking on Handheld Mobile Devices
abstract
Although eye tracking brings many benefits to users of mobile devices and developers of mobile applications, it poses significant privacy risks to both: the users of mobile devices, and the bystanders that surround users, are within the front-facing camera's field of view. Recent research demonstrates that tracking an individual's gaze reveals personal and sensitive information. This paper presents an investigation of the privacy perceptions and the subjective acceptance of users towards eye tracking on handheld mobile devices. In a four-phase user study (N=17), participants used a smartphone eye tracking app, were interviewed before and after viewing a video showing the amount of sensitive and personal data that could be derived from eye movements, and had their privacy concerns measured. Our findings 1) show factors that influence users' and bystanders' attitudes toward eye tracking on mobile devices such as the algorithms' transparency and the developers' credibility and 2) support designing mechanisms to allow for privacy-aware eye tracking solutions on mobile-devices.
Noora Alsakar, Yasmeen Abdrabou, Simone Stumpf, Mohamed Khamis
Proc. ACM Hum. Comput. Interact.3
2022 Investigating Daily Practices of Self-care to Inform the Design of Supportive Health Technologies for Living and Ageing Well with HIV
abstract
We report on a Diary Study investigating daily practices of Self-care by seven UK adults living with Human Immunodeficiency Virus (HIV), to understand their routines, experiences, needs and concerns, informing Self-care technology design to support living well. We advance a developing HCI literature evidencing how digital tools for self-managing health do not meet the complex needs of those living with long-term conditions, especially those from marginalised communities. Our evaluation of using a Self-care Diary as Design Probe responds to calls to study Self-care practices so that future digital health tools are better grounded in lived experiences of managing multi-morbidity. We contribute to HCI discourses including Personal Health Informatics, Lived Informatics and Reflection by illuminating psychosocial challenges for practicing and self-reporting on Self-care. We offer design implications from a Critical Digital Health perspective, addressing barriers to technology use related to trust, privacy, and representation, gaining new significance during the COVID-19 pandemic.
Caroline Claisse, Bakita Kasadha, Simone Stumpf, Abigail Durrant
CHI3
2022 Toward Involving End-users in Interactive Human-in-the-loop AI Fairness
abstract
Ensuring fairness in artificial intelligence (AI) is important to counteract bias and discrimination in far-reaching applications. Recent work has started to investigate how humans judge fairness and how to support machine learning experts in making their AI models fairer. Drawing inspiration from an Explainable AI approach called explanatory debugging used in interactive machine learning, our work explores designing interpretable and interactive human-in-the-loop interfaces that allow ordinary end-users without any technical or domain background to identify potential fairness issues and possibly fix them in the context of loan decisions. Through workshops with end-users, we co-designed and implemented a prototype system that allowed end-users to see why predictions were made, and then to change weights on features to “debug” fairness issues. We evaluated the use of this prototype system through an online study. To investigate the implications of diverse human values about fairness around the globe, we also explored how cultural dimensions might play a role in using this prototype. Our results contribute to the design of interfaces to allow end-users to be involved in judging and addressing AI fairness through a human-in-the-loop approach.
Yuri Nakao, Simone Stumpf, Subeida Ahmed, Aisha Naseer, Lorenzo Strappelli
ACM Trans. Interact. Intell. Syst.2
2021 Monitoring Quality of Life Indicators at Home from Sparse, and Low-Cost Sensor Data
Dympna O'Sullivan, Rilwan Remilekun Basaru, Simone Stumpf, Neil A. M. Maiden
AIME3
2021 Disability-first Dataset Creation: Lessons from Constructing a Dataset for Teachable Object Recognition with Blind and Low Vision Data Collectors
abstract
Artificial Intelligence (AI) for accessibility is a rapidly growing area, requiring datasets that are inclusive of the disabled users that assistive technology aims to serve. We offer insights from a multi-disciplinary project that constructed a dataset for teachable object recognition with people who are blind or low vision. Teachable object recognition enables users to teach a model objects that are of interest to them, e.g., their white cane or own sunglasses, by providing example images or videos of objects. In this paper, we make the following contributions: 1) a disability-first procedure to support blind and low vision data collectors to produce good quality data, using video rather than images; 2) a validation and evolution of this procedure through a series of data collection phases and 3) a set of questions to orient researchers involved in creating datasets toward reflecting on the needs of their participant community.
Lida Theodorou, Daniela Massiceti, Luisa M. Zintgraf, Simone Stumpf, Cecily Morrison, Edward Cutrell, Matthew Tobias Harris, Katja Hofmann
ASSETS4
2021 ORBIT: A Real-World Few-Shot Dataset for Teachable Object Recognition
abstract
Object recognition has made great advances in the last decade, but predominately still relies on many high-quality training examples per object category. In contrast, learning new objects from only a few examples could enable many impactful applications from robotics to user personalization. Most few-shot learning research, however, has been driven by benchmark datasets that lack the high variation that these applications will face when deployed in the real-world. To close this gap, we present the ORBIT dataset and benchmark, grounded in the real-world application of teachable object recognizers for people who are blind/low-vision. The dataset contains 3,822 videos of 486 objects recorded by people who are blind/low-vision on their mobile phones. The benchmark reflects a realistic, highly challenging recognition problem, providing a rich playground to drive research in robustness to few-shot, high-variation conditions. We set the benchmark’s first state-of-the-art and show there is massive scope for further innovation, holding the potential to impact a broad range of real-world vision applications including tools for the blind/low-vision community. We release the dataset at https://doi.org/10.25383/city.14294597 and benchmark code at https://github.com/microsoft/ORBIT-Dataset.
Daniela Massiceti, Luisa M. Zintgraf, John Bronskill, Lida Theodorou, Matthew Tobias Harris, Edward Cutrell, Cecily Morrison, Katja Hofmann, Simone Stumpf
ICCV9
2021 Interdependence in Action: People with Visual Impairments and their Guides Co-constituting Common Spaces
abstract
Prior work on AI-enabled assistive technology (AT) for people with visual impairments (VI) has treated navigation largely as an independent activity. Consequently, much effort has focused on providing individual users with wayfinding details about the environment, including information on distances, proximity, obstacles, and landmarks. However, independence is also achieved by people with VI through interacting with others, such as in collaboration with sighted guides. Drawing on the concept of interdependence, this research presents a systematic analysis of sighted guiding partnerships. Using interaction analysis as our primary mode of data analysis, we conducted an empirical, qualitative study with 4 couples, each made up of person with a vision impairment and their sighted guide. Our results show how pairs used interactional resources such as turn-taking and body movements to both co-constitute a common space for navigation, and repair moments of rupture to this space. This work is used to present an exemplary case of interdependence and draws out implications for designing AI-enabled AT that shifts the emphasis away from independent navigation, and towards the carefully coordinated actions between people navigating together.
Beatrice Vincenzi, Alex S. Taylor, Simone Stumpf
Proc. ACM Hum. Comput. Interact.3
2020 Investigating the intelligibility of a computer vision system for blind users
abstract
Computer vision systems to help blind users are becoming increasingly common, yet often these systems are not intelligible. Our work investigates the intelligibility of a wearable computer vision system to help blind users locate and identify people in their vicinity. Providing a continuous stream of information, this system allows us to explore intelligibility through interaction and instructions, going beyond studies of intelligibility that focus on explaining a decision a computer vision system might make. In a study with 13 blind users, we explored whether varying instructions (either basic or enhanced) about how the system worked would change blind users' experience of the system. We found offering a more detailed set of instructions did not affect how successful users were using the system nor their perceived workload. We did, however, find evidence of significant differences in what they knew about the system and they employed different, and potentially more effective, use strategies. Our findings have important implications for researchers and designers of computer vision systems for blind users, as well as more general implications for understanding what it means to make interactive computer vision systems intelligible.
Subeida Ahmed, Harshadha Balasubramanian, Simone Stumpf, Cecily Morrison, Abigail Sellen, Martin Grayson
IUI3
2020 Trust, Identity, Privacy, and Security Considerations for Designing a Peer Data Sharing Platform Between People Living With HIV
abstract
Resulting from treatment advances, the Human Immunodeficiency Virus (HIV) is now a long-term condition, and digital solutions are being developed to support people living with HIV in self-management. Sharing their health data with their peers may support self-management, but the trust, identity, privacy and security (TIPS) considerations of people living with HIV remain underexplored. Working with a peer researcher who is expert in the lived experience of HIV, we interviewed 26 people living with HIV in the United Kingdom (UK) to investigate how to design a peer data sharing platform. We also conducted rating activities with participants to capture their attitudes towards sharing personal data. Our mixed methods study showed that participants were highly sophisticated in their understanding of trust and in their requirements for robust privacy and security. They indicated willingness to share digital identity attributes, including gender, age, medical history, health and well-being data, but not details that could reveal their personal identity. Participants called for TIPS measures to foster and to sustain responsible data sharing within their community. These findings can inform the development of trustworthy and secure digital platforms that enable people living with HIV to share data with their peers and provide insights for researchers who wish to facilitate data sharing in other communities with stigmatised health conditions.
Adrian Bussone, Bakita Kasadha, Simone Stumpf, Abigail Durrant, Shema Tariq, Jo Gibbs, Karen C. Lloyd, Jon Bird
Proc. ACM Hum. Comput. Interact.3
2019 Co-Created Personas: Engaging and Empowering Users with Diverse Needs Within the Design Process
abstract
Personas are powerful tools for designing technology and envisioning its usage. They are widely used to imagine archetypal users around whom to orient design work. We have been exploring co-created personas as a technique to use in co-design with users who have diverse needs. Our vision was that this would broaden the demographic and liberate co-designers of their personal relationship with a health condition. This paper reports three studies where we investigated using co-created personas with people who had Parkinson's disease, dementia or aphasia. Observational data of co-design sessions were collected and analysed. Findings revealed that the co-created personas encouraged users with diverse needs to engage with co-designing. Importantly, they also afforded additional benefits including empowering users within a more accessible design process. Reflecting on the outcomes from the different user groups, we conclude with a discussion of the potential for co-created personas to be applied more broadly.
Timothy Neate, Aikaterini Bourazeri, Abi Roper, Simone Stumpf, Stephanie M. Wilson
CHI4
2019 From GenderMag to InclusiveMag: An Inclusive Design Meta-Method
abstract
How can software practitioners assess whether their software supports diverse users? Although there are empirical processes that can be used to find “inclusivity bugs” piecemeal, what is often needed is a systematic inspection method to assess software's support for diverse populations. To help fill this gap, this paper introduces InclusiveMag, a generalization of GenderMag that can be used to generate systematic inclusiveness methods for a particular dimension of diversity. We then present a multicase study covering eight diversity dimensions, of eight teams' experiences applying InclusiveMag to eight under-served populations and their “mainstream” counterparts.
Christopher J. Mendez, Lara Letaw, Margaret M. Burnett, Simone Stumpf, Anita Sarma, Claudia Hilderbrand
VL/HCC4
2019 Monitoring meaningful activities using small low-cost devices in a smart home
Jordan Tewell, Dympna O'Sullivan, Neil A. M. Maiden, James Lockerbie, Simone Stumpf
Pers. Ubiquitous Comput.5
2018 Explainable AI: The New 42?
Randy Goebel, Ajay Chander, Katharina Holzinger, Freddy Lécué, Zeynep Akata, Simone Stumpf, Peter Kieseberg, Andreas Holzinger
CD-MAKE6
2018 Welcome Letter
abstract
Welcome to this issue of the Proceedings of the ACM on Human-Computer Interaction, which will focus on contributions from the research community Engineering Interactive Computing Systems (EICS). This diverse research community explores the methods, processes, techniques and tools that support specifying, designing, developing, deploying and verifying interactive systems. Building interactive systems is a multifaceted and challenging activity, involving a plethora of different actors and roles. This is particularly true in the domain of HCI, where we continuously push the edge of what is possible, where there is a crucial need for adequate processes, tools and methods to build reliable, useful and usable systems that help people cope with the ever-increasing complexity of work and life. The contents of this issue on EICS is the sum of four separate rounds of submissions, evenly spaced from July 2017 through May 2018. In total, the rounds attracted a total of 81 submissions from Asia, Canada, Australia, Europe, Africa, and the United States. Promising submissions in a round that were not accepted were invited to resubmit to a subsequent round, and 6 of the papers appearing in this issue were accepted after at least one round of resubmission. In each round, papers were subject to a rigorous reviewing process where they were reviewed by two EICS senior editors, as well as external reviewers. At the conclusion of each round, a Virtual Committee meeting was held to discuss all of the papers and arrive at final decisions. Ultimately, 14 papers were accepted over all rounds. This issue exists because of the dedicated volunteer effort of 20 senior editors who handled two to four papers each round, and 115 expert reviewers to ensure high quality and insightful reviews for all papers in all rounds. Reviewers and committee members were kept constant as much as possible for papers that were submitted to multiple rounds. Senior members of the editorial group also helped shepherd some papers, reflecting the deep commitment of this research community. We are excited by the detailed and insightful work that resulted in this PACMHCI EICS issue and look forward to equally high quality submissions in subsequent submission cycles over the coming year. For those interested in this area, this group holds their next annual conference June 19-22, 2018 in Paris, France. That conference will provide many opportunities to share ideas with other researchers and practitioners from institutions around the world.
Simone Stumpf, Jeffrey Nichols 0001
Proc. ACM Hum. Comput. Interact.1
2016 Towards the Right Assistance at the Right Time for Using Complex Interfaces
abstract
Many users struggle when they have to use complex interfaces to complete everyday computing tasks. Offering intelligent, proactive assistance is becoming commonplace yet determining the right time to provide help is still difficult. We conducted an empirical study that aimed to uncover what user factors influenced following advice. Our results describe a user's background and expectations that appear to play a role in heeding assistance. Our work is a step towards understanding how to provide the right assistance at the right time and build proactive assistance systems that are personalized for individual users.
Blandine Ginon, Simone Stumpf, Stéphanie Jean-Daubias
AVI2
2016 Crossed Wires: Investigating the Problems of End-User Developers in a Physical Computing Task
abstract
Considerable research has focused on the problems that end users face when programming software, in order to help them overcome their difficulties, but there is little research into the problems that arise in physical computing when end users construct circuits and program them. In an empirical study, we observed end-user developers as they connected a temperature sensor to an Arduino microcontroller and visualized its readings using LEDs. We investigated how many problems participants encountered, the problem locations, and whether they were overcome. We show that most fatal faults were due to incorrect circuit construction, and that often problems were wrongly diagnosed as program bugs. Whereas there are development environments that help end users create and debug software, there is currently little analogous support for physical computing tasks. Our work is a first step towards building appropriate tools that support end-user developers in overcoming obstacles when constructing physical computing artifacts.
Tracey Booth, Simone Stumpf, Jon Bird, Sara Jones 0001
CHI2
2016 User Trust in Intelligent Systems: A Journey Over Time
abstract
Trust is a significant factor in user adoption of new systems. However, although trust is a dynamic attitude of the user towards the system and changes over time, trust in intelligent systems is typically captured as a single quantitative measure at the conclusion of a task. This paper challenges this approach. We report a case study that employed a combination of repeated quantitative and qualitative measures to examine how trust in an intelligent system evolved over time and whether this varied depending on whether the system offered explanations. We discovered different patterns in participants' trust journeys. When provided with explanations, participants' trust levels initially increased, before returning to their original level. Without explanations, participants' trust reduced over time. The qualitative data showed that perceived system ability was more important in determining trust amongst with-explanation participants and perceived transparency was a greater influence on the trust of participants who did not receive explanations. The findings provide a deeper understanding of the development of user trust in intelligent systems and indicate the value of the approach adopted.
Daniel Holliday, Stephanie M. Wilson, Simone Stumpf
IUI3
2016 GenderMag: A Method for Evaluating Software's Gender Inclusiveness
abstract
In recent years, research into gender differences has established that individual differences in how people problem-solve often cluster by gender. Research also shows that these differences have direct implications for software that aims to support users' problem-solving activities, and that much of this software is more supportive of problem-solving processes favored (statistically) more by males than by females. However, there is almost no work considering how software practitioners—such as User Experience (UX) professionals or software developers—can find gender-inclusiveness issues like these in their software. To address this gap, we devised the GenderMag method for evaluating problem-solving software from a gender-inclusiveness perspective. The method includes a set of faceted personas that bring five facets of gender difference research to life, and embeds use of the personas into a concrete process through a gender-specialized Cognitive Walkthrough. Our empirical results show that a variety of practitioners who design software—without needing any background in gender research—were able to use the GenderMag method to find gender-inclusiveness issues in problem-solving software. Our results also show that the issues the practitioners found were real and fixable. This work is the first systematic method to find gender-inclusiveness issues in software, so that practitioners can design and produce problem-solving software that is more usable by everyone.
Margaret M. Burnett, Simone Stumpf, Stephann Makri, Laura Beckwith, Irwin Kwan, Anicia N. Peters, Will Jernigan
Interact. Comput.2
2015 Principles of Explanatory Debugging to Personalize Interactive Machine Learning
abstract
How can end users efficiently influence the predictions that machine learning systems make on their behalf? This paper presents Explanatory Debugging, an approach in which the system explains to users how it made each of its predictions, and the user then explains any necessary corrections back to the learning system. We present the principles underlying this approach and a prototype instantiating it. An empirical evaluation shows that Explanatory Debugging increased participants' understanding of the learning system by 52% and allowed participants to correct its mistakes up to twice as efficiently as participants using a traditional learning system.
Todd Kulesza, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf
IUI4
2014 You Are the Only Possible Oracle: Effective Test Selection for End Users of Interactive Machine Learning Systems
abstract
How do you test a program when only a single user, with no expertise in software testing, is able to determine if the program is performing correctly? Such programs are common today in the form of machine-learned classifiers. We consider the problem of testing this common kind of machine-generated program when the only oracle is an end user: e.g., only you can determine if your email is properly filed. We present test selection methods that provide very good failure rates even for small test suites, and show that these methods work in both large-scale random experiments using a “gold standard” and in studies with real users. Our methods are inexpensive and largely algorithm-independent. Key to our methods is an exploitation of properties of classifiers that is not possible in traditional software testing. Our results suggest that it is plausible for time-pressured end users to interactively detect failures-even very hard-to-find failures-without wading through a large number of successful (and thus less useful) tests. We additionally show that some methods are able to find the arguably most difficult-to-detect faults of classifiers: cases where machine learning algorithms have high confidence in an incorrect result.
Alex Groce, Todd Kulesza, Chaoqiang Zhang, Shalini Shamasunder, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf, Shubhomoy Das, Amber Shinsel, Forrest Bice, Kevin McIntosh
IEEE Trans. Software Eng.7
2013 Too much, too little, or just right? Ways explanations impact end users' mental models
abstract
Research is emerging on how end users can correct mistakes their intelligent agents make, but before users can correctly “debug” an intelligent agent, they need some degree of understanding of how it works. In this paper we consider ways intelligent agents should explain themselves to end users, especially focusing on how the soundness and completeness of the explanations impacts the fidelity of end users' mental models. Our findings suggest that completeness is more important than soundness: increasing completeness via certain information types helped participants' mental models and, surprisingly, their perception of the cost/benefit tradeoff of attending to the explanations. We also found that oversimplification, as per many commercial agents, can be a problem: when soundness was very low, participants experienced more mental demand and lost trust in the explanations, thereby reducing the likelihood that users will pay attention to such explanations at all.
Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Sherry Yang 0002, Irwin Kwan, Weng-Keen Wong
VL/HCC2
2013 End-user feature labeling: Supervised and semi-supervised approaches based on locally-weighted logistic regression
Shubhomoy Das, Travis Moore, Weng-Keen Wong, Simone Stumpf, Ian Oberst, Kevin McIntosh, Margaret M. Burnett
Artif. Intell.4
2012 Tell me more?: the effects of mental model soundness on personalizing an intelligent agent
abstract
What does a user need to know to productively work with an intelligent agent? Intelligent agents and recommender systems are gaining widespread use, potentially creating a need for end users to understand how these systems operate in order to fix their agent's personalized behavior. This paper explores the effects of mental model soundness on such personalization by providing structural knowledge of a music recommender system in an empirical study. Our findings show that participants were able to quickly build sound mental models of the recommender system's reasoning, and that participants who most improved their mental models during the study were significantly more likely to make the recommender operate to their satisfaction. These results suggest that by helping end users understand a system's reasoning, intelligent agents may elicit more and better feedback, thus more closely aligning their output with each user's intentions.
Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Irwin Kwan
CHI2
2012 Towards recognizing "cool": can end users help computer vision recognize subjective attributes of objects in images?
abstract
Recent computer vision approaches are aimed at richer image interpretations that extend the standard recognition of objects in images (e.g., cars) to also recognize object attributes (e.g., cylindrical, has-stripes, wet). However, the more idiosyncratic and abstract the notion of an object attribute (e.g., cool car), the more challenging the task of attribute recognition. This paper considers whether end users can help vision algorithms recognize highly idiosyncratic attributes, referred to here as subjective attributes. We empirically investigated how end users recognized three subjective attributes of carscool, cute, and classic. Our results suggest the feasibility of vision algorithms recognizing subjective attributes of objects, but an interactive approach beyond standard supervised learning from labeled training examples is needed.
William Curran, Travis Moore, Todd Kulesza, Weng-Keen Wong, Sinisa Todorovic, Simone Stumpf, Rachel White, Margaret M. Burnett
IUI6
2011 End-User Feature Labeling via Locally Weighted Logistic Regression
abstract
Applications that adapt to a particular end user often make inaccurate predictions during the early stages when training data is limited. Although an end user can improve the learning algorithm by labeling more training data, this process is time consuming and too ad hoc to target a particular area of inaccuracy. To solve this problem, we propose a new learning algorithm based on Locally Weighted Logistic Regression for feature labeling by end users, enabling them to point out which features are important for a class, rather than provide new training instances. In our user study, the first allowing ordinary end users to freely choose features to label directly from text documents, our algorithm was more effective than others at leveraging end users’ feature labels to improve the learning algorithm. Our results strongly suggest that allowing users to freely choose features to label is a promising method for allowing end users to improve learning algorithms effectively.
Weng-Keen Wong, Ian Oberst, Shubhomoy Das, Travis Moore, Simone Stumpf, Kevin McIntosh, Margaret M. Burnett
AAAI5
2011 This image smells good: effects of image information scent in search engine results pages
abstract
Users are confronted with an overwhelming amount of web pages when they look for information on the Internet. Current search engines already aid the user in their information seeking tasks by providing textual results but adding images to results pages could further help the user in judging the relevance of a result. We investigated this problem from an Information Foraging perspective and we report on two empirical studies that focused on the information scent of images. Our results show that images have their own distinct "smell" which is not as strong as that of text. We also found that combining images and text cues leads to a stronger overall scent. Surprisingly, when images were added to search engine results pages, this did not lead our participants to behave significantly differently in terms of effectiveness or efficiency. Even when we added images that could confuse the participants' scent, this had no significantly detrimental impact on their behaviour. However, participants expressed a preference for results pages which included images. We discuss potential challenges and point to future research to ensure the success of adding images to textual results in search engine results pages.
Faidon Loumakis, Simone Stumpf, David Grayson
CIKM2
2011 When users generate music playlists: When words leave off, music begins?
abstract
Music systems that generate playlists are gaining increasing popularity, yet ways to select songs to be acceptable to users is still elusive. We present the results of an explorative study that focused on the language of musically untrained end users for playlist choices, in a variety of listening contexts. Our results indicate that there are a number of opportunities for playlist recommendation or retrieval systems, particularly by taking context into account.
Simone Stumpf, Sam Muscroft
ICME1
2011 End-user feature labeling: a locally-weighted regression approach
abstract
When intelligent interfaces, such as intelligent desktop assistants, email classifiers, and recommender systems, customize themselves to a particular end user, such customizations can decrease productivity and increase frustration due to inaccurate predictions - especially in early stages, when training data is limited. The end user can improve the learning algorithm by tediously labeling a substantial amount of additional training data, but this takes time and is too ad hoc to target a particular area of inaccuracy. To solve this problem, we propose a new learning algorithm based on locally weighted regression for feature labeling by end users, enabling them to point out which features are important for a class, rather than provide new training instances. In our user study, the first allowing ordinary end users to freely choose features to label directly from text documents, our algorithm was both more effective than others at leveraging end users' feature labels to improve the learning algorithm, and more robust to real users' noisy feature labels. These results strongly suggest that allowing users to freely choose features to label is a promising method for allowing end users to improve learning algorithms effectively.
Weng-Keen Wong, Ian Oberst, Shubhomoy Das, Travis Moore, Simone Stumpf, Kevin McIntosh, Margaret M. Burnett
IUI5
2011 Mini-crowdsourcing end-user assessment of intelligent assistants: A cost-benefit study
abstract
Intelligent assistants sometimes handle tasks too important to be trusted implicitly. End users can establish trust via systematic assessment, but such assessment is costly. This paper investigates whether, when, and how bringing a small crowd of end users to bear on the assessment of an intelligent assistant is useful from a cost/benefit perspective. Our results show that a mini-crowd of testers supplied many more benefits than the obvious decrease in workload, but these benefits did not scale linearly as mini-crowd size increased - there was a point of diminishing returns where the cost-benefit ratio became less attractive.
Amber Shinsel, Todd Kulesza, Margaret M. Burnett, William Curran, Alex Groce, Simone Stumpf, Weng-Keen Wong
VL/HCC6
2011 Why-oriented end-user debugging of naive Bayes text classification
abstract
Machine learning techniques are increasingly used in intelligent assistants , that is, software targeted at and continuously adapting to assist end users with email, shopping, and other tasks. Examples include desktop SPAM filters, recommender systems, and handwriting recognition. Fixing such intelligent assistants when they learn incorrect behavior, however, has received only limited attention. To directly support end-user “debugging” of assistant behaviors learned via statistical machine learning, we present a Why-oriented approach which allows users to ask questions about how the assistant made its predictions, provides answers to these “why” questions, and allows users to interactively change these answers to debug the assistant's current and future predictions. To understand the strengths and weaknesses of this approach, we then conducted an exploratory study to investigate barriers that participants could encounter when debugging an intelligent assistant using our approach, and the information those participants requested to overcome these barriers. To help ensure the inclusiveness of our approach, we also explored how gender differences played a role in understanding barriers and information needs. We then used these results to consider opportunities for Why-oriented approaches to address user barriers and information needs.
Todd Kulesza, Simone Stumpf, Weng-Keen Wong, Margaret M. Burnett, Stephen Perona, Amy J. Ko, Ian Oberst
ACM Trans. Interact. Intell. Syst.2
2010 Explanatory Debugging: Supporting End-User Debugging of Machine-Learned Programs
abstract
Many machine-learning algorithms learn rules of behavior from individual end users, such as task-oriented desktop organizers and handwriting recognizers. These rules form a “program” that tells the computer what to do when future inputs arrive. Little research has explored how an end user can debug these programs when they make mistakes. We present our progress toward enabling end users to debug these learned programs via a Natural Programming methodology. We began with a formative study exploring how users reason about and correct a text-classification program. From the results, we derived and prototyped a concept based on “explanatory debugging”, then empirically evaluated it. Our results contribute methods for exposing a learned program's logic to end users and for eliciting user corrections to improve the program's predictions.
Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Weng-Keen Wong, Yann Riche, Travis Moore, Ian Oberst, Amber Shinsel, Kevin McIntosh
VL/HCC2
2010 Explaining how to play real-time strategy games
Ronald A. Metoyer, Simone Stumpf, Christoph Neumann 0003, Jonathan Dodge, Jill Cao, Aaron Schnabel
Knowl. Based Syst.2
2009 Fixing the program my computer learned: barriers for end users, challenges for the machine
abstract
The results of a machine learning from user behavior can be thought of as a program, and like all programs, it may need to be debugged. Providing ways for the user to debug it matters, because without the ability to fix errors users may find that the learned program's errors are too damaging for them to be able to trust such programs. We present a new approach to enable end users to debug a learned program. We then use an early prototype of our new approach to conduct a formative study to determine where and when debugging issues arise, both in general and also separately for males and females. The results suggest opportunities to make machine-learned programs more effective tools.
Todd Kulesza, Weng-Keen Wong, Simone Stumpf, Stephen Perona, Rachel White, Margaret M. Burnett, Ian Oberst, Amy J. Ko
IUI3
2009 Detecting and correcting user activity switches: algorithms and interfaces
abstract
The TaskTracer system allows knowledge workers to define a set of activities that characterize their desktop work. It then associates with each user-defined activity the set of resources that the user accesses when performing that activity. In order to correctly associate resources with activities and provide useful activity-related services to the user, the system needs to know the current activity of the user at all times. It is often convenient for the user to explicitly declare which activity he/she is working on. But frequently the user forgets to do this. TaskTracer applies machine learning methods to detect undeclared activity switches and predict the correct activity of the user. This paper presents TaskPredictor2, a complete redesign of the activity predictor in TaskTracer and its notification user interface. TaskPredictor2 applies a novel online learning algorithm that is able to incorporate a richer set of features than our previous predictors. We prove an error bound for the algorithm and present experimental results that show improved accuracy and a 180-fold speedup on real user data. The user interface supports negotiated interruption and makes it easy for the user to correct both the predicted time of the task switch and the predicted activity.
Jianqiang Shen, Jed Irvine, Xinlong Bao, Michael Goodman, Stephen Kolibaba, Fredric Carl, Brenton Kirschner, Simone Stumpf, Thomas G. Dietterich
IUI9
2009 Interacting meaningfully with machine learning systems: Three experiments
Simone Stumpf, Vidya Rajaram, Lida Li, Weng-Keen Wong, Margaret M. Burnett, Thomas G. Dietterich, Erin Sullivan, Jon Herlocker
Int. J. Hum. Comput. Stud.1
2008 Integrating rich user feedback into intelligent user interfaces
abstract
The potential for machine learning systems to improve via a mutually beneficial exchange of information with users has yet to be explored in much detail. Previously, we found that users were willing to provide a generous amount of rich feedback to machine learning systems, and that the types of some of this rich feedback seem promising for assimilation by machine learning algorithms. Following up on those findings, we ran an experiment to assess the viability of incorporating real-time keyword-based feedback in initial training phases when data is limited. We found that rich feedback improved accuracy but an initial unstable period often caused large fluctuations in classifier behavior. Participants were able to give feedback by relying heavily on system communication in order to respond to changes. The results show that in order to benefit from the user's knowledge, machine learning systems must be able to absorb keyword-based rich feedback in a graceful manner and provide clear explanations of their predictions.
Simone Stumpf, Erin Sullivan, Erin Fitzhenry, Ian Oberst, Weng-Keen Wong, Margaret M. Burnett
IUI1
2007 Toward harnessing user feedback for machine learning
abstract
There has been little research into how end users might be able to communicate advice to machine learning systems. If this resource--the users themselves--could somehow work hand-in-hand with machine learning systems, the accuracy of learning systems could be improved and the users' understanding and trust of the system could improve as well. We conducted a think-aloud study to see how willing users were to provide feedback and to understand what kinds of feedback users could give. Users were shown explanations of machine learning predictions and asked to provide feedback to improve the predictions. We found that users had no difficulty providing generous amounts of feedback. The kinds of feedback ranged from suggestions for reweighting of features to proposals for new features, feature combinations, relational features, and wholesale changes to the learning algorithm. The results show that user feedback has the potential to significantly improve machine learning systems, but that learning algorithms need to be extended in several ways to be able to assimilate this feedback.
Simone Stumpf, Vidya Rajaram, Lida Li, Margaret M. Burnett, Thomas G. Dietterich, Erin Sullivan, Russell Drummond, Jon Herlocker
IUI1
2006 Predicting Task-Specific Webpages for Revisiting
Arwen Twinkle Lettkeman, Simone Stumpf, Jed Irvine, Jon Herlocker
AAAI2
2006 Supporting end-user debugging: what do users want to know?
abstract
Although researchers have begun to explicitly support end-user programmers' debugging by providing information to help them find bugs, there is little research addressing the right content to communicate to these users. The specific semantic content of these debugging communications matters because, if the users are not actually seeking the information the system is providing, they are not likely to attend to it. This paper reports a formative empirical study that sheds light on what end users actually want to know in the course of debugging a spreadsheet, given the availability of a set of interactive visual testing and debugging features. Our results provide in sights into end-user debuggers' information gaps, and further suggest opportunities to improve end-user debugging systems' support for the things end-user debuggers actually want to know.
Cory Kissinger, Margaret M. Burnett, Simone Stumpf, Neeraja Subrahmaniyan, Laura Beckwith, Sherry Yang 0002, Mary Beth Rosson
AVI3
2005 The TaskTracker System
Simone Stumpf, Xinlong Bao, Anton N. Dragunov, Thomas G. Dietterich, Jon Herlocker, Kevin Johnsrude, Lida Li, Jianqiang Shen
AAAI1