Amama Mahmood

dblp:208/3113 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-5044-1494ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Voice Assistants for Health Self-Management: Designing for and with Older Adults
Amama Mahmood, Shiye Cao, Maia Stiber, Victor Nikhil Antony, Chien-Ming Huang 0001
CHI1
2025 ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations
abstract
The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.
Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira, Wendy Ju, Micol Spitale, Hatice Gunes, Chien-Ming Huang 0001
ACM Multimedia3
2025 User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants
Amama Mahmood, Bingsheng Yao, Dakuo Wang, Chien-Ming Huang 0001
Int. J. Hum. Comput. Stud.1
2025 "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking Partner
abstract
The rapid advancement of Large Language Models (LLMs) has created numerous potentials for integration with conversational assistants (CAs) assisting people in their daily tasks, particularly due to their extensive flexibility. However, users' real-world experiences interacting with these assistants remain unexplored. In this research, we chose cooking, a complex daily task, as a scenario to explore people's successful and unsatisfactory experiences while receiving assistance from an LLM-based CA, Mango Mango . We discovered that participants value the system's ability to offer customized instructions based on context, provide extensive information beyond the recipe, and assist them in dynamic task planning. However, users expect the system to be more adaptive to oral conversation and provide more suggestive responses to keep them actively involved. Recognizing that users began treating our LLM-CA as a personal assistant or even a partner rather than just a recipe-reading tool, we propose five design considerations for future development.
Szeyi Chan, Bingsheng Yao, Amama Mahmood, Chien-Ming Huang 0001, Holly Jimison, Elizabeth D. Mynatt, Dakuo Wang
Proc. ACM Hum. Comput. Interact.4
2024 Gender Biases in Error Mitigation by Voice Assistants
abstract
Commercial voice assistants are largely feminized and associated with stereotypically feminine traits such as warmth and submissiveness. As these assistants continue to be adopted for everyday uses, it is imperative to understand how the portrayed gender shapes the voice assistant's ability to mitigate errors, which are still common in voice interactions. We report a study (N=40) that examined the effects of voice gender (feminine, ambiguous, masculine), error mitigation strategies (apology, compensation) and participant's gender on people's interaction behavior and perceptions of the assistant. Our results show that AI assistants that apologized appeared warmer than those offered compensation. Moreover, male participants preferred apologetic feminine assistants over apologetic masculine ones. Furthermore, male participants interrupted AI assistants regardless of perceived gender more frequently than female participants when errors occurred. Our results suggest that the perceived gender of a voice assistant biases user behavior, especially for male users, and that an ambiguous voice has the potential to reduce biases associated with gender-specific traits.
Amama Mahmood, Chien-Ming Huang 0001
Proc. ACM Hum. Comput. Interact.1
2023 Crowdsourcing Thumbnail Captions: Data Collection and Validation
abstract
Speech interfaces, such as personal assistants and screen readers, read image captions to users. Typically, however, only one caption is available per image, which may not be adequate for all situations (e.g., browsing large quantities of images). Long captions provide a deeper understanding of an image but require more time to listen to, whereas shorter captions may not allow for such thorough comprehension yet have the advantage of being faster to consume. We explore how to effectively collect both thumbnail captions—succinct image descriptions meant to be consumed quickly—and comprehensive captions—which allow individuals to understand visual content in greater detail. We consider text-based instructions and time-constrained methods to collect descriptions at these two levels of detail and find that a time-constrained method is the most effective for collecting thumbnail captions while preserving caption accuracy. Additionally, we verify that caption authors using this time-constrained method are still able to focus on the most important regions of an image by tracking their eye gaze. We evaluate our collected captions along human-rated axes—correctness, fluency, amount of detail, and mentions of important concepts—and discuss the potential for model-based metrics to perform large-scale automatic evaluations in the future.
Carlos A. Aguirre, Shiye Cao, Amama Mahmood, Chien-Ming Huang 0001
ACM Trans. Interact. Intell. Syst.3
2022 Owning Mistakes Sincerely: Strategies for Mitigating AI Errors
abstract
Interactive AI systems such as voice assistants are bound to make errors because of imperfect sensing and reasoning. Prior human-AI interaction research has illustrated the importance of various strategies for error mitigation in repairing the perception of an AI following a breakdown in service. These strategies include explanations, monetary rewards, and apologies. This paper extends prior work on error mitigation by exploring how different methods of apology conveyance may affect people’s perceptions of AI agents; we report an online study (N=37) that examines how varying the sincerity of an apology and the assignment of blame (on either the agent itself or others) affects participants’ perceptions and experience with erroneous AI agents. We found that agents that openly accepted the blame and apologized sincerely for mistakes were thought to be more intelligent, likeable, and effective in recovering from errors than agents that shifted the blame to others.
Amama Mahmood, Jeanie W. Fung, Isabel Won, Chien-Ming Huang 0001
CHI1
2022 Crowdsourcing Thumbnail Captions via Time-Constrained Methods
abstract
Speech interfaces, such as personal assistants and screen readers, employ captions to allow users to consume images; however, there is typically only one caption available per image, which may not be adequate for all settings (e.g., browsing large quantities of images). Longer captions require more time to consume, whereas shorter captions may hinder a user’s ability to fully understand the image’s content. We explore how to effectively collect both thumbnail captions—succinct image descriptions meant to be consumed quickly—and comprehensive captions, which allow individuals to understand visual content in greater detail. We consider text-based and time-constrained methods to collect descriptions at these two levels of detail, and find that a time-constrained method is most effective for collecting thumbnail captions while preserving caption accuracy. We evaluate our collected captions along three human-rated axes—correctness, fluency, and level of detail—and discuss the potential for model-based metrics to perform automatic evaluation.
Carlos A. Aguirre, Amama Mahmood, Chien-Ming Huang 0001
IUI2
2022 Effects of rhetorical strategies and skin tones on agent persuasiveness in assisted decision-making
abstract
Appearance and linguistic cues may influence how both people and Intelligent Virtual Agents (IVAs) are perceived and evaluated by others; appearance (e.g., skin tone) has been linked to various implicit biases such as agreeing more with stereotypical attractive faces, while particular linguistic cues may effectively increase persuasiveness. In this paper, we report an online study (N=59) evaluating how strategic linguistic cues (expertise: high vs. low) may shape the implicit disadvantages associated with ethnic stereotypes (skin tone: dark vs. light). We found that a virtual agent with a high level of expertise was considered more persuasive, dominant, intelligent, and likeable regardless of their skin tone, and that participants complied more with IVAs with a darker skin tone. Our results suggest that the design of IVAs requires the deliberate considerations of factors such as appearance and linguistic behaviors in order to achieve intended outcomes.
Amama Mahmood, Chien-Ming Huang 0001
IVA1
2020 Visual Monitoring and Servoing of a Cutting Blade during Telerobotic Satellite Servicing
abstract
We propose a system for visually monitoring and servoing the cutting of a multi-layer insulation (MLI) blanket that covers the envelope of satellites and spacecraft. The main contributions of this paper are: 1) to propose a model for relating visual features describing the engagement depth of the blade to the force exerted on the MLI blanket by the cutting tool, 2) a blade design and algorithm to reliably detect the engagement depth of the blade inside the MLI, and 3) a servoing mechanism to achieve the desired applied force by monitoring the engagement depth. We present results that validate these contributions by comparing forces estimated from visual feedback to measured forces at the blade. We also demonstrate the robustness of the blade design and vision processing under challenging conditions.
Amama Mahmood, Balázs Vágvölgyi, Will Pryor, Louis L. Whitcomb, Peter Kazanzides, Simon Léonard
IROS1