Chien-Ming Huang 0001

dblp:56/3778-1 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
33since 2021 · last 2026
0000-0002-6838-3701ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 39 · 6 first-author · 27 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 16 since 2021Systems, architecture and hardware · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 ELLA: Generative AI-Powered Social Robots for Early Language Development at Home
abstract
Early language development shapes children’s later literacy and learning, yet many families have limited access to scalable, high-quality support at home. Recent advances in generative AI make it possible for social robots to move beyond scripted interactions and engage children in adaptive, conversational activities, but it remains unclear how to design such systems for pre-schoolers and how children engage with them over time in the home. We present ELLA (Early Language Learning Agent), an autonomous, LLM-powered social robot that supports early language development through interactive storytelling, parent-selected language targets, and scaffolded dialogue. Using a multi-phased, human-centered process, we interviewed parents (n=7) and educators (n=5) and iteratively refined ELLA through twelve in-home design workshops. We then deployed ELLA with ten children for eight days. We report design insights from in-home workshops, characterize children’s engagement and behaviors during deployment, and distill design implications for generative AI–powered social robots supporting early language learning at home.
Victor Nikhil Antony, Shiye Cao, Shuning Wang, Chien-Ming Huang 0001
IDC4
2026 Dynamic Compensation Can Enhance User Engagement by Triggering Sensitivity to Financial Losses in Crowd-sourced Studies
abstract
Participation in crowd-sourced user studies is often driven by monetary incentives. However, standard payment schemes that reward completion unless responses are of poor quality may not invoke sufficient accountability. By compromising user engagement, a lack of accountability can affect data quality and the study’s ecological validity. Here, we investigate alternative compensation strategies that manipulate payment framing and evaluate their impact on engagement through task effort, outcomes, and perception. We compared a standard scheme with implicit rejection risk to a reinforced accountability condition with explicit performance-linked deductions, and two dynamic conditions that unexpectedly switched strategies. In a study with 106 Prolific participants on an image captioning task, we found that only shifting from implicit risk to reinforced accountability significantly increased engagement, likely due to loss aversion after participants had already invested time. The reverse shift decreased effort as observed in the standard group. Our results highlight the importance of carefully designing compensation schemes.
Catalina Gomez, Mung Yao Jia, Sue Min Cho, Chien-Ming Huang 0001, Mathias Unberath
CHI4
2026 Lantern: A Minimalist Robotic Object Platform
abstract
Robotic objects are simple actuated systems that subtly blend into human environments. We design and introduce Lantern, a minimalist robotic object platform to enable building simple robotic artifacts. We conducted in-depth design and engineering iterations of Lantern’s mechatronic architecture to meet specific design goals while maintaining a low build cost (~40 USD). As an extendable, open-source platform, Lantern aims to enable exploration of a range of HRI scenarios by leveraging human tendency to assign social meaning to simple forms. To evaluate Lantern’s potential for HRI, we conducted a series of explorations: 1) a co-design workshop, 2) a sensory room case study, 3) distribution to external HRI labs, 4) integration into a graduate-level HRI course, and 5) public exhibitions with older adults and children. Our findings show that Lantern effectively evokes engagement, can support versatile applications ranging from emotion regulation to focused work, and serves as a viable platform for lowering barriers to HRI as a field.
Victor Nikhil Antony, Zhili Gong 0002, Clara Jeon, Chien-Ming Huang 0001
HRI5
2026 Plant-Inspired Robot Design Metaphors for Ambient HRI
abstract
Plants offer a paradoxical model for interaction: they are ambient, low-demand presences that nonetheless shape atmosphere, routines, and relationships through temporal rhythms and subtle expressions. In contrast, most human–robot interaction (HRI) has been grounded in anthropomorphic and zoomorphic paradigms, producing overt, high-demand forms of engagement. Using a Research through Design (RtD) methodology, we explore plants as metaphoric inspiration for HRI; we conducted iterative cycles of ideation, prototyping, and reflection to investigate what design primitives emerge from plant metaphors and morphologies, and how these primitives can be combined into expressive robotic forms. We present a suite of speculative, open-source prototypes that help probe plant-inspired presence, temporality, form, and gestures. We deepened our learnings from design and prototyping through prototype-centered workshops that explored people’s perceptions and imaginaries of plant-inspired robots. This work contributes: (1) Set of plant-inspired robotic artifacts; (2) Designerly insights on how people perceive plant-inspired robots; and (3) Design consideration to inform how to use plant metaphors to reshape HRI.
Victor Nikhil Antony, Adithya R. N, Sarah Derrick, Zhili Gong 0002, Peter M. Donley, Chien-Ming Huang 0001
HRI6
2026 Reframing Conversational Design in HRI: Deliberate Design with AI Scaffolds
abstract
Large language models (LLMs) enabled conversational robots to shift toward free-form interaction. However, without context-specific adaptation, generic LLM outputs can be ineffective or inappropriate. This adaptation is often attempted through prompt engineering, which is non-intuitive and tedious. Moreover, predominant design practice in HRI relies on impression-based, trial-and-error refinement without structured methods or tools, making the process inefficient and inconsistent. To address this, we present AI-Aided Conversation Engine (ACE) to support deliberate design of human-robot conversations with three key innovations: 1) an LLM-powered voice agent that scaffolds initial prompt creation to overcome the "blank page problem," 2) an annotation interface that enables the collection of granular and grounded feedback on conversational transcripts, and 3) using LLMs to translate user feedback into prompt refinements. We evaluated ACE through two user studies, examining both designs' experience and end users' interactions with robots designed using ACE. Results show that ACE facilitates the creation of robot behavior prompts with greater clarity and specificity, and that the prompts generated with ACE lead to higher-quality human-robot conversational interactions.
Shiye Cao, Jiwon Moon 0002, Yifan Xu 0031, Anqi Liu 0001, Chien-Ming Huang 0001
HRI5
2025 Voice Assistants for Health Self-Management: Designing for and with Older Adults
Amama Mahmood, Shiye Cao, Maia Stiber, Victor Nikhil Antony, Chien-Ming Huang 0001
CHI5
2025 The Design of On-Body Robots for Older Adults
abstract
Wearable technology has significantly improved the quality of life for older adults, and the emergence of on-body, movable robots presents new opportunities to further enhance well-being. Yet, the interaction design for these robots remains under-explored, particularly from the perspective of older adults. We present findings from a two-phase co-design process involving 13 older adults to uncover design principles for on-body robots for this population. We identify a rich spectrum of potential applications and characterize a design space to inform how on-body robots should be built for older adults. Our findings highlight the importance of considering factors like co-presence, embodiment, and multi-modal communication. Our work offers design insights to facilitate the integration of on-body robots into daily life and underscores the value of involving older adults in the co-design process to promote usability and acceptance of emerging wearable robotic technologies.
Victor Nikhil Antony, Clara Jeon, Ge Gao 0001, Huaishu Peng, Anastasia K. Ostrowski, Chien-Ming Huang 0001
HRI7
2025 Xpress: A System for Dynamic, Context-Aware Robot Facial Expressions Using Language Models
abstract
Facial expressions are vital in human communication and significantly influence outcomes in human-robot interaction (HRI), such as likeability, trust, and companionship. However, current methods for generating robotic facial expressions are often labor-intensive, lack adaptability across contexts and platforms, and have limited expressive ranges-leading to repetitive behaviors that reduce interaction quality, particularly in long-term scenarios. We introduce Xpress, a system that leverages language models (LMs) to dynamically generate context-aware facial expressions for robots through a three-phase process: encoding temporal flow, conditioning expressions on context, and generating facial expression code. We demonstrated Xpress as a proof-of-concept through two user studies$(n=15\times 2)$and a case study with children and parents$(n=13)$, in storytelling and conversational scenarios to assess the system's context-awareness, expressiveness, and dynamism. Results demonstrate Xpress's ability to dynamically produce expressive and contextually appropriate facial expressions, highlighting its versatility and potential in HRI applications.
Victor Nikhil Antony, Maia Stiber, Chien-Ming Huang 0001
HRI3
2025 "See You Later, Alligator": Impacts of Robot Small Talk on Task, Rapport, and Interaction Dynamics in Human-Robot Collaboration
abstract
Small talk can foster rapport building in human-human teamwork; yet how non-anthropomorphic robots, such as collaborative manipulators commonly used in industry, may capitalize on these social communications remains unclear. This work investigates how robot-initiated small talk influences task performance, rapport, and interaction dynamics in human-robot collaboration. We developed an autonomous robot system that assists a human in an assembly task while initiating and engaging in small talk. A user study$(N=58)$was conducted in which participants worked with either a functional robot, which engaged in only task-oriented speech, or a social robot, which also initiated small talk. Our study found that participants in the social condition reported significantly higher levels of rapport with the robot. Moreover, all participants in the social condition responded to the robot's small talk attempts; 59% initiated questions to the robot, and 73% engaged in lingering conversations after requesting the final task item. Although active working times were similar across conditions, participants in the social condition recorded longer task durations than those in the functional condition. We discuss the design and implications of robot small talk in shaping human-robot collaboration.
Kaitlynn Taylor Pineda, Ethan Brown, Chien-Ming Huang 0001
HRI3
2025 ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations
abstract
The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.
Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira, Wendy Ju, Micol Spitale, Hatice Gunes, Chien-Ming Huang 0001
ACM Multimedia8
2025 User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants
Amama Mahmood, Bingsheng Yao, Dakuo Wang, Chien-Ming Huang 0001
Int. J. Hum. Comput. Stud.5
2025 "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking Partner
abstract
The rapid advancement of Large Language Models (LLMs) has created numerous potentials for integration with conversational assistants (CAs) assisting people in their daily tasks, particularly due to their extensive flexibility. However, users' real-world experiences interacting with these assistants remain unexplored. In this research, we chose cooking, a complex daily task, as a scenario to explore people's successful and unsatisfactory experiences while receiving assistance from an LLM-based CA, Mango Mango . We discovered that participants value the system's ability to offer customized instructions based on context, provide extensive information beyond the recipe, and assist them in dynamic task planning. However, users expect the system to be more adaptive to oral conversation and provide more suggestive responses to keep them actively involved. Recognizing that users began treating our LLM-CA as a personal assistant or even a partner rather than just a recipe-reading tool, we propose five design considerations for future development.
Szeyi Chan, Bingsheng Yao, Amama Mahmood, Chien-Ming Huang 0001, Holly Jimison, Elizabeth D. Mynatt, Dakuo Wang
Proc. ACM Hum. Comput. Interact.5
2024 Alchemist: LLM-Aided End-User Development of Robot Applications
abstract
Large Language Models (LLMs) have the potential to catalyze a paradigm shift in end-user robot programming---moving from the conventional process of user specifying programming logic to an iterative, collaborative process in which the user specifies desired program outcomes while LLM produces detailed specifications. We introduce a novel integrated development system, Alchemist, that leverages LLMs to empower end-users in creating, testing, and running robot programs using natural language inputs, aiming to reduce the required knowledge for developing robot applications. We present a detailed examination of our system design and provide an exploratory study involving true end-users to assess capabilities, usability, and limitations of our system. Through the design, development, and evaluation of our system, we derive a set of lessons learned from the use of LLMs in robot programming. We discuss how LLMs may be the next frontier for democratizing end-user development of robot applications.
Ulas Berk Karli, Juo-Tung Chen, Victor Nikhil Antony, Chien-Ming Huang 0001
HRI4
2024 ERR@HRI 2024 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Interactions
abstract
Despite the recent advancements in robotics and machine learning (ML), the deployment of autonomous robots in our everyday lives is still an open challenge. This is due to multiple reasons among which are their frequent mistakes, such as interrupting people or having delayed responses, as well as their limited ability to understand human speech, i.e., failure in tasks like transcribing speech to text. These mistakes may disrupt interactions and negatively influence human perception of these robots. To address this problem, robots need to have the ability to detect human-robot interaction (HRI) failures. The ERR@HRI 2024 challenge tackles this by offering a benchmark multimodal dataset of robot failures during human-robot interactions, encouraging researchers to develop and benchmark multimodal machine learning models to detect these failures. We created a dataset featuring multimodal non-verbal interaction data, including facial, speech, and pose features from video clips of interactions with a robotic coach, annotated with labels indicating the presence or absence of robot mistakes, user awkwardness, and interaction ruptures, allowing for the training and evaluation of predictive models. Challenge participants have been invited to submit their multimodal ML models for detection of robot errors, to be evaluated against various performance metrics such as accuracy, precision, recall, F1 score, with and without a margin of error reflecting the time-sensitivity of these metrics. The results of this challenge will help the research field in better understanding the robot failures in human-robot interactions and designing autonomous robots that can mitigate their own errors after successfully detecting them.
Micol Spitale, Maria Teresa Parreira, Maia Stiber, Minja Axelsson, Neval Kara, Garima Kankariya, Chien-Ming Huang 0001, Malte F. Jung, Wendy Ju, Hatice Gunes
ICMI7
2024 Reducing Performance Variability and Overcoming Limited Spatial Ability: Targeted Training for Remote Robot Teleoperation
abstract
In this paper, we present a targeted training approach for remote teleoperation aimed at achieving consistent proficiency levels across users with varying capabilities. Our approach begins by assessing users’ abilities to perform robot motion control, workspace adaptation, and gripper control. It then provides tailored training based on identified skill gaps to enhance the learning effectiveness and user experience. To demonstrate our approach, we conducted a user study, with one group undergoing conventional, free-form training and the other engaging in targeted training in accordance with their skill gaps; after the training phase, participants teleoperated a robotic arm in a simulated medication preparation task for performance evaluation. Our results show that the targeted training approach effectively reduces performance variability and mitigates the influence of spatial ability on both training and task completion time. We discuss the implications of our results for practical teleoperation training and future research.
Tsung-Chi Lin, Juo-Tung Chen, Chien-Ming Huang 0001
IROS3
2024 Designing for Appropriate Reliance: The Roles of AI Uncertainty Presentation, Initial User Decision, and User Demographics in AI-Assisted Decision-Making
abstract
Appropriate reliance is critical to achieving synergistic human-AI collaboration. For instance, when users over-rely on AI assistance, their human-AI team performance is bounded by the model's capability. This work studies how the presentation of model uncertainty may steer users' decision-making toward fostering appropriate reliance. Our results demonstrate that showing the calibrated model uncertainty alone is inadequate. Rather, calibrating model uncertainty and presenting it in a frequency format allow users to adjust their reliance accordingly and help reduce the effect of confirmation bias on their decisions. Furthermore, the critical nature of our skin cancer screening task skews participants' judgment, causing their reliance to vary depending on their initial decision. Additionally, step-wise multiple regression analyses revealed how user demographics such as age and familiarity with probability and statistics influence human-AI collaborative decision-making. We discuss the potential for model uncertainty presentation, initial user decision, and user demographics to be incorporated in designing personalized AI aids for appropriate reliance.
Shiye Cao, Anqi Liu 0001, Chien-Ming Huang 0001
Proc. ACM Hum. Comput. Interact.3
2024 Gender Biases in Error Mitigation by Voice Assistants
abstract
Commercial voice assistants are largely feminized and associated with stereotypically feminine traits such as warmth and submissiveness. As these assistants continue to be adopted for everyday uses, it is imperative to understand how the portrayed gender shapes the voice assistant's ability to mitigate errors, which are still common in voice interactions. We report a study (N=40) that examined the effects of voice gender (feminine, ambiguous, masculine), error mitigation strategies (apology, compensation) and participant's gender on people's interaction behavior and perceptions of the assistant. Our results show that AI assistants that apologized appeared warmer than those offered compensation. Moreover, male participants preferred apologetic feminine assistants over apologetic masculine ones. Furthermore, male participants interrupted AI assistants regardless of perceived gender more frequently than female participants when errors occurred. Our results suggest that the perceived gender of a voice assistant biases user behavior, especially for male users, and that an ambiguous voice has the potential to reduce biases associated with gender-specific traits.
Amama Mahmood, Chien-Ming Huang 0001
Proc. ACM Hum. Comput. Interact.2
2024 Forging Productive Human-Robot Partnerships Through Task Training
abstract
Productive human-robot partnerships are vital to successful integration of assistive robots into everyday life. Although prior research has explored techniques to facilitate collaboration during human-robot interaction, the work described here aims to forge productive partnerships prior to human-robot interaction, drawing upon team-building activities’ aid in establishing effective human teams. Through a 2 (group membership: ingroup and outgroup) ×3 (robot error: main task errors, side task errors, and no errors) online study ( N=62 ), we demonstrate that (1) a non-social pre-task exercise can help form ingroup relationships; (2) an ingroup robot is perceived as a better, more committed teammate than an outgroup robot (despite the two behaving identically); and (3) participants are more tolerant of negative outcomes when working with an ingroup robot. We discuss how pre-task exercises may serve as an active task failure mitigation strategy.
Maia Stiber, Yuxiang Gao, Russell H. Taylor, Chien-Ming Huang 0001
ACM Trans. Hum. Robot Interact.4
2024 ID.8: Co-Creating Visual Stories with Generative AI
abstract
Storytelling is an integral part of human culture and significantly impacts cognitive and socio-emotional development and connection. Despite the importance of interactive visual storytelling, the process of creating such content requires specialized skills and is labor-intensive. This article introduces ID.8, an open-source system designed for the co-creation of visual stories with generative AI. We focus on enabling an inclusive storytelling experience by simplifying the content creation process and allowing for customization. Our user evaluation confirms a generally positive user experience in domains such as enjoyment and exploration while highlighting areas for improvement, particularly in immersiveness, alignment, and partnership between the user and the AI system. Overall, our findings indicate promising possibilities for empowering people to create visual stories with generative AI. This work contributes a novel content authoring system, ID.8, and insights into the challenges and potential of using generative AI for multimedia content creation.
Victor Nikhil Antony, Chien-Ming Huang 0001
ACM Trans. Interact. Intell. Syst.2
2023 Co-Designing with Older Adults, for Older Adults: Robots to Promote Physical Activity
abstract
Lack of physical activity has severe negative health consequences for older adults and limits their ability to live independently. Robots have been proposed to help engage older adults in physical activity (PA), albeit with limited success. There is a lack of robust understanding of older adults' needs and wants from robots designed to engage them in PA. In this paper, we report on the findings of a co-design process where older adults, physical therapy experts, and engineers designed robots to promote PA in older adults. We found a variety of motivators for and barriers against PA in older adults; we, then, conceptualized a broad spectrum of possible robotic support and found that robots can play various roles to help older adults engage in PA. This exploratory study elucidated several overarching themes and emphasized the need for personalization and adaptability. This work highlights key design features that researchers and engineers should consider when developing robots to engage older adults in PA, and underscores the importance of involving various stakeholders in the design and development of assistive robots.
Victor Nikhil Antony, Sue Min Cho, Chien-Ming Huang 0001
HRI3
2023 "What If It Is Wrong": Effects of Power Dynamics and Trust Repair Strategy on Trust and Compliance in HRI
abstract
Robotic systems designed to work alongside people are susceptible to technical and unexpected errors. Prior work has investigated a variety of strategies aimed at repairing people's trust in the robot after its erroneous operations. In this work, we explore the effect of post-error trust repair strategies (promise and explanation) on people's trust in the robot under varying power dynamics (supervisor and subordinate robot). Our results show that, regardless of the power dynamics, promise is more effective at repairing user trust than explanation. Moreover, people found a supervisor robot with verbal trust repair to be more trustworthy than a subordinate robot with verbal trust repair. Our results further reveal that people are prone to complying with the supervisor robot even if it is wrong. We discuss the ethical concerns in the use of supervisor robot and potential interventions to prevent improper compliance in users for more productive human-robot collaboration.
Ulas Berk Karli, Shiye Cao, Chien-Ming Huang 0001
HRI3
2023 On Using Social Signals to Enable Flexible Error-Aware HRI
abstract
Prior error management techniques often do not possess the versatility to appropriately address robot errors across tasks and scenarios. Their fundamental framework involves explicit, manual error management and implicit domain-specific information driven error management, tailoring their response for specific interaction contexts. We present a framework for approaching error-aware systems by adding implicit social signals as another information channel to create more flexibility in application. To support this notion, we introduce a novel dataset (composed of three data collections) with a focus on understanding natural facial action unit (AU) responses to robot errors during physical-based human-robot interactions---varying across task, error, people, and scenario. Analysis of the dataset reveals that, through the lens of error detection, using AUs as input into error management affords flexibility to the system and has the potential to improve error detection response rate. In addition, we provide an example real-time interactive robot error management system using the error-aware framework.
Maia Stiber, Russell H. Taylor, Chien-Ming Huang 0001
HRI3
2023 Active Engagement with Virtual Reality Reduces Stress and Increases Positive Emotions
abstract
Stress, anxiety, and depression negatively affect productivity and the global economy with an estimated annual cost of ${\$}$1 trillion U.S. dollars, according to the World Health Organization. Moreover, prolonged daily stress—even if minor—can lead to severe health consequences, including cancer and various mental disorders. Virtual reality (VR) has been shown to be a promising tool for relieving daily stressors given its accessibility and its projected availability as compared to visiting with mental health professionals. Prior work in this area has mostly focused on the restorative effects of nature simulations, demonstrating that passively experiencing immersive nature scenes improves positive affect. However, aside from providing opportunities for exercise, little is known about how active VR engagement can improve one’s mental health. To address this research gap, this paper presents a new, active form of VR therapy and assesses its effectiveness as compared to passive VR experiences. We developed VR Drawing—inspired by art therapy, which promotes positive emotions through artistic creation—and VR Throwing—inspired by “rage rooms”, which allow people to release negative emotions via intentional destruction. In a between- participants study (n = 64), we found that both VR Drawing and VR Throwing significantly reduced participants’ stress levels and increased positive affect when compared to passively watching nature scenes in VR. Linear regression models suggest that the total number of user interactions positively affects improvement in positive emotions for VR Drawing, but has a negative impact on positive emotions for VR Throwing. This study provides empirical evidence of how active VR experiences may reduce stress and offers guidelines for creating future VR applications to promote psychological well-being.
Irene Kim, Ehsan Azimi, Peter Kazanzides, Chien-Ming Huang 0001
ISMAR4
2023 Older adults' expectations, experiences, and preferences in programming physical robot assistance
Gopika Ajaykumar, Kaitlynn Taylor Pineda, Chien-Ming Huang 0001
Int. J. Hum. Comput. Stud.3
2023 Mitigating knowledge imbalance in AI-advised decision-making through collaborative user involvement
abstract
Integrating artificial intelligence (AI) systems into decision-making tasks attempts to assist people by augmenting or complementing their abilities and ultimately improve task performance. However, when considering recommendations from modern “black box” intelligent systems, users are confronted with the decision of accepting or overriding AI’s recommendations. These decisions are even more challenging to make when there exists a significant knowledge imbalance between the users and the AI system—namely, when people lack necessary task knowledge and are therefore unable to accurately complete the task on their own. In this work, we aim to understand people’s behavior in AI-assisted decision-making tasks when faced with the challenge of knowledge imbalance and explore whether involving users in an AI’s prediction generation process makes them more willing to follow the AI’s recommendations and enhances their perception of collaboration. Our empirical study reveals that the involvement of users in generating AI recommendations during a task with notable knowledge imbalance causes them to be more willing to agree with the AI’s suggestions and to perceive the AI agent and their collaboration as a team more positively.
Catalina Gomez, Mathias Unberath, Chien-Ming Huang 0001
Int. J. Hum. Comput. Stud.3
2023 How Time Pressure in Different Phases of Decision-Making Influences Human-AI Collaboration
abstract
Human cognitive and decision-making abilities depreciate under pressure, motivating the emergence of artificial intelligence (AI) systems as decision support tools to assist people in performing tasks under stress. In this work, we study human decision-making behavior and task performance under time pressure---induced from limitedinitial observation time (time to perform the task before providing an initial response without AI input) andfinal decision time (time to weigh an AI's suggestion before reaching a collective human-AI team answer)---for spatial reasoning and count estimation tasks. Our results show that, while the impact of initial observation time on AI-assisted decision-making was dependent on task nature, participants were more likely to follow AI suggestions when they were provided with longer final decision time; moreover, although participants generally tended to adhere to their initial responses, they had more agency when they were more logically engaged in a task. Our results offer a nuanced understanding of human-AI collaboration under time pressure in different phases of the decision-making process.
Shiye Cao, Catalina Gómez Caballero, Chien-Ming Huang 0001
Proc. ACM Hum. Comput. Interact.3
2023 Crowdsourcing Thumbnail Captions: Data Collection and Validation
abstract
Speech interfaces, such as personal assistants and screen readers, read image captions to users. Typically, however, only one caption is available per image, which may not be adequate for all situations (e.g., browsing large quantities of images). Long captions provide a deeper understanding of an image but require more time to listen to, whereas shorter captions may not allow for such thorough comprehension yet have the advantage of being faster to consume. We explore how to effectively collect both thumbnail captions—succinct image descriptions meant to be consumed quickly—and comprehensive captions—which allow individuals to understand visual content in greater detail. We consider text-based instructions and time-constrained methods to collect descriptions at these two levels of detail and find that a time-constrained method is the most effective for collecting thumbnail captions while preserving caption accuracy. Additionally, we verify that caption authors using this time-constrained method are still able to focus on the most important regions of an image by tracking their eye gaze. We evaluate our collected captions along human-rated axes—correctness, fluency, amount of detail, and mentions of important concepts—and discuss the potential for model-based metrics to perform large-scale automatic evaluations in the future.
Carlos A. Aguirre, Shiye Cao, Amama Mahmood, Chien-Ming Huang 0001
ACM Trans. Interact. Intell. Syst.4
2022 Owning Mistakes Sincerely: Strategies for Mitigating AI Errors
abstract
Interactive AI systems such as voice assistants are bound to make errors because of imperfect sensing and reasoning. Prior human-AI interaction research has illustrated the importance of various strategies for error mitigation in repairing the perception of an AI following a breakdown in service. These strategies include explanations, monetary rewards, and apologies. This paper extends prior work on error mitigation by exploring how different methods of apology conveyance may affect people’s perceptions of AI agents; we report an online study (N=37) that examines how varying the sincerity of an apology and the assignment of blame (on either the agent itself or others) affects participants’ perceptions and experience with erroneous AI agents. We found that agents that openly accepted the blame and apologized sincerely for mistakes were thought to be more intelligent, likeable, and effective in recovering from errors than agents that shifted the blame to others.
Amama Mahmood, Jeanie W. Fung, Isabel Won, Chien-Ming Huang 0001
CHI4
2022 Learning a Group-Aware Policy for Robot Navigation
abstract
Human-aware robot navigation promises a range of applications in which mobile robots bring versatile assistance to people in common human environments. While prior research has mostly focused on modeling pedestrians as independent, intentional individuals, people move in groups; consequently, it is imperative for mobile robots to respect human groups when navigating around people. This paper explores learning group-aware navigation policies based on dynamic group formation using deep reinforcement learning. Through simulation experiments, we show that group-aware policies, compared to baseline policies that neglect human groups, achieve greater robot navigation performance (e.g., fewer collisions), minimize violation of social norms and discomfort, and reduce the robot's movement impact on pedestrians. Our results contribute to the development of social navigation and the integration of mobile robots into human environments.
Kapil D. Katyal, Yuxiang Gao, Jared Markowitz, Sara Pohland, Corban G. Rivera, I-Jeng Wang, Chien-Ming Huang 0001
IROS7
2022 Modeling Human Response to Robot Errors for Timely Error Detection
abstract
In human-robot collaboration, robot errors are inevitable—damaging user trust, willingness to work together, and task performance. Prior work has shown that people naturally respond to robot errors socially and that in social interactions it is possible to use human responses to detect errors. However, there is little exploration in the domain of nonsocial, physical human-robot collaboration such as assembly and tool retrieval. In this work, we investigate how people's organic, social responses to robot errors may be used to enable timely automatic detection of errors in physical human-robot interactions. We conducted a data collection study to obtain facial responses to train a real-time detection algorithm and a case study to explore the generalizability of our method with different task settings and errors. Our results show that natural social responses are effective signals for timely detection and localization of robot errors even in nonsocial contexts and that our method is robust across a variety of task contexts, robot errors, and user responses. This work contributes to robust error detection without detailed task specifications.
Maia Stiber, Russell H. Taylor, Chien-Ming Huang 0001
IROS3
2022 Crowdsourcing Thumbnail Captions via Time-Constrained Methods
abstract
Speech interfaces, such as personal assistants and screen readers, employ captions to allow users to consume images; however, there is typically only one caption available per image, which may not be adequate for all settings (e.g., browsing large quantities of images). Longer captions require more time to consume, whereas shorter captions may hinder a user’s ability to fully understand the image’s content. We explore how to effectively collect both thumbnail captions—succinct image descriptions meant to be consumed quickly—and comprehensive captions, which allow individuals to understand visual content in greater detail. We consider text-based and time-constrained methods to collect descriptions at these two levels of detail, and find that a time-constrained method is most effective for collecting thumbnail captions while preserving caption accuracy. We evaluate our collected captions along three human-rated axes—correctness, fluency, and level of detail—and discuss the potential for model-based metrics to perform automatic evaluation.
Carlos A. Aguirre, Amama Mahmood, Chien-Ming Huang 0001
IUI3
2022 Effects of rhetorical strategies and skin tones on agent persuasiveness in assisted decision-making
abstract
Appearance and linguistic cues may influence how both people and Intelligent Virtual Agents (IVAs) are perceived and evaluated by others; appearance (e.g., skin tone) has been linked to various implicit biases such as agreeing more with stereotypical attractive faces, while particular linguistic cues may effectively increase persuasiveness. In this paper, we report an online study (N=59) evaluating how strategic linguistic cues (expertise: high vs. low) may shape the implicit disadvantages associated with ethnic stereotypes (skin tone: dark vs. light). We found that a virtual agent with a high level of expertise was considered more persuasive, dominant, intelligent, and likeable regardless of their skin tone, and that participants complied more with IVAs with a darker skin tone. Our results suggest that the design of IVAs requires the deliberate considerations of factors such as appearance and linguistic behaviors in order to achieve intended outcomes.
Amama Mahmood, Chien-Ming Huang 0001
IVA2
2022 Understanding User Reliance on AI in Assisted Decision-Making
abstract
Proper calibration of human reliance on AI is fundamental to achieving complementary performance in AI-assisted human decision-making. Most previous works focused on assessing user reliance, and more broadly trust, retrospectively, through user perceptions and task-based measures. In this work, we explore the relationship between eye gaze and reliance under varying task difficulties and AI performance levels in a spatial reasoning task. Our results show a strong positive correlation between percent gaze duration on the AI suggestion and user AI task agreement, as well as user perceived reliance. Moreover, user agency is preserved particularly when the task is easy and when AI performance is low or inconsistent. Our results also reveal nuanced differences between reliance and trust. We discuss the potential of using eye gaze to gauge human reliance on AI in real-time, enabling adaptive AI assistance for optimal human-AI team performance.
Shiye Cao, Chien-Ming Huang 0001
Proc. ACM Hum. Comput. Interact.2
2020 See What I See: Enabling User-Centric Robotic Assistance Using First-Person Demonstrations
abstract
We explore first-person demonstration as an intuitive way of producing task demonstrations to facilitate user-centric robotic assistance. First-person demonstration directly captures the human experience of task performance via head-mounted cameras and naturally includes productive viewpoints for task actions. We implemented a perception system that parses natural first-person demonstrations into task models consisting of sequential task procedures, spatial configurations, and unique task viewpoints. We also developed a robotic system capable of interacting autonomously with users as it follows previously acquired task demonstrations. To evaluate the effectiveness of our robotic assistance, we conducted a user study contextualized in an assembly scenario; we sought to determine how assistance based on a first-person demonstration (user-centric assistance) versus that informed only by the cover image of the official assembly instruction (standard assistance) may shape users' behaviors and overall experience when working alongside a collaborative robot. Our results show that participants felt that their robot partner was more collaborative and considerate when it provided user-centric assistance than when it offered only standard assistance. Additionally, participants were more likely to exhibit unproductive behaviors, such as using their non-dominant hand, when performing the assembly task without user-centric assistance.
Yeping Wang, Gopika Ajaykumar, Chien-Ming Huang 0001
HRI3
2020 Intent-Aware Pedestrian Prediction for Adaptive Crowd Navigation
abstract
Mobile robots capable of navigating seamlessly and safely in pedestrian rich environments promise to bring robotic assistance closer to our daily lives. In this paper we draw on insights of how humans move in crowded spaces to explore how to recognize pedestrian navigation intent, how to predict pedestrian motion and how a robot may adapt its navigation policy dynamically when facing unexpected human movements. Our approach is to develop algorithms that replicate this behavior. We experimentally demonstrate the effectiveness of our prediction algorithm using real-world pedestrian datasets and achieve comparable or better prediction accuracy compared to several state-of-the-art approaches. Moreover, we show that confidence of pedestrian prediction can be used to adjust the risk of a navigation policy adaptively to afford the most comfortable level as measured by the frequency of personal space violation in comparison with baselines. Furthermore, our adaptive navigation policy is able to reduce the number of collisions by 43% in the presence of novel pedestrian motion not seen during training.
Kapil D. Katyal, Gregory D. Hager, Chien-Ming Huang 0001
ICRA3
2020 An Interactive Mixed Reality Platform for Bedside Surgical Procedures
Ehsan Azimi, Zhiyuan Niu, Maia Stiber, Nicholas Greene, Ruby Liu, Camilo A. Molina, Judy Huang, Chien-Ming Huang 0001, Peter Kazanzides
MICCAI (3)8
2020 Structuring Human-Robot Interactions via Interaction Conventions
abstract
Interaction conventions (e.g., using pinch gestures to zoom in and out) are designed to structure how users effectively work with an interactive technology. We contend in this paper that successful human-robot interactions may be achieved through an appropriate use of interaction conventions. We present a simple, natural interaction convention-"Put That Here"-for instructing a robot partner to perform pick-and-place tasks. This convention allows people to use common gestures and verbal commands to select objects of interest and to specify their intended location of placement. We implement an autonomous robot system capable of parsing and operating through this convention. Through a user study, we show that participants were easily able to adopt and use the convention to provide task specifications. Our results show that participants using this convention were able to complete tasks faster and experienced significantly lower cognitive load than when using only verbal commands to give instructions. Furthermore, when asked to give natural pick-and-place instructions to a human collaborator, the participants intuitively used task specification methods that paralleled our convention, incorporating both gestures and verbal commands to provide precise task-relevant information. We discuss the potential of interaction conventions in enabling productive human-robot interactions.
Gopika Ajaykumar, Chien-Ming Huang 0001
RO-MAN4
2019 PATI: a projection-based augmented table-top interface for robot programming
abstract
As robots begin to provide daily assistance to individuals in human environments, their end-users, who do not necessarily have substantial technical training or backgrounds in robotics or programming, will ultimately need to program and "re-task" their robots to perform a variety of custom tasks. In this work, we present PATI---a Projection-based Augmented Table-top Interface for robot programming---through which users are able to use simple, common gestures (e.g., pinch gestures) and tools (e.g., shape tools) to specify table-top manipulation tasks (e.g., pick-and-place) for a robot manipulator. PATI allows users to interact with the environment directly when providing task specifications; for example, users can utilize gestures and tools to annotate the environment with task-relevant information, such as specifying target landmarks and selecting objects of interest. We conducted a user study to compare PATI with a state-of-the-art, standard industrial method for end-user robot programming. Our results show that participants needed significantly less training time before they felt confident in using our system than they did for the industrial method. Moreover, participants were able to program a robot manipulator to complete a pick-and-place task significantly faster with PATI. This work indicates a new direction for end-user robot programming.
Yuxiang Gao, Chien-Ming Huang 0001
IUI2
2019 Toward Effective Robot-Child Tutoring: Internal Motivation, Behavioral Intervention, and Learning Outcomes
abstract
Personalized learning environments have the potential to improve learning outcomes for children in a variety of educational domains, as they can tailor instruction based on the unique learning needs of individuals. Robot tutoring systems can further engage users by leveraging their potential for embodied social interaction and take into account crucial aspects of a learner, such as a student’s motivation in learning. In this article, we demonstrate that motivation in young learners corresponds to observable behaviors when interacting with a robot tutoring system, which, in turn, impact learning outcomes. We first detail a user study involving children interacting one on one with a robot tutoring system over multiple sessions. Based on empirical data, we show that academic motivation stemming from one’s own values or goals as assessed by the Academic Self-Regulation Questionnaire (SRQ-A) correlates to observed suboptimal help-seeking behavior during the initial tutoring session. We then show how an interactive robot that responds intelligently to these observed behaviors in subsequent tutoring sessions can positively impact both student behavior and learning outcomes over time. These results provide empirical evidence for the link between internal motivation, observable behavior, and learning outcomes in the context of robot--child tutoring. We also identified an additional suboptimal behavioral feature within our tutoring environment and demonstrated its relationship to internal factors of motivation, suggesting further opportunities to design robot intervention to enhance learning. We provide insights on the design of robot tutoring systems aimed to deliver effective behavioral intervention during learning interactions for children and present a discussion on the broader challenges currently faced by robot--child tutoring systems.
Aditi Ramachandran, Chien-Ming Huang 0001, Brian Scassellati
ACM Trans. Interact. Intell. Syst.2
2018 Thinking Aloud with a Tutoring Robot to Enhance Learning
abstract
Thinking aloud, while requiring extra mental effort, is a metacognitive technique that helps students navigate through complex problem-solving tasks. Social robots, bearing embodied immediacy that fosters engaging and compliant interactions, are a unique platform to deliver problem-solving support such as thinking aloud to young learners. In this work, we explore the effects of a robot platform and the think-aloud strategy on learning outcomes in the context of a one-on-one tutoring interaction. Results from a 2x2 between-subjects study (n=52) indicate that both the robot platform and use of the think-aloud strategy promoted learning gains for children. In particular, the robot platform effectively enhanced immediate learning gains, measured right after the tutoring session, while the think-aloud strategy improved persistent gains as measured approximately one week after the interaction. Moreover, our results show that a social robot strengthened students» engagement and compliance with the think-aloud support while they performed cognitively demanding tasks. Our work indicates that robots can support metacognitive strategy use to effectively enhance learning and contributes to the growing body of research demonstrating the value of social robots in novel educational settings.
Aditi Ramachandran, Chien-Ming Huang 0001, Edward Gartland, Brian Scassellati
HRI2
2017 Give Me a Break!: Personalized Timing Strategies to Promote Learning in Robot-Child Tutoring
abstract
A common practice in education to accommodate the short attention spans of children during learning is to provide them with non-task breaks for cognitive rest. Holding great promise to promote learning, robots can provide these breaks at times personalized to individual children. In this work, we investigate personalized timing strategies for providing breaks to young learners during a robot tutoring interaction. We build an autonomous robot tutoring system that monitors student performance and provides break activities based on a personalized schedule according to performance. We conduct a field study to explore the effects of different strategies for providing breaks during tutoring. By comparing a fixed timing strategy with a reward strategy (break timing personalized to performance gains) and a refocus strategy (break timing personalized to performance drops), we show that the personalized strategies promote learning gains for children more effectively than the fixed strategy. Our results also reveal immediate benefits in enhancing efficiency and accuracy in completing educational problems after personalized breaks, showing the restorative effects of the breaks when administered at the right time.
Aditi Ramachandran, Chien-Ming Huang 0001, Brian Scassellati
HRI2
2016 Anticipatory Robot Control for Efficient Human-Robot Collaboration
abstract
Efficient collaboration requires collaborators to monitor the behaviors of their partners, make inferences about their task intent, and plan their own actions accordingly. To work seamlessly and efficiently with their human counterparts, robots must similarly rely on predictions of their users' intent in planning their actions. In this paper, we present an anticipatory control method that enables robots to proactively perform task actions based on anticipated actions of their human partners. We implemented this method into a robot system that monitored its user's gaze, predicted his or her task intent based on observed gaze patterns, and performed anticipatory task actions according to its predictions. Results from a human-robot interaction experiment showed that anticipatory control enabled the robot to respond to user requests and complete the task faster-2.5 seconds on average and up to 3.4 seconds-compared to a robot using a reactive control method that did not anticipate user intent. Our findings highlight the promise of performing anticipatory actions for achieving efficient human-robot teamwork.
Chien-Ming Huang 0001, Bilge Mutlu
HRI1
2015 From 9 to 90: Engaging Learners of All Ages
abstract
This paper details the creation of a two-day computer science and robotics outreach course aimed at simultaneously engaging youth (children, ages 9-14) and senior (their grandparents, ages 55+) students. Our goal is to encourage enthusiasm for science and technology in students of all ages as well as provide practical instruction regarding common computer science concepts, including variables, loops, and boolean logic. To this end, we ground our course in the emerging field of social robotics, which enables the design of several multidisciplinary hands-on activities for students. We report on a four-year experience in the development of our course, which has been offered twelve times and involved over 210 youth and senior students. Our work presents a discussion regarding the challenges in designing a course for students from diverse ages, guidelines for creating similar courses, and a reflection on how we might improve our own class. The activities and project code developed for our course are available online as open-source resources.
Allison Sauppé, Daniel Szafir, Chien-Ming Huang 0001, Bilge Mutlu
SIGCSE3
2014 Learning-based modeling of multimodal behaviors for humanlike robots
abstract
In order to communicate with their users in a natural and effective manner, humanlike robots must seamlessly integrate behaviors across multiple modalities, including speech, gaze, and gestures. While researchers and designers have successfully drawn on studies of human interactions to build models of humanlike behavior and to achieve such integration in robot behavior, the development of such models involves a laborious process of inspecting data to identify patterns within each modality or across modalities of behavior and to represent these patterns as "rules" or heuristics that can be used to control the behaviors of a robot, but provides little support for validation, extensibility, and learning. In this paper, we explore how a learning-based approach to modeling multimodal behaviors might address these limitations. We demonstrate the use of a dynamic Bayesian network (DBN) for modeling how humans coordinate speech, gaze, and gesture behaviors in narration and for achieving such coordination with robots. The evaluation of this approach in a human-robot interaction study shows that this learning-based approach is comparable to conventional modeling approaches in enabling effective robot behaviors while reducing the effort involved in identifying behavioral patterns and providing a probabilistic representation of the dynamics of human behavior. We discuss the implications of this approach for designing natural, effective multimodal robot behaviors.
Chien-Ming Huang 0001, Bilge Mutlu
HRI1
2013 Designing effective multimodal behaviors for robots: a data-driven perspective
abstract
Robots need to effectively use multimodal behaviors, including speech, gaze, and gestures, in support of their users to achieve intended interaction goals, such as improved task performance. This proposed research concerns designing effective multimodal behaviors for robots to interact with humans using a data-driven approach. In particular, probabilistic graphical models (PGMs) are used to model the interdependencies among multiple behavioral channels and generate complexly contingent multimodal behaviors for robots to facilitate human-robot interaction. This data-driven approach not only allows the investigation of hidden and temporal relationships among behavioral channels but also provides a holistic perspective on how multimodal behaviors as a whole might shape interaction outcomes. Three studies are proposed to evaluate the proposed data-driven approach and to investigate the dynamics of multimodal behaviors and interpersonal interaction. This research will contribute to the multimodal interaction community in theoretical, methodological, and practical aspects.
Chien-Ming Huang 0001
ICMI1
2013 The repertoire of robot behavior: enabling robots to achieve interaction goals through social behavior
abstract
In social interaction, people draw on a large repertoire of social acts tailoring their use of these acts to meet the demands of the social situation and to achieve the goals of the interaction. This paper presents an approach to creating such a repertoire of social acts for robots and enabling designers to specify the social situation to which robots may adapt their behaviors. Drawing on principles of Activity Theory and social-scientific findings on human social behavior, this paper introduces an implementation of this approach---the Robot Behavior Toolkit---and two studies that use a limited, proof-of-concept repertoire of specifications for gaze cues to demonstrate the feasibility of this approach for controlling robot gaze behavior. The first study assessed the feasibility of the use of this repertoire, comparing it to alternative, baseline repertoires in two human-robot interaction tasks, and found that it enabled the robot to more effectively support the interaction goals. The second study investigated the feasibility of the robot adapting its use of the repertoire to a social situation by comparing different goal specifications in two human-robot interaction tasks. The results showed that these specifications enabled the robot to achieve some of its task and communicative goals, although participant gender strongly affected whether the robot elicited these interaction outcomes.
Chien-Ming Huang 0001, Bilge Mutlu
J. Hum. Robot Interact.1
2012 Robot behavior toolkit: generating effective social behaviors for robots
abstract
Social interaction involves a large number of patterned behaviors that people employ to achieve particular communicative goals. To achieve fluent and effective humanlike communication, robots must seamlessly integrate the necessary social behaviors for a given interaction context. However, very little is known about how robots might be equipped with a collection of such behaviors and how they might employ these behaviors in social interaction. In this paper, we propose a framework that guides the generation of social behavior for humanlike robots by systematically using specifications of social behavior from the social sciences and contextualizing these specifications in an Activity-Theory-based interaction model. We present the Robot Behavior Toolkit, an open-source implementation of this framework as a Robot Operating System (ROS) module and a community-based repository for behavioral specifications, and an evaluation of the effectiveness of the Toolkit in using these specifications to generate social behavior in a human-robot interaction study, focusing particularly on gaze behavior. The results show that specifications from this knowledge base enabled the Toolkit to achieve positive social, cognitive, and task outcomes, such as improved information recall, collaborative work, and perceptions of the robot.
Chien-Ming Huang 0001, Bilge Mutlu
HRI1
2011 Effects of responding to, initiating and ensuring joint attention in human-robot interaction
abstract
Inspired by the developmental timeline of joint attention in humans, we propose a conceptual model of joint attention with three parts: responding to joint attention, initiating joint attention, and ensuring joint attention.We conduct two experiments to investigate effects of joint attention in human-robot interaction. The first experiment explores the effects of responding to joint attention. We show that a robot responding to joint attention improves task performance and is perceived as more competent and socially interactive. The second experiment studies the importance of ensuring joint attention in human-robot interaction.We find that a robot's ensuring joint attention behavior is judged as having better performance in human-robot interactive tasks and is perceived as a natural behavior.
Chien-Ming Huang 0001, Andrea Thomaz
RO-MAN1