Iolanda Leite

dblp:23/6349 · DBLP profile ↗
← Back
100ranked-venue papers
15as first author
59since 2021 · last 2026
0000-0002-2212-4325ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 79 · 15 first-author · 44 since 2021Artificial intelligence and machine learning · 77 · 8 first-author · 49 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 3 first-author · 16 since 2021Systems, architecture and hardware · 15 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Making Sense of the Felt Experience of Controlling Autonomous Systems
abstract
Methods for understanding the felt and situated experience of controlling autonomous systems are crucial for designing systems that are adjusted to the complexity of human interaction. Prior work has tended to overlook the bodily and experiential dimensions of monitoring systems at a distance, particularly in moments of losing control. We combined two approaches, ethnomethodology and conversation analysis, and soma design, to explore a case where a semi-autonomous system crashed when controlled by an inexperienced operator, captured in video ethnographic fieldwork. We report on our methodological approach, combining sequential video analysis of the unfolding sequence and interviews inspired by microphenomenology to unpack the operator’s experience of losing control. We contribute methodological considerations for interaction designers seeking to explore the felt experience of having and losing control of autonomous systems and discuss how insights gained through this combination of methods, rooted in phenomenology, support a designerly appreciation of safety operators’ work.
Hannah R. M. Pelikan, Airi Lampinen, Rachael Garrett, Emily Hofstetter, Amanda Hoskins, Hannah Kuehn, Iolanda Leite, Donald McMillan, Sergio Passero, Katie Winkle, Mathias Broth, Barry Brown 0001, Kristina Höök
DIS7
2026 Opportunities to Talk, Negotiate, and Laugh: Robot Behaviors That Shape Repeated Interactions in Groups of Older Adults
abstract
Feeling socially connected is important for personal well-being, yet many older adults report increasing loneliness and decreasing social connections. We explored how robots and their behaviors can support group interactions and foster social participation among older adults in a community center setting over repeated interactions. We developed a semi-autonomous collaborative and discussion-based variant of the game "With Other Words" for groups of three to four older adults and two robots. A facilitator robot (Furhat) mediated discussions using gaze and verbal support, while a guesser robot (Misty) attempted to guess the words that group members described 'with other words'. We invited 34 older adults aged 65+ to play the game in groups of three or four, three times over two to five weeks. An explorative mixed-method analysis, combining quantitative metrics with Ethnomethodological Conversation Analysis (EMCA), shows that robot gaze and verbal behaviors as well as negotiations around "wrangling" the guesser robot encouraged participation in the game. Further, verbal supporting behaviors elicited shared laughter but also led to breakdowns. While no direct significant improvement in social connectedness was observed, this work contributes to our understanding of how robot behaviors might shape interactions among older adults.
Sarah Gillet, Donald McMillan, Nicole Salomons, Iolanda Leite
HRI5
2026 Clarifying Constraints in Interactive Robot Learning with Language Feedback
abstract
Using non-expert language feedback in learning is crucial to making robots successful in human-centered environments. While language feedback has shown potential to teach robots complex tasks, it also brings challenges: humans leave important context, detail, and clarifications unspoken, making interactive approaches necessary to use the feedback effectively. In this work, we develop an interactive robot learning system that can ask clarifying questions to differentiate between hard and soft task constraints from user verbal feedback. The system uses feedback as either shields (hard constraints) or to shape the reward (soft constraints). We conducted a user study with 24 participants, comparing the use of both hard and soft constraints versus two baseline conditions. We show that participants significantly prefer a system using a combination of both hard and soft constraints, or only using soft constraints, compared to a system using only hard constraints. Qualitative analysis of the participants' interactions with the system revealed common feedback types: spatial, temporal and meta-level. To evaluate the learning performance of the system, we conducted simulated experiments showing that combining both hard and soft constraints performs best in terms of reaching high rewards and finding an efficient solution. Additionally, we provide demonstrations of our system on real robot hardware.
Hannah Kuehn, Leonardo Santos, Iolanda Leite
HRI3
2026 MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
abstract
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous expressions in spontaneous, multi-turn dialogue. We address this gap by introducing (1) a benchmark for referential communication in dynamic 3D environments, built from 6.7 hours of egocentric VR interaction with synchronized speech, motion, gaze, and 3D scene geometry, and (2) a two-stage grounding pipeline that explicitly resolves conversational ambiguity before visual localization. The benchmark includes over 4,200 manually verified referring expressions spanning full, partitive, and pronominal types. Our contextual rewriting approach improves grounding performance by 11-22 percentage points on average, with a pure detector (GroundingDINO) reaching 56.7% on pronominals after rewriting, nearly double the best end-to-end baseline. Results demonstrate that decoupling linguistic reasoning from visual perception is more effective than end-to-end approaches for conversational grounding.
Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan, Anna Klezovich, Lubos Marcinek, Iolanda Leite, Jonas Beskow
LREC6
2025 A Call for Deeper Collaboration Between Robotics and Game Development
abstract
While robotics and game development have independently achieved significant progress in creating interactive and intelligent systems, a deeper collaboration between these fields could be mutually beneficial. This paper argues for more collaboration, highlighting current limited interactions and proposing directions for future research. We discuss shared foundations such as Artificial Intelligence, Extended Reality, and the increasing use of common tools and standards. We then propose opportunities where game development methodologies can advance robotics (e.g., gamified data collection and richer simulation environments) and where robotics research can contribute to games (e.g., improved NPC autonomy and embodied intelligence). This cross-disciplinary interaction can accelerate innovation and lead to more intelligent and usercentered technologies in both domains.
Iolanda Leite, William Ahlberg, André Pereira 0001, Alessandro Sestini, Linus Gisslén, Konrad Tollmar
CoG1
2025 Templates and Graph Neural Networks for Social Robots Interacting in Small Groups of Varying Sizes
abstract
Social robots need to be able to interact effectively with small groups. While there is a significant interest in human-robot interaction in groups, little focus has been placed on developing autonomous social robot decision-making methods that operate smoothly with small groups of any size (e.g. 2, 3, or 4 interactants). In this work, we propose a Template- and Graph-based Modeling approach for robots interacting in small groups (TGM), enabling them to interact with groups in a way that is group-size agnostic. Critically, we separate the decision about the target of their communication, or “whom to address?” from the decision of “what to communicate?”, which allows us to use template-based actions. We further use Graph Neural Networks (GNNs) to efficiently decide on “whom“ and “what”. We evaluated TGM using imitation learning and compared the structured reasoning achieved through GNNs to unstructured approaches for this two-part decision-making problem. On two different datasets, we show that TGM outperforms the baselines encouraging future work to invest in collecting larger datasets.
Sarah Gillet, Sydney Thompson, Iolanda Leite, Marynel Vázquez
HRI3
2025 What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations
abstract
When encountering a robot in the wild, it is not inherently clear to human users what the robot's capabilities are. When encountering misunderstandings or problems in spoken interaction, robots often just apologize and move on, without additional effort to make sure the user understands what happened. We set out to compare the effect of two speech based capability communication strategies (proactive, reactive) to a robot without such a strategy, in regard to the user's rating of and their behavior during the interaction. For this, we conducted an in-person user study with 120 participants who had three speech-based interactions with a social robot in a restaurant setting. Our results suggest that users preferred the robot communicating its capabilities proactively and adjusted their behavior in those interactions, using a more conversational interaction style while also enjoying the interaction more.
Merle M. Reimann, Koen V. Hindriks, Florian Kunneman, Catharine Oertel, Gabriel Skantze, Iolanda Leite
HRI6
2025 Designing Social Behaviours for Autonomous Mobile Robots: The Role of Movement and Light in Communicating Intent
abstract
When autonomous mobile robots (AMRs) share space with humans, establishing trust becomes essential for safe, seamless, and effective interaction. Clear communication of a robot's intent is key to building trust by reducing uncertainty and enabling intuitive interaction. This study explores how AMRs can effectively communicate their intentions through simple, intuitive modalities like movement and light, making their actions more predictable and fostering trust. We designed distinct movement cues combined with light patterns to communicate two key intents; yielding (backing off) and making way (prompting humans to move), tested across four different scenarios. To evaluate the clarity and effectiveness of these behaviours, we conducted an online video study analysing qualitative feedback from open-ended responses. Additionally, we collected quantitative data assessing participants' perceptions of the safety and trustworthiness of the robot. Our findings demonstrate a strong correlation between these perceptions and the robot's ability to display socially aware behaviours.
Shashank Shirol, Joseph La Delfa, Iolanda Leite, Elmira Yadollahi
HRI3
2025 Take a Chance on Me: How Robot Performance and Risk Behaviour Affects Trust and Risk-Taking
abstract
Real-world human-robot interactions often encompass uncertainty. This uncertainty can be handled in different ways, for example by designing robot planners to be more or less risk-tolerant. However, how users actually perceive different risk-taking behaviours in robots has yet to be described. Additionally, in the absence of guarantees on optimal robot performance, the interaction between risk and performance on user perceptions is also unclear. To address this gap, we conducted a user study with 84 participants investigating how robot performance and risk behaviour affects users' trust and risk-taking decisions. Participants collaborated with a Franka robot arm to perform a block-stacking task. We compared a robot which displays consistent but sub-optimal behaviours to a robot displaying risky but occasionally optimal behaviour. Risky robot behaviour led to higher trust than consistent behaviour when the robot was on average good at stacking blocks (high expectation), but lower trust when the robot was on average bad at stacking blocks (low expectation). Individual risk-willingness also predicted likelihood of selecting the risky robot over the consistent robot for future interactions, but only when the average expectation was low. These findings have implications for risk-aware planning and decision-making in mixed human-robot systems.
Rebecca Stower, Anna Gautier, Maciej Wozniak 0001, Patric Jensfelt, Jana Tumova, Iolanda Leite
HRI6
2025 Flora: Sample-Efficient Preference-Based Rl Via Low-Rank Style Adaptation of Reward Functions
abstract
Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform the original task. However, collecting preferences for the adaptation process in robotics is often challenging and time-consuming. In this work we explore the adaptation of pre-trained robots in the low-preference-data regime. We show that, in this regime, recent adaptation approaches suffer from catastrophic reward forgetting (CRF), where the updated reward model overfits to the new preferences, leading the agent to become unable to perform the original task. To mitigate CRF, we propose to enhance the original reward model with a small number of parameters (low-rank matrices) responsible for modeling the preference adaptation. Our evaluation shows that our method can efficiently and effectively adjust robotic behavior to human preferences across simulation benchmark tasks and multiple real-world robotic tasks. We provide videos of our results and source code at https://sites.google.com/view/preflora/.
Daniel Marta, Simon Holk, Miguel Vasco, Jens Lundell, Timon Homberger, Finn Lukas Busch, Olov Andersson, Danica Kragic, Iolanda Leite
ICRA9
2025 The Impact of VR and 2D Interfaces on Human Feedback in Preference-Based Robot Learning
abstract
Aligning robot navigation with human preferences is essential for ensuring comfortable, and predictable robot movement in shared spaces. While preference-based learning methods, such as reinforcement learning from human feedback (RLHF), enable this alignment, the choice of the preference collection interface may influence the process. Traditional 2D interfaces provide structured views but lack spatial depth, whereas immersive VR offers richer perception, potentially affecting preference articulation. This study systematically examines how the interface modality impacts human preference collection and navigation policy alignment. We introduce a novel dataset of 2,325 human preference queries collected through both VR and 2D interfaces, revealing significant differences in user experience, preference consistency, and policy outcomes. Our findings highlight the trade-offs between immersion, perception, and preference reliability, emphasizing the importance of interface selection in preference-based robot learning. The dataset is available to support future research.
Jorge de Heuvel, Daniel Marta, Simon Holk, Iolanda Leite, Maren Bennewitz
IROS4
2025 STREAK: Streaming Network for Continual Learning of Object Relocations under Household Context Drifts
abstract
In real-world settings, robots are expected to assist humans across diverse tasks and still continuously adapt to dynamic changes over time. For example, in domestic environments, robots can proactively help users by fetching needed objects based on learned routines, which they infer by observing how objects move over time. However, data from these interactions are inherently non-independent and non-identically distributed (non-i.i.d.), e.g., a robot assisting multiple users may encounter varying data distributions as individuals follow distinct habits. This creates a challenge: integrating new knowledge without catastrophic forgetting. To address this, we propose STREAK (Spatio Temporal RElocation with Adaptive Knowledge retention), a continual learning framework for real-world robotic learning. It leverages a streaming graph neural network with regularization and rehearsal techniques to mitigate context drifts while retaining past knowledge. Our method is time- and memory-efficient, enabling long-term learning without retraining on all past data, which becomes infeasible as data grows in real-world interactions. We evaluate STREAK on the task of incrementally predicting human routines over 50+ days across different households. Results show that it effectively prevents catastrophic forgetting while maintaining generalization, making it a scalable solution for long-term human-robot interactions.
Ermanno Bartoli, Fethiye Irmak Dogan, Iolanda Leite
RO-MAN3
2025 The Need for (Robot) Speed: Offloading Heavy Computations Improves Response Time and User Experience in Spoken Interactions
abstract
In this work we present RoDgeR, a system that leverages edge computing to offload computationally demanding tasks for real-time human-robot interaction (HRI). We identify dialogue management as an example of a computationally intensive task and demonstrate that an edge-based Large Language Model (LLM) results in faster response times than both cloud-based and embedded LLMs. We further implement an edge-based LLM in RoDgeR to evaluate user experience with 63 participants in a simulated restaurant scenario. Our results confirm that RoDgeR outperforms embedded and cloud-based solutions, leading to improved user experience. These findings highlight the potential of edge computing for improving the quality of human-robot interactions.
Ermanno Bartoli, Rebecca Stower, Hanna Werner, Bryan Donyanvard, Jana Tumova, Iolanda Leite
RO-MAN6
2025 A Model-Agnostic Approach for Semantically Driven Disambiguation in Human-Robot Interaction
abstract
Ambiguities are inevitable in human-robot interaction, especially when a robot follows user instructions in a large, shared space. For example, if a user asks the robot to find an object in a home environment with underspecified instructions, the object could be in multiple locations depending on missing factors. For instance, a bowl might be in the kitchen cabinet or on the dining room table, depending on whether it is clean or dirty, full or empty, and the presence of other objects around it. Previous works on object search have assumed that the queried object is immediately visible to the robot or have predicted object locations using one-shot inferences, which are likely to fail for ambiguous or partially understood instructions. This paper focuses on these gaps and presents a novel model-agnostic approach leveraging semantically driven clarifications to enhance the robot’s ability to locate queried objects in fewer attempts. Specifically, we leverage different knowledge embedding models, and when ambiguities arise, we propose an informative clarification method, which follows an iterative prediction process. The user experiment evaluation of our method shows that our approach is applicable to different custom semantic encoders as well as LLMs, and informative clarifications improve performances, enabling the robot to locate objects on its first attempts. The user experiment data is publicly available at https://github.com/IrmakDogan/ExpressionDataset.
Fethiye Irmak Dogan, Maithili Patel, Iolanda Leite, Sonia Chernova
RO-MAN4
2025 Exploring Unstructured Language Feedback for Robot Learning
abstract
In this paper, we aim to explore how humans give unstructured free-form natural language feedback towards correcting robot task policies. We present a qualitative study based on crowd-sourced feedback from 66 participants. Participants give feedback on the execution of three robotic tasks in the form of mobile navigation, dexterous object manipulation and a robot arm opening a door. We find through reflexive thematic analysis which features participants reference most and what other patterns are present in the feedback, including that responders do not naturally give concrete and actionable feedback. We attribute the lack of feedback concreteness to a lack of engagement with robot behavior and false assumptions about who is receiving the feedback and how much knowledge participants have. The study presents a step towards better understanding unstructured language feedback for robotic learning.
Hannah Kuehn, William Ahlberg, Joseph La Delfa, Iolanda Leite
RO-MAN4
2025 BT-ACTION: A Test-Driven Approach for Modular Understanding of User Instruction Leveraging Behaviour Trees and LLMs
abstract
Natural language instructions are often abstract and complex, requiring robots to execute multiple subtasks even for seemingly simple queries. For example, when a user asks a robot to prepare avocado toast, the task involves several sequential steps. Moreover, such instructions can be ambiguous or infeasible for the robot or may exceed the robot’s existing knowledge. While Large Language Models (LLMs) offer strong language reasoning capabilities to handle these challenges, effectively integrating them into robotic systems remains a key challenge. To address this, we propose BT-ACTION, a test-driven approach that combines the modular structure of Behavior Trees (BT) with LLMs to generate coherent sequences of robot actions for following complex user instructions, specifically in the context of preparing recipes in a kitchen-assistance setting. We evaluated BT-ACTION in a comprehensive user study with 45 participants, comparing its performance to direct LLM prompting. Results demonstrate that the modular design of BT-ACTION helped the robot make fewer mistakes and increased user trust, and participants showed a significant preference for the robot leveraging the modular approach. The code is publicly available at https://github.com/1Eggbert7/BTLLM.
Alexander Leszczynski, Sarah Gillet, Iolanda Leite, Fethiye Irmak Dogan
RO-MAN3
2025 The Effect of Voice and Repair Strategy on Trust Formation and Repair in Human-Robot Interaction
abstract
Trust is essential for social interactions, including those between humans and social artificial agents, such as robots. Several factors and combinations thereof can contribute to the formation of trust and, importantly in the case of machines that work with a certain margin of error, to its maintenance and repair after it has been breached. In this article, we present the results of a study aimed at investigating the role of robot voice and chosen repair strategy on trust formation and repair in a collaborative task. People helped a robot navigate through a maze, and the robot made mistakes at pre-defined points during the navigation. Via in-game behaviour and follow-up questionnaires, we could measure people’s trust towards the robot. We found that people trusted the robot speaking with a state-of-the-art synthetic voice more than with the default robot voice in the game, even though they indicated the opposite in the questionnaires. Additionally, we found that three repair strategies that people use in human-human interaction (justification of the mistake, promise to be better and denial of the mistake) work also in human-robot interaction.
Marta Romeo, Ilaria Torre 0002, Sébastien Le Maguer, Alexander Sleat, Angelo Cangelosi, Iolanda Leite
ACM Trans. Hum. Robot Interact.6
2024 Join Me Here if You Will: Investigating Embodiment and Politeness Behaviors When Joining Small Groups of Humans, Robots, and Virtual Characters
abstract
Politeness and embodiment are pivotal elements in human-agent interactions. While many previous works advocate the positive role of embodiment in enhancing these interactions, it remains unclear how embodiment and politeness affect individuals joining groups. In this paper, we explore how politeness behaviors (verbal and nonverbal) exhibited by three distinct embodiments (humans, robots, and virtual characters) influence individuals’ decisions to join a group of two agents in a controlled experiment (N=54). We assessed agent effectiveness regarding persuasiveness, perceived politeness, and participants’ trajectories when joining the group. We found that embodiment does not significantly impact agent persuasiveness and perceived politeness, but politeness does. Direct and explicit politeness strategies have a higher success rate in persuading participants to join the group at the furthest side. Lastly, participants adhered to social norms when joining at the furthest side, maintained a greater physical distance from humans, chose longer paths, and walked faster when interacting with humans.
Sahba Zojaji, Andrii Matviienko, Iolanda Leite, Christopher Peters 0001
CHI3
2024 PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning
abstract
Preference-based reinforcement learning (RL) has emerged as a new field in robot learning, where humans play a pivotal role in shaping robot behavior by expressing preferences on different sequences of state-action pairs. However, formulating realistic policies for robots demands responses from humans to an extensive array of queries. In this work, we approach the sample-efficiency challenge by expanding the information collected per query to contain both preferences and optional text prompting. To accomplish this, we leverage the zero-shot capabilities of a large language model (LLM) to reason from the text provided by humans. To accommodate the additional query information, we reformulate the reward learning objectives to contain flexible highlights -- state-action pairs that contain relatively high information and are related to the features processed in a zero-shot fashion from a pretrained LLM. In both a simulated scenario and a user study, we reveal the effectiveness of our work by analyzing the feedback and its implications. Additionally, the collective feedback collected serves to train a robot on socially compliant trajectories in a simulated social navigation landscape. We provide video examples of the trained policies at https://sites.google.com/view/rl-predilect
Simon Holk, Daniel Marta, Iolanda Leite
HRI3
2024 POLITE: Preferences Combined with Highlights in Reinforcement Learning
abstract
Many solutions to address the challenge of robot learning have been devised, namely through exploring novel ways for humans to communicate complex goals and tasks in reinforcement learning (RL) setups. One way that experienced recent research interest directly addresses the problem by considering human feedback as preferences between pairs of trajectories (sequences of state-action pairs). However, when simply attributing a single preference to a pair of trajectories that contain many agglomerated steps, key pieces of information are lost in the process. We amplify the initial definition of preferences to account for highlights: state-action pairs of relatively high information (high/low reward) within a preferred trajectory. To include the additional information, we design novel regularization methods within a preference learning framework. To this extent, we present our method which is able to greatly reduce the necessary amount of preferences, by permitting the highlighting of favoured trajectories, in order to reduce the entropy of the credit assignment. We show the effectiveness of our work in both simulation and a user study, which analyzes the feedback given and its implications. We also use the total collected feedback to train a robot policy for socially compliant trajectories in a simulated social navigation environment. We release code and video examples at https://sites.google.com/view/rl-polite
Simon Holk, Daniel Marta, Iolanda Leite
ICRA3
2024 SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation
abstract
Preference-based reinforcement learning (RL) poses as a recent research direction in robot learning, by allowing humans to teach robots through preferences on pairs of desired behaviours. Nonetheless, to obtain realistic robot policies, an arbitrarily large number of queries is required to be answered by humans. In this work, we approach the sample-efficiency challenge by presenting a technique which synthesizes queries, in a semi-supervised learning perspective. To achieve this, we leverage latent variational autoencoder (VAE) representations of trajectory segments (sequences of state-action pairs). Our approach manages to produce queries which are closely aligned with those labeled by humans, while avoiding excessive uncertainty according to the human preference predictions as determined by reward estimations. Additionally, by introducing variation without deviating from the original human’s intents, more robust reward function representations are achieved. We compare our approach to recent state-of-the-art preference-based RL semi-supervised learning techniques. Our experimental findings reveal that we can enhance the generalization of the estimated reward function without requiring additional human intervention. Lastly, to confirm the practical applicability of our approach, we conduct experiments involving actual human users in a simulated social navigation setting. Videos of the experiments can be found at https://sites.google.com/view/rl-sequel
Daniel Marta, Simon Holk, Christian Pek, Iolanda Leite
ICRA4
2024 Shielding for Socially Appropriate Robot Listening Behaviors
abstract
A crucial part of traditional reinforcement learning (RL) is the initial exploration phase, in which trying available actions randomly is a critical element. As random behavior might be detrimental to a social interaction, this work proposes a novel paradigm for learning social robot behavior–the use of shielding to ensure socially appropriate behavior during exploration and learning. We explore how a data-driven approach for shielding could be used to generate listening behavior. In a video-based user study (N=110), we compare shielded exploration to two other exploration methods. We show that the shielded exploration is perceived as more comforting and appropriate than a straightforward random approach. Based on our findings, we discuss the potential for future work using shielded and socially guided approaches for learning idiosyncratic social robot behaviors through RL.
Sarah Gillet, Daniel Marta, Mohammed Akif, Iolanda Leite
RO-MAN4
2024 Let's move on: Topic Change in Robot-Facilitated Group Discussions
abstract
Robot-moderated group discussions have the potential to facilitate engaging and productive interactions among human participants. Previous work on topic management in conversational agents has predominantly focused on human engagement and topic personalization, with the agent having an active role in the discussion. Also, studies have shown the usefulness of including robots in groups, yet further exploration is still needed for robots to learn when to change the topic while facilitating discussions. Accordingly, our work investigates the suitability of machine-learning models and audiovisual non-verbal features in predicting appropriate topic changes. We utilized interactions between a robot moderator and human participants, which we annotated and used for extracting acoustic and body language-related features. We provide a detailed analysis of the performance of machine learning approaches using sequential and non-sequential data with different sets of features. The results indicate promising performance in classifying inappropriate topic changes, outperforming rule-based approaches. Additionally, acoustic features exhibited comparable performance and robustness compared to the complete set of multimodal features. Our annotated data is publicly available at https://github.com/ghadj/topic-change-robot-discussions-data-2024.
Georgios Hadjiantonis, Sarah Gillet, Marynel Vázquez, Iolanda Leite, Fethiye Irmak Dogan
RO-MAN4
2024 HRI Wasn't Built In a Day: A Call To Action For Responsible HRI Research
abstract
In recent years, the awareness of the academy around responsible research has notably increased. For instance, with advances in machine learning and artificial intelligence, recent efforts have been made to promote ethical, fair, and inclusive AI and robotics. To better understand if and to what extent HRI is incentivizing researchers to engage in responsible research, we conducted an exploratory review of the publishing guidelines for the most popular HRI conference venues. We identified 18 conferences which published at least 7 HRI papers in 2022. From these, we discuss four themes relevant to conducting responsible HRI research in line with the Responsible Research and Innovation framework: ethical and human participant considerations, transparency and reproducibility, accessibility and inclusion, and plagiarism and LLM use. We identify several gaps and room for improvement within HRI regarding responsible research. Finally, we establish a call to action to provoke conversations among HRI researchers about the importance of conducting responsible research within emerging fields like HRI.
Micol Spitale, Rebecca Stower, Maria Teresa Parreira, Elmira Yadollahi, Iolanda Leite, Hatice Gunes
RO-MAN5
2024 Smiling in the Face and Voice of Avatars and Robots: Evidence for a 'Smiling McGurk Effect'
abstract
Multisensory integration influences emotional perception, as the McGurk effect demonstrates for the communication between humans. Human physiology implicitly links the production of visual features with other modes like the audio channel: Face muscles responsible for a smiling face also stretch the vocal cords that result in a characteristic smiling voice. For artificial agents capable of multimodal expression, this linkage is modeled explicitly. In our studies, we observe the influence of visual and audio channels on the perception of the agents' emotional expression. We created videos of virtual characters and social robots either with matching or mismatching emotional expressions in the audio and visual channels. In two online studies, we measured the agents' perceived valence and arousal. Our results consistently lend support to the ‘emotional McGurk effect' hypothesis, according to which face transmits valence information, and voice transmits arousal. When dealing with dynamic virtual characters, visual information is enough to convey both valence and arousal, and thus audio expressivity need not be congruent. When dealing with robots with fixed facial expressions, however, both visual and audio information need to be present to convey the intended expression.
Ilaria Torre 0002, Simon Holk, Elmira Yadollahi, Iolanda Leite, Rachel McDonnell, Naomi Harte
IEEE Trans. Affect. Comput.4
2024 Interaction-Shaping Robotics: Robots That Influence Interactions between Other Agents
abstract
Work in Human–Robot Interaction (HRI) has investigated interactions between one human and one robot as well as human–robot group interactions. Yet the field lacks a clear definition and understanding of the influence a robot can exert on interactions between other group members (e.g., human-to-human). In this article, we define Interaction-Shaping Robotics (ISR), a subfield of HRI that investigates robots that influence the behaviors and attitudes exchanged between two (or more) other agents. We highlight key factors of interaction-shaping robots that include the role of the robot, the robot-shaping outcome, the form of robot influence, the type of robot communication, and the timeline of the robot’s influence. We also describe three distinct structures of human–robot groups to highlight the potential of ISR in different group compositions and discuss targets for a robot’s interaction-shaping behavior. Finally, we propose areas of opportunity and challenges for future research in ISR.
Sarah Gillet, Marynel Vázquez, Sean Andrist, Iolanda Leite, Sarah Sebo
ACM Trans. Hum. Robot Interact.4
2023 How do Humans take an Object from a Robot: Behavior changes observed in a User Study
abstract
To facilitate human-robot interaction and gain human trust, a robot should recognize and adapt to changes in human behavior. This work documents different human behaviors observed while taking objects from an interactive robot in an experimental study, categorized across two dimensions: pull force applied and handedness. We also present the changes observed in human behavior upon repeated interaction with the robot to take various objects.
Parag Khanna, Elmira Yadollahi, Iolanda Leite, Mårten Björkman, Christian Smith
HAI3
2023 Would You Help Me?: Linking Robot's Perspective-Taking to Human Prosocial Behavior
abstract
Despite the growing literature on human attitudes toward robots, particularly prosocial behavior, little is known about how robots' perspective-taking, the capacity to perceive and understand the world from other viewpoints, could influence such attitudes and perceptions of the robot. To make robots and AI more autonomous and self-aware, more researchers have focused on developing cognitive skills such as perspective-taking and theory of mind in robots and AI. The present study investigated whether a robot's perspective-taking choices could influence the occurrence and extent of exhibiting prosocial behavior toward the robot. We designed an interaction consisting of a perspective-taking task, where we manipulated how the robot instructs the human to find objects by changing its frame of reference and measured the human's exhibition of prosocial behavior toward the robot. In a between-subject study (N=70), we compared the robot's egocentric and addressee-centric instructions against a control condition, where the robot's instructions were object-centric. Participants' prosocial behavior toward the robot was measured using a voluntary data collection session. Our results imply that the occurrence and extent of prosocial behavior toward the robot were significantly influenced by the robot's visuospatial perspective-taking behavior. Furthermore, we observed, through questionnaire responses, that the robot's choice of perspective-taking could potentially influence the humans' perspective choices, were they to reciprocate the instructions to the robot.
João Tiago Almeida, Iolanda Leite, Elmira Yadollahi
HRI2
2023 Increasing Perceived Safety in Motion Planning for Human-Drone Interaction
abstract
Safety is crucial for autonomous drones to operate close to humans. Besides avoiding unwanted or harmful contact, people should also perceive the drone as safe. Existing safe motion planning approaches for autonomous robots, such as drones, have primarily focused on ensuring physical safety, e.g., by imposing constraints on motion planners. However, studies indicate that ensuring physical safety does not necessarily lead to perceived safety. Prior work in Human-Drone Interaction (HDI) shows that factors such as the drone's speed and distance to the human are important for perceived safety. Building on these works, we propose a parameterized control barrier function (CBF) that constrains the drone's maximum deceleration and minimum distance to the human and update its parameters on people's ratings of perceived safety. We describe an implementation and evaluation of our approach. Results of a within-subject user study (N=15) show that we can improve perceived safety of a drone by adjusting to people individually.
Sanne van Waveren, Rasmus Rudling, Iolanda Leite, Patric Jensfelt, Christian Pek
HRI3
2023 Feminist Human-Robot Interaction: Disentangling Power, Principles and Practice for Better, More Ethical HRI
abstract
Human-Robot Interaction (HRI) is inherently a human-centric field of technology. The role of feminist theories in related fields (e.g. Human-Computer Interaction, Data Science) are taken as a starting point to present a vision for Feminist HRI which can support better, more ethical HRI practice everyday, as well as a more activist research and design stance. We first define feminist design for an HRI audience and use a set of feminist principles from neighboring fields to examine existent HRI literature, showing the progress that has been made already alongside some additional potential ways forward. Following this we identify a set of reflexive questions to be posed throughout the HRI design, research and development pipeline, encouraging a sensitivity to power and to individuals' goals and values. Importantly, we do not look to present a definitive, fixed notion of Feminist HRI, but rather demonstrate the ways in which bringing feminist principles to our field can lead to better, more ethical HRI, and to discuss how we, the HRI community, might do this in practice.
Katie Winkle, Donald McMillan, Maria Arnelid, Katherine Harrison 0001, Madeline Balaam, Ericka Johnson, Iolanda Leite
HRI7
2023 Aligning Human Preferences with Baseline Objectives in Reinforcement Learning
abstract
Practical implementations of deep reinforcement learning (deep RL) have been challenging due to an amplitude of factors, such as designing reward functions that cover every possible interaction. To address the heavy burden of robot reward engineering, we aim to leverage subjective human preferences gathered in the context of human-robot interaction, while taking advantage of a baseline reward function when available. By considering baseline objectives to be designed beforehand, we are able to narrow down the policy space, solely requesting human attention when their input matters the most. To allow for control over the optimization of different objectives, our approach contemplates a multi-objective setting. We achieve human-compliant policies by sequentially training an optimal policy from a baseline specification and collecting queries on pairs of trajectories. These policies are obtained by training a reward estimator to generate Pareto optimal policies that include human preferred behaviours. Our approach ensures sample efficiency and we conducted a user study to collect real human preferences, which we utilized to obtain a policy on a social navigation environment.
Daniel Marta, Simon Holk, Christian Pek, Jana Tumova, Iolanda Leite
ICRA5
2023 Real-Time RRT* with Signal Temporal Logic Preferences
abstract
Signal Temporal Logic (STL) is a rigorous specification language that allows one to express various spatio-temporal requirements and preferences. Its semantics (called robustness) allows quantifying to what extent are the STL specifications met. In this work, we focus on enabling STL constraints and preferences in the Real-Time Rapidly Exploring Random Tree (RT-RRT*) motion planning algorithm in an environment with dynamic obstacles. We propose a cost function that guides the algorithm towards the asymptotically most robust solution, i.e. a plan that maximally adheres to the STL specification. In experiments, we applied our method to a social navigation case, where the STL specification captures spatio-temporal preferences on how a mobile robot should avoid an incoming human in a shared space. Our results show that our approach leads to plans adhering to the STL specification, while ensuring efficient cost computation.
Alexis Linard, Ilaria Torre 0002, Ermanno Bartoli, Alexander Sleat, Iolanda Leite, Jana Tumova
IROS5
2023 VARIQuery: VAE Segment-Based Active Learning for Query Selection in Preference-Based Reinforcement Learning
abstract
Human-in-the-loop reinforcement learning (RL) methods actively integrate human knowledge to create reward functions for various robotic tasks. Learning from preferences shows promise as alleviates the requirement of demonstrations by querying humans on state-action sequences. However, the limited granularity of sequence-based approaches complicates temporal credit assignment. The amount of human querying is contingent on query quality, as redundant queries result in excessive human involvement. This paper addresses the often-overlooked aspect of query selection, which is closely related to active learning (AL). We propose a novel query selection approach that leverages variational autoencoder (VAE) representations of state sequences. In this manner, we formulate queries that are diverse in nature while simultaneously taking into account reward model estimations. We compare our approach to the current state-of-the-art query selection methods in preference-based RL, and find ours to be either on-par or more sample efficient through extensive benchmarking on simulated environments relevant to robotics. Lastly, we conduct an online study to verify the effectiveness of our query selection approach with real human feedback and examine several metrics related to human effort.
Daniel Marta, Simon Holk, Christian Pek, Jana Tumova, Iolanda Leite
IROS5
2023 Generating Scenarios from High-Level Specifications for Object Rearrangement Tasks
abstract
Rearranging objects is an essential skill for robots. To quickly teach robots new rearrangements tasks, we would like to generate training scenarios from high-level specifications that define the relative placement of objects for the task at hand. Ideally, to guide the robot's learning we also want to be able to rank these scenarios according to their difficulty. Prior work has shown how generating diverse scenario from specifications and providing the robot with easy-to-difficult samples can improve the learning. Yet, existing scenario generation methods typically cannot generate diverse scenarios while controlling their difficulty. We address this challenge by conditioning generative models on spatial logic specifications to generate spatially-structured scenarios that meet the specification and desired difficulty level. Our experiments showed that generative models are more effective and data-efficient than rejection sam-pling and that the spatially-structured scenarios can drastically improve training of downstream tasks by orders of magnitude.
Sanne van Waveren, Christian Pek, Iolanda Leite, Jana Tumova, Danica Kragic
IROS3
2023 Persuasive Polite Robots in Free-Standing Conversational Groups
abstract
Politeness is at the core of the common set of behavioral norms that regulate human communication and is therefore of significant interest in the design of Human-Robot Interactions. In this paper, we investigate how the politeness behaviors of a humanoid robot impact human decisions about where to join a group of two robots. We also evaluate the resulting impact on the perception of the robot's politeness. In a study involving 59 participants, the main (Pepper) robot in the group invited participants to join using six politeness behaviors derived from Brown and Levinson's politeness theory. It requests participants to join the group at the furthest side of the group which involves more effort to reach than a closer side that is also available to the participant but would ignore the request of the robot. We evaluated the robot's effectiveness in terms of persuasiveness, politeness, and clarity. We found that more direct and explicit politeness strategies derived from the theory have a higher level of success in persuading participants to join at the furthest side of the group. We also evaluated participants' adherence to social norms i.e. not walking through the center, or o-space, of the group when joining it. Our results showed that participants tended to adhere to social norms when joining at the furthest side by not walking through the center of the group of robots, even though they were informed that the robots were fully automated.
Sahba Zojaji, Adrian Benigno Latupeirissa, Iolanda Leite, Roberto Bresin, Christopher Peters 0001
IROS3
2023 Personality-Adapted Language Generation for Social Robots
abstract
Previous works in Human-Robot Interaction have demonstrated the positive potential benefit of designing social robots which express specific personalities. In this work, we focus specifically on the adaptation of language (as the choice of words, their order, etc.) following the extraversion trait. We look to investigate whether current language models could support more autonomous generations of such personality-expressive robot output. We examine the performance of two models with user studies evaluating (i) raw text output and (ii) text output when used within multi-modal speech from the Furhat robot. We find that the ability to successfully manipulate perceived extraversion sometimes varies across different dialogue topics. We were able to achieve correct manipulation of robot personality via our language adaptation, but our results suggest further work is necessary to improve the automation and generalisation abilities of these models.
Alessio Galatolo, Iolanda Leite, Katie Winkle
RO-MAN2
2023 Effects of Explanation Strategies to Resolve Failures in Human-Robot Collaboration
abstract
DH Despite significant improvements in robot capabilities, they are likely to fail in human-robot collaborative tasks due to high unpredictability in human environments and varying human expectations. In this work, we explore the role of explanation of failures by a robot in a human-robot collaborative task. We present a user study incorporating common failures in collaborative tasks with human assistance to resolve the failure. In the study, a robot and a human work together to fill a shelf with objects. Upon encountering a failure, the robot explains the failure and the resolution to overcome the failure, either through handovers or humans completing the task. The study is conducted using different levels of robotic explanation based on the failure action, failure cause, and action history, and different strategies in providing the explanation over the course of repeated interaction. Our results show that the success in resolving the failures is not only a function of the level of explanation but also the type of failures. Furthermore, while novice users rate the robot higher overall in terms of their satisfaction with the explanation, their satisfaction is not only a function of the robot’s explanation level at a certain round but also the prior information they received from the robot.
Parag Khanna, Elmira Yadollahi, Mårten Björkman, Iolanda Leite, Christian Smith
RO-MAN4
2023 What's at Stake? Robot explanations matter for high but not low-stake scenarios
abstract
Although the field of Explainable Artificial Intelligence (XAI) in Human-Robot Interaction is gathering increasing attention, how well different explanations compare across HRI scenarios is still not well understood. We conducted an exploratory online study with 335 participants analysing the interaction between type of explanation (counterfactual, feature-based, and no explanation), the stake of the scenario (high, low) and the application scenario (healthcare, industry). Participants viewed one of 12 different vignettes depicting a combination of these three factors and rated their system understanding and trust in the robot. Compared to no explanation, both counterfactual and feature-based explanations improved system understanding and performance trust (but not moral trust). Additionally, when no explanation was present, high-stake scenarios led to significantly worse performance trust and system understanding. These findings suggest that explanations can be used to calibrate users’ perceptions of the robot in high-stake scenarios.
Gaspar Isaac Melsión, Rebecca Stower, Katie Winkle, Iolanda Leite
RO-MAN4
2023 Putting Robots in Context: Challenging the Influence of Voice and Empathic Behaviour on Trust
abstract
Trust is essential for social interactions, including those between humans and social artificial agents, such as robots. Several robot-related factors can contribute to the formation of trust. However, previous work has often treated trust as an absolute concept, whereas it is highly context-dependent, and it is possible that some robot-related features will influence trust in some contexts, but not in others. In this paper, we present the results of two video-based online studies aimed at investigating the role of robot voice and empathic behaviour on trust formation in a general context as well as in a task-specific context. We found that voice influences trust in the specific context, with no effect of voice or empathic behaviour in the general context. Thus, context mediated whether robot-related features play a role in people’s trust formation towards robots.
Marta Romeo, Ilaria Torre 0002, Sébastien Le Maguer, Angelo Cangelosi, Iolanda Leite
RO-MAN5
2023 Can a gender-ambiguous voice reduce gender stereotypes in human-robot interactions?
abstract
When deploying robots, its physical characteristics, role, and tasks are often fixed. Such factors can also be associated with gender stereotypes among humans, which then transfer to the robots. One factor that can induce gendering but is comparatively easy to change is the robot’s voice. Designing voice in a way that interferes with fixed factors might therefore be a way to reduce gender stereotypes in human-robot interaction contexts. To this end, we have conducted a video-based online study to investigate how factors that might inspire gendering of a robot interact. In particular, we investigated how giving the robot a gender-ambiguous voice can affect perception of the robot. We compared assessments (n=111) of videos in which a robot’s body presentation and occupation mis/matched with human gender stereotypes. We found evidence that a gender-ambiguous voice can reduce gendering of a robot endowed with stereotypically feminine or masculine attributes. The results can inform more just robot design while opening new questions regarding the phenomenon of robot gendering.
Ilaria Torre 0002, Erik Lagerstedt, Nathaniel Dennler, Katie Seaborn, Iolanda Leite, Éva Székely
RO-MAN5
2023 Hearing it Out: Guiding Robot Sound Design through Design Thinking
abstract
Sound can benefit human-robot interaction, but little work has explored questions on the design of nonverbal sound for robots. The unique confluence of sound design and robotics expertise complicates these questions, as most roboticists do not have sound design expertise, necessitating collaborations with sound designers. We sought to understand how roboticists and sound designers approach the problem of robot sound design through two qualitative studies. The first study followed discussions by robotics researchers in focus groups, where these experts described motivations to add robot sound for various purposes. The second study guided music technology students through a generative activity for robot sound design; these sound designers in-training demonstrated high variability in design intent, processes, and inspiration. To unify the two perspectives, we structured recommendations through the design thinking framework, a popular design process. The insights provided in this work may aid roboticists in implementing helpful sounds in their robots, encourage sound designers to enter into collaborations on robot sound, and give key tips and warnings to both.
Brian J. Zhang, Bastian Orthmann, Ilaria Torre 0002, Roberto Bresin, Jason Fick, Iolanda Leite, Naomi T. Fitter
RO-MAN6
2023 Sounding Robots: Design and Evaluation of Auditory Displays for Unintentional Human-robot Interaction
abstract
Non-verbal communication is important in HRI, particularly when humans and robots do not need to actively engage in a task together, but rather they co-exist in a shared space. Robots might still need to communicate states such as urgency or availability, and where they intend to go, to avoid collisions and disruptions. Sounds could be used to communicate such states and intentions in an intuitive and non-disruptive way. Here, we propose a multi-layer classification system for displaying various robot information simultaneously via sound. We first conceptualise which robot features could be displayed (robot size, speed, availability for interaction, urgency, and directionality); we then map them to a set of audio parameters. The designed sounds were then evaluated in five online studies, where people listened to the sounds and were asked to identify the associated robot features. The sounds were generally understood as intended by participants, especially when they were evaluated one feature at a time, and partially when they were evaluated two features simultaneously. The results of these evaluations suggest that sounds can be successfully used to communicate robot states and intended actions implicitly and intuitively.
Bastian Orthmann, Iolanda Leite, Roberto Bresin, Ilaria Torre 0002
ACM Trans. Hum. Robot Interact.2
2022 Designing Tangible Robot Mediated Co-located Games to Enhance Social Inclusion for Neurodivergent Children
abstract
Neurodivergent children with cognitive and communicative difficulties often experience a lower level of social integration in comparison to neurotypical children. Therefore it is crucial to understand social inclusion challenges and address exclusion. Since previous work shows that gamified robotic activities have a high potential to enable inclusive and collaborative environments we propose using robot-mediated games for enhancing social inclusion. In this work, we present the design of a multiplayer tangible Pacman game with three different inter-player interaction modalities: semi-dependent collaborative, dependent collaborative, and competitive. The initial usability evaluation and the observations of the experiments show the benefits of the game for creating collaborative and cooperative practices for the players and thus also potential for social interaction and social inclusion. Importantly, we observe that inter-player interaction design affects the communication between the players and their physical interaction with the game.
Arzu Güneysu, Ali Reza Majlesi, Victor Taburet, Sebastiaan A. Meijer, Iolanda Leite, Sanna Kuoppamäki
IDC5
2022 Asking Follow-Up Clarifications to Resolve Ambiguities in Human-Robot Conversation
abstract
When a robot aims to comprehend its human partner's request by identifying the referenced objects in Human-Robot Conversation, ambiguities can occur because the environment might contain many similar objects or the objects described in the request might be unknown to the robot. In the case of ambiguities, most of the systems ask users to repeat their request, which assumes that the robot is familiar with all of the objects in the environment. This assumption might lead to task failure, especially in complex real-world environments. In this paper, we address this challenge by presenting an interactive system that asks for follow-up clarifications to disambiguate the described objects using the pieces of information that the robot could understand from the request and the objects in the environment that are known to the robot. To evaluate our system while disambiguating the referenced objects, we conducted a user study with 63 participants. We analyzed the interactions when the robot asked for clarifications and when it asked users to redescribe the same object. Our results show that generating followup clarification questions helped the robot correctly identify the described objects with fewer attempts (i.e., conversational turns). Also, when people were asked clarification questions, they perceived the task as easier, and they evaluated the task understanding and competence of the robot as higher. Our code and anonymized dataset are publicly available11https://github.com/IrmakDogan/Resolving-Ambiguities.
Fethiye Irmak Dogan, Ilaria Torre 0002, Iolanda Leite
HRI3
2022 Learning Gaze Behaviors for Balancing Participation in Group Human-Robot Interactions
abstract
Robots can affect group dynamics. In particular, prior work has shown that robots that use hand-crafted gaze heuristics can influence human participation in group interactions. However, hand-crafting robot behaviors can be difficult and might have unexpected results in groups. Thus, this work explores learning robot gaze behaviors that balance human participation in conversational interactions. More specifically, we examine two techniques for learning a gaze policy from data: imitation learning (IL) and batch reinforcement learning (RL). First, we formulate the problem of learning a gaze policy as a sequential decision-making task focused on human turn-taking. Second, we experimentally show that IL can be used to combine strategies from hand-crafted gaze behaviors, and we formulate a novel reward function to achieve a similar result using batch RL. Finally, we conduct an offline evaluation of IL and RL policies and compare them via a user study (N=50). The results from the study show that the learned behavior policies did not compromise the interaction. Interestingly, the proposed reward for the RL formulation enabled the robot to encourage participants to take more turns during group human-robot interactions than one of the gaze heuristic behaviors from prior work. Also, the imitation learning policy led to more active participation from human participants than another prior heuristic behavior.
Sarah Gillet, Maria Teresa Parreira, Marynel Vázquez, Iolanda Leite
HRI4
2022 Automatic Frustration Detection Using Thermal Imaging
abstract
To achieve seamless interactions, robots have to be capable of reliably detecting affective states in real time. One of the possible states that humans go through while interacting with robots is frustration. Detecting frustration from RGB images can be challenging in some real-world situations; thus, we investigate in this work whether thermal imaging can be used to create a model that is capable of detecting frustration induced by cognitive load and failure. To train our model, we collected a data set from 18 participants experiencing both types of frustration induced by a robot. The model was tested using features from several modalities: thermal, RGB, Electrodermal Activity (EDA), and all three combined. When data from both frustration cases were combined and used as training input, the model reached an accuracy of 89% with just RGB features, 87% using only thermal features, 84% using EDA, and 86% when using all modalities. Furthermore, the highest accuracy for the thermal data was reached using three facial regions of interest: nose, forehead and lower lip.
Youssef Mohamed, Giulia Ballardini, Maria Teresa Parreira, Séverin Lemaignan, Iolanda Leite
HRI5
2022 Design Implications for Effective Robot Gaze Behaviors in Multiparty Interactions
abstract
Human-robot non-verbal communication has been a growing focus of research, as we realize its importance to achieve interaction goals (e.g. modulating turn-taking) and manage human perception of the interaction. Consequently, the development of models for robot non-verbal behavior, such as gaze, should be informed by studies of human reaction and perception to that behavior. Here, we look at data from two studies where two humans interact describing words to a robot. The robot tries to balance participation of the two players through a combination of gaze aversion, looking at the listener and looking at the speaker. We analyze how momentary gaze patterns reflect in the participant's turn length and perception of the robot, as well as in the participation imbalance. Our findings may be used as recommendations towards crafting robot gaze behaviors in multiparty interactions.
Maria Teresa Parreira, Sarah Gillet, Marynel Vázquez, Iolanda Leite
HRI4
2022 Correct Me If I'm Wrong: Using Non-Experts to Repair Reinforcement Learning Policies
abstract
Reinforcement learning has shown great potential for learning sequential decision-making tasks. Yet, it is difficult to anticipate all possible real-world scenarios during training, causing robots to inevitably fail in the long run. Many of these failures are due to variations in the robot's environment. Usually experts are called to correct the robot's behavior; however, some of these failures do not necessarily require an expert to solve them. In this work, we query non-experts online for help and explore 1) if/how non-experts can provide feedback to the robot after a failure and 2) how the robot can use this feedback to avoid such failures in the future by generating shields that restrict or correct its high-level actions. We demonstrate our approach on common daily scenarios of a simulated kitchen robot. The results indicate that non-experts can indeed understand and repair robot failures. Our generated shields accelerate learning and improve data-efficiency during retraining.
Sanne van Waveren, Christian Pek, Jana Tumova, Iolanda Leite
HRI4
2022 Norm-Breaking Responses to Sexist Abuse: A Cross-Cultural Human Robot Interaction Study
abstract
This article presents a cross-cultural replication of recent work on productively violating gender norms; specifically demonstrating that breaking norms can boost robot credibility while avoiding harmful stereotypes. In this work we demonstrate via a 3 (country) x 3 (robot behaviour) between-subject experiment that these findings replicate cross-culturally across the US, Sweden, and Japan, finding evidence that breaking gender norms boosts robot credibility regardless of gender or cultural context, and regardless of pretest gender biases. Our findings further motivate a call for feminist robots that subvert the existing gender norms of robot design.
Katie Winkle, Ryan Blake Jackson, Gaspar Isaac Melsión, Drazen Brscic, Iolanda Leite, Tom Williams 0001
HRI5
2022 Inference of Multi-Class STL Specifications for Multi-Label Human-Robot Encounters
abstract
This paper is interested in formalizing human trajectories in human-robot encounters. Inspired by robot navigation tasks in human-crowded environments, we consider the case where a human and a robot walk towards each other, and where humans have to avoid colliding with the incoming robot. Further, humans may describe different be-haviors, ranging from being in a hurry/minimizing completion time to maximizing safety. We propose a decision tree-based algorithm to extract STL formulae from multi-label data. Our inference algorithm learns STL specifications from data containing multiple classes, where instances can be labelled by one or many classes. We base our evaluation on a dataset of trajectories collected through an online study reproducing human-robot encounters.
Alexis Linard, Ilaria Torre 0002, Iolanda Leite, Jana Tumova
IROS3
2022 Safety-based Dynamic Task Offloading for Human-Robot Collaboration using Deep Reinforcement Learning
abstract
Robots with constrained hardware resources usually rely on Multi-access Edge Computing infrastructures to offload computationally expensive tasks to meet real-time and safety requirements. Offloading every task might not be the best option due to dynamic changes in the network conditions and can result in network congestion or failures. This work proposes a task offloading strategy for mobile robots in a Human-Robot Collaboration scenario that optimizes the edge resource usage and reduces network delays, leading to safety enhancement. The solution utilizes a Deep Reinforcement Learning (DRL) agent that observes safety and network metrics to dynamically decide at runtime if (i) a less accurate model should run on the robot; (ii) a more complex model should run on the edge; or (iii) the previous output should be reused through temporal coherence verification. Experiments are performed in a simulated warehouse where humans and robots have close interactions and safety needs are high. Our results show that the proposed DRL solution outperforms the baselines in several aspects. The edge is used only when the network performance is reliable, reducing the number of failures (up to 47 %). The latency is not only decreased (up to 68 %) but also adapted to the safety requirements (risk × latency reduced up to 48 %), avoiding unnecessary network congestion in safe situations and letting other devices use the network. Overall, the safety metrics get improved, such as the increased time in the safe zone by up to 3.1%.
Franco Ruggeri, Ahmad Terra, Alberto Hata, Rafia Inam, Iolanda Leite
IROS5
2022 Ice-Breakers, Turn-Takers and Fun-Makers: Exploring Robots for Groups with Teenagers
abstract
Successful, enjoyable group interactions are important in public and personal contexts, especially for teenagers whose peer groups are important for self-identity and self-esteem. Social robots seemingly have the potential to positively shape group interactions, but it seems difficult to effect such impact by designing robot behaviors solely based on related (human interaction) literature. In this article, we take a user-centered approach to explore how teenagers envisage a social robot "group assistant". We engaged 16 teenagers in focus groups, interviews, and robot testing to capture their views and reflections about robots for groups. Over the course of a two-week summer school, participants co-designed the action space for such a robot and experienced working with/wizarding it for 10+ hours. This experience further altered and deepened their insights into using robots as group assistants. We report results regarding teenagers’ views on the applicability and use of a robot group assistant, how these expectations evolved throughout the study, and their repeat interactions with the robot. Our results indicate that each group moves on a spectrum of need for the robot, reflected in use of the robot more (or less) for ice-breaking, turn-taking, and fun-making as the situation demanded.
Sarah Gillet, Katie Winkle, Giulia Belgiovine, Iolanda Leite
RO-MAN4
2022 Improving Visual Question Answering by Leveraging Depth and Adapting Explainability
abstract
During human-robot conversation, it is critical for robots to be able to answer users’ questions accurately and provide a suitable explanation for why they arrive at the answer they provide. Depth is a crucial component in producing more intelligent robots that can respond correctly as some questions might rely on spatial relations within the scene, for which 2D RGB data alone would be insufficient. Due to the lack of existing depth datasets for the task of VQA, we introduce a new dataset, VQA-SUNRGBD. When we compare our proposed model on this RGB-D dataset against the baseline VQN network on RGB data alone, we show that ours outperforms, particularly in questions relating to depth such as asking about the proximity of objects and relative positions of objects to one another. We also provide Grad-CAM activations to gain insight regarding the predictions on depth-related questions and find that our method produces better visual explanations compared to Grad-CAM on RGB data. To our knowledge, this work is the first of its kind to leverage depth and an explainability module to produce an explainable Visual Question Answering (VQA) system.
Amrita Panesar, Fethiye Irmak Dogan, Iolanda Leite
RO-MAN3
2021 Dimensional perception of a 'smiling McGurk effect'
abstract
Multisensory integration influences emotional perception, as the McGurk effect demonstrates for the communication between humans. Human physiology implicitly links the production of visual features with other modes like the audio channel: Face muscles responsible for a smiling face also stretch the vocal cords that results in a characteristic smiling voice. For artificial agents capable of multimodal expression, this linkage is modeled explicitly. In our study, we observe the influence of visual and audio channel on the perception of the agent’s emotional state. We created two virtual characters to control for anthropomorphic appearance. We record videos of these agents either with matching or mismatching emotional expression in the audio and visual channel. In an online study we measured the agent’s perceived valence and arousal. Our results show that a matched smiling voice and smiling face increase both dimensions of the Circumplex model of emotions: ratings of valence and arousal grow. When the channels present conflicting information, any type of smiling results in higher arousal rating, but only the visual channel increases the perceived valence. When engineers are constrained in their design choices, we suggest they should give precedence to convey the artificial agent’s emotional state through the visual channel.
Ilaria Torre 0002, Simon Holk, Emma Carrigan, Iolanda Leite, Rachel McDonnell, Naomi Harte
ACII4
2021 Using Explainability to Help Children UnderstandGender Bias in AI
abstract
Machine learning systems have become ubiquitous into our society. This has raised concerns about the potential discrimination that these systems might exert due to unconscious bias present in the data, for example regarding gender and race. Whilst this issue has been proposed as an essential subject to be included in the new AI curricula for schools, research has shown that it is a difficult topic to grasp by students. We propose an educational platform tailored to raise the awareness of gender bias in supervised learning, with the novelty of using Grad-CAM as an explainability technique that enables the classifier to visually explain its own predictions. Our study demonstrates that preadolescents (N=78, age 10-14) significantly improve their understanding of the concept of bias in terms of gender discrimination, increasing their ability to recognize biased predictions when they interact with the interpretable model, highlighting its suitability for educational programs.
Gaspar Isaac Melsión, Ilaria Torre 0002, Eva Vidal, Iolanda Leite
IDC4
2021 Robot Gaze Can Mediate Participation Imbalance in Groups with Different Skill Levels
abstract
Many small group activities, like working teams or study groups, have a high dependency on the skill of each group member. Differences in skill level among participants can affect not only the performance of a team but also influence the social interaction of its members. In these circumstances, an active member could balance individual participation without exerting direct pressure on specific members by using indirect means of communication, such as gaze behaviors. Similarly, in this study, we evaluate whether a social robot can balance the level of participation in a language skill-dependent game, played by a native speaker and a second language learner. In a between-subjects study (N = 72), we compared an adaptive robot gaze behavior, that was targeted to increase the level of contribution of the least active player, with a non-adaptive gaze behavior. Our results imply that, while overall levels of speech participation were influenced predominantly by personal traits of the participants, the robot's adaptive gaze behavior could shape the interaction among participants which lead to more even participation during the game.
Sarah Gillet, Ronald Cumbal, André Pereira 0001, José Lopes 0001, Olov Engwall, Iolanda Leite
HRI6
2021 Should Robots Chicken?: How Anthropomorphism and Perceived Autonomy Influence Trajectories in a Game-theoretic Problem
abstract
Two people walking towards each other in a colliding course is an everyday problem of human-human interaction. In spite of the different environmental and individual factors that might jeopardise successful human trajectories, people are generally skilled at avoiding crashing into each other. However, it is not clear if the same strategies will apply when a human is in a colliding course with a robot, nor which (if any) robot-related factors will influence the human's decision to swerve or not. In this work, we present the results of an online study where participants walked towards a virtual robot that differed in terms of anthropomorphism and perceived autonomy, and had to decide whether to swerve, or continue straight. The experiment was inspired by the game-theoretic game of chicken. We found that people performed more swerving actions when they believed the robot to be teleoperated by another participant. When they swerved, they also swerved closer to the robot with high levels of human-likeness, and farther away from the robot with low anthropomorphism score, suggesting a higher uncertainty about the mechanical-looking robot's intentions. These results are discussed in the context of socially-aware robot navigation, and will be used to design novel algorithms for robot trajectories that take robot-related differences into account.
Ilaria Torre 0002, Alexis Linard, Anders Steen, Jana Tumova, Iolanda Leite
HRI5
2021 Encoding Human Driving Styles in Motion Planning for Autonomous Vehicles
abstract
Driving styles play a major role in the acceptance and use of autonomous vehicles. Yet, existing motion planning techniques can often only incorporate simple driving styles that are modeled by the developers of the planner and not tailored to the passenger. We present a new approach to encode human driving styles through the use of signal temporal logic and its robustness metrics. Specifically, we use a penalty structure that can be used in many motion planning frameworks, and calibrate its parameters to model different automated driving styles. We combine this penalty structure with a set of signal temporal logic formula, based on the Responsibility-Sensitive Safety model, to generate trajectories that we expected to correlate with three different driving styles: aggressive, neutral, and defensive. An online study showed that people perceived different parameterizations of the motion planner as unique driving styles, and that most people tend to prefer a more defensive automated driving style, which correlated to their self-reported driving style.
Jesper Karlsson, Sanne van Waveren, Christian Pek, Ilaria Torre 0002, Iolanda Leite, Jana Tumova
ICRA5
2021 Formalizing Trajectories in Human-Robot Encounters via Probabilistic STL Inference
abstract
In this paper, we are interested in formalizing human trajectories in human-robot encounters. We consider a particular case where a human and a robot walk towards each other. A question that arises is whether, when, and how humans will deviate from their trajectory to avoid a collision. These human trajectories can then be used to generate socially acceptable robot trajectories. To model these trajectories, we propose a data-driven algorithm to extract a formal specification expressed in Signal Temporal Logic with probabilistic predicates. We evaluated our method on trajectories collected through an online study where participants had to avoid colliding with a robot in a shared environment. Further, we demonstrate that probabilistic STL is a suitable formalism to depict human behavior, choices and preferences in specific scenarios of social navigation.
Alexis Linard, Ilaria Torre 0002, Anders Steen, Iolanda Leite, Jana Tumova
IROS4
2020 Embodiment Effects in Interactions with Failing Robots
abstract
The increasing use of robots in real-world applications will inevitably cause users to encounter more failures in interactions. While there is a longstanding effort in bringing human-likeness to robots, how robot embodiment affects users' perception of failures remains largely unexplored. In this paper, we extend prior work on robot failures by assessing the impact that embodiment and failure severity have on people's behaviours and their perception of robots. Our findings show that when using a smart-speaker embodiment, failures negatively affect users' intention to frequently interact with the device, however not when using a human-like robot embodiment. Additionally, users significantly rate the human-like robot higher in terms of perceived intelligence and social presence. Our results further suggest that in higher severity situations, human-likeness is distracting and detrimental to the interaction. Drawing on quantitative findings, we discuss benefits and drawbacks of embodiment in robot failures that occur in guided tasks.
Dimosthenis Kontogiorgos, Sanne van Waveren, Olle Wallberg, André Pereira 0001, Iolanda Leite, Joakim Gustafson
CHI5
2020 Gesticulator: A framework for semantically-aware speech-driven gesture generation
abstract
During speech, people spontaneously gesticulate, which plays a key role in conveying information. Similarly, realistic co-speech gestures are crucial to enable natural and smooth interactions with social agents. Current end-to-end co-speech gesture generation systems use a single modality for representing speech: either audio or text. These systems are therefore confined to producing either acoustically-linked beat gestures or semantically-linked gesticulation (e.g., raising a hand when saying "high''): they cannot appropriately learn to generate both gesture types. We present a model designed to produce arbitrary beat and semantic gestures together. Our deep-learning based model takes both acoustic and semantic representations of speech as input, and generates gestures as a sequence of joint angle rotations as output. The resulting gestures can be applied to both virtual agents and humanoid robots. Subjective and objective evaluations confirm the success of our approach. The code and video are available at the project page svito-zar.github.io/gesticulator .
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexanderson, Iolanda Leite, Hedvig Kjellström
ICMI6
2019 Exploring Prosociality in Human-Robot Teams
abstract
This paper explores the role of prosocial behaviour when people team up with robots in a collaborative game that presents a social dilemma similar to a public goods game. An experiment was conducted with the proposed game in which each participant joined a team with a prosocial robot and a selfish robot. During 5 rounds of the game, each player chooses between contributing to the team goal (cooperate) or contributing to his individual goal (defect). The prosociality level of the robots only affects their strategies to play the game, as one always cooperates and the other always defects. We conducted a user study at the office of a large corporation with 70 participants where we manipulated the game result (winning or losing) in a between-subjects design. Results revealed two important considerations: (1) the prosocial robot was rated more positively in terms of its social attributes than the selfish robot, regardless of the game result; (2) the perception of competence, the responsibility attribution (blame/credit), and the preference for a future partner revealed significant differences only in the losing condition. These results yield important concerns for the creation of robotic partners, the understanding of group dynamics and, from a more general perspective, the promotion of a prosocial society.
Filipa Correia, Samuel Mascarenhas, Samuel Gomes, Patrícia Arriaga, Iolanda Leite, Rui Prada, Francisco S. Melo, Ana Paiva 0001
HRI5
2019 Personalization in Long-Term Human-Robot Interaction
abstract
For practical reasons, most human-robot interaction (HRI) studies focus on short-term interactions between humans and robots. However, such studies do not capture the difficulty of sustaining engagement and interaction quality across long-term interactions. Many real-world robot applications will require repeated interactions and relationship-building over the long term, and personalization and adaptation to users will be necessary to maintain user engagement and to build rapport and trust between the user and the robot. This full-day workshop brings together perspectives from a variety of research areas, including companion robots, elderly care, and educational robots, in order to provide a forum for sharing and discussing innovations, experiences, works-in-progress, and best practices which address the challenges of personalization in long-term HRI.
Bahar Irfan, Aditi Ramachandran, Samuel Spaulding, Dylan F. Glas, Iolanda Leite, Kheng Lee Koay
HRI5
2019 Comparing Human-Robot Proxemics Between Virtual Reality and the Real World
abstract
Virtual Reality (VR) can greatly benefit Human-Robot Interaction (HRI) as a tool to effectively iterate across robot designs. However, possible system limitations of VR could influence the results such that they do not fully reflect real-life encounters with robots. In order to better deploy VR in HRI, we need to establish a basic understanding of what the differences are between HRI studies in the real world and in VR. This paper investigates the differences between the real life and VR with a focus on proxemic preferences, in combination with exploring the effects of visual familiarity and spatial sound within the VR experience. Results suggested that people prefer closer interaction distances with a real, physical robot than with a virtual robot in VR. Additionally, the virtual robot was perceived as more discomforting than the real robot, which could result in the differences in proxemics. Overall, these results indicate that the perception of the robot has to be evaluated before the interaction can be studied. However, the results also suggested that VR settings with different visual familiarities are consistent with each other in how they affect HRI proxemics and virtual robot perceptions, indicating the freedom to study HRI in various scenarios in VR. The effect of spatial sound in VR drew a more complex picture and thus calls for more in-depth research to understand its influence on HRI in VR.
Marc van Almkerk, Sanne van Waveren, Elizabeth J. Carter, Iolanda Leite
HRI5
2019 Learning to Generate Unambiguous Spatial Referring Expressions for Real-World Environments
abstract
Referring to objects in a natural and unambiguous manner is crucial for effective human-robot interaction. Previous research on learning-based referring expressions has focused primarily on comprehension tasks, while generating referring expressions is still mostly limited to rule-based methods. In this work, we propose a two-stage approach that relies on deep learning for estimating spatial relations to describe an object naturally and unambiguously with a referring expression. We compare our method to the state of the art algorithm in ambiguous environments (e.g., environments that include very similar objects with similar relationships). We show that our method generates referring expressions that people find to be more accurate (~30% better) and would prefer to use (~32% more often).
Fethiye Irmak Dogan, Sinan Kalkan, Iolanda Leite
IROS3
2019 Take One For the Team: The Effects of Error Severity in Collaborative Tasks with Social Robots
abstract
We explore the effects of robot failure severity (no failure vs. low-impact vs. high-impact) on people's subjective ratings of the robot. We designed an escape room scenario in which one participant teams up with a remotely-controlled Pepper robot. We manipulated the robot's performance at the end of the game: the robot would either correctly follow the participant's instructions (control condition), the robot would fail but people could still complete the task of escaping the room (low-impact condition), or the robot's failure would cause the game to be lost (high-impact condition). Results showed no difference across conditions for people's ratings of the robot in terms of warmth, competence, and discomfort. However, people in the low-impact condition had significantly less faith in the robot's robustness in future escape room scenarios. Open-ended questions revealed interesting trends that are worth pursuing in the future: people may view task performance as a team effort and may blame their team or themselves more for the robot failure in case of a high-impact failure as compared to the low-impact failure.
Sanne van Waveren, Elizabeth J. Carter, Iolanda Leite
IVA3
2018 Using Constrained Optimization for Real-Time Synchronization of Verbal and Nonverbal Robot Behavior
abstract
Most of the motion re-targeting techniques are grounded on virtual character animation research, which means that they typically assume that the target embodiment has unconstrained joint angular velocities. However, because robots often do have such constraints, traditional re-targeting approaches can originate irregular delays in the robot motion. With the goal of ensuring synchronization between verbal and nonverbal behavior, this paper proposes an optimization framework for processing re-targeted motion sequences that addresses constraints such as joint angle and angular velocities. The proposed framework was evaluated on a humanoid robot using both objective and subjective metrics. While the analysis of the joint motion trajectories provides evidence that our framework successfully performs the desired modifications to ensure verbal and nonverbal behavior synchronization, results from a perceptual study showed that participants found the robot motion generated by our method more natural, elegant and lifelike than a control condition.
Aravind Elanjimattathil Vijayan, Simon Alexanderson, Jonas Beskow, Iolanda Leite
ICRA4
2018 A Comparison of Visualisation Methods for Disambiguating Verbal Requests in Human-Robot Interaction
abstract
Picking up objects requested by a human user is a common task in human-robot interaction. When multiple objects match the user's verbal description, the robot needs to clarify which object the user is referring to before executing the action. Previous research has focused on perceiving user's multimodal behaviour to complement verbal commands or minimising the number of follow up questions to reduce task time. In this paper, we propose a system for reference disambiguation based on visualisation and compare three methods to disambiguate natural language instructions. In a controlled experiment with a YuMi robot, we investigated realtime augmentations of the workspace in three conditions - head-mounted display, projector, and a monitor as the baseline - using objective measures such as time and accuracy, and subjective measures like engagement, immersion, and display interference. Significant differences were found in accuracy and engagement between the conditions, but no differences were found in task time. Despite the higher error rates in the head-mounted display condition, participants found that modality more engaging than the other two, but overall showed preference for the projector condition over the monitor and head-mounted display conditions.
Elena Sibirtseva, Dimosthenis Kontogiorgos, Olov Nykvist, Hakan Karaoguz, Iolanda Leite, Joakim Gustafson, Danica Kragic
RO-MAN5
2017 Persistent Memory in Repeated Child-Robot Conversations
abstract
Persistent memory is a critical mechanism in long-term human-robot interaction. In this work, we investigate how a robot can use information from prior conversations with the same child to foster a sense of relationship over time. To address this question, we conducted a repeated interaction study with three experimental conditions: a baseline control condition, in which the robot retains no information between conversations and relies on a typical elicitation-response paradigm; a persistence condition, in which children experience the same topic flow but with some robot turns that refer back to prior shared events; and a pro-active persistence condition, in which the robot attempts to offer its own feelings and opinions pro-actively and congruently with what it knows about the child. Our results indicate age differences with respect to the measures of interest. During conversations with the robot, older children who were assigned to the persistence conditions exhibited more positive affect, while younger children showed more positive affect in the control condition. Moreover, in a set of comparative judgments among robots they had played with, children in the augmented persistence condition considered PIPER to be the most intelligent and their favorite more often than children in the other conditions, overall, but the effect was more evident in the older children.
Iolanda Leite, André Pereira 0001, Jill Fain Lehman
IDC1
2017 Collaborative Storytelling between Robot and Child: A Feasibility Study
abstract
Joint storytelling is a common parent-child activity and brings multiple benefits such as improved language learning for children. Most existing storytelling robots offer rigid interaction with children and do not contribute to children's stories. In this paper, we envision a robot that collaborates with a child to create oral stories in a highly interactive manner. We performed a Wizard-of-Oz feasibility study, which involved 78 children between 4 and 10 years old, to compare two collaboration strategies: (1) inserting new story content and relating it to the existing story and (2) inserting content without relating it to the existing story. We hypothesize the first strategy can foster true collaboration and create rapport, whereas the second is a safe strategy when the robot cannot understand the story. We observed that, although the first strategy creates a heavier cognitive load, it was as enjoyable as the second. We also observed some indications that the first strategy may mitigate the difficulties in story creation for young children under the age of 7 and encourage children to speak more. This study suggests that a mixture strategy is feasible for robots in collaborative storytelling, providing sufficient cognitive challenge while concealing its shortcomings on natural language understanding.
Iolanda Leite, Jill Fain Lehman, Boyang Li 0001
IDC2
2017 Creating Prosodic Synchrony for a Robot Co-player in a Speech-controlled Game for Children
abstract
Synchrony is an essential aspect of human-human interactions. In previous work, we have seen how synchrony manifests in low-level acoustic phenomena like fundamental frequency, loudness, and the duration of keywords during the play of child-child pairs in a fast-paced, cooperative, language-based game. The correlation between the increase in such low-level synchrony and increase in enjoyment of the game suggests that a similar dynamic between child and robot co-players might also improve the child's experience. We report an approach to creating on-line acoustic synchrony by using a dynamic Bayesian network learned from prior recordings of child-child play to select from a predefined space of robot speech in response to real-time measurement of the child's prosodic features. Data were collected from 40 new children, each playing the game with both a synchronizing and non-synchronizing version of the robot. Results show a significant order effect: although all children grew to enjoy the game more over time, those that began with the synchronous robot maintained their own synchrony to it and achieved higher engagement compared with those that did not.
Najmeh Sadoughi, André Pereira 0001, Rishub Jain, Iolanda Leite, Jill Fain Lehman
HRI4
2017 Learning and Reusing Dialog for Repeated Interactions with a Situated Social Agent
James Kennedy 0001, Iolanda Leite, André Pereira 0001, Boyang Li 0001, Rishub Jain, Ricson Cheng, Eli Pincus, Elizabeth J. Carter, Jill Fain Lehman
IVA2
2017 Augmented reality dialog interface for multimodal teleoperation
abstract
We designed an augmented reality interface for dialog that enables the control of multimodal behaviors in telepresence robot applications. This interface, when paired with a telepresence robot, enables a single operator to accurately control and coordinate the robot's verbal and nonverbal behaviors. Depending on the complexity of the desired interaction, however, some applications might benefit from having multiple operators control different interaction modalities. As such, our interface can be used by either a single operator or pair of operators. In the paired-operator system, one operator controls verbal behaviors while the other controls nonverbal behaviors. A within-subjects user study was conducted to assess the usefulness and validity of our interface in both single and paired-operator setups. When faced with hard tasks, coordination between verbal and nonverbal behavior improves in the single-operator condition. Despite single operators being slower to produce verbal responses, verbal error rates were unaffected by our conditions. Finally, significantly improved presence measures such as mental immersion, sensory engagement, ability to view and understand the dialog partner, and degree of emotion occur for single operators that control both the verbal and nonverbal behaviors of the robot.
André Pereira 0001, Elizabeth J. Carter, Iolanda Leite, John Mars, Jill Fain Lehman
RO-MAN3
2017 Empathy in Virtual Agents and Robots: A Survey
abstract
This article surveys the area of computational empathy, analysing different ways by which artificial agents can simulate and trigger empathy in their interactions with humans. Empathic agents can be seen as agents that have the capacity to place themselves into the position of a user’s or another agent’s emotional situation and respond appropriately. We also survey artificial agents that, by their design and behaviour, can lead users to respond emotionally as if they were experiencing the agent’s situation. In the course of this survey, we present the research conducted to date on empathic agents in light of the principles and mechanisms of empathy found in humans. We end by discussing some of the main challenges that this exciting area will be facing in the future.
Ana Paiva 0001, Iolanda Leite, Hana Boukricha, Ipke Wachsmuth
ACM Trans. Interact. Intell. Syst.2
2016 The Robot Who Knew Too Much: Toward Understanding the Privacy/Personalization Trade-Off in Child-Robot Conversation
abstract
In human-human conversation we elicit, share and use information as a way of defining and building relationships -- how information is revealed, and by whom, matters. A similar goal of using conversation as a relationship-building mechanism in human-robot interaction might or might not require the same degree of nuance. We explore what happens in the increasingly likely situation that a robot has sensed information about a child of which the child is unaware, then discloses that information in conversation in an effort to personalize the child's experience. In a pilot study, 28 children conversed with a social robot that either told a story with characters already introduced into the conversation by the child (control) or characters hidden by the child in a treasure chest that the child was holding (experimental). Cumulative evidence showed that all participants in the experimental condition noticed the robot's violation of expectations, but younger children (4 to 6 years) exhibited more contained emotional reactions than older children (7 to 10 years), and girls expressed more negative affect than boys. Despite the immediate response, post-conversation measures suggest that the single event did not have an impact on children's ratings of robot likeability or their willingness to interact with the robot again.
Iolanda Leite, Jill Fain Lehman
IDC1
2016 2nd Workshop on Evaluating Child Robot Interaction
abstract
Many researchers have started to explore natural interaction scenarios for children. No matter if these children are normally developing or have special needs, evaluating Child-Robot Interaction (CRI) is a challenge. To find methods that work well and provide reliable data is difficult, for example because commonly used methods such as questionnaires do not work well particularly with younger children. Previous research has shown that children need support in expressing how they feel about technology. Given this, researchers often choose time-consuming behavioral measures from observations to evaluate CRI. However, these are not necessarily comparable between studies and robots.
Cristina Zaga, Manja Lohse, Vicky Charisi, Vanessa Evers, Mark A. Neerincx, Takayuki Kanda 0001, Iolanda Leite
HRI7
2016 Semi-situated learning of verbal and nonverbal content for repeated human-robot interaction
abstract
Content authoring of verbal and nonverbal behavior is a limiting factor when developing agents for repeated social interactions with the same user. We present PIP, an agent that crowdsources its own multimodal language behavior using a method we call semi-situated learning. PIP renders segments of its goal graph into brief stories that describe future situations, sends the stories to crowd workers who author and edit a single line of character dialog and its manner of expression, integrates the results into its goal state representation, and then uses the authored lines at similar moments in conversation. We present an initial case study in which the language needed to host a trivia game interaction is learned pre-deployment and tested in an autonomous system with 200 users "in the wild." The interaction data suggests that the method generates both meaningful content and variety of expression.
Iolanda Leite, André Pereira 0001, Allison Funkhouser, Boyang Li 0001, Jill Fain Lehman
ICMI1
2016 A thermal emotion classifier for improved human-robot interaction
abstract
In their expanding role as tutors, home and healthcare assistants, robots must effectively interact with individuals of varying ability and temperament. Indeed, deploying robots in long-term social engagements will almost certainly require robots to reliably detect and adapt to changes in the demeanor of social partners to promote trust and more productive collaboration. However, the recognition of emotional state typically relies on the interpretation of very subtle cues, often varying from one person to the next. In addition, while facial expressions, body posture and features of speech have been used to detect affective changes, the robustness of these measures is often hindered by cultural and age differences. Recently, infrared thermography has shown promise in detecting guilt, fear and stress, indicating that it may be a viable sensing modality for improved human-robot interaction. In this study, we evaluated the efficacy of using a far infrared (FIR) camera for detecting robot-elicited affective response compared to video-elicited affective response by tracking thermal changes in five areas of the face. Further, we analyzed localized changes in the face to assess whether thermal and electrodermal responses to emotions elicited by traditional video techniques and by robots are similar. Finally, we performed principal component analysis to reduce the dimensionality of data and evaluated the performance using machine learning techniques for classifying thermal data by emotion state, resulting in a thermal classifier with a performance accuracy of 77.5%.
Laura Boccanfuso, Quan Wang 0003, Iolanda Leite, Beibin Li, Colette Torres, Lisa Chen, Nicole Salomons, Claire E. Foster, Erin Barney, Amy Yeo-jin Ahn, Brian Scassellati, Frédérick Shic
RO-MAN3
2016 Study of children's hugging for interactive robot design
abstract
We have developed a toy sized humanoid robot with soft air-filled modules on its links which sense contact and protect the robot and any interacting humans from damaging collisions. This robot, meant for robust physical interaction, is required to endure contact with children in the form of hugs and other playful interactions. It is therefore necessary to quantify the forces exerted during these interactions so that robots can be designed to both withstand these forces, as well as interact safely and intuitively in these situations. To quantify the range of forces exerted by children when performing both soft and strong hugs, we conducted a study in which 28 children (11 boys, 17 girls) between 4 and 10 years old hugged a pressure sensing doll while the pressure was recorded. We found a child's maximum expected hugging force (2.623 psi for our setup) during free play. The data gathered in this study will guide the further development of our physically interactive robot.
Joohyung Kim, Alexander Alspach, Iolanda Leite, Katsu Yamane
RO-MAN3
2016 Autonomous disengagement classification and repair in multiparty child-robot interaction
abstract
As research on robotic tutors increases, it becomes more relevant to understand whether and how robots will be able to keep students engaged over time. In this paper, we propose an algorithm to monitor engagement in small groups of children and trigger disengagement repair interventions when necessary. We implemented this algorithm in a scenario where two robot actors play out interactive narratives around emotional words and conducted a field study where 72 children interacted with the robots three times in one of the following conditions: control (no disengagement repair), targeted (interventions addressing the child with the highest disengagement level) and general (interventions addressing the whole group). Surprisingly, children in the control condition had higher narrative recall than in the two experimental conditions, but no significant differences were found in the emotional interpretation of the narratives. When comparing the two different types of disengagement repair strategies, participants who received targeted interventions had higher story recall and emotional understanding, and their valence after disengagement repair interventions increased over time.
Iolanda Leite, Marissa McCoy, Monika Lohani, Nicole Salomons, Kara McElvaine, Charlene K. Stokes, Susan E. Rivers, Brian Scassellati
RO-MAN1
2015 Emotional Storytelling in the Classroom: Individual versus Group Interaction between Children and Robots
abstract
Robot assistive technology is becoming increasingly prevalent. Despite the growing body of research in this area, the role of type of interaction (i.e., small groups versus individual interactions) on effectiveness of interventions is still unclear. In this paper, we explore a new direction for socially assistive robotics, where multiple robotic characters interact with children in an interactive storytelling scenario. We conducted a between-subjects repeated interaction study where a single child or a group of three children interacted with the robots in an interactive narrative scenario. Results show that although the individual condition increased participant's story recall abilities compared to the group condition, the emotional interpretation of the story content seemed more dependent on the difficulty level rather than the study condition. Our findings suggest that, despite the type of interaction, interactive narratives with multiple robots are a promising approach to foster children's development of social-related skills.
Iolanda Leite, Marissa McCoy, Monika Lohani, Daniel Ullman 0002, Nicole Salomons, Charlene K. Stokes, Susan E. Rivers, Brian Scassellati
HRI1
2015 Comparing Models of Disengagement in Individual and Group Interactions
abstract
Changes in type of interaction (e.g., individual vs. group interactions) can potentially impact data-driven models developed for social robots. In this paper, we provide a first investigation in the effects of changing group size in data-driven models for HRI, by analyzing how a model trained on data collected from participants interacting individually performs in test data collected from group interactions, and vice-versa. Another model combining data from both individual and group interactions is also investigated. We perform these experiments in the context of predicting disengagement behaviors in children interacting with two social robots. Our results show that a model trained with group data generalizes better to individual participants than the other way around. The mixed model seems a good compromise, but it does not achieve the performance levels of the models trained for a specific type of interaction.
Iolanda Leite, Marissa McCoy, Daniel Ullman 0002, Nicole Salomons, Brian Scassellati
HRI1
2015 Classification of Children's Social Dominance in Group Interactions with Robots
abstract
As social robots become more widespread in educational environments, their ability to understand group dynamics and engage multiple children in social interactions is crucial. Social dominance is a highly influential factor in social interactions, expressed through both verbal and nonverbal behaviors. In this paper, we present a method for determining whether a participant is high or low in social dominance in a group interaction with children and robots. We investigated the correlation between many verbal and nonverbal behavioral features with social dominance levels collected through teacher surveys. We additionally implemented Logistic Regression and Support Vector Machines models with classification accuracies of 81% and 89%, respectively, showing that using a small subset of nonverbal behavioral features, these models can successfully classify children's social dominance level. Our approach for classifying social dominance is novel not only for its application to children, but also for achieving high classification accuracies using a reduced set of nonverbal features that, in future work, can be automatically extracted with current sensing technology.
Sarah Sebo, Iolanda Leite, Natalie Warren, Brian Scassellati
ICMI2
2014 Smart Human, Smarter Robot: How Cheating Affects Perceptions of Social Agency
Daniel Ullman 0002, Iolanda Leite, Jonathan Phillips, Julia Kim-Cohen, Brian Scassellati
CogSci2
2014 Teachers' views on the use of empathic robotic tutors in the classroom
abstract
In this paper, we describe the results of an interview study conducted across several European countries on teachers' views on the use of empathic robotic tutors in the classroom. The main goals of the study were to elicit teachers' thoughts on the integration of the robotic tutors in the daily school practice, understanding the main roles that these robots could play and gather teachers' main concerns about this type of technology. Teachers' concerns were much related to the fairness of access to the technology, robustness of the robot in students' hands and disruption of other classroom activities. They saw a role for the tutor in acting as an engaging tool for all, preferably in groups, and gathering information about students' learning progress without taking over the teachers' responsibility for the actual assessment. The implications of these results are discussed in relation to teacher acceptance of ubiquitous technologies in general and robots in particular.
Sofia Serholt, Wolmet Barendregt, Iolanda Leite, Helen Hastie, Aiden Jones, Ana Paiva 0001, Asimina Vasalou, Ginevra Castellano
RO-MAN3
2014 Context-Sensitive Affect Recognition for a Robotic Game Companion
abstract
Social perception abilities are among the most important skills necessary for robots to engage humans in natural forms of interaction. Affect-sensitive robots are more likely to be able to establish and maintain believable interactions over extended periods of time. Nevertheless, the integration of affect recognition frameworks in real-time human-robot interaction scenarios is still underexplored. In this article, we propose and evaluate a context-sensitive affect recognition framework for a robotic game companion for children. The robot can automatically detect affective states experienced by children in an interactive chess game scenario. The affect recognition framework is based on the automatic extraction of task features and social interaction-based features. Vision-based indicators of the children’s nonverbal behaviour are merged with contextual features related to the game and the interaction and given as input to support vector machines to create a context-sensitive multimodal system for affect recognition. The affect recognition framework is fully integrated in an architecture for adaptive human-robot interaction. Experimental evaluation showed that children’s affect can be successfully predicted using a combination of behavioural and contextual data related to the game and the interaction with the robot. It was found that contextual data alone can be used to successfully predict a subset of affective dimensions, such as interest toward the robot. Experiments also showed that engagement with the robot can be predicted using information about the user’s valence, interest and anticipatory behaviour. These results provide evidence that social engagement can be modelled as a state consisting of affect and attention components in the context of the interaction.
Ginevra Castellano, Iolanda Leite, André Pereira 0001, Carlos Martinho, Ana Paiva 0001, Peter W. McOwan
ACM Trans. Interact. Intell. Syst.2
2013 Fun and fair: influencing turn-taking in a multi-party game with a virtual agent
abstract
Language-based interfaces for children hold great promise in education, therapy, and entertainment. An important subset of these interfaces includes those with a virtual agent that mediates the interaction. When participants are groups of children, the agent will need to exert a certain amount of turn-taking control to ensure that all group members participate and benefit from the experience, but must do so without being so overtly directive as to undermine the children's enjoyment of and engagement in the task. We present a hierarchy of nonverbal and verbal behaviors that a virtual agent can employ flexibly when passing the conversational turn. When used effectively, these behaviors can equalize participation, and potentially decrease the amount of overlapping speech among participants, improving automatic speech recognition in turn. We evaluated the behaviors by having children play a language-based game twice, once with a flexible host and once with an inflexible host that did not have access to the behaviors. Post-game opinion cards revealed no difference between the conditions with respect to fun or likability of the host, despite the flexible agent eliciting more evenly distributed play.
Sean Andrist, Iolanda Leite, Jill Fain Lehman
IDC2
2013 Towards empathic artificial tutors
Amol A. Deshmukh, Ginevra Castellano, Arvid Kappas, Wolmet Barendregt, Fernando Nabais, Ana Paiva 0001, Tiago Ribeiro 0001, Iolanda Leite, Ruth Aylett
HRI8
2013 Sensors in the wild: exploring electrodermal activity in child-robot interaction
Iolanda Leite, Rui Henriques, Carlos Martinho, Ana Paiva 0001
HRI1
2013 Managing chaos: models of turn-taking in character-multichild interactions
abstract
Turn-taking decisions in multiparty settings are complex, especially when the participants are children. Our goal is to endow an interactive character with appropriate turn-taking behavior using visual, audio and contextual features. To that end, we investigate three distinct turn-taking models: a baseline model grounded in established turn-taking rules for adults and two machine learning models, one trained with data collected in situ and the other trained with data collected in more controlled conditions. The three models are shown to have different profiles of behavior during silences, overlapping speech, and at the end of participants' turns. An exploratory user evaluation focusing on the decision points where the models differ showed clear preference for the machine learning models over the baseline model. The results indicate that the rules for language interactions with small groups of children are not simply an extension of the rules for interacting with small groups of adults.
Iolanda Leite, Hannaneh Hajishirzi, Sean Andrist, Jill Fain Lehman
ICMI1
2013 The influence of empathy in human-robot relations
Iolanda Leite, André Pereira 0001, Samuel Mascarenhas, Carlos Martinho, Rui Prada, Ana Paiva 0001
Int. J. Hum. Comput. Stud.1
2012 Modelling empathic behaviour in a robotic game companion for children: an ethnographic study in real-world settings
abstract
The idea of autonomous social robots capable of assisting us in our daily lives is becoming more real every day. However, there are still many open issues regarding the social capabilities that those robots should have in order to make daily interactions with humans more natural. For example, the role of affective interactions is still unclear. This paper presents an ethnographic study conducted in an elementary school where 40 children interacted with a social robot capable of recognising and responding empathically to some of the children's affective states. The findings suggest that the robot's empathic behaviour affected positively how children perceived the robot. However, the empathic behaviours should be selected carefully, under the risk of having the opposite effect. The target application scenario and the particular preferences of children seem to influence the degree of empathy that social robots should be endowed with.
Iolanda Leite, Ginevra Castellano, André Pereira 0001, Carlos Martinho, Ana Paiva 0001
HRI1
2011 Automatic analysis of affective postures and body motion to detect engagement with a game companion
abstract
The design of an affect recognition system for socially perceptive robots relies on representative data: human-robot interaction in naturalistic settings requires an affect recognition system to be trained and validated with contextualised affective expressions, that is, expressions that emerge in the same interaction scenario of the target application. In this paper we propose an initial computational model to automatically analyse human postures and body motion to detect engagement of children playing chess with an iCat robot that acts as a game companion. Our approach is based on vision-based automatic extraction of expressive postural features from videos capturing the behaviour of the children from a lateral view. An initial evaluation, conducted by training several recognition models with contextualised affective postural expressions, suggests that patterns of postural behaviour can be used to accurately predict the engagement of the children with the robot, thus making our approach suitable for integration into an affect recognition system for a game companion in a real world scenario.
Jyotirmay Sanghvi, Ginevra Castellano, Iolanda Leite, André Pereira 0001, Peter W. McOwan, Ana Paiva 0001
HRI3
2011 Expressing Emotions on Robotic Companions with Limited Facial Expression Capabilities
Tiago Ribeiro 0001, Iolanda Leite, Jan Kedzierski, Adam Oleksy, Ana Paiva 0001
IVA2
2011 Using Adaptive Empathic Responses to Improve Long-Term Interaction with Social Robots
Iolanda Leite
UMAP1
2010 "Why Can't We Be Friends?" An Empathic Game Companion for Long-Term Interaction
Iolanda Leite, Samuel Mascarenhas, André Pereira 0001, Carlos Martinho, Rui Prada, Ana Paiva 0001
IVA1
2010 Inter-ACT: an affective and contextually rich multimodal video corpus for studying interaction with robots
abstract
The Inter-ACT (INTEracting with Robots - Affect Context Task) corpus is an affective and contextually rich multimodal video corpus containing affective expressions of children playing chess with an iCat robot. It contains videos that capture the interaction from different perspectives and includes synchronised contextual information about the game and the behaviour displayed by the robot. The Inter-ACT corpus is mainly intended to be a comprehensive repository of naturalistic and contextualised, task-dependent data for the training and evaluation of an affect recognition system in an educational game scenario. The richness of contextual data that captures the whole human-robot interaction cycle, together with the fact that the corpus was collected in the same interaction scenario of the target application, make the Inter-ACT corpus unique in its genre.
Ginevra Castellano, Iolanda Leite, André Pereira 0001, Carlos Martinho, Ana Paiva 0001, Peter W. McOwan
ACM Multimedia2
2009 Detecting user engagement with a robot companion using task and social interaction-based features
abstract
Affect sensitivity is of the utmost importance for a robot companion to be able to display socially intelligent behaviour, a key requirement for sustaining long-term interactions with humans. This paper explores a naturalistic scenario in which children play chess with the iCat, a robot companion. A person-independent, Bayesian approach to detect the user's engagement with the iCat robot is presented. Our framework models both causes and effects of engagement: features related to the user's non-verbal behaviour, the task and the companion's affective reactions are identified to predict the children's level of engagement. An experiment was carried out to train and validate our model. Results show that our approach based on multimodal integration of task and social interaction-based features outperforms those based solely on non-verbal behaviour or contextual information (94.79 % vs. 93.75 % and 78.13 %).
Ginevra Castellano, André Pereira 0001, Iolanda Leite, Ana Paiva 0001, Peter W. McOwan
ICMI3
2009 As Time goes by: Long-term evaluation of social presence in robotic companions
abstract
Given the recent advances in robot and synthetic character technology, many researchers are now focused on ways of establishing social relations between these agents and humans over long periods of time. Early studies have shown that the novelty effect of robots and agents quickly wears out and that people change their attitudes and preferences towards them over time. In this paper, we study the role of social presence in long-term human-robot interactions. We conducted a study where children played chess exercises with a social robot over a five week period. With this experiment, we identified possible key issues that should be considered when designing social robots for long-term interactions.
Iolanda Leite, Carlos Martinho, André Pereira 0001, Ana Paiva 0001
RO-MAN1
2008 Are emotional robots more fun to play with?
abstract
In this paper we describe a robotic game buddy whose emotional behaviour is influenced by the state of the game. Using the iCat robot and chess as the game scenario, an architecture for incorporating emotions as a result of a heuristic evaluation of the state of the game was developed. The game buddy was evaluated in two ways. First, we investigated the effects of the characterpsilas emotional behaviour on the userpsilas perception of the game state. And secondly we compared a robotic with a screen based version of the iCat in terms of their influence on userpsilas enjoyment. The results suggested that userpsilas perception of the game increases with the iCatpsilas emotional behaviour, and that the enjoyment is higher when interacting with the robotic version.
Iolanda Leite, André Pereira 0001, Carlos Martinho, Ana Paiva 0001
RO-MAN1