Angelica Lim

dblp:31/10008 · DBLP profile ↗
← Back
31ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0001-9288-0380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 1 first-author · 18 since 2021Systems, architecture and hardware · 12 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 12 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BERSting at the screams: A benchmark for distanced, emotional and shouted speech recognition
abstract
Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggle in complex, real-world, situations, such as with distanced speech. Previous challenges have released datasets to address the issue of distanced ASR, however, the focus remains primarily on distance, specifically relying on multi-microphone array systems. Here we present the B(asic) E(motion) R(andom phrase) S(hou)t(s) (BERSt) dataset. The dataset contains almost 4 h of English speech from 98 actors with varying regional and non-native accents. The data was collected on smartphones in the actors homes and therefore includes at least 98 different acoustic environments. The data also includes 7 different emotion prompts and both shouted and spoken utterances. The smartphones were places in 19 different positions, including obstructions and being in a different room than the actor. This data is publicly available for use and can be used to evaluate a variety of speech recognition tasks, including: ASR, shout detection, and speech emotion recognition (SER). We provide initial benchmarks for ASR and SER tasks, and find that ASR degrades both with an increase in distance and shout level and shows varied performance depending on the intended emotion. Our results show that the BERSt dataset is challenging for both ASR and SER tasks and continued work is needed to improve the robustness of such systems for more accurate real-world use.
Paige Tuttosi, Mantaj Dhillon, Luna Sang, Shane Eastwood, Poorvi Bhatia, Quang Minh Dinh, Avni Kapoor, Yewon Jin, Angelica Lim
Comput. Speech Lang.9
2026 Past, Present, and Future: A Survey of the Evolution of Affective Robotics for Well-Being
abstract
Recent research in affective robots has recognized their potential in supporting human well-being. Due to rapidly developing affective and artificial intelligence technologies, this field of research has undergone explosive expansion and advancement in recent years. In order to develop a deeper understanding of recent advancements, we present a systematic review of the past 10 years of research in affective robotics for wellbeing. In this review, we identify the domains of well-being that have been studied, the methods used to investigate affective robots for well-being, and how these have evolved over time. We also examine the evolution of the multifaceted research topic from three lenses: technical, design, and ethical. Finally, we discuss future opportunities for research based on the gaps we have identified in our review – proposing pathways to take affective robotics from the past and present to the future. The results of our review are of interest to human-robot interaction and affective computing researchers, as well as clinicians and well-being professionals who may wish to examine and incorporate affective robotics in their practices.
Micol Spitale, Minja Axelsson, Sooyeon Jeong, Paige Tuttosi, Caitlin A. Stamatis, Guy Laban, Angelica Lim, Hatice Gunes
IEEE Trans. Affect. Comput.7
2025 Take a Look, it's in a Book, a Reading Robot
abstract
We demonstrate EmojiVoice, a free, customizable text-to-speech (TTS) toolkit for expressive speech on social robots. We demonstrate our voices through storytelling. This task is aimed to be deployed in classrooms, or libraries where the robot can read a story out loud to children. Moreover, we introduce adaptive clarity to to noisy environments and those with reduced comprehension ability. This storytelling robot voice allows us to demonstrate how, using our light weight and customizable TTS, we are able to have a voice that is expressive, engaging, clear and socially appropriate for the task, improving interactions with and perceptions of social robots.
Paige Tuttosi, Shivam Mehta, Zachary Syvenky, Bermet Burkanova, Mohammed Hfsafsti, Yue Wang 0065, H. Henny Yeung, Gustav Eje Henter, Jean-Julien Aucouturier, Angelica Lim
HRI10
2025 Generating Consistent Prosodic Patterns from Open-Source TTS Systems
Ha Eun Shim, Olivia Yung, Paige Tuttosi, Boey Kwan, Angelica Lim, Yue Wang 0065, H. Henny Yeung
INTERSPEECH5
2025 MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
abstract
We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides fine-grained, structured descriptions that capture the full complexity of human movement—including expressive actions (e.g., emotions, stylistic walking) and interactions beyond standard motion capture datasets. MotionScript serves as both a descriptive tool and a training resource for text-to-motion models, enabling the synthesis of highly realistic and diverse human motions from text. By augmenting motion datasets with MotionScript captions, we demonstrate significant improvements in out-of-distribution motion generation, allowing large language models (LLMs) to generate motions that extend beyond existing data. Additionally, MotionScript opens new applications in animation, virtual human simulation, and robotics, providing an interpretable bridge between intuitive descriptions and motion synthesis. To the best of our knowledge, this is the first attempt to systematically translate 3D motion into structured natural language without requiring training data. Code, dataset, and examples are available at https://pjyazdian.github.io/MotionScript
Payam Jome Yazdian, Rachel Lagasse, Hamid Mohammadi, Eric Liu 0004, Angelica Lim
IROS6
2025 EmojiVoice: Towards long-term controllable expressivity in robot speech
abstract
Humans vary their expressivity when speaking for extended periods to maintain engagement with their listener. Although social robots tend to be deployed with "expressive" joyful voices, they lack this long-term variation found in human speech. Foundation model text-to-speech systems are beginning to mimic the expressivity in human speech, but they are difficult to deploy offline on robots. We present EmojiVoice, a free, customizable text-to-speech (TTS) toolkit that allows social roboticists to build temporally variable, expressive speech on social robots. We introduce emoji-prompting to allow fine-grained control of expressivity on a phase level and use the lightweight Matcha-TTS backbone to generate speech in real-time. We explore three case studies: (1) a scripted conversation with a robot assistant, (2) a storytelling robot, and (3) an autonomous speech-to-speech interactive agent. We found that using varied emoji prompting improved the perception and expressivity of speech over a long period in a storytelling task, but expressive voice was not preferred in the assistant use case.
Paige Tuttosi, Shivam Mehta, Zachary Syvenky, Bermet Burkanova, Gustav Eje Henter, Angelica Lim
RO-MAN6
2025 Systematic Review of Social Robots for Health and Wellbeing: A Personal Healthcare Journey Lens
abstract
Social robots have great potential in supporting individuals’ physical and mental health/wellbeing. While they have been increasingly evaluated in some domains, such as with children with autism, their evaluation has not been as extensive in other areas. We present a systematic review of domains in which social robots have been evaluated specifically in health/wellbeing contexts. We ask which robots have been evaluated, who the participants were, and how participants interacted with the robots. PRISMA guidelines for systematic reviews were followed. Articles with children as participants, using a purely robotic device, and in languages other than English were excluded. A total of 9,362 peer-reviewed articles (up to February 2021) from ACM DL, IEEE Xplore, Scopus, PubMed, and PsychInfo were identified. After applying the inclusion/exclusion criteria 443 articles were included in the review. The majority of studies were conducted at care centers while studies in hospitals/clinics have seen relatively limited attention. In many cases, the social robots were not programmed for specific health-related tasks, limiting their application. We also discuss robots used in real-world settings and propose a “Personal healthcare journey,” which includes different stages of one’s life which could benefit from a social robot, with the goal of increasing long-term adoption of social robots for supporting health/wellbeing.
Moojan Ghafurian, Shruti Chandra, Rebecca Hutchinson, Angelica Lim, Ishan Baliyan, Jimin Rhim, Garima Gupta, Alexander Mois Aroyo, Samira Rasouli, Kerstin Dautenhahn
ACM Trans. Hum. Robot Interact.4
2024 Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
abstract
The emotional theory of mind problem requires facial expressions, body pose, contextual information and implicit commonsense knowledge to reason about the person's emotion and its causes, making it currently one of the most difficult problems in affective computing. In this work, we propose multiple methods to incorporate the emotional reasoning capabilities by constructing “narrative captions” relevant to emotion perception, that includes contextual and physical signal descriptors that focuses on “Who”, “What”, “Where” and “How” questions related to the image and emotions of the individual. We propose two distinct ways to construct these captions using zero-shot classifiers (CLIP) and fine-tuning visual-language models (LLaVA) over human generated descriptors. We further utilize these captions to guide the reasoning of language (GPT-4) and vision-language models (LLa Va, GPT-Vision). We evaluate the use of the resulting models in an image-to-language-to-emotion task. Our experiments showed that combining the “Fast” narrative descriptors and “Slow” reasoning of language models is a promising way to achieve emotional theory of mind.
Yasaman Etesam, Özge Nilay Yalçin, Chuxuan Zhang, Angelica Lim
ACII4
2024 Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
Paige Tuttosi, H. Henny Yeung, Yue Wang 0065, Fenqi Wang, Guillaume Denis, Jean-Julien Aucouturier, Angelica Lim
INTERSPEECH7
2024 Contextual Emotion Recognition using Large Vision Language Models
abstract
How does the person in the bounding box feel?" Achieving human-level recognition of the apparent emotion of a person in real world situations remains an unsolved task in computer vision. Facial expressions are not enough: body pose, contextual knowledge, and commonsense reasoning all contribute to how humans perform this emotional theory of mind task. In this paper, we examine two major approaches enabled by recent large vision language models: 1) image captioning followed by a language-only LLM, and 2) vision language models, under zero-shot and fine-tuned setups. We evaluate the methods on the Emotions in Context (EMOTIC) dataset and demonstrate that a vision language model, fine-tuned even on a small dataset, can significantly outperform traditional baselines. The results of this work aim to help robots and agents perform emotionally sensitive decision-making and interaction in the future.
Yasaman Etesam, Özge Nilay Yalçin, Chuxuan Zhang, Angelica Lim
IROS4
2024 React to This! How Humans Challenge Interactive Agents using Nonverbal Behaviors
abstract
How do people use their faces and bodies to test the interactive abilities of a robot? Making lively, believable agents is often seen as a goal for robots and virtual agents but believability can easily break down. In this Wizard-of-Oz (WoZ) study, we observed 1169 nonverbal interactions between 20 participants and 6 types of agents. We collected the nonverbal behaviors participants used to challenge the characters physically, emotionally, and socially. The participants interacted freely with humanoid and non-humanoid forms: a robot, a human, a penguin, a pufferfish, a banana, and a toilet. We present a human behavior codebook of 188 unique nonverbal behaviors used by humans to test the virtual characters. The insights and design strategies drawn from video observations aim to help build more interaction-aware and believable robots, especially when humans push them to their limits.
Chuxuan Zhang, Bermet Burkanova, Lawrence H. Kim, Lauren Yip, Ugo Cupcic, Stéphane Lallée, Angelica Lim
IROS7
2024 Predicting Long-Term Human Behaviors in Discrete Representations via Physics-Guided Diffusion
abstract
Long-term human trajectory prediction is a challenging yet critical task in robotics and autonomous systems. Prior work that studied how to predict accurate short-term human trajectories with only unimodal features often failed in long-term prediction. Reinforcement learning provides a good solution for learning human long-term behaviors but can suffer from challenges in data efficiency and optimization. In this work, we propose a long-term human trajectory forecasting framework that leverages a guided diffusion model to generate diverse long-term human behaviors in a high-level latent action space, obtained via a hierarchical action quantization scheme using a VQ-VAE to discretize continuous trajectories and the available context. The latent actions are predicted by our guided diffusion model, which uses physics-inspired guidance at test time to constrain generated multimodal action distributions. Specifically, we use reachability analysis during the reverse denoising process to guide the diffusion steps toward physically feasible latent actions. We evaluate our framework on two publicly available human trajectory forecasting datasets: SFU-Store-Nav and JRDB, and extensive experimental results show that our framework achieves superior performance in long-term human trajectory forecasting.
Zhitian Zhang, Anjian Li, Angelica Lim, Mo Chen 0001
IROS3
2024 EmoStyle: One-Shot Facial Expression Editing Using Continuous Emotion Parameters
abstract
Recent studies have achieved impressive results in face generation and editing of facial expressions. However, existing approaches either generate a discrete number of facial expressions or have limited control over the emotion of the output image. To overcome this limitation, we introduced EmoStyle, a method to edit facial expressions based on valence and arousal, two continuous emotional parameters that can specify a broad range of emotions. EmoStyle is designed to separate emotions from other facial characteristics and to edit the face to display a desired emotion. We employ the pre-trained generator from StyleGAN2, taking advantage of its rich latent space. We also proposed an adapted inversion method to be able to apply our system on real images in a one-shot manner. The qualitative and quantitative evaluations show that our approach has the capability to synthesize a wide range of expressions to output high-resolution images.1
Bita Azari, Angelica Lim
WACV2
2023 Contextual Emotion Estimation from Image Captions
abstract
Emotion estimation in images is a challenging task, typically using computer vision methods to directly estimate people’s emotions using face, body pose and contextual cues. In this paper, we explore whether Large Language Models (LLMs) can support the contextual emotion estimation task, by first captioning images, then using an LLM for inference. First, we must understand: how well do LLMs perceive human emotions? And which parts of the information enable them to determine emotions? One initial challenge is to construct a caption that describes a person within a scene with information relevant for emotion perception. Towards this goal, we propose a set of natural language descriptors for faces, bodies, interactions, and environments. We use them to manually generate captions and emotion annotations for a subset of 331 images from the EMOTIC dataset. These captions offer an interpretable representation for emotion estimation, towards understanding how elements of a scene affect emotion perception in LLMs and beyond. Secondly, we test the capability of a large language model to infer an emotion from the resulting image captions. We find that GPT3.5, specifically the text-davinci-003 model, provides surprisingly reasonable emotion predictions consistent with human annotations, but accuracy can depend on the emotion concept. Overall, the results suggest promise in the image captioning and LLM approach.
Vera Yang, Archita Srivastava, Yasaman Etesam, Chuxuan Zhang, Angelica Lim
ACII5
2023 An MCTS-DRL Based Obstacle and Occlusion Avoidance Methodology in Robotic Follow-Ahead Applications
abstract
We propose a novel methodology for robotic follow-ahead applications that address the critical challenge of obstacle and occlusion avoidance. Our approach effectively navigates the robot while ensuring avoidance of collisions and occlusions caused by surrounding objects. To achieve this, we developed a high-level decision-making algorithm that generates short-term navigational goals for the mobile robot. Monte Carlo Tree Search is integrated with a Deep Reinforcement Learning method to enhance the performance of the decision-making process and generate more reliable navigational goals. Through extensive experimentation and analysis, we demonstrate the effectiveness and superiority of our proposed approach in comparison to the existing follow-ahead human-following robotic methods. Our code is available at https://github.com/saharLeisiazar/follow-ahead-ros.
Sahar Leisiazar, Edward J. Park, Angelica Lim, Mo Chen 0001
IROS3
2023 Read the Room: Adapting a Robot's Voice to Ambient and Social Contexts
abstract
How should a robot speak in a formal, quiet and dark, or a bright, lively and noisy environment? By designing robots to speak in a more social and ambient-appropriate manner we can improve perceived awareness and intelligence for these agents. We describe a process and results toward selecting robot voice styles for perceived social appropriateness and ambiance awareness. Understanding how humans adapt their voices in different acoustic settings can be challenging due to difficulties in voice capture in the wild. Our approach includes 3 steps: (a) Collecting and validating voice data interactions in virtual Zoom ambiances, (b) Exploration and clustering human vocal utterances to identify primary voice styles, and (c) Testing robot voice styles in recreated ambiances using projections, lighting and sound. We focus on food service scenarios as a proof-of-concept setting. We provide results using the Pepper robot's voice with different styles, towards robots that speak in a contextually appropriate and adaptive manner. Our results with N=120 participants provide evidence that the choice of voice style in different ambiances impacted a robot's perceived intelligence in several factors including: social appropriateness, comfort, awareness, human-likeness and competency.
Paige Tuttosi, Emma Hughson, Akihiro Matsufuji, Chuxuan Zhang, Angelica Lim
IROS5
2022 Inclusive HRI: Equity and Diversity in Design, Application, Methods, and Community
abstract
Discrimination and bias are pressing issues of many AI and robotics applications. These outcomes may derive from limited datasets that do not fully represent society as a whole or from the AI scientific community's western-male configuration bias. Although being a pressing issue, understanding how robotic systems can replicate and amplify inequalities and injustice among underrepresented communities is still in its infancy among social science and technical communities. This workshop contributes to filling this gap by exploring the research question: What do diversity and inclusion mean in the context of Human-Robot Interaction (HRI)? Here, attention is directed to three different levels of HRI: the technical, the community, and the target user level. Overall, this workshop will focus on the idea that AI systems can be created to be more attuned to inclusive societal needs, respect fundamental rights, and represent contemporary values in modern societies by integrating diversity and inclusion considerations.
Maartje M. A. de Graaf, Giulia Perugia, Eduard Fosch-Villaronga, Angelica Lim, Frank Broz, Elaine Short, Mark A. Neerincx
HRI4
2022 Human Navigational Intent Inference with Probabilistic and Optimal Approaches
abstract
Although human navigational intent inference has been studied in the literature, none have adequately considered both the dynamics that describe human motion and internal human parameters that may affect human navigational behaviour. In this paper, we propose a general probabilistic framework to infer the probability distribution over future navigational states of a human. Our framework incorporates an extended Dubins car dynamics to model human movement, which captures differences in human navigational behaviour depending on their position, heading, and movement speed. We assume a noisily rational model of human behaviour that incorporates a) human navigational intent that may change over time, b) how optimal a person's actions are given the navigational intent, and c) how far ahead in time a person considers when choosing navigational actions. These parameters are recursively and continuously updated in a Bayesian fashion. To make the Bayesian update and inference tractable, we exploit properties of the time-to-reach value function from optimal control and the extended Dubins car dynamics to construct a utility function on which the human policy is based, and employ particle representations of probability distributions where necessary. We demonstrate the effectiveness of our method by comparing our results with a recent approach using synthetic data and validate it on real world data.
Pedram Agand, Mahdi Taherahmadi, Angelica Lim, Mo Chen 0001
ICRA3
2022 Towards Inclusive HRI: Using Sim2Real to Address Underrepresentation in Emotion Expression Recognition
abstract
Robots and artificial agents that interact with humans should be able to do so without bias and inequity, but facial perception systems have notoriously been found to work more poorly for certain groups of people than others. In our work, we aim to build a system that can perceive humans in a more transparent and inclusive manner. Specifically, we focus on dynamic expressions on the human face, which are difficult to collect for a broad set of people due to privacy concerns and the fact that faces are inherently identifiable. Furthermore, datasets collected from the Internet are not necessarily representative of the general population. We address this problem by offering a Sim2Real approach in which we use a suite of 3D simulated human models that enables us to create an auditable synthetic dataset covering 1) underrepresented facial expressions, outside of the six basic emotions, such as confusion; 2) ethnic or gender minority groups; and 3) a wide range of viewing angles that a robot may encounter a human in the real world. By augmenting a small dynamic emotional expression dataset containing 123 samples with a synthetic dataset containing 4536 samples, we achieved an improvement in accuracy of 15% on our own dataset and 11 % on an external benchmark dataset, compared to the performance of the same model architecture without synthetic training data. We also show that this additional step improves accuracy specifically for racial minorities when the architecture's feature extraction weights are trained from scratch.
Saba Akhyani, Mehryar Abbasi Boroujeni, Mo Chen 0001, Angelica Lim
IROS4
2022 Gesture2Vec: Clustering Gestures using Representation Learning Methods for Co-speech Gesture Generation
abstract
Co-speech gestures are a principal component in conveying messages and enhancing interaction experiences between humans and critical ingredients in human-agent interaction, including virtual agents and robots. Existing machine learning approaches have yielded only marginal success in learning speech-to-motion at the frame level. Current methods generate repetitive gesture sequences that lack appropriateness with respect to the speech context. To tackle this challenge, we take inspiration from successes in natural language processing on context and long-term dependencies, and propose a new framework that views text-to-gesture as machine translation, where gestures are words in another (non-verbal) language. We propose a vector-quantized variational autoencoder structure as well as training techniques to learn a rigorous representation of gesture sequences. We then translate input text into a discrete sequence of associated gesture chunks in the learned gesture space. Ultimately, we use translated gesture tokens from the input text as an input to the autoencoder's decoder to produce gesture sequences. Subjective and objective evaluations confirm the success of our approach in terms of appropriateness, human-likeness, and diversity. We also introduce new objective metrics using the quantized gesture representation.
Payam Jome Yazdian, Mo Chen 0001, Angelica Lim
IROS3
2021 Children, Robots, and Virtual Agents: Present and Future Challenges
abstract
Research on child-agent interaction is rapidly expanding. It is, therefore, necessary to converge our collective efforts to broaden our understanding and perspectives of how virtual agents, affect and potentially improve the well-being of children. “Children, Robots and Virtual Agents: Present and Future Challenges” follows our International Conference on Social Robotics (ICSR) 2020 workshop on child-robot interactions. In this full-day workshop, we will focus on the unique technical and empirical challenges of designing and conducting child-agent interactions. In light of the current pandemic situation, we will also address the challenges and adaptations of conducting research under the “new normal” to understand how researchers overcome these challenges and what we can learn and keep in the future. We also aim to join the virtual agents and robotics communities to learn from each other and discuss both areas’ common and specific challenges. Our primary goal is to provide an opportunity for an interdisciplinary debate about the present and future of child-agent interactions. We want to bring together researchers, practitioners and pioneers from relevant disciplines and create collaboration opportunities. As part of the workshop, we will have a collaborative activity where our participants will work together and brainstorm about intelligent agents in different time frames (past, present and future). We will also have a panel of experts discussing the topics of this workshop and answering participants questions.
Elmira Yadollahi, Shruti Chandra, Marta Couto, Angelica Lim, Anara Sandygulova
IDC4
2021 The Many Faces of Anger: A Multicultural Video Dataset of Negative Emotions in the Wild (MFA-Wild)
abstract
The portrayal of negative emotions such as anger can vary widely between cultures and contexts, depending on the acceptability of expressing full-blown emotions rather than suppression to maintain harmony. The majority of emotional datasets collect data under the broad label “anger”, but social signals can range from annoyed, contemptuous, angry, furious, hateful, and more. In this work, we curated the first in-the-wild multicultural video dataset of emotions, and deeply explored anger-related emotional expressions by asking culture-fluent annotators to label the videos with 6 labels and 13 emojis in a multi-label framework. We provide a baseline multi-label classifier on our dataset, and show how emojis can be effectively used as a language-agnostic tool for annotation.
Roya Javadi, Angelica Lim
FG2
2021 A Multimodal and Hybrid Framework for Human Navigational Intent Inference
abstract
Understanding human navigational intent is essential for robots to be able to interact with and navigate around humans safely and naturally. Current methods typically perform inference through only one mode of perception such as human motion trajectory, and a single theoretical framework such as a learning-based or classical approach. In contrast, this paper studies prediction of human navigational intent using multimodal perception within a hybrid framework. Our framework consists of two modules: a) a learning-based prediction module to predict a human’s future goal position, and b) a classical control theory-inspired reconstruction module to reconstruct a possible future trajectory or a set of possible future positions using the predicted future goal position. For the prediction module, we propose an end-to-end LSTM-CNN hybrid neural network for predicting a human’s future position in the real world, given human motion, human body pose and head orientation. This visual information from an egocentric perspective is used to make predictions of a human’s future position in world space, essential for robotic navigation algorithms and planning. In the reconstruction module, we present two control theoretic methods to reconstruct possible future trajectories of human: trajectory generation for differentially flat system and reachability analysis. We evaluate the performance of our framework on a newly collected dataset called SFU-Store-Nav. Experimental results reveal that our method outperforms various baselines especially when a relatively small amount of data is available.
Zhitian Zhang, Jimin Rhim, Angelica Lim, Mo Chen 0001
IROS3
2019 The OMG-Empathy Dataset: Evaluating the Impact of Affective Behavior in Storytelling
abstract
Processing human affective behavior is important for developing intelligent agents that interact with humans in complex interaction scenarios. A large number of current approaches that address this problem focus on classifying emotion expressions by grouping them into known categories. Such strategies neglect, among other aspects, the impact of the affective responses from an individual on their interaction partner thus ignoring how people empathize towards each other. This is also reflected in the datasets used to train models for affective processing tasks. Most of the recent datasets, in particular, the ones which capture natural interactions (“in-the-wild” datasets), are designed, collected, and annotated based on the recognition of displayed affective reactions, ignoring how these displayed or expressed emotions are perceived. In this paper, we propose a novel dataset composed of dyadic interactions designed, collected and annotated with a focus on measuring the affective impact that eight different stories have on the listener. Each video of the dataset contains around 5 minutes of interaction where a speaker tells a story to a listener. After each interaction, the listener annotated, using a valence scale, how the story impacted their affective state, reflecting how they empathized with the speaker as well as the story. We also propose different evaluation protocols and a baseline that encourages participation in the advancement of the field of artificial empathy and emotion contagion.
Pablo V. A. Barros, Nikhil Churamani, Angelica Lim, Stefan Wermter
ACII3
2019 Generating robotic emotional body language with variational autoencoders
abstract
Humanoid robots in social environments can become more engaging by using their embodiment to display emotional body language. For such expressions to be effective in long term interaction, they need to be characterized by variation and complexity, so that the robot can sustain the user's interest beyond the novelty effect period. Hand-coded, pose-to-pose robotic animations can be of high quality and interpretability, but the demanding process of creating them results in limited sets; therefore, after a while, the user will realize that the behavior is repetitive. This work proposes the application of deep learning methods, and more specifically the variational autoencoder framework, for generating numerous emotional body language animations for the Pepper robot, after being trained with a few examples of hand-coded animations. Interestingly, the latent space of the model exhibits topological features that can be used to modulate the amplitude of the motion; we propose that this can be potentially useful for generating animations of specific arousal according to the dimensional theory of emotion.
Mina Marmpena, Angelica Lim, Torbjørn S. Dahl, Nikolas Hemion
ACII2
2019 Investigating Positive Psychology Principles in Affective Robotics
abstract
Positive emotions play a fundamental role in promoting well-being, social bonding, and encouraging people to flourish. We investigated an affective robot's potential to promote positive moods in human groups during a collaborative task. A between-subject experiment (N=39 teams, 78 participants) was conducted to compare the improvement in participants' mood and robot's impression after conducting a collaborative task with an affective robot showing either positive or neutral behaviors. We found that self-reported valence and arousal increased in human participants when interacting with the affective robot regardless of the robot's perceived mood. Additionally, we discovered that participants' likeability of the robot increased when interacting with a positive robot, while likeability decreased when interacting with a neutral robot. These results suggest that even if the evidence for emotional contagion between an affective robot and human group members is not conclusive, participants felt more positive after interacting with robots showing affective behaviors in human-robot team interactions.
Jimin Rhim, Anthony Cheung, David Pham, Subin Bae, Zhitian Zhang, Trista Townsend, Angelica Lim
ACII7
2019 Towards an EmoCog Model for Multimodal Empathy Prediction
abstract
This paper describes a newly proposed empathy prediction model, the EmoCog model, as our solution for the One-Minute Gradual (OMG) Empathy Challenge. The objective for the challenge was to estimate the valence (positivity/negativity) of a listener in a story-telling conversation. We implemented the EmoCog model with two approaches - one with support vector machines (SVMs) and one with neural networks (NNs). We extracted a total of six features corresponding to three categories: 1) cognitive empathy, 2) emotional empathy, 3) synchrony and used them as input to our models. On the validation set, we achieved 0.19 Concordance Correlation Coefficient (CCC) for the SVM approach and 0.25 for the NN approach. On the test set, we achieved results better than baseline, with CCC scores of 0.08 and 0.07, respectively.
Bita Azari, Zhitian Zhang, Angelica Lim
FG3
2017 UE-HRI: a new dataset for the study of user engagement in spontaneous human-robot interactions
abstract
In this paper, we present a new dataset of spontaneous interactions between a robot and humans, of which 54 interactions (between 4 and 15-minute duration each) are freely available for download and use. Participants were recorded while holding spontaneous conversations with the robot Pepper. The conversations started automatically when the robot detected the presence of a participant and kept the recording if he/she accepted the agreement (i.e. to be recorded). Pepper was in a public space where the participants were free to start and end the interaction when they wished. The dataset provides rich streams of data that could be used by research and development groups in a variety of areas.
Atef Ben Youssef, Chloé Clavel, Slim Essid, Miriam Bilac, Marine Chamoux, Angelica Lim
ICMI6
2016 International workshop on social learning and multimodal interaction for designing artificial agents (workshop summary)
abstract
The “social learning and multimodal interaction for designing artificial agents” workshop aims at presenting scientific and philosophical advances related to social learning and multimodal interaction for enhancing the design of artificial agents. Papers presented in the workshop include studies on human behavior modeling, on social robotics and on virtual agents. Our two invited speakers, Prof. Catherine Pelachaud and Prof. Louis-Philippe Morency will enrich and open the door to further discussion by bringing their widely acknowledged expertise in the field.
Mohamed Chetouani, Salvatore Maria Anzalone, Giovanna Varni, Isabelle Hupont, Ginevra Castellano, Angelica Lim, Gentiane Venture
ICMI6
2014 Making a robot dance to diverse musical genre in noisy environments
abstract
In this paper we address the problem of musical genre recognition for a dancing robot with embedded microphones capable of distinguishing the genre of a musical piece while moving in a real-world scenario. For this purpose, we assess and compare two state-of-the-art musical genre recognition systems, based on Support Vector Machines and Markov Models, in the context of different real-world acoustic environments. In addition, we compare different preprocessing robot audition variants (single channel and separated signal from multiple channels) and test different acoustic models, learned a priori, to tackle multiple noise conditions of increasing complexity in the presence of noises of different natures (e.g., robot motion, speech). The results with six different musical genres suggest improved results, in the order of 43.6pp for the most complex conditions, when recurring to Sound Source Separation and acoustic models trained in similar conditions to the testing scenarios. A robot dance demonstration session confirms the applicability of the proposed integration for genre-adaptive dancing robots in real-world noisy environments.
João Lobato Oliveira, Keisuke Nakamura, Thibault Langlois, Fabien Gouyon, Kazuhiro Nakadai, Angelica Lim, Luís Paulo Reis, Hiroshi G. Okuno
IROS6
2010 Robot musical accompaniment: integrating audio and visual cues for real-time synchronization with a human flutist
abstract
Musicians often have the following problem: they have a music score that requires 2 or more players, but they have no one with whom to practice. So far, score-playing music robots exist, but they lack adaptive abilities to synchronize with fellow players' tempo variations. In other words, if the human speeds up their play, the robot should also increase its speed. However, computer accompaniment systems allow exactly this kind of adaptive ability. We present a first step towards giving these accompaniment abilities to a music robot. We introduce a new paradigm of beat tracking using 2 types of sensory input - visual and audio - using our own visual cue recognition system and state-of-the-art acoustic onset detection techniques. Preliminary experiments suggest that by coupling these two modalities, a robot accompanist can start and stop a performance in synchrony with a flutist, and detect tempo changes within half a second.
Angelica Lim, Takeshi Mizumoto, Louis-Kenzo Cahier, Takuma Otsuka, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS1