EDBT 2026 Demo / reviewers in the wild / expert
Randy Gomez
dblp:44/2122
· DBLP profile ↗
89ranked-venue papers
28as first author
48since 2021 · last 2026
0000-0002-3191-6818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 23 first-author · 38 since 2021Human-computer interaction and ubiquitous computing · 43 · 6 first-author · 29 since 2021Systems, architecture and hardware · 25 · 8 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 13 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Pedagogical Relationships in the Presence of Social RobotsabstractAs social robots enter educational settings, how they are positioned as new social actors within existing teacher-student pedagogical relationships becomes an important yet underexplored question. They may be shaped by existing relationships, while also having the potential to reconfigure them. In this study, we explore the relational positioning of educational social robots through a drawing-based design activity with teachers. By analyzing the pedagogical configurations imagined by teachers, we reveal how teachers assign roles, responsibilities, and degrees of authority to robots, in relation to themselves and students. Building on these findings, we discuss the design implications for future educational social robots, highlighting the importance of considering how robots are added into and reshape pedagogical relationships. Zhennan Yi, Paulina Zguda, Randy Gomez, Selma Sabanovic |
DIS | 4 |
| 2026 | Adding More Value Than Work: Practical Guidelines for Integrating Robots into Intercultural Competence LearningabstractWhile social robots have demonstrated effectiveness in supporting students' intercultural competence development, it is unclear how they can effectively be adopted for integrated use in K-12 schools. We conducted two phases of design workshops with teachers, where they co-designed robot-mediated intercultural activities while considering student needs and school integration concerns. Using thematic analysis, we identify appropriate scenarios and roles for classroom robots, explore how robots could complement rather than replace teachers, and consider how to address ethical and compliance considerations. Our findings provide practical design guidelines for the HRI community to develop social robots that can effectively support intercultural education in K-12 schools. Zhennan Yi, Sophia Sakakibara Capello, Randy Gomez, Selma Sabanovic |
HRI | 3 |
| 2026 | Family Privacy Perceptions of Robots in the Home: A Multi-Site Convergence across USA, Spain, and UkraineabstractAs we introduce social robots into family homes, there is a heightened need to understand both parental and child perceptions of privacy, taking into account family perspectives. Haru4Kids (H4K) is an app-based robot simulator designed to cohabitate with children in their home. In a multi-site investigation, H4K cohabitated with 11 American, 11 Spanish, and 10 Ukrainian families. Interviews preceding and following the cohabitation period with at least one parent and all participating children ages 5–12 allowed us to gauge general expectations and experiences during cohabitation and to estimate their comfort with sharing different kinds of information with the platform. In general, children perceived Haru as a technological device that was social, yet non-human. All stakeholders were more comfortable with general information collection as opposed to information sharing with third parties. Lastly, though culture is known to affect expectations and perceptions of privacy and robots, our study shows cohabitating with H4K led to a general convergence across sites in users’ perceptions of Haru and attitudes towards preserving their privacy with robots. Leigh Levinson, Gonzalo A. García, Levko Ivanchuk, Guillermo Pérez 0001, Gloria Alvarez-Benito, J. Gabriel Amores, Manuel Castro-Malet, Randy Gomez, Selma Sabanovic |
ACM Trans. Hum. Robot Interact. | 8 |
| 2026 | Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation
Enqi Zhao, Zicheng Sun, Jianwu Fang, Eric Nichols, Randy Gomez, Bo He 0002, Jianru Xue, Guangliang Li |
IEEE Trans. Robotics | 8 |
| 2025 | Detecting lapses of attention while reading using EEG signalsabstractAttentional lapses while individuals are engaged in activities can have critical effects on their performance. EEG sensors offer the potential to monitor brain activity and detect such lapses in-situ. However, advances in automatic detection are limited, notably due to the scarcity of validated EEG data for training. In this work, we explore the design space of lapses-of-attention detectors using EEG signals, framing it as a binary classification problem. We introduce an EEG dataset with two validated attention levels acquired through a controlled experiment (N = 24) involving reading tasks with and without auditory distractions. We evaluated fifteen detectors using three different EEG feature extraction techniques, and five classifier models. Models using filterbank-CSP features yielded the highest median per-participant detection accuracy of 96%. Limited-resource analyses further indicate the Beta frequency band is the most informative for attention detection, and highlight detection can be achieved with only four EEG channels. Taken together, our findings inform on the feasibility and design of automatic attention detection for brain-computer interfaces utilizing simpler EEG devices. Eranga De Saa, Denise Alonso-Vázquez, Charles-Olivier Dufresne Camaro, Yumiko Sakamoto, Javier Mauricio Antelis, Randy Gomez, Pourang Irani |
Graphics Interface | 6 |
| 2025 | Empathetic Robots Using Empathy Classifiers in HRI SettingsabstractEmpathy is a vital part of human social interaction. It mediates emotional interaction and allows for increased rapport between individuals. We explore combining multi-modal empathy classifiers and empathetic text generation in a human-robot interaction setting. In particular, we designed a demo that uses the Haru social robot to engage in empathetic conversations with a human. We use our classifier to assess the empathy level of the user and the robot. The user's score is a metric for their experience with the robot. On the other hand, we use the robot's score to ensure highly empathetic behaviors. Our preliminary results with two participants show that our system can generate empathetic responses that we can adapt according to the user's empathy level. Christian Arzate Cruz, Edwin C. Montiel-Vázquez, Chikara Maeda, Darryl Lam, Randy Gomez |
HRI | 5 |
| 2025 | Developing Robots for SocietyabstractThe evolution of robotics has taken us from service-oriented and security applications to social robots capable of building relationships with humans and caring for individuals' well-being. Robots are now engaged in a growing range of social and psychological contexts. Furthermore, technologies like social networks have demonstrated their ability to transform human connections by creating platforms that enable global access to information, the sharing of experiences, and active participation in discussions. As societies grow more interconnected, it is high time to develop “Robots for Society”. The Honda Research Institute envisions robots for society, referred to as embodied mediators, proactively mediating relationships among individuals and groups to create a positive societal impact. These embodied mediators should foster understanding across cultural and demographic boundaries, helping to build communities where diverse groups interact under shared values and mutual respect for differences. This talk will explore various research initiatives at Honda Research Institute, conducted in collaboration with consortium partners, focusing on the core principles that define the roles, operations and boundaries of embodied mediators. Most importantly, the discussion will highlight the development of meaningful and tangible applications that extend these principles beyond theory into practice. Randy Gomez |
HRI | 1 |
| 2025 | Social Robot Haru Assisting Dynamic Group Discussion with Autonomous Eye Gaze BehaviorabstractDue to recent advances in large language models and robotics, social robots will potentially play an important role in people’s daily lives soon, and are expected to improve dynamic multi-party group discussions in social scenarios. In this paper, we developed a system to assist dynamic group discussion with our social robot Haru. Our system is composed of three modules: a Dialogue Assistance module via integrating Haru with large language models which facilitates Haru to be an embodied chatbot; a Balancing and Welcoming Behavior module to improve users’ engagement and welcome new users to join the discussion with verbal behaviors; an Autonomous Eye Gazing module to show politeness during group discussion, e.g., gazing to the talking user or the less-engaging user to encourage her, looking to the new comer when she joins the discussion, gazing via eyeball movement when the current speaking user is close to the previous one. The autonomous eye gazing behavior was first trained via deep reinforcement learning in simulation and transferred to physical Haru in the real world. Results of our user study with 50 subjects show the significant performance of our system in assisting dynamic group discussion. Mingyang Hu, Yu Fang 0007, Hongqi Yu, Eric Nichols, Randy Gomez, Guangliang Li |
IROS | 6 |
| 2025 | I Love Lemurs! What's Your Favorite Animal? : Generating Personality-Driven Conversations for the Tabletop Robot HaruabstractThe use of social robots is rapidly expanding across various domains, including education and healthcare. To achieve human-like interactions, these robots should possess well-rounded personalities. A carefully designed personality enhances a social robot’s persuasiveness and increases its appeal during interactions with humans. This paper explores how personality traits can be effectively leveraged to generate conversational responses for social robots, making human-robot interactions more engaging. We introduce a knowledge base that profiles the multifaceted dimensions of the social robot Haru. To generate contextually appropriate responses, we employ a retrieval-augmented generation (RAG) approach to retrieve relevant personality traits. Additionally, we propose a method that integrates result filtering and prompt engineering to ensure consistency in Haru’s responses. To evaluate the effectiveness of our approach, we conduct a preliminary annotation survey assessing the retrieved personality traits and generated responses. The results demonstrate that our method improves conversational flow and enhances response faithfulness to retrieved personality traits. A demo of our approach can be seen at this URL: https://www.youtube.com/watch?v=5wCQDBeSkG8. Paul Reisert, Eric Nichols, Chikara Maeda, Darryl Lam, Sarah Rose Siskind, Randy Gomez |
RO-MAN | 8 |
| 2025 | When and How to Express Empathy in Human-Robot Interaction ScenariosabstractIncorporating empathetic behavior into robots can improve their social effectiveness and interaction quality. In this paper, we present whEE (when and how to express empathy), a framework that enables social robots to detect when empathy is needed and generate appropriate responses. Using large language models, whEE identifies key behavioral empathy cues in human interactions. We evaluate it in human-robot interaction scenarios with our social robot, Haru. Results show that whEE effectively identifies and responds to empathy cues, providing valuable insights for designing social robots capable of adaptively modulating their empathy levels across various interaction contexts. Christian Arzate Cruz, Edwin C. Montiel-Vázquez, Chikara Maeda, Randy Gomez |
RO-MAN | 4 |
| 2025 | Assessment of Cancer Patients' Well-Being through Electrodermal ActivityabstractHospitalized pediatric cancer patients often experience anxiety. Social robots have been proposed as intelligent monitors, embodied mediators, and embodied companions to watch over and support children with such distress. To do so, we propose to augment social robots with biosensors. ElectroDermal Activity recordings of ±22-hour were collected from 8 in-hospital pediatric cancer patients and 6 survivors outside the hospital, together with their diaries and an anxiety questionnaire. To optimally exploit the limited data gathered, external datasets were used to build classification models, with the best performing model achieving a cross-validated F1-score of 0.59 (SD=0.12) on the test set. The models delivered promising out-of-distribution predictions of high-arousal on the pediatric recordings. The limited number of labels available did not allow validation of all high-arousal segments, confirming the challenge of gathering reliable ground truth labels in pediatric hospitals. We suggest data collection methods to foster the further development of augmented social robots. Anneloes L. Meijer, Julie Pivin-Bachler, Gloria Alvarez-Benito, J. Gabriel Amores, Randy Gomez, Egon L. van den Broek |
RO-MAN | 5 |
| 2025 | Building Friendships Across Borders: The Role of Social Robot Haru in Children Group Communication and Connection DevelopmentabstractForming friendship with peers from diverse backgrounds is key to children’s social emotional development. In this study, we explored the use of social robot, Haru, as mediator for remote communication in children group, to support connection and friendship building. We invited children from different countries aged from 10 to 15 to participate in two interaction sessions with peers from other countries, after which we conducted interviews with children from three countries, focusing on their experiences, and perceptions of the robot’s roles in the process. The findings indicated that social robot Haru effectively served as an icebreaker and entertainer; However, improvements are needed in conversation flow, transitions between different roles, and supporting children’s autonomy in guiding the conversation and the depth of their communication. Zhennan Yi, Leigh Levinson, Diego Delgado-Chaves, Jose M. Perez-Moleron, Nabil Bougria, Antonia Krummheuer, Matthias Rehm, Anders Kalsgaard Møller, Katrine Kielsholm Ramsgaard, Selma Auala, Heike Winschiers-Theophilus, Edward Nepolo, David Calero, Devis Dal Moro, Daniel Serrano, Magí Dalmau-Moreno, Randy Gomez, Luis Merino, Selma Sabanovic |
RO-MAN | 17 |
| 2025 | Haru in the Care Network: Stakeholder Perspectives on Privacy with Social Robots in PediatricsabstractSocial robots are beginning to be utilized as part of the collective networks supporting pediatric treatment, however there are few studies on children's perceptions of these agents in hospitals from a privacy and safety perspective. Through a mixed-method and value-sensitive design approach, we introduced hypothetical vignettes and engaged in discussion with 15 youth who are either receiving cancer treatments or are in remission (ages 6-25), 11 of their parents, and 5 out of 8 of their clinical staff to learn how stakeholders in pediatric oncology discuss privacy concerns regarding child-robot interactions. Our thematic analysis imparts how stakeholders perceive robots as social, non-authoritative extensions of the hospital's care network. From this privacy-sensitive perspective, this study revealed that for maximizing a robot's social utility within care systems while critically engaging with the comfort and privacy preferences of stakeholders, robots should take on the role of 1) mediators of social interaction among various stakeholders, 2) companions for children and 3) informational tools for clinicians when consent is given by the family. We emphasize how assistive technologies in pediatrics should continue to be co-designed within communities for identifying appropriate roles and returning agency to stakeholders as they navigate the blurry boundaries of privacy in healthcare. Leigh Levinson, Gloria Alvarez-Benito, J. Gabriel Amores, Deborah Szapiro, Randy Gomez, Selma Sabanovic |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Snitches Get Unplugged: Adolescents' Privacy Concerns about Robots in the Home are Relationally SituatedabstractThough teens are a population with growing agency and use of smart technologies, their concerns surrounding privacy with AI and robots are under-represented in research. Using focus group discussions and a mixed methods analysis, we present teens' comfort levels with robotic information collection and sharing during three hypothetical scenarios involving a child interacting with the Haru social robot in the home. We find participant concerns align with an access-based definition of privacy which prioritizes being in control of their information and of when the robot behaves autonomously. Responses also indicate that teens conceptualize Haru not just as an intelligent device, but also as a social entity. Their shifts in comfort and discussions reflect an engagement in social relationship management with robots in the home in cases where the robot mediates a user's responsibilities and relationships with others. Leigh Levinson, Christena Nippert-Eng, Randy Gomez, Selma Sabanovic |
HRI | 3 |
| 2024 | Design of Embodied Mediator Haru for Remote Cross Cultural CommunicationabstractSocial robots for children have focused mainly on conventional education domains such as teaching language, science, and math, while applications focusing on the enhancement of cultural competency are quite scarce. In this paper, we present a prototype of a robot-mediation framework for cross-cultural communication. This framework paves the way for a social robot to act as a mediator between groups of schoolchildren from different countries. First, we conducted a participatory design activity by an interdisciplinary team, resulting in the extraction of the design, robot’s roles, and technical requirements. Based on these requirements, we built the robot-mediation system prototype. We conducted a pilot study using the system with groups of high school children in Japan and Australia and our results show the potential of the system to drive children’s interest in communicating, sharing, and discussing cultural themes with their remote peers through the social robot. Randy Gomez, Deborah Szapiro, Sara Cooper, Nabil Bougria, Guillermo Pérez 0001, Eric Nichols, Javier Giménez-Figueroa, Jose M. Perez-Moleron, Matthew Peavy, Daniel Serrano, Luis Merino |
ICRA | 1 |
| 2024 | Assisting Group Discussions Using Desktop Robot HaruabstractSocially assistive robots are potentially to be integrated with human daily lives in the near future, and expected to be able to improve group dynamics when interacting with groups of people in social settings. In this paper, we developed a system with desktop robot Haru to assist group discussions. The system consists of three modules: a dialogue assistance module which facilitates Haru to speak to users and answer questions in a free way; a dialogue balance module to encourage participation of users in the discussion with verbal behaviors; an autonomous gazing behavior module trained via deep reinforcement learning in simulation and deployed on physical Haru in reality, which can show politeness during group discussion, e.g., gazing to the speaking member, looking to the middle when both members are talking or silent, looking at the least spoken person when encouraging her. Results of user study with 40 subjects show the significant effectiveness of our system in assisting group discussion. Chuanxiong Zheng, Hongqi Yu, Lei Zhang 0188, Eric Nichols, Randy Gomez, Guangliang Li |
ICRA | 6 |
| 2024 | Shaping Social Robot to Play Games with Human Demonstrations and Evaluative FeedbackabstractIn this paper, building on recent advances in the fields of gaming AI and social robotics, we present a new approach to facilitate the social robot Haru to imitate game strategies from human players’ demonstrated trajectories and evaluative feedback in a real-time two-player game. Our research shows that Haru is able to learn and imitate human different game strategies from human players in a human time scale. In addition, our results show that human evaluative feedback plays an important role in allowing Haru to obtain a better performance via our method than human player’s demonstrations. Finally, results of our user study indicate that Haru imitating human player’s game strategies via our method is perceived to be more human-like and have better game performance and experience than self-learning from pre-defined reward functions via traditional deep reinforcement learning. Chuanxiong Zheng, Lei Zhang 0188, Hui Wang 0141, Randy Gomez, Eric Nichols, Guangliang Li |
ICRA | 4 |
| 2024 | Autonomous Storytelling for Social Robot with Human-Centered Reinforcement LearningabstractSocial robots are gradually integrating into human’s daily lives. Storytelling by social robots could bring a different experience to users through non-verbal and emotional capabilities compared to text-only one. However, as user needs and preferences over storytelling might change over time during long-term interaction with social robots, it is important for social robots to learn from social interactions with human users in real-time. In this paper, we propose to allow our social robot Haru to learn personalized storytelling styles for different human user’s emotional states via human-centered reinforcement learning using the reward provided and delivered by directly interaction with the user explicitly. Results of our user study show that Haru can learn to adapt its storytelling style for detected human emotional states in a few number of interactions, and was perceived to have a better storytelling performance, experience and impact than a neutral one. Lei Zhang 0188, Chuanxiong Zheng, Hui Wang 0141, Randy Gomez, Eric Nichols, Guangliang Li |
IROS | 4 |
| 2024 | Identifying socio-emotional features with a mediator robotabstractIn this paper, we identify a set of socio-emotional cues and signals that are promoted by a tabletop social mediator robot in the context of a school setting. The robot adopts different roles to enhance such socio-emotional features, consequently aiding their identification by a structured annotation system. Various socio-emotional signals were observed for different robot roles, as well as different cues (gaze, speech). Future work will analyze cultural nuances as it expands the pilot to more schools worldwide. Sara Cooper, Randy Gomez, Deborah Szapiro, Luis Merino |
RO-MAN | 2 |
| 2024 | Data Augmentation for 3DMM-based Arousal-Valence Prediction for HRIabstractHumans use multiple communication channels to interact with each other. For instance, body gestures or facial expressions are commonly used to convey an intent. The use of such non-verbal cues has motivated the development of prediction models. One such approach is predicting arousal and valence (AV) from facial expressions. However, making these models accurate for human-robot interaction (HRI) settings is challenging as it requires handling multiple subjects, challenging conditions, and a wide range of facial expressions. In this paper, we propose a data augmentation (DA) technique to improve the performance of AV predictors using 3D morphable models (3DMM). We then utilize this approach in an HRI setting with a mediator robot and a group of three humans. Our augmentation method creates synthetic sequences for underrepresented values in the AV space of the SEWA dataset, which is the most comprehensive dataset with continuous AV labels. Results show that using our DA method improves the accuracy and robustness of AV prediction in realtime applications. The accuracy of our models on the SEWA dataset is 0.793 for arousal and valence. Christian Arzate Cruz, Yotam Sechayk, Takeo Igarashi, Randy Gomez |
RO-MAN | 4 |
| 2024 | Enhancing Human Perception of Direct Gaze from a Social Robot through Eye-Head CoordinationabstractThe development and integration of robots capable of expressing gaze directionality through eye-head movements are crucial for effective human-robot interaction, especially for those with eye designs on 2D screens. Our proposed mutual eye-head gaze model aligns eye movements with head/body rotation, incorporating an attention engine for estimating the most saliency location, and a retina-fovea engine for precise gaze alignment. Additionally, the eye-head engine controls head movements, enhancing the robot’s ability to perform responsive coordinated eye-head gaze behaviors. This improvement leads to enhanced human subjective perception of direct gaze from the robot, ultimately holding potential for advancing human-robot interaction in social dynamics and human-centered robot development research. Yu Fang 0007, Jose M. Perez-Moleron, Luis Merino, Randy Gomez |
RO-MAN | 4 |
| 2024 | Bow Ties & Colorful Eyes: Centering Youth Designs of Social RobotsabstractUnder UNICEF’s Policy guidance on AI for children, child-centered AI should always ‘ensure inclusion of and for children.’ To spotlight youth visions for robots, we led co-design workshops with children between 5-14 years old. Youth designs were expressive, customized, relatable, and approachable. Based on 54 drawings and descriptions of the social robot Haru, we suggest that future child-centered robots should 1) be expressive across verbal and non-verbal channels of communication, 2) allow for customization to give children more agency when interacting with the robot, 3) adapt to children’s style and hobbies to make them feel seen, and 4) aesthetically keep proportions of robot faces consistent and cartoon-like to make robots more approachable. Leigh Levinson, Randy Gomez, Selma Sabanovic |
RO-MAN | 2 |
| 2023 | Presenting Data with Social Robots: An Exploration into Conveying Data Videos using an Artificial Physical NarratorabstractData Videos (DV) have been used in a diverse set of fields. However, the possibility of utilizing them with social robots for further improving the viewer’s engagement is yet to be examined. While social robots have been used in various presentation-related applications, there is also a lack of design instructions on how to better utilize them. Hence with this early work, we explore the possibility of using social robots as potential DV presenters through; a quantitative analysis of the factors of visible presenters in DVs, and a testing phase of these factors via four group design sessions involving experienced designers. From the DV analysis, we identified 12 unique techniques across four main factors. The observations from the group design sessions show that these findings overlap with the design practices of experienced designers when designing robotic presentations. Anuradha Herath, Samar Sallam, Tanvi Vuradi, Yumiko Sakamoto, Randy Gomez, Pourang Irani |
HAI | 5 |
| 2023 | How Should a Social Robot Deliver Negative Feedback Without Creating Distance Between the Robot and Child Users?abstractResearch suggest negative feedback could guide users’ behaviours effectively in Human-AI interactions. However, providing negative feedback, relative to positive counterparts, can be more challenging in any type of communication. This paper delves into the potential of a social robot in delivering negative feedback for improving the in-class learning experience for children. With child participants (12 and younger), we conducted three co-design studies to investigate their preferred facial expressions of a social robot, Haru, which can identify them being distracted (i.e., undesirable behaviour), and redirect their attention back to their task with the facial expressions. Altogether, results indicated that children do not want to see conventional punishing expressions (e.g., angry faces) as a reaction to their undesirable behaviour. Instead, they preferred pleasant ones (e.g., funny, cute). Further, the importance of using realistic stimuli for studies and the co-design approach, as well as the challenges of interpreting children’s drawing responses, are discussed. Yumiko Sakamoto, Anuradha Herath, Tanvi Vuradi, Samar Sallam, Randy Gomez, Pourang Irani |
HAI | 5 |
| 2023 | GAN-Based Interactive Reinforcement Learning from Demonstration and Human Evaluative FeedbackabstractGenerative adversarial imitation learning (GAIL) — a general model-free imitation learning method, allows robots to directly learn policies from expert trajectories in large environments. However, GAIL shares the limitation of other imitation learning methods that they can seldom surpass the performance of demonstrations. In this paper, to address the limit of GAIL, we propose GAN-based interactive reinforcement learning (GAIRL) from demonstrations and human evaluative feedback, by combining the advantages of GAIL and interactive reinforcement learning. We test GAIRL in six physics-based control tasks, ranging from simple low-dimensional control tasks — Cart Pole, Mountain Car and Lunar Lander, to difficult high-dimensional tasks — Inverted Double Pendulum, Hopper and HalfCheetah. Our results suggest that, the GAIRL agent can generally surpass the performance of demonstrations in both low-dimensional and high-dimensional tasks and get an optimal or close to optimal policy. Jiangshan Hao, Rongshun Juan, Randy Gomez, Keisuke Nakamura, Guangliang Li |
ICRA | 4 |
| 2023 | Sim-to-Real Policy and Reward Transfer with Adaptive Forward Dynamics ModelabstractDeep reinforcement learning has shown promise in learning robust skills for robot control, but typically requires a large amount of samples to achieve good performance. Sim-to-real transfer learning has been developed to solve this problem, but the policy trained in simulation usually has unsatisfactory performance in the real world because simulators inevitably model the dynamics of reality imperfectly. To enable sample-efficient learning in the real world, we proposed progressive policy transfer with adaptive dynamics model (PPTADM). PPTADM assumes the dynamics of simulation and real world do not match but the state space is the same, transfers policy from simulation via progressive neural network (PNN) and further improves the policy with a learned forward dynamics model in reality. In addition, for real-world tasks in which reward functions are difficult or even impossible to define and verify the effectiveness, PPTADM can learn in real world solely from a transferred reward function that is estimated from simulation even though their dynamics do not match. Our results in five simulated tasks and on a real robot arm show that with PPTADM, the robot's learning efficiency and performance in the real world can be significantly improved. Rongshun Juan, Hao Ju 0003, Randy Gomez, Keisuke Nakamura, Guangliang Li |
ICRA | 4 |
| 2023 | Model-based Adversarial Imitation Learning from Demonstrations and Human RewardabstractReinforcement learning (RL) can potentially be applied to real-world robot control in complex and uncertain environments. However, it is difficult or even unpractical to design an efficient reward function for various tasks, especially those large and high-dimensional environments. Generative adversarial imitation learning (GAIL) - a general model-free imitation learning method, allows robots to directly learn policies from expert trajectories in large and high-dimensional environments. However, GAIL is still sample inefficient in terms of environmental interaction. In this paper, to solve this problem, we propose a model-based adversarial imitation learning from demonstrations and human reward (MAILDH), a novel model-based interactive imitation framework combining the advantages of GAIL, interactive RL and model-based RL. We tested our method in eight physics-based discrete and continuous control tasks for RL. Our results show that MAILDH can greatly improve the sample efficiency and robustness compared to the original GAIL. Jiangshan Hao, Rongshun Juan, Randy Gomez, Keisuke Nakarnura, Guangliang Li |
IROS | 4 |
| 2023 | How to Make a Robot Grumpy Teaching Social Robots to Stay in Character with Mood SteeringabstractConveying a robot's target mood is crucial to successful social interactions. The robot's expressive performance must be appropriate, persuasive, and consistent. However, this is challenging when interactions contain a mixture of scripted and improvised content, such as those generated by language models. In this paper, we take on the task of teaching robots to stay in character, that is to say, exhibit consistency in mood during interactions. We start by defining a communication strategy module that allows for the top-down specification of a target robot mood for a given task, goal, or context. We then propose a mood steering framework for enforcing robot mood consistency throughout an interaction that supports several target moods. Our framework consists of two components: 1. expressivity steering specifies the speech and behavior to be used by the robot to convey a target mood, and 2. language model steering ensures that improvised language is consistent with the robot's target mood. As a first step toward identifying effective communication strategies, we implement grumpy and cheerful strategies for a collaborative storytelling game and compare them to a neutral baseline. Evaluation in a collaborative storytelling game shows that our approach generates robot behavior that successfully conveys the robot's target mood throughout gameplay and language model steering generates story contributions that capture the target mood without quality degradation and raises important issues for communication strategy design. Eric Nichols, Deborah Szapiro, Yurii Vasylkiv, Randy Gomez |
IROS | 4 |
| 2023 | Exploring the Design of Social Robot User Interfaces for Presenting Data-Driven StoriesabstractTabletop social robots are becoming increasingly common, not only as social companions but as presenters and orators of information. We present an exploration of utilizing robots as a multimodal presentation tool to communicate data-driven facts. Our exploration is inspired by the wealth of research on data videos (DVs) as these have become mainstream sources for swiftly conveying data-driven information to a mass audience. We first analyze 48 DVs that contain visible narrators (presenters who are visible in the video frames) as our source for understanding the techniques used to convey factual information via presenters. Twelve dimensions across four factors (presenter-grounded; narrative-grounded; viewer-engagement-related; and data-visualization-related) were identified. These factors were carefully arranged in designing presenters to engage the audience with the video content. We adapt these findings to the design of an expressive social tabletop robot that can communicate data-driven knowledge to its audience. Supported by four design sessions with expert content creators and designers, we provide nine design implications for designing multimodal presentations with an expressive tabletop social robot. We conclude with the possible application potentials of this unique data presentation modality. Anuradha Herath, Samar Sallam, Yumiko Sakamoto, Randy Gomez, Pourang Irani |
MUM | 4 |
| 2023 | Designing Visual and Auditory Attention-Driven Movements of a Tabletop RobotabstractThis work presents a framework for a visual-auditory attention-driven robot eye-head gaze movement, which combines visual and auditory inputs to determine the direction of gaze movement for a social robot. The framework computes the most salient changes in position by considering both visual and auditory cues. The proposed system was implemented on Haru, a tabletop social robot, where eye-head gaze movement was controlled using visual input from a camera positioned above the eyes and auditory input from a seven-channel microphone. This allowed for eye movement on a two-dimensional flat screen and body rotation towards the person who is speaking. This framework provides a representation of the robot’s attentional gaze that leverages both visual and auditory cues, resulting in more natural and responsive coordinated eye-head gaze movements of the social robot. The potential benefits include improved communication, increased engagement, and a stronger sense of connection with the robot. Yu Fang 0007, Luis Merino, Serge Thill, Randy Gomez |
RO-MAN | 4 |
| 2023 | Living with Haru4Kids: Study on children's activity and engagement in a family-robot cohabitation scenarioabstractHaru4Kids (H4K) is a system that emulates the physical, social, family-oriented robot Haru, designed with the goal to cohabitate with children in their home for extended periods of time. Seven families kept H4K for two net weeks in their homes. Throughout this period of cohabitation, we collected user logs comprised of the children users ’ head angles, the rotation angles of the platform, and the actions taken by H4K as well as captured images which were afterwards hand-annotated to estimate user engagement. We report the trends of these external metrics that we collected during every session of interaction. We also developed an annotation tool and report the Engagement Level Metric we chose to estimate child engagement throughout interactions “in-the-wild.” Overall, our platform offers a feasible system that can engage with children while also allowing us to monitor their engagement and behaviour throughout each interaction. Gonzalo A. García, Guillermo Pérez 0001, Leigh Levinson, J. Gabriel Amores, Gloria Alvarez-Benito, Manuel Castro-Malet, Mario Castaño-Ocaña, Marta J. López-González de Quevedo, Ricardo Durán-Viñuelas, Randy Gomez, Selma Sabanovic |
RO-MAN | 10 |
| 2023 | Child-Robot Conversation in the Wild Wild Home: A Language Processing User StudyabstractChild-robot interaction (CRI) has been mostly studied in labs and classroom settings. In this work, we share a CRI language processing study carried out in children’s homes. Any automated system deployed “in-the-wild” faces practical problems, but when the target users are children, these problems get even more sensitive and challenging. In this work we analyse how each language processing layer performs with children at home with no researcher present. We carried out an experiment with 7 families [N=14 children, 6-13 years old] cohabiting with a simulated robot for 2 weeks in their own homes. Our goal in this study is to evaluate the performance of voice recognition, language understanding and dialogue management when children interact with a robot at home. Our results indicate that dialogue management capabilities are becoming the key element in the language processing pipeline; they also denote that the dialogue engine should include mixed-initiative capabilities and show the relative usage of different common built-in intents. Guillermo Pérez 0001, Gonzalo A. García, Manuel Castro-Malet, Mario Castaño-Ocaña, Marta J. López-González de Quevedo, Ricardo Durán, J. Gabriel Amores, Gloria Alvarez-Benito, Leigh Levinson, Selma Sabanovic, Randy Gomez |
RO-MAN | 11 |
| 2022 | Making The Unknown More Certain: A Stacked Ensemble Classifier for Open Gesture Recognition with a Social RobotabstractWe introduce a novel stacked ensemble classifier for the unconstrained recognition of known and unknown gestural input data in nonverbal communication with a social robot. The architecture utilizes three separate CNNs of different expected data input size and combines their output predictions to a unified estimate. Analysis shows that in comparison to a single CNN architecture, the combined estimate reduces prediction confidence values for unknown gestural movement segments, making the system able to identify unknown data input with higher certainty under both laboratory and real environment conditions. In a human-robot interaction experiment, we are able to improve unknown class detection accuracy by up to 40% under maintained or equal known class recognition performance, and hence considerably enhance the overall robustness of the recognition system. Heike Brock, Randy Gomez |
ICASSP | 2 |
| 2022 | Developing The Bottom-up Attentional System of A Social RobotabstractThis paper describes the development of a 3- stage signalling framework to trigger a social robot's bottom- up reactive behavior inspired by a biological model. In the first stage, low-level firing of stimuli due to external sources is constructed through perception grounding. This is followed by a saliency classifier which fires-up high level salient signals that require attention and are used to trigger the robot's reactive behavior. The whole framework evolves primarily on the knowledge ontology that defines the characteristics of the social robot and the querying mechanism that correlates the perceived stimuli with the ontology to trigger the reactive behavior. We evaluated the performance of our system with timing metrics and we achieved good results for our application. Randy Gomez, Álvaro Páez, Yu Fang 0007, Serge Thill, Luis Merino, Eric Nichols, Keisuke Nakamura, Heike Brock |
ICRA | 1 |
| 2022 | Hey Haru, Let's Be Friends! Using the Tiers of Friendship to Build Rapport through Small Talk with the Tabletop Robot HaruabstractConversation can play an essential role in forging bonds between humans and social robots, but participants need to feel like they are being listened to, remembered, and cared about in order to effectively build rapport. In this paper, we propose a novel strategy for conducting small talk with a social robot. Our approach is known as the Tiers of Friendship. It is centered around three core design elements: 1) Persuasive content and character is provided through topic modules created by professional creative writers to ensure engaging conversational content and a compelling personality for the social robot. 2) Conversational memory is achieved by allowing topic modules to specify required information that can be learned through conversation or recalled from previous interactions and organizing topic modules into a hierarchy that enforces information requirements between topics. 3) Dynamicity in conversation is promoted through topic navigation that supports fluid transitions to topics of human interest and employs elements of random ordering to create fresh conversation experiences. In this paper, we show how the Tiers of Friendship can be used to generate conversation content for a social robot that encourages the development of rapport. We describe a working implementation of a small talk system for a social robot based on the Tiers of Friendship that combines off-the-shelf ASR and NLU components and custom robot behavior components implemented via behavior trees on ROS. Finally, in order to evaluate our approach's effectiveness, we conduct an elicitation survey that evaluates conversations in terms of perceived engagement, personality traits, and rapport expectation and discuss the implications for social robotics. Eric Nichols, Sarah Rose Siskind, Levko Ivanchuk, Guillermo Pérez 0001, Waki Kamino, Selma Sabanovic, Randy Gomez |
IROS | 7 |
| 2022 | Affective Behavior Learning for Social Robot Haru with Implicit Evaluative FeedbackabstractWe propose a human-in-the-loop reinforcement learning mechanism to help robots learn emotional behavior. Unlike the previous methods of providing explicit feedback via pressing keyboard buttons or mouse clicks, we provide a more natural way for ordinary people to train social robots how to perform social tasks according to their preferences - facial expressions. The whole experiment is carried out on the desktop robot Haru, which is mainly used for the research of emotion and empathy participation. Our experimental results show that through learning from implicit feedback of facial features, Haru can quickly understand and dynamically adapt to individual preferences, and obtain a similar performance to learning from explicit feedback. In addition, we observe that the recognition error of human feedback will cause a “temporary regress” of the robot's learning performance, which is more obvious at the beginning of the training process. This phenomenon is shown to be correlated with the accuracy of recognizing negative implicit feedback. Hui Wang 0141, Jinying Lin, Yurii Vasylkiv, Heike Brock, Keisuke Nakamura, Randy Gomez, Bo He 0002, Guangliang Li |
IROS | 7 |
| 2022 | The LMA12-O Framework for Emotional Robot Eye GesturesabstractThe eyes play a significant role in how robots are perceived socially by humans due to the eye’s centrality in human communication. To date there has been no consistent or reliable system for designing and transferring affective emotional eye gestures to anthropomorphized social robots. Combining research findings from Oculesics, Laban Movement Analysis and the Twelve Principles of Animation, this paper discusses the design and evaluation of the prototype LMA12-O framework for the purpose of maximising the emotive communication potential of eye gestures in anthropomorphized social robots. Results of initial user testings evidenced LMA12-O to be effective in designing affective emotional eye gestures in the test robot with important considerations for future iterations of this framework. Kerl Galindo, Deborah Szapiro, Randy Gomez |
RO-MAN | 3 |
| 2022 | I Can't Believe That Happened! : Exploring Expressivity in Collaborative Storytelling with the Tabletop Robot HaruabstractCollaborative storytelling has long been a goal of social robotics, however, much of this research is limited in interactivity or assumes that story content is curated. In this paper, we present a working fully-automatic collaborative storytelling robot, which can collaborate with a person to create a unique, improvised story by using a large-scale neural language model to dynamically generate continuations to a story. Because effective storytelling requires engaging the emotions of participants, we explore several modalities of procedurally-generated expressivity: 1. an expressive text-to-speech voice with several delivery styles, 2. physical and verbal reactions performed by the robot, and 3. an external display used to show instructions and graphics during storytelling.To understand the issues associated with improvised collaborative storytelling with a social robot, we conduct an online survey and elicitation study with a group of online observers of collaborative storytelling gameplay, comparing several expressivity strategies in terms of storytelling-related characteristics, expressivity characteristics, and personality traits as measured by RoSAS. This evaluation showed that expressivity strategies using both emotive voice and performed reactions were perceived to be more competent storytellers and more strongly associated with positive personality traits. Eric Nichols, Deborah Szapiro, Yurii Vasylkiv, Randy Gomez |
RO-MAN | 4 |
| 2022 | Shaping Haru's Affective Behavior with Valence and Arousal Based Implicit Facial FeedbackabstractSocial robots that are able to express emotions can potentially improve human’s well-being. Whether and how they can learn from interactions between them and human being in a natural way will be key to their success and acceptance by ordinary people. In this paper, we proposed to shape social robot Haru affective behaviors with predicted continuous rewards based on received implicit facial feedback via human-centered reinforcement learning. The implicit facial feedback was estimated with the valence and arousal of received implicit facial feedback using Russell’s circumplex model, which can provide a more accurate estimation of the subtle psychological changes of human user, resulting in more effective robot behavior learning. The whole experiment is conducted on the desktop robot Haru, which is primarily used to study emotional interactions with human in different scenarios. Our experimental results show that with our proposed method, Haru can obtain a similar performance to learning from explicit feedback, eliminating the need for human users to get familiar with training interface in advance and resulting in an unobtrusive learning process. Hui Wang 0141, Randy Gomez, Keisuke Nakamura, Bo He 0002, Guangliang Li |
RO-MAN | 3 |
| 2021 | Exploring the Concept of Fairness in Everyday, Imaginary and Robot Scenarios: A Cross-Cultural Study With Children in Japan and UgandaabstractThis paper describes a cross-cultural pilot study on children’s perceptions of fairness in robot-related scenarios with children in Japan (N = 20) and Uganda (N = 24). We used storytelling to facilitate children’s narratives on fairness and to identify areas of alignment and disconnect. Initial results indicate that while both groups referred to similar aspects of fairness, namely psychological, physical and systemic, children in Tokyo focused more on psychological and mental aspects while children in Uganda emphasised on physical and material aspects. Both groups increased their emphasis on mental aspects in robot-related scenarios. All children expressed their interest to further explore fairness and unfairness as experienced by children with different cultural backgrounds and the need for inter-group contact. The results of this study will contribute to the first phase of a study with robots and children in relation to child’s fundamental rights and to the dialogue about the requirements for fairness in robot development by highlighting the importance of considering children’s perspectives especially those of typically under-represented cultural groups. Vicky Charisi, Tomoko Imai, Tiija Rinta, Joy Maliza Nakhayenze, Randy Gomez |
IDC | 5 |
| 2021 | The Effects of Robot Cognitive Reliability and Social Positioning on Child-Robot Team DynamicsabstractHuman collaboration is more likely to lead to cognitive growth when all group-members are actively involved in the collaborative process. However, there are cases that intragroup relationships need support. In this paper, we present an autonomous robotic system designed to interact with a pair of children in a problem-solving setting, aiming to understand how the robot behaviour impacts the group-members’ social dynamics. We developed an autonomous system with the Haru robot which we evaluated with an experimental study with 5-8yo children (N =84) to test the impact of the robot’s cognitive reliability and social positioning on human-to-human social dynamics, task performance and help-seeking behaviour. All participants took part in a baseline session (without the robot), an intervention (with the robot in a turn-taking setting) and an evaluation session (with a robot in a voluntary interaction setting). Results indicate that children who interacted with the reliable robot had a better task performance but children who interacted with the unreliable robot exhibited more task-related social interactions. Based on the results, we propose an interaction design concept which combines the set of the evaluated robot behaviours for an adaptive targeted support of child-robot teaming. Vicky Charisi, Luis Merino, Marina Escobar, Fernando Caballero, Randy Gomez, Emilia Gómez |
ICRA | 5 |
| 2021 | Automating Behavior Selection for Affective Telepresence RobotabstractThe tabletop robot Haru, used for affective telepresence research, enables a teleoperator to communicate affects from a distance. The robot’s expressiveness offers myriad ways of communicating affects through the execution of emotive routines. The teleoperator reacts to input modalities such as the user’s facial expression, gestures and speech-based intent as perceived by the robot’s perception system. However, due to the sheer number of routines to select from, the task of choosing the appropriate or the most preferred routine is becoming cumbersome. In this paper, we propose a human-in-the-loop reinforcement learning mechanism in which an agent learns the teleoperator’s selection preference as a function of the input modalities and aids the routine selection process by narrowing it to n-best optimal choices. Our experimental results show that with only a few number of interactions from the teleoperator, the system can learn to recommend optimal routine behaviors for all perceived modalities, which greatly reduces the workload of the teleoperator. Yurii Vasylkiv, Guangliang Li, Eleanor Sandry, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez |
ICRA | 8 |
| 2021 | Personalization of Human-Robot Gestural Communication through Voice Interaction GroundingabstractIn this paper we develop a gestural communication perception system for a social robot companion that is able to autonomously learn novel gestures on-the-fly. The system constantly tracks human gestural activities with a camera and evaluates the performed gestures under an open-set assumption. This allows for the identification of unknown gestures. Once detected, the system stores motion sequences of the novel gesture class and employs a dialogue interaction with the human to automatically label the unknown gesture. Subsequently, the gestural model is updated, grounding the unknown gesture through dialog interaction. In our experiment, we evaluate a neural network with varying threshold values for the open gesture recognition with unknown detection. Results show that the general classifier reaches an accuracy of more than 83%, and an f1-score of 0.79 in an open-ended scenario. The method is furthermore tested in a first in-lab interaction setting, which shows the system usability and its potential for future personalized human-robot gestural communication. Heike Brock, Randy Gomez |
IROS | 2 |
| 2021 | Shaping Progressive Net of Reinforcement Learning for Policy Transfer with Human Evaluative FeedbackabstractDeep reinforcement learning has achieved significant success in many fields, but will confront sampling efficiency and safety problems when applying to robot control in the real world. Sim-to-real transfer learning was proposed to make use of samples in the simulation and overcome the gap between simulation and real world. In this paper, we focus on improving Progressive Neural Network — an effective sim-to-real learning method, by proposing Interactive Progressive Network Learning (IPNL). IPNL integrates progressive network and interactive reinforcement learning (interactive RL) which learns from evaluative feedback provided by an observing human trainer. We test our method using five RL tasks with discrete or continuous actions in OpenAI Gym and a sinusoids curve following task with AUV simulator on the Gazebo platform. Our results suggest that while Progressive Network has good performance when transferring from tasks with low-dimensional state space to those with high-dimensional one but has little effect for transferring from high-dimensional tasks to low-dimensional ones, IPNL allows an agent to learn a more stable policy with better performance faster for both cases. More importantly, our further analysis indicate that there is a synergy between Progressive Network and interactive RL for improving the agent’s learning. Our results in the path following of AUV shed light on the potential of applying our method in the real world tasks. Rongshun Juan, Randy Gomez, Keisuke Nakamura, Qixin Sha, Bo He 0002, Guangliang Li |
IROS | 3 |
| 2021 | Collaborative Storytelling with Social RobotsabstractStorytelling plays a central role in human socializing and entertainment, and research on conducting storytelling with robots is gaining interest. However, much of this research assumes that story content is curated. In this paper, we expand the recently-proposed task of collaborative storytelling, where an intelligent agent and a person collaborate to create a unique story by taking turns adding to it, for application to social robot and consider the design implications that arise. Since latency can be detrimental to human-robot interaction, we examine the performance-latency trade-offs of an existing generate-and-rank-based approach to collaborative storytelling by finding the optimal ranker’s sample size that strikes the best balance between quality and computational cost. We improve on existing evaluation that was previously based on system-generated stories by having human participants play the collaborative storytelling game with our system and comparing the stories they create with our system to a naive baseline. Finally, we conduct a pilot elicitation survey that sheds light on issues to consider when adapting our collaborative storytelling system to a social robot. Our evaluation shows that participants have a positive view of collaborative storytelling with a social robot and consider rich, emoting capabilities to be key to enjoyment. Eric Nichols, Leo Gao, Yurii Vasylkiv, Randy Gomez |
IROS | 4 |
| 2021 | Developing an Engagement-Aware System for the Detection of Unfocused InteractionabstractWe introduce a perception system for social robots that is able to detect a person’s engagement in an interaction from nonverbal cues independently of principal user activity. This was achieved by the introduction of a set of proxemics, body posture and attention features relevant for human-human interaction. The features were extracted from RGB-D image data of a single Kinect and utilized to train two separate machine learning models. Multiple system configurations and feature combinations were tested, and their impact on the detection of user engagement evaluated. Combining all features, our perception system reaches an F1-score of 81% when estimating an observed person’s interaction intent through binary classification. Regression of a user’s level of availability deviates from the given ground truth values by 13.27% on average. Finally, a prototype was implemented which is able to simultaneously run both previous estimates in real-time using a shared feature vector. In the following, the proposed system shall be used to design robots whose behavior shows their awareness of user engagement. Marvin Brenner, Heike Brock, Andreas Stiegler, Randy Gomez |
RO-MAN | 4 |
| 2021 | Exploring Affective Storytelling with an Embodied AgentabstractIn this paper, we explore the storytelling potential of a robot. We exploit the use of creative contents that maximize the embodied communication affordance of the empathic robot Haru. We identify the elements in storytelling such as narration, agency, engagement and education and synthesized these into the robot. Through effective design we investigated the possible answers that could leverage the limitations and the challenges in developing storytelling applications through a robotic medium. Our preliminary findings show that the use of an embodied agent such as a robot in storytelling only has meaning when its communicative affordance (i.e. embodiment, expressiveness, and other modalities) is tapped, adding new dimension to the experience. Otherwise, traditional storytelling delivery (e.g. tablet) without the use of embodiment will suffice. Hence, robots need to be performers rather than just mere props in storytelling. Randy Gomez, Deborah Szapiro, Kerl Galindo, Luis Merino, Heike Brock, Keisuke Nakamura, Yu Fang 0007, Eric Nichols |
RO-MAN | 1 |
| 2021 | Shaping Affective Robot Haru's Reactive ResponseabstractWe describe a method of teaching a robot its empathic behavioural response from its interaction with people. We used the input modalities such as relative spatial information, facial expressions, body gestures and speech information as perception input that triggers the robot’s empathic response. First, we bootstrap the training through a pre-learning mechanism in which training is conducted by users who know the robotic system. This phase provides simulation-based training using a simple graphical user interface to simulate the input, rewards and correction feedback. In the second phase, we developed an online learning scheme for naive users to personalize their robot further, building on top of the bootstrapped model. Here, we developed a natural user interface that enables natural human-robot interaction via the suite of sensors that allows the users to provide evaluative feedback during the interaction with the robot. We evaluated the system and our results show that bootstrapping is an efficient tool to hasten the robot’s learning while online learning provided some form of personalization in the real environment with naive users. Yurii Vasylkiv, Guangliang Li, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez |
RO-MAN | 7 |
| 2020 | MoveAE: Modifying Affective Robot Movements Using Classifying Variational AutoencodersabstractWe propose a method for modifying affective robot movements using neural networks. Social robots use gestures and other movements to express their internal states. However, a robot's interactive capabilities are hindered by the predominant use of a limited set of preprogrammed or hand-animated behaviors, which can be repetitive and predictable, making sustained human-robot interactions difficult to maintain. To address this, we developed a method for modifying existing emotive robot movements by using neural networks. We use hand-crafted movement samples and a classifying variational autoencoder trained on these samples. Our method then allows for adjustment of affective movement features by using simple arithmetic in the network's latent embedding space. We present the implementation and evaluation of this approach and show that editing in the latent space can modify the emotive quality of the movements while preserving recognizability and legibility in many cases. This supports neural networks as viable tools for creating and modifying expressive robot behaviors. Michael Suguitan, Randy Gomez, Guy Hoffman |
HRI | 2 |
| 2020 | A Holistic Approach in Designing Tabletop Robot's ExpressivityabstractDefining a robot's expressivity is a difficult task that requires thoughtful consideration of the potential of various robot modalities and a model of communication that humans understand. Humanoid and zoomorphic-designed robots can easily take cues from human and animals, respectively when designing their expressivity. However, a robot design that is neither human nor animal-like does not have a clear model to follow in terms of designing expressivity. Animation presents a potential model in these circumstances as animated characters in movies take various forms, sizes, shapes and styles, and are successful in defining expressivity that is widely accepted across different languages and cultures. In this paper, we discuss the development and design of the expressivity of Haru, a table top robot that is neither human nor animal-like and the application of animation expertise to the holistic treatment of the different modalities. The method maximizes animation techniques and expertise normally applied to movies to generate expressivity that is then transferred to the robot hardware. Experimental results show that the robot's expressivity generated using our method is easily understood and are preferred to the conventional approach of generating expressions. Randy Gomez, Deborah Szapiro, Luis Merino, Keisuke Nakamura |
ICRA | 1 |
| 2020 | Collaborative Storytelling with Large-scale Neural Language ModelsabstractStorytelling plays a central role in human socializing and entertainment. However, much of the research on automatic storytelling generation assumes that stories will be generated by an agent without any human interaction. In this paper, we introduce the task of collaborative storytelling, where an artificial intelligence agent and a person collaborate to create a unique story by taking turns adding to it. We present a collaborative storytelling system which works with a human storyteller to create a story by generating new utterances based on the story so far. We constructed the storytelling system by tuning a publicly-available large scale language model on a dataset of writing prompts and their accompanying fictional works. We identify generating sufficiently human-like utterances to be an important technical issue and propose a sample-and-rank approach to improve utterance quality. Quantitative evaluation shows that our approach outperforms a baseline, and we present qualitative evaluation of our system’s capabilities. Eric Nichols, Leo Gao, Randy Gomez |
MIG | 3 |
| 2020 | Robust Real-Time Hand Gestural Recognition for Non-Verbal Communication with Tabletop Robot HaruabstractIn this paper, we present our work in close-distance non-verbal communication with tabletop robot Haru through hand gestural interaction. We implemented a novel hand gestural understanding system by training a machine-learning architecture for real-time hand gesture recognition with the Leap Motion. The proposed system is activated based on the velocity of a user's palm and index finger movement, and subsequently labels the detected movement segments under an early classification scheme. Our system is able to combine multiple gesture labels for recognition of consecutive gestures without clear movement boundaries. System evaluation is conducted on data simulating real human-robot interaction conditions, taking into account relevant performance variables such as movement style, timing and posture. Our results show robustness in hand gesture classification performance under variant conditions. We furthermore examine system behavior under sequential data input, paving the way towards seamless and natural real-time close-distance hand-gestural communication in the future. Heike Brock, Selma Sabanovic, Keisuke Nakamura, Randy Gomez |
RO-MAN | 4 |
| 2020 | Human Social Feedback for Efficient Interactive Reinforcement Agent LearningabstractAs a branch of reinforcement learning, interactive reinforcement learning mainly studies the interaction process between humans and agents, allowing agents to learn from the intentions of human users and adapt to their preferences. In most of the current studies, human users need to intentionally provide explicit feedback via pressing keyboard buttons or mouse clicks. However, in our paper, we proposed an interactive reinforcement learning method that facilitates an agent to learn from human social signals - facial feedback via a ordinary camera and gestural feedback via a leap motion sensor. Our method provides a natural way for ordinary people to train agents how to perform a task according to their preferences. We tested our method in two reinforcement learning benchmarking domains - LoopMaze and Tetris, and compared to the state of the art - the TAMER framework. Our experimental results show that when learning from facial feedback the recognition of which is very low, the TAMER agent can get a similar performance to that of learning from keypress feedback with slightly more feedback. When learning from gestural feedback with a more accurate recognition, the TAMER agent can obtain a similar performance to that of learning from keypress feedback with much less feedback received. Moreover, our results indicate that the recognition error of facial feedback has a large effect on the agent performance in the beginning training process than in the later training stage. Finally, our results indicate that with enough recognition accuracy, human social signals can effectively improve the learning efficiency of agents with less human feedback. Jinying Lin, Qilei Zhang, Randy Gomez, Keisuke Nakamura, Bo He 0002, Guangliang Li |
RO-MAN | 3 |
| 2019 | Expressivity for Sustained Human-Robot InteractionabstractExpressivity - the use of multiple, non-verbal, modalities to convey or augment the communication of internal states and intentions - is a core component of human social interactions. Studying expressivity in contexts of artificial agents has led to explicit considerations of how robots can leverage these abilities in sustained social interactions. Research on this covers aspects such as animation, robot design, mechanics, as well as cognitive science and developmental psychology. This workshop provides a forum for scientists from diverse disciplines to come together and advance the state of the art in developing expressive robots. Participants will discuss points of methodological opportunities and limitations, to develop a shared vision for next steps in expressive social robots. Vicky Charisi, Selma Sabanovic, Serge Thill, Emilia Gómez, Keisuke Nakamura, Randy Gomez |
HRI | 6 |
| 2019 | Generation of expressive motions for a tabletop robot interpolating from hand-made animationsabstractMotion is an important modality for human-robot interaction. Besides a fundamental component to carry out tasks, through motion a robot can express intentions and expressions as well. In this paper, we focus on a tabletop robot in which motion, among other modalities, is used to convey expressions. The robot incorporates a set of pre-programmed motion animations that show different expressions with various intensities. These have been created by designers with expertise in animation. The objective in the paper is to analyze if these examples can be used as demonstrations, and combined by the robot to generate additional richer expressions. Challenges are the representation space used, and the scarce number of examples. The paper compares three different learning from demonstration approaches for the task at hand. A user study is presented to evaluate the resultant new expressive motions automatically generated by combining previous demonstrations. Gonzalo Mier, Fernando Caballero, Keisuke Nakamura, Luis Merino, Randy Gomez |
RO-MAN | 5 |
| 2019 | Human-Centered Reinforcement Learning: A SurveyabstractHuman-centered reinforcement learning (RL), in which an agent learns how to perform a task from evaluative feedback delivered by a human observer, has become more and more popular in recent years. The advantage of being able to learn from human feedback for a RL agent has led to increasing applicability to real-life problems. This paper describes the state-of-the-art human centered RL algorithms and aims to become a starting point for researchers who are initiating their endeavors in human-centered RL. Moreover, the objective of this paper is to present a comprehensive survey of the recent breakthroughs in this field and provide references to the most interesting and successful works. After starting with an introduction of the concepts of RL from environmental reward, this paper discusses the origins of human-centered RL and its difference from traditional RL. Then we describe different interpretations of human evaluative feedback, which have produced many human-centered RL algorithms in the past decade. In addition, we describe research on agents learning from both human evaluative feedback and environmental rewards as well as on improving the efficiency of human-centered RL. Finally, we conclude with an overview of application areas and a discussion of future work and open questions. Guangliang Li, Randy Gomez, Keisuke Nakamura, Bo He 0002 |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2018 | Comparison of Region of Interest Segmentation Methods for Video-Based Heart Rate MeasurementsabstractConventional contact photoplethysmography (PPG) sensors are not suitable in situations of skin damage or when unconstrained movement is required. As a consequence, remote photoplethysmography (rPPG) has recently emerged because it provides remote physiological measurements without expensive hardware and improves comfort for long term monitoring. RPPG estimation methods use the spatially averaged RGB values of pixels in a Region Of Interest (ROI) to generate a temporal RGB signal. The selection of ROI is a critical first step to obtain reliable pulse signals and must contain as many skin pixels as possible with a low percentage of non-skin pixels. In this paper, we experimentally compare seven ROI segmentation methods in the perspective of heart rate (HR) measurements with dedicated metrics. The algorithms are compared using our in-house database UBFC-RPPG, comprising of 53 videos specifically geared towards rPPG analysis. Peixi Li, Yannick Benezeth, Keisuke Nakamura, Randy Gomez, Chao Li 0005, Fan Yang 0019 |
BIBE | 4 |
| 2018 | PageFlip: Leveraging Page-Flipping Gestures for Efficient Command and Value Selection on SmartwatchesabstractSelecting an item of interest on smartwatches can be tedious and time-consuming as it involves a series of swipe and tap actions. We present PageFlip, a novel method that combines into a single action multiple touch operations such as command invocation and value selection for efficient interaction on smartwatches. PageFlip operates with a page flip gesture that starts by dragging the UI from a corner of the device. We first design PageFlip by examining its key design factors such as corners, drag directions and drag distances. We next compare PageFlip to a functionally equivalent radial menu and a standard swipe and tap method. Results reveal that PageFlip improves efficiency for both discrete and continuous selection tasks. Finally, we demonstrate novel smartwatch interaction opportunities and a set of applications that can benefit from PageFlip. Teng Han, Jiannan Li, Khalad Hasan, Keisuke Nakamura, Randy Gomez, Ravin Balakrishnan, Pourang Irani |
CHI | 5 |
| 2018 | Haru: Hardware Design of an Experimental Tabletop Robot AssistantabstractThis paper discusses the design and development of an experimental tabletop robot called "Haru" based on design thinking methodology. Right from the very beginning of the design process, we have brought an interdisciplinary team that includes animators, performers and sketch artists to help create the first iteration of a distinctive anthropomorphic robot design based on a concept that leverages form factor with functionality. Its unassuming physical affordance is intended to keep human expectation grounded while its actual interactive potential stokes human interest. The meticulous combination of both subtle and pronounced mechanical movements together with its stunning visual displays, highlight its affective affordance. As a result, we have developed the first iteration of our tabletop robot rich in affective potential for use in different research fields involving long-term human-robot interaction. Randy Gomez, Deborah Szapiro, Kerl Galindo, Keisuke Nakamura |
HRI | 1 |
| 2018 | Interactive Reinforcement Learning from Demonstration and Human Evaluative FeedbackabstractPrograming robots to perform tasks is difficult in the real world because of its richness and uncertainty. For robots and agents to be more useful, they must be able to learn quickly from ordinary people via natural interactions. In this paper, we investigate how an agent can learn from demonstration and positive and negative evaluative feedback provided by a human teacher. Specifically, we proposed a model-based method-IRL-TAMER-by combining learning from demonstration via inverse reinforcement learning (IRL) and learning from human reward via the TAMER framework. We tested our method in the Grid World domain and compared with the TAMER framework using different discount factors on human reward. Our results suggest that although an agent learning via IRL can learn a useful value function indicating which states are good based on the demonstration, it cannot obtain an effective policy navigating to the goal state with one demonstration. However, learning from demonstration can reduce the number of human reward needed to obtain an optimal policy, especially the number of negative feedback. That is to say, learning from demonstration can be a jump-start for agent's learning from human reward and reduce the number of mistakes-incorrect actions. Furthermore, our results show that learning from demonstration can only be useful for agent's learning from human reward when the discount factor is small, i.e., learning from myopic human reward. Guangliang Li, Bo He 0002, Randy Gomez, Keisuke Nakamura |
RO-MAN | 3 |
| 2018 | A robust multispectral palmprint matching algorithm and its evaluation for FPGA applications
Chao Li 0005, Yannick Benezeth, Keisuke Nakamura, Randy Gomez, Fan Yang 0019 |
J. Syst. Archit. | 4 |
| 2017 | Improving separation of overlapped speech for meeting conversations using uncalibrated microphone arrayabstractIn this paper, we propose a novel approach of sound source separation for meeting conversations even when using an uncalibrated microphone array. Our method can blindly estimate three parameters for separation, namely Steering Vectors (SVs), speaker indices, and activity periods of each speaker. First, we estimate the number of speakers and SVs by clustering Time Delay Of Arrival (TDOA) of the observed signal and selecting major clusters to compute TDOA-based SVs. Then, speaker indices and activity periods are estimated by thresholding spatial spectrum using estimated SVs, whose threshold is blindly obtained. Finally, we separate overlapped speeches/noise based on dynamic design of noise correlation matrices of the minimum variance distortionless response (MVDR) beamformer using blindly estimated parameters. The proposed algorithm was evaluated in both separation objective measure and recognition correct rate and showed improvements in both single and simultaneous speech scenarios in a reverberant meeting room. Moreover, the blindly estimated parameters improved separation and recognition compared to geometrically obtained parameters. Keisuke Nakamura, Randy Gomez |
ASRU | 2 |
| 2017 | Exploring data augmentation methods in reverberant human-robot voice communicationabstractCollecting training data is not an easy task especially in situation involving robots that require tremendous physical effort. The ability to augment data through synthetic means is a convenient tool to solve this problem. Therefore it is important to evaluate the extent of the usefulness of augmented data. In this paper, we will explore data augmentation schemes in reverberant environment and investigate a method to effectively select data. We experiment in a real reverberant environment condition and investigate both the traditional automatic speech recognition (ASR) system based on gaussian mixture model-hidden markov model (GMM-HMM) and the most current system based on Deep Neural Networks (i.e, HMM-DNN). Our results show that the combination of data augmentation and data selection, further improves system performance. In our experiments, we used real test data in a reverberant hands-free human-robot communication scenario. Randy Gomez, Keisuke Nakamura |
RO-MAN | 1 |
| 2017 | SoundCraft: Enabling Spatial Interactions on Smartwatches using Hand Generated AcousticsabstractWe present SoundCraft, a smartwatch prototype embedded with a microphone array, that localizes angularly, in azimuth and elevation, acoustic signatures: non-vocal acoustics that are produced using our hands. Acoustic signatures are common in our daily lives, such as when snapping or rubbing our fingers, tapping on objects or even when using an auxiliary object to generate the sound. We demonstrate that we can capture and leverage the spatial location of such naturally occurring acoustics using our prototype. We describe our algorithm, which we adopt from the MUltiple SIgnal Classification (MUSIC) technique [31], that enables robust localization and classification of the acoustics when the microphones are required to be placed at close proximity. SoundCraft enables a rich set of spatial interaction techniques, including quick access to smartwatch content, rapid command invocation, in-situ sketching, and also multi-user around device interaction. Via a series of user studies, we validate SoundCraft's localization and classification capabilities in non-noisy and noisy environments. Teng Han, Khalad Hasan, Keisuke Nakamura, Randy Gomez, Pourang Irani |
UIST | 4 |
| 2016 | Construction of Japanese Audio-Visual Emotion Database and Its Application in Emotion Recognition
Nurul Lubis, Randy Gomez, Sakriani Sakti, Keisuke Nakamura, Koichiro Yoshino, Satoshi Nakamura 0001, Kazuhiro Nakadai |
LREC | 2 |
| 2016 | Leveraging phantom signals for improved voice-based human-robot interactionabstractVoice-based system used in human-robot interaction is susceptible to challenging environment conditions. In an enclosed environment, the speech signal is often reflected which causes smearing as it is observed in the microphone. This phenomenon creates mismatch with the acoustic model, degrading the recognition performance and the robot's ability to understand and execute commands. Moreover, phantoms increase false-alarm in robot's attention system. To address these issues, environment-matched training and model adaptation may be used. The former requires enormous amount of training data to exhaustively cover different matched conditions whereas the latter needs several adaptation data to be collected at runtime. It is important to stress that data collection and the wait time are luxuries in a robot setup. In this paper, we extend our previous work that mitigates these problem by combining environment-adaptive training, speech enhancement with phantom awareness and fast model update, respectively. As a result, we achieve a robust voice-based system that enhances the observed speech, rejects phantoms and automatically updates the model at runtime to minimize the mismatch. Results show that the proposed method significantly outperforms our previous work. Randy Gomez, Yurii Vasylkiv, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai |
RO-MAN | 1 |
| 2015 | Temporal smearing compensation in reverberant environment for speech-based human-robot interactionabstractSpeech-based human-robot interaction is often plagued with issues such as reverberation and changes in speaker position that impacts overall performance. In this paper, we show a method in compensating the joint effects of reverberation and the change in speaker position. The acoustic perturbation caused by these two takes its toll on the Automatic Speech Recognition (ASR) and then the Spoken Language Understanding (SLU). Consequently, these will lead to a failure in the human-robot interaction experience. The proposed method is specifically designed to address the challenging environment condition in which robots are deployed. First, we analyze the impact of reverberation in the form of temporal smearing per change in speaker position. Then, we extract the smearing coefficients that capture the joint dynamics between the speech signal at current position and the room acoustics as observed by the robot. These coefficients are utilized to update the room transfer function (RTF) and the suppression parameters are stored offline. Moreover, all of these processes are optimized in the context of the ASR system for robot application. In the online mode, the reverberant data at an arbitrary position is processed using the parameters pre-computed offline. This effectively compensates the joint effects of reverberation at the arbitrary speaker position. Experimental results using real data gathered in a human-robot communication setting show that the proposed method outperforms existing methods. Randy Gomez, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai |
ICRA | 1 |
| 2015 | Dereverberation for active human-robot communication robust to speaker's face orientation
Randy Gomez, Levko Ivanchuk, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai |
INTERSPEECH | 1 |
| 2015 | Utilizing visual cues in robot audition for sound source discrimination in speech-based human-robot communicationabstractIt is easy for human beings to discern whether an observed acoustic signal is a direct speech, reflected speech or noise through simple listening. Relying purely on acoustic cues is enough for human beings to discriminate between the different kinds of sound sources which is not straightforward for machines. A robot equipped with the current robot audition mechanism in most cases, will fail to differentiate a direct speech from the other sound sources because acoustic information alone is insufficient for effective discrimination. Robot audition is an important topic in speech-based human-robot communication. It enables the robot to associate the incoming speech signal to the user for an effective human-robot communication. In challenging environments, this task becomes difficult due to reflections of the direct speech signal and background noise sources. To counter this problem, a robot needs to have a minimum amount of prior information to discriminate the valid speech signal (direct speech) from the contaminants (i.e., speech reflections and background noise sources). Failure to do so would lead to false speech-to-speaker association in robot audition and will gravely impact human-robot communication experience. In this paper we propose to using visual cues to augment the traditional robot audition which relies solely on acoustic information. The proposed method significantly improves accuracy of speech-to-speaker association and machine understanding performance in real environment situation. Experimental results show that our expanded system is robust in discriminating direct speech from speech reflections and background noise sources. Randy Gomez, Levko Ivanchuk, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai |
IROS | 1 |
| 2014 | Speech-based human-robot interaction robust to acoustic reflections in real environmentabstractAcoustic reflection inside an enclosed environment is detrimental to human-robot interaction. Reflection may manifest as phantom sources emanating from unknown directions. In effect, a single speaker may falsely manifest as multiple speakers to the robot audition system, impeding the robot's ability to correctly associate the speech command to the actual speaker. Moreover, speech reflection smears the original speech signal due to reverberation. This degrades speech recognition and understanding performance. Conventional robot audition schemes that rely purely on acoustics and spatial information are very sensitive to acoustic reflection which ultimately leads to the failure in human-robot interaction. We propose a method for human-robot interaction robust to the effect of acoustic reflection. First, visual information is utilized and head tracking scheme is employed to reinforce the acoustic information with the visual presence of a prospect user. Second, we employ a model-based sound event identification scheme and scrutinize whether the acoustic information is likely to be speech or non-speech. Using all the information we have gathered, we create a simple rule construct to effectively discriminate the original source (actual speaker) from phantom sources (reflection). Consequently, the corresponding source identified as phantom (reflection) is used to estimate the unwanted smearing for effective suppression via speech enhancement. Experiments are conducted in human-robot interaction setting in which the proposed method outperforms the conventional method. Randy Gomez, Koji Inoue, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai |
IROS | 1 |
| 2014 | Multiparty Interaction Understanding Using Smart Multimodal Digital SignageabstractThis paper presents a novel multimodal system designed for multi-party human-human interaction analysis. The design of human-machine interfaces for multiple users is challenging because simultaneous processing of actions and reactions have to be consistent. The proposed system consists of a large display equipped with multiple sensing devices: microphone array, HD video cameras, and depth sensors. Multiple users positioned in front of the panel freely interact using voice or gesture while looking at the displayed content, without wearing any particular devices (such as motion capture sensors or head mounted devices). Acoustic and visual information is captured and processed jointly using established and state-of-the-art techniques to obtain individual speech and gaze direction. Furthermore, a new framework is proposed to model A/V multimodal interaction between verbal and nonverbal communication events. Dynamics of audio signals obtained from speaker diarization and head poses extracted from video images are modeled using hybrid dynamical systems (HDS). We show that HDS temporal structure characteristics can be used for multimodal interaction level estimation, which is useful feedback that can help to improve multi-party communication experience. Experimental results using synthetic and real-world datasets of group communication such as poster presentations show the feasibility of the proposed multimodal system. Tony Tung, Randy Gomez, Tatsuya Kawahara, Takashi Matsuyama |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2013 | Robustness to speaker position in distant-talking automatic speech recognitionabstractIn this paper, we show a method that significantly improved our previous work in single-channel dereverberation. The proposed method is more robust to changes in speaker position in distant talking ASR. First, we update the room transfer function (RTF) and weighting parameters for dereverberation to the target speaker position. This scheme corrects speech power variation as a function of position in the waveform level. Consequently, its impact to the acoustic model is verified. Then, we implement a fast acoustic model update reflective of the speech power level of the target speaker position. Furthermore, the scheme in updating the model is simple and precludes time-consuming model re-estimation. As a result, the proposed method can be executed online. The synergy of these corrective measures significantly minimizes the mismatch between training and testing conditions. We test our method using real reverberant data with different locations inside the room. Experimental results show that the proposed method outperforms the conventional methods in terms of ASR performance. Moreover, our fast acoustic model update scheme is at par in terms of recognition performance against time-consuming model re-estimation. Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai |
ICASSP | 1 |
| 2013 | Hands-free human-robot communication robust to speaker's radial positionabstractIn this paper we present a method in room transfer function (RTF) estimation, employed specifically for dereverberation in hands-free human-robot communication.We introduce a radial distance compensation scheme which significantly improved the RTF estimate robust to the speech power variation due to changes in speaker's radial position. The proposed method is implemented in two levels; first, waveform-level compensation is executed to reflect the change in power caused by the change of radial position to the RTF. We generated possible RTF estimates within a close neighbourhood based on curve fitting. Then, we select among these estimates the optimal RTF based on acoustic model likelihood criterion, the same criterion employed in automatic speech recognition (ASR) systems. The latter is referred to as acoustic model-level compensation, which links the generated RTF to the ASR. We note that in ASR application, both waveform and acoustic models play an important role in achieving optimal performance. Thus, the synergistic effect of the two processes guarantee ASR performance improvement when used in conjunction with our ASR-based dereverberation scheme. Experimental evaluation show robustness in recognition performance when used in hands-free human-robot communication environment. Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai, Ui-Hyun Kim, Hiroshi G. Okuno, Tatsuya Kawahara |
ICRA | 1 |
| 2013 | Dereverberation robust to speaker's azimuthal orientation in multi-channel human-robot communicationabstractThe acoustical dynamics of reverberation in an enclosed environment poses a problem to human-robot communication. Any change in the azimuthal orientation of the speaker contributes to unpredictable acoustical activity resulting in a degradation in the performance of the automatic speech recognition (ASR) system. Thus, dereverberation techniques need to address this issue prior to ASR. Dereverberation in multi-channel applications primarily evolves in the adoption of a suitable reverberant model that results to a computationally feasible solution and at the same time yields an accurate estimate of the harmful reflections (i.e., late reflection) for effective suppression. In this paper we address this problem by introducing a hybrid method based on multi-channel processing on a singlechannel reverberant model platform. The proposed method is capable of accurate signal estimation, a property inherent to a multi-channel system, and at the same time bears the computational efficiency derived from single-channel reverberant model approach. The proposed method is summarized as follows; First, multi-channel sound-source processing is employed to obtain the full reverberant and the late reflection signal estimates. Then, equalization is employed to update the late reflection estimate reflective of the change in azimuth prior to dereverberation. The equalization parameters for azimuthal change are obtained through an offline optimization procedure. Experimental evaluation in an actual human-robot communication environment shows that the proposed method outperforms existing methods in terms of robustness in the ASR performance. Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai |
IROS | 1 |
| 2013 | Real-time super-resolution three-dimensional sound source localization for robotsabstractThis paper investigates Sound Source Localization (SSL) for a robot in a real world. Previously, we focused on one-dimensional SSL for azimuth and assumed that target sources are distributed close to a horizontal plane. Without this assumption, the SSL performance is drastically degraded. Thus, three-dimensional SSL is essential to improve the localization for sound sources distributed in a three-dimensional space. Compared to one-dimensional SSL, three-dimensional SSL mainly has the following problems: 1) a massive number of Transfer Function (TF) measurements for microphone array calibration are required for three dimensions to maintain the spatial resolution of SSL sufficiently-high, 2) the computational cost for searching for sound sources drastically increases in high-dimensional spaces. For the first issue, we extend the previously-proposed one-dimensional TF interpolation method, integrating time-domain-based and frequency-domain-based interpolation, to three dimensions. The interpolation achieves three-dimensional super-resolution SSL and reduction of the number of TF measurements while maintaining the spatial resolution of SSL. For the second issue, we propose optimal hierarchical SSL, which reduces computational cost for searching for sound sources by introducing a hierarchical search algorithm instead of using greedy search in localization. We previously proposed the concept of the algorithm. This paper additionally discusses theoretical optimality in hierarchization to minimize the total computational cost of SSL. The method determines the number of hierarchies and the resolution of each hierarchy depending on desired spatial resolution. These techniques are integrated into an SSL system using a robot. The experimental result showed: 1) the proposed interpolation method achieved super-resolution SSL working better than that with pre-measured TFs, 2) the optimal hierarchical SSL drastically reduced computational cost by approximately 97%. Keisuke Nakamura, Randy Gomez, Kazuhiro Nakadai |
IROS | 2 |
| 2012 | Multi-party human-robot interaction with distant-talking speech recognitionabstractSpeech is one of the most natural medium for human communication, which makes it vital to human-robot interaction. In real environments where robots are deployed, distant-talking speech recognition is difficult to realize due to the effects of reverberation. This leads to the degradation of speech recognition and understanding, and hinders a seamless human-robot interaction. To minimize this problem, traditional speech enhancement techniques optimized for human perception are adopted to achieve robustness in human-robot interaction. However, human and machine perceive speech differently: an improvement in speech recognition performance may not automatically translate to an improvement in human-robot interaction experience (as perceived by the users). In this paper, we propose a method in optimizing speech enhancement techniques specifically to improve automatic speech recognition (ASR) with emphasis on the human-robot interaction experience. Experimental results using real reverberant data in a multi-party conversation, show that the proposed method improved human-robot interaction experience in severe reverberant conditions compared to the traditional techniques. Randy Gomez, Tatsuya Kawahara, Keisuke Nakamura, Kazuhiro Nakadai |
HRI | 1 |
| 2012 | Dereverberation based on Wavelet Packet Filtering for Robust Automatic Speech Recognition
Tatsuya Kawahara, Randy Gomez |
INTERSPEECH | 2 |
| 2011 | Denoising Using Optimized Wavelet Filtering for Automatic Speech RecognitionabstractWe present an improved denoising method based on filtering of the noisy wavelet coefficients using a Wiener gain for automatic speech recognition (ASR). We optimize the wavelet parameters for speech and different noise profiles to achieve a better estimate of the Wiener gain for effective filtering. Moreover, we introduce a scaling parameter in the Wiener gain to minimize mismatch caused by distortion during the denoising process. Experimental results in large vocabulary continuous speech recognition (LVCSR) show that the proposed method is effective and robust to different noise conditions. IndexTerms: Speech recognition, Robustness, Denoising and Wavelet Randy Gomez, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2010 | Optimizing spectral subtraction and wiener filtering for robust speech recognition in reverberant and noisy conditionsabstractSpeech enhancement is a common approach to address the effects of degradation due to noise and channel contamination. This approach is intended to suppress unwanted signal and recover the clean speech. In this paper, we focus on two simple and low-computational methods: Wiener filtering (WF) and spectral subtraction (SS). Conventionally, these are formulated with no relation with automatic speech recognition (ASR). We propose to optimize the conventional speech enhancement technique in relation with likelihood of the acoustic model. We also exploit these simple speech enhancement techniques that are originally designed for denoising, to address reverberation as well. In the experiment with real noisy and reverberant environments, we have achieved significant improvement in recognition performance using the proposed approach. Randy Gomez, Tatsuya Kawahara |
ICASSP | 1 |
| 2010 | An improved wavelet-based dereverberation for robust automatic speech recognitionabstractThis paper presents an improved wavelet-based dereverberation method for automatic speech recognition (ASR). Dereverberation is based on filtering reverberant wavelet coefficients with the Wiener gains to suppress the effect of the late reflections. Optimization of the wavelet parameters using acoustic model enables the system to estimate the clean speech and late reflections effectively. This results to a better estimate of the Wiener gains for dereverberation in the ASR application. Additional tuning of the parameters of the Wiener gain in relation with the acoustic model further improves the dereverberation process for ASR. In the experiment with real reverberant data, we have achieved a significant improvement in ASR accuracy. Randy Gomez, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2010 | Robust Speech Recognition Based on Dereverberation Parameter Optimization Using Acoustic Model LikelihoodabstractAutomatic speech recognition (ASR) in reverberant environments is a challenging task. Most dereverberation techniques address this problem through signal processing and enhances the reverberant waveform independent from the speech recognizer. In this paper, we propose a novel scheme to perform dereverberation in relation with the likelihood of the back-end ASR system. Our proposed approach effectively selects the dereverberation parameters, in the form of multiband scale factors, so that they improve the likelihood of the acoustic model. Then, the acoustic model is retrained using the optimal parameters. During the recognition phase, we implement additional optimization of the parameters. By using Gaussian mixture model (GMM), the process for selecting the scale factors become efficient. Moreover, we remove the dependency of the adopted dereverberation technique on the room impulse response (RIR) measurement, by using an artificial RIR generator and selecting based on the acoustic likelihood. Experimental results show significant improvement in recognition performance with the proposed method over the conventional approach. Randy Gomez, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Optimization of dereverberation parameters based on likelihood of speech recognizerabstractSpeech recognition under reverberant condition is a difficult task. Most dereverberation techniques used to address this problem enhance the reverberant waveform independent from that of the speech recognizer. In this paper, we improve the conventional Spectral Subtraction-based (SS) dereverberation technique. In our proposed approach, the dereverberation parameters are optimized to improve the likelihood of the acoustic model. The system is capable of adaptively fine-tuning these parameters jointly with acoustic model training. Additional optimization is also implemented during decoding of the test utterances. We have evaluated using real reverberant data and experimental results show that the proposed method significantly improves the recognition performance over the conventional approach. Randy Gomez, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2009 | Techniques in rapid unsupervised speaker adaptation based on HMM-Sufficient Statistics
Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
Speech Commun. | 1 |
| 2008 | Distant talking robust speech recognition using late reflection components of room impulse responseabstractWe propose a robust and fast dereverberation technique for real-time speech recognition application. First, we effectively identify the late reflection components of the room impulse response. We use this information together with the concept of Spectral Subtraction (SS) to remove the late reflection components of the reverberant signal. In the absence of the clean speech in actual scenario, approximation is carried out in estimating the late reflection where the estimation error is corrected through multi-band SS. The multi-band coefficients are optimized during offiine training and used in the actual online dereverberation. The proposed method performs better and faster than the relevant approach using Multi-LPC and reverberant matched model. Moreover the proposed method is robust to speaker and microphone locations. Randy Gomez, Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano |
ICASSP | 1 |
| 2008 | Rapid unsupervised speaker adaptation robust in reverberant environment conditionsabstractINTERSPEECH2008: 9th Annual Conference of the International Speech Communication Association, September 22-26, 2008, Brisbane, Australia. Randy Gomez, Jani Even, Kiyohiro Shikano |
INTERSPEECH | 1 |
| 2007 | Rapid unsupervised speaker adaptation using single utterance based on MLLR and speaker selectionabstractINTERSPEECH2007: 8th Annual Conference of the International Speech Communication Association, August 27-31, 2007, Antwerp, Belgium. Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
INTERSPEECH | 1 |
| 2006 | Improving Rapid Unsupervised Speaker Adaptation Based On Hmm Sufficient StatisticsabstractIn real-time speech recognition applications, there is a need to implement a fast and reliable adaptation algorithm. We propose a method to reduce adaptation time of the unsupervised speaker adaptation based on HMM-sufficient statistics. We use only a single arbitrary utterance without transcriptions in selecting the N-best speakers' sufficient statistics created offline to provide data for adaptation to a target speaker. Further reduction of N-best implies a reduction in adaptation time. However, it degrades recognition performance due to insufficiency of data needed to robustly adapt the model. Linear interpolation of the global HMM-sufficient statistics offsets this negative effect and achieves a 50% reduction in adaptation time without compromising the recognition performance. We have reduced the adaptation time from 10 sec to 5 sec without degradation of the word accuracy. Furthermore, we compared our method with vocal tract length normalization (VTLN), maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR). Moreover, we tested in office, car, crowd and booth noise environments in 10 dB, 15 dB, 20 dB and 25 dB SNRs Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
ICASSP (1) | 1 |
| 2005 | Rapid unsupervised speaker adaptation based on multi-template HMM sufficient statistics in noisy environmentsabstractINTERSPEECH2005: the 9th European Conference on Speech Communication and technology, September 4-8, 2005, Lisbon, Portugal. Randy Gomez, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano |
INTERSPEECH | 1 |
| 2004 | Robust speech recognition with spectral subtraction in low SNRabstractICSLP2004: the 8th International Conference on Spoken Language Processing, October 4-8, 2004, Jeju Island, Korea. Randy Gomez, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano |
INTERSPEECH | 1 |