Yurii Vasylkiv

dblp:190/3035 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2023
0000-0003-3603-1645ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 How to Make a Robot Grumpy Teaching Social Robots to Stay in Character with Mood Steering
abstract
Conveying a robot's target mood is crucial to successful social interactions. The robot's expressive performance must be appropriate, persuasive, and consistent. However, this is challenging when interactions contain a mixture of scripted and improvised content, such as those generated by language models. In this paper, we take on the task of teaching robots to stay in character, that is to say, exhibit consistency in mood during interactions. We start by defining a communication strategy module that allows for the top-down specification of a target robot mood for a given task, goal, or context. We then propose a mood steering framework for enforcing robot mood consistency throughout an interaction that supports several target moods. Our framework consists of two components: 1. expressivity steering specifies the speech and behavior to be used by the robot to convey a target mood, and 2. language model steering ensures that improvised language is consistent with the robot's target mood. As a first step toward identifying effective communication strategies, we implement grumpy and cheerful strategies for a collaborative storytelling game and compare them to a neutral baseline. Evaluation in a collaborative storytelling game shows that our approach generates robot behavior that successfully conveys the robot's target mood throughout gameplay and language model steering generates story contributions that capture the target mood without quality degradation and raises important issues for communication strategy design.
Eric Nichols, Deborah Szapiro, Yurii Vasylkiv, Randy Gomez
IROS3
2022 Affective Behavior Learning for Social Robot Haru with Implicit Evaluative Feedback
abstract
We propose a human-in-the-loop reinforcement learning mechanism to help robots learn emotional behavior. Unlike the previous methods of providing explicit feedback via pressing keyboard buttons or mouse clicks, we provide a more natural way for ordinary people to train social robots how to perform social tasks according to their preferences - facial expressions. The whole experiment is carried out on the desktop robot Haru, which is mainly used for the research of emotion and empathy participation. Our experimental results show that through learning from implicit feedback of facial features, Haru can quickly understand and dynamically adapt to individual preferences, and obtain a similar performance to learning from explicit feedback. In addition, we observe that the recognition error of human feedback will cause a “temporary regress” of the robot's learning performance, which is more obvious at the beginning of the training process. This phenomenon is shown to be correlated with the accuracy of recognizing negative implicit feedback.
Hui Wang 0141, Jinying Lin, Yurii Vasylkiv, Heike Brock, Keisuke Nakamura, Randy Gomez, Bo He 0002, Guangliang Li
IROS4
2022 I Can't Believe That Happened! : Exploring Expressivity in Collaborative Storytelling with the Tabletop Robot Haru
abstract
Collaborative storytelling has long been a goal of social robotics, however, much of this research is limited in interactivity or assumes that story content is curated. In this paper, we present a working fully-automatic collaborative storytelling robot, which can collaborate with a person to create a unique, improvised story by using a large-scale neural language model to dynamically generate continuations to a story. Because effective storytelling requires engaging the emotions of participants, we explore several modalities of procedurally-generated expressivity: 1. an expressive text-to-speech voice with several delivery styles, 2. physical and verbal reactions performed by the robot, and 3. an external display used to show instructions and graphics during storytelling.To understand the issues associated with improvised collaborative storytelling with a social robot, we conduct an online survey and elicitation study with a group of online observers of collaborative storytelling gameplay, comparing several expressivity strategies in terms of storytelling-related characteristics, expressivity characteristics, and personality traits as measured by RoSAS. This evaluation showed that expressivity strategies using both emotive voice and performed reactions were perceived to be more competent storytellers and more strongly associated with positive personality traits.
Eric Nichols, Deborah Szapiro, Yurii Vasylkiv, Randy Gomez
RO-MAN3
2021 Automating Behavior Selection for Affective Telepresence Robot
abstract
The tabletop robot Haru, used for affective telepresence research, enables a teleoperator to communicate affects from a distance. The robot’s expressiveness offers myriad ways of communicating affects through the execution of emotive routines. The teleoperator reacts to input modalities such as the user’s facial expression, gestures and speech-based intent as perceived by the robot’s perception system. However, due to the sheer number of routines to select from, the task of choosing the appropriate or the most preferred routine is becoming cumbersome. In this paper, we propose a human-in-the-loop reinforcement learning mechanism in which an agent learns the teleoperator’s selection preference as a function of the input modalities and aids the routine selection process by narrowing it to n-best optimal choices. Our experimental results show that with only a few number of interactions from the teleoperator, the system can learn to recommend optimal routine behaviors for all perceived modalities, which greatly reduces the workload of the teleoperator.
Yurii Vasylkiv, Guangliang Li, Eleanor Sandry, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez
ICRA1
2021 Collaborative Storytelling with Social Robots
abstract
Storytelling plays a central role in human socializing and entertainment, and research on conducting storytelling with robots is gaining interest. However, much of this research assumes that story content is curated. In this paper, we expand the recently-proposed task of collaborative storytelling, where an intelligent agent and a person collaborate to create a unique story by taking turns adding to it, for application to social robot and consider the design implications that arise. Since latency can be detrimental to human-robot interaction, we examine the performance-latency trade-offs of an existing generate-and-rank-based approach to collaborative storytelling by finding the optimal ranker’s sample size that strikes the best balance between quality and computational cost. We improve on existing evaluation that was previously based on system-generated stories by having human participants play the collaborative storytelling game with our system and comparing the stories they create with our system to a naive baseline. Finally, we conduct a pilot elicitation survey that sheds light on issues to consider when adapting our collaborative storytelling system to a social robot. Our evaluation shows that participants have a positive view of collaborative storytelling with a social robot and consider rich, emoting capabilities to be key to enjoyment.
Eric Nichols, Leo Gao, Yurii Vasylkiv, Randy Gomez
IROS3
2021 Shaping Affective Robot Haru's Reactive Response
abstract
We describe a method of teaching a robot its empathic behavioural response from its interaction with people. We used the input modalities such as relative spatial information, facial expressions, body gestures and speech information as perception input that triggers the robot’s empathic response. First, we bootstrap the training through a pre-learning mechanism in which training is conducted by users who know the robotic system. This phase provides simulation-based training using a simple graphical user interface to simulate the input, rewards and correction feedback. In the second phase, we developed an online learning scheme for naive users to personalize their robot further, building on top of the bootstrapped model. Here, we developed a natural user interface that enables natural human-robot interaction via the suite of sensors that allows the users to provide evaluative feedback during the interaction with the robot. We evaluated the system and our results show that bootstrapping is an efficient tool to hasten the robot’s learning while online learning provided some form of personalization in the real environment with naive users.
Yurii Vasylkiv, Guangliang Li, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez
RO-MAN1
2016 Leveraging phantom signals for improved voice-based human-robot interaction
abstract
Voice-based system used in human-robot interaction is susceptible to challenging environment conditions. In an enclosed environment, the speech signal is often reflected which causes smearing as it is observed in the microphone. This phenomenon creates mismatch with the acoustic model, degrading the recognition performance and the robot's ability to understand and execute commands. Moreover, phantoms increase false-alarm in robot's attention system. To address these issues, environment-matched training and model adaptation may be used. The former requires enormous amount of training data to exhaustively cover different matched conditions whereas the latter needs several adaptation data to be collected at runtime. It is important to stress that data collection and the wait time are luxuries in a robot setup. In this paper, we extend our previous work that mitigates these problem by combining environment-adaptive training, speech enhancement with phantom awareness and fast model update, respectively. As a result, we achieve a robust voice-based system that enhances the observed speech, rejects phantoms and automatically updates the model at runtime to minimize the mismatch. Results show that the proposed method significantly outperforms our previous work.
Randy Gomez, Yurii Vasylkiv, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
RO-MAN2