Ryo Ishii

dblp:97/5316 · DBLP profile ↗
← Back
57ranked-venue papers
29as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 37 · 20 first-author · 13 since 2021Artificial intelligence and machine learning · 31 · 16 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 GlossRefine: Gloss-Conditioned Transformer for Low-Resource Sign Language Motion Generation
Ryo Ishii, Shin'ichiro Eitoku, Junichi Sawase
FG1
2025 Investigating Role of Big Five Personality Traits in Audio-Visual Rapport Estimation
abstract
Automatic rapport estimation in social interactions is a central component of affective computing. Recent reports have shown that the estimation performance of rapport in initial interactions can be improved by using the participant’s personality traits as the model’s input. In this study, we investigate whether this findings applies to interactions between friends by developing rapport estimation models that utilize nonverbal cues (audio and facial expressions) as inputs. Our experimental results show that adding Big Five features (BFFs) to nonverbal features can improve the estimation performance of self-reported rapport in dyadic interactions between friends. Next, we demystify how BFFs improve the estimation performance of rapport through a comparative analysis between models with and without BFFs. We decompose rapport ratings into perceiver effects (people’s tendency to rate other people), target effects (people’s tendency to be rated by other people), and relationship effects (people’s unique ratings for a specific person) using the social relations model. We then analyze the extent to which BFFs contribute to capturing each effect. Our analysis demonstrates that the perceiver’s and the target’s BFFs lead estimation models to capture the perceiver and the target effects, respectively. Furthermore, our experimental results indicate that the combinations of facial expression features and BFFs achieve best estimation performances not only in estimating rapport ratings, but also in estimating three effects. Our study is the first step toward understanding why personality-aware estimation models of interpersonal perception accomplish high estimation performance.
Takato Hayashi, Ryusei Kimura, Ryo Ishii, Shogo Okada
FG3
2025 Instant 3DCG Dance Generation System Based on Music and Dance Composition
abstract
We present a novel system that automatically generates and visualizes 3DCG dance animations based on the user’s preferred music and dance composition. The key technology of the system is a transformer-based diffusion model that produces dance choreographies conditioned on arbitrary inputs of music audio and dance composition. Integrated into a user-friendly GUI, the system allows users to instantly generate and preview multiple dance sequences simply by selecting their desired music and dance composition. This capability supports both creative choreography ideation and effective dance practice.
Ryo Ishii, Shin'ichiro Eitoku, Keigo Fushio, Yoshihide Sato, Louis-Philippe Morency
FG1
2025 CDCGM: Composition-specified Dance Choreography Generation from Music
abstract
Significant research attention has recently been focused on the automatic generation of human dance choreography from music. While several generation models have been proposed, they cannot specify what kind of movements to generate, and as a result, random movements are generated. We therefore propose a generation model called Composition-specified Dance Choreography Generation from Music (CDCGM) that enables creators to specify which dance composition (i.e., type of movement) to take when generating a dance at each time step. We implemented CDCGM by first constructing a new dataset that includes motion captures of breakdancing and time-series annotation data of representative movement types. Evaluation experiments using our corpus showed that CDCGM can generate dances that faithfully reflect the specified dance composition with high quality. Compared to conventional state-of-the-art models, CDCGM is capable of generating quality dances that improve the expressiveness and the degree to which the dance matches the content and timing of the music. We also propose a new application for CDCGM in which users watch newly generated dance choreography simply by entering music and dance composition. The results of a user study evaluation of the application demonstrated that users found the experience of generating dance by specifying any dance composition for any music extremely fun, that it has the potential to greatly contribute to dance choreography and learning, and that there is a strong desire to use this application on a daily basis.
Ryo Ishii, Shin'ichiro Eitoku, Louis-Philippe Morency
FG1
2025 Impact of Personality on Generation of Co-speech Nonverbal Behaviors Represented by 3D Skeleton Pose
abstract
In this study, we examine how incorporating personality traits into a nonverbal behavior generation model for upper-body motion (head, arms, and posture) affects the quality and characteristics of the generated behaviors. We first constructed a multimodal dialogue corpus containing speech audio, transcripts, 3D upper-body skeleton data, and participants’ Big Five personality scores, and then used the corpus to develop a model that predicts 3D skeleton coordinates from speech, text, and personality traits. Objective evaluation showed that the model with personality input more accurately reproduced individualized behaviors aligned with personality traits. The generated gestures also reflected the relationship between gesture expressivity and personality. Subjective evaluation further showed that observers could reliably perceive intended differences in personality levels—specifically, high vs. low Big Five scores—based only on the generated movements. These findings demonstrate that modeling personality traits enables the generation of agent behaviors that are both personality-consistent and perceptible to users.
Ryo Ishii, Shin'ichiro Eitoku, Yoshihide Sato
HAI1
2025 Predicting End-of-turn and Backchannel Based on Multimodal Voice Activity Prediction Model
Ryo Ishii, Shin'ichiro Eitoku, Ryota Yokoyama, Junichi Sawase
ICMI1
2025 Support for Building Relationships in Speed Dating Through Observation of Pre-dialogue Simulations Using Digital Twins
Yoko Ishii, Ryo Ishii, Lidwina Andarini, Kazuya Matsuo, Atsushi Otsuka
INTERACT (2)2
2024 Prediction of Praising Skills Based on Multimodal Information
abstract
Praising behavior is an important method of communication. An existing study constructed models to predict praising skill, which indicates the degree to which the praise is done well, by using only unimodal behavior such as speech audio or visual behavior of a praiser who gives praise in dyad interactions. To improve prediction performance, a model should be constructed that uses various additional information. In this study, we propose two approaches to predict praising skill highly accurately. The first uses trimodal (multimodal) behaviors extracted from visual, acoustic, and linguistic modalities. The second uses the behaviors of the receiver of praise since the reaction of the receiver should differ depending on how good the praise is. For this study, we collect trimodal features and the degree of praising skill in each praising scene in a dialogue. We construct multiple models to predict the degree of praising skills using various combinations of the trimodal features from the praiser and receiver. The experimental results show that the model that predicts praising skill most accurately uses multiple features related to both verbal and nonverbal behaviors of the praiser and receiver. Therefore, the two approaches of using trimodal behaviors and using features from both the receiver and praiser are effective for predicting praising skills in dyad interactions.
Toshiki Onishi, Asahi Ogushi, Ryo Ishii, Atsushi Fukayama, Akihiro Miyata
ACII3
2024 Participation Role-Driven Engagement Estimation of ASD Individuals in Neurodiverse Group Discussions
abstract
Adults with autism spectrum disorder (ASD) face difficulties in communicating with neurotypical people in their daily lives and workplaces. In addition, research on modeling communication in neurodiverse groups is scarce. To recognize communication difficulties caused by neurodiversity, we first, collected a multimodal corpus for decision-making discussions in neurodiverse groups that included a person with ASD and two neurotypical participants. For corpus analysis, we investigated eye-gaze and facial expression exchanges between individuals with ASD and neurotypical participants during both listening and speaking. The findings were extended to automatically estimate the engagement of ASD individuals. To capture the effect of contingent behaviors between ASD individuals and neurotypical participants, we developed a transformer-based model that considers the participation role by changing the direction of cross-person attention depending on whether the ASD individual is listening or speaking. The proposed approach yields comparable results to the state-of-the-art for engagement estimation in neurotypical group conversations while accounting for the dynamic nature of behavior influence in face-to-face interactions. The code associated with this study is available at https://github.com/IUI-Lab/switch-attention.
Kalin Stefanov, Yukiko I. Nakano, Chisa Kobayashi, Ibuki Hoshina, Tatsuya Sakato, Fumio Nihei, Chihiro Takayama, Ryo Ishii, Masatsugu Tsujii
ICMI8
2024 GeSTICS: A Multimodal Corpus for Studying Gesture Synthesis in Two-party Interactions with Contextualized Speech
abstract
Generating natural co-speech gestures and facial expressions for effective human-agent interactions requires modeling the intricate interplay between verbal, non-verbal, and contextual cues observed in dyadic human communication. Two types of contextual cues are of particular interest: (1) individual factors of the interlocutors, such as their demographic attributes, and (2) situational factors, like the outcome of a preceding event. To facilitate their study, we introduce the GeSTICS Dataset, a novel multimodal corpus comprising 9,853 questions and 10,460 answers from audiovisual recordings of post-game sports interviews by 147 interviewees. The dataset contains speech data, including textual transcriptions, lexical descriptors, and acoustic features, as well as visual data encompassing the interviewee’s body pose and facial expressions, with an emphasis on capturing these modalities during both the question-listening and answering phases of the interview. Furthermore, GeSTICS incorporates metadata about individual factors, such as the age and cultural background of the interviewees, and situational factors, like the results of the games, which are often overlooked in existing multimodal datasets. Our preliminary analysis of GeSTICS reveals that the effects of speech features, such as loudness and lexical choice, on the production of co-speech gestures in both speaking and listening phases are moderated by situational factors and the interviewee’s individual factors. GeSTICS is designed to enhance the generation of realistic nonverbal behaviors in virtual agents, animated characters, and human-robot interaction systems, thus contributing to more engaging and effective human-agent communication. The analysis code and the dataset are available at https://gestics.github.io.
Gaoussou Youssouf Kebe, Mehmet Deniz Birlikci, Auriane Boudin, Ryo Ishii, Jeffrey M. Girard, Louis-Philippe Morency
IVA4
2023 Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance Estimation
abstract
This study investigated differences in the contributions of various features to in-person (IP) and video-mediated (VM) meetings. We focused on estimating important utterances using both an IP and a VM meeting corpora as the analysis data. A transformer model with dialogue history was used to estimate important utterances, and five types of input (text, speaker’s audio, others’ audio, speaker’s video, and others’ video) were fed to the model. A comparison of the models for IP and VM revealed that the speaker’s audio has a strong effect on the IP model, the video of the other participants strongly affects the VM model, and the text and others’ audio strongly affects both models in estimating important utterances.
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Atsushi Fukayama, Takao Nakamura
ICASSP2
2023 Continual Learning for Personalized Co-Speech Gesture Generation
abstract
Co-speech gestures are a key channel of human communication, making them important for personalized chat agents to generate. In the past, gesture generation models assumed that data for each speaker is available all at once, and in large amounts. However in practical scenarios, speaker data comes sequentially and in small amounts as the agent personalizes with more speakers, akin to a continual learning paradigm. While more recent works have shown progress in adapting to low-resource data, they catastrophically forget the gesture styles of initial speakers they were trained on. Also, prior generative continual learning works are not multimodal, making this space less studied. In this paper, we explore this new paradigm and propose C-DiffGAN: an approach that continually learns new speaker gesture styles with only a few minutes of per-speaker data, while retaining previously learnt styles. Inspired by prior continual learning works, C-DiffGAN encourages knowledge retention by 1) generating reminiscences of previous low-resource speaker data, then 2) crossmodally aligning to them to mitigate catastrophic forgetting. We quantitatively demonstrate improved performance and reduced forgetting over strong baselines through standard continual learning measures, reinforced by a qualitative user study that shows that our method produces more natural, style-preserving gestures. Code and videos can be found at https://chahuja.com/cdiffgan
Chaitanya Ahuja, Pratik Joshi, Ryo Ishii, Louis-Philippe Morency
ICCV3
2023 Prediction of Love-Like Scores After Speed Dating Based on Pre-obtainable Personal Characteristic Information
Ryo Ishii, Fumio Nihei, Yoko Ishii, Atsushi Otsuka, Kazuya Matsuo, Narichika Nomoto, Atsushi Fukayama, Takao Nakamura
INTERACT (4)1
2023 How Far ahead Can Model Predict Gesture Pose from Speech and Spoken Text?
abstract
We investigated how far into the future nonverbal behavior can be predicted from speech and speech text. Specifically, we build a model that generates future behaviors from speech and speech text information and evaluate the quality of the generated behaviors. This helps to clarify how far into the future behavior can be accurately predicted. Our experimental results show that in Gesture Pose Generation using speech and speech text, on the basis of the input speech and text, the nonverbal behavior up to at least 500 ms ahead can be predicted with objective evaluation values that are the same as those when no future prediction is made. This result shows a new possibility for Gesture Pose Generation using speech and speech text to predict the future up to at least 500 ms ahead with no performance degradation.
Ryo Ishii, Akira Morikawa, Shin'ichiro Eitoku, Atsushi Fukayama, Takao Nakamura
IVA1
2023 A Study of Prediction of Listener's Comprehension Based on Multimodal Information
abstract
During dialogues, speakers need to be able to predict whether their partners understand their message. This is important for not only for human-to-human interaction but also human-to-agent interaction. We consider that if the listener's comprehension level can be automatically predicted, interactive agents will be able to communicate appropriately according to the user's comprehension level. However, to the best of our knowledge, there is no case study that reveals how comprehension can be predicted based on multimodal information about the listener. In this study, we attempt to predict comprehension levels on the basis of the listener's multimodal information. First, we construct a dialogue corpus consisting of the listener's comprehension levels and the listener's multimodal information. Next, we construct machine learning models that predict the listener's comprehension levels on the basis of the listener's multimodal information. Our results suggest that our model was able to predict a listener's comprehension level on the basis of a listener's multimodal information. In addition, two movements, the lifting of the cheeks and the pulling up of the corners of the lips, were suggested to be important in assessing the listener's level of comprehension.
Shunichi Kinoshita, Toshiki Onishi, Naoki Azuma, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
IVA4
2023 Prediction of Various Backchannel Utterances Based on Multimodal Information
abstract
The listener's backchannels are an important part of dialogues. With appropriate backchannels, people are able to smoothly promote dialogues. Thus, backchannels are considered to be important in dialogues between not only humans but also humans and agents. Progress has been made in studying dialogue agents that perform natural affable dialogue. However, we have not clarified whether the listener's various backchannel types are predictable using the speaker's multimodal information. In this paper, we attempt to predict a listener's various backchannel types on the basis of the speaker's multimodal information in dialogues. First, we construct a dialogue corpus that consists of multimodal information of a speaker's utterances and a listener's backchannels. Second, we construct machine learning models to predict a listener's various backchannel types on the basis of a speaker's multimodal information. Our results suggest that our model was able to predict a listener's various backchannel types on the basis of a speaker's multimodal information.
Toshiki Onishi, Naoki Azuma, Shunichi Kinoshita, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
IVA4
2022 Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura
INTERSPEECH2
2022 Analysis of praising skills focusing on utterance contents
Asahi Ogushi, Toshiki Onishi, Yohei Tahara, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
INTERSPEECH4
2022 Determining most suitable listener backchannel type for speaker's utterance
abstract
A major hurdle in achieving a dialogue system that enables smooth dialogue is to determine how to generate an appropriate response to a user's utterance. Previous research has focused mainly on estimating whether to make an utterance backchannel in response to the user's utterance. We go one step further by examining, for the first time, the relationship between the type of utterance backchannel to be used and intent and type of the speaker's utterance, known as a dialogue act (DA). Specifically, we propose a new method for classifying utterance backchannels into nine types. We also created a corpus consisting of the DAs of speaker utterances and the backchannel types of listener utterances then used it to analyze the relationship between a speaker's and listener's utterances. Our findings clarify that the occurrence frequencies of a listener's backchannel types significantly depend on the DAs of the speaker's utterances. Since the goal of our research is to construct a dialogue system that generates a more natural backchannel, this classification method, which determines certain types of aids from the speaker's DA, will be beneficial to such a system.
Akira Morikawa, Ryo Ishii, Hajime Noto, Atsushi Fukayama, Takao Nakamura
IVA2
2022 A Comparison of Praising Skills in Face-to-Face and Remote Dialogues
abstract
Praising behavior is considered to an important method of communication in daily life and social activities. An engineering analysis of praising behavior is therefore valuable. However, a dialogue corpus for this analysis has not yet been developed. Therefore, we develop corpuses for face-to-face and remote two-party dialogues with ratings of praising skills. The corpuses enable us to clarify how to use verbal and nonverbal behaviors for successfully praise. In this paper, we analyze the differences between the face-to-face and remote corpuses, in particular the expressions in adjudged praising scenes in both corpuses, and also evaluated praising skills. We also compare differences in head motion, gaze behavior, facial expression in high-rated praising scenes in both corpuses. The results showed that the distribution of praising scores was similar in face-to-face and remote dialogues, although the ratio of the number of praising scenes to the number of utterances was different. In addition, we confirmed differences in praising behavior in face-to-face and remote dialogues.
Toshiki Onishi, Asahi Ogushi, Yohei Tahara, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
LREC4
2021 Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile Data
abstract
Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nick Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nicholas B. Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL/IJCNLP (1)5
2021 How People Distinguish Individuals from their Movements: Toward the Realization of Personalized Agents
abstract
Demands for agents that replicate the characteristics of specific individuals are increasing. Although ways to implement personality traits into the virtual agents’ movement have been widely researched, ways to create aspects of individuality that can be identified as belonging to specific individuals have not. To clarify how well humans can identify individuals from short movements and what elements of movement contribute to the perception of individuality, we examined the relationship between the degree of confidence in personal identification and statics of gesture movement. In the experiment, participants were asked to compare pairs of short presentation animations and give their degree of confidence that the two animations were of the same person. The animations were created with motion data from performers and with 3D-CG characters to reduce the differences in appearances and shot angles. We calculated five expressivity parameters from the wrist movement for each gesture and compared the answers from the participants. The results showed that the participants were able to distinguish individuals doing the same action and recognize the individuals by the spatial and temporal extents of their movement, which were represented by how much space they use and how fast they moved their wrists. This study clarifies the cognitive aspects of what elements need to be reproduced to develop agents with individuality.
Chihiro Takayama, Mitsuhiro Goto, Shin'ichiro Eitoku, Ryo Ishii, Hajime Noto, Shiro Ozawa, Takao Nakamura
HAI4
2021 Multimodal and Multitask Approach to Listener's Backchannel Prediction: Can Prediction of Turn-changing and Turn-management Willingness Improve Backchannel Modeling?
abstract
The listener's backchannel has the important function of encouraging a current speaker to hold their turn and continue to speak, which enables smooth conversation. The listener monitors the speaker's turn-management (a.k.a. speaking and listening) willingness and his/her own willingness to display backchannel behavior. Many studies have focused on predicting the appropriate timing of the backchannel so that conversational agents can display backchannel behavior in response to a user who is speaking. To the best of our knowledge, none of them added the prediction of turn-changing and participants' turn-management willingness to the backchannel prediction model in dyad interactions. In this paper, we proposed a novel backchannel prediction model that can jointly predict turn-changing and turn-management willingness. We investigated the impact of modeling turn-changing and willingness to improve backchannel prediction. Our proposed model is based on trimodal inputs, that is, acoustic, linguistic, and visual cues from conversations. Our results suggest that adding turn-management willingness as a prediction task improves the performance of backchannel prediction within the multi-modal multi-task learning approach, while adding turn-changing prediction is not useful for improving the performance of backchannel prediction.
Ryo Ishii, Xutong Ren, Michal Muszynski, Louis-Philippe Morency
IVA1
2020 Analyzing Nonverbal Behaviors along with Praising
abstract
In this work, as a first attempt to analyze the relationship between praising skills and human behavior in dialogue, we focus on head and face behavior. We create a new dialogue corpus including face and head behavior information of persons who give praise (praiser) and receive praise (receiver) and the degree of success of praising (praising score). We also create a machine learning model that uses features related to head and face behavior to estimate praising score, clarify which features of the praiser and receiver are important in estimating praising score. The analysis results showed that features of the praiser and receiver are important in estimating praising score and that features related to utterance, head, gaze, and chin were important. The analysis of the features of high importance revealed that the praiser and receiver should face each other without turning their heads to the left or right, and the longer the praiser's utterance, the more successful the praising.
Toshiki Onishi, Arisa Yamauchi, Ryo Ishii, Yushi Aono, Akihiro Miyata
ICMI3
2020 Impact of Personality on Nonverbal Behavior Generation
abstract
To realize natural-looking virtual agents, one key technical challenge is to automatically generate nonverbal behaviors from spoken language. Since nonverbal behavior varies depending on personality, it is important to generate these nonverbal behaviors to match the expected personality of a virtual agent. In this work, we study how personality traits relate to the process of generating individual nonverbal behaviors from the whole body, including the head, eye gaze, arms, and posture. To study this, we first created a dialogue corpus including transcripts, a broad range of labelled nonverbal behaviors, and the Big Five personality scores of participants in dyad interactions. We constructed models that can predict each nonverbal behavior label given as an input language representation from the participants' spoken sentences. Our experimental results show that personality can help improve the prediction of nonverbal behaviors.
Ryo Ishii, Chaitanya Ahuja, Yukiko I. Nakano, Louis-Philippe Morency
IVA1
2020 Can Prediction of Turn-management Willingness Improve Turn-changing Modeling?
abstract
For smooth conversation, participants must carefully monitor the turn-management (a.k.a. speaking and listening) willingness of other conversational partners and adjust turn-changing behaviors accordingly. Many studies have focused on predicting the actual moments of speaker changes (a.k.a. turn-changing), but to the best of our knowledge, none of them explicitly modeled the turn-management willingness from both speakers and listeners in dyad interactions. We address the problem of building models for predicting this willingness of both. Our models are based on trimodal inputs, including acoustic, linguistic, and visual cues from conversations. We also study the impact of modeling willingness to help improve the task of turn-changing prediction. We introduce a dyadic conversation corpus with annotated scores of speaker/listener turn-management willingness. Our results show that using all of three modalities of speaker and listener is important for predicting turn-management willingness. Furthermore, explicitly adding willingness as a prediction task improves the performance of turn-changing prediction. Also, turn-management willingness prediction becomes more accurate with this multi-task learning approach.
Ryo Ishii, Xutong Ren, Michal Muszynski, Louis-Philippe Morency
IVA1
2019 Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text Data
abstract
This paper presents a novel training method for speech-based end-of-turn detection for which not only manually annotated speech data sets but also punctuated text data sets are utilized. The speech-based end-of-turn detection estimates whether a target speaker's utterance is ended or not using speech information. In previous studies, the speech-based end-of-turn detection models were trained using only speech data sets that contained manually annotated end-of-turn labels. However, since the amounts of annotated speech data sets are often limited, the end-of-turn detection models were unable to correctly handle a wide variety of speech patterns. In order to mitigate the data scarcity problem, our key idea is to leverage punctuated text data sets for building more effective speech-based end-of-turn detection. Therefore, the proposed method introduces cross-modal representation learning to construct a speech encoder and a text encoder that can map speech and text with the same lexical information into similar vector representations. This enables us to train speech-based end-of-turn detection models from the punctuated text data sets by tackling text-based sentence boundary detection. In experiments on contact center calls, we show that speech-based end-of-turn detection models using hierarchical recurrent neural networks can be improved through the use of punctuated text data sets.
Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Takanobu Oba, Ryuichiro Higashinaka
ASRU5
2019 Determining Iconic Gesture Forms based on Entity Image Representation
abstract
Iconic gestures are used to depict physical objects mentioned in speech, and the gesture form is assumed to be based on the image of a given object in the speaker’s mind. Using this idea, this study proposes a model that learns iconic gesture forms from an image representation obtained from pictures of physical entities. First, we collect a set of pictures of each entity from the web, and create an average image representation from them. Subsequently, the average image representation is fed to a fully connected neural network to decide the gesture form. In the model evaluation experiment, our two-step gesture form selection method can classify seven types of gesture forms with over 62% accuracy. Furthermore, we demonstrate an example of gesture generation in a virtual agent system in which our model is used to create a gesture dictionary that assigns a gesture form for each entry word in the dictionary.
Fumio Nihei, Yukiko I. Nakano, Ryuichiro Higashinaka, Ryo Ishii
ICMI4
2018 Where Should Robots Talk?: Spatial Arrangement Study from a Participant Workload Perspective
abstract
Several benefits obtained using multiple robots in conversation have been reported in the human-robot interaction field. This paper first presents pre-trial results by which elderly people assigned a lower rating to a conversation with two robots than to one with a single robot. Observations of the trial suggest the hypothesis that an inappropriate spatial arrangement between robots and humans increases the workload in a conversation. Reducing the workload is important, especially when robots are used by elderly people. Therefore, we specifically examine the workload that is influenced by the spatial arrangement in group conversation. To verify the hypothesis, we use a NASA-TLX and a dual-task method to evaluate the workload and to conduct a comparative experiment in which the participant talks with two robots in two spatial arrangements. We also conduct a case study for elderly people in the same conversational conditions. From these experiments, we demonstrate that the spatial arrangement in which people cannot see both robots simultaneously increases their conversational workload and decreases their evaluation of the dialogue compared to a spatial arrangement by which people can see both robots simultaneously. We also show that the primary cause of the workload by positioning is not physical but mental.
Takahiro Matsumoto, Mitsuhiro Goto, Ryo Ishii, Tomoki Watanabe, Tomohiro Yamada, Michita Imai
HRI3
2018 Analyzing Gaze Behavior and Dialogue Act during Turn-taking for Estimating Empathy Skill Level
abstract
We explored the gaze behavior towards the end of utterances and dialogue act (DA), i.e., verbal-behavior information indicating the intension of an utterance, during turn-keeping/changing to estimate empathy skill levels in multiparty discussions. This is the first attempt to explore the relationship between such a combination. First, we collected data on Davis' Interpersonal Reactivity Index (which measures empathy skill level), utterances that include the DA categories of Provision, Self-disclosure, Empathy, Turn-yielding, and Others, and gaze behavior from participants in four-person discussions. The results of analysis indicate that the gaze behavior accompanying utterances that include these DA categories during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result was that speakers with low empathy skill levels tend to avoid making eye contact with the listener when the DA category is Self-disclosure during turn-keeping. However, they tend to maintain eye contact when the DA category is Empathy. A listener who has a high empathy skill level often looks away from the speaker during turn-changing when the DA category of a speaker's utterance is Provision or Empathy. There was also no difference in gaze behavior between empathy skill levels when the DA category of the speaker's utterance was turn-yielding. From these findings, we constructed and evaluated models for estimating empathy skill level using gaze behavior and DA information. The evaluation results indicate that using both gaze behavior and DA during turn-keeping/changing is effective for estimating an individual's empathy skill level in multi-party discussions.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Ryuichiro Higashinaka, Junji Tomita
ICMI1
2018 Generating Body Motions using Spoken Language in Dialogue
abstract
We propose a model to automatically generate whole body motions accompanying utterances at appropriate times, similar to humans, by using various types of natural-language-analysis information obtained from spoken language. Specifically, we focus on the co-occurrence relationship between various types of natural-language-analysis information such as words included in the spoken language, parts of speech, a thesaurus, word positions, dialogue acts of the spoken language, and human motions. Our model automatically generates nods, head postures, facial expressions, hand gestures, and upper-body posture using such information. We first recorded a two-person dialogue and constructed a multimodal corpus including utterance and whole body motion information. Next, using the constructed corpus, we constructed our model for generating a motion for each phrase unit using machine learning and using words, parts of speech, a thesaurus, word positions, and speech acts of the entire spoken language as inputs. These types of natural-language-analysis information were useful for motion generation. The effectiveness of our model was verified through a subjective experiment using a virtual conversational agent. As a result, the agent's body motions and impressions regarding naturalness of motion, degree of coincidence between utterance and motion, humanness of the agent, and likability of the agent improved with our model.
Ryo Ishii, Taichi Katayama, Ryuichiro Higashinaka, Junji Tomita
IVA1
2018 Automatic Generation System of Virtual Agent's Motion using Natural Language
abstract
A virtual agent in a dialogue system should express appropriate body motions according to utterances and effectively communicate with a user. We previously proposed a generation model of whole body motions such as head direction, nodding, facial expressions, hand gestures, and upper-body posture accompanying utterances at appropriate times similar to humans by using various types of natural-language-analysis information obtained from spoken language. As an attempt to promote this model, we constructed an API that can easily generate motions by using the generation model and constructed a demonstration system that can automatically control a virtual agent from only the spoken language. When inputting an arbitrary utterance language, synthesized sound and motion information are acquired from the speech synthesizer and motion-generation API, and the vocalization of the virtual agent and animated motion are generated. A dialog agent that is more attractive and can communicate smoothly by automatically generating natural motions is expected.
Ryo Ishii, Taichi Katayama, Ryuichiro Higashinaka, Junji Tomita
IVA1
2018 Predicting Nods by using Dialogue Acts in Dialogue
Ryo Ishii, Ryuichiro Higashinaka, Junji Tomita
LREC1
2018 Automatic Generation of Head Nods using Utterance Texts
abstract
We propose a model to generate head nods accompanying an utterance from natural language. To the best of our knowledge, previous models generated simple nods from the final words at the end of an utterance, i.e., using bag of words. We focused on various text analyzed using various types of language information such as dialog act, part of speech, a large-scale Japanese thesaurus, and word position in a sentence. We also generated detailed parameters of speaker's nodding presence, frequency, and depth, which was the first attempt to do so. First, we compiled a Japanese corpus of 24 dialogues including utterance and nod information. Next, using the corpus, we constructed our generation model that estimates nodding presence, frequency, and depth, during a phrase by using such various types of language information as well as bag of words. The results indicate that our model outperformed simple automatic nod-generating models using only bag of words and chance level. The results also indicate that dialog act, part of speech, the large-scale Japanese thesaurus, and word position are useful for generating nods. We also evaluated, through subjective evaluation, if our nod-generation model is useful with conversational agents. The results show that the nodding generated with our model improves user impressions of naturalness, humanness, likability, and reliability toward a conversational agent.
Ryo Ishii, Taichi Katayama, Ryuichiro Higashinaka, Junji Tomita
RO-MAN1
2018 Neural Dialogue Context Online End-of-Turn Detection
abstract
This paper proposes a fully neural network based dialogue-context online end-of-turn detection method that can utilize longrange interactive information extracted from both target speaker's and interlocutor's utterances.In the proposed method, we combine multiple time-asynchronous long short-term memory recurrent neural networks, which can capture target speaker's and interlocutor's multiple sequential features, and their interactions.On the assumption of applying the proposed method to spoken dialogue systems, we introduce target speaker's acoustic sequential features and interlocutor's linguistic sequential features, each of which can be extracted in an online manner.Our evaluation confirms the effectiveness of taking dialogue context formed by the target speaker's utterances and interlocutor's utterances into consideration.
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, Yushi Aono
SIGDIAL Conference4
2017 Computational model of idiosyncratic perception of others' emotions
abstract
This paper deals with computational modelling for predicting the idiosyncratic perception of others' emotions, namely how individual external observers will score the emotional states of others interacting with each other. We separately model the observer effect (or individual differences of observers), and the conversational-scene effect or the video-clip effect (how interlocutors are interacting), based on Bayes' theorem with the assumption of their conditional independence. The observer term describes the observer's cognitive tendency, including bias, in a probabilistic form, and does not include any clip information. In contrast, the clip term describes how a target clip is recognized by an unspecified observer. The perceived emotion is predicted to be the state that maximizes the conditional probability given the observer and target clip. An experiment with 100 observers and 97 clips demonstrated, in a leave-one-out cross-validation scenario, that 1) there is in fact no statistically and practically significant interaction between observer and clip, and 2) our Bayesian modelling achieves a 97 percent accuracy as a reference of test-retest reliability. Furthermore, when combined with existing observer and clip models that can handle unknown observers and clips, our model yielded an accuracy of around 50 percent in a more challenging leave-one-subject-and-clip-out cross-validation scenario.
Shiro Kumano, Ryo Ishii, Kazuhiro Otsuka
ACII2
2017 Comparing empathy perceived by interlocutors in multiparty conversation and external observers
abstract
This paper investigates the basic characteristics of perceived empathy in Breithaupt's three-person model to consider a way of realizing its automatic prediction or empathy reading machines. More specifically, we report the extent to which interlocutors differ from external observers in perceiving the empathy aroused during group conversation. We also evaluate the accuracies of various frequently used models, including majority voting and multiple regression, in predicting an interlocutor's ratings from those of other interlocutors and/or those of the observers. Defining empathy as the emotional congruence between pairs of interlocutors, we studied a four-person conversation, in which previously unacquainted people held a decision-making discussion. We used a 5-point Likert scale when collecting self-reports of empathy from the interlocutors, and reports from a total of forty external observers (ten for each interlocutor) who adopted a target interlocutor's perspective. We obtained three indications. First, when no empathy ratings are available from the target interlocutor for model training, it is beneficial to ask observers to take the target interlocutor's perspective. Second, when target interlocutors' self-reports are available, it is advantageous to instruct observers not to take the target interlocutor's perspective. Third, in both scenarios, it is useful to ask interlocutors to rate the pairs excluding themselves. These findings provide some insights into good rating procedure as regards studying perceived empathy.
Shiro Kumano, Ryo Ishii, Kazuhiro Otsuka
ACII2
2017 Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings
abstract
To build a conversational interface wherein an agent system can smoothly communicate with multiple persons, it is imperative to know how the timing of speaking is decided. In this research, we explore the head movements of participants as an easy-to-measure nonverbal behavior to predict the nest-utterance timing, i.e., the interval between the end of the current speaker's utterance and the start of the next speaker's utterance, in turn-changing in multi-party meetings. First, we collected data on participants' six degree-of-freedom head movements and utterances in four-person meetings. The results of the analysis revealed that the amount of head movements of current speaker, next speaker, and listeners have a positive correlation with the utterance interval. Moreover, the degree of synchrony of the head position and posture between the current speaker and next speaker is negatively correlated with the utterance interval. On the basis of these findings, we used their head movements and the synchrony of their head movements as feature values and devised several prediction models. A model using all features performed the best and was able to predict the next-utterance timing well. Therefore, this research revealed that the participants' head movement is useful for predicting the next-utterance timing in turn-changing in multi-party meetings.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
HAI1
2017 Analyzing gaze behavior during turn-taking for estimating empathy skill level
abstract
Techniques that use nonverbal behaviors to estimate communication skill in discussions have been receiving a lot of attention in recent research. In this study, we explored the gaze behavior towards the end of an utterance during turn-keeping/changing to estimate empathy skills in multiparty discussions. First, we collected data on Davis' Interpersonal Reactivity Index (IRI) (which measures empathy skill), utterances, and gaze behavior from participants in four-person discussions. The results of the analysis showed that the gaze behavior during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result is that the amount of a person's empathy skill is inversely proportional to the frequency of eye contact with the conversational partner during turn-keeping/changing. Specifically, if the current speaker has a high skill level, she often does not look at listener during turn-keeping and turn-changing. Moreover, when a person with a high skill level is the next speaker, she does not look at the speaker during turn-changing. In contrast, people who have a low skill level often continue to make eye contact with speakers and listeners. On the basis of these findings, we constructed and evaluated four models for estimating empathy skill levels. The evaluation results showed that the average absolute error of estimation is only 0.22 for the gaze transition pattern (GTP) model. This model uses the occurrence probability of GTPs when the person is a speaker and listener during turn-keeping and speaker, next-speaker, and listener during turn-changing. It outperformed the models that used the amount of utterances and duration of gazes. This suggests that the GTP during turn-keeping and turn-changing is effective for estimating an individual's empathy skills in multi-party discussions.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICMI1
2017 Online End-of-Turn Detection from Speech Based on Stacked Time-Asynchronous Sequential Networks
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Ryo Ishii, Ryuichiro Higashinaka
INTERSPEECH4
2017 Collective First-Person Vision for Automatic Gaze Analysis in Multiparty Conversations
abstract
This paper targets smallto medium-sized-group face-to-face conversations where each person wears a dual-view camera, consisting of inwardand outward-looking cameras, and presents an almost fully automatic but accurate ofline gaze analysis framework that does not require users to perform any calibration steps. Our collective first-person vision framework, where captured audio-visual signals are gathered and processed in a centralized system, jointly undertakes the fundamental functions required for group gaze analysis, including speaker detection, face tracking, and gaze tracking. Of particular note is our self-calibration of gaze trackers by exploiting a general conversation rule, namely that listeners are likely to look at the speaker. From the rough conversational prior knowledge, our system visualizes fine-grained participants' gaze behavior as a gazee-centered heat map, which quantitatively reveals what parts of the gazee's body the participant looked at and for how long while the gazer was speaking or listening. An experiment using conversations amounting to a total of 140 min, each lasting an average of 8.7 min and engaged in by 37 participants in groups of three to six, achieves a mean absolute error of 2.8° in gaze tracking. A statistical test reveals neither a group size effect nor a conversation type effect. Our method achieves F-scores of over 0.89 and 0.87 in gazee and eye contact recognition, respectively, in comparison with human annotation.
Shiro Kumano, Kazuhiro Otsuka, Ryo Ishii, Junji Yamato
IEEE Trans. Multim.3
2016 Analyzing mouth-opening transition pattern for predicting next speaker in multi-party meetings
abstract
Techniques that use nonverbal behaviors to predict turn-changing situations—e.g., predicting who will speak next and when, in multi-party meetings—have been receiving a lot of attention in recent research. In this research, we explored the transition pattern of the degree of mouth opening (MOTP) towards the end of an utterance to predict the next speaker in multiparty meetings. First, we collected data on utterances and on the degree of mouth opening (closed, slightly open, and wide open) from participants in four-person meetings. The representative results of the analysis of the MOTPs showed that the speaker often continues to open the mouth slightly in turn-keeping and starts to close the mouth from opening it slightly or continues to open the mouth largely in turn-changing. The next speaker often starts to open the mouth slightly from closing it in turn-changing. On the basis of these findings, we constructed next-speaker prediction models using the MOTPs. In addition, as a multimodal fusion, we constructed models using the MOTPs and gaze information, which is known to be one of the most useful types of information for the next-speaker prediction. The evaluation of the models suggests that the speaker's and listeners' MOTPs are effective for predicting the next speaker in multi-party meetings. It also suggests the multimodal fusion using the MOTP and gaze information is more useful for the prediction than using one or the other.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICMI1
2016 Prediction of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty Meetings
abstract
In multiparty meetings, participants need to predict the end of the speaker’s utterance and who will start speaking next, as well as consider a strategy for good timing to speak next. Gaze behavior plays an important role in smooth turn-changing. This article proposes a prediction model that features three processing steps to predict (I) whether turn-changing or turn-keeping will occur, (II) who will be the next speaker in turn-changing, and (III) the timing of the start of the next speaker’s utterance. For the feature values of the model, we focused on gaze transition patterns and the timing structure of eye contact between a speaker and a listener near the end of the speaker’s utterance. Gaze transition patterns provide information about the order in which gaze behavior changes. The timing structure of eye contact is defined as who looks at whom and who looks away first, the speaker or listener, when eye contact between the speaker and a listener occurs. We collected corpus data of multiparty meetings, using the data to demonstrate relationships between gaze transition patterns and timing structure and situations (I), (II), and (III). The results of our analyses indicate that the gaze transition pattern of the speaker and listener and the timing structure of eye contact have a strong association with turn-changing, the next speaker in turn-changing, and the start time of the next utterance. On the basis of the results, we constructed prediction models using the gaze transition patterns and timing structure. The gaze transition patterns were found to be useful in predicting turn-changing, the next speaker in turn-changing, and the start time of the next utterance. Contrary to expectations, we did not find that the timing structure is useful for predicting the next speaker and the start time. This study opens up new possibilities for predicting the next speaker and the timing of the next utterance using gaze transition patterns in multiparty meetings.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ACM Trans. Interact. Intell. Syst.1
2016 Using Respiration to Predict Who Will Speak Next and When in Multiparty Meetings
abstract
Techniques that use nonverbal behaviors to predict turn-changing situations—such as, in multiparty meetings, who the next speaker will be and when the next utterance will occur—have been receiving a lot of attention in recent research. To build a model for predicting these behaviors we conducted a research study to determine whether respiration could be effectively used as a basis for the prediction. Results of analyses of utterance and respiration data collected from participants in multiparty meetings reveal that the speaker takes a breath more quickly and deeply after the end of an utterance in turn-keeping than in turn-changing. They also indicate that the listener who will be the next speaker takes a bigger breath more quickly and deeply in turn-changing than the other listeners. On the basis of these results, we constructed and evaluated models for predicting the next speaker and the time of the next utterance in multiparty meetings. The results of the evaluation suggest that the characteristics of the speaker's inhalation right after an utterance unit—the points in time at which the inhalation starts and ends after the end of the utterance unit and the amplitude, slope, and duration of the inhalation phase—are effective for predicting the next speaker in multiparty meetings. They further suggest that the characteristics of listeners' inhalation—the points in time at which the inhalation starts and ends after the end of the utterance unit and the minimum and maximum inspiration, amplitude, and slope of the inhalation phase—are effective for predicting the next speaker. The start time and end time of the next speaker's inhalation are also useful for predicting the time of the next utterance in turn-changing.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ACM Trans. Interact. Intell. Syst.1
2015 Predicting next speaker based on head movement in multi-party meetings
abstract
We proposed a model for predicting the next speaker in multi-party meetings by focusing on the participants' head movements measured by using a six degrees-of-freedom head tracker. Results of an analysis of head movements collected from multi-party meetings revealed differences in the amounts, amplitude, and frequency of movement of the head position and rotation of the speaker near the end of an utterance in turn-keeping and turn-taking. The results also revealed the differences in the amounts of movement, amplitude, and frequency of head position movement and rotation between the listeners in turn-keeping, turn-taking, and the next speaker in turn-taking. We then built a next speaker prediction model that features two processing steps to predict whether turn-taking or turn-keeping will occur and who the next speaker will be in turn-taking. The evaluation results for the model suggest that the speaker's and listeners' head movements contribute to predicting the next speaker.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICASSP1
2015 Multimodal Fusion using Respiration and Gaze for Predicting Next Speaker in Multi-Party Meetings
abstract
Techniques that use nonverbal behaviors to predict turn-taking situations, such as who will be the next speaker and the next utterance timing in multi-party meetings are receiving a lot of attention recently. It has long been known that gaze is a physical behavior that plays an important role in transferring the speaking turn between humans. Recently, a line of research has focused on the relationship between turn-taking and respiration, a biological signal that conveys information about the intention or preliminary action to start to speak. It has been demonstrated that respiration and gaze behavior separately have the potential to allow predicting the next speaker and the next utterance timing in multi-party meetings. As a multimodal fusion to create models for predicting the next speaker in multi-party meetings, we integrated respiration and gaze behavior, which were extracted from different modalities and are completely different in quality, and implemented a model uses information about them to predict the next speaker at the end of an utterance. The model has a two-step processing. The first is to predict whether turn-keeping or turn-taking happens; the second is to predict the next speaker in turn-taking. We constructed prediction models with either respiration or gaze behavior and with both respiration and gaze behaviors as features and compared their performance. The results suggest that the model with both respiration and gaze behaviors performs better than the one using only respiration or gaze behavior. It is revealed that multimodal fusion using respiration and gaze behavior is effective for predicting the next speaker in multi-party meetings. It was found that gaze behavior is more useful for predicting turn-keeping/turn-taking than respiration and that respiration is more useful for predicting the next speaker in turn-taking.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICMI1
2015 Design and Evaluation of Mirror Interface MIOSS to Overlay Remote 3D Spaces
Ryo Ishii, Shiro Ozawa, Akira Kojima, Kazuhiro Otsuka, Yuki Hayashi, Yukiko I. Nakano
INTERACT (4)1
2014 Analysis and modeling of next speaking start timing based on gaze behavior in multi-party meetings
abstract
To realize a conversational interface where an agent system can smoothly communicate with multiple persons, it is imperative to know how the start timing of speaking is decided. In this research, we demonstrate a relationship between gaze transition patterns and the start timing of next speaking against the end of the last speaking in multi-party meetings. Then, we construct a prediction model for the start timing using gaze transition patterns near the end of an utterance. An analysis of data collected from natural multi-party meetings reveals a strong relationship between gaze transition patterns of the speaker, next speaker, and listener and the start timing of the next speaker. On the basis of the results, we used gaze transition patterns of the speaker, next speaker, and listener and mutual gaze as variables, and devised several prediction models. A model using all features performed the best and was able to predict the start timing well.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ICASSP1
2014 Analysis of Respiration for Prediction of "Who Will Be Next Speaker and When?" in Multi-Party Meetings
abstract
To build a model for predicting the next speaker and the start time of the next utterance in multi-party meetings, we performed a fundamental study of how respiration could be effective for the prediction model. The results of the analysis reveal that a speaker inhales more rapidly and quickly right after the end of a unit of utterance in turn-keeping. The next speaker takes a bigger breath toward speaking in turn-changing than listeners who will not become the next speaker. Based on the results of the analysis, we constructed the prediction models to evaluate how effective the parameters are. The results of the evaluation suggest that the speaker's inhalation right after a unit of utterance, such as the start time from the end of the unit of utterance and the slope and duration of the inhalation phase, is effective for predicting whether turn-keeping or turn-changing happen about 350 ms before the start time of the next utterance on average and that listener's inhalation before the next utterance, such as the maximal inspiration and amplitude of the inhalation phase, is effective for predicting the next speaker in turn-changing about 900 ms before the start time of the next utterance on average.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ICMI1
2013 Using a Probabilistic Topic Model to Link Observers' Perception Tendency to Personality
abstract
Targeting multiparty conversations, the present study aims to elucidate how an observer will tend to perceive others' emotional states, develops a computational model that realizes the automatic inferencing of the observer's perception tendency. This paper proposes a probabilistic model that automatically discovers the correlation between perception tendency, gender, and personality traits of a target observer. Perception tendency, a probability distribution, explains how likely the observer is to perceive a certain state/level of a target emotion. Personality traits are measured by a variety of questionnaires. The proposed model links these three factors via a latent variable and explains observer's characteristics as a mixture of prototypical characters. An experiment is conducted with fifty observers. They watch 97 short conversation videos and give their impressions about the empathy between each interacting pair. The results demonstrate that the proposed method can find a reasonable framework that underlies the factors: e.g. 1) people who have high scores in Davis's empathy measures show empathy-biased response tendency, and 2) people who have strong sense of consideration for others tend to show an extreme response tendency, and such people are likely to be females. The proposed method shows promise in estimating an observer's perception tendency from his/her gender and personality traits, even when the target perception tendency is quite different from the average perception tendency among observers.
Shiro Kumano, Kazuhiro Otsuka, Masafumi Matsuda, Ryo Ishii, Junji Yamato
ACII4
2013 Predicting next speaker and timing from gaze transition patterns in multi-party meetings
abstract
In multi-party meetings, participants need to predict the end of the speaker's utterance and who will start speaking next, and to consider a strategy for good timing to speak next. Gaze behavior plays an important role for smooth turn-taking. This paper proposes a mathematical prediction model that features three processing steps to predict (I) whether turn-taking or turn-keeping will occur, (II) who will be the next speaker in turn-taking, and (III) the timing of the start of the next speaker's utterance. For the feature quantity of the model, we focused on gaze transition patterns near the end of utterance. We collected corpus data of multi party meetings and analyzed how the frequencies of appearance of gaze transition patterns differs depending on situations of (I), (II), and (III). On the basis of the analysis, we construct a probabilistic mathematical model that uses the frequencies of appearance of all participants' gaze transition patterns. The results of an evaluation of the model show the proposed models succeed with high precision compared to ones that do not take gaze transition patterns into account.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Masafumi Matsuda, Junji Yamato
ICMI1
2013 MM+Space: n x 4 degree-of-freedom kinetic display for recreating multiparty conversation spaces
abstract
A novel system, called MM+Space, is presented for recreating multiparty face-to-face conversation scenes in the real world. It aims to display and playback pre-recorded conversations as if the people were talking in front of the viewer(s). This system consists of multiple projectors and transparent screens, which display the life-size faces of people. The key idea is the physical augmentation of human head motions, i.e. the screen pose is dynamically controlled to emulate the head motions, for boosting the viewers' perception of nonverbal behaviors and interactions. In particular, MM+Space newly introduces 2-Degree-of-Freedom (DoF) translations, in forward-backward and right-left directions, in addition to 2-DoF head rotations (nodding and shaking), which were proposed in our former MM-Space system. The full 4-DoF kinetic display is expected to enhance the expressibility of head and body motions, and to create more realistic representation of interacting people. Experiments showed that the proposed system with 4-DoF motions outperformed the rotation-only system in the increased perception of people's presence and in expressing their postures. In addition, it was reported that the proposed system allowed the viewers to experience rich emotional expressibility, immersion in conversations, and potential behavioral/emotional contagion.
Kazuhiro Otsuka, Shiro Kumano, Ryo Ishii, Maja Zbogar, Junji Yamato
ICMI3
2013 Gaze awareness in conversational agents: Estimating a user's conversational engagement from eye gaze
abstract
In face-to-face conversations, speakers are continuously checking whether the listener is engaged in the conversation, and they change their conversational strategy if the listener is not fully engaged. With the goal of building a conversational agent that can adaptively control conversations, in this study we analyze listener gaze behaviors and develop a method for estimating whether a listener is engaged in the conversation on the basis of these behaviors. First, we conduct a Wizard-of-Oz study to collect information on a user's gaze behaviors. We then investigate how conversational disengagement, as annotated by human judges, correlates with gaze transition, mutual gaze (eye contact) occurrence, gaze duration, and eye movement distance. On the basis of the results of these analyses, we identify useful information for estimating a user's disengagement and establish an engagement estimation method using a decision tree technique. The results of these analyses show that a model using the features of gaze transition, mutual gaze occurrence, gaze duration, and eye movement distance provides the best performance and can estimate the user's conversational engagement accurately. The estimation model is then implemented as a real-time disengagement judgment mechanism and incorporated into a multimodal dialog manager in an animated conversational agent. This agent is designed to estimate the user's conversational engagement and generate probing questions when the user is distracted from the conversation. Finally, we evaluate the engagement-sensitive agent and find that asking probing questions at the proper times has the expected effects on the user's verbal/nonverbal behaviors during communication with the agent. We also find that our agent system improves the user's impression of the agent in terms of its engagement awareness, behavior appropriateness, conversation smoothness, favorability, and intelligence.
Ryo Ishii, Yukiko I. Nakano, Toyoaki Nishida
ACM Trans. Interact. Intell. Syst.1
2011 Estimating a User's Conversational Engagement Based on Head Pose Information
Ryota Ooko, Ryo Ishii, Yukiko I. Nakano
IVA2
2010 Estimating user's engagement from eye-gaze behaviors in human-agent conversations
abstract
In face-to-face conversations, speakers are continuously checking whether the listener is engaged in the conversation and change the conversational strategy if the listener is not fully engaged in the conversation. With the goal of building a conversational agent that can adaptively control conversations with the user, this study analyzes the user's gaze behaviors and proposes a method for estimating whether the user is engaged in the conversation based on gaze transition 3-gram patterns. First, we conduct a Wizard-of-Oz experiment to collect the user's gaze behaviors. Based on the analysis of the gaze data, we propose an engagement estimation method that detects the user's disengagement gaze patterns. The algorithm is implemented as a real-time engagement-judgment mechanism and is incorporated into a multimodal dialogue manager in a conversational agent. The agent estimates the user's conversational engagement and generates probing questions when the user is distracted from the conversation. Finally, we conduct an evaluation experiment using the proposed engagement-sensitive agent and demonstrate that the engagement estimation function improves the user's impression of the agent and the interaction with the agent. In addition, probing performed with proper timing was also found to have a positive effect on user's verbal/nonverbal behaviors in communication with the conversational agent.
Yukiko I. Nakano, Ryo Ishii
IUI2
2008 Estimating User's Conversational Engagement Based on Gaze Behaviors
Ryo Ishii, Yukiko I. Nakano
IVA1
2006 Avatar's Gaze Control to Facilitate Conversational Turn-Taking in Virtual-Space Multi-user Voice Chat System
Ryo Ishii, Toshimitsu Miyajima, Kinya Fujita, Yukiko I. Nakano
IVA1