VLDB 2026 Research / reviewers in the wild / expert
Yukiko I. Nakano
dblp:29/5823
· DBLP profile ↗
63ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0003-1658-8219ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 43 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 39 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Impact of Cultural Differences and Politeness on Joining Small Groups of Humans, Robots, and Virtual CharactersabstractThis cross-cultural study$(\mathrm{N}=108)$examines how cultural differences between Japan and Sweden influence participants social behaviors and perceptions when joining a free-standing group of two agents. Agents within the group, embodied as humans, robots, and virtual characters, respectively, use three distinct behaviors, varying with respect to politeness strategy, to request the participant to join on a specific side and position in the group. The experimental results showed that Japanese participants, from a culture characterized by higher power distance, masculinity, uncertainty avoidance, long-term orientation, restraint, and collectivism, were more likely to comply with the agent's request regarding the joining position, compared to Swedish participants. This trend was even more pronounced when comparing different types of embodiment: Japanese participants more strictly complied with human agents than with non-human agents. Additionally, Japanese females and Swedish males adhered more to social norms by avoiding walking between group members (i.e. through the group's o-space) when joining. Second, cultural differences also significantly impacted the perception of agents' politeness behaviors, while the effect of embodiment on feelings of friendliness and closeness varied depending on the culture. We reflect on our results as a basis for highlighting key challenges involved in the design of culturally adapted agents and their behaviors toward enhancing the localization of human-agent interaction. Sahba Zojaji, Yukiko I. Nakano, Christopher Peters 0001 |
HRI | 2 |
| 2024 | Participation Role-Driven Engagement Estimation of ASD Individuals in Neurodiverse Group DiscussionsabstractAdults with autism spectrum disorder (ASD) face difficulties in communicating with neurotypical people in their daily lives and workplaces. In addition, research on modeling communication in neurodiverse groups is scarce. To recognize communication difficulties caused by neurodiversity, we first, collected a multimodal corpus for decision-making discussions in neurodiverse groups that included a person with ASD and two neurotypical participants. For corpus analysis, we investigated eye-gaze and facial expression exchanges between individuals with ASD and neurotypical participants during both listening and speaking. The findings were extended to automatically estimate the engagement of ASD individuals. To capture the effect of contingent behaviors between ASD individuals and neurotypical participants, we developed a transformer-based model that considers the participation role by changing the direction of cross-person attention depending on whether the ASD individual is listening or speaking. The proposed approach yields comparable results to the state-of-the-art for engagement estimation in neurotypical group conversations while accounting for the dynamic nature of behavior influence in face-to-face interactions. The code associated with this study is available at https://github.com/IUI-Lab/switch-attention. Kalin Stefanov, Yukiko I. Nakano, Chisa Kobayashi, Ibuki Hoshina, Tatsuya Sakato, Fumio Nihei, Chihiro Takayama, Ryo Ishii, Masatsugu Tsujii |
ICMI | 2 |
| 2024 | Modifying Gesture Style with Impression WordsabstractWhen people form impressions of others in face-to-face communication, gesture style (i.e. the way of gesturing) impacts their impressions, such as being well-mannered, honest, and enthusiastic. As a mechanism for changing the gesture style, we trained a GAN-based style transfer model using a collection of video clips of speakers. Then, we collected a new speaker dataset from YouTube videos representing three different countries, and applied them to the style encoder of our style transfer model and created a gesture style latent space. However, it is difficult to select an appropriate style from a large number of style candidates to synthesize motions that would give off a specific impression. To assist users with this, we propose a method for automatically selecting an appropriate style by fine-tuning a Large Language Model (LLM) that uses a list of impression words as input. An evaluation study found that the gesture transfer model effectively changes the impression of gesture, and styles selected by the style selection mechanism produced motions that express similar impression to those that applied ground truth styles, compared to randomly selected styles. Yoshiki Takahashi, Yukiko I. Nakano, Tatsuya Sakato, Hannes Högni Vilhjálmsson |
IVA | 3 |
| 2023 | Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance EstimationabstractThis study investigated differences in the contributions of various features to in-person (IP) and video-mediated (VM) meetings. We focused on estimating important utterances using both an IP and a VM meeting corpora as the analysis data. A transformer model with dialogue history was used to estimate important utterances, and five types of input (text, speaker’s audio, others’ audio, speaker’s video, and others’ video) were fed to the model. A comparison of the models for IP and VM revealed that the speaker’s audio has a strong effect on the IP model, the video of the other participants strongly affects the VM model, and the text and others’ audio strongly affects both models in estimating important utterances. Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Atsushi Fukayama, Takao Nakamura |
ICASSP | 3 |
| 2023 | Question Generation to Elicit Users' Food Preferences Considering the Semantic ContentabstractTo obtain a better understanding of user preferences in providing tailored services, dialogue systems have to generate semi-structured interviews that require flexible dialogue control while following a topic guide to accomplish the purpose of the interview.Toward this goal, this study proposes a semantics-aware GPT-3 fine-tuning model that generates interviews to acquire users' food preferences.The model was trained using dialogue history and semantic representation constructed from the communicative function and semantic content of the utterance.Using two baseline models: zeroshot ChatGPT and fine-tuned GPT-3, we conducted a user study for subjective evaluations alongside automatic objective evaluations.In the user study, in impression rating, the outputs of the proposed model were superior to those of baseline models and comparable to real human interviews in terms of eliciting the interviewees' food preferences. Yukiko I. Nakano, Tatsuya Sakato |
SIGDIAL | 2 |
| 2022 | Detecting Change Talk in Motivational Interviewing using Verbal and Facial InformationabstractBehavior change is one of the most important goals in psychotherapy. This study focuses on Motivational Interviewing (MI), which is collaborative communication aimed at eliciting the client’s own reasons for behavior change. To investigate the effectiveness of facial information in modeling MI, we collected an MI encounter corpus with speech and video data in the nutrition and fitness domains and annotated client utterances using the Manual for the Motivational Interviewing Skill Code (MISC). By analyzing client answers to the questions after the session, we found that clients who expressed more Change Talk were more motivated to change their behavior than those who expressed less Change Talk. We then proposed RNN-based multimodal models to detect Change Talk by setting a 2-class classification task: "Change Talk" and "not Change Talk." Our experiment showed that the best performing model was a multimodal BiLSTM model that fused language and client facial information. We also found that fusing language and facial information as context achieved better performance than the unimodal and no-context models. Moreover, we discuss the label imbalance problem and conduct an additional analysis using turns as a unit of analysis. As a result, our best model reached F1-score of 0.65 for Change Talk detection. Yukiko I. Nakano, Eri Hirose, Tatsuya Sakato, Shogo Okada, Jean-Claude Martin |
ICMI | 1 |
| 2022 | Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura |
INTERSPEECH | 3 |
| 2021 | Web-ECA: A Web-based ECA PlatformabstractRunning an embodied conversation agent (ECA) on a client–server model has the following advantages: 1) the experimenter need not bring a computer with installing the ECA to the experimental site, 2) the user need not own a high-performance computer, 3) it is easy to make changes to the ECA for system maintenance and operation, and 4) data collection using an ECA becomes easier. To realize these benefits, we propose a platform for executing ECA on a server–client model, i.e., a platform for ECA that can be accessed through the Internet. This paper describes the system configuration and an application of this proposed platform. Fumio Nihei, Yukiko I. Nakano |
ICMI | 2 |
| 2021 | Task-independent Recognition of Communication Skills in Group Interaction Using Time-series ModelingabstractCase studies of group discussions are considered an effective way to assess communication skills (CS). This method can help researchers evaluate participants’ engagement with each other in a specific realistic context. In this article, multimodal analysis was performed to estimate CS indices using a three-task-type group discussion dataset, the MATRICS corpus. The current research investigated the effectiveness of engaging both static and time-series modeling, especially in task-independent settings. This investigation aimed to understand three main points: first, the effectiveness of time-series modeling compared to nonsequential modeling; second, multimodal analysis in a task-independent setting; and third, important differences to consider when dealing with task-dependent and task-independent settings, specifically in terms of modalities and prediction models. Several modalities were extracted (e.g., acoustics, speaking turns, linguistic-related movement, dialog tags, head motions, and face feature sets) for inferring the CS indices as a regression task. Three predictive models, including support vector regression (SVR), long short-term memory (LSTM), and an enhanced time-series model (an LSTM model with a combination of static and time-series features), were taken into account in this study. Our evaluation was conducted by using the R 2 score in a cross-validation scheme. The experimental results suggested that time-series modeling can improve the performance of multimodal analysis significantly in the task-dependent setting (with the best R 2 = 0.797 for the total CS index), with word2vec being the most prominent feature. Unfortunately, highly context-related features did not fit well with the task-independent setting. Thus, we propose an enhanced LSTM model for dealing with task-independent settings, and we successfully obtained better performance with the enhanced model than with the conventional SVR and LSTM models (the best R 2 = 0.602 for the total CS index). In other words, our study shows that a particular time-series modeling can outperform traditional nonsequential modeling for automatically estimating the CS indices of a participant in a group discussion with regard to task dependency. Candy Olivia Mawalim, Shogo Okada, Yukiko I. Nakano |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Style Transfer for Co-speech Gesture Animation: A Multi-speaker Conditional-Mixture Approach
Chaitanya Ahuja, Dong Won Lee 0007, Yukiko I. Nakano, Louis-Philippe Morency |
ECCV (18) | 3 |
| 2020 | Estimating the Intensity of Facial Expressions Accompanying Feedback Responses in Multiparty Video-Mediated CommunicationabstractProviding feedback to a speaker is an essential communication signal for maintaining a conversation. In specific feedback, which indicates the listener's reaction to the speaker?s utterances, the facial expression is an effective modality for conveying the listener's reactions. Moreover, not only the type of facial expressions, but also the degree of intensity of the expressions, may influence the meaning of the specific feedback. In this study, we propose a multimodal deep neural network model that predicts the intensity of facial expressions co-occurring with feedback responses. We focus on multiparty video-mediated communication. In video-mediated communication, close-up frontal face images of each participant are continuously presented on the display; the attention of the participants is more likely to be drawn to the facial expressions. We assume that in such communication, the importance of facial expression in the listeners? feedback responses increases. We collected 33 video-mediated conversations by groups of three people and obtained audio and speech data for each participant. Using the corpus collected as a dataset, we created a deep neural network model that predicts the intensity of 17 types of action units (AUs) co-occurring with the feedback responses. The proposed method employed GRU-based model with attention mechanism for audio, visual, and language modalities. A decoder was trained to produce the intensity values for the 17 AUs frame by frame. In the experiment, unimodal and multimodal models were compared in terms of their performance in predicting salient AUs that characterize facial expression in feedback responses. The results suggest that well-performing models differ depending on the AU categories; audio information was useful for predicting AUs that express happiness, and visual and language information contributes to predicting AUs expressing sadness and disgust. Ryosuke Ueno, Yukiko I. Nakano, Fumio Nihei |
ICMI | 2 |
| 2020 | Impact of Personality on Nonverbal Behavior GenerationabstractTo realize natural-looking virtual agents, one key technical challenge is to automatically generate nonverbal behaviors from spoken language. Since nonverbal behavior varies depending on personality, it is important to generate these nonverbal behaviors to match the expected personality of a virtual agent. In this work, we study how personality traits relate to the process of generating individual nonverbal behaviors from the whole body, including the head, eye gaze, arms, and posture. To study this, we first created a dialogue corpus including transcripts, a broad range of labelled nonverbal behaviors, and the Big Five personality scores of participants in dyad interactions. We constructed models that can predict each nonverbal behavior label given as an input language representation from the participants' spoken sentences. Our experimental results show that personality can help improve the prediction of nonverbal behaviors. Ryo Ishii, Chaitanya Ahuja, Yukiko I. Nakano, Louis-Philippe Morency |
IVA | 3 |
| 2019 | Determining Iconic Gesture Forms based on Entity Image RepresentationabstractIconic gestures are used to depict physical objects mentioned in speech, and the gesture form is assumed to be based on the image of a given object in the speaker’s mind. Using this idea, this study proposes a model that learns iconic gesture forms from an image representation obtained from pictures of physical entities. First, we collect a set of pictures of each entity from the web, and create an average image representation from them. Subsequently, the average image representation is fed to a fully connected neural network to decide the gesture form. In the model evaluation experiment, our two-step gesture form selection method can classify seven types of gesture forms with over 62% accuracy. Furthermore, we demonstrate an example of gesture generation in a virtual agent system in which our model is used to create a gesture dictionary that assigns a gesture form for each entry word in the dictionary. Fumio Nihei, Yukiko I. Nakano, Ryuichiro Higashinaka, Ryo Ishii |
ICMI | 2 |
| 2017 | Predicting meeting extracts in group discussions using multimodal convolutional neural networksabstractThis study proposes the use of multimodal fusion models employing Convolutional Neural Networks (CNNs) to extract meeting minutes from group discussion corpus. First, unimodal models are created using raw behavioral data such as speech, head motion, and face tracking. These models are then integrated into a fusion model that works as a classifier. The main advantage of this work is that the proposed models were trained without any hand-crafted features, and they outperformed a baseline model that was trained using hand-crafted features. It was also found that multimodal fusion is useful in applying the CNN approach to model multimodal multiparty interaction. Fumio Nihei, Yukiko I. Nakano, Yutaka Takase |
ICMI | 2 |
| 2016 | Speakers' head and gaze dynamics weakly correlate in group conversationabstractWhen modeling natural conversational behavior of an agent, a head direction becomes an intuitive proxy to visual attention. We examine this assumption and carefully investigate the relationship between head directions and gaze dynamics through the use of eye-movement tracking. In a group conversation settings, we analyze relationships of the two nonverbal social signals - head directions and gaze dynamics - linked to influential and non-influential statements. We develop a clustering method to estimate the number of gaze targets. We employ this method to show that head and gaze dynamic behaviors are not correlated, and thus head cannot be used as a direct proxy to a person's gaze in the context of conversations. We also describe in detail how influential statements affect head and gaze behaviors. The findings have implications on methodology, modeling and design of natural conversational agents and present a supportive evidence for employing gaze-tracking into the future conversational technologies. Hana Vrzakova, Roman Bednarik, Yukiko I. Nakano, Fumio Nihei |
ETRA | 3 |
| 2016 | Generating Iconic Gestures based on Graphic Data Analysis and ClusteringabstractGesture generation is one of the most important tasks in humanoid interfaces because hand gestures by humanoid robots and animated agents are useful in improving the comprehensibility of conversation content. This study proposes a method for automatically generating iconic drawing gestures using image processing and machine learning techniques. First, we collected a set of graphic images for over 1000 objects and classified the objects into 4 types of shapes; these shapes were used as the drawing gesture shapes. By implementing a gesture shape decision mechanism, we also built a system that takes a sentence as the system input and produces hand gesture animations that are synchronized with synthetic speech. Yuki Kadono, Yutaka Takase, Yukiko I. Nakano |
HRI | 3 |
| 2016 | Assessing the Communication Attitude of the Elderly using Prosodic Information and Head MotionsabstractIn order to provide a watching service for the elderly with dementia, recognizing and assessing their cognitive and health status are indispensable. In this study, we propose a prediction model for assessing the communication attitude of the elderly during interacting with a virtual agent. We define speech features and head motion features using frequency analysis and apply them to a linear regression analysis. The coefficient of determination for the model using only speech features was 0.413 and that for a model that exploited both speech and head movement features was 0.505. This result suggests that combining speech and head movement data is useful in predicting the communication attitude of the elderly, and the model is can be applied for automatic assessment in watching services. Toshiki Yamanaka, Yutaka Takase, Yukiko I. Nakano |
HRI | 3 |
| 2016 | Meeting extracts for discussion summarization based on multimodal nonverbal informationabstractGroup discussions are used for various purposes, such as creating new ideas and making a group decision. It is desirable to archive the results and processes of the discussion as useful resources for the group. Therefore, a key technology would be a way to extract meaningful resources from a group discussion. To accomplish this goal, we propose classification models that select meeting extracts to be included in the discussion summary based on nonverbal behavior such as attention, head motion, prosodic features, and co-occurrence patterns of these behaviors. We create different prediction models depending on the degree of extract-worthiness, which is assessed by the agreement ratio among human judgments. Our best model achieves 0.707 in F-measure and 0.75 in recall rate, and can compress a discussion into 45% of its original duration. The proposed models reveal that nonverbal information is indispensable for selecting meeting extracts of a group discussion. One of the future directions is to implement the models as an automatic meeting summarization system. Fumio Nihei, Yukiko I. Nakano, Yutaka Takase |
ICMI | 2 |
| 2016 | Estimating communication skills using dialogue acts and nonverbal features in multiple discussion datasetsabstractThis paper focuses on the computational analysis of the individual communication skills of participants in a group. The computational analysis was conducted using three novel aspects to tackle the problem. First, we extracted features from dialogue (dialog) act labels capturing how each participant communicates with the others. Second, the communication skills of each participant were assessed by 21 external raters with experience in human resource management to obtain reliable skill scores for each of the participants. Third, we used the MATRICS corpus, which includes three types of group discussion datasets to analyze the influence of situational variability regarding to the discussion types. We developed a regression model to infer the score for communication skill using multimodal features including linguistic and nonverbal features: prosodic, speaking turn, and head activity. The experimental results show that the multimodal fusing model with feature selection achieved the best accuracy, 0.74 in R2 of the communication skill. A feature analysis of the models revealed the task-dependent and task-independent features to contribute to the prediction performance. Shogo Okada, Yoshihiko Ohtake, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Yutaka Takase, Katsumi Nitta |
ICMI | 3 |
| 2016 | Introduction to the Special Issue on New Directions in Eye Gaze for Interactive Intelligent SystemsabstractEye gaze has been used broadly in interactive intelligent systems. The research area has grown in recent years to cover emerging topics that go beyond the traditional focus on interaction between a single user and an interactive system. This special issue presents five articles that explore new directions of gaze-based interactive intelligent systems, ranging from communication robots in dyadic and multiparty conversations to a driving simulator that uses eye gaze evidence to critique learners’ behavior. Yukiko I. Nakano, Roman Bednarik, Hung-Hsuan Huang, Kristiina Jokinen |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2016 | Generating Robot Gaze on the Basis of Participation Roles and Dominance Estimation in Multiparty InteractionabstractGaze is an important nonverbal feedback signal in multiparty face-to-face conversations. It is well known that gaze behaviors differ depending on participation role: speaker, addressee, or side participant. In this study, we focus on dominance as another factor that affects gaze. First, we conducted an empirical study and analyzed its results that showed how gaze behaviors are affected by both dominance and participation roles. Then, using speech and gaze information that was statistically significant for distinguishing the more dominant and less dominant person in an empirical study, we established a regression-based model for estimating conversational dominance. On the basis of the model, we implemented a dominance estimation mechanism that processes online speech and head direction data. Then we applied our findings to human-robot interaction. To design robot gaze behaviors, we analyzed gaze transitions with respect to participation roles and dominance and implemented gaze-transition models as robot gaze behavior generation rules. Finally, we evaluated a humanoid robot that has dominance estimation functionality and determines its gaze based on the gaze models, and we found that dominant participants had a better impression of less dominant robot gaze behaviors. This suggests that a robot using our gaze models was preferred to a robot that was simply looking at the speaker. We have demonstrated the importance of considering dominance in human-robot multiparty interaction. Yukiko I. Nakano, Takashi Yoshino 0003, Misato Yatsushiro, Yutaka Takase |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2015 | Predicting Participation Styles using Co-occurrence Patterns of Nonverbal Behaviors in Collaborative LearningabstractWith the goal of assessing participant attitudes and group activities in collaborative learning, this study presents models of participation styles based on co-occurrence patterns of nonverbal behaviors between conversational participants. First, we collected conversations among groups of three people in a collaborative learning situation, wherein each participant had a digital pen and wore a glasses-type eye tracker. We then divided the collected multimodal data into 0.1-second intervals. The discretized data were applied to an unsupervised method to find co-occurrence behavioral patterns. As a result, we discovered 122 multimodal behavioral motifs from more than 3,000 possible combinations of behaviors by three participants. Using the multimodal behavioral motifs as predictor variables, we created regression models for assessing participation styles. The multiple correlation coefficients ranged from 0.74 to 0.84, indicating a good fit between the models and the data. A correlation analysis also enabled us to identify a smaller set of behavioral motifs (fewer than 30) that are statistically significant as predictors of participation styles. These results show that automatically discovered combinations of multiple kinds of nonverbal information with high co-occurrence frequencies observed between multiple participants as well as for a single participant are useful in characterizing the participant's attitudes towards the conversation. Yukiko I. Nakano, Sakiko Nihonyanagi, Yutaka Takase, Yuki Hayashi, Shogo Okada |
ICMI | 1 |
| 2015 | Design and Evaluation of Mirror Interface MIOSS to Overlay Remote 3D Spaces
Ryo Ishii, Shiro Ozawa, Akira Kojima, Kazuhiro Otsuka, Yuki Hayashi, Yukiko I. Nakano |
INTERACT (4) | 6 |
| 2014 | Determining robot gaze according to participation roles in multiparty conversationsabstractGaze is an important nonverbal feedback signal in multiparty face-to-face conversations. To build a robot that can convey the appropriate attentional behavior in human-robot multiparty conversations, this paper analyzes human attentional behaviors in multiparty conversations, and establishes gaze-transition models for speakers, addressees, and side participants. Further, the model was implemented in a humanoid robot that can control its gaze. Takashi Yoshino 0003, Yuki Hayashi, Yukiko I. Nakano |
HAI | 3 |
| 2014 | Gaze-in 2014: the 7th Workshop on Eye Gaze in Intelligent Human Machine InteractionabstractThis paper presents a summary of the seventh workshop on Eye Gaze in Intelligent Human Machine Interaction. The Gaze-in 2014 workshop is a part of a series of workshops held around the topics related to gaze and multimodal interaction. The workshop web-site can be found at http://hhhuang.homelinux.com/gaze_in/. Hung-Hsuan Huang, Roman Bednarik, Kristiina Jokinen, Yukiko I. Nakano |
ICMI | 4 |
| 2014 | Predicting Influential Statements in Group Discussions using Speech and Head Motion InformationabstractGroup discussions are used widely when generating new ideas and forming decisions as a group. Therefore, it is assumed that giving social influence to other members through facilitating the discussion is an important part of discussion skill. This study focuses on influential statements that affect discussion flow and highly related to facilitation, and aims to establish a model that predicts influential statements in group discussions. First, we collected a multimodal corpus using different group discussion tasks; in-basket and case-study. Based on schemes for analyzing arguments, each utterance was annotated as being influential or not. Then, we created classification models for predicting influential utterances using prosodic features as well as attention and head motion information from the speaker and other members of the group. In our model evaluation, we discovered that the assessment of each participant in terms of discussion facilitation skills by experienced observers correlated highly to the number of influential utterances by a given participant. This suggests that the proposed model can predict influential statements with considerable accuracy, and the prediction results can be a good predictor of facilitators in group discussions. Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Shogo Okada |
ICMI | 2 |
| 2014 | Estimating Collaborative Attitudes based on Non-verbal Features in Collaborative Learning InteractionabstractTo understand collaborative learning interaction, it is important to analyze not only argument processes based on verbal information but also non-verbal interaction. In order to analyze learning situations in collaborative learning, our previous work proposed an estimation method for learning attitudes based on participants’ non-verbal features. Because the method used limited features, this research enhances the method of the participants’ collaborative attitudes by analyzing non-verbal features in detail. The model also considers participants’ knowledge of their learning subject in the analysis. The estimation model detects three levels of the participants’ collaborative attitudes based on multinomial logistic regression analysis. The results of the analysis show that the speech interval feature, in particular, affects the participants’ collaborative attitudes. In addition, the results indicate that speakers with knowledge of the learning subject receive more attention from participants with insufficient knowledge. The results of the model evaluation find that the f-measure for classifying the participants’ collaborative attitudes is 0.569; for participants with knowledge, the f-measure is 0.647. Yuki Hayashi, Haruka Morita, Yukiko I. Nakano |
KES | 3 |
| 2013 | HRI face-to-face: gaze and speech communication (fifth workshop on eye-gaze in intelligent human-machine interaction)
Frank Broz, Hagen Lehmann, Bilge Mutlu, Yukiko I. Nakano |
HRI | 4 |
| 2013 | Gazein'13: the 6th workshop on eye gaze in intelligent human machine interaction: gaze in multimodal interactionabstractThis paper presents a summary of the sixth workshop in Eye Gaze in Intelligent Human Machine Interaction. The GazeIn'13 workshop is a part of a series of workshops held around the topics related to gaze and multimodal interaction. Roman Bednarik, Hung-Hsuan Huang, Yukiko I. Nakano, Kristiina Jokinen |
ICMI | 3 |
| 2013 | Implementation and evaluation of a multimodal addressee identification mechanism for multiparty conversation systemsabstractIn conversational agents with multiparty communication functionality, a system needs to be able to identify the addressee for the current floor and respond to the user when the utterance is addressed to the agent. This study proposes some addressee identification models based on speech and gaze information, and tests whether the models can be applied to different proxemics. We build an addressee identification mechanism by implementing the models and incorporate it into a fully autonomous multiparty conversational agent. The system identifies the addressee from online multimodal data and uses this information in language understanding and dialogue management. Finally, an evaluation experiment shows that the proposed addressee identification mechanism works well in a real-time system, with an F-measure for addressee estimation of 0.8 for agent-addressed utterances. We also found that our system more successfully avoided disturbing the conversation by mistakenly taking a turn when the agent is not addressed. Yukiko I. Nakano, Naoya Baba, Hung-Hsuan Huang, Yuki Hayashi |
ICMI | 1 |
| 2013 | Visualization System for Analyzing Collaborative Learning InteractionabstractIn collaborative learning, participants interact with other members by exchanging not only verbal information but also nonverbal information, such as looking at other participants and note taking, which plays an important role in facilitating effective learning. In order to utilize such non-verbal information to analyze collaborative learning interaction, we conducted experiments to collect non-verbal information (gaze direction, speech intervals, and writing actions of participants) using multimodal measurement devices. In this paper, we propose a visualization method of collaborative learning based on a multimodal corpus of collaborative learning. The method intuitively visualizes the learning interaction with respect to collaborative and learning attitudes of each participant. We introduce the prototype system and show an example of the interface. Yuki Hayashi, Yuji Ogawa, Yukiko I. Nakano |
KES | 3 |
| 2013 | Investigating culture-related aspects of behavior for virtual characters
Birgit Lugrin, Elisabeth André, Matthias Rehm, Yukiko I. Nakano |
Auton. Agents Multi Agent Syst. | 4 |
| 2013 | Gaze awareness in conversational agents: Estimating a user's conversational engagement from eye gazeabstractIn face-to-face conversations, speakers are continuously checking whether the listener is engaged in the conversation, and they change their conversational strategy if the listener is not fully engaged. With the goal of building a conversational agent that can adaptively control conversations, in this study we analyze listener gaze behaviors and develop a method for estimating whether a listener is engaged in the conversation on the basis of these behaviors. First, we conduct a Wizard-of-Oz study to collect information on a user's gaze behaviors. We then investigate how conversational disengagement, as annotated by human judges, correlates with gaze transition, mutual gaze (eye contact) occurrence, gaze duration, and eye movement distance. On the basis of the results of these analyses, we identify useful information for estimating a user's disengagement and establish an engagement estimation method using a decision tree technique. The results of these analyses show that a model using the features of gaze transition, mutual gaze occurrence, gaze duration, and eye movement distance provides the best performance and can estimate the user's conversational engagement accurately. The estimation model is then implemented as a real-time disengagement judgment mechanism and incorporated into a multimodal dialog manager in an animated conversational agent. This agent is designed to estimate the user's conversational engagement and generate probing questions when the user is distracted from the conversation. Finally, we evaluate the engagement-sensitive agent and find that asking probing questions at the proper times has the expected effects on the user's verbal/nonverbal behaviors during communication with the agent. We also find that our agent system improves the user's impression of the agent in terms of its engagement awareness, behavior appropriateness, conversation smoothness, favorability, and intelligence. Ryo Ishii, Yukiko I. Nakano, Toyoaki Nishida |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2012 | Gaze in HRI: from modeling to communicationabstractThe purpose of this half-day workshop is to explore the role of social gaze in human-robot interaction, both how to measure social gaze behavior by humans and how to implement it in robots that interact with them. Gaze directed at an interaction partner has become a subject of increased attention in human-robot interaction research. While traditional robotics research has focused work on robot gaze solely on the identification and manipulation of objects, researchers in HRI have come to recognize that gaze is a social behavior in addition to a way of sensing the world. This workshop will approach the problem of understanding the role of social gaze in human-robot interaction from the dual perspectives of investigating human-human gaze for design principles to apply to robots and of experimentally evaluating human-robot gaze interaction in order to assess how humans engage in gaze behavior with robots. Frank Broz, Hagen Lehmann, Yukiko I. Nakano, Bilge Mutlu |
HRI | 3 |
| 2012 | Listener agent for elderly people with dementiaabstractWith the goal of developing a conversational humanoid that can serve as a companion for people with dementia, we propose an autonomous virtual agent that can generate backchannel feedback, such as head nods and verbal acknowledgement, on the basis of acoustic information in the user's speech. The system is also capable of speech recognition and language understanding functionalities, which are potentially useful for evaluating the cognitive status of elderly people on a daily basis. Yoichi Sakai, Yuuko Nonaka, Kiyoshi Yasuda, Yukiko I. Nakano |
HRI | 4 |
| 2012 | Referent identification process in human-robot multimodal communicationabstractThis paper presents a communication robot that can generate a referent identification conversation with human users. First, we conduct an experiment to collect face-to-face referent identification communication and investigate how the referent is identified by exchanging multiple speech turns between the participants. On the basis of the experimental observations, we implement a communication robot that can manage a referent identification conversation with a user by integrating the linguistic information obtained from speech recognition and the vision information obtained from a robot camera. Yuta Shibasaki, Takahiro Inaba, Yukiko I. Nakano |
HRI | 3 |
| 2012 | 4th workshop on eye gaze in intelligent human machine interaction: eye gaze and multimodalityabstractThis is the fourth workshop in a series of workshops on Eye Gaze in Intelligent Human Machine Interaction, in which we have discussed a wide range of issues for eye gaze; technologies for sensing human attentional behaviors, roles of attentional behaviors as social gaze in human-human and human-humanoid interaction, attentional behaviors in problem-solving and task-performing, gaze-based intelligent user interfaces, and evaluation of gaze-based user interfaces. In addition to these topics, this year's workshop focuses on eye gaze in multimodal interpretation and generation. Since eye gaze is one of the facial communication modalities, gaze information can be combined with other modalities or bodily motions to contribute to the meaning of utterance and serve as communication signals. Yukiko I. Nakano, Kristiina Jokinen, Hung-Hsuan Huang |
ICMI | 1 |
| 2012 | Towards Assessing the Communication Responsiveness of People with Dementia
Yuuko Nonaka, Yoichi Sakai, Kiyoshi Yasuda, Yukiko I. Nakano |
IVA | 4 |
| 2011 | Making virtual conversational agent aware of the addressee of users' utterances in multi-user conversation using nonverbal informationabstractIn multi-user human-agent interaction, the agent should respond to the user when an utterance is addressed to it. To do this, the agent needs to be able to judge whether the utterance is addressed to the agent or to another user. This study proposes a method for estimating the addressee based on the prosodic features of the user's speech and head direction (approximate gaze direction). First, a WOZ experiment is conducted to collect a corpus of human-humanagent triadic conversations. Then, analysis is performed to find out whether the prosodic features as well as head direction information are correlated with the addressee-hood. Based on this analysis, a SVM classifier is trained to estimate the addressee by integrating both the prosodic features and head movement information. Finally, a prototype agent equipped with this real-time addressee estimation mechanism is developed and evaluated. Hung-Hsuan Huang, Naoya Baba, Yukiko I. Nakano |
ICMI | 3 |
| 2011 | 2nd workshop on eye gaze in intelligent human machine interactionabstractThis workshop addresses a wide range of issues concerning eye gaze: recognizing user's gaze, generating gaze behaviors in conversational humanoids, analyzing human attentional behaviors during interacting with IUIs, and evaluation of gaze-based IUIs. Through a comprehensive discussion, the workshop aims at bringing together researchers with different backgrounds, and establishing an interdisciplinary research community in attention aware interactive systems. Yukiko I. Nakano, Cristina Conati, Thomas Bader |
IUI | 1 |
| 2011 | Identifying Utterances Addressed to an Agent in Multiparty Human-Agent Conversations
Naoya Baba, Hung-Hsuan Huang, Yukiko I. Nakano |
IVA | 3 |
| 2011 | Culture-Related Topic Selection in Small Talk Conversations across Germany and Japan
Birgit Lugrin, Yukiko I. Nakano, Afia Akhter Lipi, Matthias Rehm, Elisabeth André |
IVA | 2 |
| 2011 | Estimating a User's Conversational Engagement Based on Head Pose Information
Ryota Ooko, Ryo Ishii, Yukiko I. Nakano |
IVA | 3 |
| 2010 | Estimating user's engagement from eye-gaze behaviors in human-agent conversationsabstractIn face-to-face conversations, speakers are continuously checking whether the listener is engaged in the conversation and change the conversational strategy if the listener is not fully engaged in the conversation. With the goal of building a conversational agent that can adaptively control conversations with the user, this study analyzes the user's gaze behaviors and proposes a method for estimating whether the user is engaged in the conversation based on gaze transition 3-gram patterns. First, we conduct a Wizard-of-Oz experiment to collect the user's gaze behaviors. Based on the analysis of the gaze data, we propose an engagement estimation method that detects the user's disengagement gaze patterns. The algorithm is implemented as a real-time engagement-judgment mechanism and is incorporated into a multimodal dialogue manager in a conversational agent. The agent estimates the user's conversational engagement and generates probing questions when the user is distracted from the conversation. Finally, we conduct an evaluation experiment using the proposed engagement-sensitive agent and demonstrate that the engagement estimation function improves the user's impression of the agent and the interaction with the agent. In addition, probing performed with proper timing was also found to have a positive effect on user's verbal/nonverbal behaviors in communication with the conversational agent. Yukiko I. Nakano, Ryo Ishii |
IUI | 1 |
| 2009 | The Lessons Learned in Developing Multi-user Attentive Quiz Agents
Hung-Hsuan Huang, Takuya Furukawa, Hiroki Ohashi, Aleksandra Cerekovic, Yuji Yamaoka, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 7 |
| 2009 | Information State Based Multimodal Dialogue Management: Estimating Conversational Engagement from Gaze Information
Yukiko I. Nakano, Yuji Yamaoka |
IVA | 1 |
| 2008 | Enculturating conversational interfaces by socio-cultural aspects of communicationabstractThe workshop is centered around three main research challenges: 1.) Computationally viable models of cultural aspects of conversations: Cultural norms and values penetrate all our communications and interactions by giving us heuristics how to behave and how to interpret the verbal and nonverbal behavior of others. To make such a notion like culture available for computation, we need a very specific theory of culture that takes its effects on communication and interaction into account.2.) Reliable empirical data on cultural/cross-cultural interaction: To realize technical systems that take cultural influences on behavior into account, precise data analysis on how this influence manifests itself is necessary. In the literature, this information is often given in very general forms without to the precise data on which the observations are based.3.) Enculturating conversational interfaces: Having identified cultural influences on verbal/nonverbal communicative behaviors, it remains to be shown how this can be applied to the development of human-computer interfaces, for instance in an interface reflecting cultural norms and values of communication. Matthias Rehm, Elisabeth André, Yukiko I. Nakano, Toyoaki Nishida |
IUI | 3 |
| 2008 | Estimating User's Conversational Engagement Based on Gaze Behaviors
Ryo Ishii, Yukiko I. Nakano |
IVA | 2 |
| 2008 | Enculturating Conversational Agents Based on a Comparative Corpus Study
Afia Akhter Lipi, Yuji Yamaoka, Matthias Rehm, Yukiko I. Nakano |
IVA | 4 |
| 2008 | Culture-Specific First Meeting Encounters between Virtual Agents
Matthias Rehm, Yukiko I. Nakano, Elisabeth André, Toyoaki Nishida |
IVA | 2 |
| 2007 | Towards a Multicultural ECA Tour Guide System
Aleksandra Cerekovic, Hung-Hsuan Huang, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 4 |
| 2007 | A Script Driven Multimodal Embodied Conversational Agent Based on a Generic Framework
Hung-Hsuan Huang, Aleksandra Cerekovic, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 4 |
| 2007 | A Quiz Game Console Based on a Generic Embodied Conversational Agent Framework
Hung-Hsuan Huang, Taku Inoue, Aleksandra Cerekovic, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 5 |
| 2006 | Avatar's Gaze Control to Facilitate Conversational Turn-Taking in Virtual-Space Multi-user Voice Chat System
Ryo Ishii, Toshimitsu Miyajima, Kinya Fujita, Yukiko I. Nakano |
IVA | 4 |
| 2006 | Toward a Universal Platform for Integrating Embodied Conversational Agent Components
Hung-Hsuan Huang, Tsuyoshi Masuda, Aleksandra Cerekovic, Kateryna Tarasenko, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
KES (2) | 6 |
| 2006 | Cards-to-presentation on the web: generating multimedia contents featuring agent animations
Yukiko I. Nakano, Toshihiro Murayama, Masashi Okamoto, Daisuke Kawahara, Sadao Kurohashi, Toyoaki Nishida |
J. Netw. Comput. Appl. | 1 |
| 2005 | How to Make Robot a Robust and Interactive Communicator
Yoshiyasu Ogasawara, Masashi Okamoto, Yukiko I. Nakano, Yong Xu 0012, Toyoaki Nishida |
KES (3) | 3 |
| 2005 | Generating CG Movies Based on a Cognitive Model of Shot Transition
Kazunori Okamoto, Yukiko I. Nakano, Masashi Okamoto, Hung-Hsuan Huang, Toyoaki Nishida |
KES (3) | 2 |
| 2003 | Towards a Model of Face-to-Face GroundingabstractWe investigate the verbal and nonverbal means for grounding, and propose a design for embodied conversational agents that relies on both kinds of signals to establish common ground in human-computer interaction. We analyzed eye gaze, head nods and attentional focus in the context of a direction-giving task. The distribution of nonverbal behaviors differed depending on the type of dialogue move being grounded, and the overall pattern reflected a monitoring of lack of negative feedback. Based on these results, we present an ECA that uses verbal and nonverbal grounding acts to update dialogue state. Yukiko I. Nakano, Gabe Reinstein, Tom Stocky, Justine Cassell |
ACL | 1 |
| 2003 | Embodied Conversational Agents for Presenting Intellectual Multimedia Contents
Yukiko I. Nakano, Toshihiro Murayama, Daisuke Kawahara, Sadao Kurohashi, Toyoaki Nishida |
KES | 1 |
| 2001 | Non-Verbal Cues for Discourse StructureabstractThis paper addresses the issue of designing embodied conversational agents that exhibit appropriate posture shifts during dialogues with human users. Previous research has noted the importance of hand gestures, eye gaze and head nods in conversations between embodied agents and humans. We present an analysis of human monologues and dialogues that suggests that postural shifts can be predicted as a function of discourse state in monologues, and discourse and conversation state in dialogues. On the basis of these findings, we have implemented an embodied conversational agent that uses Collagen in such a way as to generate postural shifts. Justine Cassell, Yukiko I. Nakano, Timothy W. Bickmore, Candace L. Sidner, Charles Rich |
ACL | 2 |
| 2000 | Taking Account of the User's View in 3D Multimodal Instruction Dialogue
Yukiko I. Nakano, Kenji Imamura, Hisashi Ohara |
COLING | 1 |
| 1996 | Interactive Multi modal Explanations and their Temporal Coordination
Tsuneaki Kato, Yukiko I. Nakano, H. Nakajima, Takaaki Hasegawa |
ECAI | 2 |