VLDB 2026 Research / reviewers in the wild / expert
Fumio Nihei
dblp:153/5681
· DBLP profile ↗
11ranked-venue papers
7as first author
5since 2021 · last 2024
0009-0006-8242-2240ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Participation Role-Driven Engagement Estimation of ASD Individuals in Neurodiverse Group DiscussionsabstractAdults with autism spectrum disorder (ASD) face difficulties in communicating with neurotypical people in their daily lives and workplaces. In addition, research on modeling communication in neurodiverse groups is scarce. To recognize communication difficulties caused by neurodiversity, we first, collected a multimodal corpus for decision-making discussions in neurodiverse groups that included a person with ASD and two neurotypical participants. For corpus analysis, we investigated eye-gaze and facial expression exchanges between individuals with ASD and neurotypical participants during both listening and speaking. The findings were extended to automatically estimate the engagement of ASD individuals. To capture the effect of contingent behaviors between ASD individuals and neurotypical participants, we developed a transformer-based model that considers the participation role by changing the direction of cross-person attention depending on whether the ASD individual is listening or speaking. The proposed approach yields comparable results to the state-of-the-art for engagement estimation in neurotypical group conversations while accounting for the dynamic nature of behavior influence in face-to-face interactions. The code associated with this study is available at https://github.com/IUI-Lab/switch-attention. Kalin Stefanov, Yukiko I. Nakano, Chisa Kobayashi, Ibuki Hoshina, Tatsuya Sakato, Fumio Nihei, Chihiro Takayama, Ryo Ishii, Masatsugu Tsujii |
ICMI | 6 |
| 2023 | Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance EstimationabstractThis study investigated differences in the contributions of various features to in-person (IP) and video-mediated (VM) meetings. We focused on estimating important utterances using both an IP and a VM meeting corpora as the analysis data. A transformer model with dialogue history was used to estimate important utterances, and five types of input (text, speaker’s audio, others’ audio, speaker’s video, and others’ video) were fed to the model. A comparison of the models for IP and VM revealed that the speaker’s audio has a strong effect on the IP model, the video of the other participants strongly affects the VM model, and the text and others’ audio strongly affects both models in estimating important utterances. Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Atsushi Fukayama, Takao Nakamura |
ICASSP | 1 |
| 2023 | Prediction of Love-Like Scores After Speed Dating Based on Pre-obtainable Personal Characteristic Information
Ryo Ishii, Fumio Nihei, Yoko Ishii, Atsushi Otsuka, Kazuya Matsuo, Narichika Nomoto, Atsushi Fukayama, Takao Nakamura |
INTERACT (4) | 2 |
| 2022 | Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura |
INTERSPEECH | 1 |
| 2021 | Web-ECA: A Web-based ECA PlatformabstractRunning an embodied conversation agent (ECA) on a client–server model has the following advantages: 1) the experimenter need not bring a computer with installing the ECA to the experimental site, 2) the user need not own a high-performance computer, 3) it is easy to make changes to the ECA for system maintenance and operation, and 4) data collection using an ECA becomes easier. To realize these benefits, we propose a platform for executing ECA on a server–client model, i.e., a platform for ECA that can be accessed through the Internet. This paper describes the system configuration and an application of this proposed platform. Fumio Nihei, Yukiko I. Nakano |
ICMI | 1 |
| 2020 | Estimating the Intensity of Facial Expressions Accompanying Feedback Responses in Multiparty Video-Mediated CommunicationabstractProviding feedback to a speaker is an essential communication signal for maintaining a conversation. In specific feedback, which indicates the listener's reaction to the speaker?s utterances, the facial expression is an effective modality for conveying the listener's reactions. Moreover, not only the type of facial expressions, but also the degree of intensity of the expressions, may influence the meaning of the specific feedback. In this study, we propose a multimodal deep neural network model that predicts the intensity of facial expressions co-occurring with feedback responses. We focus on multiparty video-mediated communication. In video-mediated communication, close-up frontal face images of each participant are continuously presented on the display; the attention of the participants is more likely to be drawn to the facial expressions. We assume that in such communication, the importance of facial expression in the listeners? feedback responses increases. We collected 33 video-mediated conversations by groups of three people and obtained audio and speech data for each participant. Using the corpus collected as a dataset, we created a deep neural network model that predicts the intensity of 17 types of action units (AUs) co-occurring with the feedback responses. The proposed method employed GRU-based model with attention mechanism for audio, visual, and language modalities. A decoder was trained to produce the intensity values for the 17 AUs frame by frame. In the experiment, unimodal and multimodal models were compared in terms of their performance in predicting salient AUs that characterize facial expression in feedback responses. The results suggest that well-performing models differ depending on the AU categories; audio information was useful for predicting AUs that express happiness, and visual and language information contributes to predicting AUs expressing sadness and disgust. Ryosuke Ueno, Yukiko I. Nakano, Fumio Nihei |
ICMI | 4 |
| 2019 | Determining Iconic Gesture Forms based on Entity Image RepresentationabstractIconic gestures are used to depict physical objects mentioned in speech, and the gesture form is assumed to be based on the image of a given object in the speaker’s mind. Using this idea, this study proposes a model that learns iconic gesture forms from an image representation obtained from pictures of physical entities. First, we collect a set of pictures of each entity from the web, and create an average image representation from them. Subsequently, the average image representation is fed to a fully connected neural network to decide the gesture form. In the model evaluation experiment, our two-step gesture form selection method can classify seven types of gesture forms with over 62% accuracy. Furthermore, we demonstrate an example of gesture generation in a virtual agent system in which our model is used to create a gesture dictionary that assigns a gesture form for each entry word in the dictionary. Fumio Nihei, Yukiko I. Nakano, Ryuichiro Higashinaka, Ryo Ishii |
ICMI | 1 |
| 2017 | Predicting meeting extracts in group discussions using multimodal convolutional neural networksabstractThis study proposes the use of multimodal fusion models employing Convolutional Neural Networks (CNNs) to extract meeting minutes from group discussion corpus. First, unimodal models are created using raw behavioral data such as speech, head motion, and face tracking. These models are then integrated into a fusion model that works as a classifier. The main advantage of this work is that the proposed models were trained without any hand-crafted features, and they outperformed a baseline model that was trained using hand-crafted features. It was also found that multimodal fusion is useful in applying the CNN approach to model multimodal multiparty interaction. Fumio Nihei, Yukiko I. Nakano, Yutaka Takase |
ICMI | 1 |
| 2016 | Speakers' head and gaze dynamics weakly correlate in group conversationabstractWhen modeling natural conversational behavior of an agent, a head direction becomes an intuitive proxy to visual attention. We examine this assumption and carefully investigate the relationship between head directions and gaze dynamics through the use of eye-movement tracking. In a group conversation settings, we analyze relationships of the two nonverbal social signals - head directions and gaze dynamics - linked to influential and non-influential statements. We develop a clustering method to estimate the number of gaze targets. We employ this method to show that head and gaze dynamic behaviors are not correlated, and thus head cannot be used as a direct proxy to a person's gaze in the context of conversations. We also describe in detail how influential statements affect head and gaze behaviors. The findings have implications on methodology, modeling and design of natural conversational agents and present a supportive evidence for employing gaze-tracking into the future conversational technologies. Hana Vrzakova, Roman Bednarik, Yukiko I. Nakano, Fumio Nihei |
ETRA | 4 |
| 2016 | Meeting extracts for discussion summarization based on multimodal nonverbal informationabstractGroup discussions are used for various purposes, such as creating new ideas and making a group decision. It is desirable to archive the results and processes of the discussion as useful resources for the group. Therefore, a key technology would be a way to extract meaningful resources from a group discussion. To accomplish this goal, we propose classification models that select meeting extracts to be included in the discussion summary based on nonverbal behavior such as attention, head motion, prosodic features, and co-occurrence patterns of these behaviors. We create different prediction models depending on the degree of extract-worthiness, which is assessed by the agreement ratio among human judgments. Our best model achieves 0.707 in F-measure and 0.75 in recall rate, and can compress a discussion into 45% of its original duration. The proposed models reveal that nonverbal information is indispensable for selecting meeting extracts of a group discussion. One of the future directions is to implement the models as an automatic meeting summarization system. Fumio Nihei, Yukiko I. Nakano, Yutaka Takase |
ICMI | 1 |
| 2014 | Predicting Influential Statements in Group Discussions using Speech and Head Motion InformationabstractGroup discussions are used widely when generating new ideas and forming decisions as a group. Therefore, it is assumed that giving social influence to other members through facilitating the discussion is an important part of discussion skill. This study focuses on influential statements that affect discussion flow and highly related to facilitation, and aims to establish a model that predicts influential statements in group discussions. First, we collected a multimodal corpus using different group discussion tasks; in-basket and case-study. Based on schemes for analyzing arguments, each utterance was annotated as being influential or not. Then, we created classification models for predicting influential utterances using prosodic features as well as attention and head motion information from the speaker and other members of the group. In our model evaluation, we discovered that the assessment of each participant in terms of discussion facilitation skills by experienced observers correlated highly to the number of influential utterances by a given participant. This suggests that the proposed model can predict influential statements with considerable accuracy, and the prediction results can be a good predictor of facilitators in group discussions. Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Shogo Okada |
ICMI | 1 |