Atsushi Fukayama

dblp:02/6120 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
11since 2021 · last 2024
0000-0002-2133-016XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2024 Prediction of Praising Skills Based on Multimodal Information
abstract
Praising behavior is an important method of communication. An existing study constructed models to predict praising skill, which indicates the degree to which the praise is done well, by using only unimodal behavior such as speech audio or visual behavior of a praiser who gives praise in dyad interactions. To improve prediction performance, a model should be constructed that uses various additional information. In this study, we propose two approaches to predict praising skill highly accurately. The first uses trimodal (multimodal) behaviors extracted from visual, acoustic, and linguistic modalities. The second uses the behaviors of the receiver of praise since the reaction of the receiver should differ depending on how good the praise is. For this study, we collect trimodal features and the degree of praising skill in each praising scene in a dialogue. We construct multiple models to predict the degree of praising skills using various combinations of the trimodal features from the praiser and receiver. The experimental results show that the model that predicts praising skill most accurately uses multiple features related to both verbal and nonverbal behaviors of the praiser and receiver. Therefore, the two approaches of using trimodal behaviors and using features from both the receiver and praiser are effective for predicting praising skills in dyad interactions.
Toshiki Onishi, Asahi Ogushi, Ryo Ishii, Atsushi Fukayama, Akihiro Miyata
ACII4
2023 Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance Estimation
abstract
This study investigated differences in the contributions of various features to in-person (IP) and video-mediated (VM) meetings. We focused on estimating important utterances using both an IP and a VM meeting corpora as the analysis data. A transformer model with dialogue history was used to estimate important utterances, and five types of input (text, speaker’s audio, others’ audio, speaker’s video, and others’ video) were fed to the model. A comparison of the models for IP and VM revealed that the speaker’s audio has a strong effect on the IP model, the video of the other participants strongly affects the VM model, and the text and others’ audio strongly affects both models in estimating important utterances.
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Atsushi Fukayama, Takao Nakamura
ICASSP4
2023 Prediction of Love-Like Scores After Speed Dating Based on Pre-obtainable Personal Characteristic Information
Ryo Ishii, Fumio Nihei, Yoko Ishii, Atsushi Otsuka, Kazuya Matsuo, Narichika Nomoto, Atsushi Fukayama, Takao Nakamura
INTERACT (4)7
2023 How Far ahead Can Model Predict Gesture Pose from Speech and Spoken Text?
abstract
We investigated how far into the future nonverbal behavior can be predicted from speech and speech text. Specifically, we build a model that generates future behaviors from speech and speech text information and evaluate the quality of the generated behaviors. This helps to clarify how far into the future behavior can be accurately predicted. Our experimental results show that in Gesture Pose Generation using speech and speech text, on the basis of the input speech and text, the nonverbal behavior up to at least 500 ms ahead can be predicted with objective evaluation values that are the same as those when no future prediction is made. This result shows a new possibility for Gesture Pose Generation using speech and speech text to predict the future up to at least 500 ms ahead with no performance degradation.
Ryo Ishii, Akira Morikawa, Shin'ichiro Eitoku, Atsushi Fukayama, Takao Nakamura
IVA4
2023 A Study of Prediction of Listener's Comprehension Based on Multimodal Information
abstract
During dialogues, speakers need to be able to predict whether their partners understand their message. This is important for not only for human-to-human interaction but also human-to-agent interaction. We consider that if the listener's comprehension level can be automatically predicted, interactive agents will be able to communicate appropriately according to the user's comprehension level. However, to the best of our knowledge, there is no case study that reveals how comprehension can be predicted based on multimodal information about the listener. In this study, we attempt to predict comprehension levels on the basis of the listener's multimodal information. First, we construct a dialogue corpus consisting of the listener's comprehension levels and the listener's multimodal information. Next, we construct machine learning models that predict the listener's comprehension levels on the basis of the listener's multimodal information. Our results suggest that our model was able to predict a listener's comprehension level on the basis of a listener's multimodal information. In addition, two movements, the lifting of the cheeks and the pulling up of the corners of the lips, were suggested to be important in assessing the listener's level of comprehension.
Shunichi Kinoshita, Toshiki Onishi, Naoki Azuma, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
IVA5
2023 Prediction of Various Backchannel Utterances Based on Multimodal Information
abstract
The listener's backchannels are an important part of dialogues. With appropriate backchannels, people are able to smoothly promote dialogues. Thus, backchannels are considered to be important in dialogues between not only humans but also humans and agents. Progress has been made in studying dialogue agents that perform natural affable dialogue. However, we have not clarified whether the listener's various backchannel types are predictable using the speaker's multimodal information. In this paper, we attempt to predict a listener's various backchannel types on the basis of the speaker's multimodal information in dialogues. First, we construct a dialogue corpus that consists of multimodal information of a speaker's utterances and a listener's backchannels. Second, we construct machine learning models to predict a listener's various backchannel types on the basis of a speaker's multimodal information. Our results suggest that our model was able to predict a listener's various backchannel types on the basis of a speaker's multimodal information.
Toshiki Onishi, Naoki Azuma, Shunichi Kinoshita, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
IVA5
2022 Effect of repetitive motion intervention on self-avatar on the sense of self-individuality
abstract
In recent years, the human Digital Twin has been discussed as new technology. When we discuss a world in which one’s self-avatar autonomously performs social activities in cyberspace, the questions arise whether or not the behavior of the avatars feels like one’s own, and whether or not we can approve of the self-avatars’ social activities on behalf of ourselves. We define such feeling as the sense of self-individuality. In this study, we focused on the situation in which self-avatars perform presentations on behalf of ourselves to investigate the effect of the modification experience on the presentation motions by self-avatars on the sense of self-individuality. We conducted VR-based experiments in which the motion modification intervention was performed on self-avatars over eight weeks by 24 experiment participants. As a result, we found that the sense of self-individuality was improved as the number of modifications and interventions increased. However, we found that the intensity of motion modification did not correlate with the improvement of the sense of self-individuality in this experiment condition. We also found that the sense of self-individuality was reduced when others intervened in the motion. From these results, we clarified that the experience of motion modification on self-avatars is significant when designing the behavior of avatars acting on behalf of ourselves in human Digital Twin. Further investigation is required to clarify the effect of the long-term intervention on behavior to distinguish between the mere exposure effect.
Tetsunari Inamura, Shin'ichiro Eitoku, Iwaki Toshima, Shinya Shimizu, Atsushi Fukayama, Shiro Ozawa, Takao Nakamura
HAI5
2022 Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura
INTERSPEECH6
2022 Analysis of praising skills focusing on utterance contents
Asahi Ogushi, Toshiki Onishi, Yohei Tahara, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
INTERSPEECH5
2022 Determining most suitable listener backchannel type for speaker's utterance
abstract
A major hurdle in achieving a dialogue system that enables smooth dialogue is to determine how to generate an appropriate response to a user's utterance. Previous research has focused mainly on estimating whether to make an utterance backchannel in response to the user's utterance. We go one step further by examining, for the first time, the relationship between the type of utterance backchannel to be used and intent and type of the speaker's utterance, known as a dialogue act (DA). Specifically, we propose a new method for classifying utterance backchannels into nine types. We also created a corpus consisting of the DAs of speaker utterances and the backchannel types of listener utterances then used it to analyze the relationship between a speaker's and listener's utterances. Our findings clarify that the occurrence frequencies of a listener's backchannel types significantly depend on the DAs of the speaker's utterances. Since the goal of our research is to construct a dialogue system that generates a more natural backchannel, this classification method, which determines certain types of aids from the speaker's DA, will be beneficial to such a system.
Akira Morikawa, Ryo Ishii, Hajime Noto, Atsushi Fukayama, Takao Nakamura
IVA4
2022 A Comparison of Praising Skills in Face-to-Face and Remote Dialogues
abstract
Praising behavior is considered to an important method of communication in daily life and social activities. An engineering analysis of praising behavior is therefore valuable. However, a dialogue corpus for this analysis has not yet been developed. Therefore, we develop corpuses for face-to-face and remote two-party dialogues with ratings of praising skills. The corpuses enable us to clarify how to use verbal and nonverbal behaviors for successfully praise. In this paper, we analyze the differences between the face-to-face and remote corpuses, in particular the expressions in adjudged praising scenes in both corpuses, and also evaluated praising skills. We also compare differences in head motion, gaze behavior, facial expression in high-rated praising scenes in both corpuses. The results showed that the distribution of praising scores was similar in face-to-face and remote dialogues, although the ratio of the number of praising scenes to the number of utterances was different. In addition, we confirmed differences in praising behavior in face-to-face and remote dialogues.
Toshiki Onishi, Asahi Ogushi, Yohei Tahara, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
LREC5
2011 A Smart Information Sharing Architecture in a Multi-Access Network, Multi-Service Environment
abstract
In a multi-network environment, users can receive services via several networks, protocols, and connected devices. These services are usually bound to their specific technology and often limited to a dedicated service provider domain. We are developing a platform overcoming these limitations by enabling information sharing beyond the boundaries of network technologies and services. In this paper, we propose a service architecture for the interaction of users/personas and services in a converged Internet and Telecom service space on top of multiple access networks. This is achieved by convergence of multiple separated communication channels on service layer in an IP Multimedia Subsystem based Telecom environment. We design a service platform enabling and controlling information flow across networks and services.
Niklas Blum, Junnosuke Yamada, Atsushi Fukayama, Thomas Magedanz, Naoki Uchida
ICC3
2002 Messages embedded in gaze of interface agents - impression management with agent's gaze
abstract
We propose a gaze movement model that enables an embodied interface agent to convey different impressions to users. Managing one's own impression to influence the behaviors of others plays an important role in human communications. To create a new application area which involves agents in this kind of social interaction, interface agents that manage their impressions are required. For this purpose, we build the gaze movement model based on three gaze parameters picked from a large number of psychological studies: amount of gaze, mean duration of gaze, and gaze points while averted. In this paper, we introduce the gaze movement model and gaze parameters. We then present an experiment in which subjects evaluated the impressions created by nine gaze patterns produced by altering the gaze parameters. The results indicate that reproducible relations exist between the gaze parameters and impressions, which shows the validity of the model
Atsushi Fukayama, Takehiko Ohno, Naoki Mukawa, Minako Sawaki, Norihiro Hagita
CHI1
2001 Expressing Personality of Interface Agents by Gaze
Atsushi Fukayama, Minako Sawaki, Takehiko Ohno, Hiroshi Murase, Norihiro Hagita, Naoki Mukawa
INTERACT1