EDBT 2026 Demo / reviewers in the wild / expert
Shogo Okada
dblp:05/956
· DBLP profile ↗
67ranked-venue papers
17as first author
31since 2021 · last 2025
0000-0002-9260-0403ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 31 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 29 · 11 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating Role of Big Five Personality Traits in Audio-Visual Rapport EstimationabstractAutomatic rapport estimation in social interactions is a central component of affective computing. Recent reports have shown that the estimation performance of rapport in initial interactions can be improved by using the participant’s personality traits as the model’s input. In this study, we investigate whether this findings applies to interactions between friends by developing rapport estimation models that utilize nonverbal cues (audio and facial expressions) as inputs. Our experimental results show that adding Big Five features (BFFs) to nonverbal features can improve the estimation performance of self-reported rapport in dyadic interactions between friends. Next, we demystify how BFFs improve the estimation performance of rapport through a comparative analysis between models with and without BFFs. We decompose rapport ratings into perceiver effects (people’s tendency to rate other people), target effects (people’s tendency to be rated by other people), and relationship effects (people’s unique ratings for a specific person) using the social relations model. We then analyze the extent to which BFFs contribute to capturing each effect. Our analysis demonstrates that the perceiver’s and the target’s BFFs lead estimation models to capture the perceiver and the target effects, respectively. Furthermore, our experimental results indicate that the combinations of facial expression features and BFFs achieve best estimation performances not only in estimating rapport ratings, but also in estimating three effects. Our study is the first step toward understanding why personality-aware estimation models of interpersonal perception accomplish high estimation performance. Takato Hayashi, Ryusei Kimura, Ryo Ishii, Shogo Okada |
FG | 4 |
| 2025 | Multimodal Classification of Co-speech Gesture Pragmatic Function in StorytellingabstractGestures, as essential co-speech behaviors in human communication, carry rich pragmatic functions. Accurately recognizing these functions could enhance an agent’s ability to understand communicative behavior. In spontaneous storytelling scenarios—unlike lab-controlled settings—gesture functions exhibit high variability and are strongly influenced by individual speaker differences, making it difficult for unimodal systems to reliably capture their pragmatic intent. To address this challenge, we collected and annotated naturally occurring co-speech gestures in narrative dialogues, assigning each gesture one of six pragmatic function labels. We further propose a multimodal sequential classification model that encodes skeletal motion, acoustic prosody, and facial dynamics using separate bidirectional LSTM networks. These modality-specific encodings are fused via cross-modal attention and a gated mechanism to capture temporal dependencies and complementary information across modalities. Experimental results demonstrate that our tri-modal system achieves 62.5% accuracy and a weighted F1 score of 0.62 on the six-way classification task, outperforming uni-modal and bi-modal baselines by 3–11%. Ablation analysis reveals that skeletal features provide the most discriminative power for the majority of gesture functions, acoustic features are critical for specific categories, and facial features—though weak in isolation—substantially enhance overall performance when integrated via attention. Jinqian Zhang, Sixia Li, Candy Olivia Mawalim, Shogo Okada |
HAI | 4 |
| 2025 | CCMI 2025: Cross-Cultural Multimodal Interaction
Koji Inoue, Shogo Okada, Divesh Lala, Sahba Zojaji, Nancy F. Chen, Tatsuya Kawahara |
ICMI | 2 |
| 2025 | Can Adaptive Interviewer Robot Based on Social Signals Make a Better Impression on Interviewees and Encourage Self-Disclosure?
Fuminori Nagasawa, Shogo Okada |
ICMI | 2 |
| 2025 | Robust Multilingual Audio Deepfake Detection Through Hybrid ModelingabstractThe increasing sophistication of AI-generated human voice poses a significant threat, demanding robust detection systems that can generalize effectively across diverse linguistic environments and synthesis techniques.In response to the SAFE Challenge, this paper introduces a novel approach to multilingual audio deepfake detection.Our primary contribution lies in the comprehensive study of deepfake detection using a multilingual speech corpus encompassing 17 languages and a broad spectrum of synthesis methods and acoustic conditions, designed to enable more realistic and challenging evaluations.To optimally utilize this diverse data, we propose a hybrid detection model that synergistically combines the strengths of end-to-end RawNet and AASIST architectures with language-agnostic representations learned from a multilingual selfsupervised learning model.Additionally, we explore the efficacy of RawBoost data augmentation in enhancing robustness against realworld noise.Our experimental evaluation demonstrates promising generalization in generated audio detection, achieving approximately 73% balanced accuracy across multilingual data and unseen synthesis algorithms. Candy Olivia Mawalim, Aulia Adila, Shogo Okada, Masashi Unoki |
IH&MMSec | 4 |
| 2025 | MultiMediate '25: Cross-cultural Multi-domain Engagement EstimationabstractEstimating momentary conversational engagement is central to assistive, socially aware AI systems, yet models are typically trained and evaluated within a single domain, limiting real-world robustness. The MultiMediate '25 challenge advances engagement estimation to more challenging, cross-cultural, and multi-domain settings. Building on prior challenge editions, we expand beyond NOXI as the sole training source by introducing NOXI-J, a new multilingual corpus covering Japanese and Chinese interactions, enabling both training and evaluation in diverse linguistic contexts. Although NOXI-J conceptually extends NOXI, we treat it as a distinct domain because linguistic, cultural, capture, and annotation differences induce measurable distribution shifts. In this paper, we present new annotations, precomputed multi-modal features (visual, vocal, and verbal), baseline evaluations, and an analysis of the best performing challenge solutions. Beyond accuracy, we quantify fairness using Conditional Demographic Disparity for gender and language. Our baselines confirm strong in-domain performance (e.g., paralinguistic eGeMAPS and video-transformer features) and reveal notable cross-domain drops, underscoring the challenge of cultural, linguistic, and interactional shifts. Fairness analyses indicate generally small discrepancies for our baselines. We observe the largest disparities for the proposed challenge solutions on the Chinese language test set. All annotations, features, code, and leaderboards are made publicly available to foster sustained progress on robust and fair engagement estimation. Daksitha Withanage, Marius Funk, Michal Balazia, Huajian Qiu, Shogo Okada, François Brémond, Jan Alexandersson, Andreas Bulling, Elisabeth André, Philipp Müller 0001 |
ACM Multimedia | 5 |
| 2025 | WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global CuisinesabstractGenta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Stephanie Yulia Salim, Yi Zhou 0019, Yinxuan Gui, David Ifeoluwa Adelani, Annie En-Shiun Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Wijaya, Alice Oh, Chong-Wah Ngo |
NAACL (Long Papers) | 45 |
| 2025 | Autonomous dialogue generation based on phase boundary detection within continuous motion for domestic robotabstractDialogue generation plays a key role in responding to user and providing transparency in motion execution in human-robot interaction. As motion planning is generally performed in terms of discrete motions, previous studies have focused on dialogue generation at the boundaries between motions. Recently, continuous motion generation was proposed to enable adapting actions to unique characteristics of the objects for domestic robots. Since a continuous motion generally involves physical and nonphysical phases, providing dialogues when the phase changes is crucial for decreasing users’ anxiety and guaranteeing safety. However, continuous motions lack clear phase boundaries, posing challenges for dialogue generation between phases. For this problem, we segmented continuous motions into discrete phases, and constructed a system to enable the robot to autonomously generate dialogues by detecting phase boundaries. To do so, we built phase estimation models using robot sensor data and designed system modules. Specifically, we collected data in the scenario of a robot assisting to lift the user up from bed. We segmented the continuous motion into three phases based on the user’s posture and whether the robot applied force to the human. The best phase estimation model achieved a macro F1 score of 0.894, demonstrating that phases can be estimated from sensor data. The evaluation results of our system demonstrated that the system accurately detects phase boundaries and generates appropriate dialogues corresponding to phases. Furthermore, we conducted simulations with a user agent to investigate system behaviors when the phase estimation was incorrect. The results suggested that explicitly stating the phase is important for avoiding misunderstandings and safety issues. Sixia Li, Tamon Miyake, Tetsuya Ogata, Shigeki Sugano, Shogo Okada |
RO-MAN | 5 |
| 2024 | Incremental Multimodal Sentiment Analysis for HAIs Based on Multitask Active Learning with Interannotator AgreementabstractMultimodal sentiment analysis (MSA) is critical in developing empathetic and adaptive multimodal dialogue systems or conversational agents that can naturally interact with users by recognizing sentiment and engagement. Addressing the challenges of collecting labeled data for MSA in human-agent interaction (HAI), this study introduces an innovative approach that combines active learning and multitask learning. Our efficient sentiment recognition model leverages active learning to select informative data for learning models, significantly reducing the labor-intensive data labeling process. Furthermore, we employ multitask learning to improve annotation (label) quality by evaluating alignment with true labels and interannotator agreement, thus enhancing the reliability of sentiment annotations. We evaluate the proposed multitask and active learning methods via a human-agent multimodal dialogue dataset that includes various types of sentiment annotations, which are publicly available. The experimental results demonstrate that by learning to predict the agreement score, multitask learning becomes better than singletask learning at capturing the uncertainties in the data. This study lays the groundwork for incremental learning strategies in MSA, aiming to adaptively understand user sentiments in human-agent interactions. Thus Karnjanapatchara, Sixia Li, Candy Olivia Mawalim, Kazunori Komatani, Shogo Okada |
ACII | 5 |
| 2024 | Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on EngagementabstractNon-verbal behavior is a central challenge in understanding the dynamics of a conversation and the affective states between interlocutors arising from the interaction. Although psychological research has demonstrated that non-verbal behaviors vary across cultures, limited computational analysis has been conducted to clarify these differences and assess their impact on engagement recognition. To gain a greater understanding of engagement and non-verbal behaviors among a wide range of cultures and language spheres, in this study we conduct a multilingual computational analysis of non-verbal features and investigate their role in engagement and engagement prediction. To achieve this goal, we first expanded the NoXi dataset, which contains interaction data from participants living in France, Germany, and the United Kingdom, by collecting session data of dyadic conversations in Japanese and Chinese, resulting in the enhanced dataset NoXi+J. Next, we extracted multimodal non-verbal features, including speech acoustics, facial expressions, backchanneling and gestures, via various pattern recognition techniques and algorithms. Then, we conducted a statistical analysis of listening behaviors and backchannel patterns to identify culturally dependent and independent features in each language and common features among multiple languages. These features were also correlated with the engagement shown by the interlocutors. Finally, we analyzed the influence of cultural differences in the input features of LSTM models trained to predict engagement for five language datasets. A SHAP analysis combined with transfer learning confirmed a considerable correlation between the importance of input features for a language set and the significant cultural characteristics analyzed. Marius Funk, Shogo Okada, Elisabeth André |
ICMI | 2 |
| 2024 | Do We Need To Watch It All? Efficient Job Interview Video Processing with Differentiable MaskingabstractWith technological advancements in transmitting and storing large video files, more and more organizations are incorporating asynchronous video interviews as part of their personnel selection process. Automatic evaluation of these videos is a challenging machine learning setting because the samples are composed of time series input data but only one overall label is available. It is unclear which segments of the time series input (i.e., videos) are the most important ones for prediction. Not all nonverbal features, spoken words, and utterances contribute equally to the prediction; some segments of the videos might even introduce noise to the model. Processing all multimodal information is therefore inefficient. To address this challenge, we propose a framework to model the content of the answer via the full transcription and the speaking patterns of the interviewee via short clips. Our model learns to automatically select the most informative segment by previewing the acoustic modality using a technique called differentiable masking. The results show that our method outperforms existing approaches while being more efficient since only partial multimodal data are processed, and the interpretability of the model is enhanced. Sixia Li, Candy Olivia Mawalim, Hung-Hsuan Huang, Chee Wee Leong, Shogo Okada |
ICMI | 6 |
| 2024 | Automatic mild cognitive impairment estimation from the group conversation of coimagination methodabstractThe coimagination method (Otake, 2009) is designed to prevent dementia in individuals with mild cognitive impairment (MCI) by utilizing the brain’s natural processes. This method involves participants sharing their thoughts and feelings through group conversations centered around shared photos. The coimagination method contains two phases: (1) each participant talk about their memories and experiences related to the photos they bring, and (2) other participants ask questions about the photos. Automating the MCI estimation could be helpful for assisting individuals with MCI during coimagination. However, previous MCI estimation methods rarely focused on group conversation scenarios, despite the potential of multimodal features observed in these scenarios in revealing cognitive states. This study focuses MCI individuals defined by cognitive test scores (e.g., Mini-Mental State Examination (MMSE)). We explore MCI estimation from three aspects. First, we clarify whether MCI can be effectively estimated by constructing estimation models using linguistic and acoustic features from coimagination sessions. Second, we evaluate the impact of using data from the two distinct phases, as they may activate participants’ cognitive functions differently. Finally, we analyze the effects of incorporating subtasks including participants’ conversational customary and engagement level during coimagination via multitask learning. The experimental results demonstrated that individuals with MCI can be effectively estimated from group conversations from coimagination, with the highest macro F1 score of 0.693. The results also demonstrated that the performance improved when using data from the phase that highly activates cognitive functions and when considering conversation customary as a subtask. Sixia Li, Kazumi Kumagai, Mihoko Otake, Shogo Okada |
ICMI | 4 |
| 2024 | Are Recent Deep Learning-Based Speech Enhancement Methods Ready to Confront Real-World Noisy Environments?abstractRecent advancements in speech enhancement techniques have ignited interest in improving speech quality and intelligibility.However, the effectiveness of recently proposed methods is unclear.In this paper, a comprehensive analysis of modern deep learning-based speech enhancement approaches is presented.Through evaluations using the Deep Suppression Noise and Clarity Enhancement Challenge datasets, we assess the performances of three methods: Denoiser, DeepFilterNet3, and FullSubNet+.Our findings reveal nuanced performance differences among these methods, with varying efficacy across datasets.While objective metrics offer valuable insights, they struggle to represent complex scenarios with multiple noise sources.Leveraging ASR-based methods for these scenarios shows promise but may induce critical hallucination effects.Our study emphasizes the need for ongoing research to refine techniques for diverse real-world environments. Candy Olivia Mawalim, Shogo Okada, Masashi Unoki |
INTERSPEECH | 2 |
| 2024 | MBCFNet: A Multimodal Brain-Computer Fusion Network for human intention recognition
Gaoyan Zhang, Shogo Okada, Longbiao Wang, Jianwu Dang 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Adversarial Domain Generalized Transformer for Cross-Corpus Speech Emotion RecognitionabstractSpeech emotion recognition (SER) promotes the development of intelligent devices, which enable natural and friendly human-computer interactions. However, the recognition performance of existing approaches is significantly reduced on unseen datasets, and the lack of sufficient training data limits the generalizability of deep learning models. In this work, we analyze the impact of the domain generalization method on cross-corpus SER and propose an adversarial domain generalized transformer (ADoGT), which is aimed at learning a shared feature distribution for the source and target domains. Specifically, we investigate the effect of domain adversarial learning by eliminating nonaffective information. We also combine the center loss with the softmax function as joint supervision to learn discriminative features. Moreover, we introduce unsupervised transfer learning to extract additional features, and incorporate a gated fusion model to learn the complementary information of the features learned by the supervised feature extractor and pretrained model. The proposed transformer based domain generalization method is evaluated using four emotional datasets. We also provide an ablation study of different domain adversarial model structures and feature fusion models. The results of comparative experiments demonstrate the effectiveness of the proposed ADoGT. Yuan Gao 0040, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001, Shogo Okada |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Adaptive Interview Strategy Based on Interviewees' Speaking Willingness Recognition for Interview RobotsabstractSocial signal recognition techniques based on nonverbal behavioral sensing allow conversational robots to understand the user's social signals, thereby enabling them to adopt interaction strategies based on internal states inferred from the social signals. This research investigates how the online social signal recognition and adaptive dialog strategy influences the dynamic change in a user's inner state. For this purpose, we develop a semiautonomous interview robot system with an online speaker's willingness recognition module and an adaptive question selection module based on the willingness level. The online recognition model of speaker willingness is trained from multimodal nonverbal features extracted using a novel interview corpus, which allows appropriate interview questions to be chosen based on the estimated willingness level of the user. We conduct the experiment using the system to evaluate the effectiveness of adaptive question selection based on the willingness recognition model. First, the multimodal willingness recognition model is evaluated using the interview corpus. The best recognition accuracy of willingness level (high or low) was 72:8% with the random forest classifier. Second, 27 interviewees were interviewed with the two interview robot systems: (I) with the adaptive question selection module based on willingness recognition and (II) with a random question selection strategy. The proposed adaptive question strategy significantly increased the number of utterances with high willingness compared with the baseline system (II); thus, adaptive question selection with online willingness recognition elicited the speaker's willingness even though the model cannot be estimated with near-perfect accuracy. Fuminori Nagasawa, Shogo Okada, Takuya Ishihara, Katsumi Nitta |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | FedCPC: An Effective Federated Contrastive Learning Method for Privacy Preserving Early-Stage Alzheimers Speech DetectionabstractThe early-stage Alzheimer’s disease (AD) detection has been considered an important field of medical studies. Like traditional machine learning methods, speech-based automatic detection also suffers from data privacy risks because the data of specific patients are exclusive to each medical institution. A common practice is to use federated learning to protect the patients’ data privacy. However, its distributed learning process also causes performance reduction. To alleviate this problem while protecting user privacy, we propose a federated contrastive pre-training (FedCPC) performed before federated training for AD speech detection, which can learn a better representation from raw data and enables different clients to share data in the pre-training and training stages. Experimental results demonstrate that the proposed methods can achieve satisfactory performance while preserving data privacy. Wenqing Wei, Zhengdong Yang, Yuan Gao 0040, Jiyi Li, Chenhui Chu, Shogo Okada, Sheng Li 0010 |
ASRU | 6 |
| 2023 | Analyzing Differences in Subjective Annotations by Participants and Third-party Annotators in Multimodal Dialogue CorpusabstractEstimating the subjective impressions of human users during a dialogue is necessary when constructing a dialogue system that can respond adaptively to their emotional states.However, such subjective impressions (e.g., how much the user enjoys the dialogue) are inherently ambiguous, and the annotation results provided by multiple annotators do not always agree because they depend on the subjectivity of the annotators.In this paper, we analyzed the annotation results using 13,226 exchanges from 155 participants in a multimodal dialogue corpus called Hazumi that we had constructed, where each exchange was annotated by five third-party annotators.We investigated the agreement between the subjective annotations given by the third-party annotators and the participants themselves, on both perexchange annotations (i.e., participant's sentiments) and per-dialogue (-participant) annotations (i.e., questionnaires on rapport and personality traits).We also investigated the conditions under which the annotation results are reliable.Our findings demonstrate that the dispersion of third-party sentiment annotations correlates with agreeableness of the participants, one of the Big Five personality traits. Kazunori Komatani, Ryu Takeda, Shogo Okada |
SIGDIAL | 3 |
| 2023 | Effects of Physiological Signals in Different Types of Multimodal Sentiment EstimationabstractMultimodal sentiment analysis has become a focus of research in recent years. However, most studies of multimodal sentiment analysis have considered only signals that are observable by humans, such as linguistic, audio and visual information, whereas the contribution of the multimodal fusion of such signals with unobservable signals, i.e., physiological signals, has not been comprehensively explored. In this study, we aim to investigate effects of physiological signals in multimodal sentiment analysis by evaluating all of the fusion models for different types of sentiment estimation in naturalistic human-agent interaction settings. Our results suggest that physiological features are effective in the unimodal model and that the fusion of linguistic representations with physiological features provides the best results for estimating self-sentiment labels as annotated by the users themselves. In contrast, the tensor fusion of linguistic representations with audiovisual features is effective for estimating sentiment labels as annotated by a third party in regression tasks, which can be derived from the corresponding signals that are observable by humans. A detailed analysis of the self-sentiment estimation results suggests that different modalities play different roles in sentiment estimation, and corresponding implications are discussed. Shun Katada, Shogo Okada, Kazunori Komatani |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Domain-Invariant Feature Learning for Cross Corpus Speech Emotion RecognitionabstractTo deal with speech emotion recognition (SER) in real-life applications, researchers have to focus on cross corpus SER, where the feature distribution of source and target datasets are different. In this paper, we propose an efficient domain adversarial training method to cope with the non-affective information during feature extraction. Through the proposed domain-adversarial learning, we can reduce the domain divergence between train and test data. Furthermore, we incorporate center loss with the emotion classifier to reduce the intra-class variation of features learned from the same emotion. We conduct experiments on four emotional benchmark datasets to verify the performance of the proposed method. The experimental results demonstrate that our proposed model outperform the baseline system in both cross-corpus and multi-corpus evaluation. Yuan Gao 0040, Shogo Okada, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICASSP | 2 |
| 2022 | Transformer-Based Physiological Feature Learning for Multimodal Analysis of Self-Reported SentimentabstractOne of the main challenges in realizing dialog systems is adapting to a user’s sentiment state in real time. Large-scale language models, such as BERT, have achieved excellent performance in sentiment estimation; however, the use of only linguistic information from user utterances in sentiment estimation still has limitations. In fact, self-reported sentiment is not necessarily expressed by user utterances. To mitigate the issue that the true sentiment state is not expressed as observable signals, psychophysiology and affective computing studies have focused on physiological signals that capture involuntary changes related to emotions. We address this problem by efficiently introducing time-series physiological signals into a state-of-the-art language model to develop an adaptive dialog system. Compared with linguistic models based on BERT representations, physiological long short-term memory (LSTM) models based on our proposed physiological signal processing method have competitive performance. Moreover, we extend our physiological signal processing method to the Transformer language model and propose the Time-series Physiological Transformer (TPTr), which captures sentiment changes based on both linguistic and physiological information. In ensemble models, our proposed methods significantly outperform the previous best result (p < 0.05). Shun Katada, Shogo Okada, Kazunori Komatani |
ICMI | 2 |
| 2022 | Detecting Change Talk in Motivational Interviewing using Verbal and Facial InformationabstractBehavior change is one of the most important goals in psychotherapy. This study focuses on Motivational Interviewing (MI), which is collaborative communication aimed at eliciting the client’s own reasons for behavior change. To investigate the effectiveness of facial information in modeling MI, we collected an MI encounter corpus with speech and video data in the nutrition and fitness domains and annotated client utterances using the Manual for the Motivational Interviewing Skill Code (MISC). By analyzing client answers to the questions after the session, we found that clients who expressed more Change Talk were more motivated to change their behavior than those who expressed less Change Talk. We then proposed RNN-based multimodal models to detect Change Talk by setting a 2-class classification task: "Change Talk" and "not Change Talk." Our experiment showed that the best performing model was a multimodal BiLSTM model that fused language and client facial information. We also found that fusing language and facial information as context achieved better performance than the unimodal and no-context models. Moreover, we discuss the label imbalance problem and conduct an additional analysis using turns as a unit of analysis. As a result, our best model reached F1-score of 0.65 for Change Talk detection. Yukiko I. Nakano, Eri Hirose, Tatsuya Sakato, Shogo Okada, Jean-Claude Martin |
ICMI | 4 |
| 2022 | Investigating the relationship between dialogue and exchange-level impressionabstractMultimodal dialogue systems (MDS) have recently attracted increasing attention. The automatic evaluation of user impression with spoken dialog at the dialog level plays a central role in managing dialog systems. A user usually forms an overall impression through the experience of each exchange of turns in the conversation. Thus, the user’s exchange-level sentiment should be considered when recognizing the user’s overall impression of the dialog. Previous research has focused on modeling user impressions during individual exchanges or during the overall conversation. Thus, the relationship between user sentiment at the exchange level and user impression at the dialog level is still unclear, and appropriately utilizing this relationship in impression analysis remains unexplored. In this paper, we first investigate the relation between sentiment at the exchange level and 18 labels that indicate different aspects of the user impression at the dialog level. Then, we present a multitask learning model (MTL) that uses exchange-level annotations to recognize dialog-level labels. The experimental results demonstrate that our proposed model achieves better performance at the dialog level, outperforming the single-task model by a maximum of 15.7%. Wenqing Wei, Sixia Li, Shogo Okada |
ICMI | 3 |
| 2022 | IDPS Signature Classification with a Reject Option and the Incorporation of Expert KnowledgeabstractAs the importance of intrusion detection and prevention systems (IDPSs) increases, great costs are being incurred to manually manage signatures, which are malicious communication pattern files. Network security experts need to classify signatures by importance for an IDPS to work optimally. In this study, we propose and evaluate a machine learning signature classification model with a reject option (RO) to reduce the cost of setting up an IDPS. Experts classify some signatures with predefined if-then rules. We first design two types of features, symbolic features (SFs) and keyword features (KFs), which are used in keyword matching for the if-then rules. Next, we design web information and message features (WMFs) to capture the properties of signatures that do not match the if-then rules. The WMFs are extracted as the term frequency-inverse document frequency (TF-IDF) features of the message text in each signature and the text expanded by web scraping, respectively. Because failures need to be minimized when classifying IDPS signatures, we consider introducing an RO in our proposed model. The effectiveness of the proposed classification model is evaluated in experiments conducted on two real datasets composed of signatures labeled by experts. In the experiment, the SF and WMF combination outperforms the SF and KF combination. We also show that using a deep ensemble improves the performance of the RO. An analysis shows that experts refer to the text contained in the signatures and information from the web. Hidetoshi Kawaguchi, Yuichi Nakatani, Shogo Okada |
ICMLA | 3 |
| 2022 | Multimodal Analysis for Communication Skill and Self-Efficacy Level Estimation in Job Interview ScenarioabstractAn interview for a job recruiting process requires applicants to demonstrate their communication skills. Interviewees sometimes become nervous about the interview because interviewees themselves do not know their assessed score. This study investigates the relationship between the communication skill (CS) and the self-efficacy level (SE) of interviewees through multimodal modeling. We also clarify the difference between effective features in the prediction of CS and SE labels. For this purpose, we collect a novel multimodal job interview data corpus by using a job interview agent system where users experience the interview using a virtual reality head-mounted display (VR-HMD). The data corpus includes annotations of CS by third-party experts and SE annotations by the interviewees. The data corpus also includes various kinds of multimodal data, including audio, biological (i.e., physiological), gaze, and language data. We present two types of regression models, linear regression and sequential-based regression models, to predict CS, SE, and the gap (GA) between skill and self-efficacy. Finally, we report that the model with acoustic, gaze, and linguistic features has the best regression accuracy in CS prediction (correlation coefficient r = 0.637). Furthermore, the regression model with biological features achieves the best accuracy in SE prediction (r = 0.330). Tomoya Ohba, Candy Olivia Mawalim, Shun Katada, Haruki Kuroki, Shogo Okada |
MUM | 5 |
| 2022 | Biosignal-based user-independent recognition of emotion and personality with importance weighting
Shun Katada, Shogo Okada |
Multim. Tools Appl. | 2 |
| 2021 | Multimodal Human-Agent Dialogue Corpus with Annotations at Utterance and Dialogue LevelsabstractThe behaviors of general users for a dialogue system differ greatly from those for a human interlocutor. We have been collecting a multimodal dialogue corpus between a human participant and a virtual agent operated by the Wizard-of-Oz method. This paper presents the collected corpus, Hazumi, which was released in August 2020 and March 2021. The corpus consists of three versions: Hazumi1712, Hazumi1902, and Hazumi1911. The version numbers correspond to the periods during which the data were collected. The three versions contain the dialogue data of 29, 30, and 30 participants, respectively, each of whom spoke with the agent for about 15 to 20 minutes. The corpus contains multimodal recordings, along with several annotations given to every exchange, feature files extracted from the recorded data, and the results of questionnaires conducted before and after the dialogues. The third version Hazumi1911 also contains the physiological signals of the participants during the dialogues and additional questionnaire items. We also show several analyses conducted using this corpus. We anticipate that the corpus will be useful for developing user-adaptive multimodal dialogue systems. Kazunori Komatani, Shogo Okada |
ACII | 2 |
| 2021 | CONSK-GCN: Conversational Semantic- and Knowledge-Oriented Graph Convolutional Network for Multimodal Emotion RecognitionabstractEmotion recognition in conversations (ERC) has received significant attention in recent years due to its widespread applications in diverse areas, such as social media, health care, and artificial intelligence interactions. However, different from nonconversational text, it is particularly challenging to model the effective context-aware dependence for the task of ERC. To address this problem, we propose a new Conversational Semantic- and Knowledge-oriented Graph Convolutional Network (ConSK-GCN) approach that leverages both semantic dependence and commonsense knowledge. First, we construct the contextual inter-interaction and intradependence of the interlocutors via a conversational graph-based convolutional network based on multimodal representations. Second, we incorporate commonsense knowledge to guide ConSK-GCN to model the semantic-sensitive and knowledge-sensitive contextual dependence. The results of extensive experiments show that the proposed method outperforms the current state of the art on the IEMOCAP dataset. Yahui Fu 0001, Shogo Okada, Longbiao Wang, Lili Guo 0001, Yaodong Song, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICME | 2 |
| 2021 | Recognizing Social Signals with Weakly Supervised Multitask Learning for Multimodal Dialogue SystemsabstractSocial signal processing is a methodology that is used to infer human inner states, including attitudes, sentiments and impressions, from verbal and nonverbal multimodal information. The difficulty in training a social signal recognition model is that the ground-truth (target) labels given by multiple coders often disagree because the annotation of social signals such as sentiments is a subjective and ambiguous task. We introduce weakly supervised learning (WSL) algorithms to such an inaccurate supervision setting in which the target label is not necessarily accurate. The novel challenge in this paper is to explore an effective WSL strategy for recognizing social signals. The strategy is verified through two multimodal datasets including audio, visual, and linguistic data collected in a human-agent dialogue setting. First, we clarify that the proposed WSL strategy for deep neural networks (DNNs), called tri-teaching works well in almost all classification tasks. Second, we demonstrate the effectiveness of integrating WSL and multitask learning (MTL), which exploits several label types in the datasets. Third, we show that our proposed approach achieves less accuracy degradation than an existing training algorithm for a DNN (curriculum learning) in a cross-corpus setting, with a maximum improvement of 7.2%. Yuki Hirano, Shogo Okada, Kazunori Komatani |
ICMI | 2 |
| 2021 | Multimodal User Satisfaction Recognition for Non-task Oriented Dialogue SystemsabstractMultimodal dialogue systems (MDSs) are needed to allow users to converse with virtual agents that use natural language by sensing the multimodal behavior of users. One crucial step in the development of an MDS is measuring how well the dialogue system performs. Though previous research focused on the user satisfaction modeling from linguistic modality in text-to-text dialogue systems, the user satisfaction is observed by not only spoken dialogue contents but also the acoustic and visual nonverbal behaviors of users. Multimodal social signal sensing provides a solution that automatically measures dialogue systems based on subjective evaluation. With this background, we proposed a multimodal recognition model of the user using sequence modeling algorithms (RNN, LSTM, and GRU). It is a novel challenge to recognize the user satisfaction label at the dialogue level. Each label was annotated by the user based on the overall dialogue. We extracted both verbal features and nonverbal features at the exchange level (the unit is a pair of system and user utterances) and analyzed the contributions of multimodal features and unimodal features to recognize user satisfaction labels at the dialogue level. We used a multimodal user-system dialogue data corpus with user satisfaction labels at the dialogue level. To validate the recognition accuracy of the proposed multimodal modeling approach, we compared the proposed method with two models based on human perception by external human coders and the system operator (called “Wizard”) with whom the user talks. The experimental results showed that the multimodal model achieved a better performance in both classification and regression tasks. The results indicated that the performance of the multimodal model was higher than that of the human models. Wenqing Wei, Sixia Li, Shogo Okada, Kazunori Komatani |
ICMI | 3 |
| 2021 | Task-independent Recognition of Communication Skills in Group Interaction Using Time-series ModelingabstractCase studies of group discussions are considered an effective way to assess communication skills (CS). This method can help researchers evaluate participants’ engagement with each other in a specific realistic context. In this article, multimodal analysis was performed to estimate CS indices using a three-task-type group discussion dataset, the MATRICS corpus. The current research investigated the effectiveness of engaging both static and time-series modeling, especially in task-independent settings. This investigation aimed to understand three main points: first, the effectiveness of time-series modeling compared to nonsequential modeling; second, multimodal analysis in a task-independent setting; and third, important differences to consider when dealing with task-dependent and task-independent settings, specifically in terms of modalities and prediction models. Several modalities were extracted (e.g., acoustics, speaking turns, linguistic-related movement, dialog tags, head motions, and face feature sets) for inferring the CS indices as a regression task. Three predictive models, including support vector regression (SVR), long short-term memory (LSTM), and an enhanced time-series model (an LSTM model with a combination of static and time-series features), were taken into account in this study. Our evaluation was conducted by using the R 2 score in a cross-validation scheme. The experimental results suggested that time-series modeling can improve the performance of multimodal analysis significantly in the task-dependent setting (with the best R 2 = 0.797 for the total CS index), with word2vec being the most prominent feature. Unfortunately, highly context-related features did not fit well with the task-independent setting. Thus, we propose an enhanced LSTM model for dealing with task-independent settings, and we successfully obtained better performance with the enhanced model than with the conventional SVR and LSTM models (the best R 2 = 0.602 for the total CS index). In other words, our study shows that a particular time-series modeling can outperform traditional nonsequential modeling for automatically estimating the CS indices of a participant in a group discussion with regard to task dependency. Candy Olivia Mawalim, Shogo Okada, Yukiko I. Nakano |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Is She Truly Enjoying the Conversation?: Analysis of Physiological Signals toward Adaptive Dialogue SystemsabstractIn human-agent interactions, it is necessary for the systems to identify the current emotional state of the user to adapt their dialogue strategies. Nevertheless, this task is challenging because the current emotional states are not always expressed in a natural setting and change dynamically. Recent accumulated evidence has indicated the usefulness of physiological modalities to realize emotion recognition. However, the contribution of the time series physiological signals in human-agent interaction during a dialogue has not been extensively investigated. This paper presents a machine learning model based on physiological signals to estimate a user's sentiment at every exchange during a dialogue. Using a wearable sensing device, the time series physiological data including the electrodermal activity (EDA) and heart rate in addition to acoustic and visual information during a dialogue were collected. The sentiment labels were annotated by the participants themselves and by external human coders for each exchange consisting of a pair of system and participant utterances. The experimental results showed that a multimodal deep neural network (DNN) model combined with the EDA and visual features achieved an accuracy of 63.2%. In general, this task is challenging, as indicated by the accuracy of 63.0% attained by the external coders. The analysis of the sentiment estimation results for each individual indicated that the human coders often wrongly estimated the negative sentiment labels, and in this case, the performance of the DNN model was higher than that of the human coders. These results indicate that physiological signals can help in detecting the implicit aspects of negative sentiments, which are acoustically/visually indistinguishable. Shun Katada, Shogo Okada, Yuki Hirano, Kazunori Komatani |
ICMI | 2 |
| 2019 | Dementia Scale Classification Based on Ubiquitous Daily Activity and Interaction SensingabstractThis paper investigates the integration of different approaches to automatically predict high/low-score on the dementia scale. We propose two different approaches to predict this value by capturing the following: (1) the participant's interaction behavior with a humanoid robot and (2) the indoor daily activity in the residence using ubiquitous sensors. The interaction and indoor activity data set were obtained by recording 32 participants living in common residences, including 17 with symptoms of dementia, as indicated through a cognitive test (Revised Hasegawa Dementia Scale). To obtain the interaction features, we extracted the turn-taking features of interaction with a mobile-typed humanoid robot. To extract the indoor activity features, we collected the location data of each participant in the residence using the received signal strength indicators (RSSIs) of Bluetooth signals from different access points (e.g., shared spaces or the participant's room). In the experimental evaluation, we trained binary classification models for classifying the score on the dementia scale from these datasets. The results show that the best classification accuracy (0.875) is achieved when interaction and activity features are fused using a random forest classifier. Shogo Okada, Ken Inoue, Toru Imai, Mami Noguchi, Kaiko Kuwamura |
ACII | 1 |
| 2019 | Analyzing Eye Movements in Interview Communication with Virtual Reality AgentsabstractIn human-agent interactions, human emotions and gestures expressed when interacting with agents is a high-level personally trait that quantifies human attitudes, intentions, motivations, and behaviors. The virtual reality space provides a chance to interact with virtual agents in a more immersive way. In this paper, we present a computational framework to analyze human eye movements by using a virtual reality system in a job interview scene. First, we developed a remote interview system using virtual agents and implemented the system into a virtual reality headset. Second, by tracking eye movements and collecting other multimodal data, the system could better analyze human personality traits in interview communication with virtual agents, and it could better support training in people's communication skills. In experiments, we analyzed the relationship between eye gaze feature and interview performance annotated by human experts. Experimental results showed acceptable accuracy value for the single modality of eye movement in the prediction of eye contact and total performance in job interviews. Fuhui Tian, Shogo Okada, Katsumi Nitta |
HAI | 2 |
| 2019 | Multitask Prediction of Exchange-level Annotations for Multimodal Dialogue SystemsabstractThis paper presents multimodal computational modeling of three labels that are independently annotated per exchange to implement an adaptation mechanism of dialogue strategy in spoken dialogue systems based on recognizing user sentiment by multimodal signal processing. The three labels include (1) user’s interest label pertaining to the current topic, (2) user’s sentiment label, and (3) topic continuance denoting whether the system should continue the current topic or change it. Predicting the three types of labels that capture different aspects of the user’s sentiment level and the system’s next action contribute to adopting a dialogue strategy based on the user’s sentiment. For this purpose, we enhanced shared multimodal dialogue data by annotating impressed sentiment labels and the topic continuance labels. With the corpus, we develop a multimodal prediction model for the three labels. A multitask learning technique is applied for binary classification tasks of the three labels considering the partial similarities among them. The prediction model was efficiently trained even with a small data set (less than 2000 samples) thanks to the multitask learning framework. Experimental results show that the multitask deep neural network (DNN) model trained with multimodal features including linguistics, facial expressions, body and head motions, and acoustic features, outperformed those trained as single-task DNNs by 1.6 points at the maximum. Yuki Hirano, Shogo Okada, Haruto Nishimoto, Kazunori Komatani |
ICMI | 2 |
| 2019 | Interaction Process Label Recognition in Group DiscussionabstractIn qualifying and analyzing the performance of group interaction, interaction processing analysis (IPA) defined by Bale is considered a useful approach. IPA is a system for labeling a total of 12 interaction categories for the interaction process. Automatic IPA can manually encompass the gap in spending manpower and can efficiently qualify group performance. In this paper, we present computational interaction processing analysis by developing a model to recognize categories of IPA. We extract both verbal features and nonverbal features for IPA category recognition modeling with SVM, RF, DNN and LSTM machine learning algorithms and analyze the contribution of multimodal features and unimodal features for the total data and each label. We also investigate the effect of context information by training sequences with different lengths with an LSTM and evaluating them. The results show that multimodal features achieve the best performance with an F1 score of 0.601 for the recognition of 12 IPA categories using the total data. Multimodal features are better than the unimodal features for the total data and most labels. The results of investigating context information show that a suitable length of sequence enables a longer sequence to achieve the best F1 score of 0.602 and a better performance for recognition. Sixia Li, Shogo Okada, Jianwu Dang 0001 |
ICMI | 2 |
| 2019 | Task-independent Multimodal Prediction of Group Performance Based on Product DimensionsabstractThis paper proposes an approach to develop models for predicting the performance for multiple group meeting tasks, where the model has no clear correct answer. This paper adopts ”product dimensions” [Hackman et al. 1967] (PD) which is proposed as a set of dimensions for describing the general properties of written passages that are generated by a group, as a metric measuring group output. This study enhanced the group discussion corpus called the MATRICS corpus including multiple discussion sessions by annotating the performance metric of PD. We extract group-level linguistic features including vocabulary level features using a word embedding technique, topic segmentation techniques, and functional features with dialog act and parts of speech on the word level. We also extracted nonverbal features from the speech turn, prosody, and head movement. With a corpus including multiple discussion data and an annotation of the group performance, we conduct two types of experiments thorough regression modeling to predict the PD. The first experiment is to evaluate the task-dependent prediction accuracy, in the situation that the samples obtained from the same discussion task are included in both the training and testing. The second experiments is to evaluate the task-independent prediction accuracy, in the situation that the type of discussion task is different between the training samples and testing samples. In this situation, regression models are developed to infer the performance in an unknown discussion task. The experimental results show that a support vector regression model archived a 0.76 correlation in the discussion-task-dependent setting and 0.55 in the task-independent setting. Go Miura, Shogo Okada |
ICMI | 2 |
| 2019 | Modeling Dyadic and Group Impressions with Intermodal and Interperson FeaturesabstractThis article proposes a novel feature-extraction framework for inferring impression personality traits, emergent leadership skills, communicative competence, and hiring decisions. The proposed framework extracts multimodal features, describing each participant’s nonverbal activities. It captures intermodal and interperson relationships in interactions and captures how the target interactor generates nonverbal behavior when other interactors also generate nonverbal behavior. The intermodal and interperson patterns are identified as frequent co-occurring events based on clustering from multimodal sequences. The proposed framework is applied to the SONVB corpus, which is an audiovisual dataset collected from dyadic job interviews, and the ELEA audiovisual data corpus, which is a dataset collected from group meetings. We evaluate the framework on a binary classification task involving 15 impression variables from the two data corpora. The experimental results show that the model trained with co-occurrence features is more accurate than previous models for 14 out of 15 traits. Shogo Okada, Laurent Son Nguyen, Oya Aran, Daniel Gatica-Perez |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Investigating Effectiveness of Linguistic Features Based on Speech Recognition for Storytelling Skill Assessment
Shogo Okada, Kazunori Komatani |
IEA/AIE | 1 |
| 2018 | Collection of Multimodal Dialog Data and Analysis of the Result of Annotation of Users' Interest Level
Masahiro Araki, Sayaka Tomimasu, Mikio Nakano, Kazunori Komatani, Shogo Okada, Shinya Fujie, Hiroaki Sugiyama |
LREC | 5 |
| 2017 | Recognizing Words from Gestures: Discovering Gesture Descriptors Associated with Spoken UtterancesabstractThis study investigates a new challenge: modelling the relationship between gestures and spoken words during continuous hand motions, speech and language data. The problem setting and modelling is defined as “gesture association” (associating gestures with words). We present a framework to associate spoken words with hand motion data observed from a optical motion-capture system. The framework identifies pairs of hand motions (gesture) and uttered words as training samples by autonomously aligning the speech and the co-occurring gestures. Using the samples, a supervised learning approach is undertaken to learn a model to discriminate between gestures that co-occur with a spoken word (w) and gestures when the word w is not being spoken. To detect gestures and associate them with spoken words, we extract (1) a trajectory signal feature set, (2) features of various gesture phases from hand motion data and (3) primitive gesture patterns learned by Sift Invariant Sparse Coding (SISC). Then Hidden Markov Model (HMM) and Support Vector Machine (SVM) classifiers were trained with these three-feature sets. In experiments, the classification accuracy achieved for 80 words was above 60 % (the maximum accuracy achieved was 71 %) using the proposed framework. In particular, SVMs trained with features (2) and (3) outperform an HMM trained with feature (1) (prepared as a baseline model). The results show that the gesture phase and primitive patterns trained by SISC are effective gesture features for recognizing the words accompanying gestures. Shogo Okada, Kazuhiro Otsuka |
FG | 1 |
| 2017 | Proposal of a Model to Determine the Attention Target for an Agent in Group Discussion with Non-verbal FeaturesabstractIn recent years, companies are seeking for communication skill from their employers. More and more companies adopt group discussions in employer recruitment to evaluate the ap- plicants' communication skill. However, the opportunity to improve communication skill in group discussion is limited due to the lack of partners. In order to solve this issue, our ongoing project is aiming to build a virtual agent or a robot that can participate group discussion, so that its users can re- peatedly practice group discussion with it. In this paper, we propose the models in directing the agent's attention toward the other participants in three situations:when the agent is speaking, when the agent is listening, and when no partic- ipant is speaking. First, we gathered a data corpus of the discussion of 10 four-people groups. We then use low-level non-verbal features including attention of other participant, voice prosody, head movements, and speech turn extracted in the 10-hour corpus to train support vector machine models to determine the agent's attention on the other participants, or the material. The performance of the detection models in F-measure range between 0.4 and 0.6. Seiya Kimura, Hung-Hsuan Huang, Shogo Okada, Naoki Ohta, Kazuhiro Kuwabara |
HAI | 4 |
| 2017 | Weibull partition models with applications to hidden semi-Markov modelsabstractWe develop the Weibull partition model (WPM), which defines a novel nonparametric stochastic process over distributions of partitions of sequential data, aiming at directly modeling the boundaries of segments comprising the sequence. The Weibull partition model employs a Dirichlet process mixture with a Weibull kernel. Weibull distributions having a closed-form cumulative density function plays an important role in the construction of the Weibull partition model. As an application of our model, we propose the hidden semi-Markov model based on the Weibull partition model (WPM-HSMM), along with the corresponding recursive sampling algorithm. The WPM-HSMM can be used to solve the problem of low accuracy of inference for hidden states, caused by the holding times of hidden states having a complicated distribution. In our experiments, we show that the WPM-HSMM rapidly mixes and achieves very competitive results compared to the state-of-the-art algorithms. Apart from the application to hidden Markov models, the WPM can also be found useful in any problem, where explicitly modeling segment boundaries is needed. Youwei Lu, Shogo Okada, Katsumi Nitta |
IJCNN | 2 |
| 2016 | Estimating communication skills using dialogue acts and nonverbal features in multiple discussion datasetsabstractThis paper focuses on the computational analysis of the individual communication skills of participants in a group. The computational analysis was conducted using three novel aspects to tackle the problem. First, we extracted features from dialogue (dialog) act labels capturing how each participant communicates with the others. Second, the communication skills of each participant were assessed by 21 external raters with experience in human resource management to obtain reliable skill scores for each of the participants. Third, we used the MATRICS corpus, which includes three types of group discussion datasets to analyze the influence of situational variability regarding to the discussion types. We developed a regression model to infer the score for communication skill using multimodal features including linguistic and nonverbal features: prosodic, speaking turn, and head activity. The experimental results show that the multimodal fusing model with feature selection achieved the best accuracy, 0.74 in R2 of the communication skill. A feature analysis of the models revealed the task-dependent and task-independent features to contribute to the prediction performance. Shogo Okada, Yoshihiko Ohtake, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Yutaka Takase, Katsumi Nitta |
ICMI | 1 |
| 2016 | Fashion image classification on mobile phones using layered deep convolutional neural networksabstractToward implementation of fashion recommendation system based on photos taken with mobile phones, we propose a framework to recognize hierarchical categories of a fashion item. To classify an arbitrary photo of clothes robustly, (1) we collected two kind of dataset: (I) 120K datasets of clothes images on EC sites to train the classifier and (II) the clothing image set composed of photos taken by participants with their mobile phones. (2) we proposed Layered Deep Convolutional Neural Networks (LDCNNs) which is specialized in classifying images into hierarchical categories: hoodie is a lower in hierarchy in tops category. Experimental result shows proposed LDCNNs obtained mean accuracy of 92.7% for datasets from EC sites and 96.9% for those from mobile phones. This results are better than those (84.9%, 90.6 % respectively) for MLR+CNN in classification accuracy. Kazunori Hori, Shogo Okada, Katsumi Nitta |
MUM | 2 |
| 2015 | Predicting Participation Styles using Co-occurrence Patterns of Nonverbal Behaviors in Collaborative LearningabstractWith the goal of assessing participant attitudes and group activities in collaborative learning, this study presents models of participation styles based on co-occurrence patterns of nonverbal behaviors between conversational participants. First, we collected conversations among groups of three people in a collaborative learning situation, wherein each participant had a digital pen and wore a glasses-type eye tracker. We then divided the collected multimodal data into 0.1-second intervals. The discretized data were applied to an unsupervised method to find co-occurrence behavioral patterns. As a result, we discovered 122 multimodal behavioral motifs from more than 3,000 possible combinations of behaviors by three participants. Using the multimodal behavioral motifs as predictor variables, we created regression models for assessing participation styles. The multiple correlation coefficients ranged from 0.74 to 0.84, indicating a good fit between the models and the data. A correlation analysis also enabled us to identify a smaller set of behavioral motifs (fewer than 30) that are statistically significant as predictors of participation styles. These results show that automatically discovered combinations of multiple kinds of nonverbal information with high co-occurrence frequencies observed between multiple participants as well as for a single participant are useful in characterizing the participant's attitudes towards the conversation. Yukiko I. Nakano, Sakiko Nihonyanagi, Yutaka Takase, Yuki Hayashi, Shogo Okada |
ICMI | 5 |
| 2015 | Personality Trait Classification via Co-Occurrent Multiparty Multimodal Event DiscoveryabstractThis paper proposes a novel feature extraction framework from mutli-party multimodal conversation for inference of personality traits and emergent leadership. The proposed framework represents multi modal features as the combination of each participant's nonverbal activity and group activity. This feature representation enables to compare the nonverbal patterns extracted from the participants of different groups in a metric space. It captures how the target member outputs nonverbal behavior observed in a group (e.g. the member speaks while all members move their body), and can be available for any kind of multiparty conversation task. Frequent co-occurrent events are discovered using graph clustering from multimodal sequences. The proposed framework is applied for the ELEA corpus which is an audio visual dataset collected from group meetings. We evaluate the framework for binary classification task of 10 personality traits. Experimental results show that the model trained with co-occurrence features obtained higher accuracy than previously related work in 8 out of 10 traits. In addition, the co-occurrence features improve the accuracy from 2 % up to 17 %. Shogo Okada, Oya Aran, Daniel Gatica-Perez |
ICMI | 1 |
| 2014 | Predicting Influential Statements in Group Discussions using Speech and Head Motion InformationabstractGroup discussions are used widely when generating new ideas and forming decisions as a group. Therefore, it is assumed that giving social influence to other members through facilitating the discussion is an important part of discussion skill. This study focuses on influential statements that affect discussion flow and highly related to facilitation, and aims to establish a model that predicts influential statements in group discussions. First, we collected a multimodal corpus using different group discussion tasks; in-basket and case-study. Based on schemes for analyzing arguments, each utterance was annotated as being influential or not. Then, we created classification models for predicting influential utterances using prosodic features as well as attention and head motion information from the speaker and other members of the group. In our model evaluation, we discovered that the assessment of each participant in terms of discussion facilitation skills by experienced observers correlated highly to the number of influential utterances by a given participant. This suggests that the proposed model can predict influential statements with considerable accuracy, and the prediction results can be a good predictor of facilitators in group discussions. Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Shogo Okada |
ICMI | 5 |
| 2013 | A discussion training support system and its evaluationabstractIn law schools, to teach argumentation skills, training for discussion is sometimes conducted. However, giving advice and evaluating each discussion takes long time. This paper introduces the overview of a discussion analysis tool. It helps instructors to analyze the discussion records by observing topic flow of the discussion and by comparing it with that of another record. This system translates the discussion records into a form of a time-sequenced speech acts and an argumentation diagram, and then calculates several features of the discussion record which are used to evaluate such a record. The overview of the system and its evaluation are introduced here. Shumpei Kubosawa, Kei Nishina, Masaki Sugimoto, Shogo Okada, Katsumi Nitta |
ICAIL | 4 |
| 2013 | Context-based conversational hand gesture classification in narrative interactionabstractCommunicative hand gestures play important roles in face-to-face conversations. These gestures are arbitrarily used depending on an individual; even when two speakers narrate the same story, they do not always use the same hand gesture (movement, position, and motion trajectory) to describe the same scene. In this paper, we propose a framework for the classification of communicative gestures in small group interactions. We focus on how many times the hands are held in a gesture and how long a speaker continues a hand stroke, instead of observing hand positions and hand motion trajectories. In addition, to model communicative gesture patterns, we use nonverbal features of participants addressed from participant gestures. In this research, we extract features of gesture phases defined by Kendon (2004) and co-occurring nonverbal patterns with gestures, i.e., utterance, head gesture, and head direction of each participant, by using pattern recognition techniques. In the experiments, we collect eight group narrative interaction datasets to evaluate the classification performance. The experimental results show that gesture phase features and nonverbal features of other participants improves the performance to discriminate communicative gestures that are used in narrative speeches and other gestures from 4% to 16%. Shogo Okada, Mayumi Bono, Katsuya Takanashi, Yasuyuki Sumi, Katsumi Nitta |
ICMI | 1 |
| 2013 | Semi-supervised Latent Dirichlet Allocation for Multi-label Text Classification
Youwei Lu, Shogo Okada, Katsumi Nitta |
IEA/AIE | 2 |
| 2013 | A Semantic-Similarity-Based Method for Object Description and ClusteringabstractObject recognition and clustering are useful techniques in pattern recognition and computer vision. Traditionally, these techniques have been implemented by visual-feature-based methods. However, these methods may not adequately tackle the differences in the shapes and colors of objects. In this paper, we propose an alternative method in which objects of different colors, or even different shapes, function similarly. If text strings are visible on their surfaces, we can extract the semantic features of objects, thereby recognizing and clustering them. Thus, this method is based on semantic information. The method is experimentally tested on a dataset of images containing the packing cases of commercial products. Semantic information in the dataset images is retrieved using text extraction modules, passed through an Internet data mining module and is finally described and clustered. The final clustering results are more accurate than those obtained by visual-feature-based methods. Shogo Okada, Katsumi Nitta |
SMC | 2 |
| 2012 | Analysis of the correlation between the regularity of work behavior and stress indices based on longitudinal behavioral dataabstractIncreasingly, longitudinal behavioral data captured by various sensors are being analyzed to improve workplace performance. In this paper, we analyze the correlation between the regularity of workers' behavior and their levels of stress. We used a 23-month behavioral dataset for 18 workers that recorded their use of PCs and their locations in the office. We found that the principal eigenbehaviors extracted from the dataset with PCA represented typical work behaviors such as overwork using a PC and routine times for meetings. We found that more than 80% of each of the 18 workers' individual behaviors could be reconstructed using nine principal eigenbehaviors. In addition, the deviation ranges for the reconstruction accuracies were significantly different for workers in different positions. We conducted the correlation analysis between work behaviors of the workers and their stress level. Our results show a significant negative correlation (r > 0.69, p < 0.01) between the accuracy of reconstructed work behaviors and physical stress levels; and a significant positive correlation between the accuracy of reconstructed behavior and stress dissolution abilities. Our results suggest that the correlation between the stress level of workers and the regularity of their work behavior exists. This correlation will be useful for occupational healthcare. Shogo Okada, Yusaku Sato, Yuki Kamiya, Keiji Yamada, Katsumi Nitta |
ICMI | 1 |
| 2012 | Argument Analysis with Factor Annotation ToolabstractThis paper introduces an argumentation support tool based on Toulmin Diagram. It consists of a factor-tagging editor, a semantics calculation module based on Argumentation Framework and an automated factor extractor. When we input an argumentation record, this system generates a tagged argumentation record and a set of arguments that belong to credulous or skeptical extensions. The automated factor extractor supports to extract factors from an argumentation record by using a machine learning method. Shumpei Kubosawa, Youwei Lu, Shogo Okada, Katsumi Nitta |
JURIX | 3 |
| 2012 | Formation conditions of mutual adaptation in human-agent collaborative interaction
Yong Xu 0012, Yoshimasa Ohmoto, Shogo Okada, Kazuhiro Ueda, Takanori Komatsu, Takeshi Okadome, Koji Kamei, Yasuyuki Sumi, Toyoaki Nishida |
Appl. Intell. | 3 |
| 2011 | Online incremental clustering with distance metric learning for high dimensional dataabstractIn this paper, we present a novel incremental clustering algorithm which assigns of a set of observations into clusters and learns the distance metric iteratively in an incremental manner. The proposed algorithm SOINN-AML is composed based on the Self-organizing Incremental Neural Network (Shen et al 2006), which represents the distribution of unlabeled data and reports a reasonable number of clusters. SOINN adopts a competitive Hebbian rule for each input signal, and distance between nodes is measured using the Euclidean distance. Such algorithms rely on the distance metric for the input data patterns. Distance Metric Learning (DML) learns a distance metric for the high dimensional input space of data that preserves the distance relation among the training data. DML is not performed for input space of data in SOINN based approaches. SOINN-AML learns input space of data by using the Adaptive Distance Metric Learning (AML) algorithm which is one of the DML algorithms. It improves the incremental clustering performance of the SOINN algorithm by optimizing the distance metric in the case that input data space is high dimensional. In experimental results, we evaluate the performance by using two artificial datasets, seven real datasets from the UCI dataset and three real image datasets. We have found that the proposed algorithm outperforms conventional algorithms including SOINN (Shen et al 2006) and Enhanced SOINN (Shen et al 2007). The improvement of clustering accuracy (NMI) is between 0.03 and 0.13 compared to state of the art SOINN based approaches. Shogo Okada, Toyoaki Nishida |
IJCNN | 1 |
| 2011 | Active adaptation in human-agent collaborative interaction
Yong Xu 0012, Yoshimasa Ohmoto, Kazuhiro Ueda, Takanori Komatsu, Takeshi Okadome, Koji Kamei, Shogo Okada, Yasuyuki Sumi, Toyoaki Nishida |
J. Intell. Inf. Syst. | 7 |
| 2010 | Machine Learning Approaches for Time-Series Data Based on Self-Organizing Incremental Neural Network
Shogo Okada, Osamu Hasegawa, Toyoaki Nishida |
ICANN (3) | 1 |
| 2010 | Multi Class Semi-Supervised Classification with Graph Construction Based on Adaptive Metric Learning
Shogo Okada, Toyoaki Nishida |
ICANN (2) | 1 |
| 2010 | On-Line Unsupervised Segmentation for Multidimensional Time-Series Data and Application to Spatiotemporal Gesture data
Shogo Okada, Satoshi Ishibashi, Toyoaki Nishida |
IEA/AIE (1) | 1 |
| 2009 | A Platform System for Developing a Collaborative Mutually Adaptive Agent
Yong Xu 0012, Yoshimasa Ohmoto, Kazuhiro Ueda, Takanori Komatsu, Takeshi Okadome, Koji Kamei, Shogo Okada, Yasuyuki Sumi, Toyoaki Nishida |
IEA/AIE | 7 |
| 2009 | Incremental clustering of gesture patterns based on a self organizing incremental neural networkabstractThis paper describes an incremental unsupervised clustering mechanism for sequence patterns arising from human gestures. Although self-organizing incremental neural network (SOINN) is known as a powerful tool for incremental unsupervised clustering, it is only applicable to static and fixed-length patterns. In this paper, we propose an extension to SOINN to handle dynamic sequence patterns of variable length. We use a Hidden Markov Model (HMM), as a pre-processor for SOINN, to map the variable-length patterns into fixed-length patterns. HMM contributes to robust feature extraction from sequence patterns, enabling similar statistical features to be extracted from sequence patterns of the same category. As a result of experiments with incremental clustering gesture data, we have found that HMM based SOINN (HB-SOINN) outperforms other methods. Shogo Okada, Toyoaki Nishida |
IJCNN | 1 |
| 2009 | Unsupervised simultaneous learning of gestures, actions and their associations for Human-Robot InteractionabstractHuman-robot interaction using free hand gestures is gaining more importance as more untrained humans are operating robots in home and office environments. The robot needs to solve three problems to be operated by free hand gestures: gesture (command) detection, action generation (related to the domain of the task) and association between gestures and actions. In this paper we propose a novel technique that allows the robot to solve these three problems together learning the action space, the command space, and their relations by just watching another robot operated by a human operator. The main technical contribution of this paper is the introduction of a novel algorithm that allows the robot to segment and discover patterns in its perceived signals without any prior knowledge of the number of different patterns, their occurrences or lengths. The second contribution is using a Ganger-causality based test to limit the search space for the delay between actions and commands utilizing their relations and taking into account the autonomy level of the robot. The paper also presents a feasibility study in which the learning robot was able to predict actor's behavior with 95.2% accuracy after monitoring a single interaction between a novice operator and a WOZ operated robot representing the actor. Yasser Mohammad, Toyoaki Nishida, Shogo Okada |
IROS | 3 |
| 2008 | Motion recognition based on Dynamic-Time Warping method with Self-Organizing Incremental Neural NetworkabstractThis paper presents an approach (SOINN-DTW)for recognition of motion (gesture) that is based on the Self-Organizing Incremental Neural Network (SOINN) and Dynamic Time Warping (DTW). Using SOlNN's function of eliminating noise in the input data and representing the distribution of input data, SOINN-PTW method approximates the output distribution of each state in a self-organizing manner corresponding to the input data. The proposed SOINN-DTW method enhances Stochastic Dynamic Time Warping Method (Nakagawa. 1986). Results of experiments show that SOINN-DTW outperforms HMM, CRF. and HCRF in motion data. Shogo Okada, Osamu Hasegawa |
ICPR | 1 |
| 2008 | On-line learning of sequence data based on Self-Organizing Incremental Neural NetworkabstractThis paper presents an on-line, continuously learning mechanism for sequence data. The proposed approach is based on SOINN-DTW method (Okada and Hasegawa, 2007), which is designed for learning of sequence data. It is based on self-organizing incremental neural network (SOINN) and dynamic time warping (DTW). Using SOINNpsilas function represents the topological structure of online input data, the output distribution of each states is represented and adapted in a self-organizing manner corresponding to online input data. Consequently, this method can train a network and estimate parameters of the output distribution using new (on-line) data continuously, based on scarce batch-training data. Through online learning, the recognition accuracy is improved continuously. To confirm the effectiveness of the on-line learning mechanism of SOINN-DTW, we present an extensive set of experiments that demonstrate how our method outperforms the online learning method of HMM in classifying phoneme data. Shogo Okada, Osamu Hasegawa |
IJCNN | 1 |
| 2007 | Classification of Temporal Data Based on Self-organizing Incremental Neural Network
Shogo Okada, Osamu Hasegawa |
ICANN (2) | 1 |
| 2006 | Unsupervised Learning, Recognition, and Generation of Time-series Patterns Based on Self-Organizing SegmentationabstractThis study is intended to realize a motion recognition and generation mechanism based on observation. This mechanism, which is based on imitative learning, enables unsupervised incremental learning, recognition, and generation of time-series patterns that are observed directly from motion images. The mechanism segments these patterns into primitives in a self-organized manner using mixture-of-experts (MoE) with a non-monotonous neural network (NMNN). These patterns are expressed as permutations of primitives that are output by the MoE. Applying enhanced dynamic time warping (DTW) method recognizes these permutations of primitives. In addition, we introduce a semi-supervised learning method by applying this mechanism. We confirmed the effectiveness of this mechanism through two experiments using gestures Shogo Okada, Osamu Hasegawa |
RO-MAN | 1 |