Sixia Li

dblp:251/1125 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0007-3985-6892ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Classification of Co-speech Gesture Pragmatic Function in Storytelling
abstract
Gestures, as essential co-speech behaviors in human communication, carry rich pragmatic functions. Accurately recognizing these functions could enhance an agent’s ability to understand communicative behavior. In spontaneous storytelling scenarios—unlike lab-controlled settings—gesture functions exhibit high variability and are strongly influenced by individual speaker differences, making it difficult for unimodal systems to reliably capture their pragmatic intent. To address this challenge, we collected and annotated naturally occurring co-speech gestures in narrative dialogues, assigning each gesture one of six pragmatic function labels. We further propose a multimodal sequential classification model that encodes skeletal motion, acoustic prosody, and facial dynamics using separate bidirectional LSTM networks. These modality-specific encodings are fused via cross-modal attention and a gated mechanism to capture temporal dependencies and complementary information across modalities. Experimental results demonstrate that our tri-modal system achieves 62.5% accuracy and a weighted F1 score of 0.62 on the six-way classification task, outperforming uni-modal and bi-modal baselines by 3–11%. Ablation analysis reveals that skeletal features provide the most discriminative power for the majority of gesture functions, acoustic features are critical for specific categories, and facial features—though weak in isolation—substantially enhance overall performance when integrated via attention.
Jinqian Zhang, Sixia Li, Candy Olivia Mawalim, Shogo Okada
HAI2
2025 Autonomous dialogue generation based on phase boundary detection within continuous motion for domestic robot
abstract
Dialogue generation plays a key role in responding to user and providing transparency in motion execution in human-robot interaction. As motion planning is generally performed in terms of discrete motions, previous studies have focused on dialogue generation at the boundaries between motions. Recently, continuous motion generation was proposed to enable adapting actions to unique characteristics of the objects for domestic robots. Since a continuous motion generally involves physical and nonphysical phases, providing dialogues when the phase changes is crucial for decreasing users’ anxiety and guaranteeing safety. However, continuous motions lack clear phase boundaries, posing challenges for dialogue generation between phases. For this problem, we segmented continuous motions into discrete phases, and constructed a system to enable the robot to autonomously generate dialogues by detecting phase boundaries. To do so, we built phase estimation models using robot sensor data and designed system modules. Specifically, we collected data in the scenario of a robot assisting to lift the user up from bed. We segmented the continuous motion into three phases based on the user’s posture and whether the robot applied force to the human. The best phase estimation model achieved a macro F1 score of 0.894, demonstrating that phases can be estimated from sensor data. The evaluation results of our system demonstrated that the system accurately detects phase boundaries and generates appropriate dialogues corresponding to phases. Furthermore, we conducted simulations with a user agent to investigate system behaviors when the phase estimation was incorrect. The results suggested that explicitly stating the phase is important for avoiding misunderstandings and safety issues.
Sixia Li, Tamon Miyake, Tetsuya Ogata, Shigeki Sugano, Shogo Okada
RO-MAN1
2024 Incremental Multimodal Sentiment Analysis for HAIs Based on Multitask Active Learning with Interannotator Agreement
abstract
Multimodal sentiment analysis (MSA) is critical in developing empathetic and adaptive multimodal dialogue systems or conversational agents that can naturally interact with users by recognizing sentiment and engagement. Addressing the challenges of collecting labeled data for MSA in human-agent interaction (HAI), this study introduces an innovative approach that combines active learning and multitask learning. Our efficient sentiment recognition model leverages active learning to select informative data for learning models, significantly reducing the labor-intensive data labeling process. Furthermore, we employ multitask learning to improve annotation (label) quality by evaluating alignment with true labels and interannotator agreement, thus enhancing the reliability of sentiment annotations. We evaluate the proposed multitask and active learning methods via a human-agent multimodal dialogue dataset that includes various types of sentiment annotations, which are publicly available. The experimental results demonstrate that by learning to predict the agreement score, multitask learning becomes better than singletask learning at capturing the uncertainties in the data. This study lays the groundwork for incremental learning strategies in MSA, aiming to adaptively understand user sentiments in human-agent interactions.
Thus Karnjanapatchara, Sixia Li, Candy Olivia Mawalim, Kazunori Komatani, Shogo Okada
ACII2
2024 Do We Need To Watch It All? Efficient Job Interview Video Processing with Differentiable Masking
abstract
With technological advancements in transmitting and storing large video files, more and more organizations are incorporating asynchronous video interviews as part of their personnel selection process. Automatic evaluation of these videos is a challenging machine learning setting because the samples are composed of time series input data but only one overall label is available. It is unclear which segments of the time series input (i.e., videos) are the most important ones for prediction. Not all nonverbal features, spoken words, and utterances contribute equally to the prediction; some segments of the videos might even introduce noise to the model. Processing all multimodal information is therefore inefficient. To address this challenge, we propose a framework to model the content of the answer via the full transcription and the speaking patterns of the interviewee via short clips. Our model learns to automatically select the most informative segment by previewing the acoustic modality using a technique called differentiable masking. The results show that our method outperforms existing approaches while being more efficient since only partial multimodal data are processed, and the interpretability of the model is enhanced.
Sixia Li, Candy Olivia Mawalim, Hung-Hsuan Huang, Chee Wee Leong, Shogo Okada
ICMI2
2024 Automatic mild cognitive impairment estimation from the group conversation of coimagination method
abstract
The coimagination method (Otake, 2009) is designed to prevent dementia in individuals with mild cognitive impairment (MCI) by utilizing the brain’s natural processes. This method involves participants sharing their thoughts and feelings through group conversations centered around shared photos. The coimagination method contains two phases: (1) each participant talk about their memories and experiences related to the photos they bring, and (2) other participants ask questions about the photos. Automating the MCI estimation could be helpful for assisting individuals with MCI during coimagination. However, previous MCI estimation methods rarely focused on group conversation scenarios, despite the potential of multimodal features observed in these scenarios in revealing cognitive states. This study focuses MCI individuals defined by cognitive test scores (e.g., Mini-Mental State Examination (MMSE)). We explore MCI estimation from three aspects. First, we clarify whether MCI can be effectively estimated by constructing estimation models using linguistic and acoustic features from coimagination sessions. Second, we evaluate the impact of using data from the two distinct phases, as they may activate participants’ cognitive functions differently. Finally, we analyze the effects of incorporating subtasks including participants’ conversational customary and engagement level during coimagination via multitask learning. The experimental results demonstrated that individuals with MCI can be effectively estimated from group conversations from coimagination, with the highest macro F1 score of 0.693. The results also demonstrated that the performance improved when using data from the phase that highly activates cognitive functions and when considering conversation customary as a subtask.
Sixia Li, Kazumi Kumagai, Mihoko Otake, Shogo Okada
ICMI1
2022 Investigating the relationship between dialogue and exchange-level impression
abstract
Multimodal dialogue systems (MDS) have recently attracted increasing attention. The automatic evaluation of user impression with spoken dialog at the dialog level plays a central role in managing dialog systems. A user usually forms an overall impression through the experience of each exchange of turns in the conversation. Thus, the user’s exchange-level sentiment should be considered when recognizing the user’s overall impression of the dialog. Previous research has focused on modeling user impressions during individual exchanges or during the overall conversation. Thus, the relationship between user sentiment at the exchange level and user impression at the dialog level is still unclear, and appropriately utilizing this relationship in impression analysis remains unexplored. In this paper, we first investigate the relation between sentiment at the exchange level and 18 labels that indicate different aspects of the user impression at the dialog level. Then, we present a multitask learning model (MTL) that uses exchange-level annotations to recognize dialog-level labels. The experimental results demonstrate that our proposed model achieves better performance at the dialog level, outperforming the single-task model by a maximum of 15.7%.
Wenqing Wei, Sixia Li, Shogo Okada
ICMI2
2021 Multimodal User Satisfaction Recognition for Non-task Oriented Dialogue Systems
abstract
Multimodal dialogue systems (MDSs) are needed to allow users to converse with virtual agents that use natural language by sensing the multimodal behavior of users. One crucial step in the development of an MDS is measuring how well the dialogue system performs. Though previous research focused on the user satisfaction modeling from linguistic modality in text-to-text dialogue systems, the user satisfaction is observed by not only spoken dialogue contents but also the acoustic and visual nonverbal behaviors of users. Multimodal social signal sensing provides a solution that automatically measures dialogue systems based on subjective evaluation. With this background, we proposed a multimodal recognition model of the user using sequence modeling algorithms (RNN, LSTM, and GRU). It is a novel challenge to recognize the user satisfaction label at the dialogue level. Each label was annotated by the user based on the overall dialogue. We extracted both verbal features and nonverbal features at the exchange level (the unit is a pair of system and user utterances) and analyzed the contributions of multimodal features and unimodal features to recognize user satisfaction labels at the dialogue level. We used a multimodal user-system dialogue data corpus with user satisfaction labels at the dialogue level. To validate the recognition accuracy of the proposed multimodal modeling approach, we compared the proposed method with two models based on human perception by external human coders and the system operator (called “Wizard”) with whom the user talks. The experimental results showed that the multimodal model achieved a better performance in both classification and regression tasks. The results indicated that the performance of the multimodal model was higher than that of the human models.
Wenqing Wei, Sixia Li, Shogo Okada, Kazunori Komatani
ICMI2
2020 Dimensional Emotion Prediction Based on Interactive Context in Conversation
Sixia Li, Jianwu Dang 0001
INTERSPEECH2
2019 Interaction Process Label Recognition in Group Discussion
abstract
In qualifying and analyzing the performance of group interaction, interaction processing analysis (IPA) defined by Bale is considered a useful approach. IPA is a system for labeling a total of 12 interaction categories for the interaction process. Automatic IPA can manually encompass the gap in spending manpower and can efficiently qualify group performance. In this paper, we present computational interaction processing analysis by developing a model to recognize categories of IPA. We extract both verbal features and nonverbal features for IPA category recognition modeling with SVM, RF, DNN and LSTM machine learning algorithms and analyze the contribution of multimodal features and unimodal features for the total data and each label. We also investigate the effect of context information by training sequences with different lengths with an LSTM and evaluating them. The results show that multimodal features achieve the best performance with an F1 score of 0.601 for the recognition of 12 IPA categories using the total data. Multimodal features are better than the unimodal features for the total data and most labels. The results of investigating context information show that a suitable length of sequence enables a longer sequence to achieve the best F1 score of 0.602 and a better performance for recognition.
Sixia Li, Shogo Okada, Jianwu Dang 0001
ICMI1