Geng Tu

dblp:241/1300 · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Causal-ERC: A Multimodal Framework with Causal Prompting for Emotion Recognition in Conversations with Large Language Models
abstract
The rapid advancement of large language models (LLMs) has revitalised research in Emotion Recognition in Conversation (ERC). However, existing LLM-based ERC approaches operate solely on textual input, whereas MLLM-based emotion recognition methods in non-conversational scenarios typically perform only basic multimodal fusion and fail to consider speaker-sensitive contextual dependencies, which limits their performance on ERC tasks. To integrate multimodal cues effectively and address their limitations in handling contextual dependencies, we propose a novel LLM-based framework, Causal-ERC, which captures context representations within each modality and incorporates them into the LLM. Moreover, experimental results show that LLMs perform poorly on long conversations. To improve LLMs' ability to model long conversations, we adjust corresponding causal prompts according to the causal type of each utterance. Experiments on two benchmark MERC datasets demonstrate that our Causal-ERC framework consistently outperforms existing state-of-the-art approaches and improves LLM's performance in long-context scenarios.
Ran Jing, Geng Tu, Yice Zhang, Ruifeng Xu 0001
AAAI2
2026 Consensus-Driven Multi-Agent Cognitive Reasoning for Enhancing the Emotional Intelligence of Large Language Models
abstract
Large Language Models (LLMs) have demonstrated strong performance in various NLP tasks but remain limited in emotional intelligence (EI). Benchmarks such as EmoBench attribute this gap to deficiencies in cognitively demanding tasks that require inferring others’ latent mental states, intentions, and emotions in nuanced social contexts. To address this, we propose MACRo, a Multi-Agent Cognitive Reasoning framework that generates a structured Cognitive Chain of Thought comprising Situation, Clue, Thought, Action, and Emotion. Each component is generated by a specialized agent, enabling modular, interpretable multi-step reasoning. To ensure coherence and mitigate hallucinations, a coordinator agent verifies outputs, and a consensus game mechanism enforces alignment across reasoning steps. Extensive Experiments on EmoBench show that MACRo significantly enhances both emotional understanding and application across LLMs. Further evaluations confirm its generalizability to real-world social applications such as emotional support conversations.
Geng Tu, Dingming Li, Ruifeng Xu 0001
AAAI1
2026 A Graph-Enhanced MLLM for Hierarchical Multimodal Emotion Understanding and Support in Conversations
Geng Tu, Taiyu Niu, Ruifeng Xu 0001, Min Zhang 0005
SIGIR1
2026 SocialDropout: Dynamic Agent Dropout for Social Simulation
abstract
Large language model driven multi-agent social simulation frameworks enable realistic modeling of complex societal dynamics but incur substantial computational overhead due to dense agent participation and extensive interaction costs. To address this limitation, we propose SocialDropout, a reinforcement learning–based agent selection strategy within the AgentSociety framework, inspired by the AgentDropout paradigm, which dynamically identifies and samples informative agent subsets for each simulation round. Each agent is assigned an adaptive importance weight optimized to jointly minimize agent sparsity and computational cost measured by LLM calls, token consumption, and execution time—while preserving social interaction intensity within the environment. Extensive performance evaluation demonstrates that the proposed method significantly improves simulation efficiency and scalability. Moreover, ablation studies verify that high-level behavioral realism and outcome consistency are largely maintained despite substantial agent reduction. Our approach offers a practical and general optimization mechanism for large-scale LLM-based multi-agent social simulations under constrained computational budgets.
Huajie Wang, Geng Tu, Ruifeng Xu 0001, Min Zhang 0005
SIGIR2
2026 Is multimodal conversational emotion recognition satisfactory? Exploring the gaps in performance, generalization, and confidence
Geng Tu, Ran Jing, Erik Cambria, Wenjie Li 0002, Ruifeng Xu 0001
Pattern Recognit.1
2025 BeyondGender: A Multifaceted Bilingual Dataset for Practical Sexism Detection
abstract
Sexism affects both women and men, yet research often overlooks misandry and suffers from overly broad annotations that limit AI applications. To address this, we introduce BeyondGender, a dataset meticulously annotated according to the latest definitions of misogyny and misandry. It features innovative multifaceted labels encompassing aspects of sexism, gender, phrasing, misogyny, and misandry. The dataset includes 6K English and 1.7K Chinese sexism instances, alongside 13K non-sexism examples. Our evaluations of masked language models and large language models reveal that they detect misogyny in English and misandry in Chinese more effectively, with F1-scores of 0.87 and 0.62, respectively. However, they frequently misclassify hostile and mild comments, underscoring the complexity of sexism detection. Parallel corpus experiments suggest promising data augmentation strategies to enhance AI systems for nuanced sexism detection, and our dataset can be leveraged to improve value alignment in large language models.
Han Zhang 0025, Geng Tu, Qianlong Wang 0001, Keyang Ding, Chuang Fan, Jing Li 0049, Ruifeng Xu 0001
AAAI4
2025 CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
abstract
Jingqian Zhao, Bingbing Wang, Geng Tu, Yice Zhang, Qianlong Wang, Bin Liang, Jing Li, Ruifeng Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jingqian Zhao, Geng Tu, Yice Zhang, Qianlong Wang 0001, Bin Liang 0004, Jing Li 0049, Ruifeng Xu 0001
ACL (1)3
2025 Enhancing Emotion Reasoning for Image Multi-Emotion Prediction
abstract
Image multi-emotion prediction aims to identify the emotions evoked by images in humans. In the real world, individual cognitive differences can lead to different viewers experiencing varied emotions. Most existing researchers primarily focus on analyzing image features, which are limited to the perceptual level, leading to a superficial understanding of emotions. To address this gap, we propose an Emotional Reasoning Chain (EReC) based on a multimodal large language model, which learns both perception and reasoning abilities for multi-emotion prediction. Specifically, we design a parameter-efficient fine-tuning paradigm encompassing three task instructions: perception, reasoning, and prediction instruction. This paradigm involves a progressive process to perform targeted instructions within a domain for fine-tuning, thereby optimizing the model’s capabilities in image perception, cognitive reasoning, and multi-emotion prediction. Furthermore, to alleviate hallucinations of large models, a Reason-level Alignment Score (RAS) is introduced to guide the model toward closer alignment with human cognition. Experiments on four datasets demonstrate that our EReC method, which emulates the human cognitive process of viewing and interpreting images, employs progressive tuning to integrate the perceptual, reasoning, and predictive capabilities of large models, achieving superior performance.
Geng Tu, Bin Liang 0004, Zhixin Bai, Min Yang 0007, Ruifeng Xu 0001
ICASSP2
2025 A Multi-stage and Multi-target Knowledge Distillation Framework for Multimodal Conversational Emotion Recognition
abstract
In Emotion Recognition in Conversations (ERC), one-hot labels are typically used as ground truth, but they may not fully capture all emotions conveyed in an utterance. Recent work in textual ERC has investigated self-distillation techniques for generating soft labels via single-instance generation, aiming to improve emotional understanding. However, these approaches still struggle to fully capture complex emotional expressions. In multimodal ERC (MERC), generating soft labels is even more challenging due to integrating multiple modalities, which may express distinct emotions. Based on this, we propose a Multi-stage Multi-target Knowledge Distillation Framework, consisting of two components: Multi-stage Distillation (MSD) and Multi-target Distillation (MTD). MSD focuses on multi-stage self-distillation of soft labels and utterance representations, encouraging the MERC model to refine its label predictions across stages. Building on MSD, MTD further distills soft labels and features from a feature extractor used to extract modality-common and modality-specific features, deepening the model’s understanding of emotions in multimodal scenarios. Experimental results on two datasets show that our framework significantly improves the performance of various MERC models, surpassing state-of-the-art methods.
Taiyu Niu, Geng Tu, Hui Wang 0030, Bing Qin 0001, Ruifeng Xu 0001
ICME2
2025 Multimodal Emotion Recognition in Conversations via Graph Structure Learning
abstract
Multimodal Emotion Recognition in Conversations (MERC) aims to detect emotions expressed in each utterance within conversational videos. Graph-based methods are widely employed in MERC due to their superiority in modeling intricate speaker-sensitive and context-sensitive dependencies in conversations. Despite promising advancements made, existing graph-based methods primarily suffer from two inherent issues due to their reliance on manually predefined graph structures: structural redundancy, which burdens models with irrelevant noise aggregation, and insufficient connections, which results in a lack of cross-modal contextual cues. To address the above issues, we propose a novel graph structure learning framework for MERC, which comprises two key components: Context-aware Graph Sparsification (CGS) and Implicit Graph Relation Mining (IGR). CGS employs an edge selection network to refine the manually predefined graph, filtering out noisy information caused by structural redundancy. IGR explores potential connections that are beneficial for emotional reasoning. Experimental results on two datasets show that our proposed framework significantly improves the performance of graph-based methods in MERC.
Geng Tu, Yice Zhang, Jun Wang 0012, Bin Liang 0004, Yue Yu 0001, Min Yang 0007, Ruifeng Xu 0001
ICME2
2025 Overview of the NLPCC 2025 Shared Task 8: Personalized Emotional Support Conversation
Zhengda Jin, Geng Tu, Ruifeng Xu 0001
NLPCC (4)3
2025 Meta-Learning for Incomplete Multimodal Sentiment Analysis
abstract
Modality incompleteness is a critical yet underexplored challenge in multimodal sentiment analysis (MSA). Existing efforts, trained and evaluated under fixed missing rates, struggle to adapt to real-world scenarios with varying missing rates. To address this, we propose the Missing Modality Adaptation Framework (M2AF), leveraging model-agnostic meta-learning to enhance robustness against different levels of modality incompleteness. M2AF operates in two stages: meta-training and meta-testing. In the meta-training stage, a pre-trained MSA model, initially optimized for fixed missing rates, is further adapted to different levels of missing rates-low, moderate, and high. In the meta-testing stage, the model rapidly updates its parameters using minimal training data to handle target missing scenarios. Experiments on two popular datasets demonstrate that M2AF significantly improves the performance and generalization of various MSA models, ensuring more robust sentiment analysis in real-world settings with different modality incompleteness.
Geng Tu, Tianhao Wu 0009, Wenjie Li 0002, Ruifeng Xu 0001
SIGIR1
2025 Knowing What and Why: Causal emotion entailment for emotion recognition in conversations
Hao Liu 0080, Runguo Wei, Geng Tu, Jiali Lin, Dazhi Jiang, Erik Cambria
Expert Syst. Appl.3
2025 Generalizing to Unseen Speakers: Multimodal Emotion Recognition in Conversations With Speaker Generalization
abstract
Multimodal Emotion Recognition in Conversations (MERC) aims to identify the emotion expressed in each utterance within conversational videos. Current efforts are directed toward modeling speaker-sensitive context dependencies and multimodal fusion. However, they still struggle to handle utterances from unseen speakers, hampering the model's generalizability. To tackle this challenge, we propose a Speaker Generalization Framework for MERC. Specifically, we build a prototype graph to learn Speaker-based Utterance Representations (SUR), leveraging prototypes as the bridge between seen and unseen speakers. Speaker-aware Contrastive Learning (CL) is then applied to refine SUR, pulling utterances (or prototypes) from the same speaker together while pushing those from different speakers apart. Further, we introduce a prototypical graph CL to generalize SUR to unseen speakers, ensuring that the same speakers exhibit similar graph structures, while dissimilar ones differ. To further enhance model generalization, we introduce Uncertainty-based Generalization for Speakers, randomly sampling SUR statistics from the estimated Gaussian distribution and probabilistically replacing the original SUR. Experimental findings on two datasets highlight that our framework substantially improves the generalization of various MERC models, surpassing state-of-the-art methods.
Geng Tu, Ran Jing, Bin Liang 0004, Yue Yu 0001, Min Yang 0007, Bing Qin 0001, Ruifeng Xu 0001
IEEE Trans. Affect. Comput.1
2024 Adaptive Graph Learning for Multimodal Conversational Emotion Detection
abstract
Multimodal Emotion Recognition in Conversations (ERC) aims to identify the emotions conveyed by each utterance in a conversational video. Current efforts encounter challenges in balancing intra- and inter-speaker context dependencies when tackling intra-modal interactions. This balance is vital as it encompasses modeling self-dependency (emotional inertia) where speakers' own emotions affect them and modeling interpersonal dependencies (empathy) where counterparts' emotions influence a speaker. Furthermore, challenges arise in addressing cross-modal interactions that involve content with conflicting emotions across different modalities. To address this issue, we introduce an adaptive interactive graph network (IGN) called AdaIGN that employs the Gumbel Softmax trick to adaptively select nodes and edges, enhancing intra- and cross-modal interactions. Unlike undirected graphs, we use a directed IGN to prevent future utterances from impacting the current one. Next, we propose Node- and Edge-level Selection Policies (NESP) to guide node and edge selection, along with a Graph-Level Selection Policy (GSP) to integrate the utterance representation from original IGN and NESP-enhanced IGN. Moreover, we design a task-specific loss function that prioritizes text modality and intra-speaker context selection. To reduce computational complexity, we use pre-defined pseudo labels through self-supervised methods to mask unnecessary utterance nodes for selection. Experimental results show that AdaIGN outperforms state-of-the-art methods on two popular datasets. Our code will be available at https://github.com/TuGengs/AdaIGN.
Geng Tu, Bin Liang 0004, Hongpeng Wang 0002, Ruifeng Xu 0001
AAAI1
2024 SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-Modal Intent Detection
abstract
Multi-modal intent detection aims to utilize various modalities to understand the user’s intentions, which is essential for the deployment of dialogue systems in real-world scenarios. The two core challenges for multi-modal intent detection are (1) how to effectively align and fuse different features of modalities and (2) the limited labeled multi-modal intent training data. In this work, we introduce a shallow-to-deep interaction framework with data augmentation (SDIF-DA) to address the above challenges. Firstly, SDIF-DA leverages a shallow-to-deep interaction module to progressively and effectively align and fuse features across text, video, and audio modalities. Secondly, we propose a ChatGPT-based data augmentation approach to automatically augment sufficient training data. Experimental results demonstrate that SDIF-DA can effectively align and fuse multi-modal features by achieving state-of-the-art performance. In addition, extensive analyses show that the introduced data augmentation approach can successfully distill knowledge from the large language model.
Shijue Huang, Libo Qin 0001, Geng Tu, Ruifeng Xu 0001
ICASSP4
2024 Multimodal Emotion Recognition Calibration in Conversations
abstract
Multimodal Emotion Recognition in Conversations (MERC) aims to identify the emotions conveyed by each utterance in a conversational video. Current efforts focus on modeling speaker-sensitive context dependencies and multimodal fusion. Despite the progress, the reliability of MERC methods remains largely unexplored. Extensive empirical studies reveal that current methods suffer from unreliable predictive confidence. Specifically, in some cases, the confidence estimated by these models increases when a modality or specific contextual cues are corrupted, defining these as uncertain samples. This contradicts the foundational principle in informatics, namely, the elimination of uncertainty. Based on this, we propose a novel calibration framework CMERC to calibrate MERC models without altering the model structure. It integrates curriculum learning to guide the model in progressively learning more uncertain samples; hybrid supervised contrastive learning to refine utterance representations, by pulling uncertain samples and others apart; and confidence constraint to penalize the model on uncertain samples. Experimental results on two datasets demonstrate the effectiveness and generalization capabilities of our CMERC across various MERC models, surpassing state-of-the-art methods.
Geng Tu, Bin Liang 0004, Hui Wang 0030, Ruifeng Xu 0001
ACM Multimedia1
2024 A Persona-Infused Cross-Task Graph Network for Multimodal Emotion Recognition with Emotion Shift Detection in Conversations
abstract
Recent research in Multimodal Emotion Recognition in Conversations (MERC) focuses on multimodal fusion and modeling speaker-sensitive context. In addition to contextual information, personality traits also affect emotional perception. However, current MERC methods solely consider the personality influence of speakers, neglecting speaker-addressee interaction patterns. Additionally, the bottleneck problem of Emotion Shift (ES), where consecutive utterances by the same speaker exhibit different emotions has been long neglected in MERC. Early ES research fails to distinguish diverse shift patterns and simply introduces whether shifts occur as knowledge into the MERC model without considering the complementary nature of the two tasks. Based on this, we propose a Persona-infused Cross-task Graph Network (PCGNet). It first models the speaker-addressee interactive relationships by the persona-infused refinement network. Then, it learns the auxiliary task of ES Detection and the main task of MERC using cross-task connections to capture correlations across two tasks. Finally, we introduce shift-aware contrastive learning to discern diverse shift patterns. Experimental results demonstrate that PCGNet outperforms state-of-the-art methods on two widely used datasets.
Geng Tu, Bin Liang 0004, Ruifeng Xu 0001
SIGIR1
2024 Self-supervised utterance order prediction for emotion recognition in conversations
Dazhi Jiang, Hao Liu 0080, Geng Tu, Runguo Wei, Erik Cambria
Neurocomputing3
2024 Multi-Modal Attentive Prompt Learning for Few-shot Emotion Recognition in Conversations
abstract
Emotion recognition in conversations (ERC) has emerged as an important research area in Natural Language Processing and Affective Computing, focusing on accurately identifying emotions within the conversational utterance. Conventional approaches typically rely on labeled training samples for fine-tuning pre-trained language models (PLMs) to enhance classification performance. However, the limited availability of labeled data in real-world scenarios poses a significant challenge, potentially resulting in diminished model performance. In response to this challenge, we present the Multi-modal Attentive Prompt (MAP) learning framework, tailored specifically for few-shot emotion recognition in conversations. The MAP framework consists of four integral modules: multi-modal feature extraction for the sequential embedding of text, visual, and acoustic inputs; a multi-modal prompt generation module that creates six manually-designed multi-modal prompts; an attention mechanism for prompt aggregation; and an emotion inference module for emotion prediction. To evaluate our proposed model’s efficacy, we conducted extensive experiments on two widely recognized benchmark datasets, MELD and IEMOCAP. Our results demonstrate that the MAP framework outperforms state-of-the-art ERC models, yielding notable improvements of 3.5% and 0.4% in micro F1 scores. These findings highlight the MAP learning framework’s ability to effectively address the challenge of limited labeled data in emotion recognition, offering a promising strategy for improving ERC model performance.
Xingwei Liang, Geng Tu, Jiachen Du, Ruifeng Xu 0001
J. Artif. Intell. Res.2
2024 What do they "meme"? A metaphor-aware multi-modal multi-task framework for fine-grained meme understanding
abstract
Fine-grained meme understanding aims to explore and comprehend the meanings of memes from multiple perspectives by performing various tasks, such as sentiment analysis, intention detection, and offensiveness detection. Existing approaches primarily focus on simple multi-modality fusion and individual task analysis. However, there remain several limitations that need to be addressed: (1) the neglect of incongruous features within and across modalities, and (2) the lack of consideration for correlations among different tasks. To this end, we leverage metaphorical information as text modality and propose a Metaphor-aware Multi-modal Multi-task Framework (M3F) for fine-grained meme understanding. Specifically, we create inter-modality attention enlightened by the Transformer to capture inter-modality interaction between text and image. Moreover, intra-modality attention is applied to model the contradiction between the text and metaphorical information. To learn the implicit interaction among different tasks, we introduce a multi-interactive decoder that exploits gating networks to establish the relationship between various subtasks. Experimental results on the MET-Meme dataset show that the proposed framework outperforms the state-of-the-art baselines in fine-grained meme understanding.
Shijue Huang, Bin Liang 0004, Geng Tu, Min Yang 0007, Ruifeng Xu 0001
Knowl. Based Syst.4
2023 Do Topic and Causal Consistency Affect Emotion Cognition? A Graph Interactive Network for Conversational Emotion Detection
abstract
Emotion recognition in conversations (ERC) typically requires modeling both intra- and inter-speaker context dependencies. However, when modeling inter-speaker dependencies, it may not capture differences among other participants in the conversation. Recent ERC research has attempted to improve utterance representations by utilizing speakers’ commonsense knowledge. Nonetheless, these studies ignore the causal consistency in knowledge between the two participants, which contradicts the above modeling of speaker-sensitive context dependencies. Additionally, it is observed that historical utterances from various topics are blindly leveraged in context modeling, which fails the inter- and intra-topic coherence. To address these issues, we propose the topic- and causal-aware interactive graph network (TCA-IGN). Specifically, we suggest a graph encoder to model topic-level context dependencies, achieving inter- and intra-topic coherence. The topics of utterances are derived from a context-sensitive neural topic model. Then, we present a causal-aware graph attention to keep the speaker’s causal consistency in commonsense knowledge, improving speaker-level context modeling. Finally, considering the defect of modeling inter-speaker or inter-topic context dependencies, we employ supervised contrastive learning to sweeten it. Experimental results show that TCA-IGN outperforms state-of-the-art methods on three public conversational datasets.
Geng Tu, Bin Liang 0004, Xiucheng Lyu, Lin Gui 0003, Ruifeng Xu 0001
ECAI1
2023 A Training-Free Debiasing Framework with Counterfactual Reasoning for Conversational Emotion Detection
abstract
Unintended dataset biases typically exist in existing Emotion Recognition in Conversations (ERC) datasets, including label bias, where models favor the majority class due to imbalanced training data, as well as the speaker and neutral word bias, where models make unfair predictions because of excessive correlations between specific neutral words or speakers and classes.However, previous studies in ERC generally focus on capturing context-sensitive and speaker-sensitive dependencies, ignoring the unintended dataset biases of data, which hampers the generalization and fairness in ERC.To address this issue, we propose a Training-Free Debiasing framework (TFD) that operates during prediction without additional training.To ensure compatibility with various ERC models, it does not balance data or modify the model structure.Instead, TFD extracts biases from the model by generating counterfactual utterances and contexts and mitigates them using simple yet empirically robust element-wise subtraction operations.Extensive experiments on three public datasets demonstrate that TFD effectively improves generalization ability and fairness across different ERC models 1 .
Geng Tu, Ran Jing, Bin Liang 0004, Min Yang 0007, Kam-Fai Wong, Ruifeng Xu 0001
EMNLP1
2023 Learning More from Mixed Emotions: A Label Refinement Method for Emotion Recognition in Conversations
abstract
Abstract One-hot labels are commonly employed as ground truth in Emotion Recognition in Conversations (ERC). However, this approach may not fully encompass all the emotions conveyed in a single utterance, leading to suboptimal performance. Regrettably, current ERC datasets lack comprehensive emotionally distributed labels. To address this issue, we propose the Emotion Label Refinement (EmoLR) method, which utilizes context- and speaker-sensitive information to infer mixed emotional labels. EmoLR comprises an Emotion Predictor (EP) module and a Label Refinement (LR) module. The EP module recognizes emotions and provides context/speaker states for the LR module. Subsequently, the LR module calculates the similarity between these states and ground-truth labels, generating a refined label distribution (RLD). The RLD captures a more comprehensive range of emotions than the original one-hot labels. These refined labels are then used for model training in place of the one-hot labels. Experimental results on three public conversational datasets demonstrate that our EmoLR achieves state-of-the-art performance.
Jintao Wen, Geng Tu, Dazhi Jiang, Wenhua Zhu
Trans. Assoc. Comput. Linguistics2
2023 AutoML-Emo: Automatic Knowledge Selection Using Congruent Effect for Emotion Identification in Conversations
abstract
Emotion recognition in conversations (ERC) has wide applications in medical care, human-computer interaction, and other fields. Unlike the general task of emotion analysis, humans usually rely on context and commonsense knowledge to convey emotions in conversations. Only when the model can connect and fully utilize a large-scale commonsense knowledge base, it can better understand latent contents in conversations. Unfortunately, there is no available knowledge selection mechanism to address such knowledge needs and to make sure the system is not flooded with irrelevant commonsense knowledge. Therefore, we propose an AutoML strategy based on emotion congruent effect to select suitable knowledge and models, called AutoML-Emo. Global exploration and local exploitation-based selection mechanisms (G&LESM) are used for automatic knowledge selection. The transformer-based architecture search (TAS) is applied to model selection, the selected transformer-based model is employed to incorporate knowledge and capture context information in conversations. The experimental results show that AutoML-Emo can effectively enhance external knowledge in different sizes and domain datasets. Moreover, the selected transformer-based model derived from TAS is superior to the most advanced models.
Dazhi Jiang, Runguo Wei, Jintao Wen, Geng Tu, Erik Cambria
IEEE Trans. Affect. Comput.4
2023 Sentiment- Emotion- and Context-Guided Knowledge Selection Framework for Emotion Recognition in Conversations
abstract
Emotion recognition in conversations (ERC) needs to detect the emotion of each utterance in conversations. However, it is difficult for machines to recognize the emotion of utterances like humans, partly because of the lack of commonsense knowledge. Despite existing efforts gradually incorporate knowledge in ERC, they can not adaptively adjust knowledge according to different utterances and their context. In this article, we propose a knowledge selection framework SKSEC (SelectKnowledge in light ofSentimentEmotion andContext). In the SKSEC framework, first, external knowledge is eliminated by three Knowledge Elimination (KE) modules. More concretely, In word-level KE, the concept knowledge different from the sentiment corresponding to the word in utterances is randomly eliminated. In utterance- or context-level KE, If the similarity between the knowledge representation and the emotion label representation of the current utterance or its context is less than the preset threshold, the knowledge will be eliminated. Then we refine the weight of knowledge using two Graph ATtention (GAT) mechanisms. Specifically, In Sentics GAT, we employ a dimensional emotion model to measure words in utterances and their corresponding knowledge and adjust the weight of knowledge according to their emotional similarity. In Semantics GAT, the weight of knowledge is adjusted according to the semantic similarity between context and incorporated knowledge. Finally, we feed the selected knowledge to the most advanced models to evaluate the quality of knowledge. The experimental results show that the SKSEC framework can effectively improve the performance of the model by eliminating and refining external knowledge in different size and domain datasets.
Geng Tu, Bin Liang 0004, Dazhi Jiang, Ruifeng Xu 0001
IEEE Trans. Affect. Comput.1
2022 Exploration meets exploitation: Multitask learning for emotion recognition based on discrete and dimensional models
Geng Tu, Jintao Wen, Hao Liu 0080, Sentao Chen, Lin Zheng 0003, Dazhi Jiang
Knowl. Based Syst.1
2021 Multimodality Sentiment Analysis in Social Internet of Things Based on Hierarchical Attentions and CSAT-TCN With MBM Network
abstract
Multimodality sentiment analysis in the social Internet of Things is a developing field, which is basic to empathetic mechanisms, affective computing, and artificial intelligence. Current works in this domain do not explicitly consider the influence of contextual information fusion based on correlation coefficient and memory network with branch structure for sentiment analysis. Unlike present works, this article presents a hierarchical self-attention fusion (H-SATF) model for capturing contextual information better among utterances, a contextual self-attention temporal convolutional network (CSAT-TCN) for sentiment recognition in the social Internet of Things, and a multibranch memory (MBM) network that stores self-speaker and interspeaker sentimental states into global memories. For MOSI data sets, the hybrid H-SATF-CSAT-TCN-MBM model outperforms the state-of-the-art networks and shows 0.31%-9.93% improvement.
Guorong Xiao, Geng Tu, Lin Zheng 0003, Teng Zhou, Xin Li 0102, Syed Hassan Ahmed, Dazhi Jiang
IEEE Internet Things J.2
2021 A hybrid intelligent model for acute hypotensive episode prediction with large-scale data
Dazhi Jiang, Geng Tu, Donghui Jin, Kaichao Wu, Cheng Liu 0001, Lin Zheng 0003, Teng Zhou
Inf. Sci.2
2021 A framework for designing of genetic operators automatically based on gene expression programming and differential evolution
Dazhi Jiang, Zhihang Tian, Zhihui He, Geng Tu, Ruixiang Huang
Nat. Comput.4
2019 Real-time hand gestures system based on leap motion
abstract
Summary In the three‐dimensional human‐computer interaction, the identification of dynamic and static gestures is a very important and challenging work in the field of machine vision, In this paper, we propose a new gesture recognition system. Leap Motion device is a kind of equipment, which is specially used for hand recognition, which can get the feature data to realize the gesture recognition in real time. The system is mainly composed of the following two parts. For static gestures, we use a kind of feature information based on the distance, direction, and bending degree of the fingertip, and bring the support vector machine into the training to realize the static gesture recognition. For dynamic gestures, we use gesture length as a benchmark to reject non‐key gestures and preprocess frames with abnormal gesture sequences. The average recognition rate of static gestures reaches 99.98%, and the recognition rate of dynamic gestures reaches 96.20%. The experimental results show that the algorithm has a good effect on gesture recognition, and it is suitable for the simple interaction between gestures, people and people and daily communication of daily communication barriers.
Geng Tu, Chuchu Zhao, Wenlong Yi
Concurr. Comput. Pract. Exp.2