Xincheng Ju

dblp:276/3173 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0002-6183-8096ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Information extraction and text analysis · 84% Language models and text generation · 12% Graph learning · 2%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › emotion recognition
emotion recognition in conversation
1.222024
ECFCON: Emotion Consequence Forecasting in Conversations · ACM Multimedia 2024
Transformer-based Label Set Generation for Multi-modal Multi-label Emotion Detection · ACM Multimedia 2020
Natural language and speech › Information extraction and text analysis
emotion cause analysis
0.912025
Emotion across Modalities and Cultures: Multilingual Multimodal Emotion-Cause Analysis with Memory-inspired Framework · ACM Multimedia 2025
Natural language and speech › Information extraction and text analysis › emotion cause analysis
emotion-cause pair extraction
0.912025
Enhanced Generative Framework With LLMs for Multimodal Emotion-Cause Pair Extraction in Conversations · IEEE Trans. Multim. 2025
Natural language and speech › Information extraction and text analysis › emotion recognition
multimodal emotion recognition
0.912025
Emotion across Modalities and Cultures: Multilingual Multimodal Emotion-Cause Analysis with Memory-inspired Framework · ACM Multimedia 2025
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
0.512021
Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation Detection · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
multimodal aspect-based sentiment analysis
0.512021
Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation Detection · EMNLP (1) 2021
Multimedia analysis and retrieval
affective computing
0.512021
Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing · AAAI 2021
Multimedia analysis and retrieval
multi-label classification
0.512021
Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing · AAAI 2021
Multimedia analysis and retrieval
multimodal emotion recognition
0.512021
Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing · AAAI 2021
Natural language and speech › Information extraction and text analysis
emotion recognition
0.412020
Multi-modal Multi-label Emotion Detection with Modality and Label Dependence · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › emotion recognition
multi-label emotion recognition
0.412020
Transformer-based Label Set Generation for Multi-modal Multi-label Emotion Detection · ACM Multimedia 2020
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.412020
Multi-modal Multi-label Emotion Detection with Modality and Label Dependence · EMNLP (1) 2020
Machine learning › Graph learning › graph neural network
message passing
0.112021
Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing · AAAI 2021

Methods — techniques the papers use, named apart from their topics

large language model · 1.7heterogeneous hierarchical message passing · 1.0memory bank · 0.9generative framework · 0.9few-shot learning · 0.9autoregressive aggregation · 0.9large language model prompting · 0.8clue-driven hybrid methods · 0.8hierarchical framework · 0.5cross-modal relation detection · 0.5
YearPublicationVenuePosition
2025 Emotion across Modalities and Cultures: Multilingual Multimodal Emotion-Cause Analysis with Memory-inspired Framework
abstract
Previous multimodal emotion-cause analysis in conversations (MEC-AC) has predominantly focused on English, overlooking the applicability of existing methods in multilingual contexts. To bridge this gap, we construct a Chinese contextual dataset (MEC4) to investigate how language and culture diversity influences existing MECAC approaches. Moreover, prior studies often rely on average pooling or frame sampling to extract visual and acoustic features from video and audio of long dialogues, which inevitably results in the loss of temporal dynamics and emotionally salient cues. To overcome these limitations, we propose a memory-inspired multilingual multimodal framework (M3F) based on large language model (LLM), which can effectively capture the temporal and global informative features of non-linguistic modalities through memory bank module. This module simulates the way memory is stored in human cognitive processes and incrementally aggregates past visual and acoustic features in an autoregressive manner, enabling effective reference during future sequence modeling. Through rigorous experiments and insightful analyses, we find that cultural differences cause variations in how emotional expressions in English and Chinese languages rely on modalities.
Xincheng Ju, Dong Zhang 0013, Shoushan Li, Erik Cambria, Guodong Zhou 0001
ACM Multimedia2
2025 Enhanced Generative Framework With LLMs for Multimodal Emotion-Cause Pair Extraction in Conversations
abstract
Emotion-Cause Pair Extraction (ECPE) in conversations aims to identify the emotional utterances (even their categories) along with their corresponding causal utterances, which is crucial in understanding the cause-effect relationship in dialogues. While prior studies of ECPE have predominantly focused on purely textual dialogues and neglected the exploration on the natural scenario of the dialogues with multimodal features, i.e., Multimodal Emotion-Cause Pair Extraction (MECPE) in conversations. To attempt this scenario, we propose a Generative approach for Multimodal Emotion-Cause pair extraction (GMEC) with a single stage, thus effectively reducing errors associated with the propagation and accumulation for MECPE. This approach can not only uniformly handle the information of diverse modalities, but also address all emotion and cause analysis tasks uniformly. Additionally, instead of utilizing the fixed commonsense knowledge base as previously, we resort to the Large Language Models (LLMs), which possess a powerful ability to emerge new knowledge, thereby acting as implicit knowledge engines for MECPE. We refer to this approach as enhanced GMEC. Extensive experimental results and detailed analysis demonstrate a notable improvement in the generative approach. Moreover, the integration of external knowledge from LLMs optimizes the efficiency of data utilization, particularly in few-shot scenarios. The integration of the generative model with LLMs has resulted in a cumulative enhancement of 4.94%, 10.90% on MECPE and MECPE-C (with emotion Category).
Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001
IEEE Trans. Multim.1
2024 ECFCON: Emotion Consequence Forecasting in Conversations
abstract
Conversation is a common form of human communication that includes extensive emotional interaction. Traditional approaches focused on studying emotions and their underlying causes in conversations. They try to address two issues: what emotions are present in the dialogue and what causes these emotions. However, these works often overlook the bidirectional nature of emotional interaction in dialogue: utterances can evoke emotions (cause), and emotions can also lead to certain utterances (consequence). Therefore, we propose a new issue: what consequences arise from these emotions? This leads to the introduction of a new task called Emotion Consequence Forecasting in CONversations (ECFCON). In this work, we first propose a corresponding dialogue-level dataset. Specifically, we select 2,780 video dialogues for annotation, totaling 39,950 utterances. Out of these, 12,391 utterances contain emotions, and 8,810 of these have discernible consequences. Then, we benchmark this task by conducting experiments from the perspectives of traditional methods, generalized LLMs prompting methods, and clue-driven hybrid methods. Both our dataset and benchmark codes are openly accessible to the public.
Xincheng Ju, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001
ACM Multimedia1
2023 Real-time Emotion Pre-Recognition in Conversations with Contrastive Multi-modal Dialogue Pre-training
abstract
This paper presents our pioneering effort in addressing a new and realistic scenario in multi-modal dialogue systems called Multi-modal Real-time Emotion Pre-recognition in Conversations (MREPC). The objective is to predict the emotion of a forthcoming target utterance that is highly likely to occur. We believe that this task can enhance the dialogue system's understanding of the interlocutor's state of mind, enabling it to prepare an appropriate response in advance. However, addressing MREPC poses the following challenges:1) Previous studies on emotion elicitation typically focus on textual modality and perform sentiment forecasting within a fixed contextual scenario. 2) Previous studies on multi-modal emotion recognition aim to predict the emotion of existing utterances, making it difficult to extend these approaches to MREPC due to the absence of the target utterance. To tackle these challenges, we construct two benchmark multi-modal datasets for MREPC and propose a task-specific multi-modal contrastive pre-training approach. This approach leverages large-scale unlabeled multi-modal dialogues to facilitate emotion pre-recognition for potential utterances of specific target speakers. Through detailed experiments and extensive analysis, we demonstrate that our proposed multi-modal contrastive pre-training architecture effectively enhances the performance of multi-modal real-time emotion pre-recognition in conversations.
Xincheng Ju, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001
CIKM1
2021 Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing
abstract
As an important research issue in affective computing community, multi-modal emotion recognition has become a hot topic in the last few years. However, almost all existing studies perform multiple binary classification for each emotion with focus on complete time series data. In this paper, we focus on multi-modal emotion recognition in a multi-label scenario. In this scenario, we consider not only the label-to-label dependency, but also the feature-to-label and modality-to-label dependencies. Particularly, we propose a heterogeneous hierarchical message passing network to effectively model above dependencies. Furthermore, we propose a new multi-modal multi-label emotion dataset based on partial time-series content to show predominant generalization of our model. Detailed evaluation demonstrates the effectiveness of our approach.
Dong Zhang 0013, Xincheng Ju, Junhui Li 0001, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001
AAAI2
2021 Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation Detection
abstract
Aspect terms extraction (ATE) and aspect sentiment classification (ASC) are two fundamental and fine-grained sub-tasks in aspect-level sentiment analysis (ALSA).In the textual analysis, jointly extracting both aspect terms and sentiment polarities has been drawn much attention due to the better applications than individual sub-task.However, in the multimodal scenario, the existing studies are limited to handle each sub-task independently, which fails to model the innate connection between the above two objectives and ignores the better applications.Therefore, in this paper, we are the first to jointly perform multi-modal ATE (MATE) and multi-modal ASC (MASC), and we propose a multi-modal joint learning approach with auxiliary cross-modal relation detection for multi-modal aspect-level sentiment analysis (MALSA).Specifically, we first build an auxiliary text-image relation detection module to control the proper exploitation of visual information.Second, we adopt the hierarchical framework to bridge the multi-modal connection between MATE and MASC, as well as separately visual guiding for each sub module.Finally, we can obtain all aspect-level sentiment polarities dependent on the jointly extracted specific aspects.Extensive experiments show the effectiveness of our approach against the joint textual approaches, pipeline and collapsed multi-modal approaches.
Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Shoushan Li, Min Zhang 0005, Guodong Zhou 0001
EMNLP (1)1
2020 Multi-modal Multi-label Emotion Detection with Modality and Label Dependence
abstract
As an important research issue in the natural language processing community, multi-label emotion detection has been drawing more and more attention in the last few years. However, almost all existing studies focus on one modality (e.g., textual modality). In this paper, we focus on multi-label emotion detection in a multi-modal scenario. In this scenario, we need to consider both the dependence among different labels (label dependence) and the dependence between each predicting label and different modalities (modality dependence). Particularly, we propose a multi-modal sequence-to-set approach to effectively model both kinds of dependence in multi-modal multi-label emotion detection. The detailed evaluation demonstrates the effectiveness of our approach.
Dong Zhang 0013, Xincheng Ju, Junhui Li 0001, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001
EMNLP (1)2
2020 Transformer-based Label Set Generation for Multi-modal Multi-label Emotion Detection
abstract
Multi-modal utterance-level emotion detection has been a hot research topic in both multi-modal analysis and natural language processing communities. Different from traditional single-label multi-modal sentiment analysis, typical multi-modal emotion detection is naturally a multi-label problem where an utterance often contains multiple emotions. Existing studies normally focus on multi-modal fusion only and transform multi-label emotion classification into multiple binary classification problem independently. As a result, existing studies largely ignore two kinds of important dependency information: (1) Modality-to-label dependency, where different emotions can be inferred from different modalities, that is, different modalities contribute differently to each potential emotion. (2) Label-to-label dependency, where some emotions are more likely to coexist than those conflicting emotions. To simultaneously model above two kinds of dependency, we propose a unified approach, namely multi-modal emotion set generation network (MESGN) to generate an emotion set for an utterance. Specifically, we first employ a cross-modal transformer encoder to capture cross-modal interactions among different modalities, and a standard transformer encoder to capture temporal information for each modality-specific sequence given previous interactions. Then, we design a transformer-based discriminative decoding module equipped with modality-to-label attention to handle the modality-to-label dependency. In the meanwhile, we employ a reinforced decoding algorithm with self-critic learning to handle the label-to-label dependency. Finally, we validate the proposed MESGN architecture on a word-level aligned and unaligned multi-modal dataset. Detailed experimentation shows that our proposed MESGN architecture can effectively improve the performance of multi-modal multi-label emotion detection.
Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Guodong Zhou 0001
ACM Multimedia1