VLDB 2026 Research / reviewers in the wild / expert
Dong Zhang 0013
dblp:68/3245-13
· DBLP profile ↗
39ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0002-8948-2856ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-free and zero-shot regeneration for hallucination mitigation in MLLMs: Representation understanding perspective
Dong Zhang 0013, Yuansheng Ma, Linqin Li, Shoushan Li, Erik Cambria, Guodong Zhou 0001 |
Expert Syst. Appl. | 1 |
| 2025 | Age Inference on both Textual and Social Perspectives with Semi-supervised Learning
Wu Dan, Zixuan Teng, Dong Zhang 0013 |
CogSci | 3 |
| 2025 | Learning Joint General and Specific Representation with Masked Auto-Encoder for Radiology Report Generation
Tongfei Shen, Zixuan Teng, Dong Zhang 0013 |
ICANN (4) | 3 |
| 2025 | Pathological Section Staining Transferring with Tailored Metric-based Model SelectionabstractAs the important pathological section staining, Immunohistochemistry (IHC) staining uses labeled antibodies to highlight specific antigens, providing clearer results for malignancy identification compared with Hematoxylin and Eosin (H&E) staining. However, obtaining IHC manually is labor-intensive and costly. In this study, we attempt to automatically generate IHC staining image by style transferring from H&E, which is easy to get. Due to the unalignment between single evaluation metric and real generation quality, we propose a tailored metric-based model selection (TMMS) method for pathological section staining transferring (PSST). Our method can fuse multiple traditional metrics and multiple relevant losses to measure the quality of generated IHC and select a best-performed model. Extensive experiments and analysis demonstrate the effectiveness of our method TMMS. Yiming Ji, Suyang Zhu, Dong Zhang 0013, Shoushan Li |
ICASSP | 3 |
| 2025 | Cross-Domain Constituency Parsing with Multi-LLM Debate
Qingying Sun, Haiyan Tian, Dong Zhang 0013 |
ICIC (23) | 3 |
| 2025 | Unified Option Generation for Zero- and Few-Shot Emotion and Cause Analysis in Dialogues
Qingying Sun, Dong Zhang 0013 |
ICIC (23) | 3 |
| 2025 | Emotion across Modalities and Cultures: Multilingual Multimodal Emotion-Cause Analysis with Memory-inspired FrameworkabstractPrevious multimodal emotion-cause analysis in conversations (MEC-AC) has predominantly focused on English, overlooking the applicability of existing methods in multilingual contexts. To bridge this gap, we construct a Chinese contextual dataset (MEC4) to investigate how language and culture diversity influences existing MECAC approaches. Moreover, prior studies often rely on average pooling or frame sampling to extract visual and acoustic features from video and audio of long dialogues, which inevitably results in the loss of temporal dynamics and emotionally salient cues. To overcome these limitations, we propose a memory-inspired multilingual multimodal framework (M3F) based on large language model (LLM), which can effectively capture the temporal and global informative features of non-linguistic modalities through memory bank module. This module simulates the way memory is stored in human cognitive processes and incrementally aggregates past visual and acoustic features in an autoregressive manner, enabling effective reference during future sequence modeling. Through rigorous experiments and insightful analyses, we find that cultural differences cause variations in how emotional expressions in English and Chinese languages rely on modalities. Xincheng Ju, Dong Zhang 0013, Shoushan Li, Erik Cambria, Guodong Zhou 0001 |
ACM Multimedia | 3 |
| 2025 | Ueco: Unified editing chain for efficient appearance transfer with multimodality-guided diffusion
Yuansheng Ma, Dong Zhang 0013, Shoushan Li, Erik Cambria, Guodong Zhou 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Enhanced Generative Framework With LLMs for Multimodal Emotion-Cause Pair Extraction in ConversationsabstractEmotion-Cause Pair Extraction (ECPE) in conversations aims to identify the emotional utterances (even their categories) along with their corresponding causal utterances, which is crucial in understanding the cause-effect relationship in dialogues. While prior studies of ECPE have predominantly focused on purely textual dialogues and neglected the exploration on the natural scenario of the dialogues with multimodal features, i.e., Multimodal Emotion-Cause Pair Extraction (MECPE) in conversations. To attempt this scenario, we propose a Generative approach for Multimodal Emotion-Cause pair extraction (GMEC) with a single stage, thus effectively reducing errors associated with the propagation and accumulation for MECPE. This approach can not only uniformly handle the information of diverse modalities, but also address all emotion and cause analysis tasks uniformly. Additionally, instead of utilizing the fixed commonsense knowledge base as previously, we resort to the Large Language Models (LLMs), which possess a powerful ability to emerge new knowledge, thereby acting as implicit knowledge engines for MECPE. We refer to this approach as enhanced GMEC. Extensive experimental results and detailed analysis demonstrate a notable improvement in the generative approach. Moreover, the integration of external knowledge from LLMs optimizes the efficiency of data utilization, particularly in few-shot scenarios. The integration of the generative model with LLMs has resulted in a cumulative enhancement of 4.94%, 10.90% on MECPE and MECPE-C (with emotion Category). Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Two Heads are Better than One: Zero-shot Cognitive Reasoning via Multi-LLM Knowledge FusionabstractCognitive reasoning holds a significant place within Natural Language Processing (NLP). Yet, the exploration of zero-shot scenarios, which align more closely with real-life situations than supervised scenarios, has been relatively limited. While a few studies have employed Large Language Models (LLMs) to tackle zero-shot cognitive reasoning tasks, they still grapple with two key challenges: 1) Traditional approaches rely on the chain-of-thought (CoT) mechanism, wherein LLMs are provided with a "Let's think step by step'' prompt. However, this schema may not accurately understand the meaning of a given question and ignores the possible learned knowledge (e.g., background or commonsense) of the LLMs about the questions, leading to incorrect answers. 2) Previous CoT methods normally exploit a single Large Language Model (LLM) and design many strategies to augment this LLM. We argue that the power of a single LLM is typically finite since it may not have learned some relevant knowledge about the question. To address these issues, we propose a Multi-LLM Knowledge Fusion (MLKF) approach, which resorts to heterogeneous knowledge emerging from multiple LLMs, for zero-shot cognitive reasoning tasks. Through extensive experiments and detailed analysis, we demonstrate that our MLKF can outperform the existing zero-shot or unsupervised state-of-the-art methods on four kinds of zero-shot tasks: aspect sentiment analysis, named entity recognition, question answering, and mathematical reasoning. Our code is available at https://github.com/trueBatty/MLKF Dong Zhang 0013, Shoushan Li, Guodong Zhou 0001, Erik Cambria |
CIKM | 2 |
| 2024 | Bilingual Multimodal Graph Modeling for Text-Image Relation Inference
Dong Zhang 0013, Shoushan Li, Guodong Zhou 0001 |
DASFAA (3) | 1 |
| 2024 | Cross-domain NER with Generated Task-Oriented Knowledge: An Empirical Study from Information Density PerspectiveabstractCross-domain Named Entity Recognition (CD-NER) is crucial for Knowledge Graph (KG) construction and natural language processing (NLP), enabling learning from source to target domains with limited data.Previous studies often rely on manually collected entity-relevant sentences from the web or attempt to bridge the gap between tokens and entity labels across domains.These approaches are time-consuming and inefficient, as these data are often weakly correlated with the target task and require extensive pre-training.To address these issues, we propose automatically generating task-oriented knowledge (GTOK) using large language models (LLMs), focusing on the reasoning process of entity extraction.Then, we employ taskoriented pre-training (TOPT) to facilitate domain adaptation.Additionally, current crossdomain NER methods often lack explicit explanations for their effectiveness.Therefore, we introduce the concept of information density to better evaluate the model's effectiveness before performing entity recognition.We conduct systematic experiments and analyses to demonstrate the effectiveness of our proposed approach and the validity of using information density for model evaluation † . Sophia Yat Mei Lee, Junshuang Wu, Dong Zhang 0013, Shoushan Li, Erik Cambria, Guodong Zhou 0001 |
EMNLP | 4 |
| 2024 | Generative Chain-of-Thought for Zero-Shot Cognitive Reasoning
Dong Zhang 0013, Suyang Zhu, Shoushan Li |
ICANN (5) | 2 |
| 2024 | Hair Transfer with Efficient Heuristic Chain of Editing
Yuansheng Ma, Dong Zhang 0013, Suyang Zhu, Shoushan Li |
ICANN (3) | 2 |
| 2024 | Comment-aided Video-Language Alignment via Contrastive Pre-training for Short-form Video Humor DetectionabstractThe growing importance of multi-modal humor detection within affective computing correlates with the expanding influence of short-form video sharing on social media platforms. In this paper, we propose a novel two-branch hierarchical model for short-form video humor detection (SVHD), named Comment-aided Video-Language Alignment (CVLA) via data-augmented multi-modal contrastive pre-training. Notably, our CVLA not only operates on raw signals across various modal channels but also yields an appropriate multi-modal representation by aligning the video and language components within a consistent semantic space. The experimental results on two humor detection datasets, including DY11k and UR-FUNNY, demonstrate that CVLA dramatically outperforms state-of-the-art and several competitive baseline approaches. Our dataset and code release at https://github.com/yliu-cs/CVLA. Yang Liu 0358, Tongfei Shen, Dong Zhang 0013, Qingying Sun, Shoushan Li, Guodong Zhou 0001 |
ICMR | 3 |
| 2024 | ECFCON: Emotion Consequence Forecasting in ConversationsabstractConversation is a common form of human communication that includes extensive emotional interaction. Traditional approaches focused on studying emotions and their underlying causes in conversations. They try to address two issues: what emotions are present in the dialogue and what causes these emotions. However, these works often overlook the bidirectional nature of emotional interaction in dialogue: utterances can evoke emotions (cause), and emotions can also lead to certain utterances (consequence). Therefore, we propose a new issue: what consequences arise from these emotions? This leads to the introduction of a new task called Emotion Consequence Forecasting in CONversations (ECFCON). In this work, we first propose a corresponding dialogue-level dataset. Specifically, we select 2,780 video dialogues for annotation, totaling 39,950 utterances. Out of these, 12,391 utterances contain emotions, and 8,810 of these have discernible consequences. Then, we benchmark this task by conducting experiments from the perspectives of traditional methods, generalized LLMs prompting methods, and clue-driven hybrid methods. Both our dataset and benchmark codes are openly accessible to the public. Xincheng Ju, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001 |
ACM Multimedia | 2 |
| 2024 | Response generation in multi-modal dialogues with split pre-generation and cross-modal contrasting
Linqin Li, Dong Zhang 0013, Suyang Zhu, Shoushan Li, Guodong Zhou 0001 |
Inf. Process. Manag. | 2 |
| 2023 | Real-time Emotion Pre-Recognition in Conversations with Contrastive Multi-modal Dialogue Pre-trainingabstractThis paper presents our pioneering effort in addressing a new and realistic scenario in multi-modal dialogue systems called Multi-modal Real-time Emotion Pre-recognition in Conversations (MREPC). The objective is to predict the emotion of a forthcoming target utterance that is highly likely to occur. We believe that this task can enhance the dialogue system's understanding of the interlocutor's state of mind, enabling it to prepare an appropriate response in advance. However, addressing MREPC poses the following challenges:1) Previous studies on emotion elicitation typically focus on textual modality and perform sentiment forecasting within a fixed contextual scenario. 2) Previous studies on multi-modal emotion recognition aim to predict the emotion of existing utterances, making it difficult to extend these approaches to MREPC due to the absence of the target utterance. To tackle these challenges, we construct two benchmark multi-modal datasets for MREPC and propose a task-specific multi-modal contrastive pre-training approach. This approach leverages large-scale unlabeled multi-modal dialogues to facilitate emotion pre-recognition for potential utterances of specific target speakers. Through detailed experiments and extensive analysis, we demonstrate that our proposed multi-modal contrastive pre-training architecture effectively enhances the performance of multi-modal real-time emotion pre-recognition in conversations. Xincheng Ju, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001 |
CIKM | 2 |
| 2023 | Comment-Aware Multi-Modal Heterogeneous Pre-Training for Humor Detection in Short-Form VideosabstractConventional humor analysis normally focuses on text, text-image pair, and even long video (e.g., monologue) scenarios. However, with the recent rise of short-form video sharing, humor detection in this scenario has not yet gained much exploration. To the best of our knowledge, there are two primary issues associated with short-form video humor detection (SVHD): 1) At present, there are no ready-made humor annotation samples in this scenario, and it takes a lot of manpower and material resources to obtain a large number of annotation samples; 2) Unlike the more typical audio and visual modalities, the titles (as opposed to simultaneous transcription in the lengthy film) and associated interactive comments in short-form videos may convey apparent humorous clues. Therefore, in this paper, we first collect and annotate a video dataset from DouYin (aka. TikTok in the world), namely DY24h, with hierarchical comments. Then, we also design a novel approach with comment-aided multi-modal heterogeneous pre-training (CMHP) to introduce comment modality in SVHD. Extensive experiments and analysis demonstrate that our CMHP beats several existing video-based approaches on DY24h, and that the comments modality further aids a better comprehension of humor. Our dataset, code and pre-trained models are available at https://github.com/yliu-cs/CMHP. Yang Liu 0358, Huanqin Ping, Dong Zhang 0013, Qingying Sun, Shoushan Li, Guodong Zhou 0001 |
ECAI | 3 |
| 2023 | A Unified MRC Framework with Multi-Query for Multi-modal Relation Triplets ExtractionabstractRelation triplets extraction (RTE) aims to extract all potential triplets of subject-object entities and the corresponding relation in a sentence, which is an extremely essential procedure in information extraction and knowledge graph construction. In the literature, dominant studies normally focus on text only, which neglects the fact that other modalities may have effective information available for entities and relations (e.g., the image accompanying the text on Twitter). Therefore, in this paper, we introduce a multi-modal task for relation triplets extraction (MRTE), which considers both text and image modalities. To well handle the multi-modality and achieve all potential triplets extraction, we propose a novel multi-modal approach by resorting to a unified machine reading comprehension (MRC) framework with multiple queries, namely MUMRC. Specifically, we design a unified framework to conduct multi-modal entity extraction, and multi-modal relation classification depending on all potential subject-object entity pairs, which recasts both sub-tasks as a machine reading comprehension (MRC) problem. In this unified MRC model, we incorporate multiple questions into original text data and leverage visual information to enhance the textual semantics. Extensive experiments on the MNRE dataset demonstrate that the proposed approach significantly outperforms the state-of-the-art baselines of textual RTE and multi-modal information extraction. Dong Zhang 0013, Shoushan Li, Guodong Zhou 0001 |
ICME | 2 |
| 2023 | A Text-Image Pair Is Not Enough: Language-Vision Relation Inference with Auxiliary Modality Translation
Dong Zhang 0013, Shoushan Li, Guodong Zhou 0001 |
NLPCC (2) | 2 |
| 2023 | A Benchmark for Hierarchical Emotion Cause Extraction in Spoken DialoguesabstractEmotion cause extraction (ECE) seeks to find out what causes a given emotion, which has drawn much attention in natural language and signal processing. Conventional ECE normally focuses on single level (i.e., either word-level or clause-level) in a document scenario. However, as we known, single level ECE can not satisfy the wide applications, compared with both levels. Besides, the existing dialogue systems increasingly need empathy support. Therefore, in this paper, we propose to hierarchically extract both word and utterance-level emotion causes in the spoken dialogue scenario (HECE). We first construct two datasets (i.e., HECE-DD and HECE-IE) based on previous studies, then propose a hierarchical framework, which consists of a feature extractor, an utterance-level cause extractor, and a word-level cause extractor. In this framework, utterance and word levels can naturally preform interaction. Detailed experiments on two HECE datasets demonstrate that hierarchical extraction performs better than extraction on a single level. Huanqin Ping, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Guodong Zhou 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Unified Multi-Modal Multi-Task Joint Learning for Language-Vision Relation InferenceabstractThe relationships between language and vision are valuable for natural language processing and computer vision research, where the text and image data are employed to develop computing techniques for image caption or visual grounding. Although the existing studies have been engaged in language- vision relation inference (LVRI), they are limited to low- resource of the task itself. In this paper, we mainly focus on LVRI upon text-image pairs in Twitter with a unified multi- modal multi-task joint learning approach. Different from the conventional multi-modal multi-task learning approach within the same multi-modal dataset, we leverage a relevant multi - modal task on the external dataset as an auxiliary task to facilitate the LVRI task. Systematic experiments demonstrate the effectiveness of our proposed multi-modal multi- taskjoint learning approach. Dong Zhang 0013 |
ICME | 2 |
| 2022 | Few-Shot Multi-Modal Sentiment Analysis with Prompt-Based Vision-Aware Language ModelingabstractAs a hot study topic in natural language processing, affec-tive computing and multimedia analysis, multi-modal senti-ment analysis (MSA) is widely explored on aspect-level and sentence-level tasks. However, the existing studies normally rely on a lot of annotated multi-modal data, which are difficult to collect due to the massive expenditure of manpower and re-sources, especially in some open-ended and fine-grained do-mains. Therefore, it is necessary to investigate the few-shot scenario for MSA. In this paper, we propose a prompt-based vision-aware language modeling (PVLM) approach to MSA, which only requires a few supervised data. Specifically, our PVLM can incorporate the visual information into pre-trained language model and leverage prompt tuning to bridge the gap between masked language prediction in pre-training and MSA tasks. Systematic experiments on three aspect-level and two sentence-level datasets of MSA demonstrate the effectiveness of our few-shot approach. Dong Zhang 0013 |
ICME | 2 |
| 2022 | Unified Multi-modal Pre-training for Few-shot Sentiment Analysis with Prompt-based LearningabstractMulti-modal sentiment analysis (MSA) has become more and more attractive in both academia and industry. The conventional studies normally require massive labeled data to train the deep neural models. To alleviate the above issue, in this paper, we conduct few-shot MSA with quite a small number of labeled samples. Inspired by the success of textual prompt-based fine-tuning (PF) approaches in few-shot scenario, we introduce a multi-modal prompt-based fine-tuning (MPF) approach. To narrow the semantic gap between language and vision, we propose unified pre-training for multi-modal prompt-based fine-tuning (UP-MPF) with two stages. First, in unified pre-training stage, we employ a simple and effective task to obtain coherent vision-language representations from fixed pre-trained language models (PLMs), i.e., predicting the rotation direction of the input image with a prompt phrase as input concurrently. Second, in multi-modal prompt-based fine-tuning, we freeze the visual encoder to reduce more parameters, which further facilitates few-shot MSA. Extensive experiments and analysis on three coarse-grained and three fine-grained MSA datasets demonstrate the better performance of our UP-MPF against the state-of-the-art of PF, MSA, and multi-modal pre-training approaches. Dong Zhang 0013, Shoushan Li |
ACM Multimedia | 2 |
| 2021 | Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message PassingabstractAs an important research issue in affective computing community, multi-modal emotion recognition has become a hot topic in the last few years. However, almost all existing studies perform multiple binary classification for each emotion with focus on complete time series data. In this paper, we focus on multi-modal emotion recognition in a multi-label scenario. In this scenario, we consider not only the label-to-label dependency, but also the feature-to-label and modality-to-label dependencies. Particularly, we propose a heterogeneous hierarchical message passing network to effectively model above dependencies. Furthermore, we propose a new multi-modal multi-label emotion dataset based on partial time-series content to show predominant generalization of our model. Detailed evaluation demonstrates the effectiveness of our approach. Dong Zhang 0013, Xincheng Ju, Junhui Li 0001, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
AAAI | 1 |
| 2021 | Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionabstractAspect terms extraction (ATE) and aspect sentiment classification (ASC) are two fundamental and fine-grained sub-tasks in aspect-level sentiment analysis (ALSA).In the textual analysis, jointly extracting both aspect terms and sentiment polarities has been drawn much attention due to the better applications than individual sub-task.However, in the multimodal scenario, the existing studies are limited to handle each sub-task independently, which fails to model the innate connection between the above two objectives and ignores the better applications.Therefore, in this paper, we are the first to jointly perform multi-modal ATE (MATE) and multi-modal ASC (MASC), and we propose a multi-modal joint learning approach with auxiliary cross-modal relation detection for multi-modal aspect-level sentiment analysis (MALSA).Specifically, we first build an auxiliary text-image relation detection module to control the proper exploitation of visual information.Second, we adopt the hierarchical framework to bridge the multi-modal connection between MATE and MASC, as well as separately visual guiding for each sub module.Finally, we can obtain all aspect-level sentiment polarities dependent on the jointly extracted specific aspects.Extensive experiments show the effectiveness of our approach against the joint textual approaches, pipeline and collapsed multi-modal approaches. Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Shoushan Li, Min Zhang 0005, Guodong Zhou 0001 |
EMNLP (1) | 2 |
| 2020 | Multi-modal Multi-label Emotion Detection with Modality and Label DependenceabstractAs an important research issue in the natural language processing community, multi-label emotion detection has been drawing more and more attention in the last few years. However, almost all existing studies focus on one modality (e.g., textual modality). In this paper, we focus on multi-label emotion detection in a multi-modal scenario. In this scenario, we need to consider both the dependence among different labels (label dependence) and the dependence between each predicting label and different modalities (modality dependence). Particularly, we propose a multi-modal sequence-to-set approach to effectively model both kinds of dependence in multi-modal multi-label emotion detection. The detailed evaluation demonstrates the effectiveness of our approach. Dong Zhang 0013, Xincheng Ju, Junhui Li 0001, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
EMNLP (1) | 1 |
| 2020 | Speaker Personality Recognition With Multimodal Explicit Many2many InteractionsabstractRecently, speaker personality analysis has become an increasingly popular research task in human-computer interaction. Previous studies of user personality traits recognition normally focus on leveraging static information, i.e., tweets, images and social relationships in social platforms and websites. However, in this paper, we utilize three kinds of speaking dynamic information, i.e., textual, visual and acoustic temporal sequences, for a computer to interpret human personality traits from a face-to-face monologue. Specifically, we propose an explicit many2many (many-to-many) interactive approach to help AI efficiently recognize speaker personality traits. On the one hand, we encode the long feature sequence of human speaking for each modality with bidirectional LSTM network. On the other hand, we design a many2many attention mechanism explicitly to capture the interactions across multiple modalities for multiple interactive pairs. Empirical evaluation on 12 kinds of personality traits demonstrates the effectiveness of our proposed approach to multimodal speaker personality recognition. Liangqing Wu, Dong Zhang 0013, Qiyuan Liu 0005, Shoushan Li, Guodong Zhou 0001 |
ICME | 2 |
| 2020 | Transformer-based Label Set Generation for Multi-modal Multi-label Emotion DetectionabstractMulti-modal utterance-level emotion detection has been a hot research topic in both multi-modal analysis and natural language processing communities. Different from traditional single-label multi-modal sentiment analysis, typical multi-modal emotion detection is naturally a multi-label problem where an utterance often contains multiple emotions. Existing studies normally focus on multi-modal fusion only and transform multi-label emotion classification into multiple binary classification problem independently. As a result, existing studies largely ignore two kinds of important dependency information: (1) Modality-to-label dependency, where different emotions can be inferred from different modalities, that is, different modalities contribute differently to each potential emotion. (2) Label-to-label dependency, where some emotions are more likely to coexist than those conflicting emotions. To simultaneously model above two kinds of dependency, we propose a unified approach, namely multi-modal emotion set generation network (MESGN) to generate an emotion set for an utterance. Specifically, we first employ a cross-modal transformer encoder to capture cross-modal interactions among different modalities, and a standard transformer encoder to capture temporal information for each modality-specific sequence given previous interactions. Then, we design a transformer-based discriminative decoding module equipped with modality-to-label attention to handle the modality-to-label dependency. In the meanwhile, we employ a reinforced decoding algorithm with self-critic learning to handle the label-to-label dependency. Finally, we validate the proposed MESGN architecture on a word-level aligned and unaligned multi-modal dataset. Detailed experimentation shows that our proposed MESGN architecture can effectively improve the performance of multi-modal multi-label emotion detection. Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Guodong Zhou 0001 |
ACM Multimedia | 2 |
| 2020 | Modeling both Intra- and Inter-modal Influence for Real-Time Emotion Detection in ConversationsabstractThrough much exploration in the past decade, emotion analysis in conversations was mainly conducted in textual scenario. Nowadays, with the popularization of speech and video communication, academia and industry have become gradually aware of the need in multimodal scenario. Therefore, emotion detection in conversations becomes increasingly hot not only in natural language processing (NLP) community but also in multimodal analysis community. Although previous studies normally argue that the emotion of current utterance in a conversation is much influenced by the content of historical utterances, their speakers and emotions, they model the influence derived from the history to the current utterance at the same granularity (Intra-modal influence). Intuitively, the clues of emotion detection may not exist in the history of the same modality as current utterance, but in the history of other modalities (Inter-modal influence). Besides, previous studies normally model the information propagation as the conversation flow. Intuitively, bidirectional modeling of information propagation in conversations provides rich clues for emotion detection. Therefore, this paper proposes a bidirectional dynamic dual influence network for real-time emotion detection in conversations, which can simultaneously model both intra- and inter-modal influence with bidirectional information propagation for current utterance and its historical utterances. Detailed experiments demonstrate that our approach much advances the state-of-the-art. Dong Zhang 0013, Weisheng Zhang, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
ACM Multimedia | 1 |
| 2019 | Modeling the Clause-Level Structure to Multimodal Sentiment Analysis via Reinforcement LearningabstractIn this paper, we propose a novel approach to multimodal sentiment analysis with focus on both textual and acoustic modalities. Especially, we utilize deep reinforcement learning to explore the clause-level structure in an utterance. On the basis, we perform multimodal interactions at clause-level to model hierarchical interactive representation for multimodal senitment analysis. Detailed evaluation on two benchmark datasets demonstrates the great effectiveness of our approach over several state-of-the-art baselines. Dong Zhang 0013, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
ICME | 1 |
| 2019 | Multi-Modal Language Analysis with Hierarchical Interaction-Level and Selection-Level AttentionsabstractAs an emerging research area in natural language processing, multi-modal human language analysis spans language, vision and audio modalities. Understanding multi-modal language requires not only the modeling of independent dynamics within each modality (intra-modal dynamics), but also more importantly interactive dynamics among different modalities (inter-modal dynamics). In this paper, we propose a hierarchical approach to multi-modal language analysis with two levels of attention mechanism, namely interaction-level, which captures the intra-modal and inter-modal dynamics across different modalities with multiple types of attention, and selection-level attention, which selects the effective representations for final prediction by calculating the importance of each vector obtained from interaction-level. Empirical evaluation demonstrates the effectiveness of our proposed approach to multi-modal sentiment classification, sentiment regression and emotion recognition. Dong Zhang 0013, Liangqing Wu, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
ICME | 1 |
| 2019 | Modeling both Context- and Speaker-Sensitive Dependence for Emotion Detection in Multi-speaker ConversationsabstractRecently, emotion detection in conversations becomes a hot research topic in the Natural Language Processing community. In this paper, we focus on emotion detection in multi-speaker conversations instead of traditional two-speaker conversations in existing studies. Different from non-conversation text, emotion detection in conversation text has one specific challenge in modeling the context-sensitive dependence. Besides, emotion detection in multi-speaker conversations endorses another specific challenge in modeling the speaker-sensitive dependence. To address above two challenges, we propose a conversational graph-based convolutional neural network. On the one hand, our approach represents each utterance and each speaker as a node. On the other hand, the context-sensitive dependence is represented by an undirected edge between two utterances nodes from the same conversation and the speaker-sensitive dependence is represented by an undirected edge between an utterance node and its speaker node. In this way, the entire conversational corpus can be symbolized as a large heterogeneous graph and the emotion detection task can be recast as a classification problem of the utterance nodes in the graph. The experimental results on a multi-modal and multi-speaker conversation corpus demonstrate the great effectiveness of the proposed approach. Dong Zhang 0013, Liangqing Wu, Changlong Sun, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
IJCAI | 1 |
| 2019 | Effective Sentiment-relevant Word Selection for Multi-modal Sentiment Analysis in Spoken LanguageabstractComputational modeling of human spoken language is an emerging research area in multimedia analysis spanning across the text and acoustic modalities. Multi-modal sentiment analysis is one of the most fundamental tasks in human spoken language understanding. In this paper, we propose a novel approach to selecting effective sentiment-relevant words for multi-modal sentiment analysis with focus on both the textual and acoustic modalities. Unlike the conventional soft attention mechanism, we employ a deep reinforcement learning mechanism to perform sentiment-relevant word selection and fully remove invalid words of each modality for multi-modal sentiment analysis. Specifically, we first align the raw text and audio at the word level and extract independent handcraft features for each modality to yield the textual and acoustic word sequence. Second, we establish two collaborative agents to deal with the textual and acoustic modalities in spoken language respectively. On this basis, we formulate the sentiment-relevant word selection process in a multi-modal setting as a multi-agent sequential decision problem and solve it with a multi-agent reinforcement learning approach. Detailed evaluations of multi-modal sentiment classification and emotion recognition on three benchmark datasets demonstrate the great effectiveness of our approach over several conventional competitive baselines. Dong Zhang 0013, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001 |
ACM Multimedia | 1 |
| 2019 | Hierarchical-Gate Multimodal Network for Human Communication Comprehension
Qiyuan Liu 0005, Liangqing Wu, Yang Xu 0027, Dong Zhang 0013, Shoushan Li, Guodong Zhou 0001 |
NLPCC (2) | 4 |
| 2019 | Many vs. Many Query Matching with Hierarchical BERT and Transformer
Yang Xu 0027, Qiyuan Liu 0005, Dong Zhang 0013, Shoushan Li, Guodong Zhou 0001 |
NLPCC (1) | 3 |
| 2016 | Two-View Label Propagation to Semi-supervised Reader Emotion ClassificationabstractIn the literature, various supervised learning approaches have been adopted to address the task of reader emotion classification. However, the classification performance greatly suffers when the size of the labeled data is limited. In this paper, we propose a two-view label propagation approach to semi-supervised reader emotion classification by exploiting two views, namely source text and response text in a label propagation algorithm. Specifically, our approach depends on two word-document bipartite graphs to model the relationship among the samples in the two views respectively. Besides, the two bipartite graphs are integrated by linking each source text sample with its corresponding response text sample via a length-sensitive transition probability. In this way, our two-view label propagation approach to semi-supervised reader emotion classification largely alleviates the reliance on the strong sufficiency and independence assumptions of the two views, as required in co-training. Empirical evaluation demonstrates the effectiveness of our two-view label propagation approach to semi-supervised reader emotion classification. Shoushan Li, Dong Zhang 0013, Guodong Zhou 0001 |
COLING | 3 |
| 2016 | User Classification with Multiple Textual PerspectivesabstractTextual information is of critical importance for automatic user classification in social media. However, most previous studies model textual features in a single perspective while the text in a user homepage typically possesses different styles of text, such as original message and comment from others. In this paper, we propose a novel approach, namely ensemble LSTM, to user classification by incorporating multiple textual perspectives. Specifically, our approach first learns a LSTM representation with a LSTM recurrent neural network and then presents a joint learning method to integrating all naturally-divided textual perspectives. Empirical studies on two basic user classification tasks, i.e., gender classification and age classification, demonstrate the effectiveness of the proposed approach to user classification with multiple textual perspectives. Dong Zhang 0013, Shoushan Li, Hongling Wang, Guodong Zhou 0001 |
COLING | 1 |