Chen Gong 0004

dblp:21/8587-4 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-2570-6969ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation
abstract
Multimodal emotion cause analysis in conversation aims to identify the causes of emotions by leveraging multimodal information.Existing studies mainly formulate this problem as either utterance-level emotion cause extraction, which provides clear cause localization but limited explanation, or multimodal emotion cause generation, which offers fine-grained explanations but lacks explicit traceability to source utterances.Moreover, existing datasets rely heavily on human judgment and lack well-defined structured theoretical guidance, leading to subjective and inconsistent annotations.To address these issues, we introduce joint Multimodal Emotion Cause Extraction and Summarization in conversation (MECES), a new task that simultaneously extracts emotion cause utterances and generates cause summaries, enabling both precise localization and interpretable explanations of emotion cause.We further construct a MECES dataset guided by the Activating events-Beliefs-Consequences theory from psychology.This dataset consists of 5,787 emotion utterances annotated with causes, comprising 12,231 emotion-cause pairs and 6,040 cause summaries.We also propose an effective endto-end joint learning approach for MECES task, establishing strong benchmark results for this newly introduced task and dataset.( U1,U2 , "Chuan Bai handed Guang Shi a gift, and Guang Shi didn't expect to receive one too.") Guang Shi:"I got a gift too!"
Jikun Wan, Chen Gong 0004, Guohong Fu
ACL (1)2
2026 Bridging modalities: a unified framework for textual and multimodal dialogue discourse parsing
Chen Gong 0004, Guohong Fu
Frontiers Comput. Sci.1
2026 Enhancing multi-modal aspect-based sentiment classification via emotional semantic-aware cross-modal relation inference
Chen Gong 0004, Guohong Fu
Inf. Process. Manag.2
2025 Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach
abstract
Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content.In the era of rapidly growing multimodal content and social media, MCR is particularly crucial for interpreting user interactions and bridging text-visual references to improve communication and personalization.However, MCR research for real-world dialogues remains unexplored due to the lack of sufficient data resources.To address this gap, we introduce TikTalkCoref, the first Chinese multimodal coreference dataset for social media in real-world scenarios, derived from the popular Douyin short-video platform.This dataset pairs short videos with corresponding textual dialogues from user comments and includes manually annotated coreference clusters for both person mentions in the text and the coreferential person head regions in the corresponding video frames.We also present an effective benchmark approach for MCR, focusing on the celebrity domain, and conduct extensive experiments on our dataset, providing reliable benchmark results for this newly constructed dataset.We release the TikTalk-Coref dataset to facilitate future research on MCR for real-world social media dialogues at https://github.com/lxystaruni/TikTalkCoref.
Chen Gong 0004, Guohong Fu
ACL (1)2
2025 Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation
abstract
Inspired by early research on exploring naturally annotated data for Chinese Word Segmentation (CWS), and also by recent research on integration of speech and text processing, this work for the first time proposes to explicitly mine word boundaries from parallel speech-text data. We employ the Montreal Forced Aligner (MFA) toolkit to perform character-level alignment on speech-text data, giving pauses as candidate word boundaries. Based on detailed analysis of collected pauses, we propose an effective probability-based strategy for filtering unreliable word boundaries. To more effectively utilize word boundaries as extra training data, we also propose a robust complete-then-train (CTT) strategy. We conduct cross-domain CWS experiments on two target domains, i.e., ZX and AISHELL2. We have annotated about 1K sentences as the evaluation data of AISHELL2. Experiments demonstrate the effectiveness of our proposed approach.
Zhenghua Li, Shilin Zhou 0002, Chen Gong 0004, Yang Hou 0001
COLING5
2025 Data Augmentation for Cross-domain Parsing via Lightweight LLM Generation and Tree Hybridization
abstract
Cross-domain constituency parsing remains a challenging task due to the lack of high-quality out-of-domain data. In this paper, we propose a data augmentation method via lightweight large language model (LLM) generation and tree hybridization. We utilize LLM to generate phrase structures (subtrees) for the target domain by incorporating grammar rules and lexical head information into the prompt. To better leverage LLM-generated target-domain subtrees, we hybridize them with existing source-domain subtrees to efficiently produce a large number of structurally diverse instances. Experimental results demonstrate that our method achieves significant improvements on five target domains with a lightweight LLM generation cost.
Yang Hou 0001, Chen Gong 0004, Zhenghua Li
COLING3
2025 2S-DGM4: A Two-Stage Framework for Detecting and Grounding Multi-Modal Media Manipulation
abstract
Detecting and grounding multi-modal media manipulation (DGM4) aims to identify the authenticity of media content in the form of image-text pairs, and locate forgery contents including text tokens or image regions. Existing methods adopt joint optimization frameworks to train multiple dedicated heads for different sub-tasks. This paradigm is prone to yielding models with insufficient learning regarding manipulation traces. In this paper, we propose a two-stage learning framework, named 2S-DGM4, exploiting correlations between manipulation detection and grounding for debunking manipulation content. It introduces modality-specific manipulation types into the subsequent manipulation grounding, and naturally conducts the mutual agreement among various sub-tasks. Apart from the global context in cross-modal interaction, we develop a fine-grained refinement module to capture subtle manipulation traces. Experiments demonstrate that 2S-DGM4outperforms state-of-the-art methods on the public benchmark dataset DGM4. Further analyses verify the efficacy of 2S-DGM4on manipulation reasoning.
Junjie Wu 0005, Yumeng Fu, Chen Gong 0004, Guohong Fu
ICME4
2025 Speaker Intention Enhanced Dialogue Discourse Parsing
abstract
Dialogue Discourse Parsing (DDP) task focuses on capturing the structural relations between utterances in a dialogue, represented as a dependency tree. Previous studies have utilized speaker information to enhance the understanding of dialogue semantics and improve model performance on the DDP task. However, the role of speaker intention, which reflects the psychological states of speakers, has been largely overlooked. Speaker intention is crucial for understanding dialogue semantics and has proven beneficial for various conversational tasks. To address this gap, we propose the Speaker Intention Enhanced Dialogue Discourse Parsing Model (SIEDDP), which integrates speaker intention into the DDP process. We utilize the Large Language Model (LLM) such as LLaMA3 to extract speaker intention from the dialogue and compare the intention with the traditional method leveraging the COMmonsEnse Transformer (COMET) model. Experimental results show that incorporating speaker intention significantly improves DDP performance. Further analysis demonstrate that speaker intention assists in inferring the link and relation between current utterance and the context.
Suxian Zhao, Chen Gong 0004, Guohong Fu
IJCNN3
2025 Annotation error detection in painstakingly annotated data: Part-of-speech tagging as a case study
Zhenghua Li, Chen Gong 0004, Shilin Zhou 0002, Min Zhang 0005
Expert Syst. Appl.3
2024 Improving Chinese Named Entity Recognition with Multi-grained Words and Part-of-Speech Tags via Joint Modeling
abstract
Nowadays, character-based sequence labeling becomes the mainstream Chinese named entity recognition (CNER) approach, instead of word-based methods, since the latter degrades performance due to propagation of word segmentation (WS) errors. To make use of WS information, previous studies usually learn CNER and WS simultaneously with multi-task learning (MTL) framework, or treat WS information as extra guide features for CNER model, in which the utilization of WS information is indirect and shallow. In light of the complementary information inside multi-grained words, and the close connection between named entities and part-of-speech (POS) tags, this work proposes a tree parsing approach for joint modeling CNER, multi-grained word segmentation (MWS) and POS tagging tasks simultaneously. Specifically, we first propose a unified tree representation for MWS, POS tagging, and CNER.Then, we automatically construct the MWS-POS-NER data based on the unified tree representation for model training. Finally, we present a two-stage joint tree parsing framework. Experimental results on OntoNotes4 and OntoNotes5 show that our proposed approach of jointly modeling CNER with MWS and POS tagging achieves better or comparable performance with latest methods.
Chenhui Dou, Chen Gong 0004, Zhenghua Li, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005
LREC/COLING2
2024 PromptCD: Coupled and Decoupled Prompt Learning for Vision-Language Models
abstract
Large-scale pre-trained vision-language models (VLMs), like CLIP, have presented striking generalizability for adapting to image classification in a few shot setting. Most existing methods explore a set of learnable tokens, such as prompt learning, on data-efficient utilization for task adaptation. However, they focus on either the coupled-modality property by prompt projection or decoupled-modality characteristic by prompt consistency, which ignores effective interaction between prompts. To model the deep yet sufficient cross-modal interaction and enhance the generalization between both seen and unseen tasks, in this paper, we propose a novel coupled and decoupled prompt learning framework, dubbed PromptCD, for vision-language models. Specifically, we introduce a bi-directional coupled-modality mechanism to intensify the interaction between both vision and language branches. Additionally, we propose mixture consistency to further improve the generalization and discrimination of the models on unseen tasks. The integration of such a mechanism and consistency facilitates the proposed framework adaptation for various downstream tasks. We conduct extensive experiments on 11 image classification datasets under a range of evaluation protocols, including base-to-novel and domain generalization, and cross-dataset recognition. Experimental results demonstrate that our proposed PromptCD overall outperforms state-of-the-art methods.
Junjie Wu 0005, Mingjie Sun, Chen Gong 0004, Guohong Fu
ECAI3
2023 Multi-grained Aspect Fusion for Review Response Generation
Chen Gong 0004, Dexin Kong, Guohong Fu
ICANN (9)2
2023 Discourse-Aware Causal Emotion Entailment
Dexin Kong, Chen Gong 0004, Guohong Fu
ICONIP (9)5
2023 MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with Instruction
abstract
Multimodal named entity recognition (MNER) is an essential task of vision and language, which aims to locate named entities and classify them to the predefined categories using visual scenarios. However, existing MNER studies often suffer from bias issues with fine-grained visual cue fusion, which may produce noisy coarse-grained visual cues for MNER. To accurately capture text-image relations and better refine multimodal representations, we propose a novel instruction-based Multi-granularity Cross-modality Generative framework for MNER, namely MCG-MNER. Concretely, we introduce a multi-granularity relation propagation to infer visual clues relevant to text. Then, we propose a method to jnject multi-granularity visual information into cross-modality interaction and fusion to learn a unified representation. Finally, we integrate task-specific instructions and answers for MCG-MNER. Comprehensive experimental results on three benchmark datasets, such as Twitter2015, Twitter2017 and WikiDiverse, demonstrate the superiority of our proposed method over several state-of-the-art MNER methods. We will publicly release our codes for future studies.
Junjie Wu 0005, Chen Gong 0004, Ziqiang Cao, Guohong Fu
ACM Multimedia2
2022 MuCPAD: A Multi-Domain Chinese Predicate-Argument Dataset
abstract
Yahui Liu, Haoping Yang, Chen Gong, Qingrong Xia, Zhenghua Li, Min Zhang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Haoping Yang, Chen Gong 0004, Qingrong Xia, Zhenghua Li, Min Zhang 0005
NAACL-HLT3
2022 Neural Coupled Sequence Labeling for Heterogeneous Annotation Conversion
abstract
Supervised statistical models rely on large-scale high-quality labeled data, which is important for model training but expensive to construct. Therefore, instead of constructing new dataset, researchers have attempted to make full use of various existing heterogeneous datasets to boost model performance, considering it is ubiquitous that the same task may have multiple annotated data following different and incompatible annotation guidelines. Representative methods include the guide-feature method which use the knowledge projected from the source-side to the target-side as extra features for target model guidance, and the multi-task learning (MTL) method which simultaneously train on multiple heterogeneous annotations with shared parameters to gain resource-share knowledge. Though effective, the guide-feature method fails to directly use the source-side data as training data, and the MTL method ignores the implicit mappings between heterogeneous datasets. Compared with the above methods, directly converting the heterogeneous datasets into homogeneous datasets for target model training is a more straightforward and effective way to fully exploit heterogeneous resources. In this work, we propose a neural coupled sequence labeling model for heterogeneous annotation conversion. First, for each token, we map a given one-side tag into a set of bundled tags by concatenating the tag with all the possible tags at the other side. Then, we build a neural coupled model over the bundled tag space. Finally, we convert heterogeneous annotations into homogeneous annotations by performing constraint decoding on the coupled model. We also propose a pruning strategy to address the oversize issue of the bundled tag space, which improves efficiency without hurting model performance.Experiments for part-of-speech (POS) tagging, word segmentation (WS), and WS&POS tagging tasks show that our proposed neural coupled model consistently outperforms several benchmark models for all the three tasks by large margin.
Chen Gong 0004, Zhenghua Li, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 An In-depth Study on Internal Structure of Chinese Words
abstract
Chen Gong, Saihao Huang, Houquan Zhou, Zhenghua Li, Min Zhang, Zhefeng Wang, Baoxing Huai, Nicholas Jing Yuan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Gong 0004, Saihao Huang, Houquan Zhou 0001, Zhenghua Li, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan
ACL/IJCNLP (1)1
2020 Multi-grained Chinese Word Segmentation with Weakly Labeled Data
abstract
In contrast with the traditional single-grained word segmentation (SWS), where a sentence corresponds to a single word sequence, multi-grained Chinese word segmentation (MWS) aims to segment a sentence into multiple word sequences to preserve all words of different granularities.Due to the lack of manually annotated MWS data, previous work train and tune MWS models only on automatically generated pseudo MWS data.In this work, we further take advantage of the rich word boundary information in existing SWS data and naturally annotated data from dictionary example (DictEx) sentences, to advance the state-of-the-art MWS model based on the idea of weak supervision.Particularly, we propose to accommodate two types of weakly labeled data for MWS, i.e., SWS data and DictEx data by employing a simple yet competitive graph-based parser with local loss.Besides, we manually annotate a high-quality MWS dataset according to our newly compiled annotation guideline, consisting of over 9,000 sentences from two types of texts, i.e., canonical newswire (NEWS) and non-canonical web (BAIKE) data for better evaluation.Detailed evaluation shows that our proposed model with weakly labeled data significantly outperforms the state-of-the-art MWS model by 1.12 and 5.97 on NEWS and BAIKE data in F1.
Chen Gong 0004, Zhenghua Li, Bowei Zou, Min Zhang 0005
COLING1
2020 Hierarchical LSTM with char-subword-word tree-structure representation for Chinese named entity recognition
Chen Gong 0004, Zhenghua Li, Qingrong Xia, Wenliang Chen, Min Zhang 0005
Sci. China Inf. Sci.1
2017 Multi-Grained Chinese Word Segmentation
abstract
Traditionally, word segmentation (WS) adopts the single-granularity formalism, where a sentence corresponds to a single word sequence.However, Sproat et al. (1996) show that the inter-nativespeaker consistency ratio over Chinese word boundaries is only 76%, indicating single-grained WS (SWS) imposes unnecessary challenges on both manual annotation and statistical modeling.Moreover, WS results of different granularities can be complementary and beneficial for high-level applications.This work proposes and addresses multi-grained WS (MWS).First, we build a large-scale pseudo MWS dataset for model training and tuning by leveraging the annotation heterogeneity of three SWS datasets.Then we manually annotate 1,500 test sentences with true MWS annotations.Finally, we propose three benchmark approaches by casting MWS as constituent parsing and sequence labeling.Experiments and analysis lead to many interesting findings.
Chen Gong 0004, Zhenghua Li, Min Zhang 0005, Xinzhou Jiang
EMNLP1