VLDB 2026 Research / reviewers in the wild / expert
Chong Teng
dblp:33/2162
· DBLP profile ↗
50ranked-venue papers
1as first author
43since 2021 · last 2026
0009-0008-6543-2548ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 1 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PaSE: Prototype-aligned Calibration and Shapley-based Equilibrium for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) seeks to understand human emotions by integrating textual, acoustic, and visual signals. Although multimodal fusion is designed to leverage cross-modal complementarity, real-world scenarios often exhibit modality competition: dominant modalities tend to overshadow weaker ones, leading to suboptimal performance. In this paper, we propose PaSE, a novel Prototype-aligned Calibration and Shapley-optimized Equilibrium framework, which enhances collaboration while explicitly mitigating modality competition. PaSE first applies Prototype-guided Calibration Learning (PCL) to refine unimodal representations and align them through an Entropic Optimal Transport mechanism that ensures semantic consistency. To further stabilize optimization, we introduce a Dual-Phase Optimization strategy. A prototype-gated fusion module is first used to extract shared representations, followed by Shapley-based Gradient Modulation (SGM), which adaptively adjusts gradients according to the contribution of each modality. Extensive experiments on IEMOCAP, MOSI, and MOSEI confirm that PaSE achieves the superior performance and effectively alleviates modality competition. Kang He 0005, Yuzhe Ding 0001, Fei Li 0021, Chong Teng, Donghong Ji |
AAAI | 5 |
| 2026 | Dynamic Emotion and Personality Profiling for Multimodal Deception DetectionabstractLi Zheng, Yanyi Luo, Hao Fei, Yuzhe Ding, Yujie Huang, Fei Li, Chong Teng, Donghong Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yanyi Luo, Hao Fei 0001, Yuzhe Ding 0001, Fei Li 0021, Chong Teng, Donghong Ji |
ACL (1) | 7 |
| 2026 | Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction SteeringabstractLi Zheng, Xin Zhang, Shuyi He, Fei Li, Chong Teng, Jiang-Ming Yang, Donghong Ji, Zhuang Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shuyi He, Fei Li 0021, Chong Teng, Jiang-Ming Yang, Donghong Ji, Zhuang Li 0001 |
ACL (1) | 5 |
| 2026 | Generative implicit opinion mining with term correlation prompts
Fei Li 0021, Fangfang Su, Kamran Aziz, Jingcheng Yuan, Chong Teng, Donghong Ji |
Inf. Sci. | 6 |
| 2026 | Improving Emotion and Intent Understanding in Multimodal Conversations With Progressive InteractionabstractEmotion and intent joint understanding in multimodal conversations (MC-EIU) aims to model the semantic dependencies among multimodal conversations while inferring emotion and intent information. Despite making progress, existing methods overlook the differential contributions of modalities and rely on single-round interactions for emotion and intent recognition, resulting in suboptimal model understanding performance. To overcome these limitations, we propose MEIPro, a novel progressive interaction and adaptive weight fusion based multimodal joint understanding of emotion and intent framework. We first design a hierarchical denoising module to effectively remove noise and redundant information from multimodal data. Then, we propose an adaptive weight fusion mechanism that dynamically fuses multimodal features by taking the true classification probabilities of each modality as their respective contributions, thus enhancing the fusion process. Additionally, we present a progressive dual task interaction module to capture the deep seated interactions between emotion and intent through a step-by-step multi-round iteration. Experiments on the benchmark MC-EIU bilingual dataset demonstrate that our MEI-Pro framework significantly outperforms state-of-the-art baselines in both emotion and intent tasks. Specifically, on the English dataset, the F1-scores of the multimodal emotion and intent understanding tasks have increased by 6.12% and 7.25% respectively. Tengyue Song, Yuzhe Ding 0001, Xiaorui Wu, Fei Li 0021, Dongdong Xie 0003, Chong Teng, Donghong Ji |
IEEE Trans. Affect. Comput. | 8 |
| 2025 | Multi-Granular Multimodal Clue Fusion for Meme UnderstandingabstractWith the continuous emergence of various social media platforms frequently used in daily life, the multimodal meme understanding (MMU) task has been garnering increasing attention. MMU aims to explore and comprehend the meanings of memes from various perspectives by performing tasks such as metaphor recognition, sentiment analysis, intention detection, and offensiveness detection. Despite making progress, limitations persist due to the loss of fine-grained metaphorical visual clue and the neglect of multimodal text-image weak correlation. To overcome these limitations, we propose a multi-granular multimodal clue fusion model (MGMCF) to advance MMU. Firstly, we design an object-level semantic mining module to extract object-level image feature clues, achieving fine-grained feature clue extraction and enhancing the model's ability to capture metaphorical details and semantics. Secondly, we propose a brand-new global-local cross-modal interaction model to address the weak correlation between text and images. This model facilitates effective interaction between global multimodal contextual clues and local unimodal feature clues, strengthening their representations through a bidirectional cross-modal attention mechanism. Finally, we devise a dual-semantic guided training strategy to enhance the model's understanding and alignment of multimodal representations in the semantic space. Experiments conducted on the widely-used MET-MEME bilingual dataset demonstrate significant improvements over state-of-the-art baselines. Specifically, there is an 8.14% increase in precision for offensiveness detection task, and respective accuracy enhancements of 3.53%, 3.89%, and 3.52% for metaphor recognition, sentiment analysis, and intention detection tasks. These results, underpinned by in-depth analyses, underscore the effectiveness and potential of our approach for advancing MMU. Hao Fei 0001, Zuquan Peng, Fei Li 0021, Huisheng Ma, Chong Teng, Donghong Ji |
AAAI | 7 |
| 2025 | TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data SynthesisabstractXiaorui Wu, Xiaofeng Mao, Fei Li, Xin Zhang, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaorui Wu, Xiaofeng Mao, Fei Li 0021, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li 0001 |
ACL (1) | 6 |
| 2025 | Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion KnowledgeabstractLi Zheng, Sihang Wang, Hao Fei, Zuquan Peng, Fei Li, Jianming Fu, Chong Teng, Donghong Ji. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sihang Wang, Hao Fei 0001, Zuquan Peng, Fei Li 0021, Jianming Fu, Chong Teng, Donghong Ji |
ACL (1) | 7 |
| 2025 | DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph RefinementabstractVision-Language Models (VLMs) generate discourse-level, multi-sentence visual descriptions, challenging text scene graph parsers built for single-sentence caption-to-graph mapping.Current approaches typically merge sentencelevel parsing outputs for discourse input, often missing phenomena like cross-sentence coreference, resulting in fragmented graphs and degraded downstream VLM task performance.We introduce a new task, Discourse-level text Scene Graph parsing (DiscoSG), and release DiscoSG-DS, a dataset of 400 expert-annotated and 8,430 synthesised multi-sentence captiongraph pairs.Each caption averages 9 sentences, and each graph contains at least 3× more triples than those in existing datasets.Fine-tuning GPT-4o on DiscoSG-DS yields over 40% higher SPICE metric than the best sentence-merging baseline.However, its high inference cost and licensing restrict opensource use.Smaller fine-tuned open-source models (e.g., Flan-T5) perform well on simpler graphs yet degrade on denser, more complex graphs.To bridge this gap, we introduce DiscoSG-Refiner, a lightweight open-source parser that drafts a seed graph and iteratively refines it with a novel learned graph-editing model, achieving 30% higher SPICE than the baseline while delivering 86× faster inference than GPT-4o.It generalises from simple to dense graphs, thereby consistently improving downstream VLM tasks, including discourselevel caption evaluation and hallucination detection, outperforming alternative open-source parsers. Shaoqing Lin, Chong Teng, Fei Li 0021, Donghong Ji, Lizhen Qu, Zhuang Li 0001 |
EMNLP | 2 |
| 2025 | Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability PredictionabstractNatural Language Inference (NLI) is a fundamental task in natural language processing.While NLI has developed many subdirections such as sentence-level NLI, documentlevel NLI and cross-lingual NLI, Cross-Document Cross-Lingual NLI (CDCL-NLI) remains largely unexplored.In this paper, we propose a novel paradigm: CDCL-NLI, which extends traditional NLI capabilities to multidocument, multilingual scenarios.To support this task, we construct a high-quality CDCL-NLI dataset including 25,410 instances and spanning 26 languages.To address the limitations of previous methods on CDCL-NLI task, we further propose an innovative method that integrates RST-enhanced graph fusion with interpretability-aware prediction.Our approach leverages RST (Rhetorical Structure Theory) within heterogeneous graph neural networks for cross-document context modeling, and employs a structure-aware semantic alignment based on lexical chains for crosslingual understanding.For NLI interpretability, we develop an EDU (Elementary Discourse Unit)-level attribution framework that produces extractive explanations.Extensive experiments demonstrate our approach's superior performance, achieving significant improvements over both conventional NLI models as well as large language models.Our work sheds light on the study of NLI and will bring research interest on cross-document cross-lingual context understanding, hallucination elimination and interpretability inference.Our code and dataset are available at CDCL- NLI-link. Mengying Yuan, Wenhao Wang 0002, Kangli Wei, Fei Li 0021, Chong Teng, Donghong Ji |
EMNLP | 7 |
| 2025 | Harnessing Dimensional Contrast and Information Compensation for Sentence Embedding EnhancementabstractUnsupervised sentence embedding learning excels through positive sample construction and instance-level contrastive learning (ICL). However, this approach can lead to over-compression and dimensional contamination from noisy data augmentation and unconstrained ICL processes. To mitigate these issues, we propose a novel enhancement method, MSSE, which incorporates an Information Compensation Mechanism (ICM) and a Dimensional-Level Contrastive Learning Mechanism (DCM). ICM, inspired by the information bottleneck principle, prevents excessive compression in representation learning. DCM constrains ICL and minimizes information leakage across dimensions. Experimental results show that MSSE surpasses existing competitive baselines across seven STS tasks, including unsupervised and few-shot scenarios. The source code is available at https://github.com/Hekang001/MSSE. Kang He 0005, Yuzhe Ding 0001, Bobo Li 0001, Fei Li 0021, Chong Teng, Donghong Ji |
ICASSP | 6 |
| 2025 | An Effective Multimodal Rumor Detection Model via Image Semantic Enhancement and Hierarchical FusionabstractThe rapid development of social platforms and multimedia technologies has led to the proliferation of multimodal rumors on these platforms. Therefore, automatic multimodal rumor detection has received extensive attention from researchers. Although many existing methods exhibit strong capabilities to identify multimodal rumors, they still have shortcomings in fully utilizing the information contained in images and performing fine-grained fusion of features from different modalities. In this paper, we propose an effective model via Image Semantic Enhancement and Hierarchical Fusion (ISEHF) for multimodal rumor detection. Specifically, ISEHF performs image semantic enhancement by employing a large vision-language model to obtain image captions, which is critical to fully utilizing the information contained in images. Moreover, we propose a hierarchical fusion module within the ISEHF model, which consists of a shallow fusion module and a deep fusion module, to generate richer features at different levels for fine-grained fusion of features from different modalities. Extensive experiments conducted on two public datasets demonstrate the superiority of our model in comparison with the state-of-the-art baselines. Peiwen Yan, Kang He 0005, Bobo Li 0001, Fei Li 0021, Chong Teng, Donghong Ji |
IJCNN | 6 |
| 2025 | EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious InstructionsabstractLarge language models (LLMs) frequently refuse to respond to pseudo-malicious instructions: semantically harmless input queries triggering unnecessary LLM refusals due to conservative safety alignment, significantly impairing user experience. Collecting such instructions is crucial for evaluating and mitigating over-refusals, but existing instruction curation methods, like manual creation or instruction rewriting, either lack scalability or fail to produce sufficiently diverse and effective refusal-inducing prompts. To address these limitations, we introduce EVOREFUSE, a prompt optimization approach that generates diverse pseudo-malicious instructions consistently eliciting confident refusals across LLMs. EVOREFUSE employs an evolutionary algorithm exploring the instruction space in more diverse directions than existing methods via mutation strategies and recombination, and iteratively evolves seed instructions to maximize evidence lower bound on LLM refusal probability. Using EVOREFUSE, we create two novel datasets: EVOREFUSE-TEST, a benchmark of 582 pseudo-malicious instructions that outperforms the next-best benchmark with 85.34% higher average refusal triggering rate across 9 LLMs without a safety-prior system prompt, 34.86% greater lexical diversity, and 40.03% improved LLM response confidence scores; and EVOREFUSE-ALIGN, which provides 3,000 pseudo-malicious instructions with responses for supervised and preference-based alignment training. With supervised fine-tuning on EVOREFUSE-ALIGN, LLAMA3.1-8B-INSTRUCT achieves up to 29.85% fewer over-refusals than models trained on the second-best alignment dataset, without compromising safety. Our analysis with EVOREFUSE-TEST reveals models trigger over-refusals by overly focusing on sensitive keywords while ignoring broader context. Our code and datasets are available at https://github.com/FishT0ucher/EVOREFUSE . Xiaorui Wu, Fei Li 0021, Xiaofeng Mao, Chong Teng, Donghong Ji, Zhuang Li 0001 |
NeurIPS | 7 |
| 2025 | Semantic Information Enhanced Fake News Detection
Chong Teng, Bocheng Ai, Fei Li 0021 |
NLPCC (3) | 1 |
| 2025 | STPar: A Structure-Aware Triaffine Parser for Screenplay Character Coreference ResolutionabstractAbstract Character Coreference Resolution in Movie Screenplays (MovieCoref) is a newly emerging task for understanding complex movie plots and character relationships. This task poses greater challenges than traditional coreference resolution, due to the intricate narrative structures and character interactions unique to screenplays. In light of these challenges, we introduce a novel approach: a Structure-aware Triaffine Parser (STPar) for the MovieCoref task. STPar combines discourse and syntactic structures in the feature encoding process, enabling comprehensive analysis of ternary relationships and complex interactions. During the pairing process, STPar utilizes a triaffine scorer to consider high-order relations between candidate mention pairs, thus enhancing its ability to capture detailed narrative structures. In addition, STPar incorporates multi-task learning, encompassing singleton and span detection tasks, to further improve coreference resolution performance. Our evaluations on the MovieCoref dataset demonstrate that STPar significantly outperforms the best baseline by 7.4%, 21.5%, 7.1%, and 10.2% in F1 scores of B3, CEAFe, LEA, and CoNLL. Further analysis highlights the benefits of integrating structural discourse and syntactic information as well as the combined approaches of triaffine and multi-task learning.1 Hao Fei 0001, Bobo Li 0001, Fei Li 0021, Chong Teng, Liang Zhao 0001, Donghong Ji |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation ApproachabstractDespite significant advancements in multi-label text classification, the ability of existing models to generalize to novel and seldom-encountered complex concepts, which are compositions of elementary ones, remains underexplored. This research addresses this gap. By creating unique data splits across three benchmarks, we assess the compositional generalization ability of existing multi-label text classification models. Our results show that these models often fail to generalize to compositional concepts encountered infrequently during training, leading to inferior performance on tests with these new combinations. To address this, we introduce a data augmentation method that leverages two innovative text generation models designed to enhance the classification models' capacity for compositional generalization. Our experiments show that this data augmentation approach significantly improves the compositional generalization capabilities of classification models on our benchmarks, with both generation models surpassing other text generation baselines. Our codes available at https://github.com/yychai74/LD-VAE. Yuyang Chai, Zhuang Li 0001, Fei Li 0021, Donghong Ji, Chong Teng |
AAAI | 7 |
| 2024 | Reverse Multi-Choice Dialogue Commonsense Inference with Graph-of-ThoughtabstractWith the proliferation of dialogic data across the Internet, the Dialogue Commonsense Multi-choice Question Answering (DC-MCQ) task has emerged as a response to the challenge of comprehending user queries and intentions. Although prevailing methodologies exhibit effectiveness in addressing single-choice questions, they encounter difficulties in handling multi-choice queries due to the heightened intricacy and informational density. In this paper, inspired by the human cognitive process of progressively excluding options, we propose a three-step Reverse Exclusion Graph-of-Thought (ReX-GoT) framework, including Option Exclusion, Error Analysis, and Combine Information. Specifically, our ReX-GoT mimics human reasoning by gradually excluding irrelevant options and learning the reasons for option errors to choose the optimal path of the GoT and ultimately infer the correct answer. By progressively integrating intricate clues, our method effectively reduces the difficulty of multi-choice reasoning and provides a novel solution for DC-MCQ. Extensive experiments on the CICERO and CICERO_v2 datasets validate the significant improvement of our approach on DC-MCQ task. On zero-shot setting, our model outperform the best baseline by 17.67% in terms of F1 score for the multi-choice task. Most strikingly, our GPT3.5-based ReX-GoT framework achieves a remarkable 39.44% increase in F1 score. Hao Fei 0001, Fei Li 0021, Bobo Li 0001, Lizi Liao, Donghong Ji, Chong Teng |
AAAI | 7 |
| 2024 | Revisiting Structured Sentiment Analysis as Latent Dependency Graph ParsingabstractStructured Sentiment Analysis (SSA) was cast as a problem of bi-lexical dependency graph parsing by prior studies.Multiple formulations have been proposed to construct the graph, which share several intrinsic drawbacks: (1) The internal structures of spans are neglected, thus only the boundary tokens of spans are used for relation prediction and span recognition, thus hindering the model's expressiveness; (2) Long spans occupy a significant proportion in the SSA datasets, which further exacerbates the problem of internal structure neglect.In this paper, we treat the SSA task as a dependency parsing task on partially-observed dependency trees, regarding flat spans without determined tree annotations as latent subtrees to consider internal structures of spans.We propose a twostage parsing method and leverage TreeCRFs with a novel constrained inside algorithm to model latent structures explicitly, which also takes advantages of joint scoring graph arcs and headed spans for global optimization and inference.Results of extensive experiments on five benchmark datasets reveal that our method performs significantly better than all previous bi-lexical methods, achieving new state-of-theart. Bobo Li 0001, Hao Fei 0001, Fei Li 0021, Chong Teng, Donghong Ji |
ACL (1) | 5 |
| 2024 | What Factors Influence LLMs' Judgments? A Case Study on Question AnsweringabstractLarge Language Models (LLMs) are now being considered as judges of high efficiency to evaluate the quality of answers generated by candidate models. However, their judgments may be influenced by complex scenarios and inherent biases, raising concerns about their reliability. This study aims to bridge this gap by introducing four unexplored factors and examining the performance of LLMs as judges, namely answer quantity, inducing statements, judging strategy, and judging style. Additionally, we introduce a new dimension of question difficulty to provide a more comprehensive understanding of LLMs’ judgments across varying question intricacies. We employ ChatGPT, GPT-4, Gemini, and Claude-2 as judges and conduct experiments on Vicuna Benchmark and MT-bench. Our study reveals that LLMs’ judging abilities are susceptible to the influence of these four factors, and analyzing from the newly proposed dimension of question difficulty is highly necessary. We also provide valuable insights into optimizing LLMs’ performance as judges, enhancing their reliability and adaptability across diverse evaluation scenarios. Bobo Li 0001, Zixiang Meng, Runfeng Shi, Hao Fei 0001, Fei Li 0021, Chong Teng, Donghong Ji |
LREC/COLING | 10 |
| 2024 | Enhancing Cross-Document Event Coreference Resolution by Discourse Structure and Semantic InformationabstractExisting cross-document event coreference resolution models, which either compute mention similarity directly or enhance mention representation by extracting event arguments (such as location, time, agent, and patient), lackingmthe ability to utilize document-level information. As a result, they struggle to capture long-distance dependencies. This shortcoming leads to their underwhelming performance in determining coreference for the events where their argument information relies on long-distance dependencies. In light of these limitations, we propose the construction of document-level Rhetorical Structure Theory (RST) trees and cross-document Lexical Chains to model the structural and semantic information of documents. Subsequently, cross-document heterogeneous graphs are constructed and GAT is utilized to learn the representations of events. Finally, a pair scorer calculates the similarity between each pair of events and co-referred events can be recognized using standard clustering algorithm. Additionally, as the existing cross-document event coreference datasets are limited to English, we have developed a large-scale Chinese cross-document event coreference dataset to fill this gap, which comprises 53,066 event mentions and 4,476 clusters. After applying our model on the English and Chinese datasets respectively, it outperforms all baselines by large margins. Qiang Gao 0008, Bobo Li 0001, Zixiang Meng, Fei Li 0021, Chong Teng, Donghong Ji |
LREC/COLING | 7 |
| 2024 | Self-Adaptive Fine-grained Multi-modal Data Augmentation for Semi-supervised Muti-modal Coreference ResolutionabstractCoreference resolution, an essential task in natural language processing, is particularly challenging in multi-modal scenarios where data comes in various forms and modalities. Despite advancements, limitations due to scarce labeled data and underleveraged unlabeled data persist. We address these issues with a self-adaptive fine-grained multi-modal data augmentation framework for semi-supervised MCR, focusing on enriching training data from labeled datasets and tapping into the untapped potential of unlabeled data. Regarding the former issue, we first leverage text coreference resolution datasets and diffusion models,to perform fine-grained text-to-image generation with aligned text entities and image bounding boxes. We then introduce a self-adaptive selection strategy, meticulously curating the augmented data to enhance the diversity and volume of the training set without compromising its quality. For the latter issue, we design a self-adaptive threshold strategy that dynamically adjusts the confidence threshold based on the model's learning status and performance, enabling effective utilization of valuable information from unlabeled data. Additionally, we incorporate a distance smoothing term, which smooths distances between positive and negative samples, enhancing discriminative power of the model?s feature representations and addressing noise and uncertainty in the unlabeled data. Our experiments on the widely-used CIN dataset show that our framework significantly outperforms state-of-the-art baselines by at least 9.57% on MUC F1 score and 4.92% on CoNLL F1 score. Remarkably, against weakly-supervised baselines, our framework achieves a staggering 22.24% enhancement in MUC F1 score. These results, underpinned by in-depth analyses, underscore the effectiveness and potential of our approach for advancing MCR tasks. Hao Fei 0001, Fei Li 0021, Shengqiong Wu, Lizi Liao, Donghong Ji, Chong Teng |
ACM Multimedia | 8 |
| 2024 | Generative Dialogue Sentiment and Act Recognition with Feature Denoising and Set Prediction
Bobo Li 0001, Zhuang Li 0001, Yuyang Chai, Fei Li 0021, Chong Teng, Donghong Ji |
NLPCC (5) | 6 |
| 2024 | Overview of the NLPCC 2024 Shared Task 3: Dialogue-Level Coreference Resolution and Relation Extraction
Yiyun Xiong, Fei Li 0021, Bobo Li 0001, Hao Fei 0001, Donghong Ji, Chong Teng |
NLPCC (5) | 6 |
| 2024 | MMLSCU: A Dataset for Multi-modal Multi-domain Live Streaming Comment UnderstandingabstractWith the increasing popularity of live streaming, the interactions from viewers during a live streaming can provide more specific and constructive feedback for both the streamer and platform. In such scenario, the primary and most direct feedback method from the audience is through comments. Thus, mining these live streaming comments to unearth the intentions behind them and, in turn, aiding streamers to enhance their live streaming quality is significant for the well development of live streaming ecosystem. To this end, we introduce the MMLSCU dataset, containing 50,129 intention-annotated comments across multiple modalities (text, images, vi-deos, audio) from eight streaming domains. Using multimodal pretrained large model and drawing inspiration from the Chain of Thoughts (CoT) concept, we implement an end-to-end model to sequentially perform the following tasks: viewer comment intent detection ➛ intent cause mining ➛ viewer comment explanation ➛ streamer policy suggestion. We employ distinct branches for video and audio to process their respective modalities. After obtaining the video and audio representations, we conduct a multimodal fusion with the comment. This integrated data is then fed into the large language model to perform inference across the four tasks following the CoT framework. Experimental results indicate that our model outperforms three multimodal classification baselines on comment intent detection and streamer policy suggestion, and one multimodal generation baselines on intent cause mining and viewer comment explanation. Compared to the models using only text, our multimodal setting yields superior outcomes. Moreover, incorporating CoT allows our model to enhance comment interpretation and more precise suggestions for the streamers. Our proposed dataset and model will bring new research attention on multimodal live streaming comment understanding. Zixiang Meng, Qiang Gao 0008, Bobo Li 0001, Hao Fei 0001, Shengqiong Wu, Fei Li 0021, Chong Teng, Donghong Ji |
WWW | 9 |
| 2024 | Joint rumour and stance identification based on semantic and structural information in social networks
Nanhang Luo, Dongdong Xie 0003, Yiwen Mo, Fei Li 0021, Chong Teng, Donghong Ji |
Appl. Intell. | 5 |
| 2024 | A tree-like structured perceptron for transition-based biomedical event extraction
Fangfang Su, Tao Qian 0002, Bobo Li 0001, Fei Li 0021, Chong Teng, Donghong Ji |
Knowl. Based Syst. | 6 |
| 2024 | Enhancing Chinese Event Extraction with Event Trigger StructuresabstractThe dependency syntactic structure is widely used in event extraction. However, the dependency structure reflecting syntactic features is essentially different from the event structure that reflects semantic features, leading to the performance degradation. In this article, we propose to use Event Trigger Structure for Event Extraction (ETSEE), which can compensate the inconsistency between two structures. First, we leverage the ACE2005 dataset as case study, and annotate three kinds of ETSs, that is, “light verb + trigger”, “preposition structures” and “tense + trigger”. Then we design a graph-based event extraction model that jointly identifies triggers and arguments, where the graph consists of both the dependency structure and ETSs. Experiments show that our model significantly outperforms the state-of-the-art methods. Through empirical analysis and manual observation, we find that the ETSs can bring the following benefits: (1) enriching trigger identification features by introducing structural event information; (2) enriching dependency structures with event semantic information; (3) enhancing the interactions between triggers and candidate arguments by shortening their distances in the dependency graph. Fei Li 0021, Kaifang Deng, Yiwen Mo, Yuanze Ji, Chong Teng, Donghong Ji |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2024 | Phrase-Aware Financial Sentiment Analysis Based on Constituent SyntaxabstractFinancial sentiment analysis is a fine-grained sentiment analysis task that needs to predict the sentiment value toward a given target entity. Recently, dependency-based graph neural networks have been introduced for target-based sentiment analysis. However, financial sentiment analysis with implicit sentiment expression is more challenging than target-based explicit sentiment analysis, requiring a deep understanding of the complex association between the sentiment clue in context and the target entity. In previous work related to financial sentiment analysis, most methods focused on learning the simple word-to-word relations between the contextual words and the target entity based on the dependency tree of the sentence, ignoring the exploitation of span-boundary information and phrase-level syntactic knowledge with regard to the target entity. In this paper, we perform financial implicit sentiment analysis by taking phrases as basic semantic units and proposing a graph attention network ($PhraseGAT$) based on the constituent tree to leverage the phrase syntactic knowledge. To enhance the information flow between the nodes in the graph, we construct a heterogeneous graph based on the constituent tree and encode higher-order neighbor information. In addition, we introduce a multi-edge-type graph attention network ($MET$-$GAT$) to take full consideration of syntax and semantic interactions for the final prediction on the sentiment value of the target entity. Our proposed approach achieves 85.56% and 84.37% cosine similarity on public benchmark HEADLINE and MICROBLOG datasets, outperforms several strong baselines and achieves new state-of-the-art performance, verifying its effectiveness. Chunli Xiang, Junchi Zhang, Fei Li 0021, Chong Teng, Donghong Ji |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Generative Biomedical Event Extraction With Constrained Decoding StrategyabstractCurrently, biomedical event extraction has received considerable attention in various fields, including natural language processing, bioinformatics, and computational biomedicine. This has led to the emergence of numerous machine learning and deep learning models that have been proposed and applied to tackle this complex task. While existing models typically adopt an extraction-based approach, which requires breaking down the extraction of biomedical events into multiple subtasks for sequential processing, making it prone to cascading errors. This paper presents a novel approach by constructing a biomedical event generation model based on the framework of the pre-trained language model T5. We employ a sequence-to-sequence generation paradigm to obtain events, the model utilizes constrained decoding algorithm to guide sequence generation, and a curriculum learning algorithm for efficient model learning. To demonstrate the effectiveness of our model, we evaluate it on two public benchmark datasets, Genia 2011 and Genia 2013. Our model achieves superior performance, illustrating the effectiveness of generative modeling of biomedical events. Fangfang Su, Chong Teng, Fei Li 0021, Bobo Li 0001, Donghong Ji |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | TKDP: Threefold Knowledge-Enriched Deep Prompt Tuning for Few-Shot Named Entity RecognitionabstractFew-shot named entity recognition (NER) exploits limited annotated instances to identify named mentions. Effectively transferring the internal or external resources thus becomes the key to few-shot NER. While the existing prompt tuning methods have shown remarkable few-shot performances, they still fail to make full use of knowledge. In this work, we investigate the integration of rich knowledge to prompt tuning for stronger few-shot NER. We propose incorporating the deep prompt tuning framework with threefold knowledge (namelyTKDP), including the internal 1)context knowledgeand the external 2)label knowledge& 3)sememe knowledge. TKDP encodes the three feature sources and incorporates them into soft prompt embeddings, which are further injected into an existing pre-trained language model to facilitate predictions. On five benchmark datasets, the performance of our knowledge-enriched model was boosted by at most 11.53% F1 over the raw deep prompt method, and it significantly outperforms 9 strong-performing baseline systems in 5-/10-/20-shot settings, showing great potential in few-shot NER. Our TKDP framework can be broadly adapted to other few-shot tasks without much effort. Jiang Liu 0018, Hao Fei 0001, Fei Li 0021, Bobo Li 0001, Liang Zhao 0001, Chong Teng, Donghong Ji |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion RecognitionabstractIt has been a hot research topic to enable machines to understand human emotions in multimodal contexts under dialogue scenarios, which is tasked with multimodal emotion analysis in conversation (MM-ERC). MM-ERC has received consistent attention in recent years, where a diverse range of methods has been proposed for securing better task performance. Most existing works treat MM-ERC as a standard multimodal classification problem and perform multimodal feature disentanglement and fusion for maximizing feature utility. Yet after revisiting the characteristic of MM-ERC, we argue that both the feature multimodality and conversational contextualization should be properly modeled simultaneously during the feature disentanglement and fusion steps. In this work, we target further pushing the task performance by taking full consideration of the above insights. On the one hand, during feature disentanglement, based on the contrastive learning technique, we devise a Dual-level Disentanglement Mechanism (DDM) to decouple the features into both the modality space and utterance space. On the other hand, during the feature fusion stage, we propose a Contribution-aware Fusion Mechanism (CFM) and a Context Refusion Mechanism (CRM) for multimodal and context integration, respectively. They together schedule the proper integrations of multimodal and context features. Specifically, CFM explicitly manages the multimodal feature contributions dynamically, while CRM flexibly coordinates the introduction of dialogue contexts. On two public MM-ERC datasets, our system achieves new state-of-the-art performance consistently. Further analyses demonstrate that all our proposed mechanisms greatly facilitate the MM-ERC task by making full use of the multimodal and context features adaptively. Note that our proposed methods have the great potential to facilitate a broader range of other conversational multimodal tasks. Bobo Li 0001, Hao Fei 0001, Lizi Liao, Yu Zhao 0043, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li 0021 |
ACM Multimedia | 5 |
| 2023 | DialogREC+: An Extension of DialogRE to Investigate How Much Coreference Helps Relation Extraction in Dialogs
Yiyun Xiong, Mengwei Dai, Fei Li 0021, Hao Fei 0001, Bobo Li 0001, Shengqiong Wu, Donghong Ji, Chong Teng |
NLPCC (1) | 8 |
| 2023 | UZNER: A Benchmark for Named Entity Recognition in Uzbek
Aizihaierjiang Yusufu, Liu Jiang, Abidan Ainiwaer, Chong Teng, Aizierguli Yusufu, Fei Li 0021, Donghong Ji |
NLPCC (1) | 4 |
| 2023 | A Bi-directional Multi-hop Inference Model for Joint Dialog Sentiment Classification and Act Recognition
Fei Li 0021, Yuyang Chai, Chong Teng, Donghong Ji |
NLPCC (1) | 4 |
| 2023 | Fake News Detection with Context Awareness of the PublisherabstractThe spread of fake news is a significant social problem that can have disastrous impacts on various domains, such as politics and the economy.Therefore, detecting fake news has become a major concern.However, prior research has relied solely on news text to derive news representation, which is inadequate because different news items under the same publisher are interconnected.To address this limitation, we propose an innovative approach called the Publisher-oriented Multi-view Graph Model (PMGM) that leverages the context awareness of the publisher to detect fake news.Our approach enriches the news representation by incorporating publisher profiles and text style features extracted from the news.Specifically, we construct a multi-view graph that encodes various relationships between news items from the same publisher, such as news topics and occasions in which they were released.Furthermore, we leverage a multi-layer Graph Convolutional Network in conjunction with jumping knowledge networks to model the multi-view graph and produce a publisher-oriented contextualized representation of news.Experimental results on two widely used fake news datasets, namely LIAR and Weibo21, demonstrate the effectiveness of our approach.Specifically, the PMGM model outperforms the state-ofthe-art methods significantly.Overall, our proposed model unifies various heterogeneous features and information related to news based on a publisher-oriented approach, thereby offering a novel idea to enhance fake news detection. Xingchen Ding, Chong Teng, Donghong Ji |
SEKE | 2 |
| 2023 | HGAPT: Heterogeneous Graph Augmented Prompt Tuning for Low-Resource Fake News DetectionabstractThe dissemination of disinformation on social media platforms has a significant impact on personal reputation and public trust.There has been a recent surge of interest in fake news detection.However, detecting low-resource fake news, particularly those pertaining to recent events that have not yet been disseminated by users and are typically in short text, remains challenging due to the lack of training data and prior knowledge.In this paper, we introduce a novel framework named the Heterogeneous Graph Augmented Prompt-based Tuning framework (HGAPT) that can leverage the metadata of news such as publisher and topic to construct a heterogeneous graph in same batch, which improve the performance of low-resource fake news detection.We have conducted extensive experiments on two low-resource fake news datasets that were collected from real-world sources.The results demonstrate that our proposed framework outperforms state-of-the-art methods, with superior detection performance at the zero-shot setting. Xingchen Ding, Chong Teng, Donghong Ji, Fei Li 0021 |
SEKE | 2 |
| 2023 | MOIT: A Novel task for mining opinions towards implicit targets
Fei Li 0021, Chong Teng, Yijiang Liu, Chunli Xiang, Donghong Ji |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Chinese Event Discourse Deixis Resolution: Design of the Dataset and ModelabstractAnaphora resolution is a traditional task in the natural language processing community, defined as a cohesion phenomenon where one entity points back to a previous entity. Event discourse deixis (EDD) is a kind of more complex anaphora in which the anaphors refer to event descriptions such as sentences or clauses. Event discourse deixis resolution (EDDR) is able to help machines understand the richer linguistic and semantic information in the discourse. However, compared to anaphora resolution, EDDR has received relatively less research attention. In this work, we investigate the EDDR task by designing the corresponding dataset and model. First, we manually construct a high-quality Chinese corpus for EDDR, including 4,417 documents and 5,929 event chains that consist of event antecedents and anaphors. Second, we propose a deep neural network model for EDDR, which formulates the task into two subtasks, namely event anaphor recognition and event antecedent recognition. Our model is trained under the two subtasks jointly so that the EDDR task can be performed end to end. Besides our final model, we also build seven pipeline and joint models as baselines to build comprehensive benchmarks for follow-up research. Experimental results on our EDDR dataset show that our model outperforms all the baselines and achieves about 53%, 44%, 53%, and 63% F1s using standard anaphora resolution metrics such as CoNLL, MUC, B3, and Ceafe. The performances show that EDDR is a challenging task and worth researching in the future. Our dataset and model will be released to facilitate follow-up research. Dongdong Xie 0003, Fei Li 0021, Bobo Li 0001, Chong Teng, Donghong Ji, Meishan Zhang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | TOE: A Grid-Tagging Discontinuous NER Model Enhanced by Embedding Tag/Word Relations and More Fine-Grained TagsabstractSo far, discontinuous named entity recognition (NER) has received increasing research attention and many related methods have surged such as hypergraph-based methods, span-based methods, and sequence-to-sequence (Seq2Seq) methods, etc. However, these methods more or less suffer from some problems such as decoding ambiguity and efficiency, which limit their performance. Recently, grid-tagging methods, which benefit from the flexible design of tagging systems and model architectures, have shown superiority to adapt for various information extraction tasks. In this paper, we follow the line of such methods and propose a competitive grid-tagging model for discontinuous NER. We call our model TOE because we incorporate two kinds of Tag-Oriented Enhancement mechanisms into a state-of-the-art (SOTA) grid-tagging model that casts the NER problem into word-word relationship prediction. First, we design a Tag Representation Embedding Module (TREM) to force our model to consider not only word-word relationships but also word-tag and tag-tag relationships. Concretely, we construct tag representations and embed them into TREM, so that TREM can treat tag and word representations as queries/keys/values and utilize self-attention to model their relationships. On the other hand, motivated by the Next-Neighboring-Word (NNW) and Tail-Head-Word (THW) tags in the SOTA model, we add two new symmetric tags, namely Previous-Neighboring-Word (PNW) and Head-Tail-Word (HTW), to model more fine-grained word-word relationships and alleviate error propagation from tag prediction. In the experiments of three benchmark datasets, namely CADEC, ShARe13 and ShARe14, our TOE model pushes the SOTA results by about 0.83%, 0.05% and 0.66% in F1, demonstrating its effectiveness. Jiang Liu 0018, Donghong Ji, Dongdong Xie 0003, Chong Teng, Liang Zhao 0001, Fei Li 0021 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Unified Named Entity Recognition as Word-Word Relation ClassificationabstractSo far, named entity recognition (NER) has been involved with three major types, including flat, overlapped (aka. nested), and discontinuous NER, which have mostly been studied individually. Recently, a growing interest has been built for unified NER, tackling the above three jobs concurrently with one single model. Current best-performing methods mainly include span-based and sequence-to-sequence models, where unfortunately the former merely focus on boundary identification and the latter may suffer from exposure bias. In this work, we present a novel alternative by modeling the unified NER as word-word relation classification, namely W^2NER. The architecture resolves the kernel bottleneck of unified NER by effectively modeling the neighboring relations between entity words with Next-Neighboring-Word (NNW) and Tail-Head-Word-* (THW-*) relations. Based on the W^2NER scheme we develop a neural framework, in which the unified NER is modeled as a 2D grid of word pairs. We then propose multi-granularity 2D convolutions for better refining the grid representations. Finally, a co-predictor is used to sufficiently reason the word-word relations. We perform extensive experiments on 14 widely-used benchmark datasets for flat, overlapped, and discontinuous NER (8 English and 6 Chinese datasets), where our model beats all the current top-performing baselines, pushing the state-of-the-art performances of unified NER. Hao Fei 0001, Jiang Liu 0018, Shengqiong Wu, Meishan Zhang, Chong Teng, Donghong Ji, Fei Li 0021 |
AAAI | 6 |
| 2022 | Mastering the Explicit Opinion-Role Interaction: Syntax-Aided Neural Transition System for Unified Opinion Role LabelingabstractUnified opinion role labeling (ORL) aims to detect all possible opinion structures of 'opinion-holder-target' in one shot, given a text. The existing transition-based unified method, unfortunately, is subject to longer opinion terms and fails to solve the term overlap issue. Current top performance has been achieved by employing the span-based graph model, which however still suffers from both high model complexity and insufficient interaction among opinions and roles. In this work, we investigate a novel solution by revisiting the transition architecture, and augmenting it with a pointer network (PointNet). The framework parses out all opinion structures in linear-time complexity, meanwhile breaks through the limitation of any length of terms with PointNet. To achieve the explicit opinion-role interactions, we further propose a unified dependency-opinion graph (UDOG), co-modeling the syntactic dependency structure and the partial opinion-role structure. We then devise a relation-centered graph aggregator (RCGA) to encode the multi-relational UDOG, where the resulting high-order representations are used to promote the predictions in the vanilla transition system. Our model achieves new state-of-the-art results on the MPQA benchmark. Analyses further demonstrate the superiority of our methods on both efficacy and efficiency. Shengqiong Wu, Hao Fei 0001, Fei Li 0021, Meishan Zhang, Yijiang Liu, Chong Teng, Donghong Ji |
AAAI | 6 |
| 2022 | Prompt-Based Generative Multi-label Emotion Prediction with Label Contrastive Learning
Yuyang Chai, Chong Teng, Hao Fei 0001, Shengqiong Wu, Donghong Ji, Fei Li 0021 |
NLPCC (1) | 2 |
| 2021 | Neural transition model for aspect-based sentiment triplet extraction with triplet memory
Shengqiong Wu, Bobo Li 0001, Dongdong Xie 0003, Chong Teng, Donghong Ji |
Neurocomputing | 4 |
| 2020 | A deep neural network model for speakers coreference resolution in legal texts
Donghong Ji, Hao Fei 0001, Chong Teng, Yafeng Ren |
Inf. Process. Manag. | 4 |
| 2014 | Word Sense Induction Using Lexical Chain based Hypergraph Model
Tao Qian 0002, Donghong Ji, Mingyao Zhang, Chong Teng, Congling Xia |
COLING | 4 |
| 2012 | Context-Enhanced Personalized Social Summarization
Po Hu 0001, Donghong Ji, Chong Teng, Yujing Guo |
COLING | 3 |
| 2011 | Social Summarization via Automatically Discovered Social Context
Po Hu 0001, Cheng Sun 0002, Longfei Wu, Donghong Ji, Chong Teng |
IJCNLP | 5 |
| 2009 | Query-Focused Multi-Document Summarization Using Co-Training Based Semi-Supervised Learning
Po Hu 0001, Donghong Ji, Chong Teng |
PACLIC | 4 |
| 2009 | Finding Answers to Definition Questions Using Web Knowledge Bases
Donghong Ji, Chong Teng |
PACLIC | 4 |
| 2006 | A Hybrid Sentence Ordering Strategy in Multi-document Summarization
Yanxiang He, Dexi Liu, Donghong Ji, Chong Teng, Wenqing Qi |
WISE | 5 |