Fenghuan Li

dblp:07/10130 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0003-3640-4253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Ambiguity-aware Multi-level Incongruity Fusion Network for Multi-Modal Sarcasm Detection
abstract
Multi-modal sarcasm detection aims to identify whether a given image-text pair is sarcastic. The pivotal factor of the task lies in accurately capturing incongruities from different modalities. Although existing studies have achieved impressive success, they primarily committed to fusing the textual and visual information to establish cross-modal correlations, overlooking the significance of original unimodal incongruity information at the text-level and image-level. Furthermore, the utilized fusion strategies of cross-modal information neglected the effect of inherent ambiguity within text and image modalities on multimodal fusion. To overcome these limitations, we propose a novel Ambiguity-aware Multi-level Incongruity Fusion Network (AMIF) for multi-modal sarcasm detection. Our method involves a multi-level incongruity learning module to capture the incongruity information simultaneously at the text-level, image-level and cross-modal-level. Additionally, an ambiguity-based fusion module is developed to dynamically learn reasonable weights and interpretably aggregate incongruity features from different levels. Comprehensive experiments conducted on a publicly available dataset demonstrate the superiority of our proposed model over state-of-the-art methods.
Kuntao Li, Qiaofeng Wu, Weixing Mai, Fenghuan Li, Yun Xue 0002
COLING5
2025 SACR: Self-training with Saliency-Augmented Consistency Regularization for Few-Shot Learners
abstract
Pre-trained language models have made significant strides in natural language processing tasks, enabling flexible fine-tuning for downstream applications. However, in few-shot learning scenarios, pre-trained models face challenges related to overfitting due to limited training samples, which hinders their ability to capture data diversity and robustly handle input variations. To overcome these limitations, we propose the Saliency-Augmented Consistency Regularization (SACR) framework, a novel self-training strategy designed to improve few-shot learning performance and robustness. SACR consists of three key components: (1) Saliency-guided Data Perturbation, which uses saliency analysis to identify and perturb words that significantly influence model predictions, generating semantically consistent pseudo-samples; (2) Semantic-equivalent Sample Mining, which employs K-means clustering to select semantically similar samples and prevent semantic shift; and (3) Consistency Training, where regularization ensures consistency between the semantics and prediction distributions of the original and perturbed samples. We conducted extensive experiments on 15 public datasets to evaluate our approach. SACR significantly outperforms strong baseline models and demonstrates superior generalization capabilities compared to state-of-the-art methods, achieving an average 1.0% improvement in classification accuracy.
Yanyan Feng, Yue Zhou 0012, Yun Xue 0002, Fenghuan Li, Zehong Lin
ICASSP4
2025 ReHyGen: Relational hypergraph enhanced generative aspect sentiment triplet extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE) has emerged as a pivotal task in sentiment analysis , focusing on extracting the aspect terms along with the corresponding opinion terms and the expressed sentiments. Recently, generative models have achieved significant success in ASTE task. However, existing generative approaches fail to further model the specific relations within the context for ASTE at the encoding phase, making it difficult to establish the nuanced connections between aspect and opinion terms. Additionally, these approaches rely on simple structured templates at the decoding phase to pair aspect terms with opinion terms, which fails to provide effective relation information for the decoding process. To address the aforementioned issues, we propose ReHyGen, a novel relational hypergraph enhanced framework designed to enhance the relational modeling capabilities of generative ASTE models during both the encoding and decoding phases. Specifically, ReHyGen comprises two core components: the Relational Hypergraph Enhanced Module (RHEM) and the Relational Prompt Module (RPM). RHEM leverages the hypergraph attention network and auxiliary relation classification to capture high-order word interactions and boundary-sensitive word pair relations. RPM incorporates relational information into the decoding phase by providing relation-aware prompts, guiding the generation of more accurate target sequences. Extensive experiments on benchmark datasets demonstrate that our proposed framework significantly improve the performance of generative ASTE models.
Zehong Lin, Weibo Chen, Yun Xue 0002, Fenghuan Li
Neurocomputing4
2025 Dual-level adaptive incongruity-enhanced model for multimodal sarcasm detection
Qiaofeng Wu, Wenlong Fang, Weiyu Zhong, Fenghuan Li, Yun Xue 0002, Bo Chen 0004
Neurocomputing4
2025 Semantic enhanced bi-syntactic graph convolutional network for aspect-based sentiment analysis
Junyang Xiao, Yun Xue 0002, Fenghuan Li
Inf. Sci.3
2025 Heterogeneous subgraph network with prompt learning for interpretable depression detection on social media
Chen Chen 0135, Fenghuan Li, Haopeng Chen, Yuankun Lin
Knowl. Based Syst.2
2024 Semantics-Aware Dual Graph Convolutional Networks for Argument Pair Extraction
abstract
Argument pair extraction (APE) is a task that aims to extract interactive argument pairs from two argument passages. Generally, existing works focus on either simple argument interaction or task form conversion, instead of thorough deep-level feature exploitation of argument pairs. To address this issue, a Semantics-Aware Dual Graph Convolutional Networks (SADGCN) is proposed for APE. Specifically, the co-occurring word graph is designed to tackle the lexical and semantic relevance of arguments with a pre-trained Rouge-guided Transformer (ROT). Considering the topic relevance in argument pairs, a topic graph is constructed by the neural topic model to leverage the topic information of argument passages. The two graphs are fused via a gating mechanism, which contributes to the extraction of argument pairs. Experimental results indicate that our approach achieves the state-of-the-art performance. The performance on F1 score is significantly improved by 6.56% against the existing best alternative.
Minzhao Guan, Zhixun Qiu, Fenghuan Li, Yun Xue 0002
LREC/COLING3
2024 Representation and Granularity Joint Alignment Framework for Multimodal Sarcasm Detection on Social Media
Jiapei Hu, Yun Xue 0002, Fenghuan Li, Qianhua Cai
DASFAA (7)4
2024 D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment Detection
abstract
Multimodal sentiment detection aims to classify the sentiment polarity of a given imagetext pair.Existing approaches apply the same fixed framework to all input samples, lacking the flexibility to adapt to different image-text pairs.Furthermore, the interaction patterns of these methods are overly homogenized, limiting the model's capacity to extract multimodal sentiment information effectively.In this paper, we develop a Dual-Branch Dynamic Routing Network (D 2 R), which is the first multimodal dynamic interaction model towards multimodal sentiment detection.Specifically, we design six independent units to simulate inter-and intramodal information interactions without depending on any existing fixed frameworks.Additionally, we configure a soft router in each unit to guide path generation and introduce the path regularization term to optimize these inference paths.Comprehensive experiments on three publicly available datasets demonstrate the superiority of our proposed model over state-ofthe-art methods.
Kuntao Li, Weixing Mai, Qiaofeng Wu, Yun Xue 0002, Fenghuan Li
EMNLP6
2024 Prompt-enhanced Network for Hateful Meme Classification
Junxi Liu, Yanyan Feng, Jiehai Chen, Yun Xue 0002, Fenghuan Li
IJCAI5
2024 Enhanced Syntactic and Semantic Graph Convolutional Network With Contrastive Learning for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity of a given specific aspect in the sentence. Recent studies focus on leveraging graph convolutional neural networks to encode both syntactic and semantic information. However, current syntactic parsers, which are not specifically for ABSA, introduce noise to the syntactic information. Besides, ongoing studies ignore the distinctiveness of semantics and syntax. To address these issues, we proposed an enhanced syntactic and semantic graph convolutional network (GCN) with contrastive learning in this article. An aspect-oriented syntactic graph is constructed with aspect-specific perturbed masking for reducing the syntactic noise, and a semantic graph is established with self-attention weights from bidirectional encoder representation from transformers (BERT). The semantic and syntactic representations are further enhanced by both sentiment polarity-based supervised contrastive learning and syntactic reliability-based unsupervised contrastive learning. Furthermore, label embeddings of syntactic reliability are learned to determine the weights of syntactic and semantic information. Extensive experiments on four publicly available datasets demonstrate that our model is more competitive than the state-of-the-arts.
Minzhao Guan, Fenghuan Li, Yun Xue 0002
IEEE Trans. Comput. Soc. Syst.2
2024 Dynamic Graph Construction Framework for Multimodal Named Entity Recognition in Social Media
abstract
Multimodal named entity recognition (MNER) aims to detect named entities and identify the entity types based on texts and attached images, which also generates inputs for other comprehensive tasks, such as multimodal machine translation, visual dialog, and multimodal sentiment analysis. Existing studies have limitations in text-image matching and multimodal semantic disparity reduction. For one thing, current methods fail to resolve both overall and local text-image matching issues in a self-guided way. For another, the static graphs constructed in MNER models are challenging in bridging the semantic gap between different modalities. In this work, a dynamic graph construction framework (DGCF) is proposed to solve the above-mentioned limitations. A similarity vector-based text-image matching inferring strategy is designed to obtain the overall and local matching relation between text and image while the overall matching determines the retained proportion of visual information. Then, a multimodal dynamic graph interaction module is developed. Within each layer of the module, the local matching relations and part of speech (POS)-based multihead attention are integrated to construct a dynamic cross-modal graph and a semantic graph. Lastly, a CRF layer is used to predict entity label. Extensive experiments are performed on two benchmark datasets. The experimental results reveal that our model is a competitive alternative and achieves state-of-the-art performance.
Weixing Mai, Zhengxuan Zhang, Kuntao Li, Yun Xue 0002, Fenghuan Li
IEEE Trans. Comput. Soc. Syst.5
2023 Span-Based Pair-Wise Aspect and Opinion Term Joint Extraction with Contrastive Learning
Jinjie Yang, Feipeng Dai, Fenghuan Li, Yun Xue 0002
NLPCC (2)3
2023 RBA-GCN: Relational Bilevel Aggregation Graph Convolutional Network for Emotion Recognition
abstract
Emotion recognition in conversation (ERC) has received increasing attention from researchers due to its wide range of applications. As conversation has a natural graph structure, numerous approaches used to model ERC based on graph convolutional networks (GCNs) have yielded significant results. However, the aggregation approach of traditional GCNs suffers from the node information redundancy problem, leading to node discriminant information loss. Additionally, single-layer GCNs lack the capacity to capture long-range contextual information from the graph. Furthermore, the majority of approaches are based on textual modality or stitching together different modalities, resulting in a weak ability to capture interactions between modalities. To address these problems, we present the relational bilevel aggregation graph convolutional network (RBA-GCN), which consists of three modules: the graph generation module (GGM), similarity-based cluster building module (SCBM) and bilevel aggregation module (BiAM). First, GGM constructs a novel graph to reduce the redundancy of target node information. Then, SCBM calculates the node similarity in the target node and its structural neighborhood, where noisy information with low similarity is filtered out to preserve the discriminant information of the node. Meanwhile, BiAM is a novel aggregation method that can preserve the information of nodes during the aggregation process. This module can construct the interaction between different modalities and capture long-range contextual information based on similarity clusters. On both the IEMOCAP and MELD datasets, the weighted average F1 score of RBA-GCN has a 2.17$\sim$5.21% improvement over that of the most advanced method.
Guoheng Huang, Fenghuan Li, Xiaochen Yuan, Chi-Man Pun, Guo Zhong
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Self-adaptive statistical process control for anomaly detection in time series
Dequan Zheng, Fenghuan Li, Tiejun Zhao
Expert Syst. Appl.2
2016 Corrigendum to "Self-adaptive statistical process control for anomaly detection in time series" [Expert Systems With Applications 57 (2016) 324-336]
Dequan Zheng, Fenghuan Li, Tiejun Zhao
Expert Syst. Appl.2