Yun Xue 0002

dblp:39/4919-2 · DBLP profile ↗
← Back
35ranked-venue papers
0as first author
30since 2021 · last 2026
0000-0002-4048-5298ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 DG-MCTS: Dual-Guidance Monte Carlo Tree Search for Adaptive Emotional Support Dialogue Planning
Benshuo Lin, Yuelei Li, Yun Xue 0002, Yiping Song
WWW5
2026 Multi-view dynamic perception framework for Chinese harmful meme detection
Jiapei Hu, Kuntao Li, Yun Xue 0002, Jinghua Liang
Inf. Process. Manag.5
2025 Advancing Collaborative Debates with Role Differentiation through Multi-Agent Reinforcement Learning
abstract
Multi-agent collaborative tasks exhibit exceptional capabilities in natural language applications and generation. By prompting agents to assign clear roles, it is possible to facilitate cooperation and achieve complementary capabilities among LLMs. A common strategy involves adopting a relatively general role assignment mechanism, such as introducing a “judge” or a “summarizer”. However, these approaches lack task-specific role customization based on task characteristics. Another strategy involves decomposing the task based on domain knowledge and task characteristics, followed by assigning appropriate roles according to LLMs’ respective strengths, such as programmers and testers. However, in some given tasks, obtaining domain knowledge related to task characteristics and getting the strengths of different LLMs is hard. To solve these problems, we propose a Multi-LLM Cooperation (MLC) framework with automatic role assignment capabilities. The core idea of the MLC is to initialize role assignments randomly and then allow the role embeddings to be learned jointly with the downstream task. To capture the state transitions of multiple LLMs during turn-based speaking, the role embedding is sequence-aware. At the same time, to avoid role convergence, the role differentiation module in MLC encourages behavioral differentiation between LLMs while ensuring the LLM team consistency, guiding different LLMs to develop complementary strengths from the optimization level. Our experiments on seven datasets demonstrate that MLC significantly enhances collaboration and expertise, which collaboratively addresses multi-agent tasks.
Ziyi Su, Yun Xue 0002, Zhiliang Tian, Yiping Song, Minlie Huang
ACL (1)3
2025 MSG-LLM: A Multi-scale Interactive Framework for Graph-enhanced Large Language Models
abstract
Graph-enhanced large language models (LLMs) leverage LLMs’ remarkable ability to model language and use graph structures to capture topological relationships. Existing graph-enhanced LLMs typically retrieve similar subgraphs to augment LLMs, where the subgraphs carry the entities related to our target and relations among the entities. However, the retrieving methods mainly focus solely on accurately matching subgraphs between our target subgraph and the candidate subgraphs at the same scale, neglecting that the subgraphs with different scales may also share similar semantics or structures. To tackle this challenge, we introduce a graph-enhanced LLM with multi-scale retrieval (MSG-LLM). It captures similar graph structures and semantics across graphs at different scales and bridges the graph alignment across multiple scales. The larger scales maintain the graph’s global information, while the smaller scales preserve the details of fine-grained sub-structures. Specifically, we construct a multi-scale variation to dynamically shrink the scale of graphs. Further, we employ a graph kernel search to discover subgraphs from the entire graph, which essentially achieves multi-scale graph retrieval in Hilbert space. Additionally, we propose to conduct multi-scale interactions (message passing) over graphs at various scales to integrate key information. The interaction also bridges the graph and LLMs, helping with graph retrieval and LLM generation. Finally, we employ a Chain-of-Thought-based LLM prediction to perform the downstream tasks. We evaluate our approach on two graph-based downstream tasks and the experimental results show that our method achieves state-of-the-art performance.
Zhangkai Zheng, Benshuo Lin, Yun Xue 0002, Yiping Song
COLING4
2025 Ambiguity-aware Multi-level Incongruity Fusion Network for Multi-Modal Sarcasm Detection
abstract
Multi-modal sarcasm detection aims to identify whether a given image-text pair is sarcastic. The pivotal factor of the task lies in accurately capturing incongruities from different modalities. Although existing studies have achieved impressive success, they primarily committed to fusing the textual and visual information to establish cross-modal correlations, overlooking the significance of original unimodal incongruity information at the text-level and image-level. Furthermore, the utilized fusion strategies of cross-modal information neglected the effect of inherent ambiguity within text and image modalities on multimodal fusion. To overcome these limitations, we propose a novel Ambiguity-aware Multi-level Incongruity Fusion Network (AMIF) for multi-modal sarcasm detection. Our method involves a multi-level incongruity learning module to capture the incongruity information simultaneously at the text-level, image-level and cross-modal-level. Additionally, an ambiguity-based fusion module is developed to dynamically learn reasonable weights and interpretably aggregate incongruity features from different levels. Comprehensive experiments conducted on a publicly available dataset demonstrate the superiority of our proposed model over state-of-the-art methods.
Kuntao Li, Qiaofeng Wu, Weixing Mai, Fenghuan Li, Yun Xue 0002
COLING6
2025 Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question Answering
abstract
The knowledge-based visual question answering (KB-VQA) task involves using external knowledge about the image to assist reasoning. Building on the impressive performance of multimodal large language model (MLLM), recent methods have commenced leveraging MLLM as an implicit knowledge base for reasoning. However, the direct employment of MLLM with raw external knowledge might result in reasoning errors due to misdirected knowledge information. Additionally, MLLM may lack fine-grained perception of visual features, which can result in hallucinations during reasoning. To address these challenges, we propose Notes-guided MLLM Reasoning (NoteMR), a novel framework that guides MLLM in better reasoning by utilizing knowledge notes and visual notes. Specifically, we initially obtain explicit knowledge from an external knowledge base. Then, this explicit knowledge, combined with images, is used to assist the MLLM in generating knowledge notes. These notes are designed to filter explicit knowledge and identify relevant internal implicit knowledge within the MLLM. We then identify highly correlated regions between the images and knowledge notes, retaining them as image notes to enhance the model’s fine-grained perception, thereby mitigating MLLM induced hallucinations. Finally, both notes are fed into the MLLM, enabling a more comprehensive understanding of the image-question pair and enhancing the model’s reasoning capabilities. Our method achieves state-of-the-art performance on the OK-VQA and A-OKVQA datasets, demonstrating its robustness and effectiveness across diverse VQA scenarios.
Wenlong Fang, Qiaofeng Wu, Yun Xue 0002
CVPR4
2025 SACR: Self-training with Saliency-Augmented Consistency Regularization for Few-Shot Learners
abstract
Pre-trained language models have made significant strides in natural language processing tasks, enabling flexible fine-tuning for downstream applications. However, in few-shot learning scenarios, pre-trained models face challenges related to overfitting due to limited training samples, which hinders their ability to capture data diversity and robustly handle input variations. To overcome these limitations, we propose the Saliency-Augmented Consistency Regularization (SACR) framework, a novel self-training strategy designed to improve few-shot learning performance and robustness. SACR consists of three key components: (1) Saliency-guided Data Perturbation, which uses saliency analysis to identify and perturb words that significantly influence model predictions, generating semantically consistent pseudo-samples; (2) Semantic-equivalent Sample Mining, which employs K-means clustering to select semantically similar samples and prevent semantic shift; and (3) Consistency Training, where regularization ensures consistency between the semantics and prediction distributions of the original and perturbed samples. We conducted extensive experiments on 15 public datasets to evaluate our approach. SACR significantly outperforms strong baseline models and demonstrates superior generalization capabilities compared to state-of-the-art methods, achieving an average 1.0% improvement in classification accuracy.
Yanyan Feng, Yue Zhou 0012, Yun Xue 0002, Fenghuan Li, Zehong Lin
ICASSP3
2025 ReHyGen: Relational hypergraph enhanced generative aspect sentiment triplet extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE) has emerged as a pivotal task in sentiment analysis , focusing on extracting the aspect terms along with the corresponding opinion terms and the expressed sentiments. Recently, generative models have achieved significant success in ASTE task. However, existing generative approaches fail to further model the specific relations within the context for ASTE at the encoding phase, making it difficult to establish the nuanced connections between aspect and opinion terms. Additionally, these approaches rely on simple structured templates at the decoding phase to pair aspect terms with opinion terms, which fails to provide effective relation information for the decoding process. To address the aforementioned issues, we propose ReHyGen, a novel relational hypergraph enhanced framework designed to enhance the relational modeling capabilities of generative ASTE models during both the encoding and decoding phases. Specifically, ReHyGen comprises two core components: the Relational Hypergraph Enhanced Module (RHEM) and the Relational Prompt Module (RPM). RHEM leverages the hypergraph attention network and auxiliary relation classification to capture high-order word interactions and boundary-sensitive word pair relations. RPM incorporates relational information into the decoding phase by providing relation-aware prompts, guiding the generation of more accurate target sequences. Extensive experiments on benchmark datasets demonstrate that our proposed framework significantly improve the performance of generative ASTE models.
Zehong Lin, Weibo Chen, Yun Xue 0002, Fenghuan Li
Neurocomputing3
2025 Dual-level adaptive incongruity-enhanced model for multimodal sarcasm detection
Qiaofeng Wu, Wenlong Fang, Weiyu Zhong, Fenghuan Li, Yun Xue 0002, Bo Chen 0004
Neurocomputing5
2025 Knowledge based attribute completion for heterogeneous graph node classification
Zhangkai Zheng, Yun Xue 0002, Yiping Song, Zhuoming Liang
Neurocomputing3
2025 Semantic enhanced bi-syntactic graph convolutional network for aspect-based sentiment analysis
Junyang Xiao, Yun Xue 0002, Fenghuan Li
Inf. Sci.2
2025 VS-MRC: A visual semantics-guided machine reading comprehension framework for multimodal named entity recognition with multiple images
Jiapei Hu, Yun Xue 0002
Knowl. Based Syst.3
2024 Semantics-Aware Dual Graph Convolutional Networks for Argument Pair Extraction
abstract
Argument pair extraction (APE) is a task that aims to extract interactive argument pairs from two argument passages. Generally, existing works focus on either simple argument interaction or task form conversion, instead of thorough deep-level feature exploitation of argument pairs. To address this issue, a Semantics-Aware Dual Graph Convolutional Networks (SADGCN) is proposed for APE. Specifically, the co-occurring word graph is designed to tackle the lexical and semantic relevance of arguments with a pre-trained Rouge-guided Transformer (ROT). Considering the topic relevance in argument pairs, a topic graph is constructed by the neural topic model to leverage the topic information of argument passages. The two graphs are fused via a gating mechanism, which contributes to the extraction of argument pairs. Experimental results indicate that our approach achieves the state-of-the-art performance. The performance on F1 score is significantly improved by 6.56% against the existing best alternative.
Minzhao Guan, Zhixun Qiu, Fenghuan Li, Yun Xue 0002
LREC/COLING4
2024 Representation and Granularity Joint Alignment Framework for Multimodal Sarcasm Detection on Social Media
Jiapei Hu, Yun Xue 0002, Fenghuan Li, Qianhua Cai
DASFAA (7)3
2024 D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment Detection
abstract
Multimodal sentiment detection aims to classify the sentiment polarity of a given imagetext pair.Existing approaches apply the same fixed framework to all input samples, lacking the flexibility to adapt to different image-text pairs.Furthermore, the interaction patterns of these methods are overly homogenized, limiting the model's capacity to extract multimodal sentiment information effectively.In this paper, we develop a Dual-Branch Dynamic Routing Network (D 2 R), which is the first multimodal dynamic interaction model towards multimodal sentiment detection.Specifically, we design six independent units to simulate inter-and intramodal information interactions without depending on any existing fixed frameworks.Additionally, we configure a soft router in each unit to guide path generation and introduce the path regularization term to optimize these inference paths.Comprehensive experiments on three publicly available datasets demonstrate the superiority of our proposed model over state-ofthe-art methods.
Kuntao Li, Weixing Mai, Qiaofeng Wu, Yun Xue 0002, Fenghuan Li
EMNLP5
2024 Prompt-enhanced Network for Hateful Meme Classification
Junxi Liu, Yanyan Feng, Jiehai Chen, Yun Xue 0002, Fenghuan Li
IJCAI4
2024 GeDa: Improving training data with large language models for Aspect Sentiment Triplet Extraction
Weixing Mai, Zhengxuan Zhang, Kuntao Li, Yun Xue 0002
Knowl. Based Syst.5
2024 Enhanced Syntactic and Semantic Graph Convolutional Network With Contrastive Learning for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity of a given specific aspect in the sentence. Recent studies focus on leveraging graph convolutional neural networks to encode both syntactic and semantic information. However, current syntactic parsers, which are not specifically for ABSA, introduce noise to the syntactic information. Besides, ongoing studies ignore the distinctiveness of semantics and syntax. To address these issues, we proposed an enhanced syntactic and semantic graph convolutional network (GCN) with contrastive learning in this article. An aspect-oriented syntactic graph is constructed with aspect-specific perturbed masking for reducing the syntactic noise, and a semantic graph is established with self-attention weights from bidirectional encoder representation from transformers (BERT). The semantic and syntactic representations are further enhanced by both sentiment polarity-based supervised contrastive learning and syntactic reliability-based unsupervised contrastive learning. Furthermore, label embeddings of syntactic reliability are learned to determine the weights of syntactic and semantic information. Extensive experiments on four publicly available datasets demonstrate that our model is more competitive than the state-of-the-arts.
Minzhao Guan, Fenghuan Li, Yun Xue 0002
IEEE Trans. Comput. Soc. Syst.3
2024 End-to-End Visual Grounding Framework for Multimodal NER in Social Media Posts
abstract
Multimodal named entity recognition (MNER) for social media aims to detect named entities in user-generated posts with the aid of visual information from attached images. Existing methods use pretrained visual models or visual grounding (VG) toolkits to learn visual information. However, they still suffer from the mismatch issue, where the visual features extracted from visual encoder are inconsistent with actual requirements for cross-modal interaction. In an ideal scenario, the visual encoder should actively extract visual information guided by the text, which inherently provides the blueprint of desired visual features. In this article, we present an end-to-end VG framework for MNER task (VG-MNER), which adaptively learns the text-related visual features. Specifically, we introduce a backbone network with a feature fusion module to learn and aggregate multisize visual representations. We then develop a text-related visual attention to refine the visual features. Notably, entity-image contrast loss is designed to guide the training of visual encoder. The proposed model outperforms several state-of-the-art methods, achieving F1 scores of 75.62% and 88.11% on two benchmark datasets. Experimental results reveal the effectiveness of leveraging text-related visual information in the MNER task.
Jiapei Hu, Yun Xue 0002, Qianhua Cai
IEEE Trans. Comput. Soc. Syst.3
2024 Dynamic Graph Construction Framework for Multimodal Named Entity Recognition in Social Media
abstract
Multimodal named entity recognition (MNER) aims to detect named entities and identify the entity types based on texts and attached images, which also generates inputs for other comprehensive tasks, such as multimodal machine translation, visual dialog, and multimodal sentiment analysis. Existing studies have limitations in text-image matching and multimodal semantic disparity reduction. For one thing, current methods fail to resolve both overall and local text-image matching issues in a self-guided way. For another, the static graphs constructed in MNER models are challenging in bridging the semantic gap between different modalities. In this work, a dynamic graph construction framework (DGCF) is proposed to solve the above-mentioned limitations. A similarity vector-based text-image matching inferring strategy is designed to obtain the overall and local matching relation between text and image while the overall matching determines the retained proportion of visual information. Then, a multimodal dynamic graph interaction module is developed. Within each layer of the module, the local matching relations and part of speech (POS)-based multihead attention are integrated to construct a dynamic cross-modal graph and a semantic graph. Lastly, a CRF layer is used to predict entity label. Extensive experiments are performed on two benchmark datasets. The experimental results reveal that our model is a competitive alternative and achieves state-of-the-art performance.
Weixing Mai, Zhengxuan Zhang, Kuntao Li, Yun Xue 0002, Fenghuan Li
IEEE Trans. Comput. Soc. Syst.4
2024 Modeling Inter-Aspect Relations With Clause and Contrastive Learning for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task that aims to identify the sentiment polarity of the given aspect. Recent studies fail to establish the relation among multiple aspects in one sentence. To address this issue, a clause-level relational graph attention network with contrastive learning (CLRCL) model is proposed. Specifically, the given sentence is segmented into clauses to obtain the relation between two aspects based on clause-level interaction. Then, to integrate multiple-aspect information, a clause-level relational graph which contains all aspects and inter-aspect relations is developed. Notably, to precisely learn the inter-aspect relations, the supervised contrastive learning strategy is used. Experimental results reveal that the proposed model is a competitive alternative compared with the state-of-the-art methods.
Zhixun Qiu, Kehai Chen, Yun Xue 0002, Zhengxuan Zhang
IEEE Trans. Comput. Soc. Syst.3
2023 A Token-wise Graph-based Framework for Multimodal Named Entity Recognition
abstract
Multimodal Named Entity Recognition (MNER) on social media posts is a leading but challenging task. However, most existing MNER methods fail to effectively exploit the visual information from the image. Besides, the multimodal interaction and alignment remains unsettled. In this paper, we propose a novel token-wise graph-based framework to deal with the MNER task. Specifically, a token-wise image processing manner is established. A muti-modal graph is constructed based on the textual token derived from BERT and the visual token derived from SwinT. Then, the muti-modal graph is fed into a multi-layer Transformer-based module for intra- and inter-modal information fusion. In addition, multiple contrastive learning is devised to perform the global and local alignment between textual and visual nodes. Experimental results on two benchmark multimodal datasets indicate that our model achieves state-of-the-art performance in MNER tasks.
Zhengxuan Zhang, Weixing Mai, Haoliang Xiong, Chuhan Wu, Yun Xue 0002
ICME5
2023 Span-Based Pair-Wise Aspect and Opinion Term Joint Extraction with Contrastive Learning
Jinjie Yang, Feipeng Dai, Fenghuan Li, Yun Xue 0002
NLPCC (2)4
2022 Exploring fine-grained syntactic information for aspect-based sentiment classification with dual graph neural networks
Luwei Xiao, Yun Xue 0002, Hua Wang 0002, Donghong Gu, Yongsheng Zhu
Neurocomputing2
2022 Aspect-Level Sentiment Analysis with Local Semantic and Global Syntactic Features Integration
abstract
Aspect-level sentiment analysis aims to predict the sentiment polarity toward a specific aspect in a sentence. Most current approaches are based on deep learning and the attention mechanism. However, these models cannot simultaneously include the context semantic information carried by local words and the global syntactic information possibly carried by remote words. In this paper, we propose a local semantic and global syntactic integration scheme, which employs a local focus mechanism over local context words and exploits improved graph convolutional networks over dependency tree to encode global syntactic information. Moreover, multi-head attention is used to capture both the semantic information and also the interactive information between semantics and syntactic features. Experimental results on five datasets show the effectiveness of our model over a series of latest models.
Luwei Xiao, Yue-Cai Huang, Yun Xue 0002, Haoliang Zhao
Int. J. Pattern Recognit. Artif. Intell.4
2022 Multi-head self-attention based gated graph convolutional networks for aspect-based sentiment classification
Luwei Xiao, Yun Xue 0002, Bingliang Chen, Donghong Gu, Bixia Tang
Multim. Tools Appl.4
2021 Aspect-Based Sentiment Analysis Using Graph Convolutional Networks and Co-attention Mechanism
Zhaowei Chen, Yun Xue 0002, Luwei Xiao, Hao Lan Zhang 0001
ICONIP (6)2
2021 Bilateral-brain-like Semantic and Syntactic Cognitive Network for Aspect-level Sentiment Analysis
abstract
Aspect-level sentiment analysis (ALSA) is a fine-grained task for classifying the sentiment polarity of a specific aspect in a sentence. In spite of the progress in deep-learning algorithms within natural language processing domain, the methods in line with human cognition are absent. Inspired by the processing principle of our brain, we propose a Bilateral-brain-like Semantic and Syntactic Cognitive Network (BSSCN) for aspect-level sentiment analysis. There are four major modules established in BSSCN, which are left hemisphere semantic activation (LH-SA), semantic selection (SS) module, semantic integration (SI) module and right hemisphere semantic activation (RH-SA). Experimental results on a variety of datasets show that our model stably outperforms the widely-used methods, which establishes a strong evidence of the effectiveness in ALSA tasks.
Zuohua Huang, Yun Xue 0002
IJCNN3
2021 A Novel Bi-Branch Graph Convolutional Neural Network for Aspect Level Sentiment Classification
abstract
Aspect-level sentiment classification is a fine-grained task in sentiment analysis whose main purpose is to identify the sentiment polarity of a specific aspect. Current Graph Convolutional Network (GCN) has its distinctive superiority in tackling sentiment classification both semantically and syntactically. However, GCN still has deficiencies in introducing the noise during processing and dealing with sentences of complex structure. To address these issues, we propose a novel Bi-branch GCN (Bi-B GCN). In our model, an attention weight graph, by employing the attention mechanism, is constructed to substitute the basic syntax dependency tree and thus to remove the irrelevant information. Furthermore, a semantic dependency graph is devised to supplement the semantic information to the syntax dependency tree, based on which the connection between different words can be captured. In addition, on the task of sentiment classification, the integration of semantic information and the syntactic information is conducted by using a combinational gated mechanism. Substantial experiments to validate the working performance of Bi-B GCN are performed on a variety of datasets. The encouraging results establish a strong evidence of the high accuracy of the proposed model.
Bingliang Chen, Guojun Lu, Yun Xue 0002, Qianhua Cai
IJCNN3
2021 SIntactical Distance Attention Guided Graph Convolutional Network for aspect-based sentiment analIsis
abstract
Aspect-based sentiment analIsis (ABSA) aims to detect the sentiment polaritI of a specific aspect in an opinionated sentence. Current work focuses on exploiting the sIntactic tree to shorten the distance between the aspect term and context words. However, the “hard-pruning” strategI on the sIntactic tree maI lead to the reduction of importa nt sIntactic information. In this paper, we propose a novel sInt actical distance attention guided graph convolutional network (SDGCN) for ABSA. Our model is capable of fullI exploiting the sIntactic knowledge with a “soft pruning” strategI and learning crucial fine-grain sIntactic distance info rmation. AdditionallI, an effective denselI connected graph convolutional laIer is applied to avoid the over-sm oothing problem of standard GCN. Experiments conducted on three benchmark datasets show that our model achieves promising results comparing to the baseline models.
Luwei Xiao, Donghong Gu, Yun Xue 0002, Yongsheng Zhu
IJCNN3
2019 Co-attention Networks for Aspect-Level Sentiment Analysis
Haihui Li, Yun Xue 0002, Hongya Zhao, Sancheng Peng
NLPCC (2)2
2019 Shared-Private LSTM for Multi-domain Text Classification
Haiming Wu, Yue Zhang 0004, Yun Xue 0002, Ziwen Wang 0004
NLPCC (2)4
2019 A novel feature extraction methodology for sentiment analysis of product reviews
Xin Chen 0046, Yun Xue 0002, Hongya Zhao
Neural Comput. Appl.2
2016 Classification of Epileptic EEG Signals with Stacked Sparse Autoencoder Based on Deep Learning
Qin Lin 0002, Shuqun Ye, Xiu-mei Huang, Si-you Li, Meizhen Zhang, Yun Xue 0002
ICIC (3)6
2011 Spoken arabic digits recognition based on wavelet neural networks
abstract
The paper describes a novel method for discrete speech recognition based on spoken Arabic digit recognition by means of wavelet neural network in which Morlet wavelet is introduced to the hidden layer. The speech signal is extracted by means of Mel Frequency Cepstral Coefficients (MFCCs) and followed by vector quantization (VQ). The experimental results obtained on a spoken Arabic digit dataset proved that it could achieve better accuracy and need less learning time than the proposed method.
Lvjun Zhan, Yun Xue 0002, Weixing Zhou, Liangjun Zhang
SMC3