EDBT 2026 Demo / reviewers in the wild / expert
Xiaocui Yang
dblp:219/8174
· DBLP profile ↗
30ranked-venue papers
4as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response GenerationabstractZhuoyue Gao, Xiaohui Wang, Xiaocui Yang, Wen Zhang, Daling Wang, Shi Feng, Yifei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhuoyue Gao, Xiaocui Yang, Daling Wang, Shi Feng 0001, Yifei Zhang 0003 |
ACL (1) | 3 |
| 2026 | Cat-MoD: Accelerating Multimodal Alignment via Caption Token Guided Asymmetric Mixture-of-DepthsabstractYiJie Huang, Xiaocui Yang, Shi Feng, Wen Zhang, Kaisong Song, Yifei Zhang, Daling Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaocui Yang, Shi Feng 0001, Kaisong Song, Yifei Zhang 0003, Daling Wang |
ACL (1) | 2 |
| 2026 | SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement LearningabstractPeidong Wang, Zhiming Ma, Xin Dai, YongKang Liu, Shi Feng, Xiaocui Yang, Wenxing Hu, Zhihao Wang, Mingjun Pan, Li Yuan, Daling Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peidong Wang 0001, Zhiming Ma, Yongkang Liu 0002, Shi Feng 0001, Xiaocui Yang, Wenxing Hu, Mingjun Pan, Li Yuan 0007, Daling Wang |
ACL (1) | 6 |
| 2026 | GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective TestingabstractMing Wang, Shuang Wu, Bixuan Wang, Lu Lin, Yuxin Chen, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang, Yufan Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ming Wang 0006, Bixuan Wang, Xiaocui Yang, Daling Wang, Shi Feng 0001, Yifei Zhang 0003, Yufan Sun |
ACL (1) | 6 |
| 2026 | Why Do More Experts Fail? A Theoretical Analysis of Model MergingabstractZijing Wang, Xingle Xu, YongKang Liu, Yiqun Zhang, Peiqin Lin, Shi Feng, Daling Wang, Xiaocui Yang, Hinrich Schuetze. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xingle Xu, Yongkang Liu 0002, Peiqin Lin, Shi Feng 0001, Daling Wang, Xiaocui Yang, Hinrich Schütze |
ACL (1) | 8 |
| 2026 | CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question AnsweringabstractZili Wei, Yilin Wang, Xiaocui Yang, Shi Feng, Weidong Bao, Daling Wang, Zihan Wang, Yifei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zili Wei, Xiaocui Yang, Shi Feng 0001, Weidong Bao 0005, Daling Wang, Yifei Zhang 0003 |
ACL (1) | 3 |
| 2026 | MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint EmbeddingsabstractYiqun Zhang, Hao Li, Zihan Wang, Shi Feng, Xiaocui Yang, Daling Wang, Bo Zhang, Lei Bai, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Li 0069, Shi Feng 0001, Xiaocui Yang, Daling Wang, Bo Zhang 0069, Lei Bai 0001, Shuyue Hu |
ACL (1) | 5 |
| 2026 | Nature-Inspired Population-Based Evolution of Large Language ModelsabstractYiqun Zhang, Peng Ye, Xiaocui Yang, Shi Feng, Shufei Zhang, Lei Bai, Wanli Ouyang, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peng Ye 0006, Xiaocui Yang, Shi Feng 0001, Shufei Zhang, Lei Bai 0001, Wanli Ouyang, Shuyue Hu |
ACL (1) | 3 |
| 2026 | HIPPO: Enhancing the Table Understanding Capability of LLMs Through Hybrid-Modal Preference Optimization
Haolan Wang, Zhenghao Liu 0001, Xiaocui Yang, Yu Gu 0002, Yukun Yan, Qi Shi 0002, Fangfang Li 0002, Ge Yu 0001 |
DASFAA (4) | 4 |
| 2026 | Enhancing LLM-Based Recommendation with Semantic-Aligned Collaborative Knowledge
Jinghao Lin, Xiaocui Yang, Yongkang Liu 0002, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Ge Yu 0001 |
DASFAA (1) | 3 |
| 2026 | Advancing referring image segmentation with bidirectional feature enhancement and adaptive multimodal fusion
Wen Qu, Xiaocui Yang, Yonggong Ren |
Neurocomputing | 3 |
| 2026 | Affective computing in the era of large language models: A survey from the NLP perspective
Xiaocui Yang, Xingle Xu, Zeran Gao, Shiyi Mu, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Kaisong Song, Ge Yu 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Pixel-Level Reasoning Segmentation via Multi-turn ConversationsabstractDexian Cai, Xiaocui Yang, YongKang Liu, Daling Wang, Shi Feng, Yifei Zhang, Soujanya Poria. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Dexian Cai, Xiaocui Yang, Yongkang Liu 0002, Daling Wang, Shi Feng 0001, Yifei Zhang 0003, Soujanya Poria |
ACL (1) | 2 |
| 2025 | MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction-Tuning for Emotion-Cause Pair Extraction
Shiyi Mu, Yongkang Liu 0002, Shi Feng 0001, Xiaocui Yang, Daling Wang, Yifei Zhang 0003 |
CogSci | 4 |
| 2025 | TOOL-ED: Enhancing Empathetic Response Generation with the Tool Calling Capability of LLMabstractEmpathetic conversation is a crucial characteristic in daily conversations between individuals. Nowadays, Large Language models (LLMs) have shown outstanding performance in generating empathetic responses. Knowledge bases like COMET can assist LLMs in mitigating illusions and enhancing the understanding of users’ intentions and emotions. However, models remain heavily reliant on fixed knowledge bases and unrestricted incorporation of external knowledge can introduce noise. Tool learning is a flexible end-to-end approach that assists LLMs in handling complex problems. In this paper, we propose Emotional Knowledge Tool Calling (EKTC) framework, which encapsulates the commonsense knowledge bases as empathetic tools, enabling LLMs to integrate external knowledge flexibly through tool calling. In order to adapt the models to the new task, we construct a novel dataset TOOL-ED based on the EMPATHETICDIALOGUE (ED) dataset. We validate EKTC on the ED dataset, and the experimental results demonstrate that our framework can enhance the ability of LLMs to generate empathetic responses effectively. Our code is available at https://anonymous.4open.science/r/EKTC-3FEF. Huiying Cao, Shi Feng 0001, Xiaocui Yang, Daling Wang, Yifei Zhang 0003 |
COLING | 4 |
| 2025 | Enhancing Zero-Shot Emotion Perception in Conversation Through the Internal-to-External Chain-of-Thought
Xingle Xu, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Xiaocui Yang |
DASFAA (2) | 5 |
| 2025 | Language Models as Continuous Self-Evolving Data EngineersabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities, yet their further evolution is often hampered by the scarcity of high-quality training data and the heavy reliance of traditional methods on expert-labeled data.This reliance sets a ceiling on LLM performance and is particularly challenging in low data resource scenarios where extensive supervision is unavailable.To address this issue, we propose a novel paradigm named LANCE (LANguage models as Continuous self-Evolving data engineers) that enables LLMs to train themselves by autonomously generating, cleaning, reviewing, and annotating data with preference information.Our approach demonstrates that LLMs can serve as continuous self-evolving data engineers, significantly reducing the time and cost of post-training data construction.Through iterative fine-tuning on Qwen2 series models, we validate the effectiveness of LANCE across various tasks, showing that it can maintain high-quality data generation and continuously improve model performance.Across multiple benchmark dimensions, LANCE results in an average score enhancement of 3.64 for Qwen2-7B and 1.75 for Qwen2-7B-Instruct.This autonomous data construction paradigm not only lessens reliance on human experts or external models but also ensures data aligns with human preferences, offering a scalable path for LLM self-improvement, especially in contexts with limited supervisory data. Peidong Wang 0001, Ming Wang 0006, Zhiming Ma, Xiaocui Yang, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Kaisong Song |
EMNLP | 4 |
| 2025 | Generative Emotion Cause Explanation in Multimodal ConversationsabstractMultimodal conversation, a crucial form of human communication, carries rich emotional content, making the exploration of the causes of emotions within it a research endeavor of significant importance. However, existing research on the causes of emotions typically employs an utterance selection method within a single textual modality to locate causal utterances. This approach remains limited to coarse-grained assessments, lacks nuanced explanations of emotional causation, and demonstrates inadequate capability in identifying multimodal emotional triggers. Therefore, we introduce a task-Multimodal Emotion Cause Explanation in Conversation (MECEC). This task aims to generate a summary based on the multimodal context of conversations, clearly and intuitively describing the reasons that trigger a given emotion. To adapt to this task, we develop a new dataset (ECEM) based on the MELD dataset. ECEM combines video clips with detailed explanations of character emotions, helping to explore the causal factors behind emotional expression in multimodal conversations. A novel approach, FAME-Net, is further proposed, that harnesses the power of Large Language Models (LLMs) to analyze visual data and accurately interpret the emotions conveyed through facial expressions in videos. By exploiting the contagion effect of facial emotions, FAME-Net effectively captures the emotional causes of individuals engaged in conversations. Our experimental results on the newly constructed dataset show that FAME-Net outperforms several excellent baselines. Code and dataset are available at https://github.com/3222345200/FAME-Net. Lin Wang 0063, Xiaocui Yang, Shi Feng 0001, Daling Wang, Yifei Zhang 0003 |
ICMR | 2 |
| 2025 | A Two-Stage Full Fine-Tuning and LLM Post-processing Framework for MCABSAabstractIn recent years, Multimodal Sentiment Analysis (MSA) has attracted growing attention for its ability to interpret human emotions by integrating information across multiple modalities. Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) extends this research frontier by incorporating multi-party conversational contexts and requiring comprehensive extraction of sentiment elements. MCABSA presents substantial challenges, including the need to understand complex conversational contexts, integrate heterogeneous multimodal signals, and identify causal reasoning at the cognitive level. To address these challenges, we propose a two-stage Full Fine-tuning and LLM Post-processing (FLP) framework. In the first stage, we develop a multimodal caption-enhanced full fine-tuning pipeline that performs structured extraction of sextuples and sentiment flip tuples. The second stage introduces paraphrase-based sextuple verification to identify and filter low-quality sextuples for Panoptic Sentiment Sextuple Extraction (Task-1), while implementing trigger classification with a distribution alignment mechanism to determine trigger types for sentiment flipping and enhance output consistency for Sentiment Flipping Analysis (Task-2). Comprehensive experiments on both MCABSA challenge subtasks demonstrate the effectiveness of our approach, achieving 1st place on Task-1 and 3rd place on Task-2. Deyuan Chen, Xiaocui Yang, Shi Feng 0001, Daling Wang, Yifei Zhang 0003 |
ACM Multimedia | 2 |
| 2025 | Adaptive Persona Context Modulation for Personalized Emotional Support Conversation
Xiaocui Yang, Daling Wang, Shi Feng 0001, Yifei Zhang 0003 |
PRICAI (4) | 2 |
| 2025 | Is Mamba effective for time series forecasting?
Fanheng Kong, Shi Feng 0001, Ming Wang 0006, Xiaocui Yang, Daling Wang, Yifei Zhang 0003 |
Neurocomputing | 5 |
| 2024 | Few-Shot Multimodal Named Entity Recognition Based on Mutlimodal Causal Intervention GraphabstractMultimodal Named Entity Recognition (MNER) models typically require a significant volume of labeled data for effective training to extract relations between entities. In real-world scenarios, we frequently encounter unseen relation types. Nevertheless, existing methods are predominantly tailored for complete datasets and are not equipped to handle these new relation types. In this paper, we introduce the Few-shot Multimodal Named Entity Recognition (FMNER) task to address these novel relation types. FMNER trains in the source domain (seen types) and tests in the target domain (unseen types) with different distributions. Due to limited available resources for sampling, each sampling instance yields different content, resulting in data bias and alignment problems of multimodal units (image patches and words). To alleviate the above challenge, we propose a novel Multimodal causal Intervention graphs (MOUSING) model for FMNER. Specifically, we begin by constructing a multimodal graph that incorporates fine-grained information from multiple modalities. Subsequently, we introduce the Multimodal Causal Intervention Strategy to update the multimodal graph. It aims to decrease spurious correlations and emphasize accurate correlations between multimodal units, resulting in effectively aligned multimodal representations. Extensive experiments on two multimodal named entity recognition datasets demonstrate the superior performance of our model in the few-shot setting. Feihong Lu, Xiaocui Yang, Qian Li 0033, Qingyun Sun, Cheng Ji 0001, Jianxin Li 0002 |
LREC/COLING | 2 |
| 2024 | PAPER: A Persona-Aware Chain-of-Thought Learning Framework for Personalized Dialogue Response Generation
Yameng Li, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Xiaocui Yang |
NLPCC (1) | 5 |
| 2024 | Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet ExtractionabstractDocument-level Relation Triplet Extraction (DocRTE) is a fundamental task in information systems that aims to simultaneously extract entities with semantic relations from a document. Existing methods heavily rely on a substantial amount of fully labeled data. However, collecting and annotating data for newly emerging relations is time-consuming and labor-intensive. Recent advanced Large Language Models (LLMs), such as ChatGPT and LLaMA, exhibit impressive long-text generation capabilities, inspiring us to explore an alternative approach for obtaining auto-labeled documents with new relations. In this paper, we propose a Zero-shot Document-level Relation Triplet Extraction (ZeroDocRTE) framework, which Generates labeled data by Retrieval and Denoising Knowledge from LLMs, called GenRDK. Specifically, we propose a chain-of-retrieval prompt to guide ChatGPT to generate labeled long-text data step by step. To improve the quality of synthetic data, we propose a denoising strategy based on the consistency of cross-document knowledge. Leveraging our denoised synthetic data, we proceed to fine-tune the LLaMA2-13B-Chat for extracting document-level relation triplets. We perform experiments for both zero-shot document-level relation and triplet extraction on two public datasets. The experimental results illustrate that our GenRDK framework outperforms strong baselines. Xiaocui Yang, Rong Tong, Soujanya Poria |
WWW | 3 |
| 2023 | Uncertainty Guided Label Denoising for Document-level Distant Relation ExtractionabstractDocument-level relation extraction (DocRE)aims to infer complex semantic relations among entities in a document.Distant supervision (DS) is able to generate massive auto-labeled data, which can improve DocRE performance.Recent works leverage pseudo labels generated by the pre-denoising model to reduce noise in DS data.However, unreliable pseudo labels bring new noise, e.g., adding false pseudo labels and losing correct DS labels.Therefore, how to select effective pseudo labels to denoise DS data is still a challenge in document-level distant relation extraction.To tackle this issue, we introduce uncertainty estimation technology to determine whether pseudo labels can be trusted.In this work, we propose a Documentlevel distant Relation Extraction framework with Uncertainty Guided label denoising, UG-DRE.Specifically, we propose a novel instancelevel uncertainty estimation method, which measures the reliability of the pseudo labels with overlapping relations.By further considering the long-tail problem, we design dynamic uncertainty thresholds for different types of relations to filter high-uncertainty pseudo labels.We conduct experiments on two public datasets.Our framework outperforms strong baselines by 1.91 F 1 and 2.28 Ign F 1 on the RE-DocRED dataset. Xiaocui Yang, Pengfei Hong, Soujanya Poria |
ACL (1) | 3 |
| 2023 | Multiple Contrastive Learning for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis has received extensive attention with the explosion of multimodal data. For multimodal data, representations should have disparate distributions in the feature space under different labels. The paired multi-modal image-text posts should be closer than unpaired. We propose Multimodal fine-grained interaction with the Multiple Contrastive Learning (M2CL) model for image-text multi-modal sentiment detection. Specifically, we first obtain the reinforced global representation of one modality with the assistance of fine-grained information from another via the Multimodal Interaction Component. Then, we introduce the Multiple Contrastive Learning Component, including Supervised Contrastive Learning (SCL) and Dual Multimodal Contrastive Learning (DMCL). SCL accomplishes pushing the posts with the same sentiment closer and pulling the instances of different sentiments apart within each modality. DMCL pushes the paired image-text features together and pulls the unpaired apart with multiple stages. Extensive experiments conducted on three datasets confirm the effectiveness of our approach. Xiaocui Yang, Shi Feng 0001, Daling Wang, Pengfei Hong, Soujanya Poria |
ICASSP | 1 |
| 2023 | Few-shot Multimodal Sentiment Analysis Based on Multimodal Probabilistic Fusion PromptsabstractMultimodal sentiment analysis has gained significant attention due to the proliferation of multimodal content on social media. However, existing studies in this area rely heavily on large-scale supervised data, which is time-consuming and labor-intensive to collect. Thus, there is a need to address the challenge of few-shot multimodal sentiment analysis. To tackle this problem, we propose a novel method called Multimodal Probabilistic Fusion Prompts (MultiPoint) that leverages diverse cues from different modalities for multimodal sentiment detection in the few-shot scenario. Specifically, we start by introducing a Consistently Distributed Sampling approach called CDS, which ensures that the few-shot dataset has the same category distribution as the full dataset. Unlike previous approaches primarily using prompts based on the text modality, we design unified multimodal prompts to reduce discrepancies between different modalities and dynamically incorporate multimodal demonstrations into the context of each multimodal instance. To enhance the model's robustness, we introduce a probabilistic fusion method to fuse output predictions from multiple diverse prompts for each input. Our extensive experiments on six datasets demonstrate the effectiveness of our approach. First, our method outperforms strong baselines in the multimodal few-shot setting. Furthermore, under the same amount of data (1% of the full dataset), our CDS-based experimental results significantly outperform those based on previously sampled datasets constructed from the same number of instances of each class. Xiaocui Yang, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Soujanya Poria |
ACM Multimedia | 1 |
| 2021 | Multimodal Sentiment Detection Based on Multi-channel Graph Neural NetworksabstractXiaocui Yang, Shi Feng, Yifei Zhang, Daling Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xiaocui Yang, Shi Feng 0001, Yifei Zhang 0003, Daling Wang |
ACL/IJCNLP (1) | 1 |
| 2021 | Image-Text Multimodal Emotion Classification via Multi-View Attentional NetworkabstractCompared with single-modal content, multimodal data can express users’ feelings and sentiments more vividly and interestingly. Therefore, multimodal sentiment analysis has become a popular research topic. However, most existing methods either learn modal sentiment feature independently, without considering their correlations, or they simply integrate multimodal features. In addition, most publicly available multimodal datasets are labeled by sentiment polarities, while the emotions expressed by users are specific. Based on this observation, in this paper, we build a large-scale image-text emotion dataset (i.e., labeled by different emotions), called TumEmo, with more than190,000instances from Tumblr.1We further propose a novel multimodal emotion analysis model based on the Multi-view Attentional Network (MVAN), which utilizes a memory network that is continually updated to obtain the deep semantic features of image-text. The model includes three stages: feature mapping, interactive learning, and feature fusion. In the feature mapping stage, we leverage image features from an object viewpoint and a scene viewpoint to capture effective information for multimodal emotion analysis. Then, an interactive learning mechanism is adopted that uses the memory network; this mechanism extracts single-modal emotion features and interactively models the cross-view dependencies between the image and text. In the feature fusion stage, multiple features are deeply fused using a multilayer perceptron and a stacking-pooling module. The experimental results on the MVSA-Single, MVSA-Multiple, and TumEmo datasets show that the proposed MVAN outperforms strong baseline models by large margins. Xiaocui Yang, Shi Feng 0001, Daling Wang, Yifei Zhang 0003 |
IEEE Trans. Multim. | 1 |
| 2021 | SINN: A speaker influence aware neural network model for emotion detection in conversations
Shi Feng 0001, Daling Wang, Xiaocui Yang, Zhenfei Yang, Yifei Zhang 0003, Ge Yu 0001 |
World Wide Web | 4 |