Caixia Yuan

dblp:69/9013 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021
YearPublicationVenuePosition
2026 Partial-Correlation Learning for Large Language Models with Skip-Tuning
Yuheng Lu, Zuhe Song, Caixia Yuan
ICPR (15)3
2025 Data with High and Consistent Preference Difference Are Better for Reward Model
abstract
Reinforcement Learning from Human Feedback (RLHF) is a commonly used alignment method for Large Language Models (LLMs). This method relies on a reward model trained on a preference dataset to provide scalar rewards. However, the human-annotated preference data is often sparse, noisy, and costly to obtain, necessitating more efficient utilization. This paper proposes a new metric for better preference data utilization from both theoretical and empirical perspectives. Starting with the Bradley-Terry model, we compute the Mean Square Error (MSE) between the expected loss and empirical loss of the reward model. Our findings reveal that data with higher and more consistent difference result in lower MSE. We therefore propose the Preference Difference (PD), the reward difference between two samples, as a filter for preference data. Experimental results on three open-source models show that reward models trained by filtered data with PD achieve higher calibrated accuracy, as well as better RLHF alignment performance. The conclusion remains consistent when we extend the experiments and theoretical derivations to implicit reward alignment algorithms, such as Direct Preference Optimization (DPO).
Hengtong Lu, Caixia Yuan, Huixing Jiang
AAAI3
2025 A Systematic Exploration of Knowledge Graph Alignment with Large Language Models in Retrieval Augmented Generation
abstract
Retrieval Augmented Generation (RAG) with Knowledge Graphs (KGs) is an effective way to enhance Large Language Models (LLMs). Due to the natural discrepancy between structured KGs and sequential LLMs, KGs must be linearized to text before being inputted into LLMs, leading to the problem of KG Alignment with LLMs (KGA). However, recent KG+RAG methods only consider KGA as a simple step without comprehensive and in-depth explorations, leaving three essential problems unclear: (1) What are the factors and their effects in KGA? (2) How do LLMs understand KGs? (3) How to improve KG+RAG by KGA? To fill this gap, we conduct systematic explorations on KGA, where we first define the problem of KGA and subdivide it into the graph transformation phase (graph-to-graph) and the linearization phase (graph-to-text). In the graph transformation phase, we study graph features at the node, edge, and full graph levels from low to high granularity. In the linearization phase, we study factors on formats, orders, and templates from structural to token levels. We conduct substantial experiments on 15 typical LLMs and three common datasets. Our main findings include: (1) The centrality of the KG affects the final generation; formats have the greatest impact on KGA; orders are model-dependent, without an optimal order adapting for all models; the templates with special token separators are better. (2) LLMs understand KGs by a unique mechanism, different from processing natural sentences, and separators play an important role. (3) We achieved 7.3% average performance improvements on four common LLMs on the KGQA task by combining the optimal factors to enhance KGA.
Shiyu Tian, Shuyue Xing, Xingrui Li, Yangyang Luo, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006
AAAI5
2025 Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language Models
abstract
Large language models (LLMs) exhibit remarkable capabilities in natural language processing but face catastrophic forgetting when learning new tasks, where adaptation to a new domain leads to a substantial decline in performance on previous tasks. In this paper, we propose Controlled LoRA (CLoRA), a subspace regularization method on LoRA structure. Aiming to reduce the scale of output change while introducing minimal constraint on model capacity, CLoRA imposes constraints on the direction of updating matrix’s null space. Experimental results on one-stage LLM finetuning tasks and continual learning settings highlight the superiority of CLoRA as an effective parameter-efficient finetuning method with catastrophic forgetting mitigating. Further investigation for model parameters indicates that CLoRA effectively balances the trade-off between model capacity and degree of forgetting. The code for implementing CLoRA will be publicly available.
Yuheng Lu, Bingshuo Qian, Caixia Yuan, Huixing Jiang
ACL (1)3
2025 Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
abstract
Large Language Models (LLMs) based agent systems have made great strides in realworld applications beyond traditional NLP tasks.This paper proposes a new LLMbased Multi-Agent System (LLM-MAS) benchmark, Collab-Overcooked, built on the popular Overcooked-AI game with more applicable and challenging tasks in interactive environments.Collab-Overcooked extends existing benchmarks in two novel ways.First, it provides a multi-agent framework supporting diverse tasks and objectives and encourages collaboration through natural language communication.Second, it introduces a spectrum of process-oriented evaluation metrics to assess the fine-grained collaboration capabilities of different LLM agents, a dimension often overlooked in prior work.We conduct extensive experiments with 13 popular LLMs and show that, while the LLMs exhibit a strong ability in goal interpretation, there are significant shortcomings in active collaboration and continuous adaptation, which are critical for efficiently fulfilling complex tasks.Notably, we highlight the strengths and weaknesses of LLM-MAS and provide insights for improving and evaluating LLM-MAS on a unified and open-source benchmark.The environments, 30 open-ended tasks, and the evaluation package are publicly available at https://github. com/YusaeMeow/Collab-Overcooked.
Lujie Niu, Fangkun Zhao, Caixia Yuan
EMNLP8
2025 VL-DynaRefine: A Vision-Language Dynamic Refinement Approach for Visual Reasoning
abstract
Visual reasoning is a key capability that significantly impacts the performance of multimodal tasks, such as compositional visual question answering and visual grounding. These tasks often require complex, multi-step reasoning processes. In recent years, several training-free methods for Vision-Language Models (VLMs) have emerged, with visual programming methods being proposed to enhance the capability of VLMs in visual reasoning tasks. While these methods have made some progress, they still face two primary challenges due to the lack of verification and refinement mechanisms for each action's output during the reasoning process: error accumulation and feedback delay, as well as insufficient utilization of multimodal contextual information. To address these challenges, we propose VL-DynaRefine, a training-free approach consisting of three modules: a planner, a verifier, and a refiner. The planner generates programmatic actions to solve the problem and executes each action in sequence, which is inspected by a verifier that reassesses the actions via confidence scores and determines whether refinement is necessary based on the evaluation results. In the refiner module, we incorporate a context-aware local refinement mechanism and a global refinement mechanism based on visual and action trajectories to reduce the impact of reasoning errors on the outcome. We evaluate our approach on multiple visual reasoning datasets, and the experimental results show that our method outperforms existing visual programming methods in both reasoning accuracy and efficiency, further validating its effectiveness in visual reasoning tasks.
Zeyuan Zang, Fangxiang Feng, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006
ACM Multimedia5
2025 Multi-task Contrastive Learning Enhanced Instruction Tuning for Dialog Understanding
Zimeng Bai, Xiying Zhao, Zhuoxin Han, Lujie Niu, Caixia Yuan, Xiaojie Wang 0006
NLPCC (3)6
2025 CGT-Corrector: Chinese Government Text Correction with Knowledge Bases
Fangkun Zhao, Zimeng Bai, Lujie Niu, Caixia Yuan
NLPCC (4)5
2024 Enhancing Document Information Selection Through Multi-Granularity Responses for Dialogue Generation
abstract
Abstract Document information selection is an essential part of document-grounded dialogue tasks, and more accurate information selection results can provide more appropriate dialogue responses. Existing works have achieved excellent results by employing multi-granularity of dialogue history information, indicating the effectiveness of multi-level historical information. However, these works often focus on exploring the hierarchical information of dialogue history, while neglecting the multi-granularity utilization in response, important information that holds an impact on the decoding process. Therefore, this paper proposes a model for document information selection based on multi-granularity responses. By integrating the document selection results at the response word level and semantic unit level, the model enhances its capability in knowledge selection and produces better responses. For the division at the semantic unit level of the response, we propose two semantic unit division methods, static and dynamic. Experiments on two public datasets show that our models combining static or dynamic semantic unit levels significantly outperform baseline models.
Kangyu Qiao, Shuyue Xing, Caixia Yuan, Xiaojie Wang 0006
Neural Process. Lett.4
2023 SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph
abstract
Existing multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In this paper, we propose a Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph (SPRING) with abilities of reasoning multi-hops spatial relations and connecting them with visual attributes in crowded situated scenarios. Specifically, we design two types of Multimodal Question Answering (MQA) tasks to pretrain the agent. All QA pairs utilized during pretraining are generated from novel Increment Layout Graphs (ILG). QA pair difficulty labels automatically annotated by ILG are used to promote MQA-based Curriculum Learning. Experimental results verify the SPRING's effectiveness, showing that it significantly outperforms state-of-the-art approaches on both SIMMC 1.0 and SIMMC 2.0 datasets. We release our code and data at https://github.com/LYX0501/SPRING.
Yuxing Long, Binyuan Hui, Fulong Ye, Yanyang Li, Zhuoxin Han, Caixia Yuan, Yongbin Li 0001, Xiaojie Wang 0006
AAAI6
2023 Joint Modeling for ASR Correction and Dialog State Tracking
abstract
In spoken dialog system, transcription errors in Automated Speech Recognition (ASR) impact downstream task, especially dialog state tracking (DST). Approaches to alleviate such errors involve using richer information such as word-lattices and word confusion networks. However, in some cases, this information may not be easily obtained. In addition, the large pre-trained language model is trained on plain text, leading to the gap between spoken DST and original pretrained model. In this paper, we propose a multi-task method which performs DST jointly with ASR correction to improve the performance of both tasks. To do so, we build a MultiWOZ-ASR dataset containing ASR noise in DST and mitigate the gap by utilizing a multi-task pre-training framework. Moreover, curriculum learning is adopted to alleviate the phenomenon that the correction task is difficult to converge at the initial stage of pre-training. Experimental results show that our model achieves significant improvements on DSTC2 and MultiWOZ-ASR dataset.
Deyuan Wang, Caixia Yuan, Xiaojie Wang 0006
ICASSP3
2023 An Asynchronous Updating Reinforcement Learning Framework for Task-Oriented Dialog System
abstract
Reinforcement learning has been applied to train the dialog systems in many works. Previous approaches divide the dialog system into multiple modules including DST (dialog state tracking) and DP (dialog policy), and train these modules simultaneously. However, different modules influence each other during training. The errors from DST might misguide the dialog policy, and the system action brings extra difficulties for the DST module. To alleviate this problem, we propose Asynchronous Updating Reinforcement Learning framework (AURL) that updates the DST module and the DP module asynchronously under a cooperative setting. Furthermore, curriculum learning is implemented to address the problem of unbalanced data distribution during reinforcement learning sampling, and multiple user models are introduced to increase the dialog diversity. Results on the public SSD-PHONE dataset show that our method achieves a compelling result with a 31.37% improvement on the dialog success rate. The code is publicly available via https://github.com/shunjiu/AURL.
Xiaojie Wang 0006, Caixia Yuan
ICASSP4
2023 Introducing Natural Language-based Instruction Protocol for Intelligent Machine Collaborations
abstract
The rapid development of various autonomous unmanned systems has increased the endogenous intelligence of machines, which empowers functional collaborations among intelligent machines. However, traditional protocols for intelligent machine collaboration have limitations regarding functionality, efficiency, and scalability. Moreover, existing research on intelligent machine collaboration is not yet approaching a unified paradigm that facilitates interactions among machines and between machines and humans. Therefore, aiming to enable a more efficient functional collaboration among intelligent machines, we propose a natural language-based instruction (NLI) protocol, which enjoys the advantages of autonomy, robustness and efficiency. In particular, we specify the NLI protocol architecture by introducing a corpus of NLIs, an NLI generation module, and an NLI parsing module, wherein the corpus contains control instructions, intents and slots, the generation module is used to generate NLIs, and the parsing module is used to parse the intents and slots in the instruction. Furthermore, we case-study the proposed NLI protocol in an electromagnetic interference avoidance scenario based on semi-physical simulation with software-defined radio. Simulation results show that the NLI protocol is more robust and efficient than traditional control protocol in severe wireless channels, which validates the feasibility and effectiveness of implementing intelligent machine collaboration based on NLI.
Haobing Gong, Hui Gao 0001, Caixia Yuan
IWCMC4
2023 A Task-Oriented Dialog Model with Task-Progressive and Policy-Aware Pre-training
Lucen Zhong, Hengtong Lu, Caixia Yuan, Xiaojie Wang 0006, Jiashen Sun, Guanglu Wan
NLPCC (1)3
2023 Hierarchical history based information selection for document grounded dialogue generation
Shiyu Tian, Ziwei Bai, Caixia Yuan, Xiaojie Wang 0006
Appl. Intell.4
2022 Learn to Adapt for Generalized Zero-Shot Text Classification
abstract
Generalized zero-shot text classification aims to classify textual instances from both previously seen classes and incrementally emerging unseen classes.Most existing methods generalize poorly since the learned parameters are only optimal for seen classes rather than for both classes, and the parameters keep stationary in predicting procedures.To address these challenges, we propose a novel Learn to Adapt (LTA) network using a variant meta-learning framework.Specifically, LTA trains an adaptive classifier by using both seen and virtual unseen classes to simulate a generalized zero-shot learning (GZSL) scenario in accordance with the test time, and simultaneously learns to calibrate the class prototypes and sample representations to make the learned parameters adaptive to incoming unseen classes.We claim that the proposed model is capable of representing all prototypes and samples from both classes to a more consistent distribution in a global space.Extensive experiments on five text classification datasets show that our model outperforms several competitive previous approaches by large margins.
Caixia Yuan, Xiaojie Wang 0006, Ziwei Bai
ACL (1)2
2022 Domain adaptive multi-task transformer for low-resource machine reading comprehension
Ziwei Bai, Baoxun Wang, Zongsheng Wang, Caixia Yuan, Xiaojie Wang 0006
Neurocomputing4
2021 Converse, Focus and Guess - Towards Multi-Document Driven Dialogue
Caixia Yuan, Xiaojie Wang 0006, Yushu Yang, Huixing Jiang, Zhongyuan Wang 0006
AAAI2
2020 Label-Wise Document Pre-training for Multi-label Text Classification
Caixia Yuan, Xiaojie Wang 0006
NLPCC (1)2
2019 MrMep: Joint Extraction of Multiple Relations and Multiple Entity Pairs Based on Triplet Attention
abstract
This paper focuses on how to extract multiple relational facts from unstructured text.Neural encoder-decoder models have provided a viable new approach for jointly extracting relations and entity pairs.However, these models either fail to deal with entity overlapping among relational facts, or neglect to produce the whole entity pairs.In this work, we propose a novel architecture that augments the encoder and decoder in two elegant ways.First, we apply a binary CNN classifier for each relation, which identifies all possible relations maintained in the text, while retaining the target relation representation to aid entity pair recognition.Second, we perform a multihead attention over the text and a triplet attention with the target relation interacting with every token of the text to precisely produce all possible entity pairs in a sequential manner.Experiments on three benchmark datasets show that our proposed method successfully addresses the multiple relations and multiple entity pairs even with complex overlapping and significantly outperforms the state-of-theart methods.All source code and documentations are available at https://github. com/chenjiayu1502/MrMep.
Caixia Yuan, Xiaojie Wang 0006, Ziwei Bai
CoNLL2
2019 Hierarchical Dialog State Tracking with Unknown Slot Values
Guohua Yang, Xiaojie Wang 0006, Caixia Yuan
Neural Process. Lett.3
2016 Image color harmony modeling through neighbored co-occurrence colors
Peng Lu 0007, Xujun Peng, Caixia Yuan, Ruifan Li, Xiaojie Wang 0006
Neurocomputing3
2015 Recognition of Person Relation Indicated by Predicates
abstract
This paper focuses on recognizing person relations indicated by predicates from large scale of free texts. In order to determine whether a sentence contains a potential relation between persons, we cast this problem to a classification task. Dynamic Convolution Neural Network (DCNN) is improved for this task. It uses frame convolution for making uses of more features efficiently. Experimental results on Chinese person relation recognition show that the proposed model is superior when compared to the original DCNN and several strong baseline models. We also explore employing large scale unlabeled data to achieve further improvements.
Zhongping Liang, Caixia Yuan, Bing Leng, Xiaojie Wang 0006
NLPCC2
2015 Stochastic Language Generation Using Situated PCFGs
abstract
This paper presents a purely data-driven approach for generating natural language (NL) expressions from its corresponding semantic representations. Our aim is to exploit a parsing paradigm for natural language generation (NLG) task, which first encodes semantic representations with a situated probabilistic context-free grammar (PCFG), then decodes and yields natural sentences at the leaves of the optimal parsing tree. We deployed our system in two different domains, one is response generation for a Chinese spoken dialogue system, and the other is instruction generation for a virtual environment in English language, obtaining results comparable to state-of-the-art systems both in terms of BLEU scores and human evaluation.
Caixia Yuan, Xiaojie Wang 0006, Ziming Zhong
NLPCC1
2012 A Linear Max K-min classifier
Mingzhi Dong, Weihong Deng, Qiang Wang 0048, Caixia Yuan, Jun Guo 0002, Liwei Ma
ICPR5
2008 BUPT Systems in the SIGHAN Bakeoff 2007
Caixia Yuan, Jiashen Sun, Xiaojie Wang 0006
IJCNLP2