Youcheng Pan

dblp:162/8127 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-8270-5455ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
abstract
As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual modalities with a primary emphasis on English, which creates a gap in evaluation when processing multilingual input, especially in speech. To bridge this gap, we propose a novel Cross-lingual and Cross-modal Factuality benchmark (CCFQA). Specifically, the CCFQA benchmark contains parallel speech-text factual questions across 8 languages, designed to systematically evaluate MLLMs' cross-lingual and cross-modal factuality capabilities. Our experimental results demonstrate that current MLLMs still face substantial challenges on the CCFQA benchmark. Furthermore, we propose a few-shot transfer learning strategy that effectively transfers the Question Answering (QA) capabilities of LLMs in English to multilingual Spoken Question Answering (SQA) tasks, achieving competitive performance with GPT-4o-mini-Audio using just 5-shot training. We release CCFQA as a foundational research resource to promote the development of MLLMs with more robust and reliable speech understanding capabilities.
Yexing Du, Youcheng Pan, Bo Yang 0006, Ming Liu 0004, Yang Xiang 0003
AAAI3
2025 Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
abstract
Yexing Du, Youcheng Pan, Ziyang Ma, Bo Yang, Yifan Yang, Keqi Deng, Xie Chen, Yang Xiang, Ming Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yexing Du, Youcheng Pan, Ziyang Ma 0001, Bo Yang 0006, Yifan Yang 0005, Keqi Deng, Xie Chen 0001, Yang Xiang 0003, Ming Liu 0004, Bing Qin 0001
ACL (1)2
2025 From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement
abstract
Chain-of-Thought (CoT) reasoning improves performance on complex tasks but introduces significant inference latency due to its verbosity.In this work, we propose Multiround Adaptive Chain-of-Thought Compression (MACC), a framework that leverages the token elasticity phenomenon-where overly small token budgets may paradoxically increase output length-to progressively compress CoTs via multiround refinement.This adaptive strategy allows MACC to dynamically determine the optimal compression depth for each input.Our method achieves an average accuracy improvement of 5.6% over state-of-the-art baselines, while also reducing CoT length by an average of 47 tokens and significantly lowering latency.Furthermore, we show that test-time performance-accuracy and token length-can be reliably predicted using interpretable features like perplexity and compression rate on training set.Evaluated across different models, our method enables efficient model selection and forecasting without repeated fine-tuning, demonstrating that CoT compression is both effective and predictable.Our code will be released in https://github.com/Leon221220/ MACC.
Jianzhi Yan, Youcheng Pan, Zike Yuan, Yang Xiang 0003, Buzhou Tang
EMNLP3
2025 ReFEdit: Rehearsal-Free Lifelong Knowledge Editing for Large Language Models
abstract
Knowledge editing has emerged as a promising strategy for updating obsolete or inaccurate knowledge embedded within large language models (LLMs) without costly fine-tuning. The widely adopted locating-then-editing paradigm first locates parameters responsible for knowledge storage and then modifies them to integrate updated knowledge. However, in lifelong knowledge editing scenarios, catastrophic forgetting poses a significant challenge. Existing methods often rely on rehearsal-based techniques, such as storing a feature covariance matrix of previously preserved knowledge to constrain errors, thus raising efficiency and privacy issues. To address this, we introduce ReFEdit, a Rehearsal-Free Lifelong Knowledge Editing framework that enforces an orthogonality restriction on parameter modifications. By aligning the update direction orthogonally to both the latest and initial parameters, ReFEdit minimizes the interference between sequentially edited knowledge while mitigating the impact on previously preserved knowledge, thereby effectively addressing catastrophic forgetting. Extensive evaluations on multiple representative LLMs, including LLaMA3, GPT-J, and GPT2-XL, demonstrate that ReFEdit significantly outperforms most existing rehearsal-based knowledge editing methods while eliminating the need for the rehearsal phase, marking a substantial advancement toward more reliable and flexible lifelong knowledge editing. Our code is available at: https://github.com/Cedric-Mo/ReFEdit
Xianjie Mo, Youcheng Pan, Yongshuai Hou, Ping Luo 0001, Yang Xiang 0003
ICME2
2024 ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order Optimization
abstract
Lowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference. However, the simulated perturbation stochastic approximation for gradient estimate in MeZO leads to severe oscillations and incurs a substantial time overhead. Moreover, without momentum regularization, MeZO shows severe over-fitting problems. Lastly, the perturbation-irrelevant momentum on ZO-SGD does not improve the convergence rate. This study proposes ZO-AdaMU to resolve the above problems by adapting the simulated perturbation with momentum in its stochastic approximation. Unlike existing adaptive momentum methods, we relocate momentum on simulated perturbation in stochastic gradient approximation. Our convergence analysis and experiments prove this is a better way to improve convergence stability and rate in ZO-SGD. Extensive experiments demonstrate that ZO-AdaMU yields better generalization for LLMs fine-tuning across various NLP tasks than MeZO and its momentum variants.
Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang 0003, Yukang Lin, Xiangping Wu 0001, Chuanyi Liu, Xiaobao Song
AAAI3
2024 Linguistic Rule Induction Improves Adversarial and OOD Robustness in Large Language Models
abstract
Ensuring robustness is especially important when AI is deployed in responsible or safety-critical environments. ChatGPT can perform brilliantly in both adversarial and out-of-distribution (OOD) robustness, while other popular large language models (LLMs), like LLaMA-2, ERNIE and ChatGLM, do not perform satisfactorily in this regard. Therefore, it is valuable to study what efforts play essential roles in ChatGPT, and how to transfer these efforts to other LLMs. This paper experimentally finds that linguistic rule induction is the foundation for identifying the cause-effect relationships in LLMs. For LLMs, accurately processing the cause-effect relationships improves its adversarial and OOD robustness. Furthermore, we explore a low-cost way for aligning LLMs with linguistic rules. Specifically, we constructed a linguistic rule instruction dataset to fine-tune LLMs. To further energize LLMs for reasoning step-by-step with the linguistic rule, we construct the task-relevant LingR-based chain-of-thoughts. Experiments showed that LingR-induced LLaMA-13B achieves comparable or better results with GPT-3.5 and GPT-4 on various adversarial and OOD robustness evaluations.
Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Yukang Lin
LREC/COLING4
2024 A Lifelong Multilingual Multi-granularity Semantic Alignment Approach via Maximum Co-occurrence Probability
abstract
Cross-lingual pre-training methods mask and predict tokens in multilingual text to generalize diverse multilingual information. However, due to the lack of sufficient aligned multilingual resources in the pre-training process, these methods may not fully explore the multilingual correlation of masked tokens, resulting in the limitation of multilingual information interaction. In this paper, we propose a lifelong multilingual multi-granularity semantic alignment approach, which continuously extracts massive aligned linguistic units from noisy data via a maximum co-occurrence probability algorithm. Then, the approach releases a version of the multilingual multi-granularity semantic alignment resource, supporting seven languages, namely English, Czech, German, Russian, Romanian, Hindi and Turkish. Finally, we propose how to use this resource to improve the translation performance on WMT14 18 benchmarks in twelve directions. Experimental results show an average of 0.3 1.1 BLEU improvements in all translation benchmarks. The analysis and discussion also demonstrate the superiority and potential of the proposed approach. The resource used in this work will be publicly available.
Shaojie Dai, Youcheng Pan
LREC/COLING5
2024 Mitigating Knowledge Conflicts in Data-to-Text Generation via the Internalization of Fact Extraction
abstract
Large Language Models (LLMs) have made remarkable advancements in Natural Language Generation. Nonetheless, LLMs are prone to encountering knowledge conflicts, scenarios where the generated statement exhibits either twisted facts that contradict the source or unverified facts that lack substantiation from the source. We posit that LLMs should possess the capability to differentiate between facts grounded in evidence and those derived from inference. A promising avenue for mitigating knowledge conflicts involves prompting LLMs to externalize the multi-step reasoning during text generation within the chain-of-thought paradigm. Nevertheless, such externalization of reasoning has only shown benefits for sufficiently large LLMs (e.g., those exceeding 100B parameters). Despite attempts to distill this ability into smaller LLMs, the reliance on large and inefficient teacher models remains a necessity. Thus, we introduce the Internalization of Fact Extraction (IFE), which optimizes instruction-tuning, empowering smaller LLMs to internalize the fact extraction process during text generation within the auxiliary learning paradigm. Consequently, models can directly generate statements where knowledge conflicts are mitigated. Specifically, we design an auxiliary fact extraction task built upon the primary text generation task, requiring models to identify evidential facts and filter out inferential facts after text generation. The auxiliary training dataset for the auxiliary fact extraction task was constructed using word-stemming and string-matching techniques. By leveraging auxiliary learning, the models’ ability to internalize the fact extraction process emerges, thereby enhancing its performance in the primary text generation task. Extensive experiments and analyses are conducted on three benchmark data-to-text generation datasets (i.e., FeTaQA, DialogSum, and SAMSum), covering sub-tasks including table-based text generation and dialogue-based text generation. Experimental results show that our IFE instruction-tuned models outperform the vanilla instruction-tuned baselines in terms of both factuality and quality without introducing any inference latency, demonstrating the superiority of IFE instruction-tuning in mitigating knowledge conflicts.
Xianjie Mo, Yang Xiang 0003, Youcheng Pan, Yongshuai Hou, Ping Luo 0001
IJCNN3
2024 MGCoT: Multi-Grained Contextual Transformer for table-based text generation
Xianjie Mo, Yang Xiang 0003, Youcheng Pan, Yongshuai Hou, Ping Luo 0001
Expert Syst. Appl.3
2024 Confounder balancing in adversarial domain adaptation for pre-trained large models fine-tuning
abstract
The excellent generalization, contextual learning, and emergence abilities in the pre-trained large models (PLMs) handle specific tasks without direct training data, making them the better foundation models in the adversarial domain adaptation (ADA) methods to transfer knowledge learned from the source domain to target domains. However, existing ADA methods fail to account for the confounder properly, which is the root cause of the source data distribution that differs from the target domains. This study proposes a confounder balancing method in adversarial domain adaptation for PLMs fine-tuning (CadaFT), which includes a PLM as the foundation model for a feature extractor, a domain classifier and a confounder classifier, and they are jointly trained with an adversarial loss. This loss is designed to improve the domain-invariant representation learning by diluting the discrimination in the domain classifier. At the same time, the adversarial loss also balances the confounder distribution among source and unmeasured domains in training. Compared to newest ADA methods, CadaFT can correctly identify confounders in domain-invariant features, thereby eliminating the confounder biases in the extracted features from PLMs. The confounder classifier in CadaFT is designed as a plug-and-play and can be applied in the confounder measurable, unmeasurable, or partially measurable environments. Empirical results on natural language processing and computer vision downstream tasks show that CadaFT outperforms the newest GPT-4, LLaMA2, ViT and ADA methods.
Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001, Yukang Lin
Neural Networks4
2024 BaSFormer: A Balanced Sparsity Regularized Attention Network for Transformer
abstract
Attention networks often make decisions relying solely on a few pieces of tokens, even if those reliances are not truly indicative of the underlying meaning or intention of the full context. This can lead to over-fitting in transformers and hinder their ability to generalize. Attention regularization and sparsity-based methods have been used to overcome this issue. However, these methods cannot guarantee that all tokens have sufficient receptive fields for global information inference. Thus, the impact of individual biases cannot be effectively reduced. As a result, the generalization of these approaches improved slightly from the training data to new data. To address these limitations, we propose a balanced sparsity (BaS) regularized attention network on top of the transformers, called BaSFormer. BaS regularization introduces the K-regular graph constraint on self-attention connections, which replaces SoftMax with SparseMax in the attention transformation. In BaS-regularized self-attention, SparseMax assigns zero attention scores to low-scoring connections, highlighting influential and meaningful contexts. The K-regular graph constraint ensures that all tokens have an equal-sized receptive field to aggregate information, which facilitates the involvement of global tokens in the feature update of each layer and reduces the impact of individual biases. Given that there is no continuous loss can be used for the K-regular graph regularization, we propose an exponential extremum loss with an augmented Lagrangian function. The experimental results showed that BaSFormer improved the effectiveness of debiasing compared to that of the newest LLMs, such as the GPT-3.5, GPT-4 and LLaMA. In addition, BaSFormer achieves new state-of-the-art (SOTA) results in text generation tasks. Interestingly, this work also shows that BaSFormer can learn hierarchical linguistic dependencies in gradient attributions, which improves interpretability and adversarial robustness.
Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Learning to Improve Out-of-Distribution Generalization via Self-Adaptive Language Masking
abstract
Although the pre-trained Transformers learned general linguistic knowledge from large-scale corpus, they still over-fit on the lexical biases when fine-tuning on specific datasets. This problem limits the generalizability of pre-trained models, particularly when learning over out-of-distribution (OOD) data. To address this issue, this paper proposes a self-adaptive language masking (AdaLMask) paradigm to fine-tune the pre-trained Transformers. AdaLMask obviates lexical biases by eliminating the dependence on semantically inessential words. Specifically, AdaLMask learns a Gumbel-Softmax distribution to determine the desired masking positions, and the distribution parameters are optimized via a representation-invariant (RInv) objective to ensure the masked positions are semantically lossless. Four natural language processing tasks are chosen to evaluate the effectiveness of the proposed method on the robustness of lexical biases and OOD generalization. All empirical results demonstrate that the AdaLMask paradigm substantially improves the OOD generalization of pre-trained Transformers.
Shuoran Jiang, Youcheng Pan, Qingcai Chen, Yang Xiang 0003, Xiangping Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Learning to generate complex question with intent prediction from long passage
Youcheng Pan, Baotian Hu, Shiyue Wang, Xiaolong Wang 0001, Qingcai Chen, Zenglin Xu, Min Zhang 0005
Appl. Intell.1
2022 Multi-Role Event Argument Extraction as Machine Reading Comprehension with Argument Match Optimization
abstract
Extracting arguments for the pre-defined roles is a crucial step for event extraction. Recently, there are some insightful works that view it as a machine reading comprehension problem and achieve significant progress. However, most of them need multi-turns to extract the arguments of each role independently, which ignores the relationships among roles in the same event. To alleviate this problem, we propose a novel Multi-Role Argument Extraction method named MRAE which can exploit the relationship of event roles by extracting all arguments for an event simultaneously. To force MRAE to locate more arguments accurately, we propose an argument match optimization loss based on the minimum risk training to exploit sentence-level F1 score. We conduct experiments on the widely used ACE2005 dataset. The experimental results demonstrate that MRAE outperforms the competitor methods by at least +1.2% F1 score on argument extraction, and also shows superiority on data scarce scenarios.
Jingcong Tao, Youcheng Pan, Baotian Hu, Weihua Peng, Cuiyun Han, Xiaolong Wang 0001
ICASSP2
2021 Enriching BERT With Knowledge Graph Embedding For Industry Classification
Shiyue Wang, Youcheng Pan, Zhenran Xu, Baotian Hu, Xiaolong Wang 0001
ICONIP (6)2
2020 MedWriter: Knowledge-Aware Medical Text Generation
abstract
To exploit the domain knowledge to guarantee the correctness of generated text has been a hot topic in recent years, especially for high professional domains such as medical.However, most of recent works only consider the information of unstructured text rather than structured information of the knowledge graph.In this paper, we focus on the medical topic-to-text generation task and adapt a knowledge-aware text generation model to the medical domain, named MedWriter, which not only introduces the specific knowledge from the external MKG but also is capable of learning graph-level representation.We conduct experiments on a medical literature dataset collected from medical journals, each of which has a set of topic words, an abstract of medical literature and a corresponding knowledge graph from CMeKG.Experimental results demonstrate incorporating knowledge graph into generation model can improve the quality of the generated text and has robust superiority over the competitor methods.
Youcheng Pan, Qingcai Chen, Weihua Peng, Xiaolong Wang 0001, Baotian Hu, Xin Liu 0054, Wenxiu Zhou
COLING1
2020 Learning to Generate Diverse Questions from Keywords
abstract
Diverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., generating a series of questions for each set of keywords (one-to-many). As an effort towards this, we propose a novel neural generative model, which incorporates context information and control signal to produce multiple diverse questions from a given fixed set of keywords. The control signal is designed to increase the diversity of questions by capturing the diverse patterns from the entire dataset. The context information is used to guarantee the generated questions are highly related to the given keywords. To evaluate the effectiveness of the proposed model, we collect a dataset which contains 62835 questions with respect to 12567 sets of keywords.1To the best of our knowledge, it's the first Chinese financial dataset for diverse question generation. The experimental results show that our model outperforms the competitor methods in terms of BLEU and Distinct. The qualitative evaluation indicates that our model is able to generate diverse and meaningful questions.
Youcheng Pan, Baotian Hu, Qingcai Chen, Yang Xiang 0003, Xiaolong Wang 0001
ICASSP1
2017 Answer Selection in Community Question Answering by Normalizing Support Answers
Zhihui Zheng, Daohe Lu, Qingcai Chen, Yang Xiang 0003, Youcheng Pan
NLPCC6