Yafu Li

dblp:293/9896 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-7895-9997ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 18 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
abstract
Instruction-following is essential for aligning large language models (LLMs) with user intent.Yet recent reasoning-oriented models, despite their strong performance on complex mathematical problems, often fail to comply with simple natural language directives.In this work, we analyze the interaction between reasoning ability and instruction adherence in large reasoning models (LRMs).Using a controlled evaluation framework (MathIF), we uncover a persistent trade-off: as models scale reasoning capacity through long chains-of-thought or reinforcement learning on reasoning traces, their obedience to instructions degrades, particularly when generation length grows.We further show that interventions such as constraining or repeating instructions can partially restore compliance, but typically at the expense of reasoning performance.Taken together, our findings expose a dilemma between intelligence and obedience in current training paradigms and underscore the need for instruction-aware approaches to developing controllable reasoning models.
Tingchen Fu, Yafu Li, Jiawei Gu, Xiaoye Qu, Yu Cheng 0001
ACL (1)2
2026 Heterogeneous Text Style Control Using Prompts
abstract
Advancements in natural language processing (NLP) have markedly improved paraphrase generation, an essential task for numerous applications. However, current methods face limitations due to model and constraint specificity, which hinder their flexibility and practical deployment. In this work, we introduce a unified prompt-driven approach to paraphrase generation that leverages diverse prompts, enabling fine-grained user control over aspects such as syntax and sentiment. Moreover, we incorporate translation to enable sophisticated cross-lingual text controls. Our system employs a data-centric paradigm which organizes prompts with natural language instructions. The proposed method is compatible with various sequence-to-sequence architectures and utilizes a novel training strategy to address the versatility of prompt combinations. Empirical results show that our approach not only demonstrates its capacity to adhere to multiple user-defined constraints but also maintains high performance in generation tasks without prompts. Moreover, extensive analysis shows that the model exhibits robustness to prompt variance such as language and quantity.
Yafu Li, Jiahao Gai, Yongjing Yin, Jianhao Yan, Yue Zhang 0004
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2025 Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
abstract
Large language models (LLMs) have achieved remarkable success in machine translation, demonstrating impressive performance across diverse languages. However, translationese—characterized by overly literal and unnatural translations—remains a persistent challenge in LLM-based translation systems. Despite their pre-training on vast corpora of natural utterances, LLMs exhibit translationese errors and generate unexpected unnatural translations, stemming from biases introduced during supervised fine-tuning (SFT). In this work, we systematically evaluate the prevalence of translationese in LLM-generated translations and investigate its roots during supervised training. We introduce methods to mitigate these biases, including polishing golden references and filtering unnatural training instances. Empirical evaluations demonstrate that these approaches significantly reduce translationese while improving translation naturalness, validated by human evaluations and automatic metrics. Our findings highlight the need for training-aware adjustments to optimize LLM translation outputs, paving the way for more fluent and target-language-consistent translations.
Yafu Li, Ronghao Zhang, Zhilin Wang, Leyang Cui, Yongjing Yin, Tong Xiao 0001, Yue Zhang 0004
ACL (1)1
2025 Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
abstract
Dynamical systems theory provides a framework for analyzing iterative processes and evolution over time.Within such systems, repetitive transformations can lead to stable configurations, known as attractors, including fixed points and limit cycles.Applying this perspective to large language models (LLMs), which iteratively map input text to output text, provides a principled approach to characterizing long-term behaviors.Successive paraphrasing serves as a compelling testbed for exploring such dynamics, as paraphrases re-express the same underlying meaning with linguistic variation.Although LLMs are expected to explore a diverse set of paraphrases in the text space, our study reveals that successive paraphrasing converges to stable periodic states, such as 2period attractor cycles, limiting linguistic diversity.This phenomenon is attributed to the selfreinforcing nature of LLMs, as they iteratively favour and amplify certain textual forms over others.This pattern persists with increasing generation randomness or alternating prompts and LLMs.These findings underscore inherent constraints in LLM generative capability, while offering a novel dynamical systems perspective for studying their expressive potential.Our code is available here.
Zhilin Wang, Yafu Li, Jianhao Yan, Yu Cheng 0001, Yue Zhang 0004
ACL (1)2
2025 Keys to Robust Edits: From Theoretical Insights to Practical Advances
abstract
Large language models (LLMs) struggle with maintaining accurate knowledge due to conflicting/outdated parametric memories.While locate-and-edit methods address this, their reliance on models' internal representations leads to robustness failures in long-context reasoning and paraphrased queries.We identify a fundamental limitation of locate-and-edit methods: existing semantic keys (for memory localization) cannot simultaneously satisfy robustness (context-invariant activation) and specificity (precise knowledge discrimination).Through theoretical error-bound analysis, we establish formal criteria for effective editing.Our solution introduces Robust Edit Pathway (REP), a plug-and-play module that: (1) disentangles editing keys from native model representations; (2) dynamically adjusts keys via contrastive learning to achieve robustness-specificity balance.Extensive experiments across various editing methods (ROME/MEMIT/R-ROME/EMMET), existing LLMs (LLaMA2, QWen, Mistral), and datasets (CounterFact, ZsRE) show that REP improves success rate over robustness tests by up-to 66.4% while maintaining the success rate unaffected.
Jianhao Yan, Futing Wang, Yafu Li
ACL (1)4
2025 Keyphrase Generation Based on the Fusion of Sequence and Word Graph Features
Yafu Li, Wu Zhuang
ICIC (21)2
2025 Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
abstract
Large language models (LLMs) have presented impressive performance but often lack the flexibility to adapt to human preferences quickly without retraining. Inspired by the recent efforts on test-time scaling, we make the first attempt to propose Test-time Preference Optimization (TPO), a framework that aligns LLM outputs with human preferences during inference, eliminating the need to update model parameters. Instead of relying on purely numerical rewards, TPO translates reward signals into \emph{textual} critiques and uses them as textual rewards to iteratively refine its response. Evaluations on benchmarks covering instruction following, preference alignment, safety, and mathematics reveal that TPO progressively improves alignment with human preferences. Notably, after only a few TPO steps, the initially unaligned Llama-3.1-70B-SFT model can surpass the aligned counterpart, Llama-3.1-70B-Instruct. Furthermore, TPO scales efficiently with both the search width and depth of the inference process. Through case studies, we illustrate how TPO exploits the innate capacity of LLM to interpret and act upon reward signals. Our findings establish TPO as a practical, lightweight alternative for test-time preference optimization, achieving alignment on the fly.
Yafu Li, Xuyang Hu, Xiaoye Qu, Yu Cheng 0001
ICML1
2025 Learning to Reason under Off-Policy Guidance
abstract
Recent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning with verifiable rewards~(RLVR). However, existing RLVR approaches are inherently ``on-policy'', limiting learning to a model's own outputs and failing to acquire reasoning abilities beyond its initial capabilities. To address this issue, we introduce LUFFY (Learning to reason Under oFF-policY guidance), a framework that augments RLVR with off-policy reasoning traces. LUFFY dynamically balances imitation and exploration by combining off-policy demonstrations with on-policy rollouts during training. Specifically, LUFFY combines the Mixed-Policy GRPO framework, which has a theoretically guaranteed convergence rate, alongside policy shaping via regularized importance sampling to avoid superficial and rigid imitation during mixed-policy training. Compared with previous RLVR methods, LUFFY achieves an over +6.4 average gain across six math benchmarks and an advantage of over +6.2 points in out-of-distribution tasks. Most significantly, we show that LUFFY successfully trains weak models in scenarios where on-policy RLVR completely fails. These results provide compelling evidence that LUFFY transcends the fundamental limitations of on-policy RLVR and demonstrates the great potential of utilizing off-policy guidance in RLVR.
Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang 0001, Ganqu Cui, Xiaoye Qu, Yu Cheng 0001, Yue Zhang 0004
NeurIPS2
2025 MCRanker: Generating Diverse Criteria On-the-Fly to Improve Pointwise LLM Rankers
abstract
The most recent pointwise Large Language Model (LLM) rankers have achieved remarkable ranking results. However, these rankers are hindered by two major drawbacks: (1) they fail to follow a standardized comparison guidance during the ranking process, and (2) they struggle with comprehensive considerations when dealing with diverse semantics of the query and complicated info in the passages. To address these shortcomings, we propose to build a zero-shot pointwise ranker that first recruits a virtual annotation team to generate query-based criteria from various perspectives and then uses these criteria to conduct an ensemble passage evaluation. Additionally, we are among the first to explore how criteria can be generated automatically and used in text ranking tasks. Our method, tested on eight datasets from the BEIR benchmark, demonstrates that incorporating this multi-perspective criteria ensemble approach significantly enhanced the performance of pointwise LLM rankers.
Honglei Zhuang, Yafu Li, Qi Zhu 0008, Yue Zhang 0004
WSDM5
2024 MAGE: Machine-generated Text Detection in the Wild
abstract
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi 0001, Yue Zhang 0004
ACL (1)1
2024 Understanding In-Context Learning from Repetitions
abstract
This paper explores the elusive mechanism underpinning in-context learning in Large Language Models (LLMs). Our work provides a novel perspective by examining in-context learning via the lens of surface repetitions. We quantitatively investigate the role of surface features in text generation, and empirically establish the existence of \emph{token co-occurrence reinforcement}, a principle that strengthens the relationship between two tokens based on their contextual co-occurrences. Furthermore, we find similar reinforcements lie behind the pretraining corpus, revealing the existence is due to LLMs' efforts to maximize the likelihood. By investigating the dual impacts of these features, our research illuminates the internal workings of in-context learning and expounds on the reasons for its failures. This paper provides an essential contribution to the understanding of in-context learning and its potential limitations, providing a fresh perspective on this exciting capability.
Jianhao Yan, Chiyu Song, Chenming Wu, Yafu Li, Yue Zhang 0004
ICLR5
2023 Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved Annotation
abstract
Yulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai, Yueguan Wang, Ming Zhong, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulong Chen 0001, Xuefeng Bai 0001, Yueguan Wang, Ming Zhong 0005, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang 0004
ACL (1)8
2023 Explicit Syntactic Guidance for Neural Text Generation
abstract
Most existing text generation models follow the sequence-to-sequence paradigm.Generative Grammar suggests that humans generate natural language texts by learning language grammar.We propose a syntax-guided generation schema, which generates the sequence guided by a constituency parse tree in a topdown direction.The decoding process can be decomposed into two parts: (1) predicting the infilling texts for each constituent in the lexicalized syntax context given the source sentence;(2) mapping and expanding each constituent to construct the next-level syntax context.Accordingly, we propose a structural beam search method to find possible syntax structures hierarchically.Experiments on paraphrase generation and machine translation show that the proposed method outperforms autoregressive baselines, while also demonstrating effectiveness in terms of interpretability, controllability, and diversity.
Yafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin, Wei Bi, Shuming Shi 0001, Yue Zhang 0004
ACL (1)1
2023 Consistency Regularization Training for Compositional Generalization
abstract
Existing neural models have difficulty generalizing to unseen combinations of seen components.To achieve compositional generalization, models are required to consistently interpret (sub)expressions across contexts.Without modifying model architectures, we improve the capability of Transformer on compositional generalization through consistency regularization training, which promotes representation consistency across samples and prediction consistency for a single sample.Experimental results on semantic parsing and machine translation benchmarks empirically demonstrate the effectiveness and generality of our method.In addition, we find that the prediction consistency scores on in-distribution validation sets can be an alternative for evaluating models during training, when commonly-used metrics are not informative.
Yongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004
ACL (1)3
2022 Categorizing Semantic Representations for Neural Machine Translation
abstract
Modern neural machine translation (NMT) models have achieved competitive performance in standard benchmarks. However, they have recently been shown to suffer limitation in compositional generalization, failing to effectively learn the translation of atoms (e.g., words) and their semantic composition (e.g., modification) from seen compounds (e.g., phrases), and thus suffering from significantly weakened translation performance on unseen compounds during inference. We address this issue by introducing categorization to the source contextualized representations. The main idea is to enhance generalization by reducing sparsity and overfitting, which is achieved by finding prototypes of token representations over the training set and integrating their embeddings into the source encoding. Experiments on a dedicated MT dataset (i.e., CoGnition) show that our method reduces compositional generalization error rates by 24% error reduction. In addition, our conceptually simple method gives consistently better results than the Transformer baseline on a range of general MT datasets.
Yongjing Yin, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004
COLING2
2022 Multi-Granularity Optimization for Non-Autoregressive Translation
abstract
Despite low latency, non-autoregressive machine translation (NAT) suffers severe performance deterioration due to the naive independence assumption.This assumption is further strengthened by cross-entropy loss, which encourages a strict match between the hypothesis and the reference token by token.To alleviate this issue, we propose multi-granularity optimization for NAT, which collects model behaviors on translation segments of various granularities and integrates feedback for backpropagation.Experiments on four WMT benchmarks show that the proposed method significantly outperforms the baseline models trained with cross-entropy loss, and achieves the best performance on WMT'16 En⇔Ro and highly competitive results on WMT'14 En⇔De for fully non-autoregressive translation.
Yafu Li, Leyang Cui, Yongjing Yin, Yue Zhang 0004
EMNLP1
2022 Label Attention Network for Structured Prediction
abstract
Sequence labeling assigns a label to each token in a sequence, which is a fundamental problem in natural language processing (NLP). Many NLP tasks, including part-of-speech tagging and named entity recognition, can be solved in a form of sequence labeling problem. Other tasks such as constituency parsing and non-autoregressive machine translation can also be transformed into sequence labeling tasks. Neural models have been shown powerful for sequence labeling by employing a multi-layer sequence encoding network. Conditional random field (CRF) is proposed to enrich information over label sequences, yet it suffers large computational complexity and over-reliance on Marko assumption. To this end, we propose label attention network (LAN) to hierarchically refine representation of marginal label distributions bottom-up, enabling higher layers to learn more informed label sequence distribution based on information from lower layers. We demonstrate the effectiveness of LAN through extensive experiments on various NLP tasks including POS tagging, NER, CCG supertagging, constituency parsing and non-autoregressive machine translation. Empirical results show that LAN not only improves the overall tagging accuracy with similar number of parameters, but also significantly speeds up the training and testing compared to CRF.
Leyang Cui, Yafu Li, Yue Zhang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 On Compositional Generalization of Neural Machine Translation
abstract
Yafu Li, Yongjing Yin, Yulong Chen, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yafu Li, Yongjing Yin, Yulong Chen 0001, Yue Zhang 0004
ACL/IJCNLP (1)1
2021 Sentence-State LSTMs For Sequence-to-Sequence Learning
Xuefeng Bai 0001, Yafu Li, Zhirui Zhang, Mingzhou Xu, Boxing Chen, Weihua Luo, Derek F. Wong, Yue Zhang 0004
NLPCC (1)2