EDBT 2026 Demo / reviewers in the wild / expert
Jianhao Yan
dblp:242/4255
· DBLP profile ↗
16ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Heterogeneous Text Style Control Using PromptsabstractAdvancements in natural language processing (NLP) have markedly improved paraphrase generation, an essential task for numerous applications. However, current methods face limitations due to model and constraint specificity, which hinder their flexibility and practical deployment. In this work, we introduce a unified prompt-driven approach to paraphrase generation that leverages diverse prompts, enabling fine-grained user control over aspects such as syntax and sentiment. Moreover, we incorporate translation to enable sophisticated cross-lingual text controls. Our system employs a data-centric paradigm which organizes prompts with natural language instructions. The proposed method is compatible with various sequence-to-sequence architectures and utilizes a novel training strategy to address the versatility of prompt combinations. Empirical results show that our approach not only demonstrates its capacity to adhere to multiple user-defined constraints but also maintains high performance in generation tasks without prompts. Moreover, extensive analysis shows that the model exhibits robustness to prompt variance such as language and quantity. Yafu Li, Jiahao Gai, Yongjing Yin, Jianhao Yan, Yue Zhang 0004 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2026 | Benchmarking LLMs Against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise LevelsabstractThis study presents a comprehensive evaluation of the translation capabilities of existing LLMs, such as GPT-4, ALMA-R, and Deepseek-R1, compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translations across three language pairs (Chinese$\longleftrightarrow$English, Russian$\longleftrightarrow$English, and Chinese$\longleftrightarrow$Hindi) and three domains (News, Technology, and Biomedical). Our findings reveal that LLMs achieve performance comparable to junior-level translators in terms of total errors, while still lagging behind senior translators. Unlike traditional Neural Machine Translation systems, which show significant performance degradation in resource-poor language directions, LLMs like GPT-4 maintain consistent translation quality across all evaluated language pairs. Through qualitative analysis, we identify distinctive patterns in translation approaches: GPT-4 tends toward overly literal translations and exhibits lexical inconsistency, while human translators sometimes over-interpret context and introduce hallucinations. This study presents a systematic comparison between LLMs and human translators across different proficiency levels, providing valuable insights into the current capabilities and limitations of LLM-based translation systems. Jianhao Yan, Pingchuan Yan, Yulong Chen 0001, Xianchao Zhu, Yue Zhang 0004 |
IEEE Trans. Big Data | 1 |
| 2025 | Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive ParaphrasingabstractDynamical systems theory provides a framework for analyzing iterative processes and evolution over time.Within such systems, repetitive transformations can lead to stable configurations, known as attractors, including fixed points and limit cycles.Applying this perspective to large language models (LLMs), which iteratively map input text to output text, provides a principled approach to characterizing long-term behaviors.Successive paraphrasing serves as a compelling testbed for exploring such dynamics, as paraphrases re-express the same underlying meaning with linguistic variation.Although LLMs are expected to explore a diverse set of paraphrases in the text space, our study reveals that successive paraphrasing converges to stable periodic states, such as 2period attractor cycles, limiting linguistic diversity.This phenomenon is attributed to the selfreinforcing nature of LLMs, as they iteratively favour and amplify certain textual forms over others.This pattern persists with increasing generation randomness or alternating prompts and LLMs.These findings underscore inherent constraints in LLM generative capability, while offering a novel dynamical systems perspective for studying their expressive potential.Our code is available here. Zhilin Wang, Yafu Li, Jianhao Yan, Yu Cheng 0001, Yue Zhang 0004 |
ACL (1) | 3 |
| 2025 | Keys to Robust Edits: From Theoretical Insights to Practical AdvancesabstractLarge language models (LLMs) struggle with maintaining accurate knowledge due to conflicting/outdated parametric memories.While locate-and-edit methods address this, their reliance on models' internal representations leads to robustness failures in long-context reasoning and paraphrased queries.We identify a fundamental limitation of locate-and-edit methods: existing semantic keys (for memory localization) cannot simultaneously satisfy robustness (context-invariant activation) and specificity (precise knowledge discrimination).Through theoretical error-bound analysis, we establish formal criteria for effective editing.Our solution introduces Robust Edit Pathway (REP), a plug-and-play module that: (1) disentangles editing keys from native model representations; (2) dynamically adjusts keys via contrastive learning to achieve robustness-specificity balance.Extensive experiments across various editing methods (ROME/MEMIT/R-ROME/EMMET), existing LLMs (LLaMA2, QWen, Mistral), and datasets (CounterFact, ZsRE) show that REP improves success rate over robustness tests by up-to 66.4% while maintaining the success rate unaffected. Jianhao Yan, Futing Wang, Yafu Li |
ACL (1) | 1 |
| 2025 | Dynamics of Instruction Fine-Tuning for Chinese Large Language ModelsabstractInstruction tuning is a burgeoning method to elicit the general intelligence of Large Language Models (LLMs). While numerous studies have examined the impact of factors such as data volume and model size on English models, the scaling properties of instruction tuning in other languages remain largely unexplored. In this work, we systematically investigate the effects of data quantity, model size, and data construction methods on instruction tuning for Chinese LLMs. We utilize a newly curated dataset, DoIT, which includes over 40,000 high-quality instruction instances covering ten underlying abilities, such as creative writing, code generation, and logical reasoning. Our experiments, conducted on models ranging from 7b to 33b parameters, yield three key findings: (i) While these factors directly affect overall model performance, some abilities are more responsive to scaling, whereas others demonstrate significant resistance. (ii) The scaling sensitivity of different abilities to these factors can be explained by two features: Complexity and Transference. (iii) By tailoring training strategies to their varying sensitivities, specific abilities can be efficiently learned, enhancing performance on two public benchmarks. Chiyu Song, Zhanchao Zhou, Jianhao Yan, Yuejiao Fei, Zhen-Zhong Lan, Yue Zhang 0004 |
COLING | 3 |
| 2025 | ELICIT: LLM Augmentation Via External In-context CapabilityabstractEnhancing the adaptive capabilities of large language models is a critical pursuit in both research and application.
Traditional fine-tuning methods require substantial data, computational resources, and specific capabilities, while in-context learning is limited by the need for appropriate demonstrations and efficient token usage.
Inspired by the expression of in-context learned capabilities through task vectors and the concept of modular capability or knowledge, we propose ELICIT, a framework consisting of two modules designed to effectively store and reuse task vectors to enhance the diverse adaptive capabilities of models without additional training or inference tokens.
Our comprehensive experiments and analysis demonstrate that our pipeline is highly transferable across different input formats, tasks, and model architectures.
Externally storing and reusing vectors that represent in-context learned capabilities not only shows the potential to extract modular capabilities but also significantly enhances the performance, versatility, adaptability, and scalability of large language models, paving the way for more efficient and effective use of these models in a wide range of applications. Futing Wang, Jianhao Yan, Yue Zhang 0004 |
ICLR | 2 |
| 2025 | Learning to Reason under Off-Policy GuidanceabstractRecent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning with verifiable rewards~(RLVR).
However, existing RLVR approaches are inherently ``on-policy'', limiting learning to a model's own outputs and failing to acquire reasoning abilities beyond its initial capabilities.
To address this issue, we introduce LUFFY (Learning to reason Under oFF-policY guidance), a framework that augments RLVR with off-policy reasoning traces.
LUFFY dynamically balances imitation and exploration by combining off-policy demonstrations with on-policy rollouts during training.
Specifically, LUFFY combines the Mixed-Policy GRPO framework, which has a theoretically guaranteed convergence rate, alongside policy shaping via regularized importance sampling to avoid superficial and rigid imitation during mixed-policy training.
Compared with previous RLVR methods, LUFFY achieves an over +6.4 average gain across six math benchmarks and an advantage of over +6.2 points in out-of-distribution tasks.
Most significantly, we show that LUFFY successfully trains weak models in scenarios where on-policy RLVR completely fails. These results provide compelling evidence that LUFFY transcends the fundamental limitations of on-policy RLVR and demonstrates the great potential of utilizing off-policy guidance in RLVR. Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang 0001, Ganqu Cui, Xiaoye Qu, Yu Cheng 0001, Yue Zhang 0004 |
NeurIPS | 1 |
| 2024 | DC-MBR: Distributional Cooling for Minimum Bayesian Risk DecodingabstractMinimum Bayesian Risk Decoding (MBR) emerges as a promising decoding algorithm in Neural Machine Translation. However, MBR performs poorly with label smoothing, which is surprising as label smoothing provides decent improvement with beam search and improves generality in various tasks. In this work, we show that the issue arises from the inconsistency of label smoothing on the token-level and sequence-level distributions. We demonstrate that even though label smoothing only causes a slight change in the token level, the sequence-level distribution is highly skewed. We coin the issue autoregressive over-smoothness. To address this issue, we propose a simple and effective method, Distributional Cooling MBR (DC-MBR), which manipulates the entropy of output distributions by tuning down the Softmax temperature. We theoretically prove the equivalence between the pre-tuning label smoothing factor and distributional cooling. Extensive experiments on NMT benchmarks validate that distributional cooling improves MBR in various settings. Jianhao Yan, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004 |
LREC/COLING | 1 |
| 2024 | Understanding In-Context Learning from RepetitionsabstractThis paper explores the elusive mechanism underpinning in-context learning in Large Language Models (LLMs). Our work provides a novel perspective by examining in-context learning via the lens of surface repetitions. We quantitatively investigate the role of surface features in text generation, and empirically establish the existence of \emph{token co-occurrence reinforcement}, a principle that strengthens the relationship between two tokens based on their contextual co-occurrences. Furthermore, we find similar reinforcements lie behind the pretraining corpus, revealing the existence is due to LLMs' efforts to maximize the likelihood. By investigating the dual impacts of these features, our research illuminates the internal workings of in-context learning and expounds on the reasons for its failures. This paper provides an essential contribution to the understanding of in-context learning and its potential limitations, providing a fresh perspective on this exciting capability. Jianhao Yan, Chiyu Song, Chenming Wu, Yafu Li, Yue Zhang 0004 |
ICLR | 1 |
| 2023 | Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved AnnotationabstractYulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai, Yueguan Wang, Ming Zhong, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yulong Chen 0001, Xuefeng Bai 0001, Yueguan Wang, Ming Zhong 0005, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang 0004 |
ACL (1) | 7 |
| 2023 | Explicit Syntactic Guidance for Neural Text GenerationabstractMost existing text generation models follow the sequence-to-sequence paradigm.Generative Grammar suggests that humans generate natural language texts by learning language grammar.We propose a syntax-guided generation schema, which generates the sequence guided by a constituency parse tree in a topdown direction.The decoding process can be decomposed into two parts: (1) predicting the infilling texts for each constituent in the lexicalized syntax context given the source sentence;(2) mapping and expanding each constituent to construct the next-level syntax context.Accordingly, we propose a structural beam search method to find possible syntax structures hierarchically.Experiments on paraphrase generation and machine translation show that the proposed method outperforms autoregressive baselines, while also demonstrating effectiveness in terms of interpretability, controllability, and diversity. Yafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin, Wei Bi, Shuming Shi 0001, Yue Zhang 0004 |
ACL (1) | 3 |
| 2022 | Digging Errors in NMT: Evaluating and Understanding Model Errors from Partial Hypothesis SpaceabstractSolid evaluation of neural machine translation (NMT) is key to its understanding and improvement.Current evaluation of an NMT system is usually built upon a heuristic decoding algorithm (e.g., beam search) and an evaluation metric assessing similarity between the translation and golden reference.However, this system-* Equal contribution. Jianhao Yan, Chenming Wu, Fandong Meng, Jie Zhou 0016 |
EMNLP | 1 |
| 2022 | Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationabstractWhile large-scale neural language models, such as GPT2 and BART,have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e.g.}, greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in the human corpus (e.g., 0.02\% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probability of repetitive tokens and their previous repetitions in context. Through our quantitative experiments, we find that 1) Models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a \textit{self-reinforcement effect}: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method \textbf{DITTO} (Pseu\underline{D}o-Repet\underline{IT}ion Penaliza\underline{T}i\underline{O}n), where the model learns to penalize probabilities of sentence-level repetitions from synthetic repetitive data. Although our method is motivated by mitigating repetitions, our experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method. Jin Xu 0010, Xiaojiang Liu, Jianhao Yan, Deng Cai 0002, Jian Li 0015 |
NeurIPS | 3 |
| 2022 | Affective word embedding in affective explanation generation for fine art paintings
Jianhao Yan, Wenmin Wang 0001 |
Pattern Recognit. Lett. | 1 |
| 2021 | Selective Knowledge Distillation for Neural Machine TranslationabstractFusheng Wang, Jianhao Yan, Fandong Meng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fusheng Wang 0008, Jianhao Yan, Fandong Meng, Jie Zhou 0016 |
ACL/IJCNLP (1) | 2 |
| 2020 | Multi-Unit Transformers for Neural Machine TranslationabstractTransformer models (Vaswani et al., 2017) achieve remarkable success in Neural Machine Translation.Many efforts have been devoted to deepening the Transformer by stacking several units (i.e., a combination of Multihead Attentions and FFN) in a cascade, while the investigation over multiple parallel units draws little attention.In this paper, we propose the Multi-Unit TransformErs (MUTE), which aim to promote the expressiveness of the Transformer by introducing diverse and complementary units.Specifically, we use several parallel units and show that modeling with multiple units improves model performance and introduces diversity.Further, to better leverage the advantage of the multi-unit setting, we design biased module and sequential dependency that guide and encourage complementariness among different units.Experimental results on three machine translation tasks, the NIST Chinese-to-English, WMT'14 English-to-German and WMT'18 Chinese-to-English, show that the MUTE models significantly outperform the Transformer-Base, by up to +1.52, +1.90 and +1.10 BLEU points, with only a mild drop in inference speed (about 3.1%).In addition, our methods also surpass the Transformer-Big model, with only 54% of its parameters.These results demonstrate the effectiveness of the MUTE, as well as its efficiency in both the inference process and parameter usage. 1 Jianhao Yan, Fandong Meng, Jie Zhou 0016 |
EMNLP (1) | 1 |