Tingchen Fu

dblp:318/0986 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-3692-729XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
abstract
Instruction-following is essential for aligning large language models (LLMs) with user intent.Yet recent reasoning-oriented models, despite their strong performance on complex mathematical problems, often fail to comply with simple natural language directives.In this work, we analyze the interaction between reasoning ability and instruction adherence in large reasoning models (LRMs).Using a controlled evaluation framework (MathIF), we uncover a persistent trade-off: as models scale reasoning capacity through long chains-of-thought or reinforcement learning on reasoning traces, their obedience to instructions degrades, particularly when generation length grows.We further show that interventions such as constraining or repeating instructions can partially restore compliance, but typically at the expense of reasoning performance.Taken together, our findings expose a dilemma between intelligence and obedience in current training paradigms and underscore the need for instruction-aware approaches to developing controllable reasoning models.
Tingchen Fu, Yafu Li, Jiawei Gu, Xiaoye Qu, Yu Cheng 0001
ACL (1)1
2025 Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
abstract
Insensitivity to semantically-preserving variations of prompts (paraphrases) is crucial for reliable behavior and real-world deployment of large language models.However, language models exhibit significant performance degradation with semantically equivalent but differently phrased prompts, and existing solutions either depend on trial-and-error prompt engineering or require computationally expensive inference-time algorithms.In this study, built on the key insight that worst-case prompts exhibit a drift in embedding space, we present Latent Adversarial Paraphrasing (LAP), a dualloop adversarial framework that optimizes a trainable perturbation as "latent continuous paraphrase" and language model performance on these perturbations iteratively.Extensive experiments are conducted to demonstrate the effectiveness of LAP across multiple backbones on the RobustAlpaca benchmark with a 0.5% ∼ 4% absolution improvement on worstcase win-rate.
Tingchen Fu, Fazl Barez
EMNLP1
2025 PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference Data
abstract
Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBench, a benchmark for evaluating large language models' susceptibility to data poisoning during preference learning. Data poisoning attacks can manipulate large language model responses to include hidden malicious content or biases, potentially causing the model to generate harmful or unintended outputs while appearing to function normally. We deploy two distinct attack types across eight realistic scenarios, assessing 22 widely-used models. Our findings reveal concerning trends: (1) Scaling up parameter size does not always enhance resilience against poisoning attacks and the influence on model resilience varies among different model suites. (2) There exists a log-linear relationship between the effects of the attack and the data poison ratio; (3) The effect of data poisoning can generalize to extrapolated triggers that are not included in the poisoned data. These results expose weaknesses in current preference learning techniques, highlighting the urgent need for more robust defenses against malicious models and data manipulation.
Tingchen Fu, Mrinank Sharma, Philip Torr 0001, Shay B. Cohen, David Krueger 0001, Fazl Barez
ICML1
2025 Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts
abstract
Tingchen Fu, Yupeng Hou, Julian McAuley, Rui Yan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tingchen Fu, Yupeng Hou, Julian J. McAuley, Rui Yan 0001
NAACL (Long Papers)1
2024 The Reasonableness Behind Unreasonable Translation Capability of Large Language Model
abstract
Multilingual large language models trained on non-parallel data yield impressive translation capabilities. Existing studies demonstrate that incidental sentence-level bilingualism within pre-training data contributes to the LLM's translation abilities. However, it has also been observed that LLM's translation capabilities persist even when incidental sentence-level bilingualism are excluded from the training corpus. In this study, we comprehensively investigate the unreasonable effectiveness and the underlying mechanism for LLM's translation abilities, specifically addressing the question why large language models learn to translate without parallel data, using the BLOOM model series as a representative example. Through extensive experiments, our findings suggest the existence of unintentional bilingualism in the pre-training corpus, especially word alignment data significantly contributes to the large language model's acquisition of translation ability. Moreover, the translation signal derived from word alignment data is comparable to that from sentence-level bilingualism. Additionally, we study the effects of monolingual data and parameter-sharing in assisting large language model to learn to translate. Together, these findings present another piece of the broader puzzle of trying to understand how large language models acquire translation capability.
Tingchen Fu, Lemao Liu, Deng Cai 0002, Guoping Huang, Shuming Shi 0001, Rui Yan 0001
ICLR1
2023 On the Compositional Generalization in Versatile Open-domain Dialogue
abstract
Previous research has demonstrated the potential of multi-task learning to foster a conversational agent's ability to acquire a variety of skills.However, these approaches either suffer from interference among different datasets (also known as negative transfer), or fail to effectively reuse knowledge and skills learned from other datasets.In contrast to previous works, we develop a sparsely activated modular network: (1) We propose a wellrounded set of operators and instantiate each operator with an independent module; (2) We formulate dialogue generation as the execution of a generated programme which recursively composes and assembles modules.Extensive experiments on 9 datasets verify the efficacy of our methods through automatic evaluation and human evaluation.Notably, our model outperforms state-of-the-art supervised approaches on 4 datasets with only 10% training data thanks to the modular architecture and multi-task learning.1 † Tingchen Fu and Xueliang Zhao contribute equally to this work.
Tingchen Fu, Xueliang Zhao, Lemao Liu, Rui Yan 0001
ACL (1)1
2023 Delving into Global Dialogue Structures: Structure Planning Augmented Response Selection for Multi-turn Conversations
abstract
Retrieval-based dialogue systems are a crucial component of natural language processing, employing information retrieval techniques to select responses from a predefined pool of candidates. The advent of pre-trained language models (PLMs) has significantly advanced the field, with a prevailing paradigm that involves post-training PLMs on specific dialogue corpora, followed by fine-tuning for the response selection (RS) task. This post-training process aims to capture dialogue-specific features, as most PLMs are originally trained on plain text. However, prior approaches predominantly rely on self-supervised tasks or session-level graph neural networks during post-training, focusing on capturing underlying patterns of coherent dialogues without explicitly refining the global pattern across the entire dialogue corpus. Consequently, the learned knowledge for organizing coherent dialogues remains isolated, heavily reliant on specific contexts. Additionally, interpreting or visualizing the implicit knowledge acquired through self-supervised tasks proves challenging. In this study, we address these limitations by explicitly refining the knowledge required for response selection and structuring it into a coherent global flow, known as "dialogue structure." This structure captures the inter-dependency of utterances and topic shifts, thereby enhancing the response selection task. To achieve this, we propose a novel structure model comprising a state recognizer and a structure planner. This model effectively captures the flow within the utterance history and plans the trajectory of future utterances. Importantly, the structure model operates orthogonally to the retrieval model, enabling seamless integration with existing retrieval models and facilitating collaborative training. Extensive experiments conducted on three benchmark datasets demonstrate the superior performance of our method over a wide range of competitive baselines, establishing a new state-of-the-art in the field.
Tingchen Fu, Xueliang Zhao, Rui Yan 0001
KDD1
2022 There Are a Thousand Hamlets in a Thousand People's Eyes: Enhancing Knowledge-grounded Dialogue with Personal Memory
abstract
Knowledge-grounded conversation (KGC) shows great potential in building an engaging and knowledgeable chatbot, and knowledge selection is a key ingredient in it.However, previous methods for knowledge selection only concentrate on the relevance between knowledge and dialogue context, ignoring the fact that age, hobby, education and life experience of an interlocutor have a major effect on his or her personal preference over external knowledge.Without taking the personalization issue into account, it is difficult to select the proper knowledge and generate persona-consistent responses.In this work, we introduce personal memory into knowledge selection in KGC to address the personalization issue.We propose a variational method to model the underlying relationship between one's personal memory and his or her selection of knowledge, and devise a learning scheme in which the forward mapping from personal memory to knowledge and its inverse mapping is included in a closed loop so that they could teach each other.Experiment results show that our method outperforms existing KGC methods significantly on both automatic evaluation and human evaluation.
Tingchen Fu, Xueliang Zhao, Chongyang Tao, Ji-Rong Wen, Rui Yan 0001
ACL (1)1
2022 There Is No Standard Answer: Knowledge-Grounded Dialogue Generation with Adversarial Activated Multi-Reference Learning
abstract
Knowledge-grounded conversation (KGC) shows excellent potential to deliver an engaging and informative response.However, existing approaches emphasize selecting one golden knowledge given a particular dialogue context, overlooking the one-to-many phenomenon in dialogue.As a result, the existing paradigm limits the diversity of knowledge selection and generation.To this end, we establish a multireference KGC dataset and propose a series of metrics to systematically assess the one-tomany efficacy of existing KGC models.Furthermore, to extend the hypothesis space of knowledge selection to enhance the mapping relationship between multiple knowledge and multiple responses, we devise a span-based variational model and optimize the model in a wake-sleep style with an ameliorated evidence lower bound objective to learn the oneto-many generalization.Both automatic and human evaluations demonstrate the efficacy of our approach.
Xueliang Zhao, Tingchen Fu, Chongyang Tao, Rui Yan 0001
EMNLP2
2022 Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent Structure
abstract
With the availability of massive generaldomain dialogue data, pre-trained dialogue generation appears to be super appealing to transfer knowledge from the general domain to downstream applications.In most existing work, such transferable ability is mainly obtained by fitting a large model with hundreds of millions of parameters on massive data in an exhaustive way, leading to inefficient running and poor interpretability.This paper proposes a novel dialogue generation model with a latent structure that is easily transferable from the general domain to downstream tasks in a lightweight and transparent way.Experiments on two benchmarks validate the effectiveness of the proposed model.Thanks to the transferable latent structure, our model is able to yield better dialogue responses than four strong baselines in terms of both automatic and human evaluations, and our model with about 22% parameters particularly delivers a 5x speedup in running time compared with the strongest baseline.Moreover, the proposed model is explainable by interpreting the discrete latent variables.
Xueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001
EMNLP3
2022 Learning to Express in Knowledge-Grounded Conversation
abstract
Xueliang Zhao, Tingchen Fu, Chongyang Tao, Wei Wu, Dongyan Zhao, Rui Yan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xueliang Zhao, Tingchen Fu, Chongyang Tao, Wei Wu 0014, Dongyan Zhao 0001, Rui Yan 0001
NAACL-HLT2