Chak Tou Leong

dblp:358/9146 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-6124-1890ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
abstract
Large reasoning models (LRMs) have demonstrated remarkable proficiency in tackling complex tasks through step-by-step thinking.However, this lengthy reasoning process incurs substantial computational and latency overheads, hindering the practical deployment of LRMs.This work presents a new approach to mitigating overthinking in LRMs via black-box persuasive prompting.By treating LRMs as black-box communicators, we investigate how to persuade them to generate concise responses without compromising accuracy.We introduce WHISPER, an iterative refinement framework that generates high-quality persuasive prompts from diverse perspectives.Experiments across multiple benchmarks demonstrate that WHIS-PER consistently reduces token usage while preserving performance.Notably, WHISPER achieves a 3× reduction in average response length on simple GSM8K questions for the Qwen3 series and delivers an average ∼40% token reduction overall.For closed-source APIs, WHISPER reduces token usage on MATH-500 by 46% for Claude-3.7 and 50% for Gemini-2.5.Further analysis reveals the broad applicability of WHISPER across data domains, model scales, and families, underscoring the potential of black-box persuasive prompting as a practical strategy for enhancing LRM efficiency. 1 * Work done during the author's internship at Sea AI Lab.† Corresponding authors. 1 We release our code and top-performing prompt candidates at https://github.com/hemingkx/Whisper.Question: John arm wrestles 20 people.He beats 80%.How many people did he lose to?
Heming Xia, Cunxiao Du, Rui Li 0094, Chak Tou Leong, Yongqi Li 0001, Wenjie Li 0002
ACL (1)4
2025 Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
abstract
Large Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised text (rewriting, expansion, and polishing), which can have only minor changes from its original human prompt. As the content of text may originate from human prompts, detecting machine-revised text often involves identifying distinctive machine styles, e.g., worded favored by LLMs. However, existing methods struggle to detect machine-style phrasing hidden within the content contributed by humans. We propose the “Imitate Before Detect” (ImBD) approach, which first imitates the machine-style token distribution, and then compares the distribution of the text to be tested with the machine-style distribution to determine whether the text has been machine-revised. To this end, we introduce Style Preference Optimization (SPO), which aligns a scoring LLM model to the preference of text styles generated by machines. The aligned scoring model is then used to calculate the style-conditional probability curvature (Style-CPC), quantifying the log probability difference between the original and conditionally sampled texts for effective detection. We conduct extensive comparisons across various scenarios, encompassing text revisions by six LLMs, four distinct text domains, and three machine revision types. Compared to existing state-of-the-art methods, our method yields a 13% increase in AUC for detecting text revised by open-source LLMs, and improves performance by 5% and 19% for detecting GPT-3.5 and GPT-4o revised text, respectively. Notably, our method surpasses the commercially trained GPT-Zero with just 1,000 samples and five minutes of SPO, demonstrating its efficiency and effectiveness.
Xiaoye Zhu, Yiwen Yuan, Chak Tou Leong, Zuchao Li, Tang Long, Chenyu Yan, Guanghao Mei, Lefei Zhang
AAAI7
2025 Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
abstract
The safety alignment of large language models (LLMs) remains vulnerable, as their initial behavior can be easily jailbroken by even relatively simple attacks.Since infilling a fixed template between the input instruction and initial model output is a common practice for existing LLMs, we hypothesize that this template is a key factor behind their vulnerabilities: LLMs' safety-related decision-making overly relies on the aggregated information from the template region, which largely influences these models' safety behavior.We refer to this issue as template-anchored safety alignment.In this paper, we conduct extensive experiments and verify that template-anchored safety alignment is widespread across various aligned LLMs.Our mechanistic analyses demonstrate how it leads to models' susceptibility when encountering inference-time jailbreak attacks.Furthermore, we show that detaching safety mechanisms from the template region is promising in mitigating vulnerabilities to jailbreak attacks.We encourage future research to develop more robust safety alignment techniques that reduce reliance on the template region.
Chak Tou Leong, Qingyu Yin, Jian Wang 0054, Wenjie Li 0002
ACL (1)1
2025 Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
abstract
Kaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002
ACL (1)5
2025 Symbolic Representation for Any-to-Any Generative Tasks
abstract
We propose a symbolic generative task description language and a corresponding inference engine capable of representing arbitrary multimodal tasks as structured symbolic flows. Unlike conventional generative models that rely on large-scale training and implicit neural representations to learn cross-modal mappings—often at high computational cost and with limited flexibility—our framework introduces an explicit symbolic representation comprising three core primitives: $\color{blue}{\text{functions}}$, $\color{green}{\text{parameters}}$, and $\color {Purple}{\text{topological}}\,{\text{logic}}$. Leveraging a pre-trained language model, our inference engine maps natural language instructions directly to symbolic workflows in a training-free manner. Our framework successfully performs over 12 diverse multimodal generative tasks, demonstrating strong performance and flexibility without the need for task-specific tuning. Experiments show that our method not only matches or outperforms existing state-of-the-art unified models in content quality, but also offers greater efficiency, editability, and interruptibility. We believe that symbolic task representations provide a cost-effective and extensible foundation for advancing the capabilities of generative AI.
Xiaoye Zhu, Tianyang Liu 0003, Chak Tou Leong, Yifei Ke, Yiwen Yuan, Julian J. McAuley, Li-jia Li
CVPR7
2025 Video-Bench: Human-Aligned Video Generation Benchmark
abstract
Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories: traditional benchmarks, which use metrics and embeddings to evaluate generated video quality across multiple dimensions but often lack alignment with human judgments; and large language model (LLM)-based benchmarks, though capable of human-like reasoning, are constrained by a limited understanding of video quality metrics and cross-modal consistency. To address these challenges and establish a benchmark that better aligns with human preferences, this paper introduces Video-Bench, a comprehensive benchmark featuring a rich prompt suite and extensive evaluation dimensions. This benchmark represents the first attempt to systematically leverage MLLMs across all dimensions relevant to video generation assessment in generative models. By incorporating few-shot scoring and chain-of-query techniques, Video-Bench provides a structured, scalable approach to generated video evaluation. Experiments on advanced models including Sora demonstrate that Video-bench achieve superior alignment with human preferences across all dimensions. Moreover, in instances where our framework’s assessments diverge from human evaluations, it consistently offers more objective and accurate insights, suggesting an even greater potential advantage over traditional human judgment.
Yiwen Yuan, Yuling Wu, Yufan Deng, Chak Tou Leong, Hanwen Du, Junchen Fu, Youhua Li, Chi Zhang 0007, Li-jia Li, Yongxin Ni
CVPR7
2025 Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
abstract
Dingwei Chen, Ziqiang Liu, Feiteng Fang, Chak Tou Leong, Shiwen Ni, Ahmadreza Argha, Hamid Alinejad-Rokny, Min Yang, Chengming Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Dingwei Chen, Feiteng Fang, Chak Tou Leong, Shiwen Ni, Ahmadreza Argha, Hamid Alinejad-Rokny, Min Yang 0007, Chengming Li 0004
EMNLP4
2025 TokenSkip: Controllable Chain-of-Thought Compression in LLMs
abstract
Chain-of-Thought (CoT) has been proven effective in enhancing the reasoning capabilities of large language models (LLMs).Recent advancements, such as OpenAI's o1 and DeepSeek-R1, suggest that scaling up the length of CoT sequences during inference could further boost LLM reasoning performance.However, due to the autoregressive nature of LLM decoding, longer CoT outputs lead to a linear increase in inference latency, adversely affecting user experience, particularly when the CoT exceeds 10,000 tokens.To address this limitation, we analyze the semantic importance of tokens within CoT outputs and reveal that their contributions to reasoning vary.Building on this insight, we propose TokenSkip, a simple yet effective approach that enables LLMs to selectively skip less important tokens, allowing for controllable CoT compression.Extensive experiments across various models and tasks demonstrate the effectiveness of TokenSkip in reducing CoT token usage while preserving strong reasoning performance.Notably, when applied to Qwen2.5-14B-Instruct,TokenSkip reduces reasoning tokens by 40% (from 313 to 181) on GSM8K, with less than a 0.4% performance drop.We release our code and checkpoints in https: //github.com/hemingkx/TokenSkip.
Heming Xia, Chak Tou Leong, Wenjie Wang 0007, Yongqi Li 0001, Wenjie Li 0002
EMNLP2
2025 Constrain Alignment with Sparse Autoencoders
abstract
The alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have achieved notable success, they often experience computational inefficiencies and training instability. In this paper, we propose Feature-level constrained Preference Optimization (FPO), a novel method designed to simplify the alignment process while ensuring stability. FPO leverages pre-trained Sparse Autoencoders (SAEs) and introduces feature-level constraints, allowing for efficient, sparsity-enforced alignment. Our approach enjoys efficiency by using sparse features activated in a well-trained sparse autoencoder and the quality of sequential KL divergence by using the feature-level offline reference. Experimental results on benchmark datasets demonstrate that FPO achieves an above 5% absolute improvement in win rate with much lower computational cost compared to state-of-the-art baselines, making it a promising solution for efficient and controllable LLM alignments.
Qingyu Yin, Chak Tou Leong, Minjun Zhu, Hanqi Yan, Qiang Zhang 0026, Yulan He 0001, Wenjie Li 0002, Jun Wang 0012, Yue Zhang 0004, Linyi Yang
ICML2
2024 Cooper: Coordinating Specialized Agents towards a Complex Dialogue Goal
abstract
In recent years, there has been a growing interest in exploring dialogues with more complex goals, such as negotiation, persuasion, and emotional support, which go beyond traditional service-focused dialogue systems. Apart from the requirement for much more sophisticated strategic reasoning and communication skills, a significant challenge of these tasks lies in the difficulty of objectively measuring the achievement of their goals in a quantifiable way, making it difficult for existing research to directly optimize the dialogue procedure towards them. In our work, we emphasize the multifaceted nature of complex dialogue goals and argue that it is more feasible to accomplish them by comprehensively considering and jointly promoting their different aspects. To this end, we propose a novel dialogue framework, Cooper, which coordinates multiple specialized agents, each dedicated to a specific dialogue goal aspect separately, to approach the complex objective. Through this divide-and-conquer manner, we make complex dialogue goals more approachable and elicit greater intelligence via the collaboration of individual agents. Experiments on persuasion and emotional support dialogues demonstrate the superiority of our method over a set of competitive baselines. Our codes are available at https://github.com/YiCheng98/Cooper.
Wenge Liu, Jian Wang 0054, Chak Tou Leong, Wenjie Li 0002, Xian Wu 0001, Yefeng Zheng 0001
AAAI4
2024 Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
abstract
Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents.Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interactive process that dialogues ought to be.Such a manner often leads to unsatisfactory chat consistency for the built agent.In this work, we emphasize the interactive, communicative nature of dialogue and argue that it is more feasible to model the speaker roles of agent and user separately, enabling the agent to adhere to its role consistently.With this in mind, we propose an efficient Multi-round Interactive Dialogue Tuning (MIDI-Tuning) framework 1 .It models the agent and user individually with two adapters built upon large language models.The adapters make use of respective utterances round by round in alternating order and they are tuned via a round-level memory caching mechanism.Extensive experiments demonstrate that, our framework performs superior to traditional finetuning and harbors the tremendous potential for improving dialogue consistency.
Jian Wang 0054, Chak Tou Leong, Jiashuo Wang, Dongding Lin, Wenjie Li 0002, Xiaoyong Wei
ACL (1)2
2024 SCREEN: A Benchmark for Situated Conversational Recommendation
abstract
Engaging in conversational recommendations within a specific scenario represents a promising paradigm in the real world. Scenario-relevant situations often affect conversations and recommendations from two closely related aspects: varying the appealingness of items to users, namely situated item representation, and shifting user interests in the targeted items, namely situated user preference. We highlight that considering those situational factors is crucial, as this aligns with the realistic conversational recommendation process in the physical world. However, it is challenging yet under-explored. In this work, we are pioneering to bridge this gap and introduce a novel setting: Situated Conversational Recommendation Systems (SCRS). We observe an emergent need for high-quality datasets, and building one from scratch requires tremendous human effort. To this end, we construct a new benchmark, named SCREEN, via a role-playing method based on multimodal large language models. We take two multimodal large language models to play the roles of a user and a recommender, simulating their interactions in a co-observed scene. Our SCREEN comprises over 20k dialogues across 1.5k diverse situations, providing a rich foundation for exploring situational influences on conversational recommendations. Based on the SCREEN, we propose three worth-exploring subtasks and evaluate several representative baseline models. Our evaluations suggest that the benchmark is high quality, establishing a solid experimental basis for future research. The code and data are available at https://github.com/DongdingLin/SCREEN.
Dongding Lin, Jian Wang 0054, Chak Tou Leong, Wenjie Li 0002
ACM Multimedia3
2023 Self-Detoxifying Language Models via Toxification Reversal
abstract
Language model detoxification aims to minimize the risk of generating offensive or harmful content in pretrained language models (PLMs) for safer deployment.Existing methods can be roughly categorized as finetuning-based and decoding-based.However, the former is often resource-intensive, while the latter relies on additional components and potentially compromises the generation fluency.In this paper, we propose a more lightweight approach that enables the PLM itself to achieve "selfdetoxification".Our method is built upon the observation that prepending a negative steering prompt can effectively induce PLMs to generate toxic content.At the same time, we are inspired by the recent research in the interpretability field, which formulates the evolving contextualized representations within the PLM as an information stream facilitated by the attention layers.Drawing on this idea, we devise a method to identify the toxification direction from the normal generation process to the one prompted with the negative prefix, and then steer the generation to the reversed direction by manipulating the information movement within the attention layers.Experimental results show that our approach, without any fine-tuning or extra components, can achieve comparable performance with state-of-the-art methods. 1 A simple approach to controlled text generation.
Chak Tou Leong, Jiashuo Wang, Jian Wang 0054, Wenjie Li 0002
EMNLP1
2023 Target-oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset Curation
abstract
Target-oriented dialogue systems, designed to proactively steer conversations toward predefined targets or accomplish specific system-side goals, are an exciting area in conversational AI.In this work, by formulating a pair as the conversation target, we explore a novel problem of personalized targetoriented dialogue by considering personalization during the target accomplishment process.However, there remains an emergent need for high-quality datasets, and building one from scratch requires tremendous human effort.To address this, we propose an automatic dataset curation framework using a role-playing approach.Based on this framework, we construct a large-scale personalized target-oriented dialogue dataset, TOPDIAL 1 , which comprises about 18K multi-turn dialogues.The experimental results show that this dataset is of high quality and could contribute to exploring personalized target-oriented dialogue.
Jian Wang 0054, Dongding Lin, Chak Tou Leong, Wenjie Li 0002
EMNLP4