Chujie Zheng

dblp:242/8504 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
abstract
Binghai Wang, Yantao Liu, Yuxuan Liu, Tianyi Tang, Shenzhi Wang, Chang Gao, Chujie Zheng, Yichang Zhang, Le Yu, Shixuan Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Bowen Yu, Fei Huang, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Binghai Wang, Yantao Liu, Shenzhi Wang, Chujie Zheng, Yichang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Bowen Yu 0002, Fei Huang 0002, Junyang Lin
ACL (1)7
2025 Model Extrapolation Expedites Alignment
abstract
Given the high computational cost of preference alignment training of large language models (LLMs), exploring efficient methods to reduce the training overhead remains an important and compelling research problem.Motivated by the observation that alignment training typically involves only small parameter changes without injecting new knowledge into models, we propose a straightforward method called EXPO (model extrapolation) to expedite LLMs' alignment with human preferences.Given a partially-trained model and its initial SFT checkpoint, EXPO improves the implicit optimization objective of alignment training by simply amplifying the parameter change based on a first-order approximation, without any additional training overhead.Through controlled experiments, we demonstrate that EXPO boosts a DPO model trained with only 20% steps to outperform the fullytrained one.Moreover, we show that EXPO notably improves existing open-source LLMs (ranging from 1.8B to 70B parameters) on the leading AlpacaEval 2.0 and MT-Bench benchmarks, which highlights EXPO's broader utility in efficiently enhancing LLM alignment.
Chujie Zheng, Ziqi Wang 0003, Heng Ji 0001, Minlie Huang, Nanyun Peng 0001
ACL (1)1
2025 ProcessBench: Identifying Process Errors in Mathematical Reasoning
abstract
Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chujie Zheng, Zhenru Zhang, Runji Lin, Keming Lu, Bowen Yu 0002, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin
ACL (1)1
2025 Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), yet its underlying mechanisms remain insufficiently understood. In this work, we undertake a pioneering exploration of RLVR through the novel perspective of token entropy patterns, comprehensively analyzing how different tokens influence reasoning performance. By examining token entropy patterns in Chain-of-Thought (CoT) reasoning, we observe that only a small fraction (approximately 20\%) of tokens exhibit high entropy, and these tokens semantically act as critical forks that steer the model toward diverse reasoning pathways. We further demonstrate that moderately increasing the entropy of these high-entropy tokens via decoding temperature adjustments leads to improved performance, quantitatively confirming their role as decision points in reasoning. We ultimately refine RLVR by restricting policy gradient updates to these forking tokens. Despite utilizing only 20\% of tokens, our approach achieves comparable performance to full-gradient updates on the Qwen3-8B base model. Moreover, it demonstrates remarkable improvements on the larger Qwen3-32B base model, boosting AIME'25 scores by 11.04 and AIME'24 scores by 7.71. In contrast, training exclusively on the 80\% lowest-entropy tokens leads to a marked decline in performance. These findings indicate that the efficacy of RLVR primarily arises from optimizing the high-entropy tokens that dictate key reasoning directions. Collectively, our results suggest promising avenues for optimizing RLVR algorithms by strategically leveraging the potential of these high-entropy minority tokens to further enhance the reasoning abilities of LLMs.
Shenzhi Wang, Chujie Zheng, Rui Lu 0001, Kai Dang, Xiong-Hui Chen, Jianxin Yang, Zhenru Zhang, Yuqiong Liu, An Yang, Andrew Zhao, Shiji Song, Bowen Yu 0002, Gao Huang 0001, Junyang Lin
NeurIPS4
2025 Retrieval for Semantic People Search
abstract
These days we have a large number of open-source embedding LLMs that can be leveraged in a bi-encoder architecture for retrieval. However, they do not perform very well when leveraged as is for people search because the (query, document) pairs encountered in people search are much more complex than the text pairs that these open-source embedding LLMs are trained on. In this paper we present our experiments, learnings and solution to the problem of retrieval for people search. Our solution involves query simplification, document simplification, fine-tuning of a 7B parameter embedding LLM, and compression of embeddings through Matryoshka learning.
Rupesh Gupta, Chujie Zheng
SIGIR2
2024 Large Language Models Are Not Robust Multiple Choice Selectors
abstract
Multiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs). This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent “selection bias”, namely, they prefer to select specific option IDs as answers (like “Option A”). Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs’ token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs. To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model’s prior bias for option IDs from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples, and then applies the estimated prior to debias the remaining samples. We demonstrate that it achieves interpretable and transferable debiasing with high computational efficiency. We hope this work can draw broader research attention to the bias and robustness of modern LLMs.
Chujie Zheng, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Minlie Huang
ICLR1
2024 On Prompt-Driven Safeguarding for Large Language Models
abstract
Prepending model inputs with safety prompts is a common practice for safeguarding large language models (LLMs) against queries with harmful intents. However, the underlying working mechanisms of safety prompts have not been unraveled yet, restricting the possibility of automatically optimizing them to improve LLM safety. In this work, we investigate how LLMs’ behavior (i.e., complying with or refusing user queries) is affected by safety prompts from the perspective of model representation. We find that in the representation space, the input queries are typically moved by safety prompts in a "higher-refusal" direction, in which models become more prone to refusing to provide assistance, even when the queries are harmless. On the other hand, LLMs are naturally capable of distinguishing harmful and harmless queries without safety prompts. Inspired by these findings, we propose a method for safety prompt optimization, namely DRO (Directed Representation Optimization). Treating a safety prompt as continuous, trainable embeddings, DRO learns to move the queries’ representations along or opposite the refusal direction, depending on their harmfulness. Experiments with eight LLMs on out-of-domain and jailbreak benchmarks demonstrate that DRO remarkably improves the safeguarding performance of human-crafted safety prompts, without compromising the models’ general performance.
Chujie Zheng, Fan Yin, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Kai-Wei Chang 0001, Minlie Huang, Nanyun Peng 0001
ICML1
2023 CASE: Aligning Coarse-to-Fine Cognition and Affection for Empathetic Response Generation
abstract
Empathetic conversation is psychologically supposed to be the result of conscious alignment and interaction between the cognition and affection of empathy.However, existing empathetic dialogue models usually consider only the affective aspect or treat cognition and affection in isolation, which limits the capability of empathetic response generation.In this work, we propose the CASE model for empathetic dialogue generation.It first builds upon a commonsense cognition graph and an emotional concept graph and then aligns the user's cognition and affection at both the coarse-grained and fine-grained levels.Through automatic and manual evaluation, we demonstrate that CASE outperforms state-of-the-art baselines of empathetic dialogues and can generate more empathetic and informative responses.1
Jinfeng Zhou, Chujie Zheng, Bo Wang 0011, Zheng Zhang 0020, Minlie Huang
ACL (1)2
2022 CEM: Commonsense-Aware Empathetic Response Generation
abstract
A key trait of daily conversations between individuals is the ability to express empathy towards others, and exploring ways to implement empathy is a crucial step towards human-like dialogue systems. Previous approaches on this topic mainly focus on detecting and utilizing the user’s emotion for generating empathetic responses. However, since empathy includes both aspects of affection and cognition, we argue that in addition to identifying the user’s emotion, cognitive understanding of the user’s situation should also be considered. To this end, we propose a novel approach for empathetic response generation, which leverages commonsense to draw more information about the user’s situation and uses this additional information to further enhance the empathy expression in generated responses. We evaluate our approach on EMPATHETICDIALOGUES, which is a widely-used benchmark dataset for empathetic response generation. Empirical results demonstrate that our approach outperforms the baseline models in both automatic and human evaluations and can generate more informative and empathetic responses. Our code is available at https://github.com/Sahandfer/CEM.
Sahand Sabour, Chujie Zheng, Minlie Huang
AAAI2
2022 COLD: A Benchmark for Chinese Offensive Language Detection
abstract
Offensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models.However, this task in Chinese is still under exploration due to the scarcity of reliable datasets.To this end, we propose a benchmark -COLD for Chinese offensive language analysis, including a Chinese Offensive Language Dataset -COLDATASET and a baseline detector -COLDETECTOR which is trained on the dataset.We show that the COLD benchmark contributes to Chinese offensive language detection which is challenging for existing resources.We then deploy the COLDETECTOR and conduct detailed analyses on popular Chinese pre-trained language models.We first analyze the offensiveness of existing generative models and show that these models inevitably expose varying degrees of offensive issues.Furthermore, we investigate the factors that influence the offensive generations, and we find that anti-bias contents and keywords referring to certain groups or revealing negative attitudes trigger offensive outputs easier.
Jiawen Deng 0006, Jingyan Zhou, Hao Sun 0012, Chujie Zheng, Fei Mi, Helen M. Meng, Minlie Huang
EMNLP4
2022 CDConv: A Benchmark for Contradiction Detection in Chinese Conversations
abstract
Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Zhen Guo, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Minlie Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Wenquan Wu, Zhengyu Niu, Hua Wu 0003, Minlie Huang
EMNLP1
2021 Towards Emotional Support Dialog Systems
abstract
Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Siyang Liu 0003, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li 0013, Zhou Yu 0005, Yong Jiang 0001, Minlie Huang
ACL/IJCNLP (1)2
2021 Enhanced Seq2Seq Autoencoder via Contrastive Learning for Abstractive Text Summarization
abstract
In this paper, we present a denoising sequence-to-sequence (seq2seq) autoencoder via contrastive learning for abstractive text summarization. Our model adopts a standard Transformer-based architecture with a multi-layer bi-directional encoder and an auto-regressive decoder. To enhance its denoising ability, we incorporate self-supervised contrastive learning along with various sentence-level document augmentation. These two components, seq2seq autoencoder and contrastive learning, are jointly trained through fine-tuning, w hich i mproves the performance of text summarization with regard to ROUGE scores and human evaluation. We conduct experiments on two datasets and demonstrate that our model outperforms many existing benchmarks and even achieves comparable performance to the state-of-the-art abstractive systems trained with more complex architecture and extensive computation resources.
Chujie Zheng, Kunpeng Zhang 0001, Harry J. Wang, Ling Fan
IEEE BigData1
2020 KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven Conversation
abstract
The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consists of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese multi-domain knowledge-driven conversation dataset, KdConv, which grounds the topics in multi-turn conversations to knowledge graphs. Our corpus contains 4.5K conversations from three domains (film, music, and travel), and 86K utterances with an average turn number of 19.0. These conversations contain in-depth discussions on related topics and natural transition between multiple topics. To facilitate the following research on this corpus, we provide several benchmark models. Comparative results show that the models can be enhanced by introducing background knowledge, yet there is still a large space for leveraging knowledge to model multi-turn conversations for further research. Results also show that there are obvious performance differences between different domains, indicating that it is worth further explore transfer learning and domain adaptation. The corpus and benchmark models are publicly available.
Hao Zhou 0012, Chujie Zheng, Kaili Huang, Minlie Huang, Xiaoyan Zhu 0001
ACL2
2019 ChID: A Large-scale Chinese IDiom Dataset for Cloze Test
abstract
Cloze-style reading comprehension in Chinese is still limited due to the lack of various corpora.In this paper we propose a large-scale Chinese cloze test dataset ChID, which studies the comprehension of idiom, a unique language phenomenon in Chinese.In this corpus, the idioms in a passage are replaced by blank symbols and the correct answer needs to be chosen from well-designed candidate idioms.We carefully study how the design of candidate idioms and the representation of idioms affect the performance of state-of-the-art models.Results show that the machine accuracy is substantially worse than that of human, indicating a large space for further research.
Chujie Zheng, Minlie Huang, Aixin Sun
ACL (1)1