VLDB 2026 Research / reviewers in the wild / expert
Zhiheng Xi
dblp:333/4268
· DBLP profile ↗
33ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0001-8299-8449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 6 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyabstractSpeech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency. Xiaoran Fan, Yangfan Gao, Jingfei Xiong, Hang Yan 0001, Yifei Cao, Zhihao Zhang 0002, Zhiheng Xi, Yuhao Zhou 0005, Senjie Jin, Changhao Jiang, Junjie Ye 0005, Ming Zhang 0030, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Gui |
AAAI | 10 |
| 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data ContaminationabstractReasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model’s performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods. Mingqi Wu, Zhihao Zhang 0002, Qiaole Dong, Zhiheng Xi, Jun Zhao 0019, Senjie Jin, Xiaoran Fan, Yuhao Zhou 0005, Huijie Lv, Ming Zhang 0030, Yanwei Fu 0001, Qin Liu 0010, Songyang Zhang 0001, Qi Zhang 0001 |
AAAI | 4 |
| 2026 | MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningabstractOutcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B. Zhiheng Xi, Yiwen Ding, Senjie Jin, Shichun Liu, Jixuan Huang, Dingwen Yang, Jiafu Tang, Boyang Hong, Junjie Ye 0005, Shihan Dou, Ming Zhang 0030, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 1 |
| 2026 | Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancingabstractXin Guo, Zhiheng Xi, Yiwen Ding, Yitao Zhai, Xiaowei Shi, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiheng Xi, Yiwen Ding, Yitao Zhai, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 2 |
| 2026 | Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-TrainingabstractChanghao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 7 |
| 2026 | Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationabstractSearch agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and longtail information unavailable during pretraining.While reinforcement learning has been widely adopted for training such agents, existing approaches face key limitations: process supervision often suffers from unstable value estimation, whereas outcome supervision struggles with credit assignment due to sparse, trajectory-level rewards.To bridge this gap, we propose Contribution-Weighted GRPO (CW-GRPO), a framework that integrates process supervision into group relative policy optimization.Instead of directly optimizing process rewards, CW-GRPO employs an LLM judge to assess the retrieval utility and reasoning correctness at each search round, producing per-round contribution weights.These weights are used to rescale outcome-based advantages along the trajectory, enabling finegrained credit assignment without sacrificing optimization stability.Experiments on multiple knowledge-intensive benchmarks show that CW-GRPO outperforms standard GRPO by 5.0% on Qwen3-8B and 6.3% on Qwen3-1.7B,leading to more effective search behaviors.Additional analysis reveals that successful trajectories exhibit concentrated contributions in specific rounds, providing empirical insight into search agent tasks.Our code is available at https://github.com/zsxmwjz/CW-GRPO. Junzhe Wang 0001, Zhiheng Xi, Yajie Yang, Shihan Dou, Tao Gui, Qi Zhang 0001 |
ACL (1) | 2 |
| 2026 | AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsabstractZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2026 | Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentabstractYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2026 | LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsabstractMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang, Junzhe Wang, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ming Zhang 0030, Yujiong Shen, Jingyi Deng, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang 0004, Junzhe Wang 0001, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang 0002, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 18 |
| 2026 | VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-TrainingabstractDingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang, Chenhao Huang, Yunke Zhang, Yuran Wang, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiazheng Zhang, Junjie Ye 0005, Mingxu Chai, Enyu Zhou, Ming Zhang 0030, Caishuang Huang, Chenhao Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2026 | AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and ProgressabstractDespite rapid development, large language models (LLMs) still encounter challenges in multi-turn decision-making tasks (i.e., agent tasks) like web shopping and browser navigation, which require making a sequence of intelligent decisions based on environmental feedback. Previous work for LLM agents typically relies on elaborate prompt engineering or fine-tuning with expert trajectories to improve performance. In this work, we take a different perspective: we explore constructing process reward models (PRMs) to evaluate each decision and guide the agent's decision-making process. Unlike LLM reasoning, where each step is scored based on correctness, actions in agent tasks do not have a clear-cut correctness. Instead, they should be evaluated based on their proximity to the goal and the progress they have made. Building on this insight, we propose a re-defined PRM for agent tasks, named AgentPRM, to capture both the interdependence between sequential decisions and their contribution to the final goal. This enables better progress tracking and exploration-exploitation balance. To scalably obtain labeled data for training AgentPRM, we employ a Temporal Difference-based (TD-based) estimation method combined with Generalized Advantage Estimation (GAE), which proves more sample-efficient than prior methods. Extensive experiments across different agentic tasks show that AgentPRM is over 8× more compute-efficient than baselines, and it demonstrates robust improvement when scaling up test-time compute. Moreover, we perform detailed analyses to show how our method works and offer more insights, e.g., applying AgentPRM to the reinforcement learning of LLM agents. Zhiheng Xi, Chenyang Liao, Zhihao Zhang 0002, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou 0005, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
WWW | 1 |
| 2026 | What is wrong with your code generated by large language models? An extensive study
Shihan Dou, Haoxiang Jia, Shenxi Wu, Huiyuan Zheng, Muling Wu, Yunbo Tao, Ming Zhang 0030, Mingxu Chai, Jessica Fan, Zhiheng Xi, Yueming Wu 0001, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 10 |
| 2025 | CritiQ: Mining Data Quality Criteria from Human PreferencesabstractHonglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui |
ACL (1) | 5 |
| 2025 | AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsabstractZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Xin Guo, Dingwen Yang, Chenyang Liao, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang 0001, Dingwen Yang, Chenyang Liao, Wei He 0024, Songyang Gao, Lu Chen 0001, Yicheng Zou, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
ACL (1) | 1 |
| 2025 | ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool UseabstractJunjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, Jiecao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junjie Ye 0005, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Jiecao Chen |
ACL (1) | 9 |
| 2025 | Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for ReasoningabstractSenjie Jin, Lu Chen, Zhiheng Xi, Yuhui Wang, Sirui Song, Yuhao Zhou, Xinbo Zhang, Peng Sun, Hong Lu, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Senjie Jin, Lu Chen 0001, Zhiheng Xi, Sirui Song, Yuhao Zhou 0005, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 3 |
| 2025 | LoRACoE: Improving Large Language Model via Composition-based LoRA ExpertabstractThe Mixture of Experts (MoE) architecture improves large language models (LLMs) by utilizing sparsely activated expert sub-networks with a routing module, yet it typically demands high training cost.Previous work introduces parameter-efficient fine-tuning (PEFT) modules, e.g., LoRA, to achieve a lightweight MoE for efficiency.However, they construct static experts by manually splitting the LoRA parameters into fixed groups, which limits flexibility and dynamism.Furthermore, this manual partitioning also hinders the effective utilization of well-initialized LoRA modules.To tackl the challenges, we first delve into the parameter patterns in LoRA modules, revealing that there exists task-relevant parameters that are concentrated along the rank dimension.Based on this, we redesign the construction of experts and propose the LoRACoE (LoRA Composition of Experts) method.Specifically, when confronted with a task, it dynamically builds experts based on rank-level parameter composition, i.e., experts can flexibly combine rank-level parameters in LoRA module.Extensive experiments demonstrate that compared to other LoRA-based MoE methods, our method achieves better task performance across a broader range of tasks. Zhiheng Xi, Zhihao Zhang 0002, Boyang Hong, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 2 |
| 2025 | Have the VLMs Lost Confidence? A Study of Sycophancy in VLMsabstractIn the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue. Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang 0001, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 7 |
| 2025 | RMB: Comprehensively benchmarking reward models in LLM alignmentabstractReward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization.
We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. We will release our evaluation code and datasets upon publication. Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Limao Xiong, Jessica Fan, Yurong Mou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 4 |
| 2025 | Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided SamplingabstractYiwen Ding, Zhiheng Xi, Wei He, Lizhuoyuan Lizhuoyuan, Yitao Zhai, Shi Xiaowei, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yiwen Ding, Zhiheng Xi, Wei He 0024, Lizhuoyuan Lizhuoyuan, Yitao Zhai, Shi Xiaowei, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NAACL (Long Papers) | 2 |
| 2025 | Pre-Trained Policy Discriminators are General Reward ModelsabstractWe offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance.
For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines.
POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks.
Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99.
The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models. Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026 |
NeurIPS | 15 |
| 2025 | BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning DatasetabstractIn this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community. Zhiheng Xi, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 1 |
| 2025 | Visual Sketchbook: Enhancing Chart-to-Code Generation via Reflective Refinement
Junzhe Wang 0001, Zhiheng Xi, Wei He 0024, Dingwei Zhu, Shihan Dou, Tao Gui, Qi Zhang 0001 |
NLPCC (2) | 2 |
| 2025 | The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui |
Sci. China Inf. Sci. | 1 |
| 2024 | StepCoder: Improving Code Generation with Reinforcement Learning from Compiler FeedbackabstractShihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Yan Liu 0002, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang 0001, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 10 |
| 2024 | LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style PluginabstractShihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Enyu Zhou, Yan Liu 0002, Songyang Gao, Limao Xiong, Yuhao Zhou 0005, Xiao Wang 0001, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2024 | RoCoIns: Enhancing Robustness of Large Language Models through Code-Style InstructionsabstractLarge Language Models (LLMs) have showcased remarkable capabilities in following human instructions. However, recent studies have raised concerns about the robustness of LLMs for natural language understanding (NLU) tasks when prompted with instructions combining textual adversarial samples. In this paper, drawing inspiration from recent works that LLMs are sensitive to the design of the instructions, we utilize instructions in code style, which are more structural and less ambiguous, to replace typically natural language instructions. Through this conversion, we provide LLMs with more precise instructions and strengthen the robustness of LLMs. Moreover, under few-shot scenarios, we propose a novel method to compose in-context demonstrations using both clean and adversarial samples (adversarial context method) to further boost the robustness of the LLMs. Experiments on eight robustness datasets show that our method consistently outperforms prompting LLMs with natural language, for example, with gpt-3.5-turbo on average, our method achieves an improvement of 5.68% in test set accuracy and a reduction of 5.66 points in Attack Success Rate (ASR). Yuansen Zhang, Xiao Wang 0001, Zhiheng Xi, Han Xia 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
LREC/COLING | 3 |
| 2024 | Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean SignalsabstractDeep neural networks (DNNs) are notoriously vulnerable to adversarial attacks that place carefully crafted perturbations on normal examples to fool DNNs. To better understand such attacks, a characterization of the features carried by adversarial examples is needed. In this paper, we tackle this challenge by inspecting the subspaces of sample features through spectral analysis. We first empirically show that the features of either clean signals or adversarial perturbations are redundant and span in low-dimensional linear subspaces respectively with minimal overlap, and the classical low-dimensional subspace projection can suppress perturbation features out of the subspace of clean signals. This makes it possible for DNNs to learn a subspace where only features of clean signals exist while those of perturbations are discarded, which can facilitate the distinction of adversarial examples. To prevent the residual perturbations that is inevitable in subspace learning, we propose an independence criterion to disentangle clean signals from perturbations. Experimental results show that the proposed strategy enables the model to inherently suppress adversaries, which not only boosts model robustness but also motivates new directions of effective adversarial defense. Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
LREC/COLING | 3 |
| 2024 | ORTicket: Let One Robust BERT Ticket Transfer across Different TasksabstractPretrained language models can be applied for various downstream tasks but are susceptible to subtle perturbations. Most adversarial defense methods often introduce adversarial training during the fine-tuning phase to enhance empirical robustness. However, the repeated execution of adversarial training hinders training efficiency when transitioning to different tasks. In this paper, we explore the transferability of robustness within subnetworks and leverage this insight to introduce a novel adversarial defense method ORTicket, eliminating the need for separate adversarial training across diverse downstream tasks. Specifically, (i) pruning the full model using the MLM task (the same task employed for BERT pretraining) yields a task-agnostic robust subnetwork(i.e., winning ticket in Lottery Ticket Hypothesis); and (ii) fine-tuning this subnetwork for downstream tasks. Extensive experiments demonstrate that our approach achieves comparable robustness to other defense methods while retaining the efficiency of traditional fine-tuning.This also confirms the significance of selecting MLM task for identifying the transferable robust subnetwork. Furthermore, our method is orthogonal to other adversarial training approaches, indicating the potential for further enhancement of model robustness. Yuhao Zhou 0005, Wenxiang Chen, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
LREC/COLING | 4 |
| 2024 | Improving Discriminative Capability of Reward Models in RLHF Using Contrastive LearningabstractLu Chen, Rui Zheng, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye, Zhihao Zhang, Yuhao Zhou, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lu Chen 0001, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye 0005, Zhihao Zhang 0002, Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 9 |
| 2024 | Improving Generalization of Alignment with Human Preferences through Group Invariant LearningabstractThe success of AI assistants based on language models (LLMs) hinges crucially on Reinforcement Learning from Human Feedback (RLHF), which enables the generation of responses more aligned with human preferences.
As universal AI assistants, there's a growing expectation for them to perform consistently across various domains.
However, previous work shows that Reinforcement Learning (RL) often exploits shortcuts to attain high rewards and overlooks challenging samples.
This focus on quick reward gains undermines both the stability in training and the model's ability to generalize to new, unseen data.
In this work, we propose a novel approach that can learn a consistent policy via RL across various data groups or domains.
Given the challenges associated with acquiring group annotations, our method automatically classifies data into different groups, deliberately maximizing performance variance.
Then, we optimize the policy to perform well on challenging groups.
Lastly, leveraging the established groups, our approach adaptively adjusts the exploration space, allocating more learning capacity to more challenging data and preventing the model from over-optimizing on simpler data. Experimental results indicate that our approach significantly enhances training stability and model generalization. Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou 0005, Zhiheng Xi, Xiao Wang 0001, Haoran Huang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 7 |
| 2024 | Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningabstractIn this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$. Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICML | 1 |
| 2022 | Efficient Adversarial Training with Robust Early-Bird TicketsabstractAdversarial training is one of the most powerful methods to improve the robustness of pretrained language models (PLMs).However, this approach is typically more expensive than traditional fine-tuning because of the necessity to generate adversarial examples via gradient descent.Delving into the optimization process of adversarial training, we find that robust connectivity patterns emerge in the early training phase (typically 0.15 ∼ 0.3 epochs), far before parameters converge.Inspired by this finding, we dig out robust early-bird tickets (i.e., subnetworks) to develop an efficient adversarial training method: (1) searching for robust tickets with structured sparsity in the early stage; (2) fine-tuning robust tickets in the remaining time.To extract the robust tickets as early as possible, we design a ticket convergence metric to automatically terminate the searching process.Experiments show that the proposed efficient adversarial training method can achieve up to 7× ∼ 13× training speedups while maintaining comparable or even better robustness compared to the most competitive state-of-the-art adversarial training methods. Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 1 |