EDBT 2026 Demo / reviewers in the wild / expert
Zhiwei He 0002
dblp:52/6077-2
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0002-4807-0062ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form GenerationabstractConventional speculative decoding (SD) methods utilize a predefined length policy for proposing drafts, which implies the premise that the target model smoothly accepts the proposed draft tokens.However, reality deviates from this assumption: the oracle draft length varies significantly, and the fixed-length policy hardly satisfies such a requirement.Moreover, such discrepancy is further exacerbated in scenarios involving complex reasoning and long-form generation, particularly under testtime scaling for reasoning-specialized models.Through both theoretical and empirical estimation, we establish that the discrepancy between the draft and target models can be approximated by the draft model's prediction entropy: a high entropy indicates a low acceptance rate of draft tokens, and vice versa.Based on this insight, we propose SVIP: Self-Verification Length Policy for Long-Context Speculative Decoding, which is a training-free dynamic length policy for speculative decoding systems that adaptively determines the lengths of draft sequences by referring to the draft entropy.Experimental results on mainstream SD benchmarks as well as reasoning-heavy benchmarks demonstrate the superior performance of SVIP, achieving up to 17% speedup on MT-Bench at 8K context compared with fixed draft lengths, and 22% speedup for QwQ in long-form reasoning. Ziyin Zhang, Zhiwei He 0002, Rui Wang 0015, Zhaopeng Tu |
EMNLP | 5 |
| 2025 | RaSA: Rank-Sharing Low-Rank AdaptationabstractLow-rank adaptation (LoRA) has been prominently employed for parameter-efficient fine-tuning of large language models (LLMs). However, the limited expressive capacity of LoRA, stemming from the low-rank constraint, has been recognized as a bottleneck, particularly in rigorous tasks like code generation and mathematical reasoning. To address this limitation, we introduce Rank-Sharing Low-Rank Adaptation (RaSA), an innovative extension that enhances the expressive capacity of LoRA by leveraging partial rank sharing across layers. By forming a shared rank pool and applying layer-specific weighting, RaSA effectively increases the number of ranks without augmenting parameter overhead. Our theoretically grounded and empirically validated approach demonstrates that RaSA not only maintains the core advantages of LoRA but also significantly boosts performance in challenging code and math tasks. Code, data and scripts are available at: https://github.com/zwhe99/RaSA. Zhiwei He 0002, Zhaopeng Tu, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015 |
ICLR | 1 |
| 2025 | Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned ModelabstractAligning language models (LMs) with human preferences has become a key area of research, enabling these models to meet diverse user needs better. Inspired by weak-to-strong generalization, where a strong LM fine-tuned on labels generated by a weaker model can consistently outperform its weak supervisor, we extend this idea to model alignment. In this work, we observe that the alignment behavior in weaker models can be effectively transferred to stronger models and even exhibit an amplification effect. Based on this insight, we propose a method called Weak-to-Strong Preference Optimization (WSPO), which achieves strong model alignment by learning the distribution differences before and after the alignment of the weak model. Experiments demonstrate that WSPO delivers outstanding performance, improving the win rate of Qwen2-7B-Instruct on Arena-Hard from 39.70 to 49.60, achieving a remarkable 47.04 length-controlled win rate on AlpacaEval 2, and scoring 7.33 on MT-bench. Our results suggest that using the weak model to elicit a strong model with a high alignment ability is feasible. The code is available at https://github.com/zwhong714/weak-to-strong-preference-optimization. Zhiwei He 0002, Rui Wang 0015 |
ICLR | 2 |
| 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning ModelsabstractThe remarkable performance of long reasoning models can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple strategies to enhance problem-solving capabilities. However, a critical question remains: How to intelligently and efficiently scale computational resources during testing. This paper presents the first comprehensive study on the prevalent issue of overthinking in these models, where long reasoning models generate redundant solutions that contribute minimally to accuracy and diversity, thereby wasting computational resources on simple problems with minimal benefit. We introduce novel efficiency metrics from both outcome and process perspectives to evaluate the rational use of computational resources by long reasoning models. Using a self-training paradigm, we propose strategies to mitigate overthinking, simplifying reasoning processes without compromising accuracy. Experimental results show that our approach successfully reduces computational overhead while preserving model performance across a range of testsets with varying difficulty levels, such as GSM8K, MATH500, GPQA, and AIME. Our code is open-source and available at https://github.com/galaxyChen/overthinking. Zhiwei He 0002, Jianhui Pang, Dian Yu 0001, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
ICML | 4 |
| 2025 | Gloss Matters: Unlocking the Potential of Non-Autoregressive Sign Language TranslationabstractWhile non-autoregressive sign language translation (NASLT) has the advantage in inference speed, the translation quality of NASLT models lags significantly behind that of the state-of-the-art (SOTA) autoregressive sign language translation (ASLT) models. To bridge the quality gap, we exploit glosses to unlock the potential of NASLT models. Concretely, we propose Gloss-enhanced Levenshtein Transformer (GLevT) for sign language translation (SLT), which takes glosses as initial sequences for editing into texts. In particular, to alleviate the inconsistency between training and inference of GLevT, which is introduced by glosses, we propose a dual-centric learning policy and a keyframe-based gloss replacement method for training, further improving the translation quality of GLevT. Experiments on CSL-Daily demonstrate that GLevT outperforms other NASLT models by approximately 4 points in BLEU and ROUGE scores, while achieving performance comparable to the SOTA ASLT models with a 3.46~5.26× inference speed-up. Furthermore, we extend GLevT to gloss-free SLT, achieving performance comparable to SOTA large models, despite having only 49M parameters. We release code at https://github.com/XMUDeepLIT/GLevT. Zhiwei He 0002, Kangjie Zheng, Liangying Shao, Junfeng Yao, Jinsong Su |
ACM Multimedia | 3 |
| 2025 | The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning ModelsabstractImproving the reasoning capabilities of large language models (LLMs) typically requires supervised fine-tuning with labeled data or computationally expensive sampling. We introduce Unsupervised Prefix Fine-Tuning (UPFT), which leverages the observation of Prefix Self-Consistency -- the shared initial reasoning steps across diverse solution trajectories -- to enhance LLM reasoning efficiency. By training exclusively on the initial prefix substrings (as few as 8 tokens), UPFT removes the need for labeled data or exhaustive sampling. Experiments on reasoning benchmarks show that UPFT matches the performance of supervised methods such as Rejection Sampling Fine-Tuning, while reducing training time by 75\% and sampling cost by 99\%. Further analysis reveals that errors tend to appear in later stages of the reasoning process and that prefix-based training preserves the model’s structural knowledge. This work demonstrates how minimal unsupervised fine-tuning can unlock substantial reasoning gains in LLMs, offering a scalable and resource-efficient alternative to conventional approaches. Ke Ji, Qiuzhi Liu, Zhiwei He 0002, Benyou Wang, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 5 |
| 2025 | Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsabstractLarge Language Models (LLMs) show great promise in complex reasoning, with Reinforcement Learning with Verifiable Rewards (RLVR) being a key enhancement strategy. However, a prevalent issue is ``superficial self-reflection'', where models fail to robustly verify their own outputs. We introduce RISE (Reinforcing Reasoning with Self-Verification), a novel online RL framework designed to tackle this. RISE explicitly and simultaneously trains an LLM to improve both its problem-solving and self-verification abilities within a single, integrated RL process. The core mechanism involves leveraging verifiable rewards from an outcome verifier to provide on-the-fly feedback for both solution generation and self-verification tasks. In each iteration, the model generates solutions, then critiques its own on-policy generated solutions, with both trajectories contributing to the policy update.
Extensive experiments on diverse mathematical reasoning benchmarks show that RISE consistently improves model's problem-solving accuracy while concurrently fostering strong self-verification skills. Our analyses highlight the advantages of online verification and the benefits of increased verification compute. Additionally, RISE models exhibit more frequent and accurate self-verification behaviors during reasoning. These advantages reinforce RISE as a flexible and effective path towards developing more robust and self-aware reasoners. Zhiwei He 0002, Wenxuan Wang 0001, Pinjia He, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 3 |
| 2025 | Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional TrainingabstractMixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning depth and efficiency without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed cognitive experts that orchestrate meta-level reasoning operations characterized by tokens like <think>. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks (AIME and GPQA Diamond) demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models. Yue Wang 0039, Zhiwei He 0002, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang 0001, Ruotian Ma, Haitao Mi, Ningyu Zhang 0001, Zhaopeng Tu, Dong Yu 0001 |
NeurIPS | 4 |
| 2025 | Thoughts Are All Over the Place: On the Underthinking of Long Reasoning ModelsabstractLong reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking. Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 6 |
| 2024 | Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language ModelsabstractZhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, Rui Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhiwei He 0002, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang 0007, Zhaopeng Tu, Zhuosheng Zhang 0001, Rui Wang 0015 |
ACL (1) | 1 |
| 2024 | Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateabstractModern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies.Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively.However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution.Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation.Experiment results on two challenging datasets, commonsense machine translation and counterintuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework.Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance.Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents.Code is available at https://github. com/Skytliang/Multi-Agents-Debate. Zhiwei He 0002, Wenxiang Jiao, Xing Wang 0007, Yan Wang 0060, Rui Wang 0015, Yujiu Yang 0001, Shuming Shi 0001, Zhaopeng Tu |
EMNLP | 2 |
| 2024 | Improving Open-Ended Text Generation via Adaptive DecodingabstractCurrent language models decode text token by token according to probabilistic distribution, and determining the appropriate candidates for the next token is crucial to ensure generation quality. This study introduces adaptive decoding, a mechanism that dynamically empowers language models to ascertain a sensible candidate set during generation. Specifically, we introduce an entropy-based metric called confidence and conceptualize determining the optimal candidate set as a confidence-increasing process. The rationality of including a token in the candidate set is assessed by leveraging the increment of confidence. Experimental results reveal that our method balances diversity and coherence well. The human evaluation shows that our method can generate human-preferred text. Additionally, our method can potentially improve the reasoning ability of language models. Hongkun Hao, Zhiwei He 0002, Yiming Ai, Rui Wang 0015 |
ICML | 3 |
| 2024 | Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward ModelabstractZhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, Zhaopeng Tu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhiwei He 0002, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu |
NAACL-HLT | 1 |
| 2024 | Exploring Human-Like Translation Strategy with Large Language ModelsabstractAbstract Large language models (LLMs) have demonstrated impressive capabilities in general scenarios, exhibiting a level of aptitude that approaches, in some aspects even surpasses, human-level intelligence. Among their numerous skills, the translation abilities of LLMs have received considerable attention. Compared to typical machine translation that focuses solely on source-to-target mapping, LLM-based translation can potentially mimic the human translation process, which might take preparatory steps to ensure high-quality translation. This work explores this possibility by proposing the MAPS framework, which stands for Multi-Aspect Prompting and Selection. Specifically, we enable LLMs first to analyze the given source sentence and induce three aspects of translation-related knowledge (keywords, topics, and relevant demonstrations) to guide the final translation process. Moreover, we employ a selection mechanism based on quality estimation to filter out noisy and unhelpful knowledge. Both automatic (3 LLMs × 11 directions × 2 automatic metrics) and human evaluation (preference study and MQM) demonstrate the effectiveness of MAPS. Further analysis shows that by mimicking the human translation process, MAPS reduces various translation errors such as hallucination, ambiguity, mistranslation, awkward style, untranslated text, and omission. Source code is available at https://github.com/zwhe99/MAPS-mt. Zhiwei He 0002, Wenxiang Jiao, Zhuosheng Zhang 0001, Yujiu Yang 0001, Rui Wang 0015, Zhaopeng Tu, Shuming Shi 0001, Xing Wang 0007 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Discrete limited attentional collaborative filtering for fast social recommendation
Zhibin Hu, Xuebin Zhou, Zhiwei He 0002, Zehang Yang, Jian Chen 0011, Jin Huang 0007 |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationabstractBack-translation is a critical component of Unsupervised Neural Machine Translation (UNMT), which generates pseudo parallel data from target monolingual data.A UNMT model is trained on the pseudo parallel data with translated source, and translates natural source sentences in inference.The source discrepancy between training and inference hinders the translation performance of UNMT models.By carefully designing experiments, we identify two representative characteristics of the data gap in source: (1) style gap (i.e., translated vs. natural text style) that leads to poor generalization capability; (2) content gap that induces the model to produce hallucination content biased towards the target language.To narrow the data gap, we propose an online self-training approach, which simultaneously uses the pseudo parallel data {natural source, translated target} to mimic the inference scenario.Experimental results on several widelyused language pairs show that our approach outperforms two strong baselines (XLM and MASS) by remedying the style and content gaps. 1 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 38.4 33.6 29.5 33.9 33.7 32.5 33.6 XLM 37.4 34.5 27.2 34.3 34.6 32.7 33.5 MASS 37.8 34.9 27.1 35.2 35.1 33.4 33.9 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 37.3 33.4 29.7 33.8 33.8 32.4 33.4 XLM 36.3 34.3 27.4 34.1 34.8 32.4 33.2 MASS 36.6 34.7 27.3 35.1 35.2 33.0 33.7 Zhiwei He 0002, Xing Wang 0007, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu |
ACL (1) | 1 |