Binghai Wang

dblp:352/2462 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 57% Language models and text generation · 32% Trustworthy machine learning · 11%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
1.012026
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress · WWW 2026
Natural language and speech › Language models and text generation › alignment
reasoning alignment
1.012026
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
RMB: Comprehensively benchmarking reward models in LLM alignment · ICLR 2025
Machine learning › Reinforcement learning › reward learning › reward modeling
reward model evaluation
0.912025
RMB: Comprehensively benchmarking reward models in LLM alignment · ICLR 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning · EMNLP 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
0.812024
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning · EMNLP 2024
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.312026
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
0.312026
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress · WWW 2026

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 1.0reward modeling · 1.0generalized advantage estimation · 1.0pairwise evaluation · 0.9majority voting · 0.9best-of-n evaluation · 0.9contrastive learning · 0.8
YearPublicationVenuePosition
2026 Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
abstract
Binghai Wang, Yantao Liu, Yuxuan Liu, Tianyi Tang, Shenzhi Wang, Chang Gao, Chujie Zheng, Yichang Zhang, Le Yu, Shixuan Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Bowen Yu, Fei Huang, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Binghai Wang, Yantao Liu, Shenzhi Wang, Chujie Zheng, Yichang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Bowen Yu 0002, Fei Huang 0002, Junyang Lin
ACL (1)1
2026 AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
abstract
Despite rapid development, large language models (LLMs) still encounter challenges in multi-turn decision-making tasks (i.e., agent tasks) like web shopping and browser navigation, which require making a sequence of intelligent decisions based on environmental feedback. Previous work for LLM agents typically relies on elaborate prompt engineering or fine-tuning with expert trajectories to improve performance. In this work, we take a different perspective: we explore constructing process reward models (PRMs) to evaluate each decision and guide the agent's decision-making process. Unlike LLM reasoning, where each step is scored based on correctness, actions in agent tasks do not have a clear-cut correctness. Instead, they should be evaluated based on their proximity to the goal and the progress they have made. Building on this insight, we propose a re-defined PRM for agent tasks, named AgentPRM, to capture both the interdependence between sequential decisions and their contribution to the final goal. This enables better progress tracking and exploration-exploitation balance. To scalably obtain labeled data for training AgentPRM, we employ a Temporal Difference-based (TD-based) estimation method combined with Generalized Advantage Estimation (GAE), which proves more sample-efficient than prior methods. Extensive experiments across different agentic tasks show that AgentPRM is over 8× more compute-efficient than baselines, and it demonstrates robust improvement when scaling up test-time compute. Moreover, we perform detailed analyses to show how our method works and offer more insights, e.g., applying AgentPRM to the reinforcement learning of LLM agents.
Zhiheng Xi, Chenyang Liao, Zhihao Zhang 0002, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou 0005, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
WWW6
2025 RMB: Comprehensively benchmarking reward models in LLM alignment
abstract
Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization. We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. We will release our evaluation code and datasets upon publication.
Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Limao Xiong, Jessica Fan, Yurong Mou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICLR3
2024 Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
abstract
Lu Chen, Rui Zheng, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye, Zhihao Zhang, Yuhao Zhou, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lu Chen 0001, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye 0005, Zhihao Zhang 0002, Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP3