Yufei Wang 0005

dblp:61/5568-5 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021
YearPublicationVenuePosition
2026 RubricBench: Aligning Model-Generated Rubrics with Human Standards
abstract
Junyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junyi Zhou 0007, Qiyuan Zhang 0001, Yufei Wang 0005, Fuyuan Lyu, Yidong Ming, Can Xu 0002, Qingfeng Sun, Kai Zheng 0001, Xue (Steve) Liu, Chen Ma 0001
ACL (1)3
2025 Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
abstract
Qiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Qiyuan Zhang 0001, Yufei Wang 0005, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ACL (1)2
2025 From General Reward to Targeted Reward: Improving Open-ended Long-context Generation Models
abstract
Current research on long-form context in Large Language Models (LLMs) primarily focuses on the understanding of long-contexts, the Openended Long Text Generation (Open-LTG) remains insufficiently explored.Training a longcontext generation model requires curation of gold-standard reference data, which is typically nonexistent for informative Open-LTG tasks.However, previous methods only utilize general assessments as reward signals, which limits accuracy.To bridge this gap, we introduce ProxyReward, an innovative reinforcement learning (RL) based framework, which includes a dataset and a reward signal computation method.Firstly, ProxyReward Dataset generation is accomplished through simple prompts that enables the model to create automatically, obviating extensive labeled data or significant manual effort.Secondly, ProxyReward Signal offers a targeted evaluation of information comprehensiveness and accuracy for specific questions.The experimental results indicate that our method Prox-yReward surpasses even GPT-4-Turbo.It can significantly enhance performance by 20% on the Open-LTG task when training widely used open-source models, while also surpassing the LLM-as-a-Judge approach.Our work presents effective methods to enhance the ability of LLMs to address complex open-ended questions posed by humans.
Zhihan Guo, Jiele Wu, Wenqian Cui, Minda Hu, Yufei Wang 0005, Irwin King
EMNLP6
2025 NILE: Internal Consistency Alignment in Large Language Models
abstract
Minda Hu, Qiyuan Zhang, Yufei Wang, Bowei He, Hongru Wang, Jingyan Zhou, Liangyou Li, Yasheng Wang, Chen Ma, Irwin King. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Minda Hu, Qiyuan Zhang 0001, Yufei Wang 0005, Bowei He, Hongru Wang 0003, Jingyan Zhou, Liangyou Li, Yasheng Wang, Chen Ma 0001, Irwin King
EMNLP3
2025 Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP4
2025 G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
abstract
Large language models (LLMs) have shown remarkable proficiency in human-level reasoning and generation capabilities, which encourages extensive research on their application in mathematical problem solving. However, current work has been largely focused on text-based mathematical problems, with limited investigation in problems involving multi-modal geometric information. Addressing this gap, we aim to enable LLMs to solve geometric problems by understanding image input. We first identify the limitations of current Multimodal Large Language Models (MLLMs) in this area: they struggle to accurately comprehend basic geometric elements and their relationships. To address these challenges, we leverage the inherent attribute of logical structure compactness in geometric figures, utilizing text-only Large Language Models (LLMs) to curate a comprehensive multimodal geometry dataset. This dataset, named Geo170k, contains more than 170K geometric image-caption and question-answer pairs. Utilizing the Geo170k dataset, we introduce G-LLaVA, a model that demonstrates exceptional performance in solving geometric problems. It significantly outperforms GPT4-V on the geometry task of MathVista benchmark with only 7B parameters.
Jiahui Gao 0002, Renjie Pi, Jiacheng Ye, Wanjun Zhong, Yufei Wang 0005, Lanqing Hong, Jianhua Han, Hang Xu 0004, Zhenguo Li, Lingpeng Kong
ICLR6
2025 Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
abstract
Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are typically isolated, leading to weak correlations between them as well as suboptimal alignment performance. To address this issue, we propose an effective framework for Bridging and Modeling Correlations in pairwise data, named BMC. Firstly, we increase the consistency and informativeness of the pairwise preference signals through targeted modifications, synthesizing a pseudo-winning response by improving the losing response with the winning response as a reference. Secondly, we identify that DPO alone is insufficient to model these correlations and capture nuanced variations. Therefore, we propose learning token-level correlations by dynamically leveraging the policy model's confidence during training. Comprehensive experiments on QA, math, and instruction-following tasks demonstrate the effectiveness of our approach, significantly surpassing competitive baselines, including DPO. Additionally, our in-depth quantitative analysis reveals the reasons behind our method's superior performance over DPO and showcases its versatility to other DPO variants.
Bo Huang 0017, Yufei Wang 0005, Xingshan Zeng, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Wei Wang 0011
ICLR3
2025 RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
abstract
With significant efforts in recent studies, LLM-as-a-Judge has become a cost-effective alternative to human evaluation for assessing text generation quality in a wide range of tasks. However, there still remains a reliability gap between LLM-as-a-Judge and human evaluation. One important reason is the lack of guided oracles in the evaluation process. Motivated by the role of reference pervasively used in classic text evaluation, we introduce RevisEval, a novel text generation evaluation paradigm via the response-adapted references. RevisEval is driven by the key observation that an ideal reference should maintain the necessary relevance to the response to be evaluated. Specifically, RevisEval leverages the text revision capabilities of large language models (LLMs) to adaptively revise the response, then treat the revised text as the reference (response-adapted reference) for the subsequent evaluation. Extensive experiments demonstrate that RevisEval outperforms traditional reference-free and reference-based evaluation paradigms that use LLM-as-a-Judge across NLG tasks and open-ended instruction-following tasks. More importantly, our response-adapted references can further boost the classical text metrics, e.g., BLEU and BERTScore, compared to traditional references and even rival the LLM-as-a-Judge. A detailed analysis is also conducted to confirm RevisEval's effectiveness in bias reduction, the impact of inference cost, and reference relevance.
Qiyuan Zhang 0001, Yufei Wang 0005, Tiezheng Yu, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ICLR2
2024 FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
abstract
Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011
ACL (1)2
2024 Learning to Edit: Aligning LLMs with Knowledge Editing
abstract
Yuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao 0002, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Qun Liu 0001, Wei Wang 0011
ACL (1)2
2024 M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun, Liangyou Li, Yuxin Jiang, Lifeng Shang, Qun Liu, Kam-Fai Wong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Yusen Sun, Liangyou Li, Lifeng Shang, Qun Liu 0001, Kam-Fai Wong
ACL (1)3
2024 MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Liangyou Li, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP4