Weizhe Yuan

dblp:207/1964 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-6117-9417ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
abstract
Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason E Weston, Sainbayar Sukhbaatar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Tianhao Wu 0002, Weizhe Yuan, Olga Golovneva, Jing Xu 0014, Yuandong Tian, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar
EMNLP2
2025 Following Length Constraints in Instructions
abstract
Aligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral.
Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014
EMNLP1
2025 Thinking LLMs: General Instruction Following with Thought Generation
abstract
LLMs are typically trained to answer user questions or follow instructions similarly to how human experts respond. However, in the standard alignment framework they lack the basic ability of explicit thinking before answering. Thinking is important for complex questions that require reasoning and planning – but can be applied to any task. We propose a training method for equipping existing LLMs with such thinking abilities for general instruction following without use of additional human data. We achieve this by an iterative search and optimization procedure that explores the space of possible thought generations, allowing the model to learn how to think without direct supervision. For each instruction, the thought candidates are scored using a judge model to evaluate their responses only, and then optimized via preference optimization. We show that this procedure leads to superior performance on AlpacaEval and Arena-Hard, and shows gains from thinking on non-reasoning categories such as marketing, health and general knowledge, in addition to more traditional reasoning & problem-solving tasks.
Tianhao Wu 0002, Janice Lan, Weizhe Yuan, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar
ICML3
2025 Self-Consistency Preference Optimization
abstract
Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku.
Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu 0014, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, Jane Dwivedi-Yu
ICML2
2025 R.I.P.: Better Models by Survival of the Fittest Prompts
abstract
Training data quality is one of the most important drivers of final model quality. In this work, we introduce a method for evaluating data integrity based on the assumption that low-quality input prompts result in high variance and low quality responses. This is achieved by measuring the rejected response quality and the reward gap between the chosen and rejected preference pair. Our method, Rejecting Instruction Preferences (RIP) can be used to filter prompts from existing training sets, or to make high quality synthetic datasets, yielding large performance gains across various benchmarks compared to unfiltered data. Using Llama 3.1-8B-Instruct, RIP improves AlpacaEval2 LC Win Rate by 9.4%, Arena-Hard by 8.7%, and WildBench by 9.9%. Using Llama 3.3-70B-Instruct, RIP improves Arena-Hard from 67.5 to 82.9, from 18th place to 6th overall in the leaderboard.
Weizhe Yuan, Olga Golovneva, Tianhao Wu 0002, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014
ICML2
2025 NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
abstract
Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding.
Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003
NeurIPS1
2024 System-Level Natural Language Feedback
abstract
Natural language (NL) feedback offers rich insights into user experience.While existing studies focus on an instance-level approach, where feedback is used to refine specific examples, we introduce a framework for system-level use of NL feedback.We show how to use feedback to formalize system-level design decisions in a human-in-the-loop-process -in order to produce better models.In particular this is done through: (i) metric design for tasks; and (ii) language model prompt design for refining model responses.We conduct two case studies of this approach for improving search query and dialog response generation, demonstrating the effectiveness of system-level feedback.We show the combination of system-level and instancelevel feedback brings further gains, and that human written instance-level feedback results in more grounded refinements than GPT-3.5 written ones, underlying the importance of human feedback for building systems.We release our code and data at https://github.com/ yyy-Apple/Sys-NL-Feedback.
Weizhe Yuan, Kyunghyun Cho, Jason Weston
EACL (1)1
2024 Generative Judge for Evaluating Alignment
abstract
The rapid development of Large Language Models (LLMs) has substantially expanded the range of tasks they can address. In the field of Natural Language Processing (NLP), researchers have shifted their focus from conventional NLP tasks (e.g., sequence tagging and parsing) towards tasks that revolve around aligning with human needs (e.g., brainstorming and email writing). This shift in task distribution imposes new requirements on evaluating these aligned models regarding *generality* (i.e., assessing performance across diverse scenarios), *flexibility* (i.e., examining under different protocols), and *interpretability* (i.e., scrutinizing models with explanations). In this paper, we propose a generative judge with 13B parameters, **Auto-J**, designed to address these challenges. Our model is trained on user queries and LLM-generated responses under massive real-world scenarios and accommodates diverse evaluation protocols (e.g., pairwise response comparison and single-response evaluation) with well-structured natural language critiques. To demonstrate the efficacy of our approach, we construct a new testbed covering 58 different scenarios. Experimentally, **Auto-J** outperforms a series of strong competitors, including both open-source and closed-source models, by a large margin. We also provide detailed analysis and case studies to further reveal the potential of our method and make a variety of resources public at https://github.com/GAIR-NLP/auto-j.
Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao 0001, Pengfei Liu 0003
ICLR3
2024 Self-Rewarding Language Models
abstract
We posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these reward models require additional human preferences data to further improve.In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training, not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes.
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 0003, Sainbayar Sukhbaatar, Jing Xu 0014, Jason Weston
ICML1
2024 Iterative Reasoning Preference Optimization
abstract
Iterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks. In this work we develop an iterative approach that optimizes the preference between competing generated Chain-of-Thought (CoT) candidates by optimizing for winning vs. losing reasoning steps. We train using a modified DPO loss with an additional negative log-likelihood term, which we find to be crucial. We show reasoning improves across repeated iterations of this scheme. While only relying on examples in the training set, our approach results in increasing accuracy on GSM8K, MATH, and ARC-Challenge for Llama-2-70B-Chat, outperforming other Llama-2-based models not relying on additionally sourced datasets. For example, we see a large improvement from 55.6% to 81.6% on GSM8K and an accuracy of 88.7% with majority voting out of 32 samples.
Richard Yuanzhe Pang, Weizhe Yuan, He He 0001, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston
NeurIPS2
2023 CFFMixer: Multi-Dimensional Feature Fusion for Object Detection
abstract
Object detection is a fundamental task in the field of computer vision, and one of its essential requirements is high-quality feature fusion. Previous works have made various efforts in this regard: CNN-based detectors use convolutional blocks to fuse local features and dense prior knowledge to predict objects, while query-based detectors fuse global features by self-attention then decode features with object queries. However, their feature fusion methods are relatively monotonous. Considering that different modules are applicable to different dimensions, we proposed an object detector named CFFMixer which used hybrid architecture to achieve multi-dimensional feature fusion. The sampling strategy to extract abundant local and global features was first introduced then the Comprehensive Feature Fusion Network (CFFN) was proposed to integrate them. CFFN not only achieved local and global features interaction in the spatial dimension, but also fused semantics in the channel dimension. Furthermore, we conducted experiments and made a comparison with competitive models, our model finally got 43.0 mAP on COCO 2017 dataset within 12 epochs. Experimental results showed that the model’s accuracy benefits from the powerful feature fusion capability of CFFN. Besides, we performed ablation studies on our modules to evaluate their effectiveness.
Weizhe Yuan, Bin Kang, Songlin Du
ICASSP2
2022 KID-Review: Knowledge-Guided Scientific Review Generation with Oracle Pre-training
abstract
The surge in the number of scientific submissions has brought challenges to the work of peer review. In this paper, as a first step, we explore the possibility of designing an automated system, which is not meant to replace humans, but rather providing a first-pass draft for a machine-assisted human review process. Specifically, we present an end-to-end knowledge-guided review generation framework for scientific papers grounded in cognitive psychology research that a better understanding of text requires different types of knowledge. In practice, we found that this seemingly intuitive idea suffered from training difficulties. In order to solve this problem, we put forward an oracle pre-training strategy, which can not only make the Kid-Review better educated but also make the generated review cover more aspects. Experimentally, we perform a comprehensive evaluation (human and automatic) from different perspectives. Empirical results have shown the effectiveness of different types of knowledge as well as oracle pre-training. We make all code, relevant dataset available: https://github.com/Anonymous4nlp233/KIDReview as well as the Kid-Review system: http://nlpeer.reviews.
Weizhe Yuan, Pengfei Liu 0003
AAAI1
2022 Can We Automate Scientific Reviewing?
abstract
The rapid development of science and technology has been accompanied by an exponential growth in peer-reviewed scientific publications. At the same time, the review of each paper is a laborious process that must be carried out by subject matter experts. Thus, providing high-quality reviews of this growing number of papers is a significant challenge. In this work, we ask the question “can we automate scientific reviewing? ”, discussing the possibility of using natural language processing (NLP) models to generate peer reviews for scientific papers. Because it is non-trivial to define what a “good” review is in the first place, we first discuss possible evaluation metrics that could be used to judge success in this task. We then focus on the machine learning domain and collect a dataset of papers in the domain, annotate them with different aspects of content covered in each review, and train targeted summarization models that take in papers as input and generate reviews as output. Comprehensive experimental results on the test set show that while system-generated reviews are comprehensive, touching upon more aspects of the paper than human-written reviews, the generated texts are less constructive and less factual than human-written reviews for all aspects except the explanation of the core ideas of the papers, which are largely factually correct. Given these results, we pose eight challenges in the pursuit of a good review generation system together with potential solutions, which, hopefully, will inspire more future research in this direction. We make relevant resource publicly available for use by future research: https://github. com/neulab/ReviewAdvisor. In addition, while our conclusion is that the technology is not yet ready for use in high-stakes review settings we provide a system demo, ReviewAdvisor (http://review.nlpedia.ai/), showing the current capabilities and failings of state-of-the-art NLP models at this task (see demo screenshot in A.2). A review of this paper written by the system proposed in this paper can be found in A.1.
Weizhe Yuan, Pengfei Liu 0003, Graham Neubig
J. Artif. Intell. Res.1
2021 BARTScore: Evaluating Generated Text as Text Generation
abstract
A wide variety of NLP applications, such as machine translation, summarization, and dialog, involve text generation. One major challenge for these applications is how to evaluate whether such generated texts are actually fluent, accurate, or effective. In this work, we conceptualize the evaluation of generated text as a text generation problem, modeled using pre-trained sequence-to-sequence models. The general idea is that models trained to convert the generated text to/from a reference output or the source text will achieve higher scores when the generated text is better. We operationalize this idea using BART, an encoder-decoder based pre-trained model, and propose a metric BARTScore with a number of variants that can be flexibly applied in an unsupervised fashion to evaluation of text from different perspectives (e.g. informativeness, fluency, or factuality). BARTScore is conceptually simple and empirically effective. It can outperform existing top-scoring metrics in 16 of 22 test settings, covering evaluation of 16 datasets (e.g., machine translation, text summarization) and 7 different perspectives (e.g., informativeness, factuality). Code to calculate BARTScore is available at https://github.com/neulab/BARTScore, and we have released an interactive leaderboard for meta-evaluation at http://explainaboard.nlpedia.ai/leaderboard/task-meval/ on the ExplainaBoard platform, which allows us to interactively understand the strengths, weaknesses, and complementarity of each metric.
Weizhe Yuan, Graham Neubig, Pengfei Liu 0003
NeurIPS1
2017 LiveJack: Integrating CDNs and Edge Clouds for Live Content Broadcasting
abstract
Emerging commercial live content broadcasting platforms are facing great challenges to accommodate large scale dynamic viewer populations. Existing solutions constantly suffer from balancing the cost of deploying at the edge close to the viewers and the quality of content delivery. We propose LiveJack, a novel network service to allow CDN servers to seamlessly leverage ISP edge cloud resources. LiveJack can elastically scale the serving capacity of CDN servers by integrating Virtual Media Functions (VMF) in the edge cloud to accommodate flash crowds for very popular contents. LiveJack introduces minor application layer changes for streaming service providers and is completely transparent to end users. We have prototyped LiveJack in both LAN and WAN environments. Evaluations demonstrate that LiveJack can increase CDN server capacity by more than six times, and can effectively accommodate highly dynamic workloads with an improved service quality.
Bo Yan 0004, Shu Shi, Yong Liu 0013, Weizhe Yuan, Haoqin He, Rittwik Jana, Yang Xu 0010, H. Jonathan Chao
ACM Multimedia4