EDBT 2026 Demo / reviewers in the wild / expert
Weizhe Yuan
dblp:207/1964
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-6117-9417ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeabstractTianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason E Weston, Sainbayar Sukhbaatar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Tianhao Wu 0002, Weizhe Yuan, Olga Golovneva, Jing Xu 0014, Yuandong Tian, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar |
EMNLP | 2 |
| 2025 | Following Length Constraints in InstructionsabstractAligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral. Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014 |
EMNLP | 1 |
| 2025 | Thinking LLMs: General Instruction Following with Thought GenerationabstractLLMs are typically trained to answer user questions or follow instructions similarly to how human experts respond. However, in the standard alignment framework they lack the basic ability of explicit thinking before answering. Thinking is important for complex questions that require reasoning and planning – but can be applied to any task. We propose a training method for equipping existing LLMs with such thinking abilities for general instruction following without use of additional human data. We achieve this by an iterative search and optimization procedure that explores the space of possible thought generations, allowing the model to learn how to think without direct supervision. For each instruction, the thought candidates are scored using a judge model to evaluate their responses only, and then optimized via preference optimization. We show that this procedure leads to superior performance on AlpacaEval and Arena-Hard, and shows gains from thinking on non-reasoning categories such as marketing, health and general knowledge, in addition to more traditional reasoning & problem-solving tasks. Tianhao Wu 0002, Janice Lan, Weizhe Yuan, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar |
ICML | 3 |
| 2025 | Self-Consistency Preference OptimizationabstractSelf-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku. Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu 0014, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, Jane Dwivedi-Yu |
ICML | 2 |
| 2025 | R.I.P.: Better Models by Survival of the Fittest PromptsabstractTraining data quality is one of the most important drivers of final model quality. In this work, we introduce a method for evaluating data integrity based on the assumption that low-quality input prompts result in high variance and low quality responses. This is achieved by measuring the rejected response quality and the reward gap between the chosen and rejected preference pair. Our method, Rejecting Instruction Preferences (RIP) can be used to filter prompts from existing training sets, or to make high quality synthetic datasets, yielding large performance gains across various benchmarks compared to unfiltered data. Using Llama 3.1-8B-Instruct, RIP improves AlpacaEval2 LC Win Rate by 9.4%, Arena-Hard by 8.7%, and WildBench by 9.9%. Using Llama 3.3-70B-Instruct, RIP improves Arena-Hard from 67.5 to 82.9, from 18th place to 6th overall in the leaderboard. Weizhe Yuan, Olga Golovneva, Tianhao Wu 0002, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014 |
ICML | 2 |
| 2025 | NaturalReasoning: Reasoning in the Wild with 2.8M Challenging QuestionsabstractScaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding. Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003 |
NeurIPS | 1 |
| 2024 | System-Level Natural Language FeedbackabstractNatural language (NL) feedback offers rich insights into user experience.While existing studies focus on an instance-level approach, where feedback is used to refine specific examples, we introduce a framework for system-level use of NL feedback.We show how to use feedback to formalize system-level design decisions in a human-in-the-loop-process -in order to produce better models.In particular this is done through: (i) metric design for tasks; and (ii) language model prompt design for refining model responses.We conduct two case studies of this approach for improving search query and dialog response generation, demonstrating the effectiveness of system-level feedback.We show the combination of system-level and instancelevel feedback brings further gains, and that human written instance-level feedback results in more grounded refinements than GPT-3.5 written ones, underlying the importance of human feedback for building systems.We release our code and data at https://github.com/ yyy-Apple/Sys-NL-Feedback. Weizhe Yuan, Kyunghyun Cho, Jason Weston |
EACL (1) | 1 |
| 2024 | Generative Judge for Evaluating AlignmentabstractThe rapid development of Large Language Models (LLMs) has substantially expanded the range of tasks they can address. In the field of Natural Language Processing (NLP), researchers have shifted their focus from conventional NLP tasks (e.g., sequence tagging and parsing) towards tasks that revolve around aligning with human needs (e.g., brainstorming and email writing). This shift in task distribution imposes new requirements on evaluating these aligned models regarding *generality* (i.e., assessing performance across diverse scenarios), *flexibility* (i.e., examining under different protocols), and *interpretability* (i.e., scrutinizing models with explanations). In this paper, we propose a generative judge with 13B parameters, **Auto-J**, designed to address these challenges. Our model is trained on user queries and LLM-generated responses under massive real-world scenarios and accommodates diverse evaluation protocols (e.g., pairwise response comparison and single-response evaluation) with well-structured natural language critiques. To demonstrate the efficacy of our approach, we construct a new testbed covering 58 different scenarios. Experimentally, **Auto-J** outperforms a series of strong competitors, including both open-source and closed-source models, by a large margin. We also provide detailed analysis and case studies to further reveal the potential of our method and make a variety of resources public at https://github.com/GAIR-NLP/auto-j. Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao 0001, Pengfei Liu 0003 |
ICLR | 3 |
| 2024 | Self-Rewarding Language ModelsabstractWe posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these reward models require additional human preferences data to further improve.In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training, not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes. Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 0003, Sainbayar Sukhbaatar, Jing Xu 0014, Jason Weston |
ICML | 1 |
| 2024 | Iterative Reasoning Preference OptimizationabstractIterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks. In this work we develop an iterative approach that optimizes the preference between competing generated Chain-of-Thought (CoT) candidates by optimizing for winning vs. losing reasoning steps. We train using a modified DPO loss with an additional negative log-likelihood term, which we find to be crucial. We show reasoning improves across repeated iterations of this scheme. While only relying on examples in the training set, our approach results in increasing accuracy on GSM8K, MATH, and ARC-Challenge for Llama-2-70B-Chat, outperforming other Llama-2-based models not relying on additionally sourced datasets. For example, we see a large improvement from 55.6% to 81.6% on GSM8K and an accuracy of 88.7% with majority voting out of 32 samples. Richard Yuanzhe Pang, Weizhe Yuan, He He 0001, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston |
NeurIPS | 2 |
| 2023 | CFFMixer: Multi-Dimensional Feature Fusion for Object DetectionabstractObject detection is a fundamental task in the field of computer vision, and one of its essential requirements is high-quality feature fusion. Previous works have made various efforts in this regard: CNN-based detectors use convolutional blocks to fuse local features and dense prior knowledge to predict objects, while query-based detectors fuse global features by self-attention then decode features with object queries. However, their feature fusion methods are relatively monotonous. Considering that different modules are applicable to different dimensions, we proposed an object detector named CFFMixer which used hybrid architecture to achieve multi-dimensional feature fusion. The sampling strategy to extract abundant local and global features was first introduced then the Comprehensive Feature Fusion Network (CFFN) was proposed to integrate them. CFFN not only achieved local and global features interaction in the spatial dimension, but also fused semantics in the channel dimension. Furthermore, we conducted experiments and made a comparison with competitive models, our model finally got 43.0 mAP on COCO 2017 dataset within 12 epochs. Experimental results showed that the model’s accuracy benefits from the powerful feature fusion capability of CFFN. Besides, we performed ablation studies on our modules to evaluate their effectiveness. Weizhe Yuan, Bin Kang, Songlin Du |
ICASSP | 2 |
| 2022 | KID-Review: Knowledge-Guided Scientific Review Generation with Oracle Pre-trainingabstractThe surge in the number of scientific submissions has brought challenges to the work of peer review. In this paper, as a first step, we explore the possibility of designing an automated system, which is not meant to replace humans, but rather providing a first-pass draft for a machine-assisted human review process. Specifically, we present an end-to-end knowledge-guided review generation framework for scientific papers grounded in cognitive psychology research that a better understanding of text requires different types of knowledge. In practice, we found that this seemingly intuitive idea suffered from training difficulties. In order to solve this problem, we put forward an oracle pre-training strategy, which can not only make the Kid-Review better educated but also make the generated review cover more aspects. Experimentally, we perform a comprehensive evaluation (human and automatic) from different perspectives. Empirical results have shown the effectiveness of different types of knowledge as well as oracle pre-training. We make all code, relevant dataset available: https://github.com/Anonymous4nlp233/KIDReview as well as the Kid-Review system: http://nlpeer.reviews. Weizhe Yuan, Pengfei Liu 0003 |
AAAI | 1 |
| 2022 | Can We Automate Scientific Reviewing?abstractThe rapid development of science and technology has been accompanied by an exponential growth in peer-reviewed scientific publications. At the same time, the review of each paper is a laborious process that must be carried out by subject matter experts. Thus, providing high-quality reviews of this growing number of papers is a significant challenge. In this work, we ask the question “can we automate scientific reviewing? ”, discussing the possibility of using natural language processing (NLP) models to generate peer reviews for scientific papers. Because it is non-trivial to define what a “good” review is in the first place, we first discuss possible evaluation metrics that could be used to judge success in this task. We then focus on the machine learning domain and collect a dataset of papers in the domain, annotate them with different aspects of content covered in each review, and train targeted summarization models that take in papers as input and generate reviews as output. Comprehensive experimental results on the test set show that while system-generated reviews are comprehensive, touching upon more aspects of the paper than human-written reviews, the generated texts are less constructive and less factual than human-written reviews for all aspects except the explanation of the core ideas of the papers, which are largely factually correct. Given these results, we pose eight challenges in the pursuit of a good review generation system together with potential solutions, which, hopefully, will inspire more future research in this direction. We make relevant resource publicly available for use by future research: https://github. com/neulab/ReviewAdvisor. In addition, while our conclusion is that the technology is not yet ready for use in high-stakes review settings we provide a system demo, ReviewAdvisor (http://review.nlpedia.ai/), showing the current capabilities and failings of state-of-the-art NLP models at this task (see demo screenshot in A.2). A review of this paper written by the system proposed in this paper can be found in A.1. Weizhe Yuan, Pengfei Liu 0003, Graham Neubig |
J. Artif. Intell. Res. | 1 |
| 2021 | BARTScore: Evaluating Generated Text as Text GenerationabstractA wide variety of NLP applications, such as machine translation, summarization, and dialog, involve text generation. One major challenge for these applications is how to evaluate whether such generated texts are actually fluent, accurate, or effective. In this work, we conceptualize the evaluation of generated text as a text generation problem, modeled using pre-trained sequence-to-sequence models. The general idea is that models trained to convert the generated text to/from a reference output or the source text will achieve higher scores when the generated text is better. We operationalize this idea using BART, an encoder-decoder based pre-trained model, and propose a metric BARTScore with a number of variants that can be flexibly applied in an unsupervised fashion to evaluation of text from different perspectives (e.g. informativeness, fluency, or factuality). BARTScore is conceptually simple and empirically effective. It can outperform existing top-scoring metrics in 16 of 22 test settings, covering evaluation of 16 datasets (e.g., machine translation, text summarization) and 7 different perspectives (e.g., informativeness, factuality). Code to calculate BARTScore is available at https://github.com/neulab/BARTScore, and we have released an interactive leaderboard for meta-evaluation at http://explainaboard.nlpedia.ai/leaderboard/task-meval/ on the ExplainaBoard platform, which allows us to interactively understand the strengths, weaknesses, and complementarity of each metric. Weizhe Yuan, Graham Neubig, Pengfei Liu 0003 |
NeurIPS | 1 |
| 2017 | LiveJack: Integrating CDNs and Edge Clouds for Live Content BroadcastingabstractEmerging commercial live content broadcasting platforms are facing great challenges to accommodate large scale dynamic viewer populations. Existing solutions constantly suffer from balancing the cost of deploying at the edge close to the viewers and the quality of content delivery. We propose LiveJack, a novel network service to allow CDN servers to seamlessly leverage ISP edge cloud resources. LiveJack can elastically scale the serving capacity of CDN servers by integrating Virtual Media Functions (VMF) in the edge cloud to accommodate flash crowds for very popular contents. LiveJack introduces minor application layer changes for streaming service providers and is completely transparent to end users. We have prototyped LiveJack in both LAN and WAN environments. Evaluations demonstrate that LiveJack can increase CDN server capacity by more than six times, and can effectively accommodate highly dynamic workloads with an improved service quality. Bo Yan 0004, Shu Shi, Yong Liu 0013, Weizhe Yuan, Haoqin He, Rittwik Jana, Yang Xu 0010, H. Jonathan Chao |
ACM Multimedia | 4 |