Yifu Huo

dblp:354/6320 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 49% Language models and text generation · 22% Trustworthy machine learning · 14%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reinforcement learning from human feedback
3.152026
GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025
ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation · AAAI 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
1.922026
GRAM-R²: Self-Training Generative Foundation Reward Models for Reward Reasoning · AAAI 2026
GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025
Natural language and speech › Language models and text generation
alignment
1.012026
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model
1.012026
GRAM-R²: Self-Training Generative Foundation Reward Models for Reward Reasoning · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026
Machine learning › Reinforcement learning › reward learning › reward modeling
reward model evaluation
1.012026
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
reward model interpretability
1.012026
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026
Natural language and speech › Language models and text generation › alignment
preference alignment
0.912025
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025
Computer vision › Vision and language › vision-language model
vision-language model alignment
0.912025
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.812024
ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation · AAAI 2024
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312025
GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.212024
ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation · AAAI 2024

Methods — techniques the papers use, named apart from their topics

self-training · 1.0probing · 1.0preference learning · 1.0policy optimization · 1.0multi-objective optimization · 1.0optimal transport · 0.9label smoothing · 0.9generative model · 0.9direct preference optimization · 0.9best-of-n sampling · 0.9
YearPublicationVenuePosition
2026 Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
abstract
Previous methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with LLM alignment performance, supporting it as a reliable reference for developing advanced reward models. By analyzing the evaluation results on MRMBench, we reveal that reward models struggle to simultaneously capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Furthermore, our results demonstrate that the proposed inference-time probing method provides a reliable metric for assessing the confidence of reward predictions, leading to improved alignment of large language models.
Chenglong Wang 0002, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu 0001, Tong Xiao 0001
AAAI2
2026 GRAM-R²: Self-Training Generative Foundation Reward Models for Reward Reasoning
abstract
Major progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs to generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short in instilling explicit reasoning capabilities into reward models. To bridge this gap, we propose a self-training approach that can leverage unlabeled data to scale up reward reasoning in reward models. Based on this approach, we develop GRAM-R² a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R² can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as policy optimization and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R² consistently delivers strong performance, outperforming several strong discriminative and generative baselines.
Chenglong Wang 0002, Yongyu Mu, Yifu Huo, Jiali Zeng, Murun Yang, Xiaoyang Hao, Chunliang Zhang, Fandong Meng, Tong Xiao 0001
AAAI4
2025 RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
abstract
Large vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization.
Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao 0001, Chunliang Zhang, Tongran Liu
AAAI3
2025 GRAM: A Generative Foundation Reward Model for Reward Generalization
abstract
In aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward models using both unlabeled and labeled data. Building on the generative models in LLMs, we develop a generative reward model that is first trained via large-scale unsupervised learning and then fine-tuned via supervised learning. We also show that by using label smoothing, we are in fact optimizing a regularized pairwise ranking loss. This result, in turn, provides a new view of training reward models, which links generative models and discriminative models under the same class of training objectives. The outcome of these techniques is a foundation reward model, which can be applied to a wide range of tasks with little or no further fine-tuning effort. Extensive experiments show that this model generalizes well across several tasks, including response ranking, reinforcement learning from human feedback, and task adaptation with fine-tuning, achieving significant performance improvements over several strong baseline models.
Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Tong Xiao 0001, Chunliang Zhang, Tongran Liu
ICML3
2024 ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation
abstract
Applying Reinforcement Learning (RL) to sequence generation models enables the direct optimization of long-term rewards (e.g., BLEU and human feedback), but typically requires large-scale sampling over a space of action sequences. This is a computational challenge as presented by the practice of sequence generation problems, such as machine translation, where we often deal with a large action space (e.g., a vocabulary) and a long action sequence (e.g., a translation). In this work, we introduce two-stage sampling and dynamic sampling approaches to improve the sampling efficiency during training sequence generation models via RL. We experiment with our approaches on the traditional sequence generation tasks, including machine translation and abstractive summarization. Furthermore, we evaluate our approaches in RL from human feedback (RLHF) through training a large language model using the reward model. Experimental results show that the efficient sampling-based RL, referred to as ESRL, can outperform all baselines in terms of both training efficiency and memory consumption. Notably, ESRL yields consistent performance gains over the strong REINFORCE, minimum risk training, and proximal policy optimization methods. The code is available at https://github.com/wangclnlp/DeepSpeed-Chat-Extension/examples/esrl.
Chenglong Wang 0002, Yimin Hu, Yifu Huo, Tongran Liu, Tong Xiao 0001
AAAI4