Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiapeng Sun

dblp:327/9641 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 70% Language models and text generation · 23% Vision and language · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model
0.912025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.9grouped comparison · 0.9generative reward modeling · 0.9
YearPublicationVenuePosition
2025 Generative RLHF-V: Learning Principles from Multi-modal Human Preference
abstract
Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that integrates GRMs with multi-modal RLHF. We propose a two-stage pipeline: \textbf{multi-modal generative reward modeling from RL}, where RL guides GRMs to actively capture human intention, then predict the correct pair-wise scores; and \textbf{RL optimization from grouped comparison}, which enhances multi-modal RL scoring precision by grouped responses comparison. Experimental results demonstrate that, besides out-of-distribution generalization of RM discrimination, our framework improves 4 MLLMs' performance across 7 benchmarks by 18.1\%, while the baseline RLHF is only 5.3\%. We further validate that Generative RLHF-V achieves a near-linear improvement with an increasing number of candidate responses.
Jiaming Ji, Boyuan Chen 0008, Jiapeng Sun, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang 0001
NeurIPS4