VLDB 2026 Research / reviewers in the wild / expert
Jiapeng Sun
dblp:327/9641
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 70% Language models and text generation · 23% Vision and language · 7% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model |
0.9 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.9grouped comparison · 0.9generative reward modeling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceabstractTraining multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that integrates GRMs with multi-modal RLHF. We propose a two-stage pipeline: \textbf{multi-modal generative reward modeling from RL}, where RL guides GRMs to actively capture human intention, then predict the correct pair-wise scores; and \textbf{RL optimization from grouped comparison}, which enhances multi-modal RL scoring precision by grouped responses comparison. Experimental results demonstrate that, besides out-of-distribution generalization of RM discrimination, our framework improves 4 MLLMs' performance across 7 benchmarks by 18.1\%, while the baseline RLHF is only 5.3\%. We further validate that Generative RLHF-V achieves a near-linear improvement with an increasing number of candidate responses. Jiaming Ji, Boyuan Chen 0008, Jiapeng Sun, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 4 |