VLDB 2026 Research / reviewers in the wild / expert
Daniel Sohn
dblp:232/4980
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 70% Reinforcement learning · 30% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
reward hacking |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Machine learning › Reinforcement learning › reward learning
reward model training |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
data augmentation · 0.9causal framework · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RRM: Robust Reward Model Training Mitigates Reward HackingabstractReward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%. Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh |
ICLR | 11 |