Daniel Sohn

dblp:232/4980 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 70% Reinforcement learning · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Natural language and speech › Language models and text generation › alignment
reward hacking
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Machine learning › Reinforcement learning › reward learning
reward model training
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.312025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025

Methods — techniques the papers use, named apart from their topics

data augmentation · 0.9causal framework · 0.9
YearPublicationVenuePosition
2025 RRM: Robust Reward Model Training Mitigates Reward Hacking
abstract
Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%.
Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh
ICLR11