Research2026-08-10

Researchers led by Nathan Lambert, with Valentina Pyatkin, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi and eight others, published RewardBench, a benchmark dataset and code base for evaluating reward models used in RLHF alignment. The dataset consists of prompt-chosen-rejected trios spanning chat, reasoning and safety, including comparison sets where the preferred answer is verifiably better for reasons such as code bugs or incorrect facts. An accompanying leaderboard evaluates reward models trained by different methods, including direct MLE classifier training and implicit reward models derived from Direct Preference Optimization. The paper reports findings on refusal propensity, reasoning limitations and instruction-following shortcomings across evaluated reward models. It was submitted to arXiv on 20 March 2024 and last revised on 8 June 2024, running 44 pages with 19 figures and 12 tables.

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.