Arena (formerly LMArena) launched AutoEval scores on its public leaderboards on 30 July 2026, providing provisional model rankings on the day a model launches rather than after days of accumulated human votes. AutoEval trains a pointwise reward model on Arena's human preference dataset of millions of pairwise comparisons, then substitutes reward-model "soft" votes for human votes while using the same Arena ranking methodology; entries are labeled "AutoEval" until validated by live votes. Arena reports its text reward model predicts human preferences 8-10% more accurately than frontier LLM judges (Gemini-3-flash/pro, GPT-5), and that a holdout test trained on data through April 2026 produced rank correlation above 0.98 with live scores, with over 90% head-to-head accuracy when gaps exceed 10 points across more than 40 test models.
- arena.ai2026-07-30