Princeton researchers tested an autonomous research agent by asking it to tackle the research questions behind two machine-learning papers, then having those papers' original authors grade the results. The agent handled the engineering well, running hundreds of experiments over days without getting stuck or faking results, but its papers scored 2 out of 6 and 1 out of 6. It tended to commit to one idea too early and never revisit it.
What changed
Earlier claims for automated AI research rested on conference peer review, which the authors call unreliable.
What it unlocks
A tougher way to test research agents: have the original authors of a paper grade an agent's attempt at the same research question.
- scores of 2/6 and 1/6
- 6 days per paper
- $3,000 compute credits per paper
- 2 papers evaluated
- nature.com2026-08-13