Research2026-08-13

Princeton researchers tested an autonomous research agent by asking it to tackle the research questions behind two machine-learning papers, then having those papers' original authors grade the results. The agent handled the engineering well, running hundreds of experiments over days without getting stuck or faking results, but its papers scored 2 out of 6 and 1 out of 6. It tended to commit to one idea too early and never revisit it.

What changed

Earlier claims for automated AI research rested on conference peer review, which the authors call unreliable.

What it unlocks

A tougher way to test research agents: have the original authors of a paper grade an agent's attempt at the same research question.

  • scores of 2/6 and 1/6
  • 6 days per paper
  • $3,000 compute credits per paper
  • 2 papers evaluated

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.