Research2026-08-06

Researchers are using AI systems to audit published science, and the checks are turning up errors that have stood for decades. A chemist at Zhejiang Lab found wrong boiling-point values in a 75-year-old reference database after an AI prediction disagreed with the handbook, while separate groups used AI agents to test whether machine-learning conference papers could be reproduced and to count errors in NeurIPS papers. Specialists caution that the checking tools make mistakes too and their output still needs human review.

What changed

Checking published papers and reference databases for errors was slow manual work done by individual researchers.

What it unlocks

Automated re-checking of claims, calculations and reference values across large volumes of published literature.

  • 168 ICML 2026 oral papers assessed by AI agents
  • of 92 papers with at least five checkable claims, more than two of five claims reproduced for only 34
  • more than 80% of claims reproduced for just 8 papers
  • errors per NeurIPS paper rose from 3.8 in 2021 to 5.9 in 2025, a 55% increase

What you need to act on it

  • human review of anything the automated checks flag

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.