Researchers are using AI systems to audit published science, and the checks are turning up errors that have stood for decades. A chemist at Zhejiang Lab found wrong boiling-point values in a 75-year-old reference database after an AI prediction disagreed with the handbook, while separate groups used AI agents to test whether machine-learning conference papers could be reproduced and to count errors in NeurIPS papers. Specialists caution that the checking tools make mistakes too and their output still needs human review.
What changed
Checking published papers and reference databases for errors was slow manual work done by individual researchers.
What it unlocks
Automated re-checking of claims, calculations and reference values across large volumes of published literature.
- 168 ICML 2026 oral papers assessed by AI agents
- of 92 papers with at least five checkable claims, more than two of five claims reproduced for only 34
- more than 80% of claims reproduced for just 8 papers
- errors per NeurIPS paper rose from 3.8 in 2021 to 5.9 in 2025, a 55% increase
What you need to act on it
- human review of anything the automated checks flag
- nature.com2026-08-06