An AI model has exposed errors baked into a 75-year-old chemistry reference database — trusted boiling-point values that turned out to be wrong all along. Theoretical chemist Sebastian Pios of Zhejiang Lab in Hangzhou initially assumed his model was mistaken when it clashed with the handbook numbers; manual checks of the original literature showed the database was wrong. The same system caught a typo in a paper and incorrect century-old boiling-point measurements, errors that likely caused "a lot of trouble" for researchers who relied on them.

Pios is part of a growing movement using AI agents to audit science. In an analysis posted on 22 July, SAI Labs used agents to assess 168 papers selected for presentation at ICML 2026: of 92 papers with at least five assessable claims, agents reproduced more than two of five claims in only 34 papers. Separately, a Stanford team's arXiv study found objective errors per paper at NeurIPS rose from 3.8 in 2021 to 5.9 in 2025 — a 55% increase. Researchers caution the tools still "make mistakes like humans do" and need human oversight, but the advantage is scale: "The biggest difference is to be able to do this at a scale that was not possible before," says Stanford's James Zou.