The dream of fully automated science — machines that dream up hypotheses, run the experiments, and write the papers — is not yet reality, according to a new stress test published this week. The study, posted on the preprint server arXiv in late July and covered by Nature, asked the original authors of two computer-science papers to grade the output of an AI research system. The verdict was harsh: overall scores of 2/6 and 1/6.

The test was designed by Sayash Kapoor and colleagues at Princeton University, who argued that peer review is an unreliable yardstick for judging AI-generated research, especially in fast-moving fields like machine learning. Their method, called "shadow evaluation," picks papers already submitted to a major conference, gives an AI tool the research question from each paper, and lets the human authors scrutinize the result — assuming those authors would bring far more expertise and dedication to the task than a harried reviewer.

The AI system under examination was built by "harnessing" the large language model Claude Opus 4.8 inside a modified version of the agentic framework OpenClaw, wrapped in a broader scaffold with general instructions and progress checks. The harness gave the model an arsenal of tools: it could spawn sub-agents, browse the Internet, use software libraries, rent computer processors for running experiments, and consult software that simulates peer review. For each paper, the system had six days and US$3,000 in computing credits. Its two assignments were to design a method for precise control of chatbot personality and to build a failure detector for a certain class of neural network.

The authors admitted to being surprised by its engineering stamina. It ran hundreds of experiments over several days without getting stuck in an endless loop of unresolvable errors, produced solid literature reviews, made some minor findings, caught its own false claims — hallucinations — and did not try to cut corners, or "reward hack," despite predictions that it would.

But it mostly failed at the actual research. The typical failure mode: it would pick a few hypotheses, settle on one far too early, and never backtrack when its approach stopped working. Its subsequent self-review was not negative enough, so it persisted along the initial path, whittling down its claims until it had little of interest left to say.

"I don't think full automation of open-ended research is on the horizon right now," Kapoor said. The effort to automate the whole scientific pipeline was pioneered by Sakana AI of Tokyo with its "The AI Scientist" system, unveiled in 2024 and refined in a paper published in Nature in March of this year. That system was tasked with studying pitfalls in machine learning; of three papers it produced for a conference workshop, one scored high enough to be accepted — a result Kapoor's team regards as an unreliable measure of true understanding.

The study suggests that today's AI scientists are strong engineers but weak scientists: superb at grinding through experiments, still poor at the harder judgment calls of knowing when an idea is dead and when to change course. The preprint is available on arXiv under identifier 2607.27191.