When ICML 2026 accepted 6,352 papers — roughly double the previous year — the question was no longer whether AI research is reproducible, but whether anyone could even check. Hugging Face decided to find out by turning the community loose on the problem.

Between July 15 and August 2, 2026, 1,221 volunteers brought their own coding agents — Claude Code, Codex, Cursor, and others — and attempted to reproduce 2,226 ICML papers, claim by claim. The result, published on the HuggingFace blog, is the largest open, claim-by-claim audit of a scientific conference ever conducted.

The numbers tell a complex story. Out of 35,908 individual claims judged:

51% of examined papers had at least one claim independently verified. Of those, 266 papers were fully reproduced with every extracted claim confirmed, and 632 more had partial evidence with nothing falsified. In total, 3,978 individual claims were verified with real experiments.

23% of examined papers had at least one claim falsified or contested. That includes 49 papers where all claims were falsified and nothing could be verified, plus 242 papers where independent reproduction teams reached opposite verdicts on the same claims.

The rest fell into a middle ground: 502 papers with toy-scale evidence only, and 280 where nothing could be established either way, mostly because of missing artifacts — unreleased checkpoints, proprietary datasets, or incomplete code.

Among the most striking findings:

A paging algorithm paper claimed its robustness was bounded by a constant term. A participant found the additive term actually grew logarithmically. The organizers extended the test to k = 1,024 and confirmed the growth at roughly nine sigma. The true robustness is worse than claimed.

In a paper on attention mechanisms and Frank-Wolfe optimization, three independent teams found counterexamples to a central theorem, with violations first appearing at step 224 — explaining why previous reviewers, who stopped checking earlier, had missed the flaw.

A transformer evaluation paper had roughly 66% of its evaluated label positions filled with EOS padding tokens that trained to near-zero loss, artificially deflating perplexity. The papers headline claim of a 3.1% quality cost for 50% cache reduction became roughly 9.4% once corrected.

In one case, the challenge also caught a false falsification: a participant claimed a method was 2x slower than baseline, but had inadvertently compared per-trajectory time against per-batch time. After normalization, the original paper held up.

The strongest results came not from fully autonomous agent runs but from human-directed workflows. Agents got stuck in local loops, misread scale-dependent behavior, and occasionally built entire falsifications on arithmetic errors. The most reliable reproductions involved a human steering the agent — re-pointing it, questioning assumptions, and deciding when an experiments premise was wrong before burning compute on it.

Hugging Face has begun contacting authors of every confirmed finding. Early responses have been positive: several authors have confirmed the results, two arXiv corrections are in flight, and in one case an author had quietly fixed the error a month before the challenge found it.

The project raises a fundamental question about the future of peer review. With ICML 2026 accepting twice as many papers as the year before, traditional volunteer-based reviewing cannot keep pace. The challenge demonstrates that coding agents can help, but only when paired with human judgment — particularly for qualitative evaluation, experimental design, and catching the kinds of subtle errors that require understanding context rather than just running code.