For years, AI-detection software had a reputation problem. Early tools relied on simple statistical measures — perplexity and burstiness — and frequently flagged human-written text as machine-generated, frustrating academics and publishers alike. Now, a new breed of detector is rewriting the rules.
Pangram Labs, a New York City startup co-founded by Stanford graduates Max Spero and Bradley Emi, has built a tool that claims 99.98% accuracy on English text. Rather than counting linguistic tells, Pangram trains a machine-learning model on millions of human-written texts and their AI-generated mirrors, learning to distinguish the two at a level that resists human-readable description. An independent test by Epoch AI found zero false positives across 495 human-written texts. The tool is already reshaping academic publishing. NeurIPS rejected 18% of its submissions after screening with Pangram. The preprint server arXiv now surfaces Pangram scores through a mirror site called alphaXiv. The University of Chicago has begun using it to vet student coursework. And in one memorable case, a peer reviewer — confronted by AACR journal operations director Daniel Evanko — admitted to using an LLM after running out of time, responding: Wow, you guys are good!
Competitor GPTZero, also based in New York, mirrors the approach and reports 99% accuracy. Five computer-science conferences and three universities have signed up so far.
But the tools biggest challenge lies ahead. As AI-assisted writing becomes the norm — where a human drafts and an AI polishes, or vice versa — detection becomes a spectrum, not a binary. Pangrams latest model, Pangram 4, released in July 2026, divides text into finer-grained chunks as small as 30 to 40 words and attempts to classify segments as human, AI, or mixed. It reports a 0.34% false-negative rate overall.
Still, the grey zone remains thorny. When Nature ran a Scholarly Kitchen blog post through Pangram, the tool initially flagged it as 100% AI. The author insisted the ideas and first draft were entirely human-created and that AI was used only to polish — a practice most consider acceptable. After Pangram 4s release, the score dropped to 96% AI, but the tension persists.
There is a lot of room for improvement in judging real-world cases where AI and human writing blend homogeneously, Spero acknowledges. Small edits can shift scores dramatically — a problem known as jitter. And when consumer AI is used to substantially modify human-written essays, Pangram still labels 41% of results as fully human.
The stakes are rising. Spero has personally called out journalists and even flagged the Popes social-media posts as AI-written. Substack integrated Pangram across its platform in July, letting readers see whether posts are deemed AI-generated. If it is taboo to call out that somebody is using AI to write, then I think we are going to see a lot more people shirking their jobs, Spero says.
Independent researchers urge caution. Tim Requarth of NYU Langone says the tools are good for screening out places that are pumping out slop, but their percentage scores for AI-assisted text should not be taken too literally. The priority, both firms agree, should be detecting fully AI-generated content — the clearest case of misconduct — rather than policing the continuum of human-AI collaboration.
The detection arms race is far from over. Both Pangram and GPTZero release dozens of model updates per year as new LLMs emerge. Third-party evaluations will always lag behind. For now, the tools work best as starting points for investigation, not as final verdicts.




