A major career feature published today in Nature has cast serious doubt on the reliability of AI-detection tools that universities worldwide are increasingly relying on to catch student cheating.

The investigation, by journalist Anna McKie, reveals that tools such as GPTZero, Copyleaks, and Turnitin's AI detector — now used by over 60% of higher education institutions — frequently produce false positives, flagging human-written essays as AI-generated.

One chemistry undergraduate, Lauren Jager, told Nature she ran her PhD application personal statement through several detectors and found they rated it "almost 100% AI" — even though she had written it entirely herself. After rewriting her statement to sound "less perfect" to evade the detectors, the score dropped to 30% AI. She was accepted to her PhD program.

Most AI detectors rely on a metric called "perplexity," which measures how predictable each word in a sequence is. Because AI-generated text tends to follow statistically more predictable patterns, lower perplexity signals machine authorship. But human writing — especially by meticulous, rule-following writers — can also score low.

The consequences are severe. Some PhD application portals now warn that AI-flagged statements will be "disregarded entirely." A 2023 Stanford study found that AI detectors mislabeled over 61% of essays written by Chinese English-learners as AI-generated, compared to accurate classification for US students' work — raising concerns about systemic bias against non-native English speakers.

Mike Perkins, a researcher at British University Vietnam studying AI's impact on academia, told Nature: "The short answer is no, they don't work reliably. The long answer is yes, they can work — but there are so many concerns about false positives that they shouldn't be used for anything sensitive for a student."

Nature itself tested ZeroGPT with the US Declaration of Independence and was told the 1776 text was 95–100% AI-generated.

The article notes that students can evade detection by running AI-generated text through "humanizer" tools or having one AI rewrite another's output — creating an arms race between detection and evasion that "doesn't really help anyone."