The single number that should have ended the "just trust the detector" era
If you write English as a second language and a detector told you that an essay you wrote yourself is "98% AI," you are not imagining the unfairness. There is peer-reviewed evidence for it, and it is precise.
Liang, Yuksekgonul, Mao, Wu, and Zou published "GPT detectors are biased against non-native English writers" in Patterns, volume 4, issue 7, article 100779. The digital object identifier is 10.1016/j.patter.2023.100779. The study fed 91 TOEFL essays, written by non-native speakers, into seven widely used GPT detectors. The average false-positive rate across those detectors was 61.3%. The same tools flagged native-speaker essays at a far lower rate; the study's preprint put that baseline at roughly 5.19%.
Sit with the gap. A detector that wrongly accuses one native writer in twenty wrongly accuses three non-native writers in five. That is not noise. That is a systematic failure aimed at one population, and it was statistically clean enough to publish in a respected journal. When you cite this study, you are not citing an opinion. You are citing the closest thing the field has to a controlled experiment on detector fairness.
Why careful second-language writing reads as "machine"
Most people assume a false positive means the writing was somehow flawed. The truth is the opposite, and it matters because it changes what you do next. ESL writers get flagged precisely because they write carefully.
Second-language academic writing tends to lean on a small set of safe, formal moves. Transition phrases like "in addition," "furthermore," and "on the other hand." Hedging like "it can be argued that" and "one might suggest." Parallel three-item lists. A controlled vocabulary, because a writer reaches for the words they trust rather than gambling on a rare synonym. A steady, formal register with no slang, no contractions, no casual breaks in rhythm.
Every one of those habits is what good language instruction teaches. Every one of them is also what a large language model produces by default. The writer and the model converge on the same surface, for completely different reasons. A detector trained to separate "human" from "AI" on a binary contrast cannot tell the difference, so it punishes the writer for sounding deliberate. The skill that earned a high TOEFL band is the same skill that trips the flag.
The two statistics that actually move the score
To understand why this misfire is built in rather than accidental, you have to look at what these detectors measure. They do not score vocabulary, which is why swapping a few flagged words for synonyms never reliably beats them. They score statistical properties of the prose, and two carry most of the weight.
The first is perplexity, which is roughly how surprised a language model is by your next word. Human writing surprises the model often, because people reach for the unexpected phrasing, the idiom that fits the moment, the slightly-off word that a model would never pick. Model writing is low-perplexity by design, because it selects the safe, probable word at every step. Second-language writers, working inside a smaller comfort zone of vocabulary, also produce lower-perplexity prose. To the detector, "careful and predictable" and "machine-generated" look like the same thing.
The second is burstiness, the variance in sentence length. A native writer in full flow fires a three-word fragment, then a forty-word subordinate-clause marathon, then something in between, because thought is uneven. Many ESL writers, taught to write clear and uniform sentences, produce a steadier rhythm. Models produce an even steadier one. Low burstiness plus low perplexity is the exact fingerprint these tools were tuned to call "AI," and second-language academic prose lands there honestly. The result is a 61.3% false-positive rate that no one designed on purpose but everyone should have predicted.
Did the detectors quietly fix this after 2023?
A fair question is whether any of this still matters three years on. Detectors retrain, vendors publish fairness statements, and a 2023 number could in principle be ancient history. The honest answer is that the bias narrowed in places and persisted in most.
A handful of vendors responded directly. The clearest example is a dedicated lower-aggression mode shipped by one major detector in early 2026, built to reduce false positives on non-native and edited text. That helps, but only when a reviewer actually selects it, and most institutional accounts run the default setting that was never calibrated for ESL writing. Independent audits through 2024 and 2025 kept finding elevated false-positive rates on non-native samples across the tools that never shipped a calibration at all. The pattern held: where a vendor explicitly addressed the bias, the gap shrank; where they did not, careful second-language prose kept getting flagged.
So the Stanford number is not a relic. It is a baseline that some tools improved on and many did not. Treating "the detector has surely been fixed by now" as a safe assumption is exactly the mistake that lands an innocent writer in a hearing.
What a wrong flag actually costs
It is tempting to treat this as an academic curiosity. For the people on the receiving end, it is not. The cost of a false positive scales with the stakes of the document, and for international students the stakes are unusually high.
A flagged essay can trigger an academic-integrity referral. For a domestic student that is serious. For an international student it can cascade into a hold on enrollment, a threat to a scholarship tied to good standing, or, in the worst case, a visa status that depends on staying enrolled. The student who is least equipped to argue the case in a second language, in an unfamiliar disciplinary process, under time pressure, is the same student the detector is most likely to accuse. The unfairness compounds at exactly the point where a person can least afford it.
This is why "the detector said so" is a dangerous standard. A tool with a known 61.3% error rate on a specific population should never be the sole evidence in a decision that can alter someone's legal residency or financial aid. The Stanford figure is not just a research result. It is the reason a single automated score should never carry a disciplinary outcome on its own.
The institutional precedent you can actually cite
The good news is that you are not the first person to make this argument, and several institutions have already conceded it in writing. That gives a flagged writer something stronger than a complaint: a precedent.
Vanderbilt University disabled Turnitin's AI detector in 2023, the earliest high-profile institutional restriction, and said plainly that the tool was not reliable enough to use. Yale's Poorvu Center for Teaching and Learning, the University of Waterloo's Office of the Associate Vice-President, Academic, and Curtin University followed between 2023 and early 2026, disabling the feature outright or discouraging the use of AI-detection scores as primary evidence in academic-integrity cases. (A frequently misreported fifth entry deserves precision: at UC San Diego, it was the continuing-education arm, UC San Diego Extended Studies, that switched off Turnitin's AI indicator on April 7, 2025, and campus instructors otherwise make their own detector choices.) The chain spans private and public universities across the United States, Canada, and Australia, which means a student can usually find a comparable institution to point to regardless of where they study.
You do not need to argue from first principles. You can write, in a calm appeal, that named universities have formally restricted these tools for exactly the reliability reason the Stanford study documents. For the longer appeal-pathway walkthrough in an academic-integrity context, see our Turnitin AI bypass guide, where Step 4 covers the ESL false-positive scenario specifically.
How to read a flag honestly, without panic and without denial
A detector score is a probability estimate produced by a heuristic, not a verdict. Reading it honestly means refusing two opposite mistakes: treating a high score as proof of guilt, and treating any score as meaningless.
If you wrote the text yourself and got flagged, do three things. First, keep your evidence of process: drafts, version history, notes, search history, anything that shows the work happened over time. A document's edit timeline is far harder to fake than a detector score and far more persuasive to a human reviewer. Second, run the same passage through more than one detector. Disagreement between tools is itself an argument, because it shows the result is unstable rather than authoritative. Third, attach the Stanford reference, Liang et al. 2023, DOI 10.1016/j.patter.2023.100779, with the 61.3% TOEFL figure named directly. The combination of your process evidence plus a published false-positive rate is what moves a reasonable reviewer.
What you should not do is rewrite the essay into something you would not say, just to lower a number. The goal is to defend authentic work, not to disown it.
Where checking your own writing fits, and where it does not
Knowing the mechanism is the foundation. Having a way to test your own draft before you submit it is the practical layer on top, and it is where a calibration-aware detector earns its place.
Our own checker at /detect was rebuilt to weight non-native passages differently from native ones, so that the careful, formal habits described above do not automatically inflate the score. Across our internal benchmark of verified human ESL writing, the calibrated false-positive band moved from the Stanford-style 61.3% range down to 4-9%, with native English sitting around 1-3%. We mention that figure once and do not oversell it, because the honest framing is the point: it is a meaningful improvement, not a cure, and no detector in 2026 is reliable enough to be the final word on its own.
What a calibration-aware check buys you is information before the moment of risk. Paste a paragraph you wrote yourself and see whether it reads as predictable to a machine, then decide for yourself whether to add the burstiness and specificity that authentic editing naturally produces. The free tier covers four checks a day with no signup, which is enough to test a draft before you hand it in. And if you draft with ChatGPT first, the ESL prompt template in our two-layer workflow guide is calibrated for exactly the flagging pattern this study documents.
How HumanizeMyAI's own work measures against the same tools
Because we publish numbers for ourselves under the same conditions, here is the canonical reference row so you can judge our calibration claims against our actual output rather than a marketing line. Across the same six-detector panel in May 2026, HumanizeMyAI's measured results were 4% GPTZero / 8% Turnitin / 8% Originality AI / 6% Copyleaks / 30/30 QuillBot pass / 3% ZeroGPT.
That row exists for transparency, not as a promise about your specific text. It shows that prose built on real human writing, rather than synthetic loops, sits low on the same detectors that flag careful ESL writing high. The relevance to a non-native writer is the underlying point of the Stanford study: detectors react to statistical texture, not to honesty, and the way to read low is to write the way real people in your field actually write. For the full hub on what a flag proves and what to do about it, see our ESL and AI detection explainer.
The honest bottom line for non-native writers
The Stanford 2023 study did not prove that AI detectors are useless. It proved something narrower and more important: that on careful second-language writing, the current generation of detectors is wrong far more often than its confident percentages suggest. A 61.3% false-positive rate on a population is not a rounding error. It is the headline.
If you have been flagged and you did the work, you are holding stronger evidence than the detector is. You have your drafting history, a published study with a citable DOI, and a list of universities that have already restricted these tools for the exact reason you are about to argue. Use them in that order, stay calm, and make the reviewer engage with the evidence rather than the number. The score is a starting point for a conversation, not the end of one. Anyone who tells a second-language writer otherwise has not read the research.
By Fırat Mıhcı (ResearchGate). $0 affiliate stake.