HomeAI DetectorESL & AI Detection

AI Detection and ESL Writers: Why Non-Native English Gets Falsely Flagged

By Fırat Mıhcı. Built HumanizeMyAI on a published 2,590-essay corpus of student writing that includes non-native English voices. ResearchGate profile. Nobody named here pays us; this is independent research. Last updated July 31, 2026.

TL;DR: ESL Writers and AI Detectors

AI detectors flag non-native English essays as AI far more often than native ones, 61.3% versus 5.19% in Stanford’s 2023 study, usually just for the formal register second-language teaching rewards. If your own work got flagged, verify it on our free detector, and humanize it to read naturally if you need to.

If you write English as a second language and an AI detector has flagged work you wrote yourself, you are not imagining the pattern, and it is not your writing’s fault. The bias is real, it is documented in peer-reviewed research, and it is one of the most consistent failure modes of the whole detection approach. This page is the hub for that one problem: what the research actually shows, why it happens, what a flag does and does not prove, which universities have restricted detector evidence over exactly this issue, and the concrete steps to take if you have been caught in it.

If you came here holding a specific number and wanting to know whether it is a bad one, skip to what percentage is actually acceptable. The short version is that no vendor and no university publishes a threshold, the 20% figure you have probably seen quoted is a display rule rather than a safe zone, and none of that depends on whether English is your first language.

The deeper methodology, including how we recalibrated our own detector to reduce the false-positive rate on non-native prose, lives in a separate Stanford 2023 deep-dive. This hub stays focused on the situation an ESL writer is actually in and what to do about it.

1. What the Research Actually Found (Stanford 2023)

In July 2023, a Stanford team led by Weixin Liang published a study in the Cell Press journal Patterns titled “GPT detectors are biased against non-native English writers.” The full citation is Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023), Patterns 4(7), article 100779, DOI 10.1016/j.patter.2023.100779. It is the canonical reference for this problem, and it is the single most useful thing an ESL writer can cite.

Our own lab has published on the same terrain: a study of 3,300 learner essays from the public W&I+LOCNESS dataset, What Actually Changes as English Proficiency Grows (DOI 10.13140/RG.2.2.35861.90086), maps how advanced learner writing converges on native accuracy while its complexity moves into the noun phrase, the very register features detectors tend to misread as machine output. Method, code and data are open.

The team ran a set of TOEFL essays written by non-native English speakers through seven commercial GPT detectors, alongside a control set of essays written by native-English speakers. The headline result:

  • The detectors flagged 61.3% of the non-native TOEFL essays as AI-generated, even though every one was written by a human.
  • On the native-English control set, the false-positive rate was roughly 5.19%.

That is a large gap. The detectors were not detecting “AI.” They were detecting the statistical signature of second-language English and mislabeling it. Any number you read off a detector has to be understood against that backdrop, because competitor detectors carry a structural bias toward exactly the writers least able to absorb the consequences of a wrongful flag.

One update matters, and citing a 2023 paper in 2026 without it would be misleading. In 2024, researchers at ETS, the organisation behind TOEFL and the GRE, published a study in Computers & Education (Jiang, Y., Hao, J., Fauss, M., & Li, C., volume 217, article 105070, DOI 10.1016/j.compedu.2024.105070) asking the same question of a detector they built themselves, on real large-scale writing-assessment data. They found no disadvantage against non-native writers.

Be careful about what that does and does not mean, because it is easy to over-read in either direction. It is not evidence that the detector your institution runs has fixed anything. It is evidence that the bias is an engineering choice rather than a law of nature: a classifier built with fairness in mind can avoid it. Whether the specific tool judging your essay was built that way is a separate question, and one that no vendor has answered with comparable evidence. So the 2023 finding remains the right thing to cite when you are challenging a flag, and the 2024 finding is the right thing to cite when someone tells you this is simply how detectors have to work.

2. Why ESL Writing Trips Perplexity and Burstiness Classifiers

Most AI detectors lean on two statistical signals, and both of them happen to align with how careful second-language English reads. Understanding the mechanism is what turns an anxiety-inducing percentage into something you can explain and act on.

Perplexity measures how predictable each word is given the words before it. Language models are built to pick the most probable next word, so their output tends to read smoothly and predictably, which shows up as low perplexity. Second-language writers often produce low perplexity too, but for a different reason: you tend to stick to the vocabulary and constructions you are confident in rather than reaching for unusual, low-frequency words. Careful, controlled word choice reads as “too predictable” to a detector that was trained to treat predictability as a machine signal.

Burstiness measures how much sentence length and rhythm vary across a passage. Native casual prose is bursty: a four-word sentence lands after a twenty-eight-word compound clause, then a short callback. Default model output clusters tightly around medium sentence lengths with little variation. Formal ESL academic prose also tends toward even, controlled sentence rhythm, partly because TOEFL and academic-English instruction explicitly drills consistent structure. Flat variance reads as machine-like to the classifier, whether a model or a careful human produced it.

Stack a third factor on top: the specific moves that formal English-as-a-second-language instruction teaches, such as formal transitions (“in addition,” “furthermore,” “on the other hand”), hedged claims (“it can be argued that”), parallel-list construction, and a consistent register with no slang or contractions. These are exactly the conventions a language learner is taught to use, and they are also the conventions a large language model produces by default. A detector trained on a native-versus-AI contrast reads the overlap as AI. The writer is penalized for writing carefully. None of this is a flaw in non-native writing; it is a property of the stylometric approach itself.

3. What a Flag Does and Does Not Prove

A detector score is a probability estimate, not a verdict. Reading it correctly is the difference between a manageable conversation and a panic response.

What a high score does suggest: the text carries statistical patterns the classifier associates with machine writing, which can include genuine AI output and can equally include careful human prose in a formal register. It is a reason to look more closely, nothing more.

What a high score does not prove: that you used AI, that the percentage equals the share of your essay that is machine-written, or that the number is even stable. A 40% score does not mean 40% of your essay came from a model; it means the classifier is roughly 40% confident about the whole passage. Run the same text twice and most detectors will shift a point or two, because the score is probabilistic, not deterministic. And for a non-native writer the score carries less evidentiary weight than the same number on a native writer’s paper, because the false-positive base rate is higher to begin with.

The honest framing, which I hold to even though we build a detector ourselves: a detector score is a signal to weigh, not a verdict, especially on writing in a formal or second-language register. Read the per-pattern breakdown rather than the headline number. Our detector names which specific patterns fired so you can see whether a flag is driven by two stray sentences or by an even, register-wide signal.

4. Is the Acceptable AI Percentage Different for ESL Writers?

No acceptable percentage has ever been published for anyone, and writing English as a second language does not move a line that was never drawn. What changes for you is the base rate sitting underneath the figure: the same seven detectors that flagged 61.3% of the TOEFL essays described above read a second-language register as machine-like far more often, so an identical number carries a materially higher chance of being wrong on your paper than on a native classmate’s. Treat your score as more likely to be an artifact of the instrument, never as a lower bar you personally have to clear.

The general answer (what the figure measures, what Turnitin prints and what it withholds, what colleges actually do with it, and when a number turns into a review) is set out in full on what percentage of AI is acceptable.

5. The Institutional Precedent: Universities That Restricted Detector Evidence

The Stanford finding was not just an academic result. It fed an institutional-policy chain that an ESL writer facing review can cite directly. A growing list of universities have disabled AI detection or barred its use as standalone disciplinary evidence, and the documented cases span several countries, which makes the precedent citable across institutional contexts.

  • Vanderbilt University disabled Turnitin’s AI detector in 2023, citing a false-positive rate on legitimate student writing too high to use as disciplinary evidence. This is the precedent-setting case the rest are often measured against.
  • Yale University (Poorvu Center for Teaching and Learning) advised against relying on Turnitin’s AI indicator, citing accuracy and fairness concerns.
  • University of Waterloo (Office of the Associate Vice-President, Academic) discontinued the AI-writing indicator for all users effective September 2025, citing reliability concerns and bias against non-native English writers.
  • Curtin University (Western Australia) disabled Turnitin AI Detection in its academic-integrity workflow in January 2026, citing the gap between a probability score and the evidentiary standard for a case.
  • University of California San Diego moved to withhold the AI-detection indicator following Academic Senate and faculty deliberation across May 2024 to February 2026.

The chain spans US private, US public, Canadian public, and Australian public institutions. The common thread is not “the technology is broken” but that the evidentiary weight of a single probability score is lower than the cost of being wrong about a student’s work, a calculus that lands hardest on non-native writers. The faculty-facing version of this story, with the detection mechanics behind it, is covered in the Turnitin AI checker guide.

6. What to Do If You Are an ESL Writer Who Got Flagged

A flag on work you wrote yourself is the start of a conversation, not the end of one. Here is the concrete sequence, in the order that actually helps.

  1. Gather your draft history first. The single strongest piece of evidence is proof that the writing developed over time. A Google Docs revision timeline, Word version history, tracked changes, email drafts to yourself, or notes and outlines all show the work taking shape. Collect this before you respond to anyone, because it is much harder to reconstruct after the fact.
  2. Cite Liang et al. 2023. Bring the Stanford study to the conversation: DOI 10.1016/j.patter.2023.100779. The 61.3% versus 5.19% finding is peer-reviewed and directly on point: detectors carry a documented bias against non-native writers, so a score on your paper should not be read at face value. The university-restriction precedent in the section above is supporting context for why a score alone is insufficient.
  3. Request a human review. Ask for your work to be evaluated by a person rather than by a number. Most institutions require corroborating evidence beyond a detector score before any sanction, and non-native writers have strong grounds to ask for review of a flag that has no corroboration. The same logic applies outside the classroom: non-native professionals accused of using AI at work can point a client to their draft history and the Stanford bias finding rather than accept a score at face value. Be specific: ask what evidence beyond the score exists, and offer your draft history and a conversation about the substance of your argument.
  4. Pre-check future drafts with a free detector. Going forward, you can see how your own honest writing reads before anyone else runs it. Paste a passage into our free detector and open the per-sentence breakdown. It shows which specific patterns are firing, usually transition clustering, hedge density, or low sentence-length variance, so you can adjust those clauses without changing your meaning. This is a self-check, not a guarantee, and it does not store your text.

If the institution you are dealing with uses Turnitin specifically, the Turnitin AI checker guide walks through how the score is generated and what the bands mean. For the full appeal-pathway treatment and the underlying research, the Stanford 2023 deep-dive is the deeper methodology read.

7. How HumanizeMyAI Helps Without Disguising Authorship

There is an honest version of what a humanizer does for a second-language writer, and a dishonest version, and the difference matters. Let me be precise about which one this is.

When you genuinely wrote a passage yourself and a detector flags it because of the register, the useful move is to round off the specific ESL-shaped patterns that classifiers misread, the transition clustering, the flat sentence rhythm, the narrow academic vocabulary, while keeping your meaning, your argument, and your authorship intact. That is not laundering machine output past a checker; it is smoothing the statistical edges that wrongly read as machine writing in the first place. Our humanizer is trained on a 2,590-essay corpus of real student writing that deliberately includes non-native English voices, so the rewrites it produces stay in a natural human register rather than swapping in a thesaurus.

What this is not for: taking AI-generated work you did not write and disguising it to pass off as your own under a policy that forbids it. No tool changes the underlying academic-integrity contract, and if your assignment forbids AI assistance, the right move is to write the work yourself. Most universities and clients in 2026 accept disclosed AI-assisted drafting; the cleanest path through any of this is to say what you used.

The natural loop is to pair the two tools. Run a flagged passage through the humanizer, then re-check it on the detector and read the per-sentence view. If a couple of sentences still flag, the breakdown tells you which ones to adjust by hand. The detector tells you what reads as machine-shaped, the humanizer smooths it, and the detector confirms. Both tools are free to start, with no account required for detection.

Affiliate Transparency

HumanizeMyAI runs no paid placements or affiliate commissions on this page. The Stanford 2023 figures are reproducible from the published paper, the university-restriction cases are limited to ones with public documentation, and the only commercial product on this site is HumanizeMyAI itself (Basic $18/mo, Pro $27/mo, Ultra $48/mo). The detector is free with no signup needed, as a deliberate product decision: a first-pass read is too important to put behind a paywall.

Fırat Mıhcı built HumanizeMyAI on a published 2,590-essay corpus of real student writing, with an academic background in second-language English research that informs the ESL calibration behind the detector. His ResearchGate profile includes the underlying research. For the deeper methodology of how the detector was recalibrated against the Stanford baseline, see the Stanford 2023 deep-dive.