HomeBypass AI DetectorsBypass Pangram

Pangram AI Detector (2026): How It Works & What It Checks

By Fırat Mıhcı, who built HumanizeMyAI on a published 2,590-essay corpus. ResearchGate profile. Last updated August 31, 2026. Refreshed monthly.

TL;DR

Pangram is the hardest AI detector to clear in 2026: a University of Chicago evaluation reports it catching humanized text 97% of the time where GPTZero caught 46%, with roughly one false positive in ten thousand. Ours is trained on 2,590 real student essays, and on August 31, 2026 it read 0.3% AI as a mean across the detectors that hand back a score. Try it free and judge for yourself.

If your writing was flagged by Pangram AI Detector, this guide explains what the classifier actually measures and what your honest options are. Pangram works differently from GPTZero, Turnitin, or Copyleaks, and the older advice written for those tools mostly backfires here. Read the first section before anything else, because it decides whether this guide fits your situation at all. If you want to check a draft yourself first, our free detector reads it on the spot and asks for nothing in return.

Who This Guide Is For (And What It Is Not)

Pangram AI Detector sits in front of a specific set of readers, so this guide is written for three of them and explicitly not for a fourth.

The first reader is the college student whose course or learning-management system has started running Pangram, who wrote the work themselves, and who received an AI flag they believe is wrong. Pangram is a newer arrival than Turnitin, so a flag from it often lands on students who have never heard the name and have no idea what it checked. Understanding the classifier is the first step to responding to a flag calmly instead of in a panic.

The second reader is the non-native English (ESL) writer flagged on authentic, self-authored work. This cohort has the most at stake with any detector, because the writing patterns that older tools mislabel as “AI” overlap heavily with second-language phrasing. Pangram’s published record on this is unusually strong, and the ESL section below walks through the exact numbers so you know where you stand.

The third reader is the graduate student producing thesis-length or publication-track work who needs to understand Pangram before submitting a long document. Pangram returns a four-tier verdict rather than a single yes-or-no, and one of those tiers, “Lightly AI-Assisted,” is genuinely ambiguous. Knowing what that label means, and what it does not mean, matters before a committee or an editor reads it.

This guide is not a manual for passing off wholly AI-generated coursework as your own original work where your institution prohibits AI. That is an academic-integrity violation no matter what any detector reads, and no tool changes that policy. Everything technical that follows presumes the argument is yours, and that any AI you leaned on was either allowed outright or confined to drafting under a disclosure your course accepts. Legitimate reasons to read on include verifying your own writing before submission, understanding a flag you think is mistaken, and researching detector accuracy for a policy decision. Where a course bans AI outright, there is one honest response, and it is to do the writing yourself.

What Is Pangram AI Detector

Pangram AI Detector is an AI-text classifier built by Pangram Labs, a US company founded in 2023 by Max Spero and Bradley Emi, both former machine-learning engineers. It reads a block of text and estimates how much of it was written or edited by a large language model, then returns a verdict. Pangram is sold to schools, publishers, and recruiters, and it is reachable through a web app, a browser extension, and an API. Individual plans run roughly $20 to $35 a month, while schools negotiate per-institution contracts rather than per-student licenses, which is why your university’s deployment may look different from the public site. A free check is available on Pangram for short samples.

What makes Pangram worth a separate guide is not its interface but its accuracy profile. Most detectors trade false positives against catch rate. Pangram’s published numbers claim to win on both at once, and an independent University of Chicago evaluation broadly backs that claim. The next section explains how it gets there.

How Pangram Detects AI Writing: Mirror Prompts, Hard Negative Mining, EditLens

Pangram AI Detector reaches its accuracy through three design choices that set it apart from perplexity-based tools like GPTZero. Understanding them tells you why the surface edits that slipped past older detectors do so little here.

Mirror Prompts. Pangram’s training data is built in matched pairs. For every real human document, the team generates an AI counterpart on the same topic, at the same length, sometimes seeded with the same opening sentence. The classifier then learns the difference between a genuine human text and its closest possible AI twin, rather than the difference between human writing and some generic AI sample. This is why surface edits help so little: the model was trained against AI text that already looks like yours.

Hard Negative Mining. Pangram does not stop at the first round of training. It runs the classifier, finds the examples it got wrong (the AI text it scored as human, and the human text it scored as AI), and retrains specifically on those hardest cases. Repeat that loop enough times and the model concentrates its attention exactly where other detectors fail: on the borderline, lightly edited, paraphrased text that older tools wave through. Each cycle makes the remaining blind spots smaller.

EditLens. Pangram’s newer engine, EditLens, was published at ICLR 2026 (arXiv:2510.03154). Instead of a single AI-or-human label, EditLens estimates how much of a document was AI-edited on a continuous scale, then maps that score to four tiers: Fully Human, Lightly AI-Assisted, Moderately AI-Assisted, and Fully AI-Generated. This is the verdict you actually see. The “Lightly AI-Assisted” tier is the one to understand, because it is where a human draft that got a grammar pass, or a non-native writer who leaned on a language tool, can land. That tier is a description of edit extent, not a misconduct ruling. What it means for you depends entirely on your institution’s policy, not on Pangram. A Lightly AI-Assisted verdict does not automatically start an academic-integrity case: many institutions treat the four-tier output as information for the instructor rather than an automatic finding. If you receive that verdict and your syllabus says nothing about AI editing, the practical step is to ask your instructor what the tier means in your course context before assuming it counts as a violation.

Two definitions help here. Humanized text means AI output that has been run through a rewriting tool to read more naturally, which is the thing most “bypass” guides are really about. A false positive means human writing that a detector wrongly labels as AI. Pangram is built to be hard on the first and rare on the second, and the next two sections show how well it delivers on each. For a side-by-side of how a perplexity-based detector reasons differently, our GPTZero guide covers that classifier’s lexical-predictability approach.

How Accurate Is Pangram

Pangram AI Detector’s published false-positive rate is roughly one in ten thousand, and the University of Chicago evaluation described it as the only detector in its four-tool test that stayed robust against humanizing tools.

Pangram AI Detector was, by the strongest independent measurement available, the most accurate of the four detectors tested, within the scope of that one evaluation. A University of Chicago Becker Friedman Institute working paper (WP_2025-116, Jabarian & Imas, September 2025) tested four detectors against a policy cap that allowed no more than one false positive in two hundred (BFI evaluation). Pangram was the only one of the four that met that cap, and the paper described it as the only detector in the group robust to humanizing tools. That is a narrow, specific finding (“most accurate of the four tested under a strict false-positive cap”), not a universal crown, and it is worth quoting precisely rather than rounding up.

The number that matters most to an innocent writer is the false-positive rate, and Pangram’s is roughly one in ten thousand. Put that in classroom terms. Imagine a university that processes ten thousand papers in a semester. At Pangram’s published rate, that is about one paper wrongly flagged across the entire term. Run the same ten thousand papers through a detector with a one-percent false-positive rate, like several older tools, and you would expect around a hundred innocent students flagged instead. The gap between one and a hundred is the whole reason Pangram’s accuracy claim matters to people who did nothing wrong.

One caution on version numbers. Pangram moves fast, and the figures above attach to specific releases: version 3.0 shipped in December 2025, version 3.2 in February 2026 added support for samples as short as fifty words, and version 3.3 in May 2026 claimed to roughly double its catch rate on humanized text. When you read any Pangram statistic, including the ones in this guide, check which version it came from, because a number from one release can be stale a quarter later.

Does Pangram Falsely Flag ESL Writers

Pangram AI Detector posts the best published false-positive record on non-native English writing of any detector I have reviewed, and for ESL students that single fact reframes the whole question. Across 25,021 samples drawn from four standard second-language corpora (ELLIPSE, ICNALE, PELIC, and TOEFL writing), Pangram’s false-positive rate was 0.012%. On the PELIC subset it was 0.019%, and on a Stanford-derived TOEFL benchmark it reported zero false positives.

Set that against Turnitin, which has been measured at roughly 1.4% on second-language writing. That is not a small edge. 1.4% against 0.012% is roughly a hundred times lower, two orders of magnitude. The reason this gap exists at all traces back to a peer-reviewed warning that older detectors never fully fixed. The Stanford 2023 study (arXiv:2304.02819) found that GPT detectors of that era flagged 61.3% of real TOEFL essays written by non-native speakers as AI, because the simpler writing patterns that ESL students use looked, to a perplexity-based model, like machine text. Pangram’s matched-pair training is specifically designed to stop confusing “simple” with “synthetic.”

If you are a non-native writer who was flagged on your own work, the practical takeaway is twofold. First, on Pangram specifically, a flag is far less likely to be a language-pattern error than it would be on an older tool. Second, if you were flagged on a different detector and are trying to understand whether the result is trustworthy, our guide for ESL writers facing AI detection covers the evidence a student can actually assemble: drafts in progress, version history, and a clear reply to a flag you know is wrong. Knowing which detector flagged you changes how much weight the flag deserves.

Does Pangram Catch Humanized Text

Pangram AI Detector is the detector that humanizing tools have the hardest time with, and the honest answer to the question in this heading is: against every humanizer publicly tested so far, yes. In an August 2025 benchmark, Pangram ran nineteen humanizing tools through its classifier and detected all of them at over ninety percent. QuillBot’s humanizer was caught 99% of the time, StealthGPT 95.6%, and Undetectable AI 90.3%. Every one of those nineteen tools is what I call paraphraser-class: they work by swapping synonyms and reshuffling sentence structure, which is exactly the pattern Pangram’s Mirror Prompts and Hard Negative Mining are tuned to catch.

Two things must be said plainly about that benchmark. First, it is from August 2025 and predates Pangram’s version 3.3, which itself claimed to roughly double humanized-text catch rate, so if anything the live numbers for those nineteen tools are higher now, not lower. Second, and more important for honesty: HumanizeMyAI was not in that benchmark. We are a different architecture class. Instead of swapping synonyms, our system is corpus-trained, trained on 2,590 real student essays, and 58% of that corpus came from non-native English writers, which means the output carries the rhythm and word-choice variance of real human academic writing rather than the predictable substitution pattern that defines paraphraser-class tools.

What this page carries is the record we read straight off each detector’s own interface on August 31, 2026:

DetectorHumanizeMyAI result (Aug 31, 2026)
GPTZero0% AI
TurnitinHuman (no score under 20%)
Originality AIHuman (15% or less, the lowest the free tier lets you measure)
Copyleaks0% AI
QuillBot AI Detector0% AI
ZeroGPT0-3% AI
Mean, detectors that return a score0.3% AI
Pangram AI Detectornot measured on our side

Read that table for what it is. A 0.3% mean across the detectors that return a score is a measured result, and it is different in kind from the 90-99% range every paraphraser-class tool posted on Pangram. The two classifiers that withhold a percentage, Turnitin and Originality AI, both returned a Human verdict on the same output. The Pangram row stays empty until that run lands.

Pangram vs Turnitin: Which Should You Worry About

Pangram AI Detector and Turnitin solve the same problem from opposite ends of the workflow, so which one you should worry about depends on how your institution checks work. Turnitin is LMS-integrated: it lives inside Canvas, Blackboard, Moodle, or D2L, runs automatically when you submit, and most students never see its interface directly. Pangram is more often standalone: an instructor pastes your text into the web app or browser extension, or a department runs it as a separate check. If your submissions go through your university’s LMS, Turnitin is the detector you will face by default; if an individual professor is doing the checking, it is increasingly Pangram.

On raw catch rate against humanized text, the cited gap is large. The same University of Chicago context reports Pangram detecting humanized text 97% of the time against GPTZero’s 46% (a figure attributed to Russell et al.), and Turnitin’s AI-paraphrase detection has historically sat well below Pangram’s. The practical implication is that a tool which clears Turnitin’s AI flag cannot be assumed to clear Pangram, because they are tuned differently and trained on different data. For the Turnitin side specifically, our Turnitin AI checker accuracy guide covers how its August 2025 classifier reads burstiness and lexical fingerprints, and our bypass-Turnitin guide walks through that detector’s three-signal architecture in full. Read both detectors as separate problems, because they are. And if the standalone check you face is Winston AI rather than Pangram, our Winston AI detector review breaks down that tool’s accuracy claims the same way.

Which Colleges Use Pangram

Pangram AI Detector’s campus footprint is real but smaller and more nuanced than the round numbers floating around online suggest, so here is only what is actually documented. Wellesley College has been reported as piloting Pangram, per the Wellesley News in December 2025 (Wellesley News coverage). Imperial College London and Stanford University appear in research partnerships with Pangram Labs rather than as adopters of the detector for grading. Researchers at the University of Chicago and the University of Maryland have used Pangram in studies, which is not the same as their institutions deploying it. Delaware County Community College and Stony Brook University appear on Pangram’s own use-case materials. Pangram also has an education partnership with Wiki Education.

You may see claims that “twelve or more universities switched from Turnitin to Pangram.” I could not verify any such named switch from a primary source, so I will not repeat the figure. The same goes for any specific “University X dropped Turnitin for Pangram” story; none is documented by name anywhere I could check. The reliable way to know whether your own work will face Pangram is not a list online; it is your syllabus, your course’s academic-integrity statement, and your LMS settings. If those do not say, ask your instructor directly which AI detector the course uses. A two-line email removes all the guesswork.

5 Steps to Prepare a Draft for a Pangram Check

Pangram AI Detector rewards writing that carries genuine specificity and a consistent personal voice, so these five steps focus on getting your own work into that state rather than on gaming the classifier. They assume you wrote the work yourself and want to verify it before submission.

  1. 1Run a baseline check on your own draft first

    Before changing anything, paste your draft into a detector so you know where you actually stand. If it reads as human, you may be done. Our free AI detector gives you a same-day baseline read so you are not guessing at the problem. Knowing your starting point stops you from rewriting passages that were fine.

  2. 2Use the flagged-sentence highlight view

    Most detectors, Pangram included, can show you which sentences drove the score rather than a single document-level number. Work at the sentence level. A draft is rarely uniformly flagged; usually a handful of generic, low-specificity sentences carry most of the weight, and those are the ones to focus on.

  3. 3Rewrite the flagged sentences with real specificity

    The single most effective edit is adding detail that only you could write: a specific example from your reading, a number from your own data, a concrete observation, a personal reason for an argument. Generic claims read as machine-like to every detector because they are the easiest sentences for an AI to generate. Specificity is the opposite of what Mirror Prompt training learned to flag.

  4. 4Humanize, then re-check, and read the result honestly

    If a passage still reads as machine-generated after manual editing, run it through a humanizer and check it again. Our free tool handles up to 250 words per run on a free account, which is enough to test your single highest-risk paragraph first before committing to a longer document. Read the rewrite next to your original and keep whichever sentences sound most like you. If you are working on a thesis-length document and need more words per run, our pricing page lays out the paid tiers, since four free runs cover 1,000 words in all.

  5. 5Read the whole thing aloud for voice consistency

    EditLens scores edit extent, which means a document that lurches between your natural voice and a polished machine register can read as more edited than one that is consistently yours. Read the full piece aloud. Where the voice suddenly shifts, rewrite that passage in your own words until it matches the rest. Consistency of voice is what a Fully Human verdict actually looks like.

Is Pangram Open Source

Pangram AI Detector has an open-source side that almost no other detector offers, and it is worth knowing about. The team has released EditLens model weights and code under an Open Pangram banner, available on GitHub and Hugging Face under a CC BY-NC-SA license. The models are small enough to run locally; the published versions will load on a MacBook without specialized hardware. For a researcher, an academic-integrity office, or a curious graduate student, that means you can inspect how the classifier scores text yourself instead of trusting a black-box API. It is the closest thing to a reproducible, auditable AI detector currently available. The non-commercial license means you cannot build a paid product on it, but for understanding what a flag was based on, it is a genuinely useful resource that no competing detector matches.

Is Pangram Worth Worrying About? The Verdict

Pangram AI Detector is the hardest AI classifier to clear in 2026, and that conclusion is honest rather than promotional. Its matched-pair training, its hard-negative retraining loop, and its EditLens edit-extent engine make it strong exactly where older detectors are weak, on lightly edited and paraphrased text. Its roughly one-in-ten-thousand false-positive rate, independently backed by the University of Chicago evaluation, makes it unusually safe for innocent writers, including non-native English speakers, where it flagged just 0.012% of 25,000-plus second-language samples. If your school uses Pangram, those are the numbers to keep in mind.

For our own tool, the record stands where the readings put it. HumanizeMyAI learns from a corpus of 2,590 real student essays instead of swapping synonyms around, and on August 31, 2026 its mean was 0.3% AI across the detectors that return a score, with Turnitin and Originality AI both reading Human. Our Pangram-specific measurement is still in progress, and when that test is done, this page will carry the number, whatever it is. You can check any draft yourself today with our free detector, see how our approach compares against other tools in our best AI humanizer roundup, or try the humanizer on your highest-risk paragraph. Whatever tool you use, the first and last rule is the same one this guide opened with: no detector and no humanizer changes whether your institution permits AI, the syllabus decides that, and your instructor can settle any doubt in a sentence.

Editorial note: No money passes between HumanizeMyAI and Pangram Labs, or any detector or competing tool named on this page. Detector figures are sourced to the University of Chicago BFI working paper, Pangram’s published benchmarks, and the cited arXiv papers; the HumanizeMyAI figures are ours, and a reader can re-derive each one on the same free detectors. Last reviewed August 31, 2026, and refreshed monthly. By Fırat Mıhcı, ResearchGate.

Run a Paragraph Before a Pangram Check

Put a flagged paragraph in and read the result. Registering is free and brings four runs of 250 words each.

Type oryour AI-generated text or
0/250