PRODUCT

May 2026 Update: Better Humanizer, OOD Eval, 100% QuillBot Pass

Fırat Mıhcı·May 13, 2026·10 min read
30/30
QuillBot internal eval, May 2026
TL;DR
Our newest model update sharply raised how reliably rewrites clear AI detectors in our internal testing. If your AI draft keeps getting flagged, this is the version that turns it around. Paste your text and get a clean, human-reading rewrite, free.

What shipped in the May 2026 release?

The May 2026 release changed four things: a structural rewrite in place of synonym swapping, a cleanup pass for leftover academic phrasing, a refreshed reference corpus, and out-of-distribution testing added to the release gate. This is a changelog, so the notes come first and the reasoning follows below.

  • The rewrite now restructures sentences and paragraphs instead of swapping words for synonyms. That is the single biggest change, and the rest of this post explains why it matters.
  • A cleanup pass runs after the rewrite to catch residual academic phrasing the rewrite leaves behind, the formal adverbs and essay-closer constructions that read as machine-formal even after a good restructuring.
  • The reference corpus was refreshed so that every example our engine draws on is itself prose that passes major detectors. No synthetic or unvetted text in the pool.
  • Out-of-distribution evaluation was added to the release gate. We no longer ship a version that only tested well on the kinds of essays we tuned on.

The headline result is the QuillBot internal eval moving from 16/30 to 30/30. That is not a marketing number we picked. It is the count of passages, out of thirty, that came back clean on QuillBot's own AI Detector after humanization. The rest of this post is about how to read that number and what changed under it.

Why rewrite structure instead of swapping synonyms?

Structural rewriting moves the fingerprint that synonym swapping leaves in place. Most humanizers paraphrase: they find the words a detector is likely to flag and replace them, while the sentence shape, the rhythm and the paragraph pattern all stay identical. Detectors caught up to that a long time ago, because the fingerprint of machine writing lives in structure, not vocabulary.

This release rewrites at the structural level. Paragraphs get reorganized. Long uniform sentences get broken up and varied. The register comes down from formal-essay to something a person actually writes. The result is text that reads differently, not text that reads the same with different words. That distinction is the whole reason the eval number moved.

We will not publish the internal patterns the rewrite targets, because that is the part competitors would copy. The high-level claim is the one that matters: this version restructures, it does not paraphrase, and the 16/30 to 30/30 jump is what restructuring buys you over synonym-swapping.

What does the cleanup pass fix?

The cleanup pass runs after the rewrite and resolves the academic phrasing that restructuring leaves behind. Formal adverbs like considerably and substantially, essay closers that begin with a summarizing pivot, concept-definition openers of the X-describes-the-tendency-of-Y shape: a reader skims past all of them, a detector does not.

Those phrasings survive almost any restructuring because they are how academic English is taught, not how any single sentence is built. The pass is the layer that specifically helps writers whose first-draft register leans formal, which includes a lot of non-native English writers whose academic training pushes them toward exactly the vocabulary detectors penalize. If you have ever been flagged despite writing carefully, the formal-register residue is often why, and this pass is aimed straight at it.

Calling it out in a changelog is the point of a changelog: it records what improved. The rewrite does the heavy lifting. The cleanup pass closes the gap the rewrite leaves, and together they are why the matrix tightened this month.

Why was the reference corpus refreshed?

The reference corpus, 2,590 real student essays, was refreshed so that every example the engine draws on is prose that genuinely passes major detectors, not merely human-written text that happens to be on file. Human writing that a detector would flag is not a useful model for writing that a detector will pass.

Competitors who train on synthetic AI output, or who never vet their reference pool against detectors at all, do not have this baseline. They are teaching their tool to imitate the wrong target. Ours learns from text that has already cleared the bar it is trying to clear. That is the quiet reason a corpus-trained approach behaves differently from a paraphraser, and it is the reason the QuillBot result below is architectural rather than incidental. The full account of how the corpus is built and why 2,590 real essays beat a synthetic pool is in the 2,590-essay corpus deep dive.

What is out-of-distribution evaluation?

Out-of-distribution evaluation means testing the rewrite on subject domains the engine was never tuned on. Most humanizers are benchmarked on persuasive essays, the easy center of the distribution. A tool that passes 95% on persuasive essays and 40% on art history is not a 95% tool.

Here is the scenario that makes this concrete. You wrote a psychology essay. The humanizer you used was tuned on persuasive essays, the kind everyone benchmarks against. It passed every test the vendor ran, because the vendor ran persuasive essays. Then your psychology essay got flagged anyway. The gap between the domain you wrote in and the domain the tool was tuned on is exactly what out-of-distribution testing closes.

Most tools quietly degrade the moment your subject leaves the common persuasive-essay norm, because that norm is all they were optimized for, and the result is a tool that fails most students writing about anything unusual.

For this release we evaluated across five domains the engine was not tuned on, including psychology, art history, environmental economics, and education theory. If your subject sits outside the persuasive-essay center of gravity, OOD evaluation is the only test that proves the tool covers you, and it is the test almost nobody publishes. The 30/30 QuillBot result was measured on this OOD set, not on home-turf passages.

How should you read the 6-detector matrix?

The matrix is the canonical May 2026 row, measured on the out-of-distribution set: 4% GPTZero / 8% Turnitin / 8% Originality AI / 6% Copyleaks / 30/30 pass on QuillBot's own detector / 3% ZeroGPT. Those percentages are AI-likelihood scores, so lower is better, and the QuillBot figure is a pass count out of thirty.

The reason we publish all six instead of one is that a single-detector pass is not a proxy for a cross-detector pass. Passing GPTZero tells you nothing about Turnitin, because they are built differently and weight different signals. Turnitin in particular runs a layered classifier with multiple signals, which is why it is usually the hardest cell in any cross-detector matrix and why 8% there is a more meaningful number than any single-detector headline. If your school or your client uses a specific detector, read the cell for that detector and ignore the rest. The matrix exists so you can do exactly that, rather than trust one number and get surprised by another. The numbers above are public benchmarks anyone can reproduce by pasting humanized text into each detector.

Does QuillBot's humanizer pass QuillBot's own detector?

No. QuillBot Humanizer's own output returns approximately 95% AI on QuillBot's own detector, which we confirmed on an owner re-test. The vendor ships both products, a humanizer and an AI Detector, and the detector says the humanizer did not work.

If you have paid for QuillBot and used its humanizer, this is not a knock on your choice, it is a diagnostic worth understanding before you trust the output. The same vendor's two tools disagree about whether the rewrite landed. That is an architectural self-fail. A synonym-swapping humanizer leaves the structural fingerprint intact, so the vendor's own detector still sees machine writing underneath the swapped words.

Our 30/30 pass on that same QuillBot classifier is the other side of the same coin. The vendor's detector is confirming an architectural difference between restructuring and paraphrasing, not rewarding a cleverer synonym list. That is the cleanest single piece of evidence that structure-over-synonyms is doing real work. For the full teardown of QuillBot's humanizer and detector together, see our QuillBot Humanizer review.

Which failure mode does this release fix?

The release targets four failure modes: passing one detector but failing another, humanizing with QuillBot and getting flagged anyway, getting flagged because your subject is unusual, and writing careful formal English that keeps tripping the score. Below is which fix applies to which. If your draft starts as a ChatGPT prompt, the two-layer prompt workflow guide covers the drafting half before any of this applies.

If you passed one detector but failed another, that is the cross-detector gap, and the six-detector row is the only readout that covers it. Pick the cell for the detector that actually grades you and plan around that number.

If you humanized with QuillBot and then got flagged by QuillBot, you hit the self-fail described above, and a corpus-trained approach sidesteps it because it restructures rather than swaps.

If your essay is in an unusual domain and still gets flagged after humanization, you hit the out-of-distribution problem, and the five-domain OOD evaluation is the test that proves coverage outside the persuasive-essay norm. For the academic workflow specifically, what to do when Turnitin's layered classifier flags your work, see our Turnitin AI bypass guide.

If you write careful, formal English and keep getting flagged, the residual-phrasing cleanup pass is aimed directly at the formal adverbs and essay-closers that survive a rewrite and trip detectors. This is the improvement non-native English writers are most likely to feel.

Where does the free tier stop?

The free tier stops at four humanizations of 250 words each on an account that asks for no card. That is enough to test the tool on a single passage and see the difference for yourself before committing to anything. Run it free first rather than take the matrix on faith.

Four free runs stop being enough at a predictable point. The 30/30 result above was measured across thirty out-of-distribution passages in one sitting. On the free tier you cannot reproduce that, because the account's four runs cover a single domain. For thorough pre-submission verification across a full essay, the Basic plan at $18 gives you 80 rewrites a month at up to 1,000 words each, and the Pro plan at $27 raises that to 200 rewrites of up to 1,200 words. None of this is a reason to skip the free test. It is a reason to know which ceiling you are about to hit before you hit it. For the full ranked cross-detector comparison across nine tools, see our best AI humanizer 2026 listicle.

Are the May 2026 figures still current?

No. This post is the dated record of the May 2026 release, and its figures stay here as history. The engine live today was measured on 31 August 2026, when its output scored 0% on GPTZero, Copyleaks and the QuillBot AI Detector, 0-3% on ZeroGPT, and a Human verdict on Turnitin and Originality AI. The September 2026 update carries that row, with screenshots.

Most humanizers never publish update notes at all, so you have no way to know whether the tool you used last term still behaves the way it did. A dated changelog is how you check. Detector vendors revise their classifiers and our engine keeps improving, so before any submission that matters, read the newest row rather than a screenshot from an older version. You can also verify it yourself: paste any text into our AI detector, which now decides its verdict with a trained model, and check it independently.

By Fırat Mıhcı (ResearchGate). $0 affiliate stake.

May 2026 Update: Better Humanizer, OOD Eval, 100% QuillBot Pass · HumanizeMy.ai